Visual odometer for mixed reality devices
By managing the estimated state and tuning measurements in the buffers of the mixed reality system, the depth uncertainty problem in pose estimation is solved, the stability of hologram anchoring and pose estimation accuracy are improved, and the user experience is improved.
Patent Information
- Application Number
- CN202380082539.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-30
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-08
AI Technical Summary
When determining the equipment posture, existing mixed reality systems are difficult to effectively manage depth uncertainty, resulting in unstable hologram anchoring and affecting user experience.
By managing multiple estimated states in the buffer, the representative measurements are tuned after partial states are popped up to keep the depth uncertainty within the threshold range, and the pose estimation is optimized using pseudo-measurements and representative measurements.
Improves the stability of hologram anchoring and the pose estimation accuracy of mixed reality systems, improving user experience.
Smart Images

Figure CN120283263A_ABST
Abstract
Description
Background Art
[0001] Mixed reality (MR) systems, which include virtual reality (VR) and augmented reality (AR) systems, have received significant attention due to their ability to create truly unique experiences for their users. For reference, traditional VR systems create a fully immersive experience by restricting the user's field of view to only a virtual environment. This is typically achieved by using a head-mounted device (HMD) that completely blocks any view of the real world. Thus, the user is fully immersed within the virtual environment. In contrast, traditional AR systems create an augmented reality experience by visually presenting virtual objects that are placed in or interact with the real world.
[0002] As used herein, VR and AR systems are described and referred to interchangeably. Unless otherwise noted, the descriptions herein apply equally to all types of MR systems, which (as detailed above) include AR systems, VR reality systems, and / or any other similar systems capable of displaying virtual content.
[0003] MR systems can be used to display various different types of information to a user. Some of this information is displayed in the form of augmented reality or virtual reality content, which may also be referred to as a "hologram". That is, as used herein, the term "hologram" generally refers to image content displayed by an MR system. In some instances, a hologram may have the appearance of a three-dimensional (3D) object, while in other instances, a hologram may have the appearance of a two-dimensional (2D) object.
[0004] Typically, holograms are displayed in a manner as if they are part of the actual physical world. For example, a hologram of a vase may be displayed on a table in the real world. In this scenario, the hologram can be considered to be "locked" or "anchored" to the real world. Such a hologram may be referred to as a "world-locked" hologram or a "space-locked" hologram, which is spatially anchored to the real world. Regardless of the user's movement, the world-locked hologram will be displayed as if it is anchored or associated with the real world. A state estimator, such as a Kalman filter, is typically used to facilitate the display of world-locked holograms. The state estimator enables the projection of content to a known location or scene despite various movements. The state estimator can provide a transformation matrix that is used to project the content and display the hologram.
[0005] In contrast, a field of view (FOV)-locked hologram is a type of hologram that is continuously displayed at a specific location within the user's FOV regardless of any movement of the user's FOV. For example, an FOV-locked hologram may be continuously displayed in the upper right corner of the user's FOV.
[0006] To properly display world-locked holograms, the task of the MR system is to obtain a spatial understanding of its environment and its pose relative to that environment. This spatial understanding is typically achieved via the use of the MR system's camera and inertial measurement unit (IMU), which includes various accelerometers, gyroscopes, and magnetometers. The MR system provides the data generated from these subsystems to a motion model and then relies on that motion model to anchor the holograms to positions in the real world. With this understanding, there is a need in the art to improve how the pose of the MR system is determined or estimated.
[0007] The subject matter claimed herein is not limited to embodiments that solve any disadvantages such as those described above or that operate only in environments such as those described above. Rather, this background is provided merely to illustrate an exemplary technical area in which some embodiments described herein may be practiced. Summary of the Invention
[0008] The embodiments disclosed herein relate to systems, devices, and methods for: (i) determining a depth uncertainty level for three-dimensional (3D) features represented in a set of images based on a plurality of estimated states included in a buffer; (ii) popping one of the estimated states from the size-limited buffer, thereby causing a modification to the available data for determining the depth of the 3D features; and (iii) attempting to maintain the depth uncertainty level even after the one estimated state has been popped.
[0009] Some embodiments access a buffer of estimated states generated based on visual observations derived from a plurality of images. The estimated states include pose information as reflected in the plurality of images. Identify 3D features co-represented in the images. The embodiments also identify two-dimensional (2D) feature points within each of the images as observations of the 3D features. Thus, a plurality of 2D feature points are identified. The images provide data usable to determine the depth of the 3D features. The embodiments determine a first estimated state to be popped from the buffer. The embodiments also calculate, for at least one of the 2D feature points, a corresponding representative measurement including a corresponding depth and a corresponding uncertainty value for that depth. Use at least these representative measurements to determine a first joint uncertainty for the depth of the 3D features. The embodiments pop the first estimated state from the buffer, thereby causing a modification to the data available for determining the depth of the 3D features. After the first estimated state is popped, tune the representative measurements until the resulting second joint uncertainty based on the tuned representative measurements is within a similarity threshold level of the first joint uncertainty, despite the modification to the data available for determining the depth of the 3D features.
[0010] Some embodiments attempt to maintain an uncertainty level for a depth even after an image in an image set is subsequently no longer available to assist in calculating a depth of a three-dimensional (3D) feature represented within the image. For example, some embodiments identify 3D features represented in an image included in a buffer. The image includes a first image and a second image, and the image provides data that can be used to determine a depth for the 3D feature. The embodiment identifies two-dimensional (2D) feature points within the image that are observations of the 3D feature. The embodiment calculates a pseudo-measurement for the 2D feature point that includes a depth and an uncertainty value for the depth. The pseudo-measurement is used to determine a first combined uncertainty for the 3D feature depth. The embodiment pops the first image from the buffer, thereby at least temporarily causing a reduction in the amount of data available to determine a depth for the 3D feature. The pseudo-measurement is tuned until a resulting second combined uncertainty based on the pseudo-measurement is within a similarity threshold level of the first combined uncertainty.
[0011] This Summary is provided to introduce a series of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to assist in determining the scope of the claimed subject matter.
[0012] Additional features and advantages will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of the teachings herein. The features and advantages of the invention may be realized and obtained by means of the instrumentalities and combinations particularly pointed out in the appended claims. The features of the present invention will become more readily apparent from the following description and appended claims, or may be learned by the practice of the invention set forth hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To describe the manner in which the above-recited and other advantages and features can be obtained, a more particular description of the subject matter briefly described above will be rendered by reference to specific embodiments that are illustrated in the appended drawings. It is to be understood that these drawings depict only typical embodiments and are therefore not to be considered limiting of its scope, and embodiments will be described and explained with additional specificity and detail through the use of the drawings in which:
[0014] Figure 1 An example head-mounted device (HMD) configured to perform the disclosed operations is shown.
[0015] Figure 2 Another configuration of the HMD is shown.
[0016] Figure 3 How the HMD can include an inertial measurement unit (IMU) is shown.
[0017] Figure 4 Shows an example architecture that can be implemented to improve how pose is estimated.
[0018] Figure 5 Shows an example environment.
[0019] Figure 6 Shows an example image of the environment.
[0020] Figure 7 Shows a buffer.
[0021] Figure 8A 、 Figure 8B and Figure 8C Shows various operations implementing the disclosed principles.
[0022] Figure 9 Shows techniques for observing features.
[0023] Figure 10 Shows improved techniques for determining the depth for a feature.
[0024] Figure 11 Shows an overview of some of the various processes performed to estimate pose.
[0025] Figure 12 Shows a flowchart of an example method for improving how pose is determined.
[0026] Figure 13 Shows another flowchart of an example method for improving how pose is determined.
[0027] Figure 14 Shows an example computer system that can be configured to perform any of the disclosed operations. Detailed Description
[0028] The disclosed embodiments (i) determine a depth uncertainty level for a three-dimensional (3D) feature represented in a set of images based on a plurality of estimated states included in a buffer; (ii) pop one of the estimated states from the buffer, thereby causing a modification to the available data that can be used to determine the depth for the 3D feature; and (iii) attempt to maintain the depth uncertainty level even after the one estimated state has been popped.
[0029] For example, some embodiments access a buffer of estimated states that are generated based on visual observations derived from multiple images. The estimated states include pose information as reflected in the multiple images. Identify 3D features that are jointly represented in the images. The embodiments also identify 2D feature points that are observations of the 3D features. The images provide data that can be used to determine the depth for the 3D features. The embodiments determine a first estimated state to be popped from the buffer. The embodiments also calculate, for at least one of the 2D feature points, a corresponding representative measurement that includes a depth and a corresponding uncertainty value for that depth. Use at least these representative measurements to determine a first joint uncertainty for the depth of the 3D feature. After the first estimated state is popped, tune the representative measurements until the resulting second joint uncertainty based on the tuned representative measurements is within a similarity threshold level of the first joint uncertainty.
[0030] Some embodiments identify 3D features that are represented in images included in the buffer. The images include a first image and a second image, and the images provide data that can be used to determine the depth for the 3D features. The embodiments identify 2D feature points within the images that are observations of the 3D features. The embodiments calculate a pseudo-measurement for the 2D feature point that includes a depth and an uncertainty value for that depth. Use the pseudo-measurement to determine a first joint uncertainty for the depth of the 3D feature. Pop the first image from the buffer, thereby at least temporarily causing a reduction in the amount of data available to determine the depth for the 3D features. Tune the pseudo-measurement until the resulting second joint uncertainty based on the pseudo-measurement is within a similarity threshold level of the first joint uncertainty.
[0031] Examples of technical benefits, improvements, and practical applications
[0032] The following section outlines some example improvements and practical applications provided by the disclosed embodiments. However, it will be understood that these are merely examples and the embodiments are not limited to just these improvements.
[0033] The disclosed embodiments provide significant benefits, advantages, and practical applications regarding how to determine the pose of a device. By improving the pose determination process, the embodiments also improve how visual images are rendered and displayed for a user to view and interact with. Thus, the embodiments not only improve the visual display of information, but they also improve how a user interacts with a computer system. Accordingly, these and many other benefits will now be described in more detail in the remainder of the present disclosure.
[0034] Example MR system and HMD
[0035] Attention will now be turned to Figure 1, which shows an example of a head-mounted device (HMD) 100. The HMD 100 can be any type of MR system 100A, including a VR system 100B or an AR system 100C. It should be noted that while most of this disclosure focuses on the use of an HMD, the embodiments are not limited to being practiced only using an HMD. For example, the disclosed operations can optionally be performed by a cloud service that communicates with the HMD.
[0036] The HMD 100 is shown as including one or more scanning sensors 105 (i.e., a type of scanning or camera system), and the HMD 100 can use the one or more scanning sensors 105 to scan the environment, map the environment, capture environmental data, and / or generate any kind of environmental image. The one or more scanning sensors 105 can include any number or any type of scanning device without limitation.
[0037] In some embodiments, the one or more scanning sensors 105 include one or more visible light cameras 110, one or more low light cameras 115, one or more thermal imaging cameras 120, optionally (but not necessarily, as Figure 1 represented by the dashed box) one or more ultraviolet (UV) cameras 125, optionally (but not necessarily, as Figure 1 represented by the dashed box) a point illuminator 130, and even an infrared camera 135. The ellipsis 140 shows how any other type of camera or camera system (e.g., depth camera, time-of-flight camera, virtual camera, depth laser, etc.) can be included among the one or more scanning sensors 105.
[0038] It should be noted that any number of cameras can be provided on the HMD 100 for each different camera type (i.e., modality). That is, the one or more visible light cameras 110 can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 cameras. However, typically, the number of cameras is at least 2, so that the HMD 100 can perform through-pass image generation and / or stereo depth matching. Similarly, the one or more low light cameras 115, the one or more thermal imaging cameras 120, and the one or more UV cameras 125 can each separately include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 corresponding cameras. The HMD 100 is also shown as including an inertial measurement unit (IMU) 145. Further details regarding this feature will be provided later.
[0039] Figure 2 Shows an example HMD 200, which represents from Figure 1HMD 100. HMD 200 is shown as including a plurality of different cameras, including camera 205, camera 210, camera 215, camera 220, and camera 225. Cameras 205 through 225 represent any number or combination of (multiple) visible light cameras 110, (multiple) low light cameras 115, (multiple) thermal imaging cameras 120, and (multiple) UV cameras 125 from Figure 1 Although only five cameras are shown in Figure 2 , HMD 200 may include more or fewer than five cameras. Any one of these cameras may be referred to as a "system camera".
[0040] Figure 3 An example of HMD 300 representing the HMD and MR systems discussed so far is shown. The descriptions of "MR device" and "MR system" may be used interchangeably with each other. In some cases, HMD 300 itself is considered an MR device. Thus, references to HMD, MR device, or MR system are generally related to each other and may be used interchangeably.
[0041] In accordance with the disclosed principles, HMD 300 is capable of using IMU data and a motion model to stabilize the visual placement of any number of holograms (e.g., 1, 2, 3, 4, 5, 10, 20, 30, 40, 50, or more than 50 holograms) rendered by the display of HMD 300. This stabilization can occur even when (it is used for visual placement while specific location data is being collected in an environment where HMD 300 is moving and it has conflicting or conflicting information).
[0042] HMD 300 is shown as including IMU 305, which represents IMU 145 from Figure 1 . IMU 305 is a type of device that measures force, angular adjustment / rate, orientation, acceleration, velocity, gravity, and sometimes even measures magnetic fields. To this end, IMU 305 may include any number of data acquisition devices, which include any number of accelerometers, gyroscopes, and even magnetometers.
[0043] The IMU 305 can be used to measure the roll rate 305A, pitch rate 305B, and yaw rate 305C. The IMU 305 can be used to measure the sum of the gravitational acceleration and body acceleration in the inertial frame. The IMU 305 can also measure angular rate and possibly absolute positioning. However, it will be understood that a motion sensor that may include the IMU 305 can measure changes in any one of the six degrees of freedom 310. The six degrees of freedom 310 refer to the ability of the body to move in three-dimensional space. For example, assume the HMD 300 is operating in the cockpit of an airplane that is rolling along a runway. Here, the cockpit can be considered the "first" environment, and the runway can be considered the "second" environment. The first environment is moving relative to the second environment. Regardless of which environment the HMD 300 is operating in, the movement of one environment relative to the other environment (as recorded or monitored by at least some of the data collection devices in the data collection device of the HMD 300) can be detected or measured in any one or more of the six degrees of freedom 310.
[0044] The six degrees of freedom 310 include surge 310A (e.g., moving forward / backward), heave 310B (e.g., moving up / down), sway 310C (e.g., moving left / right), pitch 310D (e.g., moving along the lateral axis), roll 310E (e.g., moving along the longitudinal axis), and yaw 310F (e.g., moving along the normal axis). Correlatively, the 3 DOF characteristics only include pitch 310D, roll 310E, and yaw 310F. Embodiments are capable of using 6 DOF information or 3 DOF information.
[0045] Accordingly, the IMU 305 can be used to measure changes in force and changes in movement, including any acceleration changes of the HMD 300. The data collected can be used to help determine the position, orientation, and / or perspective of the HMD 300 relative to its environment. To improve position and orientation determination, the data generated by the IMU 305 can enhance or supplement the data collected by the head tracking (HT) system. The orientation information is used to display holograms in the scene.
[0046] Figure 3 Also shown is a first HT camera 315 and its corresponding field of view (FOV) 320 (i.e., the observable area of the HT camera 315, or more precisely, the observable angle through which the HT camera 315 can capture electromagnetic radiation), as well as a second HT camera 325 and its corresponding FOV 330. Although only two HT cameras are shown, it will be understood that any number of HT cameras (e.g., 1, 2, 3, 4, 5, or more than 5 cameras) can be used on the HMD 300. Additionally, these cameras can be included as part of the HT system 335 implemented on the HMD 300.
[0047] The HT cameras 315 and 325 can be any type of HT camera. In some cases, the HT cameras 315 and 325 can be stereo HT cameras, where portions of the FOVs 320 and 330 overlap each other to provide stereo HT operation. In other embodiments, the HT cameras 315 and 325 are other types of HT cameras. In some cases, the HT cameras 315 and 325 are capable of capturing electromagnetic radiation in the visible light spectrum and generating visible light images. In other cases, the HT cameras 315 and 325 are capable of capturing electromagnetic radiation in the infrared (IR) spectrum and generating IR light images. In some cases, the HT cameras 315 and 325 include a combination of a visible light sensor and an IR light sensor. In some cases, the HT cameras 315 and 325 include or are associated with a depth detection function for detecting depth in the environment.
[0048] Accordingly, the HMD 300 can use the display orientation information generated by the IMU 305 and the display orientation information generated by the HT system 335 to determine the position and orientation of the HMD 300. Then this position and orientation information will enable the HMD 300 to accurately render holograms within the MR scene provided by the HMD 300. For example, if a hologram is to be fixedly displayed on a wall of a room, the position and orientation of the HMD 300 are used during the placement operation of the hologram to ensure that the hologram is rendered / placed at the appropriate wall location.
[0049] More specifically, to complete the hologram placement operation, a motion model (such as a Kalman filter) can be used to combine the information from the HT cameras and the information from the (multiple) IMUs to provide a robust head tracking position and orientation estimate, and the position and orientation information is used to perform hologram placement. As used herein, a "Kalman" filter is a type of combination algorithm where multiple sensor inputs (which are collected over a defined time period and which are collected using the (multiple) IMUs and HT cameras) are combined together to provide more accurate display orientation information than can be achieved by using either sensor alone. This combination can occur even in the face of statistical noise and / or other inaccuracies. The combined data is used during hologram placement. The disclosed embodiments are designed to improve how the orientation is estimated. Accordingly, the remainder of the present disclosure will now further detail various techniques for estimating the orientation of a device.
[0050] (Multiple) Example architectures
[0051] Attention will now be turned to Figure 4, which shows an example architecture 400 that can provide the benefits mentioned earlier. Architecture 400 is shown as including a service 405. The service 405 can be any type of service. As used herein, the term "service" refers to a computer program whose task is to perform automated operations or events based on an input. The service 405 can optionally be a cloud-based service operating in a cloud environment. Alternatively, the service 405 can be a local service operating on a local device. In some cases, the service 405 can be a hybrid that includes cloud-based components and local components. Generally, the task of the service 405 is to perform a number of operations. One operation involves determining the depth uncertainty level of three-dimensional (3D) features (e.g., objects in the environment) represented in a number of images. These images are stored in a buffer with a limited size. At least one of these images will eventually be popped from the buffer, resulting in a modification of the available data that can be used to determine the depth of the 3D features. A further task of the service 405 is to attempt to maintain the depth uncertainty level even after the image has been popped.
[0052] Architecture 400 shows how a set of images 410 is fed as an input to the service 405. These images can be generated by any of the camera types or modalities discussed earlier. For example, the images 410 can optionally be generated by an MR system. The images 410 are images of an environment, and the images 410 typically include relevant content.
[0053] As an example, Figure 5 shows an example of an environment 500; in this case, a bedroom. Of course, any type of environment can be representative. Figure 5 Also shown is a particular 3D feature 505 included in the environment 500; in this case, the 3D feature 505 is the corner of a bed. Additionally, the 3D feature 505 has a determined depth 510 relative to the position of a camera that can take an image of the environment 500.
[0054] Figure 6 shows Figure 5 three different images of the environment 500. For example, Figure 6 shows a first image 600, a second image 605, and a third image 610. Note that all three images include Figure 5Observation of the 3D feature 505. That is, the two-dimensional (2D) feature points 615 refer to a set of one or more image pixels that represent detectable features or objects included in the environment 500. In this case, the 2D feature points 615 are pixel-based observations of the 3D feature 505. Similarly, the 2D feature points 620 are a set of one or more pixels in the image 605; these one or more pixels also represent an observation of the 3D feature 505. To complete this example, the 2D feature points 625 are also a set of one or more pixels in the image 610, and these pixels also represent the 3D feature 505.
[0055] Returning to Figure 4 , the image 410 can represent the images 600, 605, and 610 from Figure 6 . These images 410 are fed as inputs to the service 405.
[0056] In some embodiments, the motion data 415 can also be fed as an input to the service 405. The motion data 415 can include data generated by an IMU (such as the IMU discussed previously).
[0057] The service 405 causes the image 410 to be stored in the buffer 420. The buffer 420 can also be referred to as a size-limited buffer 420, a queue, or even as a "sliding window" of images. The size of the buffer 420 is typically set such that it can store a selected number of images and / or a selected number of estimated states, as shown by the state 420A. Typically, the number of images is between 2 images and 16 images. In some cases, the number of images can exceed 16. The state 420A is generated based on visual observations derived from the images. The state 420A can include pose information as reflected in the images. The state 420A can include other variable information, such as timing information, speed information, and / or intrinsic or extrinsic parameters of the sensors.
[0058] Another task of the service 405 is to perform feature detection 425 on the image 410. By "feature detection", it means that the service 405 is able to compute various abstractions of the image and then make a centralized or localized decision for each pixel to determine whether there is a specific type of content represented by that pixel. In other words, feature detection refers to the process of classifying or assigning a type to each pixel, where the assigned class is based on the content represented by that pixel. For example, if a group of pixels in an image shows a dog, then each pixel in these pixels can be classified as the "dog" class or type. Feature detection performs this classification. Specific types of classes can also be customized such that feature detection searches for specific types of content, such as performing corner, edge, or other distinguishable objects.
[0059] Referring to Figure 6, Service 405 performed feature detection on images 600, 605, and 610. Service 405 identified 2D feature points 615, 620, and 625 as observations 630 of 3D feature 505 from Figure 5 . Each of these feature points 615, 620, and 625 can have a determined corresponding depth 635 value, where the depth 635 is an attempt to approximate the actual depth 510 of 3D feature 505 relative to the camera that generated images 600, 605, and 610.
[0060] Returning to Figure 4 And as previously mentioned, Service 405 caused image 410 to be stored in buffer 420. In some cases, feature detection 425 occurs when image 410 is stored in buffer 420. In other cases, feature detection 425 occurs in time before image 410 is stored in buffer 420.
[0061] Figure 7 Buffer 700 is shown, which represents buffer 420. As previously mentioned, the size 705 of buffer 700 is limited such that it supports or can include a set number of images.
[0062] Figure 7 Many images (e.g., rectangles) stored in buffer 700 are shown. For example, in particular, buffer 700 is currently storing images 710, 715, and 720. Currently, buffer 700 has reached its maximum capacity with respect to its storage ability. If a new image (e.g., image 725) is to be injected or inserted into buffer 700 (e.g., injection 730), then an existing image (e.g., image 735) will need to be popped out of buffer 700 (e.g., pop 740).
[0063] Recall that the various different images in the buffer include observations of 3D features. These observations are used to determine the depth of the 3D feature. Depth data is useful because it helps to localize the MR system by enabling the MR system to determine its pose relative to the environment. The display of holograms depends on the pose of the MR system. Therefore, in order for the MR system to provide high-quality image content, it is desirable to have accurate and robust depth data.
[0064] In addition, generally, older images in buffer 700 (e.g., as determined by the corresponding timestamp of each image) include a rich amount of data that can be used to determine the depth of 3D features. For example, generally, older images in buffer 700 observe 3D features from a favorable position or perspective different from that of newer images. In some cases, older images observe 3D features from a different depth than newer images. Thus, the combination of older and newer images provides a greater baseline (i.e., representing the difference in the perceived perspective in the images) based on which the depth of 3D features is calculated, and typically, a greater baseline enables more accurate depth determination because a greater baseline results in less uncertainty regarding depth calculation.
[0065] When an image in the buffer 700 is popped (especially one of the older images), the amount of available data or observations 745 that can be used to calculate depth is typically reduced or at least modified. Typically, the oldest image is popped first, such as in a first-in, first-out (FIFO) type of queue. Of course, buffer 700 can optionally be configured in any way and is not limited to the FIFO scheme.
[0066] When compared and contrasted with newer images, the combination of older and newer images provides a robust mechanism for determining depth. However, when the older image is popped, traditionally, the amount of data now available for determining depth is significantly reduced. Due to this reduction, the resulting depth calculation will also be less robust. The disclosed embodiments are designed to mitigate the impact that occurs when an image is popped from buffer 700, such that robust depth calculation and resulting pose estimation can be performed.
[0067] Returning to Figure 4 , the task of service 405 is to identify 2D feature points 430 from the image, as discussed. For example, 2D feature points 615, 2D feature points 620, and 2D feature points 625 from Figure 6 represent 2D feature points 430.
[0068] The task of service 405 is also responsible for calculating the so-called "representative measurements" or "pseudo-measurements", as shown by representative measurement 435. In some cases, a representative measurement 435 is calculated for each 2D feature point in the image. As an example, a first representative measurement is calculated for 2D feature point 615, a second representative measurement is calculated for 2D feature point 620, and a third representative measurement is calculated for 2D feature point 625. Each of these representative measurements is associated with or related to 3D feature 505 from Figure 5 .
[0069] Note that some embodiments avoid or do not compute representative measurements for each 2D feature point. In some cases, the number of representative measurements may be less than the number of 2D feature points for a particular 3D feature. For example, assume that five images jointly observe the same 3D feature. The service may identify five 2D feature points for the 3D feature (i.e., one 2D feature point in each image). Then, some embodiments may generate five different representative measurements. However, other embodiments may generate fewer than five representative measurements. For example, some embodiments may generate 1, 2, 3, or 4 representative measurements.
[0070] The representative measurement 435 includes a depth 435A (relative to the camera that generated the image) that is determined for the 3D feature 505. The representative measurement 435 also includes a determined uncertainty value 435B for the determined depth 435A. The uncertainty value 435B reflects a corresponding measure of how much depth information is available (within a particular image) to determine the depth 435A for the 3D feature. Some images may have more available depth information than other images, and thus some uncertainty values may be higher or lower for certain images. Generally, the uncertainty value 435B reflects the likelihood regarding the accuracy of the computed depth 435A.
[0071] In some embodiments, representative measurements are computed only when an image is to be popped from the buffer 420. Thus, in some embodiments, representative measurements are not computed while the buffer 420 is being filled. Popping an image from the buffer 420 can be a trigger event such that the representative measurements will be computed, including for the image that is to be popped.
[0072] The service 405 also computes a so-called joint uncertainty 440 for a set of images 410. The joint uncertainty 440 is computed based on a combination of representative measurements.
[0073] After computing the joint uncertainty 440, the service 405 allows one of the images in the buffer 420 to be popped. Optionally, a new image can be injected into the buffer 420. The joint uncertainty 440 will now be used as a "check" or "intermediate quantity" that is used to optimize the resulting representative measurements in the form of a tuning operation, as shown by the tuning 445.
[0074] That is, the service 405 performs a tuning operation to tune the uncertainty of the pseudo-measurements (i.e., representative measurements) in an attempt to minimize the information lost due to popping frames from the buffer. An example would be helpful.
[0075] In one scenario, the pseudo-measurement includes an average depth and some positive uncertainty about that depth of a single scene point relative to the camera. The depth of this pseudo-measurement can be fixed to the latest estimate (e.g., 3.0 m). In such a pseudo-measurement in square meters, the uncertainty value is expected to be determined in the form of variance. After determining the depth value and the uncertainty, it can be included in any subsequent estimation process, including pose estimation.
[0076] The pseudo-measurement uncertainty is calculated as follows. Given all the available observations / measurements in buffer 420 (e.g., a sliding window) related to the image / frame to be popped, including: (i) previous pseudo-measurements, (ii) visual observations, (iii) IMU data, and (iv) possibly other data, service 405 can calculate a measure of the joint uncertainty 440, including the frame to be popped, the points observed in that frame, and possibly other camera poses.
[0077] Since this calculation involves correlations among multiple variables, it can be represented as a matrix. An embodiment can pop a frame and merge the pseudo-measurements. Then, the embodiment can define the uncertainty about the remaining variables after this process, and it is expected that this uncertainty remains invariant through the popping process. An exact measure and process for this will be provided later. Before this involved explanation, simplified examples will be provided in Figure 8A 、 Figure 8B and Figure 8C as follows.
[0078] Figure 8A Figure 800 shows a sliding window with 4 frames (labeled T1 to T4). Figure 8A It also shows two points X and Y. Measurements are shown using solid lines. At this stage, there are no previous pseudo-measurements, but there may be in other scenarios. Frame-to-frame measurements may also exist (e.g., possibly from an IMU), but this is not strictly necessary. Frame T1 is popped from the sliding window 800, as shown by the frame to be popped 805.
[0079] As Figure 8A shown, in step "A", the service can ignore all measurements except those related to T1 because, at this step, the service is focused on retaining the information that will be lost by popping T1. This leaves the service with two measurements. The service calculates the joint uncertainty 810 about point X, frame T1, and frame T2. Generally, this joint uncertainty 810 consists of all variables related to T1 through any measurement and / or pseudo-measurement.
[0080] As Figure 8B shown, in step "B", the service will calculate the uncertainty of a single pseudo-measurement about point X in frame T2. Figure 8BShows the target uncertainty 815 and the actual uncertainty 820. Note that the target uncertainty 815 corresponds to the boxed value in the joint uncertainty 810 of Figure 8A the
[0081] It is desirable that the uncertainty about (T2, X) generated by this pseudo-measurement (the dashed line between point X and frame T2) is close to some target determined by the previously calculated uncertainty (e.g., the target uncertainty 815). Preferably, the value in the target uncertainty 815 is simply copied to the actual uncertainty 820, but the service only has one number to choose from (plotted as σ^2 in Figure 8B ). The pseudo-measurement is not completely general, so all that can be achieved by choosing σ^2 is Figure 8B the rightmost matrix in
[0082] As Figure 8C shown, in step "C", the service merges this pseudo-measurement back into the sliding window of the popped frame T1. Now, the service no longer considers the previously calculated joint uncertainty. The service can use the new frame to enhance the sliding window and repeat the process, treating it as any other tracking system. Here, pseudo-measurement generation and merging can be used as a direct replacement for an alternative method of removing the oldest frame and losing information.
[0083] In general, there can be multiple points and multiple frames viewing point X. The detailed description to be provided later is a way to generalize these scenarios and to quantify the approximations in detail. Leaving these details aside, Figure 8A 、 Figure 8B and Figure 8C represent the concepts and intuitive perceptions of the method. The disclosed embodiments are beneficially and uniquely customized to provide a formula where the embodiments limit the approximations to ignore measurements that do not involve the frame to be removed.
[0084] Another simplified example would be helpful. Suppose the service has three pseudo-measurements -- (1) for a first feature point in a first image, the depth is calculated to be 3.0 meters and the uncertainty is 3%; (2) for a second feature point in a second image, the depth is calculated to be 3.1 meters and the uncertainty is 3.2%; (3) for a third feature point in a third image, the depth is calculated to be 2.9 meters and the uncertainty is 3.1%. For these three images, the initial combined uncertainty can be calculated to be 3.05% (based on these percentage values and other data, such as the actual depth data). Then, the third image including the third feature point is popped out, so the data for the first and second feature points remain. Then, the service modifies or "tunes" the uncertainties of 3% and 3.2% (for the first and second feature points respectively) to "maintain" the combined uncertainty of 3.05%.
[0085] The above description is a possible implementation of the process in a special case. However, in most cases, the camera frames do not directly observe depth and only have 2D observations in the images. Therefore, the depth uncertainty may not be directly modified. To solve this problem, the pseudo-measurements can optionally be combined into completely new measurements. If there are already direct depth measurements, they may be combined in the above manner.
[0086] Thus, in this way, the "combined uncertainty" is used to fine-tune the pseudo-measurement data. The pseudo-measurement data are the actual data used to calculate the estimated pose of the device; the combined uncertainty is not used during this calculation. Therefore, the combined uncertainty is a "check" or "intermediate quantity" used to optimize the pseudo-measurement data.
[0087] Briefly return Figure 4 , the (multiple) "tuned" representative measurements are tuned in this way such that the resulting combined uncertainty is within the threshold 450 degrees relative to the combined uncertainty calculated when the buffer 420 was full. The tuned representative measurements can then be used to calculate or estimate the pose 455 of the device.
[0088] Detailed examples
[0089] Inside-out tracking is a technique used in MR devices to present virtual content that appears stationary to the user. State-of-the-art tracking is implemented using efficient and accurate Visual Inertial Odometry (VIO) modules that meet the unique computational constraints of the MR system. Typically, the VIO modules take as input timestamped images and data from an IMU. The system uses the images, existing 3D points, and some knowledge about the state of the device to compute 2D feature observations corresponding to 3D world points. This incorporates any of several feature detection and tracking techniques. The system can then decide whether it should compute the "state" of the rig / MR system. This "state" includes the pose of the device in a common coordinate system and can include differentials (e.g., velocity, acceleration, etc.) and / or IMU state (e.g., accelerometer bias, gyroscope bias, and gyroscope scale factor). If the decision is "no", it waits for the next frame. Otherwise, it constructs a non-linear least squares optimization problem that includes: 1) the keyframe / image state and the 3D positions of the points; 2) constraints from the available sensors, including feature observations, previously generated pseudo-measurements, and previous marginalization priors.
[0090] The state of the device at the image timestamp is incorporated into the new keyframe state. This solves the problem such that it can provide the state of the device to the display and other subsystems that require an updated head pose. Then, some process removes the previous keyframe / image, such as the oldest keyframe / image, the details of which are described herein.
[0091] Figure 9 Technique 900 for maintaining state by removing the state of the oldest keyframe is shown. There are various measurements or constraints included in the optimization problem. Each of them encodes some information about the variables it is connected to. The triangle represents the state of the device (e.g., the MR system), and the map consists of X1, X2, and X3. In reality, many more keyframes and points would be included, but they are omitted for illustrative purposes.
[0092] The middle square is the 2D observation (e.g., feature observation) of the 3D point derived from the image. The bottom square is the "marginalization prior", which encodes information about the relationship between keyframes, including their relative pose and optional constraints from inertial measurements. The transformation applied is called "marginalization", which includes certain numerical calculations required to remove the oldest keyframe. Before marginalization, typically the feature observations visible to both the oldest and the newest keyframes are discarded. This results in loss of information, which in turn reduces tracking accuracy.
[0093] Figure 10 An overview of an improved technique 1000 in accordance with the disclosed details is presented. The transformation is more than Figure 9The transformation is more complex and will be described in more detail later. It is worth noting that even in the worst-case scenario, the overall computational (time and space) complexity of the algorithm is linear with the size of the map. The method is customized to track MR applications with unique computational constraints.
[0094] It is worth noting that although the disclosed technology is applicable to VIO systems used on MR devices, the disclosed principles can also be practiced in a sliding-window-based visual odometry system that does not include an IMU. If additional sensors and constraints do not include the correlation between map points, they can complement the operation. Some examples include satellite measurement data from the Global Positioning System (GPS), magnetometers, and sensor-independent motion constraints (trajectory priors). The disclosed principles are applicable to trackers mounted on devices other than head-mounted displays, such as motion controllers. The details of the concurrency and timing of the operation may vary slightly. For example, the marginalization process can be performed concurrently with feature detection and tracking.
[0095] A general summary of some of the disclosed operations will now be provided below.
[0096] An inertial measurement unit including a gyroscope and an accelerometer is embedded in the mobile device. One or more visible light cameras are attached to the device, which can view the environment. A sliding-window-based VIO method based on the state of the device is used to fuse visual and inertial measurements, thereby generating an estimate of the motion of the device. The state of the device is estimated by solving a non-linear least squares optimization problem. The updated instantaneous state of the device is returned to other subsystems.
[0097] Through a process of generating pseudo-measurements, the previous state is removed from the sliding window, which incorporates the pseudo-measurements into the subsequent optimization problem. These pseudo-measurements mitigate information loss and thus improve tracking accuracy.
[0098] Calculating the uncertainty of a pseudo measurement
[0099] The embodiment first only considers Figure 9 a part of the graph in , which only includes the variable T1 to be marginalized, any error terms involving this variable, and any variables involving these error terms. The latter includes N map points X1, X2,..., X N and the states of other key frames aggregated in T2. The naive marginalization of T1 will generate new constraints on the remaining variables. Therefore, the Gaussian distribution describing the modeling of the remaining variables will have a dense information matrix. Due to the lack of structure, this problem is usually too costly to solve. Therefore, it is desirable to approximate this information using pseudo-measurements associated with a sparse information matrix ( Figure 11 ).
[0100] Figure 11Shows various subgraphs, including: a) the marginalized variables and associated error terms; b) the information generated by the ordinary marginalization process; c) the information generated by the traditional method of discarding observations; and d) the sparse information that approximates the target information using pseudo-measurements. The sparse information has a similar expected structure to the traditional method but retains more information.
[0101] The embodiments only consider these subgraphs and not the rest of the graph because: 1) the rest of the graph undergoes changes, and the approximation criteria selected by the embodiments are independent of the rest of the graph; 2) the marginalization of T1 has a local effect in the graph, which is limited to these variables; 3) in this subgraph, given the variables to be marginalized, T2 and the map are conditionally independent. This structure leads to an efficient linear-time algorithm.
[0102] Given Figure 11 the structure in (a) shown in, the embodiments can partition the information matrix into blocks.
[0103]
[0104] Among them, the block with B corresponds to X1, X2, X3, and T2. The block with M corresponds to T1. Direct marginalization will result in Figure 11 the (dense) target information in (b) of.
[0105]
[0106] Applying the Woodbury matrix identity gives the inverse matrix.
[0107]
[0108] Λ BB is a block diagonal matrix, so its inverse matrix can be calculated efficiently. The embodiments now use this matrix to calculate the sparse information matrix. Let H be the vertically stacked Jacobian matrix of the pseudo-measurements with respect to the variables in the original subgraph. H has the following block structure, where the last column corresponds to T2 and each of the previous columns corresponds to one of the map points.
[0109]
[0110] Minimizing the Kullback - Liebler divergence from the sparse information to the target information gives a closed - form solution. The means of the two distributions are equal, and the i - th diagonal block of the sparse information matrix provides the desired uncertainty and is given by:
[0111] (Λ s ) i =(H∑ t H T ) i-1
[0112] This expression can be expanded to:
[0113]
[0114] where the subscripts are zero-based block indices in their respective matrices. These two sums involve at most two terms, which makes computing the blocks of the sparse information a constant-time operation. The blocks of the covariance can also be computed in constant time. Since there is one block per map point, the total cost is linear with the size of the map.
[0115] The blocks of the covariance matrix are computed as follows.
[0116] First, the embodiment computes the Cholesky factorization of the following expression:
[0117]
[0118] Next, the embodiment defines:
[0119]
[0120] The embodiment has ∑ t = D + UU T . The embodiment can compute the blocks. The blocks of D are zero for non-diagonal blocks and are the pseudo-inverse matrices of small matrices for diagonal blocks.
[0121] Example methods
[0122] The following discussion now pertains to multiple methods and method acts that can be performed. Although method acts may be discussed in a particular order or shown in a flowchart as being in a particular order, no particular order is required unless specifically stated or required because an act depends on another act that is completed before that act is performed.
[0123] Attention is now directed to Figure 12 , which shows a flowchart of an example method 1200 for the following operations: (i) determining a depth uncertainty level for three-dimensional (3D) features represented in a set of images based on a plurality of estimated states (e.g., state 420A from Figure 4 ) included in a buffer, (ii) popping one of the estimated states from the buffer, thereby causing a modification to the available data that can be used to determine the depth for the D features, and (iii) attempting to maintain the depth uncertainty level even after one of the estimated states has been popped. Method 1200 can be in Figure 4is implemented within the architecture 400. Additionally, the method 1200 can be implemented by the service 405.
[0124] The method 1200 includes an action (action 1205) of accessing a size - limited buffer of estimated states that are generated based on visual observations derived from multiple images. In some embodiments, the buffer can also include multiple images, including a first image and a second image. In some implementations, the first image and / or the first estimated state can be the oldest image or the oldest state in the size - limited buffer. In some cases, the first image and / or the first estimated state may not be the oldest image. Optionally, the number of images and / or states included in the size - limited buffer is set such that the size - limited buffer reaches its maximum capacity. In some cases, after performing certain actions (e.g., extracting 2D feature points), an image can be discarded from the buffer. Optionally, the number of estimated states included in the buffer is set such that the buffer reaches its maximum capacity. The estimated states include pose information reflected in the multiple images. In addition to the estimated states, the buffer can also include visual observations derived from the images. The estimated states can also include other variables associated with the time instance at which the images were acquired. For example, the other variables can include velocity and the intrinsic or extrinsic parameters of sensors on the device. The buffer maintains estimates of these states, which can be updated given new information. When a state is finally popped from the buffer, the system can commit to its corresponding estimate.
[0125] Action 1210 includes identifying 3D features that are co - represented in the images. For example, 3D feature 505 can be identified in each of image 600, image 605, and image 610.
[0126] Action 1215 includes identifying two - dimensional (2D) feature points within each of the images as observations of the 3D features. Thus, multiple 2D feature points are identified. Notably, the images provide data that can be used to determine the depth for the 3D features.
[0127] Action 1220 includes determining that a first estimated state will be popped from the size - limited buffer. Optionally, the first image can additionally or alternatively be popped from the buffer. This can occur because a new image is now available for insertion into the buffer, or because a new estimated state will be inserted into the buffer.
[0128] For at least one of the 2D feature points, action 1225 includes calculating a corresponding representative measurement (also known as a pseudo - measurement), which includes a corresponding depth and a corresponding uncertainty value for that depth. Thus, one or more representative measurements are calculated. Each of the uncertainty values reflects how much depth information is available for determining the corresponding measure of the depth for the 3D feature.
[0129] More specifically, for a measurement with N values, the uncertainty values are an N×N covariance matrix that is Gaussian distributed. It can also be the inverse of this matrix or other equivalents.
[0130] Each uncertainty value in the uncertainty values is also typically based on the pixel resolution of the image. Compared with an image having a lower pixel resolution, an image having a higher pixel resolution provides a better framework for determining depth.
[0131] Action 1230 includes determining a first joint uncertainty of depth for a 3D feature using at least one or more representative measurements. In some implementations, the first joint uncertainty takes into account both the actual depth measurement data and one or more representative measurements.
[0132] Action 1235 includes popping a first image and / or a first estimated state from a size-limited buffer, thereby causing a modification to the data available for determining the depth of the 3D feature.
[0133] After the first image and / or the first estimated state is popped, action 1240 includes tuning one or more representative measurements until the resulting second joint uncertainty based on the tuned one or more representative measurements is within a similarity threshold level of the first joint uncertainty, despite the modification to the data available for determining the depth of the 3D feature. In other words, the uncertainty of the representative measurements is tuned. The depth value is typically fixed. Subsequently, the uncertainties associated with the measurements and the pseudo-measurements / representative measurements are used in pose estimation. The measurements and uncertainties can be used for estimation purposes using weighted least squares, Kalman filtering, or similar methods.
[0134] The method can also include an action of estimating the pose of the MR system or the computer system using at least the modified representative measurements. The pose estimation is performed using the uncertainties of the representative measurements and those pseudo-measurements, rather than using the joint uncertainty previously used in calculating those pseudo-measurements. Pose estimation typically requires measurements and an uncertainty associated with each measurement, so only the pseudo-measurements and uncertainties are used.
[0135] Optionally, the motion data can also be used to facilitate pose estimation. In some cases, the pose of the computer system / MR system includes one or more of differential information or motion state information. Additionally, the pose can be estimated relative to an initial baseline pose. For example, when the MR system is first initialized, its pose can be determined. Subsequent poses can be determined based on this initial pose, thereby tracking or reflecting how the MR system has moved over time relative to the initial baseline pose. In some cases, the pose estimation is performed relative to a previous pose, which was determined earlier in time relative to the time at which the pose is being estimated. That is, the previous pose is not necessarily the initial baseline pose.
[0136] Figure 13 Another flowchart of a method is shown that attempts to maintain an uncertainty level for a depth even after one of the images in an image set is subsequently no longer available to assist in calculating the depth of a three-dimensional (3D) feature represented within the image set. Method 1300 can also be implemented by Figure 4 service 405.
[0137] Method 1300 includes an action of identifying 3D features represented in an image included in a buffer (action 1305). The image includes a first image and a second image, and the images provide data that can be used to determine the depth for the 3D feature.
[0138] Action 1310 includes identifying two-dimensional (2D) feature points within the image that are observations of the 3D feature.
[0139] Action 1315 includes, for the 2D feature points, calculating a pseudo-measurement that includes a depth and an uncertainty value for the depth. Advantageously, the pseudo-measurement approximates an observation that has been removed or will subsequently be removed due to popping the first image from the buffer. The pseudo-measurement can be one of a plurality of pseudo-measurements. The pseudo-measurement can include a displacement metric that includes any one of a 1 degree of freedom (DOF) measurement, a 3 DOF measurement, a 6 DOF measurement, or even a 9 DOF measurement. Generally, the pseudo-measurement does not exceed 9 DOF. It is typically a measurement involving a 3D point and a 6 DOF pose. The formula for the measurement will be fixed, and only its uncertainty will be tuned.
[0140] Action 1320 includes using the pseudo-measurement to determine a first joint uncertainty for the depth of the 3D feature.
[0141] Action 1325 includes popping the first image from the buffer, thereby at least temporarily causing a reduction in the amount of data available to determine the depth for the 3D feature. The reduction in the amount of data available to determine the depth for the 3D feature includes a reduction in the number of observations available within the image and relative to the 3D feature.
[0142] Action 1330 includes tuning the pseudo-measurements until the resulting second joint uncertainty based on the pseudo-measurements is within a similarity threshold level of the first joint uncertainty. Tuning the pseudo-measurements can include modifying the uncertainty value for the depth that is part of the pseudo-measurements. In some implementations, the uncertainty value is only determined once and is not modified based on some previous value.
[0143] Method 1300 can also include injecting a new image into the buffer. The process can then repeat itself using the new image. In some cases, method 1300 can include identifying redundant pseudo-measurements. "Redundant" refers to that one pseudo-measurement may have the same value as another pseudo-measurement. Alternatively, "redundant" refers to that one pseudo-measurement may have a value within the similarity threshold level of another pseudo-measurement. Then, the method can include the action of retaining non-redundant pseudo-measurements and possibly eliminating redundant pseudo-measurements. Embodiments can also merge pseudo-measurements with existing measurements by only reducing the uncertainty of the existing measurements. This may result in some additional computational benefits.
[0144] Accordingly, the disclosed embodiments are directed to various techniques for improving how to estimate the pose of a device based on depth data. The embodiments are able to provide highly accurate estimates even when the observation data used to calculate the depth data is lost when an image is popped from a sliding window or buffer.
[0145] Example computer / computer system
[0146] Attention is now turned to Figure 14 , which shows an example computer system 1400, which can include and / or be used to perform any of the operations described herein. Computer system 1400 can implement Figure 4 service 405.
[0147] Computer system 1400 can take various different forms. For example, computer system 1400 can be embodied as a tablet computer, a desktop computer, a laptop computer, a mobile device, or a stand-alone device, such as those described throughout this disclosure. Computer system 1400 can also be a distributed system that includes one or more connected computing components / devices that communicate with computer system 1400.
[0148] In its most basic configuration, computer system 1400 includes various different components. Figure 14 It is shown that computer system 1400 includes one or more processors 1405 (also known as "hardware processing units") and a storage device 1410.
[0149] Regarding the (multiple) processors 1405, it will be understood that the functions described herein can be performed, at least in part, by one or more hardware logic components (e.g., the (multiple) processors 1405). For example, but not limited to, illustrative types of hardware logic components / processors that can be used include field programmable gate arrays (“FPGAs”), application specific integrated circuits (“ASICs”), application specific standard products (“ASSPs”), system on chips (“SOCs”), complex programmable logic devices (“CPLDs”), central processing units (“CPUs”), graphics processing units (“GPUs”), or any other type of programmable hardware.
[0150] As used herein, the terms “executable module,” “executable component,” “component,” “module,” “service,” or “engine” can refer to a hardware processing unit or a software object, routine, or method that can be executed on a computer system 1400. The different components, modules, engines, and services described herein can be implemented as objects or processors that execute on the computer system 1400 (e.g., as separate threads).
[0151] The storage device 1410 can be a physical system memory, which can be volatile, non-volatile, or some combination of the two. The term “memory” can also be used herein to refer to non-volatile mass storage devices such as physical storage media. If the computer system 1400 is distributed, then the processing, memory, and / or storage capabilities can also be distributed.
[0152] The storage device 1410 is shown as including executable instructions 1415. The executable instructions 1415 represent instructions executable by the (multiple) processors 1405 of the computer system 1400 to perform the disclosed operations (such as those described in the various methods).
[0153] The disclosed embodiments may include or utilize a special purpose or general purpose computer including computer hardware, such as, for example, one or more processors (such as (a) processors 1405) and system memory (such as storage device 1410), as discussed in more detail below. Embodiments also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. A computer-readable media that stores computer-executable instructions in the form of data is a "physical computer storage medium" or "hardware storage device". Additionally, computer-readable storage media (including physical computer storage media and hardware storage devices) excludes signals, carriers, and propagated signals. On the other hand, a computer-readable media that carries computer-executable instructions is a "transmission medium" and includes signals, carriers, and propagated signals. Thus, by way of example and not limitation, the current embodiments can include at least two distinctly different kinds of computer-readable media: computer storage media and transmission media.
[0154] Computer storage media (aka "hardware storage device") is computer-readable hardware storage devices such as RAM, ROM, EEPROM, CD-ROM, RAM-based solid state drives ("SSD"), flash memory, phase change memory ("PCM"), or other types of memory, or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code units in the form of computer-executable instructions, data, or data structures and that can be accessed by a general purpose or special purpose computer.
[0155] Computer system 1400 may also be connected (via wired or wireless connections) to external sensors (e.g., one or more remote cameras) or devices via network 1420. For example, computer system 1400 may communicate with any number of devices or cloud services to obtain or process data. In some cases, network 1420 itself may be a cloud network. Additionally, computer system 1400 may also be connected to (a) remote / standalone computer system(s) via one or more wired or wireless networks, the (a) remote / standalone computer system(s) being configured to perform any of the processing described with respect to computer system 1400.
[0156] A "network", such as network 1420, is defined as one or more data links and / or data switches that enable the transfer of electronic data between computer systems, modules, and / or other electronic devices. When information is transmitted or provided to a computer via a network (wired, wireless, or a combination of wired and wireless), the computer appropriately views such a connection as a transmission medium. Computer system 1400 will include one or more communication channels for communicating with network 1420. A transmission medium includes a network that can be used to carry data or desired program code units in the form of computer-executable instructions or in the form of data structures. In addition, these computer-executable instructions can be accessed by a general or special-purpose computer. The above combinations should also be included within the scope of computer-readable media.
[0157] Upon reaching various computer system components, program code units in the form of computer-executable instructions or data structures can automatically transfer from the transmission medium to computer storage media (and vice versa). For example, computer-executable instructions or data structures received via a network or data link can be buffered in RAM within a network interface module (e.g., a network interface card or "NIC") and then ultimately transferred to the computer system RAM and / or more non-volatile computer storage media at the computer system. Thus, it should be understood that computer storage media can be included in computer system components that also (or even primarily) utilize the transmission medium.
[0158] Computer-executable (or computer-interpretable) instructions include, for example, instructions that cause a general-purpose computer, special-purpose computer, or special-purpose processing device to perform a certain function or set of functions. For example, computer-executable instructions can be binary files, intermediate format instructions (such as assembly language), or even source code. Although the subject matter has been described in language specific to structural features and / or method acts, it will be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Instead, the described features and acts are disclosed as example forms for implementing the claims.
[0159] Those skilled in the art will understand that embodiments may be practiced in network computing environments with many types of computer system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, and the like. Embodiments may also be practiced in distributed system environments where local and remote computer systems that are linked through a network (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) each perform tasks (such as cloud computing, cloud services, etc.). In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0160] The present invention may be embodied in other specific forms without departing from its characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. Thus, the scope of the present invention is indicated by the appended claims rather than by the foregoing description. All changes that fall within the meaning and scope of the claims will be embraced within their scope.
Claims
1. A computer system that (i) determines a depth uncertainty level for three-dimensional (3D) features represented in a set of images based on multiple estimated states included in a buffer, (ii) pops one of the estimated states from the buffer, thereby causing a modification to the available data that can be used to determine the depth for the 3D features, and (iii) attempts to maintain the depth uncertainty level even after the one estimated state has been popped, the computer system comprising: At least one processor; And At least one hardware storage device storing instructions executable by the at least one processor to cause the computer system to: Access a buffer of estimated states generated based on visual observations derived from multiple images, wherein the estimated states include pose information as reflected in the multiple images; Identify 3D features co-represented in the images; Within each of the images, identify two-dimensional (2D) feature points that are observations of the 3D features such that a plurality of 2D feature points are identified, wherein the images provide data that can be used to determine the depth for the 3D features; Determine a first estimated state to be popped from the buffer; For at least one of the 2D feature points, compute a corresponding representative measurement including a corresponding depth and a corresponding uncertainty value for the depth such that the one or more representative measurements are computed; Use at least the one or more representative measurements to determine a first joint uncertainty for the depth of the 3D features; Pop the first estimated state from the buffer, thereby causing a modification to the data that can be used to determine the depth for the 3D features; And After the first estimated state is popped, tune the one or more representative measurements until a resulting second joint uncertainty based on the tuned one or more representative measurements is within a similarity threshold level of the first joint uncertainty despite the modification to the data that can be used to determine the depth for the 3D features.
2. The computer system according to claim 1, wherein the number of estimated states included in the buffer is set such that the buffer reaches a maximum capacity.
3. The computer system according to claim 1, wherein the first joint uncertainty takes into account both actual depth measurement data and the one or more representative measurements.
4. The computer system according to claim 1, wherein the execution of the instructions further causes the computer system to estimate a pose for the computer system using at least the modified representative measurements.
5. The computer system according to claim 4, wherein motion data is also used to estimate the pose of the computer system.
6. The computer system according to claim 4, wherein the pose of the computer system includes one or more of the following: differential information or motion state information.
7. The computer system according to claim 1, wherein the pose is estimated relative to an initial baseline pose.
8. The computer system according to claim 1, wherein the pose is estimated relative to a previous pose, the previous pose being determined earlier in time relative to the time at which the pose is estimated.
9. The computer system according to claim 1, wherein the first estimated state is the oldest estimated state in the buffer.
10. The computer system according to claim 1, wherein each of the uncertainty values reflects how much depth information is available for determining a corresponding measure of the depth for the 3D feature.
11. The computer system according to claim 1, wherein each of the uncertainty values is based on the pixel resolution of the image.
12. A method for attempting to maintain an uncertainty level for depth even after an image in an image set is no longer available to assist in determining the computed depth of a three-dimensional (3D) feature represented within the image set, the method comprising: Identifying 3D features represented in images included in a buffer, wherein the images include a first image and a second image, and wherein the images provide data that can be used to determine the depth for the 3D feature; Identifying two-dimensional (2D) feature points within the images that are observations of the 3D feature; Calculating, for the 2D feature points, a pseudo-measurement that includes a depth and an uncertainty value for the depth; Using the pseudo-measurement to determine a first combined uncertainty for the depth of the 3D feature; Popping the first image from the buffer, thereby at least temporarily causing a reduction in the amount of data that can be used to determine the depth for the 3D feature; And Tuning the pseudo-measurement until a resulting second combined uncertainty based on the pseudo-measurement is within a similarity threshold level of the first combined uncertainty.
13. The method according to claim 12, wherein the method further comprises injecting a new image into the buffer.
14. The method according to claim 12, wherein tuning the pseudo-measurement includes modifying the uncertainty value for the depth that is part of the pseudo-measurement.
15. The method according to claim 12, wherein the reduction in the amount of data that can be used to determine the depth for the 3D feature includes a reduction in the number of observations that are available in the image and relative to the 3D feature.
16. The method according to claim 15, wherein the pseudo-measurement is near the observations that have been removed due to the first image being popped from the buffer.
17. The method according to claim 12, wherein the pseudo-measurement is one of a plurality of pseudo-measurements.
18. The method according to claim 12, wherein the pseudo-measurement includes a displacement metric, the displacement metric including any one of: a three degrees of freedom (DOF) measurement or a one DOF measurement.
19. A computer system that attempts to maintain an uncertainty level for a depth even after an image in an image set is subsequently no longer available to assist in calculating a depth of a three-dimensional (3D) feature represented within the image set, the computer system comprising: at least one processor; and at least one hardware storage device storing instructions executable by the at least one processor to cause the computer system to: identify 3D features represented in images included in a buffer, where the images include a first image and a second image, and where the images provide data that can be used to determine a depth for the 3D features; identify two-dimensional (2D) feature points within the images that are observations of the 3D features; calculate, for the 2D feature points, pseudo-measurements that include a depth and an uncertainty value for the depth; use the pseudo-measurements to determine a first joint uncertainty for the depth of the 3D features; pop the first image from the buffer, thereby at least temporarily causing a reduction in the amount of data that can be used to determine the depth for the 3D features; and tune the pseudo-measurements until a resulting second joint uncertainty based on the pseudo-measurements is within a similarity threshold level of the first joint uncertainty.
20. The computer system of claim 19, wherein execution of the instructions further causes the computer system to: identify redundant pseudo-measurements; and retain non-redundant pseudo-measurements and eliminate redundant pseudo-measurements.