Improved pose determination for display devices - Patents.com

The system addresses the challenges of presenting virtual content in AR systems by using imaging devices and processors to match prominent points in real-world environments, enabling accurate orientation determination and improved user experience.

JP7674454B2Active Publication Date: 2025-05-09MAGIC LEAP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023206824
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-01-30
Filing Date
2023-12-07
Publication Date
2025-05-09
Estimated Expiration
2038-12-14

AI Technical Summary

Technical Problem

Existing augmented reality (AR) display systems face challenges in providing a comfortable, natural-like, and rich presentation of virtual image elements within real-world environments, due to the complexity of the human visual perceptual system.

Method used

A system comprising imaging devices, processors, and computer storage media that acquires current images of real-world environments, projects prominent points from previous images onto corresponding points in the current image, extracts new prominent points, and matches them with real-world locations defined in a descriptor-based map, to determine the orientation of imaging devices within the environment.

Benefits of technology

This approach enables accurate determination of the orientation of imaging devices, allowing for improved tracking and rendering of virtual content in AR systems, thereby enhancing user experience by providing a more realistic and comfortable AR experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007674454000001
    Figure 0007674454000001
  • Figure 0007674454000002
    Figure 0007674454000002
  • Figure 0007674454000003
    Figure 0007674454000003
Patent Text Reader

Abstract

To provide enhanced pose determination for a display device.SOLUTION: A head-mounted display system having an imaging device can obtain a current image of a real-world environment, with points corresponding to salient points which may be used to determine the head pose, so as to determine the head pose of a user. The salient points are patch-based and include a first salient point being projected onto the current image from a previous image, and a second salient point included in the current image being extracted from the current image. Each salient point is subsequently matched with real-world points based on descriptor-based map information indicating locations of salient points in the real-world environment. The orientation of the imaging device is determined based on the matching and based on the relative positions of the salient points in a view captured in the current image. The orientation may be used to extrapolate the head pose of a wearer of the head-mounted display system.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] (CROSS REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 62 / 599,620, filed December 15, 2017, and U.S. Provisional Application No. 62 / 623,606, filed January 30, 2018, each of which is incorporated by reference in its entirety herein.

[0002] This application further incorporates by reference the entirety of each of the following patent applications: U.S. Patent Application No. 14 / 555,585, filed November 27, 2014, and published on July 23, 2015 as U.S. Patent Publication No. 2015 / 0205126; U.S. Patent Application No. 14 / 690,401, filed April 18, 2015, and published on October 22, 2015 as U.S. Patent Publication No. 2015 / 0302652; U.S. Patent Application No. 14 / 212,961, filed March 14, 2014, and published on August 16, 2016, now U.S. Patent No. 9,417,452; No. 14 / 331,218, filed July 14, 2014 and published on October 29, 2015 as U.S. Patent Publication No. 2015 / 0309263; U.S. Patent Application No. 14 / 205,126, filed March 11, 2014 and published on October 16, 2014 as U.S. Patent Publication No. 2014 / 0306866; U.S. Patent Application No. 15 / 597,694, filed May 17, 2017; and U.S. Patent Application No. 15 / 717747, filed September 27, 2017.

[0003] (Field) The present disclosure relates to display systems, and more particularly to augmented reality display systems. [Background technology]

[0004] Description of Related Art Modern computing and display technologies have facilitated the development of systems for so-called "virtual reality" or "augmented reality" experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that appears or may be perceived as real. Virtual reality, or "VR", scenarios typically involve the presentation of digital or virtual image information without transparency to other actual real-world visual inputs, while augmented reality or "AR" scenarios typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the real world around the user. Mixed reality or "MR" scenarios are a type of AR scenario that typically involve virtual objects that are integrated into and responsive to the natural world. For example, in MR scenarios, AR image content may be perceived as appearing blocked by or otherwise interacting with objects in the real world.

[0005] 1, an augmented reality scene 10 is depicted in which a user of the AR technology sees a real-world park-like setting 20 featuring people, trees, and buildings in the background, and a concrete platform 30. In addition to these items, the user of the AR technology also perceives that they are "seeing" "virtual content," such as a robotic figure 40 standing on the real-world platform 30 and a flying, cartoon-like avatar character 50 that appears to be an anthropomorphic bumblebee, even though these elements 40, 50 do not exist in the real world. The human visual perception system is complex, making it difficult to create AR technology that facilitates a comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or real-world image elements.

[0006] The systems and methods disclosed herein address various challenges associated with AR or VR technology. Summary of the Invention [Means for solving the problem]

[0007] Some non-limiting embodiments include a system comprising one or more imaging devices, one or more processors, and one or more computer storage media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including: acquiring a current image of a real-world environment via the one or more imaging devices, the current image including a plurality of points for determining a pose, projecting patch-based first salient points from a previous image onto corresponding ones of the plurality of points in the current image, extracting second salient points from the current image, providing individual descriptors for the salient points, matching salient points associated with the current image with real-world locations defined in a descriptor-based map of the real-world environment, and determining a pose associated with the system based on the matching, the pose indicating at least an orientation of the one or more imaging devices within the real-world environment.

[0008] In the above embodiment, the operation may further include adjusting a position of the patch-based first salient point on the current image, the adjusting step including obtaining a first patch associated with the first salient point, the first patch including a portion of the previous image encompassing the first salient point and an area of ​​the previous image surrounding the first salient point, and locating a second patch in the current image similar to the first patch, the first salient point being positioned in a similar location to the first patch in the second patch. Locating the second patch may include minimizing a difference between the first patch in the previous image and the second patch in the current image. The step of projecting the patch-based first salient point onto the current image may be based at least in part on information from an inertial measurement unit of the system. Extracting the second salient points may include determining that an image area of ​​the current image has less than a threshold number of salient points projected from a previous image, and extracting one or more descriptor-based salient points from the image area, the extracted salient points including the second salient points. The image area may comprise the entirety of the current image, or the image may comprise a subset of the current image. The image area may comprise a subset of the current image, and the system may be configured to adjust a size associated with the subset based on one or more of a processing constraint or a difference between one or more previously determined poses. Matching the salient points associated with the current image with real-world locations defined in a map of the real-world environment may include accessing map information, the map information comprising real-world locations of the salient points and associated descriptors, and matching the descriptors for the salient points of the current image with the descriptors for the salient points at the real-world locations. The operations may further include a step of projecting salient points provided in the map information onto the current image, the projection being based on one or more of an inertial measurement unit, an extended Kalman filter, or visual inertial odometry.The system may be configured to generate the map using at least one or more imaging devices. The determining pose may be based on real-world locations of the salient points in a view captured in the current image and relative positions of the salient points. The operations may further include generating, for a subsequent image of the current image, a patch associated with each salient point extracted from the current image such that the patch may comprise salient points available for projecting onto the subsequent image. The providing of the descriptor may include generating a descriptor for each salient point.

[0009] In another embodiment, an augmented reality display system is provided. The augmented reality display device includes one or more imaging devices and one or more processors configured to obtain a current image of a real-world environment, perform frame / frame tracking on the current image such that patch-based salient points contained in a previous image are projected onto the current image, perform map / frame tracking on the current image such that descriptor-based salient points contained in a map database are matched with salient points of the current image, and determine a pose associated with the display device.

[0010] In the above embodiment, the frame / frame tracking may further include refining the location of the projected patch using photometric error optimization. The map / frame tracking may further include determining patch-based salient point descriptors and matching the salient point descriptors with descriptor-based salient points in the map database. The one or more processors may further be configured to generate the map database using at least one or more imaging devices. The augmented reality display system may further comprise a plurality of waveguides configured to output light with different wavefront divergences corresponding to different depth planes, the output light being located at least in part based on a pose associated with the display device.

[0011] In another embodiment, a method is provided that includes acquiring a current image of a real-world environment via one or more imaging devices, the current image including a plurality of points for determining a pose, projecting patch-based first salient points from the previous image onto corresponding ones of the plurality of points in the current image, extracting second salient points from the current image, providing individual descriptors for the salient points, matching the salient points associated with the current image with real-world locations defined in a descriptor-based map of the real-world environment, and determining a pose associated with a display device based on the matching, the pose indicating at least an orientation of the one or more imaging devices within the real-world environment.

[0012] In these embodiments, the method may further include adjusting a position of the patch-based first salient point on the current image, the adjusting the position includes obtaining a first patch associated with the first salient point, the first patch including a portion of the previous image encompassing the first salient point and an area of ​​the previous image surrounding the first salient point, and locating a second patch in the current image similar to the first patch, the first salient point being positioned in a similar location to the first patch in the second patch. Locating the second patch may include determining a patch in the current image with a minimum difference from the first patch. Projecting the patch-based first salient point onto the current image may be based at least in part on information from an inertial measurement unit of the display device. Extracting the second salient points may include determining that an image area of ​​the current image has less than a threshold number of salient points projected from a previous image, and extracting one or more descriptor-based salient points from the image area, the extracted salient points including the second salient points. The image area may comprise the entirety of the current image, or the image may comprise a subset of the current image. The image area may comprise the subset of the current image, and the processor may be configured to adjust a size associated with the subset based on one or more of a processing constraint or a difference between one or more previously determined poses. Matching the salient points associated with the current image with real-world locations defined in a map of the real-world environment may include accessing map information, the map information comprising real-world locations of the salient points and associated descriptors, and matching the descriptors for the salient points of the current image with the descriptors for the salient points at the real-world locations. The method may further include a step of projecting salient points provided in the map information onto the current image, the projection being based on one or more of an inertial measurement unit, an extended Kalman filter, or visual inertial odometry.The step of determining the pose may be based on real world locations of the salient points in a view captured in the current image and the relative positions of the salient points. The method may further include generating, for a subsequent image of the current image, patches associated with the respective salient points extracted from the current image such that the patches comprise salient points available for projecting onto the subsequent image. The step of providing a descriptor may include generating a descriptor for each salient point. The method may further include generating the map using at least one or more imaging devices. The present invention provides, for example, the following: (Item 1) 1. A system comprising: one or more imaging devices; one or more processors; One or more computer storage media having instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to: acquiring a current image of a real-world environment via the one or more imaging devices, the current image including a plurality of points for determining a pose; projecting a patch-based first salient point from a previous image onto a corresponding one of the points in the current image; Extracting a second salient point from the current image; and providing individual descriptors of said salient features; matching salient points associated with the current image with real-world locations defined within a descriptor-based map of the real-world environment; determining a pose associated with the system based on the match, the pose indicating at least an orientation of the one or more imaging devices within the real-world environment; one or more computer storage media for performing operations including: A system comprising: (Item 2) The operations further include adjusting a position of the patch-based first salient point on the current image, the adjusting the position comprising: obtaining a first patch associated with the first salient point, the first patch including a portion of the previous image that includes the first salient point and an area of ​​the previous image surrounding the first salient point; locating a second patch in the current image similar to the first patch, the first salient point being located in a similar location in the second patch as the first patch; 2. The system according to item 1, comprising: (Item 3) 3. The system of claim 2, wherein locating the second patch includes minimizing a difference between the first patch in the previous image and the second patch in the current image. (Item 4) 3. The system of claim 2, wherein projecting the patch-based first salient point onto the current image is based, at least in part, on information from an inertial measurement unit of the system. (Item 5) Extracting the second salient point includes: determining that an image area of ​​the current image has less than a threshold number of salient points projected from the previous image; extracting one or more descriptor-based salient points from the image area, the extracted salient points including the second salient point; 2. The system according to item 1, comprising: (Item 6) 6. The system of claim 5, wherein the image area comprises the entirety of the current image, or the image area comprises a subset of the current image. (Item 7) 6. The system of claim 5, wherein the image area comprises a subset of the current image, and the system is configured to adjust a size associated with the subset based on one or more of a processing constraint or a difference between one or more previously determined poses. (Item 8) Matching salient points associated with the current image with real-world locations defined within a map of the real-world environment includes: accessing map information, the map information comprising real-world locations of salient points and associated descriptors; Matching the salient feature descriptors of the current image with the salient feature descriptors of real-world locations; and 2. The system according to item 1, comprising: (Item 9) The operation further comprises: projecting salient points provided in the map information onto the current image, the projection being based on one or more of an inertial measurement unit, an extended Kalman filter, or visual inertial odometry; 8. The system according to item 7, comprising: (Item 10) 2. The system of claim 1, wherein the system is configured to generate the map using at least the one or more imaging devices. (Item 11) 2. The system of claim 1, wherein determining the pose is based on a real-world location of the salient point and a relative position of the salient point within a view captured in the current image. (Item 12) The operation further comprises: generating, for a subsequent image of the current image, patches associated with respective salient points extracted from the current image, such that the patches comprise salient points available for projecting onto the subsequent image; 2. The system according to item 1, comprising: (Item 13) 2. The system of claim 1, wherein providing a descriptor comprises generating a descriptor for each of the salient points. (Item 14) 1. An augmented reality display system, comprising: one or more imaging devices; One or more processors, the processors comprising: Obtaining a current image of a real-world environment; performing frame / frame tracking on the current image such that patch-based salient points contained in a previous image are projected onto the current image; performing map / frame tracking on the current image such that descriptor-based salient features contained in a map database are matched with salient features of the current image; determining a pose associated with the display device; one or more processors configured to An augmented reality display system comprising: (Item 15) 1. A method comprising: acquiring a current image of a real-world environment via one or more imaging devices, the current image including a plurality of points for determining a pose; projecting a patch-based first salient point from a previous image onto a corresponding one of the points in the current image; Extracting a second salient point from the current image; and providing individual descriptors of said salient features; matching salient points associated with the current image with real-world locations defined within a descriptor-based map of the real-world environment; determining a pose associated with a display device based on the match, the pose indicating at least an orientation of the one or more imaging devices within the real-world environment; A method comprising: (Item 16) and adjusting a position of the patch-based first salient point on the current image, the adjusting the position comprising: obtaining a first patch associated with the first salient point, the first patch including a portion of the previous image that includes the first salient point and an area of ​​the previous image surrounding the first salient point; locating a second patch in the current image similar to the first patch, the first salient point being located in a similar location in the second patch as the first patch; Item 16. The method according to item 15, comprising: (Item 17) Item 17. The method of item 16, wherein locating the second patch includes determining a patch in the current image with a minimum difference from the first patch. (Item 18) Item 17. The method of item 16, wherein projecting the patch-based first salient point onto the current image is based, at least in part, on information from an inertial measurement unit of the display device. (Item 19) Extracting the second salient point includes: determining that an image area of ​​the current image has less than a threshold number of salient points projected from the previous image; extracting one or more descriptor-based salient points from the image area, the extracted salient points including the second salient point; Item 16. The method according to item 15, comprising: (Item 20) 20. The method of claim 19, wherein the image area comprises the entirety of the current image, or the image area comprises a subset of the current image. (Item 21) 20. The method of claim 19, wherein the image area comprises a subset of the current image, and the processor is configured to adjust a size associated with the subset based on one or more of processing constraints or differences between one or more previously determined poses. (Item 22) Matching salient points associated with the current image with real-world locations defined within a map of the real-world environment includes: accessing map information, the map information comprising real-world locations of salient points and associated descriptors; Matching the salient feature descriptors of the current image with the salient feature descriptors of real-world locations; and Item 16. The method according to item 15, comprising: (Item 23) projecting salient points provided in the map information onto the current image, the projection being based on one or more of an inertial measurement unit, an extended Kalman filter, or visual inertial odometry; 23. The method of claim 22, further comprising: (Item 24) Item 16. The method of item 15, wherein determining the pose is based on real-world locations of the salient points and relative positions of the salient points within a view captured in the current image. (Item 25) generating, for a subsequent image of the current image, patches associated with respective salient points extracted from the current image, such that the patches comprise salient points available for projecting onto the subsequent image; Item 16. The method of item 15, further comprising: (Item 26) Item 16. The method of item 15, wherein providing a descriptor comprises generating a descriptor for each of the salient points. (Item 27) Item 16. The method of item 15, further comprising generating the map using at least the one or more imaging devices. [Brief description of the drawings]

[0013] [Figure 1] FIG. 1 illustrates a user's view of an augmented reality (AR) device.

[0014] [Diagram 2] FIG. 2 illustrates a conventional display system for simulating a three-dimensional image for a user.

[0015] [Diagram 3] 3A-3C illustrate the relationship between the radius of curvature and the radius of focus.

[0016] [Figure 4A] FIG. 4A illustrates a representation of the accommodation-vergence response of the human visual system.

[0017] [Figure 4B] FIG. 4B illustrates examples of different accommodation and convergence states of a pair of a user's eyes.

[0018] [Figure 4C] FIG. 4C illustrates an example of a top-down view representation of a user viewing content through a display system.

[0019] [Figure 4D] FIG. 4D illustrates another example of a top-down view representation of a user viewing content through a display system.

[0020] [Diagram 5] FIG. 5 illustrates aspects of an approach for simulating a three-dimensional image by modifying wavefront divergence.

[0021] [Figure 6] FIG. 6 illustrates an embodiment of a waveguide stack for outputting image information to a user.

[0022] [Figure 7] FIG. 7 illustrates an example of an output beam output by a waveguide.

[0023] [Figure 8] FIG. 8 illustrates an example of a stacked waveguide assembly where each depth plane contains an image formed using multiple different primary colors.

[0024] [Figure 9A] FIG. 9A illustrates a cross-sectional side view of an example of a set of stacked waveguides, each of which includes an internal coupling optical element.

[0025] [Figure 9B] FIG. 9B illustrates a perspective view of the multiple stacked waveguide embodiment of FIG. 9A.

[0026] [Figure 9C] FIG. 9C illustrates a top-down plan view of the multiple stacked waveguide embodiment of FIGS. 9A and 9B.

[0027] [Figure 9D] FIG. 9D illustrates an example of a wearable display system.

[0028] [Figure 10A] FIG. 10A illustrates a flowchart of an exemplary process for determining the pose of a display system and the pose of a user's head.

[0029] [Figure 10B] FIG. 10B illustrates an example image area of ​​the current image.

[0030] [Figure 11] FIG. 11 illustrates a flow chart of an exemplary process for frame / frame tracking.

[0031] [Figure 12A] FIG. 12A illustrates an example of a previous image and a current image.

[0032] [Figure 12B]FIG. 12B illustrates an example of a patch projected onto the current image of FIG. 12A.

[0033] [Figure 13] FIG. 13 illustrates a flow chart of an exemplary process for map / frame tracking.

[0034] [Figure 14A] FIG. 14A illustrates an example of a previous image and a current image.

[0035] [Figure 14B] FIG. 14B illustrates an example of frame / frame tracking after map / frame tracking.

[0036] [Figure 15] FIG. 15 illustrates a flowchart of an exemplary process for determining head pose. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0037] Display systems such as augmented reality (AR) or virtual reality (VR) display systems can present content to a user (or viewer) in different areas of the user's field of view. For example, an augmented reality display system may present virtual content to a user, which may appear to the user as placed in a real-world environment. As another example, a virtual reality display system can present content via a display such that the content may appear to the user as being three-dimensional and placed in a three-dimensional environment. For example, the placement of this content relative to the user may have a positive or negative impact on the realism associated with the presented content and the comfort of the user wearing the display system. Because the placement of content may depend on the head pose of the user of the display system, these display systems can be improved through the use of accurate schemes for determining head pose, as will be described below.

[0038] The pose of the user's head may be understood to be the orientation of the user's head (e.g., head pitch, yaw, and / or roll) relative to the real-world environment, e.g., relative to a coordinate system associated with the real-world environment. In some embodiments, the display system may also have a pose that corresponds to a particular orientation of the display system (e.g., an AR or VR display device) or a portion of the display system, relative to the real-world environment, e.g., relative to a coordinate system associated with the real-world environment. The pose may optionally generally represent an orientation in the real-world environment relative to the coordinate system. For example, when a user rotates a display system mounted on their head (e.g., by rotating their head), the pose of both the user's head and the display system may be adjusted according to the rotation. Thus, the content being presented to the user may be adjusted based on the pose of the user's head, which may also change the pose of the display of the display system mounted on the user's head. In some embodiments, the pose of the display system may be determined and the user's head pose may be extrapolated from this display system pose. By determining head pose as a user moves around a real-world environment, content can be realistically adjusted in location and orientation based on the determined pose of the user's head. Several examples are described below.

[0039] With regard to augmented reality (AR) and virtual reality (VR) display systems, realism can be enhanced if the user can move around the presented virtual content and the presented virtual content can appear to remain substantially in a fixed real-world location. For example, the robot 40 image illustrated in FIG. 1 can be the virtual content presented to the user. As the user walks towards or around the robot 40, the augmented reality scene 10 would appear more realistic to the user if the robot 40 appeared to remain in the same location in the park. Thus, the user can see different perspectives of the robot 40, different parts of the robot 40, etc. To ensure that the robot 40 appears as a fixed, realistic image, the display system can utilize the determined pose when rendering the virtual content. For example, the display system can obtain information indicating that the user has rotated its head at a particular angle. This rotation can inform the placement of the virtual content such that the robot 40 will appear to remain standing upright as an image.

[0040] As another example, a user may play a first-person video game while wearing a display system. In this example, the user may dive their head or rotate their head to move to avoid a virtual enemy object jumping on the user, as presented to the user via the display system. This movement (e.g., head dive or rotation) can be tracked and the user's head pose can be determined. In this way, the display system can determine whether the user has successfully avoided the enemy object.

[0041] Systems for determining pose can be complex. Exemplary schemes for determining head pose can utilize sensors and emitters of light. For example, an infrared emitter can emit pulses of infrared light from a fixed location in the real-world environment (e.g., the emitter can be in a room surrounding the device). A display device worn by the user can include a sensor to detect these pulses. The display device can thus determine its orientation relative to the fixed emitter. Similarly, the display device can determine its position in the real-world environment based on the fixed emitter. As another example, the display device may include a fixed emitter of light (e.g., visible or infrared light), and one or more cameras may be positioned in the real-world environment that track the emission of the light. In this example, as the display device rotates, the camera can detect that the emission of the light is rotating from an initial position. These exemplary schemes may therefore require complex hardware to determine the pose of the display device.

[0042] A display system described herein (e.g., display system 60 illustrated in FIG. 9D) can determine an accurate pose estimate without the complexity and rigidity of a fixed emitter of light. In some embodiments, the pose of a display system worn by a user (e.g., display 70 coupled to frame 80 as illustrated in FIG. 9D) may be determined. From the display system pose, the user's head pose may be determined. The display system can determine its pose without requiring the user to set up complex hardware in a room and without requiring the user to remain in the room. As will be described, the display system can determine its pose through the use of an imaging device on the display system (e.g., an optical imaging device such as a camera). The imaging device can acquire images of the real-world environment and based on these images, the display system can determine its pose. Through the techniques described herein, the display system can advantageously determine pose while limiting processing and memory requirements. In this manner, display systems with limited memory and computational budgets, such as AR and MR display systems, can efficiently determine pose estimates and increase realism for the user.

[0043] To determine pose, the display system may utilize both (1) patch-based tracking of distinguishable points (e.g., distinctly different isolated portions of the images) between successive images of the environment captured by the display system (referred to herein as “frame-frame tracking”), and (2) matching points of interest in the current image with a descriptor-based map of known real-world locations of corresponding points of interest (referred to herein as “map-frame tracking”). In frame-frame tracking, the display system may track specific points of interest (referred to herein as “salient points”), such as corners between captured images of the real-world environment. For example, the display system may identify locations of visual points of interest in the current image that were contained (e.g., located within) a previous image. This identification may be accomplished, for example, using a photometric error minimization process. In map-frame tracking, the display system may access map information indicating the real-world locations (e.g., three-dimensional coordinates) of the points of interest, and match points of interest contained in the current image to points of interest indicated in the map information. Information about the points of interest may be stored as a descriptor, for example in a map database. The display system can then calculate its pose based on the matched visual features. The step of generating the map information will be described in more detail below with respect to FIG. 10A. As used herein, in some embodiments, a point may refer to a discrete pixel of an image or a set of pixels corresponding to an area of ​​an image.

[0044] As described above, to determine pose, the display system can utilize distinguishable visual features, referred to herein as "salient points." As used herein, a salient point corresponds to any unique portion of a real-world environment that can be tracked. For example, a salient point can be a corner. A corner can represent a substantially perpendicular intersection of two lines and can include a scratch on a desk, a mark on a wall, a keyboard number "7," etc. As will be described, corners can be detected from images acquired by an imaging device according to a corner detection scheme. Exemplary corner detection schemes can include Harris corner detection, features from accelerated segment test (FAST) corner detection, etc.

[0045] With regard to frame / frame tracking, the display system can track salient points from a previous image to a current image via projecting each tracked salient point from the previous image onto the current image. For example, the display system can determine an optical flow between the current image and the previous image using trajectory prediction or, optionally, using information from an inertial measurement unit. The optical flow can represent the movement of the user from the time the previous image was acquired to the time the current image was acquired. The trajectory prediction can inform the location in the current image to which the salient points contained in the previous image correspond. The display system can then obtain image portions, known herein as "patches," that surround each salient point in the previous image and determine matching image portions in the current image. A patch can be, for example, an M×N pixel area that surrounds each salient point in the previous image, where M and N are positive integers. To match a patch from a previous image to a current image, the display system can identify a patch in the current image in which the photometric error between the patch and the previous image patch is reduced (e.g., minimized). A salient point may be understood to be located at a particular consistent two-dimensional image location within a patch. For example, the centroid of a matching patch in the current image may correspond to the tracked salient point. Thus, a projection from a previous image onto the current image roughly locates the salient point and associated patch in the current image, the location of which may be refined using, for example, photometric error minimization to determine a location that minimizes pixel intensity differences between the patch and a particular area of ​​the current image.

[0046] With respect to map / frame tracking, the display system can extract salient points from the current image (e.g., identify locations in the current image that correspond to new salient points). For example, the display can extract salient points from image areas of the current image that have less than a threshold number of tracked salient points (e.g., determined from frame / frame tracking). The display system can then match the salient points (e.g., newly extracted salient points, tracked salient points) in the current image to distinct real-world locations based on the descriptor-based map information. As described herein, the display system can generate a descriptor for each salient point that uniquely describes the (e.g., visual) attributes of the salient point. The map information can similarly store the descriptors for the real-world salient points. Based on the matching descriptors, the display system can determine the real-world locations of the salient points contained within the current image. Thus, the display system can determine its orientation relative to the real-world location and determine its pose, which can then be used to determine head pose.

[0047] It should be appreciated that the use of a photometric error minimization scheme may enable highly accurate tracking of salient points between images, for example, through comparison of patches as described above. Indeed, sub-pixel accuracy in tracking salient points between a previous image and a current image may be achieved. In contrast, a descriptor may be less accurate in tracking salient points between images, but would utilize less memory than a patch for photometric error minimization. Because the descriptor may be less accurate in tracking salient points between images, the determined pose estimate may vary more than if photometric error minimization were used. Although accurate, the use of a patch may require storing a patch for each salient point. Because a descriptor may be an alphanumeric value that describes the visual characteristics of a salient point and / or an image area surrounding a salient point, such as a histogram, a descriptor may be an order of magnitude or more smaller than a patch.

[0048] Thus, as described herein, the display device may take advantage of patch-based photometric error minimization and descriptors to enable a robust and memory-efficient pose determination process. For example, frame / frame tracking may take advantage of patch-based photometric error minimization to accurately track salient points between images. In this way, salient points may be tracked, for example, with sub-pixel accuracy. However, over time (e.g., across multiple frames or images), slight errors may be introduced such that a drift caused by cumulative errors in tracking salient points over a threshold number of images may become evident. This drift may reduce the accuracy of pose determination. Thus, in some embodiments, map / frame tracking may be utilized to link each salient point to a real-world location. For example, in map / frame tracking, salient points are matched to salient points stored in map information. Thus, the real-world coordinates of each salient point may be identified.

[0049] If photometric error minimization is utilized for map / frame tracking, the map information would store a patch for each salient point identified in the real-world environment. Since there may be thousands, hundreds of thousands, etc. of salient points represented in the map information, memory requirements would be enormous. Advantageously, using a descriptor can reduce memory requirements associated with map / frame tracking. For example, the map information can store the real-world coordinates of each salient point along with a descriptor for the salient point. In some embodiments, the descriptor can be at least an order of magnitude smaller in size than the patch, so the map information can be significantly reduced.

[0050] As will be described below, the display system can therefore leverage both patch-based frame-to-frame tracking and descriptor-based map-to-frame tracking. For example, the display system can track salient points between successive images acquired of a real-world environment. As described above, tracking salient points can include projecting salient points from a previous image onto a current image. Through the use of patch-based photometric error minimization, the locations of the tracked salient points can be determined within the current image with high accuracy. The display system can then identify image areas of the current image that contain tracked salient points below a threshold measurement. For example, the current image can be separated into different image areas, each image area being ¼, ⅛, 1 / 16, 1 / 32, a user-selectable size, etc., of the current image. As another example, the display system can analyze the sparseness of the current image with respect to the tracked salient points. In this example, the display system can determine whether any area of ​​the image (e.g., an area of ​​a threshold size) contains less than a threshold number of salient points or a density of salient points below a threshold. Optionally, the image area can be the entire current image, such that the display system can identify whether the entire current image contains tracked salient points below the threshold measurement. The display system can then extract new salient points from the identified image area and generate a descriptor for each salient point of the current image (e.g., tracked salient points and newly extracted salient points). Through matching each generated descriptor to the descriptors of salient points indicated in the map information, a real-world location of each salient point in the current image can be identified. Thus, the pose of the display system can be determined. The salient points contained in the current image can then be tracked in subsequent images, for example, as described herein.

[0051] Salient point tracking may utilize a potentially large amount of identical tracked salient points between successive image frames, since new salient points may only be extracted in image areas with tracked salient points below the threshold measurement. As explained above, tracking is performed via photometric error minimization, which may ensure highly accurate localization of salient points between images. In addition, jitter in pose determination may also be reduced, since these identical tracked salient points will be matched to map information in successive image frames. Furthermore, processing requirements may also be reduced, since the display system may only be required to extract new salient points in specific image areas. In addition, drift in pose determination may be reduced, since salient points in the current image are matched to map information. Optionally, map / frame tracking may not be required for some current images. For example, a user may be viewing a substantially similar real-world area, such that the display system may maintain a similar pose. In this embodiment, the current image may not include image areas with tracked salient points below the threshold measurement. Thus, only frame / frame tracking may be utilized to determine the pose of the display system. Optionally, map / frame tracking may be utilized even if an image area does not contain less than a threshold number of tracked salient points. For example, descriptors can be generated for tracked salient points and compared to the map information without extracting new salient points. In this way, the display system can perform less processing, thus saving processing resources and reducing energy consumption.

[0052] Reference is now made to the drawings in which like reference numbers refer to like parts throughout. Unless specifically indicated otherwise, the drawings are schematic and are not necessarily drawn to scale.

[0053] FIG. 2 illustrates a conventional display system for simulating a three-dimensional image for a user. It should be understood that when a user's eyes are spaced apart and viewing a real object in space, each eye may have a slightly different view of the object and form an image of the object at a different location on the retina of each eye. This may be referred to as binocular disparity and may be exploited by the human visual system to provide a perception of depth. Conventional display systems simulate binocular disparity by presenting two distinctly different images 190, 200 with slightly different views of the same virtual object, one for each eye 210, 220, corresponding to the view of the virtual object as it would appear by each eye as if the virtual object were a real object at a desired depth. These images provide binocular cues that the user's visual system may interpret to derive a perception of depth.

[0054] Continuing with reference to FIG. 2, the images 190, 200 are spaced from the eyes 210, 220 by a distance 230 on the z-axis, which is parallel to the optical axis of the viewer when that eye is fixating on an object at optical infinity directly in front of the viewer. The images 190, 200 are flat and at a fixed distance from the eyes 210, 220. Based on slightly different views of the virtual object in the images presented to the eyes 210, 220, respectively, the eyes may naturally rotate such that the image of the object falls on a corresponding point on the eye's respective retina, maintaining single binocular vision. This rotation may cause the line of sight of each of the eyes 210, 220 to converge on a point in space where the virtual object is perceived to reside. As a result, providing three-dimensional images traditionally involves manipulating the convergence and divergence movements of the user's eyes 210, 220 to provide binocular cues that the human visual system interprets to provide the perception of depth.

[0055] However, creating a realistic and comfortable perception of depth is difficult. It should be understood that light from an object at different distances from the eye has a wavefront with different amounts of divergence. Figures 3A-3C illustrate the relationship between distance and divergence of light rays. The distance between the object and the eye 210 is represented in the order of decreasing distances R1, R2, and R3. As shown in Figures 3A-3C, the light rays become more divergent as the distance to the object decreases. Conversely, as the distance increases, the light rays become more collimated. In other words, the light field generated by a point (an object or part of an object) can be said to have a spherical wavefront curvature that is a function of the distance the point is away from the user's eye. As the curvature increases, the distance between the object and the eye 210 decreases. Although only a single eye 210 is illustrated in Figures 3A-3C and various other figures herein for clarity of illustration, the discussion regarding the eye 210 may apply to both eyes 210 and 220 of a viewer.

[0056] 3A-3C, light from an object that a viewer's eye is fixating on may have different wavefront divergences. Due to the different wavefront divergences, the light may be focused differently by the eye's lens, which in turn may require the lens to assume a different shape and form a focused image on the eye's retina. If a focused image is not formed on the retina, the resulting retinal blur acts as an accommodative cue, causing the shape of the eye's lens to change until a focused image is formed on the retina. For example, the accommodative cue may trigger the relaxation or contraction of the ciliary muscle surrounding the eye's lens, thereby modulating the force applied to the suspensory ligament that holds the lens, thus changing the shape of the eye's lens, thereby forming a focused image of the fixated object on the eye's retina (e.g., fovea) until retinal blur of the fixated object is eliminated or minimized. The process by which the eye lens changes shape may be referred to as accommodation, and the shape of the eye lens required to form a focused image of a fixated object on the eye's retina (e.g., the fovea) may be referred to as the state of accommodation.

[0057] Referring now to FIG. 4A, a representation of the accommodation-vergence response of the human visual system is illustrated. A movement of the eye to fixate an object causes the eye to receive light from the object, which forms an image on each of the eye's retinas. The presence of retinal blur in the image formed on the retina may provide a cue for accommodation, and the relative location of the image on the retina may provide a cue for vergence. A cue for accommodation causes accommodation, resulting in the eye adopting a particular accommodation state in which the eye's lens forms a focused image of the object on the eye's retina (e.g., the fovea). On the other hand, a cue for vergence causes a vergence movement (eye rotation) such that the image formed on each retina of each eye is at a corresponding retinal point that maintains single binocular vision. In these positions, the eye may be said to adopt a particular vergence state. Continuing with reference to FIG. 4A, accommodation can be understood to be the process by which the eye achieves a particular accommodation state, and convergence can be understood to be the process by which the eye achieves a particular convergence state. As shown in FIG. 4A, the accommodation and convergence state of the eye can change when the user fixates a different object. For example, the accommodated state can change when the user fixates a new object at a different depth on the z-axis.

[0058] Without being limited by theory, it is believed that a viewer of an object may perceive the object as being "three-dimensional" due to a combination of vergence-divergence and accommodation. As previously mentioned, vergence-divergence movements of the two eyes relative to one another (e.g., rotation of the eyes such that the pupils move toward or away from one another, converging the gaze of the eyes to fixate on an object) are closely linked to accommodation of the eye's lenses. Under normal conditions, changing the shape of the eye's lenses and shifting focus from one object to another at a different distance will automatically produce a matching change in vergence to the same distance, in a relationship known as the "accommodation-vergence-divergence reflex." Similarly, a change in vergence will trigger a matching change in lens shape under normal conditions.

[0059] 4B, an example of different accommodation and convergence-divergence states of the eyes is illustrated. Paired eye 222a fixates an object at optical infinity, while paired eye 222b fixates an object 221 at less than optical infinity. Notably, the convergence-divergence states of each pair of eyes are different, paired eye 222a points straight ahead, while paired eye 222 converges on object 221. The accommodation states of the eyes forming each pair of eyes 222a and 222b are also different, as represented by the different shapes of the lenses 210a, 220a.

[0060] Unfortunately, many users of conventional "3-D" display systems may find such conventional systems uncomfortable or may not perceive any depth sensation due to a mismatch between accommodation and vergence states in these displays. As previously mentioned, many stereoscopic or "3-D" display systems display a scene by providing a slightly different image to each eye. Such systems are uncomfortable for many viewers because, among other things, they simply provide a different presentation of the scene, causing changes in the vergence states of the eyes, but without a corresponding change in the accommodation states of those eyes. Rather, images are presented by the display at a fixed distance from the eyes, such that the eyes view all image information in a single accommodation state. Such an arrangement counters the "accommodation-vergence-divergence reflex" by causing changes in vergence states without a matching change in accommodation state. This mismatch is believed to cause viewer discomfort. By providing better matching between accommodation and convergence-divergence movements, display systems may form a more realistic and comfortable simulation of three-dimensional images.

[0061] Without being limited by theory, it is believed that the human eye is typically capable of interpreting a finite number of depth planes to provide depth perception. As a result, a highly realistic simulation of perceived depth may be achieved by providing the eye with different presentations of images corresponding to each of these limited number of depth planes. In some embodiments, the different presentations may provide both vergence cues and matching cues for accommodation, thereby providing physiologically correct accommodation-vergence divergence matching.

[0062] 4B, two depth planes 240 are illustrated, corresponding to different distances in space from the eyes 210, 220. For a given depth plane 240, vergence-divergence cues may be provided by displaying appropriately different perspective images for each eye 210, 220. Additionally, for a given depth plane 240, the light forming the image provided to each eye 210, 220 may have a wavefront divergence corresponding to the light field generated by a point at the distance of that depth plane 240.

[0063] In the illustrated embodiment, the distance along the z-axis of the depth plane 240 containing the point 221 is 1 m. As used herein, the distance or depth along the z-axis may be measured with the zero point located at the exit pupil of the user's eye. Thus, the depth plane 240 located at a depth of 1 m corresponds to a distance of 1 m away from the exit pupil of the user's eye on the optical axis of the eye with the eye pointed towards optical infinity. As an approximation, the depth or distance along the z-axis may be measured from the display (e.g., the surface of the waveguide) in front of the user's eye, and a value for the distance between the device and the exit pupil of the user's eye may be added. That value may be called the pupil distance and may correspond to the distance between the exit pupil of the user's eye and the display worn by the user in front of the eye. In practice, the value for the pupil distance may be a normalized value that is generally used for all viewers. For example, the pupil distance may be assumed to be 20 mm, and the depth plane at a depth of 1 m may be at a distance of 980 mm in front of the display.

[0064] 4C and 4D, examples of matched accommodation-vergence-divergence distances and mismatched accommodation-vergence-divergence distances are illustrated, respectively. As illustrated in FIG. 4C, the display system may provide an image of a virtual object to each eye 210, 220. The image may cause the eyes 210, 220 to assume a convergence-divergence state in which the eyes converge on point 15 on the depth plane 240. In addition, the image may be formed by light having a wavefront curvature corresponding to a real object in that depth plane 240. As a result, the eyes 210, 220 assume an accommodation state in which the image is focused on the retinas of those eyes. Thus, the user may perceive the virtual object as being at point 15 on the depth plane 240.

[0065] It should be appreciated that each of the accommodation and convergence states of the eyes 210, 220 is associated with a particular distance on the z-axis. For example, an object at a particular distance from the eyes 210, 220 will cause those eyes to assume a particular accommodation state based on the distance of the object. The distance associated with a particular accommodation state is referred to as the accommodation distance A. d Similarly, a particular convergence-divergence distance V associated with the eyes in a particular convergence-divergence state may be referred to as d or positions relative to each other. When the accommodation distance and the vergence distance match, the relationship between accommodation and vergence is said to be physiologically correct. This is considered to be the most comfortable scenario for the viewer.

[0066] However, in a stereoscopic display, the accommodation distance and the vergence distance may not always match. For example, as illustrated in FIG. 4D, the images displayed to the eyes 210, 220 may be displayed with a wavefront divergence corresponding to the depth plane 240, and the eyes 210, 220 may be in a particular accommodation state in which the points 15a, 15b on the depth plane are in focus. However, the images displayed to the eyes 210, 220 may provide a cue for vergence that causes the eyes 210, 220 to converge on a point 15 that is not located on the depth plane 240. As a result, the accommodation distance corresponds in some embodiments to the distance from the exit pupils of the eyes 210, 220 to the depth plane 240, while the vergence distance corresponds to a larger distance from the exit pupils of the eyes 210, 220 to the point 15. The accommodation distance is different from the vergence distance. As a result, there is an accommodation-vergence-divergence mismatch. Such a mismatch may be considered undesirable and may cause discomfort to the user. The mismatch may be related to distance (e.g., V d -A d ) and can be characterized in terms of diopters.

[0067] It should be understood that in some embodiments, reference points other than the exit pupils of the eyes 210, 220 may be utilized to determine distances for determining accommodation-vergence-divergence mismatch, so long as the same reference points are utilized for accommodation distance and vergence-divergence distance. For example, distances may be measured from the cornea to the depth plane, from the retina to the depth plane, from the eyepiece (e.g., a waveguide in a display device) to the depth plane, etc.

[0068] Without being limited by theory, it is believed that a user may still perceive an accommodation-vergence-divergence mismatch of up to about 0.25 diopters, up to about 0.33 diopters, and up to about 0.5 diopters as physiologically correct, without the mismatch itself causing significant discomfort. In some embodiments, a display system disclosed herein (e.g., display system 250, FIG. 6) presents an image to a viewer having an accommodation-vergence-divergence mismatch of about 0.5 diopters or less. In some other embodiments, the accommodation-vergence-divergence mismatch of the image provided by the display system is about 0.33 diopters or less. In still other embodiments, the accommodation-vergence-divergence mismatch of the image provided by the display system is about 0.25 diopters or less, including about 0.1 diopters or less.

[0069] FIG. 5 illustrates aspects of an approach for simulating a three-dimensional image by modifying wavefront divergence. The display system includes a waveguide 270 configured to receive light 770 encoded with image information and output the light to a user's eye 210. The waveguide 270 may output light 650 with a defined amount of wavefront divergence that corresponds to the wavefront divergence of a light field generated by a point on a desired depth plane 240. In some embodiments, the same amount of wavefront divergence is provided for all objects presented on that depth plane. In addition, the user's other eye will be illustrated as being provided with image information from a similar waveguide.

[0070] In some embodiments, a single waveguide may be configured to output light with a set wavefront divergence corresponding to a single or limited number of depth planes, and / or the waveguide may be configured to output light of a limited range of wavelengths. As a result, in some embodiments, multiple or stacked waveguides may be utilized to provide different wavefront divergence for different depth planes and / or output light of different ranges of wavelengths. As used herein, it will be understood that a depth plane may follow the contour of a flat or curved surface. In some embodiments, advantageously and for convenience, a depth plane may follow the contour of a flat surface.

[0071] 6 illustrates an example of a waveguide stack for outputting image information to a user. The display system 250 includes a stack of waveguides or a stacked waveguide assembly 260 that can be utilized to provide a three-dimensional perception to the eye / brain using multiple waveguides 270, 280, 290, 300, 310. It should be understood that the display system 250 may be considered a light field display in some embodiments. Additionally, the waveguide assembly 260 may also be referred to as an eyepiece.

[0072] In some embodiments, the display system 250 may be configured to provide a substantially continuous cue for convergence and multiple discrete cues for accommodation. The cues for convergence may be provided by displaying different images to each of the user's eyes, and the cues for accommodation may be provided by outputting light forming images with selectable discrete amounts of wavefront divergence. In other words, the display system 250 may be configured to output light with variable levels of wavefront divergence. In some embodiments, each discrete level of wavefront divergence corresponds to a particular depth plane and may be provided by a particular one of the waveguides 270, 280, 290, 300, 310.

[0073] Continuing with reference to FIG. 6, the waveguide assembly 260 may also include a number of features 320, 330, 340, 350 between the waveguides. In some embodiments, the features 320, 330, 340, 350 may be one or more lenses. The waveguides 270, 280, 290, 300, 310 and / or the number of lenses 320, 330, 340, 350 may be configured to transmit image information to the eye with various levels of wavefront curvature or light beam divergence. Each waveguide level may be associated with a particular depth plane and may be configured to output image information corresponding to that depth plane. The image input devices 360, 370, 380, 390, 400 may act as light sources for the waveguides and may be utilized to input image information into the waveguides 270, 280, 290, 300, 310, each of which may be configured to disperse incident light across each individual waveguide for output towards the eye 210, as described herein. Light exits output surfaces 410, 420, 430, 440, 450 of the image input devices 360, 370, 380, 390, 400 and is input into corresponding input surfaces 460, 470, 480, 490, 500 of the waveguides 270, 280, 290, 300, 310. In some embodiments, each of the input surfaces 460, 470, 480, 490, 500 may be an edge of the corresponding waveguide or may be a portion of the major surface of the corresponding waveguide (i.e., one of the waveguide surfaces that directly faces the world 510 or the viewer's eye 210). In some embodiments, a single beam of light (e.g., a collimated beam) may be launched into each waveguide, outputting a total field of cloned collimated beams that are directed towards the eye 210 at a particular angle (and divergence) corresponding to the depth plane associated with the particular waveguide. In some embodiments, a single one of the image launch devices 360, 370, 380, 390, 400 may be associated with and launch light into multiple (e.g., three) waveguides 270, 280, 290, 300, 310.

[0074] In some embodiments, each of the image input devices 360, 370, 380, 390, 400 is a discrete display that generates image information for input into the respective waveguides 270, 280, 290, 300, 310. In some other embodiments, the image input devices 360, 370, 380, 390, 400 are the output of a single multiplexed display that may, for example, send image information to each of the image input devices 360, 370, 380, 390, 400 via one or more optical conduits (such as fiber optic cables). It should be understood that the image information provided by the image input devices 360, 370, 380, 390, 400 may include light of different wavelengths or colors (e.g., different primary colors as discussed herein).

[0075] In some embodiments, the light injected into the waveguides 270, 280, 290, 300, 310 is provided by a light projector system 520, which includes a light module 530, which may include a light emitter such as a light emitting diode (LED). The light from the light module 530 may be directed and modified by a light modulator 540, e.g., a spatial light modulator, via a beam splitter 550. The light modulator 540 may be configured to vary the perceived intensity of the light injected into the waveguides 270, 280, 290, 300, 310 and encode the light with image information. Examples of spatial light modulators include liquid crystal displays (LCDs), including liquid crystal on silicon (LCOS) displays. It should be understood that image injection devices 360, 370, 380, 390, 400 are illustrated diagrammatically and in some embodiments these image injection devices may represent different light paths and locations within a common projection system configured to output light into associated ones of waveguides 270, 280, 290, 300, 310. In some embodiments, the waveguides of the waveguide assembly 260 may act as ideal lenses, relaying light injected into the waveguides to the user's eye. In this concept, the object may be a spatial light modulator 540 and the image may be an image on a depth plane.

[0076] In some embodiments, the display system 250 may be a scanning fiber display comprising one or more scanning fibers configured to project light in various patterns (e.g., raster scan, spiral scan, Lissajous pattern, etc.) into one or more waveguides 270, 280, 290, 300, 310 and ultimately to the viewer's eye 210. In some embodiments, the illustrated image injection devices 360, 370, 380, 390, 400 may diagrammatically represent a single scanning fiber or a bundle of scanning fibers configured to inject light into one or more waveguides 270, 280, 290, 300, 310. In some other embodiments, the illustrated image injection devices 360, 370, 380, 390, 400 may diagrammatically represent multiple scanning fibers or multiple bundles of scanning fibers, each configured to inject light into an associated one of the waveguides 270, 280, 290, 300, 310. It should be understood that one or more optical fibers may be configured to transmit light from the optical module 530 to one or more of the waveguides 270, 280, 290, 300, 310. It should be understood that one or more intervening optical structures may be provided between the scanning fiber or fibers and one or more of the waveguides 270, 280, 290, 300, 310, for example, to redirect light exiting the scanning fiber into one or more of the waveguides 270, 280, 290, 300, 310.

[0077] The controller 560 controls the operation of one or more of the stacked waveguide assemblies 260, including the operation of the image input devices 360, 370, 380, 390, 400, the light source 530, and the light module 540. In some embodiments, the controller 560 is part of the local data processing module 140. The controller 560 includes programming (e.g., instructions in a non-transient medium) that coordinates the timing and provisioning of image information to the waveguides 270, 280, 290, 300, 310, for example, according to any of the various schemes disclosed herein. In some embodiments, the controller may be a single integrated device or may be a distributed system connected by wired or wireless communication channels. The controller 560 may be part of the processing module 140 or 150 (FIG. 2) in some embodiments.

[0078] Continuing with reference to FIG. 6, the waveguides 270, 280, 290, 300, 310 may be configured to propagate light within each individual waveguide by total internal reflection (TIR). Each of the waveguides 270, 280, 290, 300, 310 may be planar or have another shape (e.g., curved) with major top and bottom surfaces and edges extending between the major top and bottom surfaces. In the illustrated configuration, each of the waveguides 270, 280, 290, 300, 310 may include an outcoupling optical element 570, 580, 590, 600, 610 configured to extract light from the waveguide by redirecting light propagating within each individual waveguide out of the waveguide and outputting image information to the eye 210. The extracted light may also be referred to as outcoupling light, and the outcoupling optical element light may also be referred to as a light extraction optical element. The extracted beam of light may be output by the waveguide at the location where the light propagating within the waveguide strikes the light extraction optical element. The outcoupling optical element 570, 580, 590, 600, 610 may be a grating, for example, including diffractive optical features as further discussed herein. Although shown disposed on the bottom major surface of the waveguide 270, 280, 290, 300, 310 for ease of explanation and clarity of the drawings, in some embodiments the outcoupling optical element 570, 580, 590, 600, 610 may be disposed on the top major surface and / or bottom major surface and / or directly within the volume of the waveguide 270, 280, 290, 300, 310, as further discussed herein. In some embodiments, the outcoupling optical elements 570, 580, 590, 600, 610 may be formed within a layer of material that is attached to a transparent substrate and forms the waveguide 270, 280, 290, 300, 310. In some other embodiments, the waveguide 270, 280, 290, 300, 310 may be a monolithic piece of material and the outcoupling optical elements 570, 580, 590, 600, 610 may be formed on a surface of and / or within that piece of material.

[0079] With continued reference to FIG. 6, as discussed herein, each waveguide 270, 280, 290, 300, 310 is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 270 closest to the eye may be configured to deliver collimated light (injected into such waveguide 270) to the eye 210. The collimated light may represent an optical infinity focal plane. The next upper waveguide 280 may be configured to send collimated light that passes through a first lens 350 (e.g., a negative lens) before it can reach the eye 210. Such first lens 350 may be configured to generate a slight convex wavefront curvature such that the eye / brain interprets the light originating from the next upper waveguide 280 as originating from a first focal plane closer inward from optical infinity toward the eye 210. Similarly, the third upper waveguide 290 passes its output light through both the first lens 350 and the second lens 340 before reaching the eye 210. The combined refractive powers of the first lens 350 and the second lens 340 may be configured to produce another incremental amount of wavefront curvature such that the eye / brain interprets the light emerging from the third waveguide 290 as emerging from a second focal plane that is closer inward from optical infinity towards the person than was the light from the next upper waveguide 280.

[0080] The other waveguide layers 300, 310 and lenses 330, 320 are similarly configured, with the highest waveguide 310 in the stack sending its output through all of the lenses between it and the eye for an aggregate focal power that represents the focal plane closest to the person. To compensate the stack of lenses 320, 330, 340, 350 when viewing / interpreting light originating from the world 510 on the other side of the stacked waveguide assembly 260, a compensating lens layer 620 may be placed on top of the stack to compensate for the aggregate power of the lens stacks 320, 330, 340, 350 below. Such a configuration provides as many perceived focal planes as there are waveguide / lens pairs available. Both the outcoupling optical elements of the waveguides and the focusing sides of the lenses may be static (i.e., not dynamic or electroactive). In some alternative embodiments, one or both may be dynamic using electroactive features.

[0081] In some embodiments, two or more of the waveguides 270, 280, 290, 300, 310 may have the same associated depth plane. For example, multiple waveguides 270, 280, 290, 300, 310 may be configured to output images set at the same depth plane, or multiple subsets of the waveguides 270, 280, 290, 300, 310 may be configured to output images set at the same depth planes, with one set per depth plane. This may provide the advantage of forming tiled images to provide an extended field of view at those depth planes.

[0082] Continuing with reference to FIG. 6, the outcoupling optical elements 570, 580, 590, 600, 610 may be configured to redirect light from its respective waveguide and output the light with an appropriate amount of divergence or collimation for a particular depth plane associated with the waveguide. As a result, waveguides with different associated depth planes may have different configurations of outcoupling optical elements 570, 580, 590, 600, 610 that output light with different amounts of divergence depending on the associated depth plane. In some embodiments, the light extraction optical elements 570, 580, 590, 600, 610 may be volume or surface features that may be configured to output light at specific angles. For example, the light extraction optical elements 570, 580, 590, 600, 610 may be volume holograms, surface holograms, and / or diffraction gratings. In some embodiments, features 320, 330, 340, 350 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers and / or structures for forming air gaps).

[0083] In some embodiments, the outcoupling optical elements 570, 580, 590, 600, 610 are diffractive features or "diffractive optical elements" (also referred to herein as "DOEs") that form a diffraction pattern. Preferably, the DOEs have a sufficiently low diffraction efficiency such that only a portion of the light in the beam is deflected toward the eye 210 at each intersection of the DOE, while the remainder continues to travel through the waveguide via TIR. The light carrying the image information is thus split into several related exit beams that exit the waveguide at various locations, resulting in a very uniform pattern of exit emission toward the eye 210 for this particular collimated beam bouncing within the waveguide.

[0084] In some embodiments, one or more DOEs may be switchable between an "on" state in which they actively diffract and an "off" state in which they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer dispersed liquid crystal in which microdroplets comprise a diffractive pattern in a host medium, and the refractive index of the microdroplets may be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdroplets may be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts incident light).

[0085] In some embodiments, a camera assembly 630 (e.g., a digital camera, including visible and infrared light cameras) may be provided to capture images of the eye 210 and / or tissue surrounding the eye 210, for example, to detect user input and / or monitor a physiological condition of the user. As used herein, a camera may be any image capture device. In some embodiments, the camera assembly 630 may include an image capture device and a light source that projects light (e.g., infrared light) onto the eye, which may then be reflected by the eye and detected by the image capture device. In some embodiments, the camera assembly 630 may be mounted to the frame 80 (FIG. 9D) and may be in electrical communication with the processing modules 140 and / or 150, which may process image information from the camera assembly 630. In some embodiments, one camera assembly 630 may be utilized per eye, monitoring each eye separately.

[0086] 7, an example of an output beam output by a waveguide is shown. Although one waveguide is shown, it should be understood that other waveguides in the waveguide assembly 260 (FIG. 6) may function similarly and that the waveguide assembly 260 includes multiple waveguides. Light 640 is launched into the waveguide 270 at the input surface 460 of the waveguide 270 and propagates within the waveguide 270 by TIR. At the point where the light 640 impinges on the DOE 570, a portion of the light exits the waveguide as an output beam 650. The output beam 650 is shown as approximately parallel, but may be redirected to propagate to the eye 210 at an angle (e.g., divergent output beam formation) as discussed herein and depending on the depth plane associated with the waveguide 270. It should be understood that a nearly collimated exit beam may refer to a waveguide with outcoupling optics that outcouples light to form an image that appears to be set at a depth plane at a large distance (e.g., optical infinity) from the eye 210. Other waveguides or other sets of outcoupling optics may output an exit beam pattern that is more divergent, which would require the eye 210 to accommodate to a closer distance and focus on the retina, and would be interpreted by the brain as light from a distance closer to the eye 210 than optical infinity.

[0087] In some embodiments, a full color image may be formed at each depth plane by overlaying an image on each of the primary colors, for example, three or more primary colors. FIG. 8 illustrates an example of a stacked waveguide assembly, where each depth plane includes an image formed using multiple different primary colors. The illustrated embodiment shows depth planes 240a-240f, but more or less depths are also contemplated. Each depth plane may have three or more primary color images associated with it, including a first image of a first color G, a second image of a second color R, and a third image of a third color B. The different depth planes are indicated in the diagram by different numbers for diopters (dpt) following the letters G, R, and B. Merely by way of example, the numbers following each of these letters indicate the diopters (1 / m), i.e., the inverse distance of the depth plane from the viewer, and each box in the diagram represents an individual primary color image. In some embodiments, the exact locations of the depth planes for the different primary colors may be varied to account for differences in the eye's focusing of light of different wavelengths. For example, different primary color images for a given depth plane may be placed on depth planes that correspond to different distances from the user. Such an arrangement may increase visual acuity and user comfort, and / or reduce chromatic aberration.

[0088] In some embodiments, the light of each primary color may be output by a single dedicated waveguide, such that each depth plane may have multiple waveguides associated with it. In such embodiments, each box in the figure containing the letter G, R, or B may be understood to represent an individual waveguide, and three waveguides may be provided per depth plane, with three primary color images provided per depth plane. The waveguides associated with each depth plane are shown adjacent to each other in this drawing for ease of illustration, but it should be understood that in a physical device, the waveguides may all be arranged in a stack with one waveguide per level. In some other embodiments, multiple primary colors may be output by the same waveguide, such that, for example, only a single waveguide may be provided per depth plane.

[0089] 8, in some embodiments, G is green, R is red, and B is blue. In some other embodiments, other colors associated with other wavelengths of light, including magenta and cyan, may be used in addition to or in place of one or more of the red, green, or blue colors.

[0090] It should be understood that references throughout this disclosure to a given color of light will be understood to encompass light of one or more wavelengths within the range of wavelengths of light that are perceived by a viewer to be of that given color. For example, red light may include one or more wavelengths of light within a range of about 620-780 nm, green light may include one or more wavelengths of light within a range of about 492-577 nm, and blue light may include one or more wavelengths of light within a range of about 435-493 nm.

[0091] In some embodiments, the light source 530 (FIG. 6) may be configured to emit light at one or more wavelengths outside the range of visual perception of a viewer, e.g., infrared and / or ultraviolet wavelengths. Additionally, the waveguide in-coupling, out-coupling, and other light redirecting structures of the display 250 may be configured to direct and emit this light from the display toward the user's eye 210, e.g., for imaging and / or user stimulation applications.

[0092] 9A, in some embodiments, light impinging on a waveguide may need to be redirected to in-couple the light into the waveguide. An in-coupling optical element may be used to redirect and in-couple the light into its corresponding waveguide. FIG. 9A illustrates a cross-sectional side view of an example of a plurality or set 660 of stacked waveguides, each including an in-coupling optical element. Each of the waveguides may be configured to output light of one or more different wavelengths or one or more different wavelength ranges. The stack 660 may correspond to the stack 260 (FIG. 6), and the illustrated waveguides of the stack 660 may correspond to a portion of the plurality of waveguides 270, 280, 290, 300, 310, but it should be understood that light from one or more of the image input devices 360, 370, 380, 390, 400 is launched into the waveguide from a position that requires the light to be redirected for in-coupling.

[0093] The illustrated set 660 of stacked waveguides includes waveguides 670, 680, and 690. Each waveguide includes an associated incoupling optical element (which may also be referred to as the light input area on the waveguide), for example, incoupling optical element 700 is disposed on a major surface (e.g., the top major surface) of waveguide 670, incoupling optical element 710 is disposed on a major surface (e.g., the top major surface) of waveguide 680, and incoupling optical element 720 is disposed on a major surface (e.g., the top major surface) of waveguide 690. In some embodiments, one or more of the incoupling optical elements 700, 710, 720 may be disposed on a bottom major surface of an individual waveguide 670, 680, 690 (particularly, one or more of the incoupling optical elements are reflective polarizing optical elements). As shown, the incoupling optical elements 700, 710, 720 may be disposed on the upper major surface of the respective waveguide 670, 680, 690 (or on top of the next lower waveguide), and in particular, the incoupling optical elements are transmissive deflecting optical elements. In some embodiments, the incoupling optical elements 700, 710, 720 may be disposed within the body of the respective waveguide 670, 680, 690. In some embodiments, as discussed herein, the incoupling optical elements 700, 710, 720 are wavelength selective such that they selectively redirect one or more wavelengths of light while transmitting other wavelengths of light. Although illustrated on one side or corner of the respective waveguide 670, 680, 690, it should be understood that the incoupling optical elements 700, 710, 720 may be disposed within other areas of the respective waveguide 670, 680, 690 in some embodiments.

[0094] As shown, the in-coupling optical elements 700, 710, 720 may be laterally offset from one another. In some embodiments, each in-coupling optical element may be offset to receive light without that light passing through another in-coupling optical element. For example, each in-coupling optical element 700, 710, 720 may be configured to receive light from different image input devices 360, 370, 380, 390, and 400 as shown in FIG. 6 and may be separated (e.g., laterally spaced) from the other in-coupling optical elements 700, 710, 720 so as to not substantially receive light from others of the in-coupling optical elements 700, 710, 720.

[0095] Each waveguide also includes an associated optically dispersive element, for example, optically dispersive element 730 is disposed on a major surface (e.g., a top major surface) of waveguide 670, optically dispersive element 740 is disposed on a major surface (e.g., a top major surface) of waveguide 680, and optically dispersive element 750 is disposed on a major surface (e.g., a top major surface) of waveguide 690. In some other embodiments, optically dispersive elements 730, 740, 750 may be disposed on the bottom major surfaces of associated waveguides 670, 680, 690, respectively. In some other embodiments, optically dispersive elements 730, 740, 750 may be disposed on both the top and bottom major surfaces of associated waveguides 670, 680, 690, respectively, or optically dispersive elements 730, 740, 750 may be disposed on different ones of the top and bottom major surfaces in different associated waveguides 670, 680, 690, respectively.

[0096] The waveguides 670, 680, 690 may be spaced apart and separated, for example, by gas, liquid and / or solid layers of material. For example, as shown, layer 760a may separate the waveguides 670 and 680, and layer 760b may separate the waveguides 680 and 690. In some embodiments, layers 760a and 760b are formed from a low index material (i.e., a material having a lower index of refraction than the material forming the immediate neighbors of the waveguides 670, 680, 690). Preferably, the index of refraction of the material forming layers 760a, 760b is 0.05 or greater, or 0.10 or less, relative to the index of refraction of the material forming the waveguides 670, 680, 690. Advantageously, the lower refractive index layers 760a, 760b may act as cladding layers to promote total internal reflection (TIR) ​​of light through the waveguides 670, 680, 690 (e.g., TIR between the top and bottom major surfaces of each waveguide). In some embodiments, the layers 760a, 760b are formed from air. Although not shown, it should be understood that the top and bottom of the illustrated set of waveguides 660 may include immediate cladding layers.

[0097] Preferably, for ease of manufacturing and other considerations, the materials forming the waveguides 670, 680, 690 are similar or the same, and the materials forming the layers 760a, 760b are similar or the same. In some embodiments, the materials forming the waveguides 670, 680, 690 may differ between one or more of the waveguides, and / or the materials forming the layers 760a, 760b may differ while still maintaining the various refractive index relationships discussed above.

[0098] 9A, light rays 770, 780, 790 are incident on the set of waveguides 660. It should be understood that light rays 770, 780, 790 may be injected into the waveguides 670, 680, 690 by one or more image injection devices 360, 370, 380, 390, 400 (FIG. 6).

[0099] In some embodiments, the light beams 770, 780, 790 have different properties, e.g., different wavelengths or different wavelength ranges, that may correspond to different colors. Each of the incoupling optical elements 700, 710, 720 deflects the incident light such that the light propagates through a respective one of the waveguides 670, 680, 690 by TIR. In some embodiments, each of the incoupling optical elements 700, 710, 720 selectively deflects one or more particular wavelengths of light while transmitting other wavelengths to the underlying waveguide and associated incoupling optical element.

[0100] For example, incoupling optical element 700 may be configured to deflect light beam 770 having a first wavelength or wavelength range while transmitting light beams 780 and 790 having different second and third wavelengths or wavelength ranges, respectively. The transmitted light beam 780 impinges on and is deflected by incoupling optical element 710, which is configured to deflect light of the second wavelength or wavelength range. Light beam 790 is deflected by incoupling optical element 720, which is configured to selectively deflect light of the third wavelength or wavelength range.

[0101] 9A, the deflected light rays 770, 780, 790 are deflected to propagate through the corresponding waveguides 670, 680, 690. That is, the incoupling optical element 700, 710, 720 of each waveguide deflects the light into its corresponding waveguide 670, 680, 690 and incouplings the light into the corresponding waveguide. The light rays 770, 780, 790 are deflected at an angle that causes the light to propagate through the respective waveguides 670, 680, 690 by TIR. The light rays 770, 780, 790 propagate through the respective waveguides 670, 680, 690 by TIR until they impinge on the corresponding optical dispersive element 730, 740, 750 of the waveguide.

[0102] 9B, a perspective view of the multiple stacked waveguide embodiment of FIG. 9A is illustrated. As previously described, the in-coupled light rays 770, 780, 790 are deflected by the in-coupling optical elements 700, 710, 720, respectively, and then propagate by TIR within the waveguides 670, 680, 690, respectively. The light rays 770, 780, 790 then impinge on the optically dispersive elements 730, 740, 750, respectively. The optically dispersive elements 730, 740, 750 deflect the light rays 770, 780, 790 to propagate towards the out-coupling optical elements 800, 810, 820, respectively.

[0103] In some embodiments, the optically dispersive elements 730, 740, 750 are orthogonal pupil expanders (OPEs). In some embodiments, the OPEs deflect or disperse the light to the out-coupling optical elements 800, 810, 820, and in some embodiments may also increase the beam or spot size of this light as it propagates to the out-coupling optical elements. In some embodiments, the optically dispersive elements 730, 740, 750 may be omitted and the in-coupling optical elements 700, 710, 720 may be configured to deflect the light directly to the out-coupling optical elements 800, 810, 820. For example, referring to FIG. 9A, the optically dispersive elements 730, 740, 750 may be replaced with the out-coupling optical elements 800, 810, 820, respectively. In some embodiments, the outcoupling optical elements 800, 810, 820 are exit pupils (EPs) or exit pupil expanders (EPEs) that direct light to the viewer's eye 210 (FIG. 7). It should be understood that the OPEs may be configured to increase the size of the eyebox in at least one axis, and the EPEs may increase the eyebox in an axis that intersects, e.g., orthogonal to, the axis of the OPE. For example, each OPE may be configured to redirect a portion of the light striking the OPE to an EPE of the same waveguide, while allowing the remaining portion of the light to continue propagating down the waveguide. In response to striking the OPE, again, another portion of the remaining light is redirected to the EPE, which continues to propagate further down the waveguide, etc. Similarly, in response to striking the EPE, a portion of the striking light is directed out of the waveguide towards the user, and the remaining portion of the light continues to propagate through the waveguide until it strikes the EP again, at which point another portion of the striking light is directed out of the waveguide, etc. As a result, a single beam of internally coupled light may be "replicated" each time a portion of that light is redirected by an OPE or EPE, thereby forming a cloned beam field of light, as shown in Figure 6. In some embodiments, the OPE and / or EPE may be configured to modify the size of the beam of light.

[0104] Thus, referring to Figures 9A and 9B, in some embodiments, a set of waveguides 660 includes, for each primary color, a waveguide 670, 680, 690, an in-coupling optical element 700, 710, 720, an optical dispersion element (e.g., OPE) 730, 740, 750, and an out-coupling optical element (e.g., EP) 800, 810, 820. The waveguides 670, 680, 690 may be stacked with an air gap / cladding layer between each one. The in-coupling optical element 700, 710, 720 redirects or deflects the incoming light into its waveguide (with different in-coupling optical elements receiving different wavelengths of light). The light then propagates at an angle that will result in TIR within the individual waveguides 670, 680, 690. In the illustrated embodiment, light ray 770 (e.g., blue light) is deflected by the first in-coupling optical element 700 in the manner described above, then continues to bounce back down the waveguide, interacting with the optically dispersive element (e.g., OPE) 730 and then the out-coupling optical element (e.g., EP) 800. Light rays 780 and 790 (e.g., green and red light, respectively) pass through the waveguide 670, where light ray 780 impinges on and is deflected by the in-coupling optical element 710. Light ray 780 will then proceed via TIR back down the waveguide 680, to its optically dispersive element (e.g., OPE) 740 and then the out-coupling optical element (e.g., EP) 810. Finally, light ray 790 (e.g., red light) passes through the waveguide 690 and impinges on the optically in-coupling optical element 720 of the waveguide 690. The light in-coupling optical element 720 deflects the light beam 790 so that it propagates by TIR through a light dispersive element (e.g., OPE) 750 and then by TIR to an out-coupling optical element (e.g., EP) 820. The out-coupling optical element 820 then finally out-couples the light beam 790 to a viewer, who also receives the out-coupled light from the other waveguides 670, 680.

[0105] FIG. 9C illustrates a top-down plan view of an example of the multiple stacked waveguides of FIGS. 9A and 9B. As shown, the waveguides 670, 680, 690 may be vertically aligned with each waveguide's associated light dispersive element 730, 740, 750 and associated out-coupling optical elements 800, 810, 820. However, as discussed herein, the in-coupling optical elements 700, 710, 720 are not vertically aligned. Rather, the in-coupling optical elements are preferably non-overlapping (e.g., laterally spaced apart, as seen in the top-down view). As discussed further herein, this non-overlapping spatial arrangement facilitates the injection of light from different sources into different waveguides on a one-to-one basis, thereby allowing a specific light source to be uniquely coupled to a specific waveguide. In some embodiments, an arrangement including non-overlapping spatially separated in-coupling optical elements may be referred to as a shifted pupil system, and the in-coupling optical elements in these arrangements may correspond to sub-pupils.

[0106] 9D illustrates an example of a wearable display system 60 into which the various waveguide and associated systems disclosed herein may be integrated. In some embodiments, the display system 60 is the system 250 of FIG. 6, which diagrammatically illustrates some portions of the system 60 in greater detail. For example, the waveguide assembly 260 of FIG. 6 may be part of the display 70.

[0107] 9D , the display system 60 includes a display 70 and various mechanical and electronic modules and systems to support the functionality of the display 70. The display 70 may be coupled to a frame 80, which is wearable by a display system user or viewer 90 and configured to position the display 70 in front of the eye of the user 90. The display 70 may be considered an eyepiece in some embodiments. In some embodiments, a speaker 100 is coupled to the frame 80 and configured to be positioned adjacent to the ear canal of the user 90 (in some embodiments, another speaker, not shown, may also be optionally positioned adjacent the user's other ear canal to provide stereo / shapeable sound control). The display system 60 may also include one or more microphones 110 or other devices to detect sound. In some embodiments, the microphones may be configured to allow the user to provide input or commands to the system 60 (e.g., voice menu command selections, natural language queries, etc.) and / or enable audio communication with other persons (e.g., other users of a similar display system). The microphone may further be configured as an ambient sensor and collect audio data (e.g., sounds from the user and / or the environment). In some embodiments, the display system 60 may further include one or more outwardly directed environmental sensors 112 configured to detect objects, stimuli, people, animals, places, or other aspects of the world around the user. For example, the environmental sensors 112 may include one or more cameras, which may be positioned, for example, facing outward, to capture images similar to at least a portion of the user 90's normal field of view. In some embodiments, the display system may also include an ambient sensor 120a, which may be separate from the frame 80 and mounted on the body of the user 90 (e.g., the head, torso, limbs, etc. of the user 90). The ambient sensor 120a may, in some embodiments, be configured to obtain data characterizing a physiological state of the user 90. For example, the sensor 120a may be an electrode.

[0108] 9D, the display 70 is operably coupled to the local data processing module 140 by a communication link 130, such as wired or wireless connectivity, which may be mounted in a variety of configurations, such as fixedly attached to the frame 80, fixedly attached to a helmet or hat worn by the user, embedded within headphones, or otherwise removably attached to the user 90 (e.g., in a backpack-type configuration, a belt-type configuration). Similarly, the sensor 120a may be operably coupled to the local processor and data module 140 by a communication link 120b, such as wired or wireless connectivity. The local processing and data module 140 may comprise a hardware processor and digital memory, such as non-volatile memory (e.g., flash memory or hard disk drive), both of which may be utilized to aid in processing, caching, and storing data. Optionally, the local processing and data module 140 may include one or more central processing units (CPUs), graphics processing units (GPUs), dedicated processing hardware, and the like. The data may include a) data captured from sensors (e.g., image capture devices (such as cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, gyroscopes, and / or other sensors disclosed herein (e.g., which may be operatively coupled to the frame 80 or otherwise attached to the user 90)) and / or b) data acquired and / or processed using the remote processing module 150 and / or remote data repository 160 (including data related to virtual content), possibly for passing to the display 70 after processing or retrieval. The local processing and data module 140 may be operatively coupled to the remote processing module 150 and the remote data repository 160 by communication links 170, 180, such as via wired or wireless communication links, such that these remote modules 150, 160 are operatively coupled to each other and available as resources to the local processing and data module 140.In some embodiments, local processing and data module 140 may include one or more of an image capture device, a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope. In some other embodiments, one or more of these sensors may be mounted on frame 80 or may be a stand-alone structure that communicates with local processing and data module 140 by a wired or wireless communication path.

[0109] 9D , in some embodiments, remote processing module 150 may comprise one or more processors configured to analyze and process data and / or image information, and may include, for example, one or more central processing units (CPUs), graphics processing units (GPUs), special purpose processing hardware, etc. In some embodiments, remote data repository 160 may comprise digital data storage facilities that may be available through the Internet or other networking configurations in a "cloud" resource configuration. In some embodiments, remote data repository 160 may include one or more remote servers, which provide information to local processing and data module 140 and / or remote processing module 150, for example, information for generating augmented reality content. In some embodiments, all data is stored and all computations are performed within the local processing and data module, allowing for fully autonomous use from the remote modules. Optionally, an external system (e.g., one or more processors, one or more computer systems), including a CPU, GPU, etc., may perform at least a portion of the processing (e.g., generating image information, processing data) and provide information to and receive information from modules 140, 150, 160, e.g., via a wireless or wired connection.

[0110] (Posture determination) As described herein, a display device (e.g., display system 60, illustrated in FIG. 9D) can present virtual content to a user (e.g., a user wearing a display device, such as wearing a display 70 coupled to a frame 80). During presentation of the virtual content, the display device can determine a pose of the display device and the user's head. The pose can identify the orientation of the display device and / or the user's head, and optionally the position of the display device and / or the user's head, as described above. For example, the display device can present virtual content comprising a virtual document on a real-world desk. As the user rotates their head about the document or moves closer or further away from the document, the display device can determine their head pose. In this manner, the display device can adjust the presented virtual content such that the virtual document appears on the real-world desk as a real document. Although this description refers to virtual content such as augmented reality, the display device may be a virtual reality display system and utilize the techniques described herein.

[0111] The display device can utilize an imaging device, such as the environmental sensor 112 described in FIG. 9D above, to determine the pose. The imaging device may be, for example, an outward-facing camera fixed on the display device. The imaging device can thus capture images of the real-world environment, and the display device can use these images to determine the pose. The imaging device may capture images in response to the passage of time (e.g., every 1 / 10, 1 / 15, 1 / 30, 1 / 60 of a second) or in response to detecting that the display device has moved an amount greater than a threshold. For example, the imaging device can capture live images of the real-world environment (e.g., the camera's sensor may be configured to capture image information at all times). In this example, the display device can determine that the incoming image information has changed an amount greater than a threshold or is changing at a rate greater than a threshold. As another example, the display device may include sensors (e.g., magnetometers, gyroscopes, etc.), such as in an inertial measurement unit, and the display device can use these sensors to identify whether the display device has moved an amount greater than a threshold. In this embodiment, the display device can acquire images based on information detected by the inertial measurement unit. A current image may thus be acquired via the imaging device and may differ from a previous image based on the user's movement. For example, the display device can acquire successive images as the user looks around a room.

[0112] As described herein, the display device can track salient points between successive images acquired by the imaging device. In some embodiments, the display device can be configured to perform a patch-based frame-to-frame tracking process. The salient points can represent distinguishable visual points, such as corners, as described above. To track the salient points from a previous image to a current image, the display device can project a patch on the current image that surrounds the salient points in the previous image. As described herein, the patch can be an M×N image area that surrounds the salient points. For example, the salient points can correspond to two-dimensional locations in the current image, and the patch can be an M×N image area that surrounds the two-dimensional locations. The display device can then adjust the location of the projected patch to minimize the error or aggregate the difference in pixel intensity between the projected patch and the corresponding image area in the current image. Exemplary error minimization processes can include Levenberg-Marquardt, conjugate gradient, and the like. A consistent, selected location within the patch, e.g., the centroid of the projected patch, can be understood to be the location of the tracked salient point in the current image. In this way, the display device can identify the movement of specific visual points of interest (eg, salient points such as corners) from the previous frame to the current frame.

[0113] The display device may also be configured to utilize a descriptor-based map / frame tracking. As described herein, map / frame tracking utilizes map information indicating real-world locations (e.g., three-dimensional locations) of salient points and associated descriptors. For example, the map information may indicate three-dimensional coordinates for a particular corner in a real-world environment. When a particular corner is imaged by the display device and thus represented in the current image, the display device may match the representation in the current image to its corresponding real-world location. The step of generating the map information will be described in more detail below with respect to FIG. 10A. In map / frame tracking, new salient points may be extracted from image areas of the current image with less than a threshold measurement (e.g., less than a threshold number or density) of tracked salient points. A descriptor may be generated for each salient point, and the descriptor may be matched to the descriptors of the salient points shown in the map information. As used herein, a descriptor may describe one or more visual elements associated with a salient point. For example, the descriptor may indicate the shape, color, texture, etc. of the salient point. Additionally, the descriptor can describe the area surrounding each salient point (e.g., an M×N pixel area surrounding the salient point), the descriptor can represent, for example, a histogram of the area surrounding the salient point (e.g., an alphanumeric value associated with the histogram), a hash of the area (e.g., a cryptographic hash calculated from the values ​​of each pixel in the M×N pixel area), etc.

[0114] Based on the generated descriptors, the display device can then match the generated descriptors for each salient point with the descriptors of the salient points shown in the map information. In this way, the display device can identify real-world locations (e.g., 3D coordinates) that correspond to each salient point in the current image. Thus, the salient points in the current image can represent the projection of the corresponding 3D real-world coordinates onto the 2D image.

[0115] The display device can determine its pose according to these matches. For example, the display device can implement an exemplary pose estimation process such as n-point perspective (pnp), efficient pnp, pnp with random sample sharing terms, etc. The display device can then track salient points in subsequent images. For example, the display device can project salient points in a current image onto a subsequent image as described above.

[0116] 10A illustrates a flowchart of an example process 1000 for determining a pose of a display system and a pose of a user's head. In some embodiments, the process 1000 may be described as being performed by a display device (e.g., an augmented reality display system 60, which may include processing hardware and software and may optionally provide information to one or more computational external systems for processing, e.g., offload processing loads to the external systems, and receive information from the external systems). In some embodiments, the display device may be a virtual reality display device, which includes one or more processors.

[0117] A display device acquires a current image of a real-world environment (block 1002). The display device may acquire the current image from an imaging device, such as an outward-facing camera fixed on the display device. For example, an outward-facing camera may be positioned in front of the display device and acquire a view similar to that seen by a user (e.g., a forward-facing view). As described above with respect to FIG. 9D, the display device may be worn by a user. The display device may optionally utilize two or more imaging devices to acquire images from each simultaneously (e.g., substantially simultaneously). The imaging devices may thus be configured to acquire stereoscopic images of the real-world environment, which can be utilized to determine the depth of locations within the images.

[0118] The display device may trigger or otherwise cause the imaging device to acquire a current image based on a threshold amount of time having elapsed since a previously acquired image. The imaging device may thus acquire images at a particular frequency, such as 10 times per second, 15 times per second, 30 times per second, etc. Optionally, the particular frequency may be adjusted based on the processing workload of the display device. For example, the particular frequency may be adaptively reduced if the processor of the display device is utilized above one or more threshold percentages. Additionally or alternatively, the display device may adjust the frequency based on movement of the display device. For example, the display device may acquire information indicative of a threshold number of previously determined poses and determine an amount of variance between the poses. Based on the amount of variance, the display device may increase the frequency with which the display device acquires images, for example, until a measure of central tendency falls below a particular threshold. In some embodiments, the display device may utilize sensors, such as those included in an inertial measurement unit, to increase or decrease the frequency according to the estimated movement of the user. In some embodiments, in addition to acquiring a current image based on the passage of a threshold amount of time, the display device may acquire a current image based on an estimation that the user has moved more than a threshold amount (e.g., a threshold distance about one or more three-dimensional axes). For example, the display device may utilize an inertial measurement unit to estimate the user's movement. In some embodiments, the display device may utilize one or more other sensors, such as sensors that detect light, chromatic dispersion, etc., to determine that information detected by the sensor has changed (e.g., indicates movement) more than a threshold amount within a threshold amount of time.

[0119] The current image can thus be associated with the user's current view. The display device can store the current image, for example in a volatile or non-volatile memory, for processing. In addition, the display device can store an image acquired before the current image. As will be described, the current image can be compared to the previous image, and salient points can be tracked from the previous image to the current image. Thus, the display device can store information associated with each salient point in the previous image. For example, the information can include a patch for each salient point, and optionally a location in the previous image in which the patch appeared (e.g., pixel coordinates of the salient point). In some embodiments, instead of storing the complete previous image, the display device can store a patch for each salient point included in the previous image.

[0120] As explained above, a patch can represent an M×N sized image area surrounding a salient point (e.g., a salient point as imaged). For example, the salient point can be the centroid of the patch. Since a salient point can be a point of visual interest such as a corner, the corner can be larger than a single pixel in some embodiments. The patch can thus surround, for example, a location of visual interest where two lines intersect (e.g., on a keyboard "7", the patch can surround the intersection of a horizontal line and a tilted vertical line). For example, the display device can select a particular pixel as being a salient point, and the patch can surround this particular pixel. In addition, two or more pixels may be selected, and the patch can surround these two or more pixels. As will be explained below, a patch of a previous image can be utilized to track associated salient points in a current image.

[0121] The display device projects the tracked salient points from the previous image onto the current image (block 1004). As described above, the display device may store information associated with the salient points contained in the previous image. Exemplary information may include a patch surrounding the salient point along with information identifying the location of the patch in the previous image. The display device may project each salient point from the previous image onto the current image. As an example, a pose associated with the previous image may be utilized to project each salient point onto the current image. As will be described below, a pose estimate, such as optical flow, may be determined by the display device. This pose estimate may adjust the pose determined for the previous image, and thus an initial projection of the tracked salient points onto the current image may be obtained. As will be described, this initial projection may be refined.

[0122] The display device can determine a pose estimate, sometimes also referred to as an initial value, based on trajectory prediction (e.g., based on a previously determined pose) and / or based on an inertial measurement unit, an extended Kalman filter, visual inertial odometry, etc. With respect to trajectory prediction, the display device can determine a direction in which the user is likely moving. For example, if a threshold number of previous pose determinations indicate that the user is rotating their head downward in a particular manner, the trajectory prediction can extend this rotation. With respect to an inertial measurement unit, the display device can obtain information indicating an adjustment to the orientation and / or position as measured by a sensor of the inertial measurement unit. The pose estimate can thus enable the determination of an initial estimated location corresponding to each tracked salient point in the current image. In addition to the pose estimate, the display device can utilize the real-world location of each salient point as indicated in the map information to project the salient points. For example, the pose estimate can inform the estimated movement of each salient point from a 2D location in a previous image to a 2D location in the current image. This new 2D location can be compared to the map information and the estimated locations of the salient points can be determined.

[0123] A patch for each salient point in the previous image can be compared to an identically sized M×N pixel area of ​​the current image. For example, the display device can adjust the location of the patch projected onto the current image until the photometric error between the patch and the identically sized M×N pixel area of ​​the current image onto which the patch is projected is minimized (e.g., substantially minimized, such as an error below a local or global minimum or a user-selectable threshold). In some embodiments, the centroid of the M×N pixel area of ​​the current image can be designated to correspond to the tracked salient point. The steps of projecting the tracked salient points are described in more detail below with respect to FIGS. 11-12B.

[0124] Optionally, to determine the pose estimate, the display device can minimize a combined photometric cost function of all projected patches by varying the pose of the current image. For example, the display device can project a patch associated with each salient point in the previous image onto the current image (e.g., based on an initial pose estimate, as described above). The display device can then globally adjust the patches, e.g., via modifying this initial pose estimate, until the photometric cost function is minimized. In this way, a more accurate refined pose estimate can be obtained. As will be described below, this refined pose estimate can be used as an initial value, i.e., a regularization value, when determining the pose of the display device. For example, the refined pose estimate can be associated with a cost function such that deviations from the refined pose estimate have an associated cost.

[0125] Thus, the current image may include tracked salient points from the previous image. As will be explained below, the display device may identify image areas of the current image with tracked salient points below a threshold measurement. This may represent, for example, that the user is moving their head to a new location in the real-world environment. In this manner, the new image area of ​​the current image that images the new location may not include tracked salient points from the previous image.

[0126] The display device determines whether an image area of ​​the current image contains less than a threshold measurement of tracked salient points (block 1006). As described above, the display device can determine its pose according to patch-based frame / frame tracking, e.g., via projection of tracked salient points onto successive images, optionally in combination with map / frame tracking. Map / frame tracking can be utilized if one or more image areas of the current image contain less than a threshold measurement of tracked salient points within the image area, e.g., a threshold number of salient points or a threshold density of salient points.

[0127] 10B illustrates an example image area of ​​an example current image. In the example of FIG. 10B, current images 1020A, 1020B, 1020C, and 1020D are illustrated. These current images may be acquired via a display device, for example, as described above with respect to block 1002. Current image 1020A is illustrated with example image area 1022. As described above, an image area may encompass the entirety of the current image. Thus, the display device can determine whether current image 1020A as a whole includes less than a threshold measurement of tracked salient points.

[0128] In some other embodiments, the current image may be subdivided into distinct portions. For example, the current image 1020B in the example of FIG. 10B is separated into a 5×5 grid, with each area of ​​the grid forming a distinct portion. The exemplary image area 1024 is thus one of these portions. The display device may determine whether any of these portions include tracked salient points below the threshold measurement. In this manner, as new locations of the real-world environment are included within the current image 1020B, one or more of the portions may include tracked salient points below the threshold measurement. The size of the grid may be adjustable by the user and / or by the display system. For example, the grid may be selected to be 3×3, 7×7, 2×4, etc. Optionally, the grid may be adjusted during operation of the display system, e.g., the sizes of various portions of the image vary as the user utilizes the display system (e.g., substantially in real time, according to processing constraints, accuracy thresholds, differences in pose estimates between images, etc.).

[0129] A current image 1020C is illustrated with exemplary tracked salient points. In this example, image areas may be determined according to the sparseness of the tracked salient points. For example, image areas 1026A and 1026B are illustrated surrounding a single tracked salient point. The size of the image areas may be user selectable or a fixed system determined size (e.g., M×N pixel areas). The display device may analyze the tracked salient points and determine whether image areas with measurements below the threshold may be located within the current image 1020C. For example, image areas 1026A and 1026B have been identified by the display device as including measurements below the threshold. Optionally, the display device may identify image areas including measurements above the threshold and identify the remaining image areas as including measurements below the threshold. For example, image areas 1028 and 1030 have been identified as including tracked salient points above the threshold measurement. Thus, in this example, the display device may identify anywhere outside of images 1028 and 1030 as having tracked salient points below the threshold measurement. The display device can then extract new salient points within these outer image areas. Optionally, the display device can determine a clustering measure for locations within the current image. For example, the clustering measure can indicate the average distance a location originates from tracked salient points. In addition, the clustering measure can indicate the average number of tracked salient points within a threshold distance of the location. If the clustering measure falls below one or more thresholds, the display device can extract new salient points at these locations. Optionally, the display device can extract new salient points within an M×N area surrounding each location.

[0130] The current image 1020D is illustrated with an exemplary image area 1032. In this example, the image area 1032 can be located within a particular location of the current image 1020D, such as the center of the current image 1020D. In some embodiments, the exemplary image area 1032 can represent a particular field of view of the user. The image area 1032 can be a particular shape or polygon, such as a circle, an oval, a rectangle, etc. In some embodiments, the image area 1032 can be based on an accuracy associated with a lens of an imaging device. For example, the image area 1032 can represent the center of the lens with substantially no distortion introduced at the edge of the lens. Thus, the display device can identify whether the image area 1032 includes less than a threshold measurement of tracked salient points.

[0131] Referring again to FIG. 10A, the display device performs map / frame tracking (block 1008). The display device can identify whether any image area of ​​the current image contains tracked salient points below a threshold measurement. As described above, the display device can extract new salient points within the identified image area. For example, the display device can identify 2D locations of the new salient points within the image area. The display device can then receive descriptors for the newly extracted salient points and the existing tracked salient points. For example, the display device can provide descriptors for the newly extracted tracked salient points, for example, by generating these descriptors. Optionally, the display device can provide descriptors by generating descriptors for the newly extracted salient points and by receiving (e.g., retrieving from memory) descriptors for the tracked salient points. As an example, the display device can utilize previously generated descriptors for the tracked salient points. For example, descriptors may have been generated when each tracked salient point was newly extracted from the image. These descriptors can be matched by the display device with descriptors stored in the map information. The map information stores the real-world coordinates of salient points so that the display device can identify the real-world coordinates of each salient point tracked in the current image. The steps for performing map / frame tracking will be described in more detail below with respect to FIG.

[0132] The map information can be generated by a display device as utilized herein. For example, the display device can utilize a stereoscopic imaging device, a depth sensor, a lidar, etc. to determine depth information associated with locations in the real-world environment. The display device can update the map information periodically, for example, every threshold number of seconds or minutes. In addition, the map information can be updated, for example, based on identifying that a current image as obtained from a stereoscopic imaging device is a key frame. This can be identified according to time, as described above, and optionally according to the difference between the current image and a previous (e.g., most recent) key frame. For example, if the current image has changed by more than a threshold, the current image can be identified as a key frame. These key frames can then be analyzed to update the map information.

[0133] For the stereoscopic imaging devices, the display device can generate a descriptor for salient points in each stereoscopic image. Using known collateral calibration information, e.g., the relative pose between the two imaging devices, depth information can be identified. Based on the descriptors matching salient points between the stereoscopic images and the depth information, real-world coordinates (e.g., relative to a coordinate reference frame) can be determined for each salient point. One or more of the generated descriptors for each matched salient point can then be stored. Thus, during map / frame tracking, these stored descriptors for real-world salient points can be matched to the descriptors of salient points contained in a captured image (e.g., the current image). As an example, one of the stereoscopic imaging devices may acquire a current image (e.g., as described in block 1002). The display device can access the map information and, in some embodiments, match the descriptors generated for this same imaging device to the descriptors of salient points contained in the current image. Optionally, patch-based photometric error minimization may be utilized to match salient points between stereoscopic images and thus determine real-world coordinates to be stored in the map information. The display device may then generate individual descriptors for the salient points (e.g., from one or more of the stereoscopic images), which may be utilized to perform map / frame tracking. Further description of generating map information is included at least in FIG. 16 and associated description in U.S. Patent Publication No. 2014 / 0306866, which is incorporated herein by reference in its entirety.

[0134] Continuing with reference to FIG. 10A, the display device determines a pose based on the descriptor matches (block 1012). As described above, the display device can identify real-world coordinates (e.g., 3D coordinates relative to a particular coordinate reference frame) for each salient point contained in the current image. The display device can then determine its pose, for example, utilizing an n-point perspective algorithm. Information associated with the imaging device, such as intrinsic camera parameters, can be utilized to determine the pose. Thus, the pose determined by the display device can represent a camera pose. The display device can adjust this camera pose to determine a user's pose, a pose associated with the front (e.g., center) of the display device, etc. (e.g., block 1016). For example, the display device can linearly transform the camera pose according to a known translation or rotation offset of the camera from the user.

[0135] In some embodiments, the display device can utilize information obtained from an IMU to determine the attitude. For example, the information can be utilized as an initial value, i.e., a regularization value, to determine the attitude. The display device can then use the inertial measurement unit information as a cost function associated with the determination. As an example, the amount of divergence from the inertial measurement unit information can be associated with a cost. In this manner, the inertial measurement information can be taken into account and the accuracy of the resulting attitude determination can be improved. Similarly, the display device may utilize information associated with an extended Kalman filter and / or visual inertial odometry.

[0136] Similarly, the display device can utilize information acquired during frame / frame tracking as a similarity initial value, i.e., regularization value. As explained above, the display device can minimize the combined photometric cost function of all patches by varying the pose of the current image. In this way, a refined pose estimate can be identified. This refined pose estimate can be utilized as an initial value, optionally in combination with an IMU, an extended Kalman filter, and / or visual inertial odometry.

[0137] Following the pose determination, the display device can generate a patch for each salient point contained in the current image. For example, the display device can generate a patch for the newly extracted salient point from the current image and also for the tracked salient points from the previous image. Generating the patches can include obtaining an M×N pixel area surrounding each salient point in the current image. Optionally, for the tracked salient points from the previous image, the display device can utilize the patch associated with the previous image. That is, when tracking the salient points in the subsequent image, the patch from the previous image (e.g., not the current image) can be utilized in the frame / frame tracking. The display device can then acquire the subsequent image, and blocks 1002-1016 can be repeated for this subsequent image.

[0138] Continuing with reference to FIG. 10A, the display device determines a pose from the tracked salient points (block 1014). If the image area of ​​the current image does not include less tracked salient points than the threshold measurement, the display device can optionally determine its pose based on the tracked salient points. For example, the display device can determine its pose using, for example, an n-point perspective algorithm. Optionally, the display device can use an IMU, an extended Kalman filter, or visual inertial odometry information as an initial or regularization value. Optionally, the display device can perform block 1012 and not perform block 1014 if the image area does not include less tracked salient points than the threshold measurement.

[0139] The display device determines a pose of the display device user (block 1016). The pose of the display device may represent a camera pose, e.g., a pose associated with the imaging device. The display device may adjust this pose based on a known offset of the user from the camera. Optionally, the display device may perform initial training when the user wears the display device to, e.g., determine an appropriate offset. This training may inform the user's line of sight relative to the imaging device and may be utilized to determine the pose of the display device user. Some examples of methods for performing initial training may be found in U.S. Patent Application No. 15 / 717,747, filed Sep. 27, 2017, which is incorporated herein by reference in its entirety.

[0140] 11 illustrates a flowchart of an example process 1100 for frame / frame tracking. In some embodiments, the process 1100 may be described as being performed by a display device (e.g., display system 60, which may include processing hardware and software and may optionally provide information for processing to one or more computational external systems, e.g., offload processing to the external systems and receive information from the external systems). In some embodiments, the display device may be a virtual reality display device, including one or more processors.

[0141] The display device obtains patches associated with each salient point from the previous image (block 1102). As described above with respect to FIG. 10A, the display device can store a patch for each salient point that is being tracked. Thus, when the current image is obtained, the display device can obtain (e.g., from a stored memory) patch information associated with salient points contained in the previous image.

[0142] The display device projects each acquired patch onto the current image (block 1104). Reference is now made to Figure 12A, which illustrates an example of a previous image (e.g., image A 1202) and a current image (e.g., image B 1204). Each image is illustrated as including salient points that are being tracked.

[0143] As described above with respect to FIG. 10A, the display device can determine a pose estimate associated with the current image B 1204. In addition, each salient point included in the previous image A 1202 has a known real-world location or coordinates (e.g., based on map / frame tracking previously performed for the current image). Thus, based on these real-world locations, the pose of the previous image A 1202, and the pose estimate, a projection of the salient points onto the image B 1204 can be determined. For example, the pose of the previous image A 1202 can be adjusted according to the pose estimate and based on the real-world locations of the salient points in the image A 1202, and the salient points can be projected onto the image B 1204 at the 2D location of the image 1204. As shown, the tracked salient points 1208 are associated with a real-world location 1206. Based on the current real-world location 1206, the display device has determined that the salient points 1208 are located in the image B 1204 at an initial estimated 2D location. As explained above, optionally, the attitude estimate can be refined via information from an inertial measurement unit, an extended Kalman filter, visual inertial odometry, etc.

[0144] 12B illustrates a patch 1212 associated with a salient point 1208 being projected onto image B 1204. As will be explained below, the projected patch 1212 can be adjusted in location on image B 1204. The adjustment can be based on reducing an error associated with the projected patch 1212 and the corresponding image area of ​​image B 1204. The error can be, for example, a difference in pixel values ​​(e.g., intensity values) between pixels of the projected patch 1212 and pixels in the corresponding image area of ​​image B 1204, and the position of the patch 1208 is adjusted to minimize the difference in values.

[0145] Referring again to FIGS. 11 and 12A, the display device determines an image area in the current image that matches the projected patch (block 1106). As illustrated in FIG. 12B, a patch 1212 is projected onto image B 1204. The projection may represent, for example, an initial estimated location. The display device may then adjust the location of the projected patch 1212 on image B 1204 to refine the estimate. For example, the location of the projected patch 1212 may be moved horizontally or vertically from the initial estimated location by one or more pixels. For each adjusted location, the display device may determine the difference between the projected patch 1212 and the same M×N image area of ​​image B 1204 on which the patch 1212 is located. For example, the difference in individual pixel values ​​(e.g., intensity values) may be calculated. The display device may adjust the patch 1212 based on photometric error optimization. For example, Levenberg-Marquardt, conjugate gradient, etc. may be utilized to reach a local error minimum, a global error minimum, an error below a threshold, etc.

[0146] The display device identifies a tracked salient point in the current image (block 1108). For example, a tracked salient point 1208 can be identified as having a 2D location that corresponds to the centroid of a patch to be adjusted 1212 on image B 1204. Thus, as shown, the salient point 1208 is being tracked from image A 1202 to image B 1204.

[0147] 13 illustrates a flowchart of an example process 1300 for map / frame tracking. In some embodiments, the process 1300 may be described as being performed by a display device (e.g., an augmented reality display system 60, which may include processing hardware and software and may optionally provide information for processing to one or more computational external systems, e.g., offload processing to the external systems and receive information from the external systems). In some embodiments, the display device may be a virtual reality display device, including one or more processors.

[0148] The display device extracts new salient points from the current image (block 1302). As described above with respect to FIG. 10A, the display device can identify image areas of the current image that have tracked salient points below a threshold measurement. Exemplary image areas are illustrated in FIG. 10B. For these identified image areas, the display device can extract new salient points (e.g., identify locations in the current image that illustrate the salient points). For example, for salient points that are corners, the display device can perform Harris corner detection, features from Accelerated Segment Test (FAST) corner detection, etc. on the identified image areas.

[0149] The display device generates a descriptor for each salient point (block 1304). The display device may generate descriptors for (1) tracked salient points (e.g., tracked salient points from a previous image) and (2) newly extracted salient points. As described above, the descriptors may be generated to describe the visual focus of the salient points (e.g., as imaged in the current image) or the M×N image area surrounding the salient points. For example, the descriptors may indicate shape, color, texture, etc. associated with the salient points. As another example, the descriptors may indicate histogram information associated with the salient points.

[0150] The display device projects the real-world locations of the salient points onto the current image (block 1306). Reference is now made to FIG. 14A, which illustrates current image B 1204 along with newly extracted salient points (e.g., salient point 1402). For example, the display device determined that a lower portion of image B 1204 contains tracked salient points below the threshold measurement, and extracted salient points from this lower portion. Thus, in the example of FIG. 14A, image B 1204 contains seven salient points, i.e., four tracked salient points and three newly extracted salient points.

[0151] The display device identifies real-world locations that correspond to salient points contained within image B 1204. This identification can be an initial estimate of the real-world locations for the salient points contained within image B 1204. As will be explained, this estimate can be refined based on the descriptor matches such that each real-world location of the salient points in image B 1204 can be precisely determined.

[0152] With respect to the tracked salient point 1208, the display device can identify that the tracked salient point 1208 is likely within a threshold real-world distance of the real-world location 1206. Because the salient point 1208 was tracked from the previous image A 1202 (e.g., as illustrated in FIGS. 12A-12B), the display device has access to the real-world location of the salient point 1208. That is, map / frame tracking has already been performed on the previous image A 1202, and thus the real-world location 1206 is identified. As illustrated in FIGS. 11-12B, the salient point has been accurately tracked from the previous image A 1202 to the current image B 1204. Thus, the display device can identify that the tracked salient point 1208 corresponds to the same real-world location 1206 as its matching salient point in the previous image A 1202 (e.g., the salient point 1208 illustrated in FIG. 12A). The display device can therefore compare the generated descriptor for the tracked salient point 1208 with descriptors of real-world salient points within a threshold distance of the location 1206. In this manner, the real-world salient points can be matched to the tracked salient points 1208.

[0153] For the newly extracted salient point 1402, the display device can identify that the salient point 1402 is likely within a threshold real-world distance of the real-world location 1404. For example, the display device can utilize the map information, optionally together with a pose estimate for image B 1204, to identify an initial estimate for the real-world location 1402 of the salient point. That is, the display device can access information indicating the pose of the previous image A 1202 and adjust the pose according to the pose estimate. Optionally, the pose estimate can be refined according to the technique described in FIG. 10A. The display device can then determine an initial estimate for a real-world location corresponding to the 2D location of the salient point 1402 based on the adjusted pose. Through descriptor matching, as will be described below, the real-world location for the salient point 1402 can be determined. The initial estimate (e.g., real-world location 1404) can thus enable a reduction in the number of comparisons between the descriptor for the salient point 1402 and the descriptors of the salient points shown in the map information.

[0154] 14A, the display device matches the salient point descriptors (block 1308). The display device can compare the salient point descriptors indicated in the map information with the generated descriptors for each salient point in the current image (e.g., as described in block 1304) to find a suitable match.

[0155] As described above, an initial projection of salient points shown in the map information onto the current image can be identified. As an example, a number of real-world salient points may be proximate to the real-world location 1404. The display device can compare the descriptors for these multiple salient points with the descriptors generated for the tracked salient points 1402. Thus, the initial projection can enable a reduction in comparisons that need to be performed to enable the display device to identify likely real-world locations for the salient points 1402. The display device can match the descriptors that are mostly similar, for example, based on one or more similarity measures (e.g., differences in histograms, shape, color, texture, etc.). In this way, the display device can determine a real-world location that corresponds to each salient point included within the current image B 1204.

[0156] The display device can then determine its pose, as described in Figure 10A. The display device can then generate a patch for each salient point contained in the current image B 1204. As these salient points are tracked in subsequent images, the patches will be utilized to perform frame / frame tracking, as described above.

[0157] For example, FIG. 14B illustrates frame / frame tracking after map / frame tracking. In an exemplary illustration, image B 1204 (e.g., the current image in FIGS. 12A-B and 14A-B) now represents the previous image, and image C 1410 represents the current image. In this example, salient point 1412 is projected onto current image C 1410. That is, a patch associated with salient point 1412, generated, for example, following the display device determining its pose as described above, can be projected onto current image C 1410. Optionally, as described in FIG. 10A, the patch associated with salient point 1412 can be the same patch as that obtained in previous image A 1202.

[0158] Thus, frame / frame tracking can be performed by the display device. Similar to the above description, the current image C1410 can then be analyzed and any image areas of the current image C1410 with tracked salient points below a threshold measurement can be identified. Map / frame tracking can then be performed and a new pose can be determined.

[0159] 15 illustrates a flow chart of an exemplary process for determining head pose. For convenience, process 1500 may be described as being performed by a display device (e.g., display system 60), which may include processing hardware and software and may optionally provide information to one or more computational external systems for processing, e.g., offload processing to the external systems, and receive information from the external systems. In some embodiments, the display device may be a virtual reality display device, including one or more processors.

[0160] The display device projects the tracked salient points onto the current image in block 1502. As described above with respect to FIG. 10A, the display device can track two-dimensional image locations of the salient points between images acquired via one or more imaging devices. For example, a particular corner may be included (e.g., illustrated) in the first image at a particular two-dimensional (2D) location (e.g., one or more pixels of the first image) of the first image. Similarly, a particular corner may be included in the second image at a different 2D location of the second image. As described above, the display device can determine that a particular 2D location of the first image corresponds to a different 2D location of the second image. For example, each of these different 2D locations illustrates a particular corner, and the particular corner is therefore tracked between the first image and the second image. Correspondence between the first image and the second image can be determined according to these tracked salient points.

[0161] As illustrated in FIG. 15, the output of block 1514, in which the head pose is calculated with respect to the previous image, is obtained by the display device for use in block 1502. In addition, the matched salient points are obtained by the display device for use in block 1502. Thus, in this embodiment, the display device has access to the user's previous calculated head pose and information associated with the salient points to be tracked from the previous image to the current image. The information may include the real world location of the salient points and a separate patch associated with the salient points (e.g., an M×N image area of ​​the previous image surrounding the salient points), as described above. The information may be utilized to track the salient points, as will be described.

[0162] In block 1502, the display device acquires a current image (e.g., as described in FIG. 10A above) and projects salient points contained in the previous image (e.g., as shown in the figure) onto the current image, or alternatively, projects salient points from the map onto the current image. The display device may identify an initial estimated location in the current image to which each salient point corresponds. For example, FIG. 14B illustrates an exemplary salient point 1412, represented as a 2D location in previous image B 1204. This exemplary salient point 1412 is then projected onto current image C 1410. For example, the display device may utilize a calculated pose for the previous image (e.g., block 1514), a pose estimate for the current image, and real-world locations of salient points contained in the previous image (e.g., as shown in the figure). The pose estimate may be based on a trajectory prediction and / or inertial measurement unit (IMU) information, as described above. The IMU information is optional, but in some embodiments may improve the pose estimate for the trajectory prediction or the trajectory prediction itself. In some embodiments, the pose estimate may be the same pose derived from the previous image. The display device may utilize the pose estimate to adjust the calculated pose for the previous image. Because the display device has access to the real-world locations of the salient points, the display device may project these real-world locations onto the two-dimensional current image based on the adjusted pose.

[0163] Thus, the display device can estimate 2D locations of the current image that correspond to individual salient points. As described above, the display device can store a patch for each tracked salient point. The patch can be an M×N image area that surrounds the 2D location of the image illustrating the salient point. For example, the patch can extend a set number of pixels along the horizontal direction of the image from the 2D location of the salient point. Similarly, the patch can extend a set number of pixels along the vertical direction of the image from the 2D location of the salient point. The display device can acquire a patch associated with each salient point, e.g., an M×N image area of ​​the previous image that surrounds each patch. Each acquired patch can then be projected onto the current image. As an example, a patch associated with a particular salient point may be acquired. The patch can be projected onto the current image to surround the estimated 2D location of the particular salient point. As described above, the 2D location of the projected patch can be adjusted based on photometric error minimization. For a particular salient point example, the display device can determine the error between the patch and the M×N area of ​​the current image onto which the patch is projected. The display device can then adjust the location of the patch (e.g., along the vertical and / or horizontal directions) as disclosed herein until the error is reduced (e.g., minimized).

[0164] The display device may optionally refine the pose estimate at block 1504. An initial pose estimate may be determined as described above, but optionally, the display device may refine the pose estimate. The display device may utilize the refined pose estimate as an initial value when calculating the head pose (e.g., the refined pose estimate may be associated with a cost function).

[0165] As illustrated in FIG. 10A, the display device can minimize the combined photometric cost function of all projected patches by varying the pose estimate of the current image. Due to the varying pose estimate, the estimated 2D locations of the current image corresponding to individual salient points will be adjusted. Thus, the projected patches on the current image may be adjusted globally according to the varying pose estimate. The display device varies the pose estimate until a minimum combined error (e.g., a global or local minimum, a minimum below a threshold, etc.) between the projected patch and the corresponding image area of ​​the current image is identified. As illustrated in process 1500, inertial measurement unit information, extended Kalman information (EKF), visual inertial odometry (VIO) information, etc. may be utilized as predictions when refining the pose estimate.

[0166] The display device refines the 2D location of the projected salient point in block 1506. As described above, the display device can project a patch (e.g., an image area of ​​the previous image surrounding the salient point) onto the current image. The display device can then compare (1) the patch and (2) the M×N image area of ​​the current image onto which the patch was projected. First, the display device can compare the patch associated with the salient point with the M×N image area of ​​the current image surrounding the salient point. Then, the display device can adjust the M×N image area along a vertical direction (e.g., upward or downward in the current image) and / or a horizontal direction (e.g., left or right in the current image). After each adjustment, the patch can be compared to the new M×N image area and an error can be determined. For example, the error can represent a sum of pixel intensity differences between corresponding pixels in the patch and the M×N image area (e.g., the difference between the top left pixel of the patch and the top left pixel of the image area can be calculated, and so on). As described above, according to an error minimization scheme such as Levenberg-Marquardt, the display device can identify an M×N image area of ​​the current image that minimizes the error with the patch. 2D locations of the current image that are encompassed by the identified M×N image area can be identified as salient points associated with the patch. Thus, the display device can track the 2D locations of the salient points between the previous image and the current image.

[0167] The display device extracts salient points in image areas with tracked salient points below the threshold measurement at block 1508. As described above with respect to Figures 10A-10B, the display device can identify image areas of the current image where new salient points should be identified. An extraction process, such as Harris angle detection, can be applied to the current image and 2D locations of the current image that correspond to the new salient points can be identified.

[0168] The display device then generates, at block 1510, a descriptor for the salient points contained within the current image. The display device may generate the descriptor based on 2D locations in the current image that correspond to the salient points. The salient points include salient points tracked from a previous image to the current image and newly identified salient points in the current image. As an example, a descriptor for a particular salient point may be generated based on pixels associated with the 2D location of the particular salient point or based on the image area surrounding the 2D location.

[0169] The display device matches the descriptors included in the map information to the generated descriptors in block 1512. As described in FIG. 13 above, the display device may access the map information and match the descriptors included in the map information to the generated descriptors. The map information may include real-world coordinates (e.g., 3D coordinates) of salient points along with the descriptors associated with those real-world coordinates. Thus, a match between the map information descriptor and the generated descriptor indicates the real-world coordinates of the salient points associated with the generated descriptor.

[0170] To match the descriptors, the display device can compare each generated descriptor for a salient point contained in the current image with a descriptor contained in the map information. To limit the number of comparisons performed, the display device can estimate the real-world location of the salient point contained in the current image. For example, the salient points tracked from the previous image to the current image have known real-world coordinates. As another example, the real-world coordinates of a newly identified salient point in the current image can be estimated according to the pose estimate of the display device. The display device can then use these estimated real-world coordinates to identify the portion of the real-world environment in which each salient point is estimated to be contained. For example, a particular salient point contained in the current image can be determined to have estimated real-world coordinates. The display device can compare the generated descriptor for this particular salient point with a descriptor contained in the map information associated with real-world coordinates within a threshold distance of the estimated real-world coordinates. Thus, the number of comparisons between the descriptors contained in the map information and the generated descriptors can be reduced because the display device can focus the comparison.

[0171] The display device calculates the head pose in block 1514. As described above, the display device can calculate the head pose based on the real-world coordinates of salient points contained in the current image and their corresponding 2D locations in the current image. For example, the display device can implement an n-point perspective algorithm using the camera information (e.g., intrinsic camera parameters) of the imaging device. In this manner, the display device can determine the camera pose of the imaging device. The display can then linearly transform this camera pose to determine the user's head pose. For example, the translation and / or rotation of the user's head relative to the camera pose can be calculated. The user's head pose can then be utilized by the display device for subsequent images, for example, the head pose can be utilized in block 1502.

[0172] Optionally, the display device can use the refined pose estimate as an initial value when calculating the computing head pose, as described in block 1504. In addition, the display device can use inertial measurement unit information, extended Kalman filter information, inertial visual odometry information, etc. as an initial value.

[0173] Computer vision for detecting objects in the environment As discussed above, the display system may be configured to detect objects or properties thereof in the environment surrounding the user. Detection may be accomplished using a variety of techniques, including various environmental sensors (e.g., cameras, audio sensors, temperature sensors, etc.), as discussed herein. For example, the object may represent a salient point (e.g., a corner).

[0174] In some embodiments, objects present in the environment may be detected using computer vision techniques. For example, as disclosed herein, a front-facing camera of a display system may be configured to image the surrounding environment, and the display system may be configured to perform image analysis on the images to determine the presence of objects in the surrounding environment. The display system may analyze images obtained by an outward-facing imaging system and perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, image restoration, and the like. As another example, the display system may be configured to perform face and / or eye recognition to determine the presence and location of faces and / or human eyes within the user's field of view. One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale Invariant Feature Transform (SIFT), Speed ​​Up Robust Features (SURF), Orientation FAST and Rotation BRIEF (ORB), Binary Robust Invariant Scalable Key Points (BRISK), Fast Retinal Key Points (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous localization and mapping (vSLAM) techniques, sequential Bayes estimators (e.g., Kalman filter, extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), iterative nearest neighbor (ICP), semi-global matching (SGM), semi-global block matching (SGBM), feature point histograms, various machine learning algorithms (e.g., support vector machines, k-nearest neighbor algorithms, naive Bayes, neural networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), and the like.

[0175] One or more of these computer vision techniques may also be used in conjunction with data obtained from other environmental sensors (e.g., microphones, etc.) to detect and determine various properties of objects detected by the sensors.

[0176] As discussed herein, objects within the surrounding environment may be detected based on one or more criteria. When the display system detects the presence or absence of a criterion within the surrounding environment using computer vision algorithms or using data received from one or more sensor assemblies (which may or may not be part of the display system), the display system may then signal the presence of the object.

[0177] (Machine Learning) Various machine learning algorithms may be used to learn to identify the presence of objects in the surrounding environment. Once trained, the machine learning algorithms may be stored by the display system. Some examples of machine learning algorithms may include supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., a priori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural networks, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., stacked generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models may be customized for individual data sets. For example, the wearable device may generate or store a base model. The base model may be used as a starting point to generate additional models specific to a data type (e.g., a particular user), a data set (e.g., a set of additional images acquired), a conditional situation, or other variations. In some embodiments, the display system may be configured to utilize multiple techniques to generate models for analysis of the aggregated data. Other techniques may include using predefined thresholds or data values.

[0178] The criteria for detecting the object may include one or more threshold conditions. If the analysis of the data obtained by the environmental sensors indicates that the threshold conditions are passed, the display system may provide a signal indicating the detection of the presence of the object in the surrounding environment. The threshold conditions may involve quantitative and / or qualitative measurements. For example, the threshold conditions may include a score or percentage associated with the likelihood that the object is present in the environment. The display system may compare the score calculated from the data of the environmental sensors to the threshold score. If the score is higher than the threshold level, the display system may detect the reflection and / or the presence of the object. In some other embodiments, the display system may signal the presence of the object in the environment if the score is lower than the threshold. In some embodiments, the threshold conditions may be determined based on the user's emotional state and / or the user's interaction with the surrounding environment.

[0179] It should be understood that each of the processes, methods, and algorithms described herein and / or depicted in the figures may be embodied in code modules executed by one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, and thereby be fully or partially automated. For example, a computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer programmed with specific computer instructions, special-purpose circuits, etc. The code modules may be written in a programming language that may be compiled and linked into an executable program, installed in a dynamic link library, or interpreted. In some implementations, certain operations and methods may be performed by circuitry specific to a given function.

[0180] Furthermore, certain implementations of the functionality of the present disclosure may be sufficiently mathematically, computationally, or technically complex that special purpose hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may be required to implement the functionality, e.g., due to the amount or complexity of the calculations involved, or to provide results in substantially real-time. For example, a video may contain many frames, each frame may have millions of pixels, and specifically programmed computer hardware may be required to process the video data to provide the desired image processing task or application in a commercially reasonable amount of time.

[0181] The code modules or any type of data may be stored on any type of non-transitory computer readable medium, such as physical computer storage devices, including hard drives, solid state memory, random access memory (RAM), read only memory (ROM), optical disks, volatile or non-volatile storage devices, combinations of the same, and / or the like. In some embodiments, the non-transitory computer readable medium may be part of one or more of the local processing and data module (140), the remote processing module (150), and the remote data repository (160). The methods and modules (or data) may also be transmitted as a data signal generated (e.g., as part of a carrier wave or other analog or digital propagating signal) over a variety of computer readable transmission media, including wireless-based and wired / cable-based media, and may take a variety of forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed processes or process steps may be stored persistently or otherwise in any type of non-transitory tangible computer storage device, or communicated via a computer readable transmission medium.

[0182] Any process, block, state, step, or functionality in the flow diagrams described herein and / or depicted in the accompanying figures should be understood as potentially representing a code module, segment, or portion of code that includes one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in the process. Various processes, blocks, states, steps, or functionality may be combined, rearranged, added, deleted, modified, or otherwise altered from the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the blocks, steps, or states associated therewith may be performed in other sequences as appropriate, e.g., serially, in parallel, or in some other manner. Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all embodiments. It should be understood that the program components, methods, and systems described may generally be integrated together in a single computer product or packaged into multiple computer products.

[0183] The foregoing specification has been described with reference to specific embodiments thereof. It will, however, be apparent that various modifications and changes can be made therein without departing from the broader spirit and scope of the disclosure. The specification and drawings are, therefore, to be regarded in an illustrative rather than a restrictive sense.

[0184] Indeed, it should be understood that the systems and methods of the present disclosure each have several innovative aspects, none of which is solely responsible for or required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and subcombinations are intended to fall within the scope of the present disclosure.

[0185] Certain features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as acting in a combination and may even be initially claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or variation of the subcombination. No single feature or group of features is necessary or essential to every embodiment.

[0186] In particular, it should be understood that conditional statements used herein, such as "can," "could," "might," "may," "eg," and the like, are generally intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not, unless specifically stated otherwise or understood otherwise within the context as used. Thus, such conditional statements are generally not intended to suggest that features, elements, and / or steps are in any way required for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether those features, elements, and / or steps should be included or performed in any particular embodiment, with or without authorial input or prompting. The terms "comprise," "include," "have," and the like, are synonymous and used inclusively in a non-limiting manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not its exclusive sense), thus, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. In addition, the articles "a," "an," and "the," as used in this application and the appended claims, should be interpreted to mean "one or more" or "at least one," unless otherwise specified. Similarly, although operations may be depicted in the figures in a particular order, this should be recognized that such operations need not be performed in the particular order depicted, or in sequential order, or that all of the depicted operations need not be performed to achieve desirable results. Additionally, the figures may diagrammatically depict one or more exemplary processes in the form of a flow chart. However, other operations not depicted may also be incorporated within the diagrammatically depicted exemplary methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or during any of the depicted operations.In addition, operations may be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged into multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve desirable results.

[0187] Thus, the claims are not intended to be limited to the implementations shown in the specification but are to be accorded the widest scope consistent with the disclosure, the principles and novel features disclosed herein.

Claims

1. 1. A system comprising: one or more imaging devices; one or more processors; one or more computer storage media having instructions stored thereon; Equipped with The instructions, when executed by the one or more processors, acquiring a current image of a real-world environment via the one or more imaging devices, the current image including a plurality of points for determining a pose; projecting a patch-based first salient point from a previous image onto a corresponding one of the points in the current image; determining that an image area of ​​the current image has less than a threshold number of salient points projected from the past image and extracting one or more descriptor-based salient points from the image area, the extracted salient points including a second salient point from the current image, the image area comprising a subset of the current image, the system being configured to adjust a size associated with the subset during operation; providing respective descriptors for said salient features; matching salient points associated with the current image with real-world locations defined within a descriptor-based map of the real-world environment; determining a pose associated with the system based on the matching, the pose indicating at least an orientation of the one or more imaging devices within the real-world environment; The system further comprises:

2. The operations further include adjusting a position of the patch-based first salient point on the current image; Adjusting the position comprises: obtaining a first patch associated with the first salient point, the first patch including a portion of the past image that includes the first salient point and an area of ​​the past image surrounding the first salient point; locating a second patch in the current image similar to the first patch, the first salient point being located in a similar location in the second patch as the first patch; The system of claim 1 , comprising:

3. The system of claim 2 , wherein locating the second patch comprises minimizing a difference between the first patch in the past image and the second patch in the current image.

4. The system of claim 2 , wherein projecting the patch-based first salient point onto the current image is based at least in part on information from an inertial measurement unit of the system.

5. The system of claim 1 , wherein the system is configured to adjust a size associated with the subset based on one or more of a processing constraint or a difference between one or more previously determined poses.

6. Matching salient points associated with the current image with real-world locations defined within a map of the real-world environment includes: accessing map information, the map information comprising real-world locations of salient points and associated descriptors; Matching the salient feature descriptors of the current image with the salient feature descriptors of real-world locations; and The system of claim 1 , comprising:

7. The operations further include projecting salient points provided in the map information onto the current image; The system of claim 6 , wherein the projecting is based on one or more of an inertial measurement unit, an extended Kalman filter, or visual inertial odometry.

8. The system of claim 1 , wherein the system is configured to generate the map using at least the one or more imaging devices.

9. The system of claim 1 , wherein determining the pose is based on real-world locations of the salient points and relative positions of the salient points within a view captured in the current image.

10. The operations further include generating a patch associated with each salient point extracted from the current image; The system of claim 1 , wherein for a subsequent image of the current image, the patch includes salient points available for projecting onto the subsequent image.

11. The system of claim 1 , wherein providing a descriptor comprises generating a descriptor for each of the salient points.

12. 1. A system comprising: one or more outward-facing coupling devices wearable by a user, the one or more outward-facing coupling devices configured to capture images of a real-world environment in a vicinity of the user; one or more processors; one or more computer storage media having instructions stored thereon; Equipped with The instructions, when executed by the one or more processors, identifying, in a current image of the real world environment, a first salient point included in a previous image of the real world environment, the first salient point representing a first feature of the real world environment; determining that an image area of ​​the current image has less than a threshold number of salient points projected from the previous image and extracting one or more descriptor-based salient points from the image area, the extracted salient points including a second salient point from the current image, the image area comprising a subset of the current image, the system being configured to adjust a size associated with the subset during operation; accessing respective descriptors for the first salient point and the second salient point identified in the current image, the second salient point representing a second feature of the real-world environment, the system storing a descriptor-based map of the real-world environment indicating real-world locations associated with the first feature and the second feature; determining a pose of the system based on the accessed descriptors and the descriptor-based map; and The system further comprises:

13. 13. The system of claim 12, wherein identifying the first salient point in the current image of the real-world environment is accomplished by matching a portion of the current image with a patch stored by the system.

14. The system of claim 13 , wherein the portion of the current image is matched to the patches stored by the system based on minimizing a cost function.

15. The system of claim 13 , wherein the portion of the current image is matched to the patch using information from an inertial measurement unit of the system.

16. The system of claim 13 , wherein matching the patch with the portion of the current image includes projecting the patch onto the current image and refining a location associated with the patch.

17. The system of claim 12 , wherein the descriptor is generated based on respective image pixels associated with respective locations of the first salient point and the second salient point in the current image.

18. The system of claim 12 , wherein the descriptors are generated based on respective image areas associated with respective locations of the first salient point and the second salient point in the current image.

19. The system of claim 12 , wherein the descriptor-based map comprises a plurality of descriptors associated with a plurality of features, the plurality of features comprising the first feature and the second feature.

20. 20. The system of claim 19, wherein the operations further include identifying at least a first descriptor and a second descriptor by matching a subset of the plurality of descriptors with the accessed descriptor, the first descriptor and the second descriptor matching the accessed descriptor.

21. The descriptor-based map includes three-dimensional coordinates of salient points associated with the plurality of descriptors, and the pose is (1) a three-dimensional location associated with the first and second descriptors; and (2) two-dimensional locations of the first salient point and the second salient point within the current image; The system of claim 20, based on comparing:

22. The operations further include projecting a plurality of salient points contained within the descriptor-based map onto the current image; The system of claim 12 , wherein the projection is based on one or more of an inertial measurement unit, an extended Kalman filter, or visual inertial odometry.

23. 1. A method implemented by a system, the system including one or more outward-facing coupling devices wearable by a user, the one or more outward-facing coupling devices configured to capture images of a real-world environment in a vicinity of the user; The method comprises: identifying, in a current image of the real world environment, a first salient point included in a previous image of the real world environment, the first salient point representing a first feature of the real world environment; determining that an image area of ​​the current image has less than a threshold number of salient points projected from the previous image and extracting one or more descriptor-based salient points from the image area, the extracted salient points including a second salient point from the current image, the image area comprising a subset of the current image, the system being configured to adjust a size associated with the subset during operation; accessing respective descriptors for the first salient point and the second salient point identified in the current image, the second salient point representing a second feature of the real-world environment, the system storing a descriptor-based map of the real-world environment indicating real-world locations associated with the first feature and the second feature; determining a pose of the system based on the descriptors and the descriptor-based map; A method comprising:

24. 24. The method of claim 23, wherein identifying the first salient point in the current image of the real-world environment is accomplished by matching a portion of the current image with a patch stored by the system.

25. 25. The method of claim 24, wherein the portion of the current image is matched to the patches stored by the system based on minimizing a cost function.

26. The method of claim 24 , wherein the portion of the current image is matched to the patch using information from an inertial measurement unit of the system.

27. 25. The method of claim 24, wherein matching the patch with the portion of the current image comprises projecting the patch onto the current image and refining a location associated with the patch.

28. The descriptor is a representation of the first salient point and the second salient point in the current image.

24. The method of claim 23, wherein the image is generated based on respective image pixels associated with each location.

29. The method of claim 23, wherein the descriptors are generated based on respective image areas associated with respective locations of the first salient point and the second salient point in the current image.

30. The method of claim 23 , wherein the descriptor-based map comprises a plurality of descriptors associated with a plurality of features, the plurality of features comprising the first feature and the second feature.

31. 31. The method of claim 30, further comprising matching a subset of the plurality of descriptors with the accessed descriptor to identify at least a first descriptor and a second descriptor, the first descriptor and the second descriptor matching the accessed descriptor.

32. The descriptor-based map includes three-dimensional coordinates of salient points associated with the plurality of descriptors, and the pose is (1) a three-dimensional location associated with the first and second descriptors; and (2) two-dimensional locations of the first salient point and the second salient point within the current image; The method of claim 31 , based on comparing:

33. The method further comprises projecting a plurality of salient points contained in the descriptor-based map onto the current image; The method of claim 23 , wherein the projection is based on one or more of an inertial measurement unit, an extended Kalman filter, or visual inertial odometry.

34. 1. An augmented reality display system, comprising: one or more outward-facing coupling devices wearable by a user, the one or more outward-facing coupling devices configured to capture images of a real-world environment in a vicinity of the user; One or more processors Equipped with The one or more processors: acquiring a current image of the real world environment; performing frame-to-frame tracking on the current image such that patch-based salient points contained in the past image are matched with locations in the current image; extracting one or more descriptor-based salient points from an image area having a number of salient points less than a threshold number projected from the past image, the image area comprising a subset of the current image, the augmented reality display system being configured to adjust a size associated with the subset during operation; performing map / frame tracking on the current image, the map / frame tracking including matching the patch-based salient point descriptors with map-based descriptors stored in a descriptor-based map of the real-world environment; determining a pose of the augmented reality display system; 16. An augmented reality display system configured to:

35. 35. The augmented reality display system of claim 34, wherein matching patch-based salient points with locations in the current image is based on minimizing a cost function.

36. 35. The augmented reality display system of claim 34, wherein a subset of the map-based descriptors is determined to match the descriptors for the patch-based salient points.

37. 37. The augmented reality display system of claim 36, wherein the pose is based on comparing (1) real-world locations associated with the subset of the map-based descriptors and (2) two-dimensional locations of the patch-based salient points in the current image.

38. 35. The augmented reality display system of claim 34, further comprising an inertial measurement unit (IMU), the one or more processors configured to perform frame-to-frame tracking by using the IMU.

Citation Information

Patent Citations

  • Detecting device for object in front of vehicle

    JP1997095194A

  • Robust Tracking Using Point and Line Features

    JP2016521885A

  • Device and method of detecting gradual shot transition in moving picture

    US20070248243A1

  • Image-based localization

    US20140010407A1