Multi-camera cross-reality device

A lightweight XR headset with a calibration routine and efficient sensor management addresses weight and power issues, ensuring accurate alignment of virtual and physical objects for enhanced realism.

JP2026012401APending Publication Date: 2026-01-23MAGIC LEAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025185585
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-12-21
Filing Date
2025-11-04
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing wearable cross-reality (XR) systems face challenges with headset weight, sensor positional shifts, high power consumption, and inaccurate stereoscopic imaging due to camera misalignment, which affect user enjoyment and realism of virtual and physical object interaction.

Method used

A lightweight headset with two cameras and an inertial measurement unit (IMU) performs a calibration routine to determine relative orientation, uses grayscale cameras with global shutters, and selectively activates sensors to maintain accurate stereoscopic depth information with reduced power consumption.

Benefits of technology

The system provides accurate, low-power, and lightweight XR experiences with enhanced realism by compensating for camera shifts and optimizing sensor usage, ensuring virtual objects align correctly with physical objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012401000001_ABST
    Figure 2026012401000001_ABST
Patent Text Reader

Abstract

To provide a suitable multi-camera cross reality device.SOLUTION: A wearable display system with a limited number of cameras. The two cameras can be arranged to provide an overlapping central field of view associated with one of the two cameras and a peripheral field of view. A third camera can be arranged to provide a color field of view that overlaps the central field of view. The wearable display system may be coupled to a processor configured to use the two cameras to generate a world model and track motion of the hand in the central field of view. The processor may be configured to perform a calibration routine to compensate for distortion during use of the wearable display system.SELECTED DRAWING: Figure 22
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application generally relates to a wearable cross-reality display system that includes a camera. [Background technology]

[0002] A computer may control a human user interface and create an X-reality (XR or cross-reality) environment in which part or all of the XR environment is generated by the computer as perceived by the user. These XR environments may be virtual reality (VR), augmented reality (AR), or mixed reality (MR) environments in which part or all of the XR environment may be generated by the computer using data that describes the environment. This data may, for example, describe virtual objects that may be rendered in a way that a user senses or perceives as part of the physical world so that the user may interact with the virtual objects. The user may experience these virtual objects as a result of data being rendered and presented through a user interface device, such as a head-mounted display device. The data may control audio that may be displayed to the user to see or played to the user to hear, or may control a tactile (or haptic) interface, allowing the user to experience touch sensations that the user senses or perceives as they feel the virtual objects.

[0003] XR systems can be useful for many applications, ranging from scientific visualization, medical training, engineering design and prototyping, remote manipulation and telepresence, and personal entertainment. AR and MR, in contrast to VR, involve one or more virtual objects in association with real objects in the physical world. Experiencing virtual objects interacting with real objects generally enhances the user's enjoyment when using XR systems and opens up possibilities for a variety of applications that present realistic and easily understandable information about how the physical world can be altered. Summary of the Invention [Means for solving the problem]

[0004] Aspects of the present application relate to a wearable cross-reality display system that includes a camera. The techniques described herein may be used together, separately, or in any suitable combination.

[0005] According to some embodiments, a wearable display system is provided, the wearable display system may include: a headset; two first cameras mechanically coupled to the headset, wherein a first inertial measurement unit is mechanically coupled to a first of the two first cameras and a second inertial measurement unit is mechanically coupled to a second of the two first cameras; and a processor operably coupled to the two first cameras and configured to perform a calibration routine configured to determine a relative orientation of the two first cameras using images acquired by the two first cameras and outputs of the first inertial measurement unit and the second inertial measurement unit.

[0006] In some embodiments, the calibration routine, when executed, may further determine the relative positions of the two first cameras. In some embodiments, performing the calibration routine may include identifying corresponding features in images acquired using each of the first and second first cameras; calculating an error for each of a plurality of estimated relative orientations of the two first cameras, the error indicating a difference between the corresponding feature as it appears in the images acquired using each of the two first cameras and an estimate of the identified feature calculated based on the estimated relative orientations of the two first cameras; and selecting a relative orientation of the plurality of estimated relative orientations based on the calculated error as the determined relative orientation. In some embodiments, the calibration routine may further include selecting at least an initial estimate of the plurality of estimated relative orientations based, in part, on outputs of the first inertial measurement unit and the second inertial measurement unit.

[0007] In some embodiments, the wearable display system may further include a color camera coupled to the headset, and the processor may be further configured to create a world model using the two first cameras and the color camera, update the world model at a first rate using the two first cameras, and update the world model at a second rate slower than the first rate using the two first cameras and the color camera. In some embodiments, the two first cameras may be mechanically coupled to the headset to provide a central field of view associated with both of the first cameras and a peripheral field of view associated with a first of the two first cameras, and the processor may further be configured to track hand movements within the central field of view using depth information determined from images obtained by the two first cameras.

[0008] In some embodiments, tracking hand movement within the central visual field using depth information may include selecting a point within the central visual field, stereoscopically determining depth information for the selected point using images acquired by the two first cameras, generating a depth map using the stereoscopically determined depth information, and matching portions of the depth map to corresponding portions of a model of the hand, including both shape and movement constraints. In some embodiments, the processor may be further configured to track hand movement within the peripheral visual field using one or more images acquired from the two first cameras by matching portions of the images to corresponding portions of a model of the hand, including both shape and movement constraints.

[0009] In some embodiments, the headset may be a lightweight headset weighing between 30 and 300 grams. In some embodiments, the headset may further comprise a processor. In some embodiments, the headset may further comprise a battery pack. In some embodiments, the two first cameras may be configured to obtain grayscale images. In some embodiments, the processor may be mechanically coupled to the headset. In some embodiments, the headset may comprise a display device mechanically coupled to the processor. In some embodiments, a local data processing module may comprise the processor, the local data processing module may be operably coupled to the display device through a communications link, and the headset may comprise the display device.

[0010] According to some embodiments, a wearable display system is provided, the wearable display system may include: two first cameras mechanically coupled to the frame to provide a central field of view associated with both the frame and the cameras and a first peripheral field of view associated with a first of the two first cameras; a color camera mechanically coupled to the frame to provide a color field of view overlapping the central field of view; and a processor operably coupled to the two first cameras and the color camera and configured to track hand movement within the central field of view using stereoscopically determined depth information from first images acquired by the two first cameras, track hand movement within the first peripheral field of view using one or more second images acquired by a first of the two first cameras, create a world model using the two first cameras and the color camera, and update the world model using the two first cameras.

[0011] In some embodiments, the two first cameras may be configured with a global shutter. In some embodiments, the wearable display system may further include a hardware accelerator for stereoscopically determining depth information using the first grayscale images acquired by the two first cameras. In some embodiments, the two first cameras may have equidistant lenses. In some embodiments, the two first cameras may each have a horizontal field of view between 90 degrees and 175 degrees. In some embodiments, the central field of view may have an angular range between 40 and 80 degrees. In some embodiments, the processor may further be configured to perform a calibration routine to determine a relative orientation of the two first cameras.

[0012] In some embodiments, the calibration routine may include the steps of identifying corresponding features in images acquired using each of the two first cameras; calculating an error for each of a plurality of estimated relative orientations of the two first cameras, the error indicating the difference between the corresponding feature as it appears in the images acquired using each of the two first cameras and an estimate of the identified feature calculated based on the estimated relative orientations of the two first cameras; and selecting a relative orientation from the plurality of estimated relative orientations as the determined relative orientation based on the calculated error.

[0013] In some embodiments, the wearable display system may further include a first inertial measurement unit mechanically coupled to a first of the two first cameras and a second inertial measurement unit mechanically coupled to a second of the two first cameras, and the calibration routine may further include selecting at least one of a plurality of estimated relative orientations based in part on outputs of the first inertial measurement unit and the second inertial measurement unit.

[0014] In some embodiments, the processor may be further configured to repeatedly perform the calibration routine while the wearable display system is worn, such that the calibration routine compensates for distortion in the frame during use of the wearable display system. In some embodiments, the calibration routine may compensate for distortion in the frame caused by changes in temperature. In some embodiments, the calibration routine may compensate for distortion in the frame caused by mechanical distortion. In some embodiments, the two first cameras may be configured to obtain grayscale images. In some embodiments, the wearable display system may have one color camera. In some embodiments, the processor may be mechanically coupled to the frame. In some embodiments, the display device may be mechanically coupled to the frame, and the display device comprises the processor. In some embodiments, the local data processing module may comprise the processor, and the local data processing module may be operably coupled to the display device through a communications link, and the display device may be mechanically coupled to the frame. In some embodiments, the first image may comprise one or more second images.

[0015] According to some embodiments, a wearable display system is provided, which may include a frame, two first cameras mechanically coupled to the frame, a color camera mechanically coupled to the frame, and a processor operably coupled to the two first cameras and the color camera and configured to create a world model using a first grayscale image acquired by the two first cameras and one or more color images acquired by the color camera, update the world model using a second grayscale image acquired by the two first cameras, determine portions of the world model that include incomplete depth information, and update the world model with additional depth information for portions of the world model that include incomplete depth information.

[0016] In some embodiments, the wearable display system may further include one or more emitters, and additional depth information may be obtained using the one or more emitters, and the processor may be further configured to enable the one or more emitters in response to determining that the portion of the world model includes incomplete depth information. In some embodiments, the one or more emitters may include an infrared emitter, and the two first cameras may include filters configured to pass infrared light. In some embodiments, the infrared emitter may emit light having a wavelength between 900 nanometers and 1 micrometer.

[0017] In some embodiments, the processor may be further configured to detect a planar surface in the physical world in response to determining that the portion of the world model includes incomplete depth information, and may estimate additional depth information based on the detected planar surface. In some embodiments, the processor may be further configured to detect an object in the portion of the world model including incomplete depth information, identify an object template corresponding to the detected object, construct an instance of the object template based on an image of the object in the updated world model, and estimate additional depth information based on the constructed instance of the object template. In some embodiments, the two first cameras may be mechanically coupled to the frame to provide a central field of view associated with both of the first cameras and a peripheral field of view associated with a first of the two first cameras, and the processor may be further configured to track hand movements within the central field of view using depth information determined from images obtained by the two first cameras.

[0018] In some embodiments, tracking hand movement within the central visual field using depth information may include selecting a point within the central visual field, stereoscopically determining depth information for the selected point using images acquired by the two first cameras, generating a depth map using the stereoscopically determined depth information, and matching portions of the depth map to corresponding portions of a hand model, including both shape and movement constraints. In some embodiments, the processor may be further configured to track hand movement within the peripheral visual field using one or more images acquired from a first of the two first cameras by matching portions of the images to corresponding portions of a hand model, including both shape and movement constraints. In some embodiments, the processor may be mechanically coupled to the frame. In some embodiments, a display device mechanically coupled to the frame may comprise the processor. In some embodiments, a local data processing module may comprise the processor, and the local data processing module may be operably coupled to the display device through a communications link, and the display device may be mechanically coupled to the frame.

[0019] According to some embodiments, a wearable display system is provided, the wearable display system may include a frame; two grayscale cameras mechanically coupled to the frame, the two grayscale cameras being a first grayscale camera having a first field of view and a second grayscale camera having a second field of view, the first grayscale camera and the second grayscale camera being positioned to provide a central field of view where the first field of view overlaps the second field of view, and a first peripheral field of view that is within the first field of view and outside the second field of view; and a color camera mechanically coupled to the frame to provide a color field of view that overlaps the central field of view.

[0020] In some embodiments, the two grayscale cameras may have a global shutter. In some embodiments, the color camera may have a rolling shutter. In some embodiments, the wearable display system may include two inertial measurement units mechanically coupled to the frame and one or more emitters mechanically coupled to the frame. In some embodiments, the one or more emitters may be infrared emitters, and the two grayscale cameras may include a filter configured to pass infrared light. In some embodiments, the infrared emitter may emit light having a wavelength between 900 nanometers and 1 micrometer. In some embodiments, the illumination field of the one or more emitters may overlap with the central field of view. In some embodiments, a first inertial measurement unit of the two inertial measurement units may be mechanically coupled to a first grayscale camera, and a second inertial measurement unit of the two inertial measurement units may be mechanically coupled to a second grayscale camera of the two grayscale cameras. In some embodiments, the two inertial measurement units may be configured to measure tilt, acceleration, velocity, or any combination thereof. In some embodiments, the two grayscale cameras may have a horizontal field of view between 90 degrees and 175 degrees. In some embodiments, the central field of view may have an angular range between 40 and 80 degrees.

[0021] The foregoing description is provided by way of illustration and is not intended to be limiting. The present specification also provides, for example, the following items: (Item 1) 1. A wearable display system, comprising: A headset and two first cameras mechanically coupled to the headset, a first inertial measurement unit mechanically coupled to a first of the two first cameras and a second inertial measurement unit mechanically coupled to a second of the two first cameras; a processor operably coupled to the two first cameras, the processor comprising: performing a calibration routine configured to determine a relative orientation of the two first cameras using images acquired by the two first cameras and outputs of the first inertial measurement unit and the second inertial measurement unit; a processor configured to: A wearable display system comprising: (Item 2) Item 1. The wearable display system of item 1, wherein the calibration routine, when executed, further determines the relative positions of the two first cameras. (Item 3) performing the calibration routine identifying corresponding features in images acquired using each of the first and second first cameras; calculating an error for each of a plurality of estimated relative orientations of the two first cameras, the error indicating a difference between a corresponding feature as it appears in an image acquired using each of the two first cameras and an estimate of the identified feature calculated based on the estimated relative orientations of the two first cameras; selecting a relative orientation from the plurality of estimated relative orientations based on the calculated error as a determined relative orientation; and Item 1. The wearable display system of item 1, comprising: (Item 4) The wearable display system of item 2, wherein the calibration routine further includes selecting at least an initial estimate of the plurality of estimated relative orientations based, in part, on outputs of the first inertial measurement unit and the second inertial measurement unit. (Item 5) The wearable display system further comprises: a color camera coupled to the headset; The processor further comprises: creating a world model using the two first cameras and the color camera; updating the world model at a first rate using the two first cameras; updating the world model at a second rate slower than the first rate using the two first cameras and the color camera; Item 1. The wearable display system of item 1, configured to perform the following: (Item 6) the two first cameras are mechanically coupled to the headset to provide a central field of view associated with both first cameras and a peripheral field of view associated with a first of the two first cameras; the processor is further configured to track hand movement within the central field of view using depth information determined from images acquired by the two first cameras. Item 1. The wearable display system of item 1. (Item 7) Tracking hand movement within the central visual field using depth information includes: selecting a point within the central visual field; stereoscopically determining depth information for the selected point using images acquired by the two first cameras; generating a depth map using the stereoscopically determined depth information; and Matching portions of the depth map to corresponding portions of a hand model including both shape and motion constraints; Item 7. The wearable display system of item 6, comprising: (Item 8) Item 7. The wearable display system of item 6, wherein the processor is further configured to track hand movement within the peripheral vision using one or more images obtained from the two first cameras by matching portions of the images to corresponding portions of a model of a hand that includes both shape and movement constraints. (Item 9) Item 1. The wearable display system according to item 1, wherein the headset is a lightweight headset weighing 30 to 300 grams. (Item 10) Item 10. The wearable display system of item 9, wherein the headset further comprises the processor. (Item 11) Item 10. The wearable display system of item 9, wherein the headset further comprises a battery pack. (Item 12) Item 1. The wearable display system of item 1, wherein the two first cameras are configured to obtain grayscale images. (Item 13) Item 1. The wearable display system of item 1, wherein the processor is mechanically coupled to the headset. (Item 14) Item 1. The wearable display system of item 1, wherein the headset comprises a display device mechanically coupled to the processor. (Item 15) Item 1. The wearable display system of item 1, wherein a local data processing module comprises the processor, the local data processing module is operably coupled to a display device via a communication link, and the headset comprises the display device. (Item 16) 1. A wearable display system, comprising: The frame and two first cameras, the two first cameras comprising: a central field of view associated with both cameras; a first peripheral field of view associated with a first of the two first cameras; two first cameras mechanically coupled to the frame to provide a color camera mechanically coupled to the frame to provide a color field of view overlapping the central field of view; a processor operably coupled to the two first cameras and the color camera, the processor comprising: tracking hand movements within the central field of view using stereoscopically determined depth information from first images acquired by the two first cameras; tracking hand movement within the first peripheral field of view using one or more second images acquired by a first of the two first cameras; creating a world model using the two first cameras and the color camera; updating the world model using the two first cameras; and a processor configured to: A wearable display system comprising: (Item 17) 1. A wearable display system, comprising: The frame and two first cameras mechanically coupled to the frame; a color camera mechanically coupled to the frame; a processor operably coupled to the two first cameras and the color camera, the processor comprising: creating a world model using first grayscale images acquired by the two first cameras and one or more color images acquired by the color camera; updating the world model using second grayscale images acquired by the two first cameras; determining portions of the world model that contain incomplete depth information; updating the world model with additional depth information for portions of the world model that include incomplete depth information; a processor configured to: A wearable display system comprising: (Item 18) 1. A wearable display system, comprising: The frame and two grayscale cameras mechanically coupled to the frame, the two grayscale cameras comprising a first grayscale camera having a first field of view and a second grayscale camera having a second field of view, the first grayscale camera and the second grayscale camera comprising: a central visual field where the first visual field overlaps with the second visual field; a first peripheral field of view within the first field of view and outside the second field of view; two grayscale cameras positioned to provide a color camera mechanically coupled to the frame to provide a color field of view overlapping the central field of view; A wearable display system comprising: [Brief explanation of the drawings]

[0022] The accompanying drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component illustrated in various figures is represented by a like numeral. For purposes of clarity, not every component is labeled in every drawing.

[0023] [Figure 1] FIG. 1 is a sketch illustrating an example of a simplified augmented reality (AR) scene, according to some embodiments.

[0024] [Figure 2] FIG. 2 is a schematic diagram illustrating an example of an AR display system, according to some embodiments.

[0025] [Figure 3A] FIG. 3A is a schematic diagram illustrating a user wearing an AR display system that renders AR content as the user moves through a physical world environment, according to some embodiments.

[0026] [Figure 3B] FIG. 3B is a schematic diagram illustrating a viewing optics assembly and associated components, according to some embodiments.

[0027] [Figure 4] FIG. 4 is a schematic diagram illustrating an image sensing system, according to some embodiments.

[0028] [Figure 5A] FIG. 5A is a schematic diagram illustrating the pixel cell in FIG. 4 according to some embodiments.

[0029] [Figure 5B] FIG. 5B is a schematic diagram illustrating output events of the pixel cell of FIG. 5A, according to some embodiments.

[0030] [Figure 6] FIG. 6 is a schematic diagram illustrating an image sensor, according to some embodiments.

[0031] [Figure 7] FIG. 7 is a schematic diagram illustrating an image sensor, according to some embodiments.

[0032] [Figure 8] FIG. 8 is a schematic diagram illustrating an image sensor, according to some embodiments.

[0033] [Figure 9] FIG. 9 is a simplified flowchart of a method for image sensing, according to some embodiments.

[0034] [Figure 10] FIG. 10 is a simplified flowchart of the patch identification acts of FIG. 9 according to some embodiments.

[0035] [Figure 11] FIG. 11 is a simplified flowchart of the patch trajectory estimation acts of FIG. 9 according to some embodiments.

[0036] [Figure 12] FIG. 12 is a schematic diagram illustrating the patch trajectory estimation of FIG. 11 for one viewpoint, according to some embodiments.

[0037] [Figure 13] FIG. 13 is a schematic diagram illustrating the patch trajectory estimation of FIG. 11 over viewpoint changes, according to some embodiments.

[0038] [Figure 14] FIG. 14 is a schematic diagram illustrating an image sensing system, according to some embodiments.

[0039] [Figure 15] FIG. 15 is a schematic diagram illustrating the pixel cell in FIG. 14 according to some embodiments.

[0040] [Figure 16] FIG. 16 is a schematic diagram of a pixel sub-array, according to some embodiments.

[0041] [Figure 17A] FIG. 17A is a cross-sectional view of a plenoptic device with an angle-of-arrival / intensity converter in the form of two aligned stacked transmissive diffractive masks (TDMs) according to some embodiments.

[0042] [Figure 17B] FIG. 17B is a cross-sectional view of a plenoptic device with an angle-of-arrival / intensity converter in the form of two misaligned stacked TDMs, according to some embodiments.

[0043] [Figure 18A]FIG. 18A is a pixel sub-array with color pixel cells and arrival corner pixel cells according to some embodiments.

[0044] [Figure 18B] FIG. 18B is a pixel sub-array with color pixel cells and arrival corner pixel cells according to some embodiments.

[0045] [Figure 18C] FIG. 18C is a pixel sub-array with white pixel cells and corner-of-arrival pixel cells according to some embodiments.

[0046] [Figure 19A] FIG. 19A is a top view of a photodetector array with a single TDM according to some embodiments.

[0047] [Figure 19B] FIG. 19B is a side view of a photodetector array with a single TDM according to some embodiments.

[0048] [Figure 20A] FIG. 20A is a top view of a photodetector array with multiple angle-of-arrival / intensity converters in the form of a TDM, according to some embodiments.

[0049] [Figure 20B] FIG. 20B is a side view of a photodetector array with multiple TDMs according to some embodiments.

[0050] [Figure 20C] FIG. 20C is a side view of a photodetector array with multiple TDMs according to some embodiments.

[0051] [Figure 21] FIG. 21 is a schematic diagram of a headset including three cameras and associated components, according to some embodiments.

[0052] [Figure 22] FIG. 22 is a simplified flowchart of a calibration routine, according to some embodiments.

[0053] [Figure 23A] 23A-23C are exemplary field of view diagrams associated with the headset of FIG. 21, according to some embodiments. [Figure 23B] 23A-23C are exemplary field of view diagrams associated with the headset of FIG. 21, according to some embodiments. [Figure 23C] 23A-23C are exemplary field of view diagrams associated with the headset of FIG. 21, according to some embodiments.

[0054] [Figure 24] FIG. 24 is a simplified flowchart of a method for updating a world model, according to some embodiments.

[0055] [Figure 25] FIG. 25 is a simplified flowchart of a process for hand tracking, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0056] The inventors have recognized and appreciated design and operational techniques for wearable XR display systems that improve the enjoyment and usability of such systems. These design and / or operational techniques may enable the use of a limited number of cameras to obtain information for performing multiple functions, including hand tracking, head pose tracking, and world reconstruction, which may be used to realistically render virtual objects so that they appear to interact realistically with physical objects. Wearable cross-reality display systems may be lightweight and consume low power during operation. The systems may use a specific configuration of sensors to obtain image information about physical objects in the physical world with low latency. The systems may implement various routines to improve the accuracy and / or realism of the displayed XR environment. Such routines may include calibration routines to improve the accuracy of stereoscopic depth measurements, even when a lightweight frame distorts during use, and routines to detect and address incomplete depth information in a model of the physical world around the user.

[0057] The weight of known XR system headsets can limit user enjoyment. Such XR headsets can weigh more than 340 grams (sometimes even more than 700 grams). Glasses, in contrast, can weigh less than 50 grams. Wearing such relatively heavy headsets for extended periods of time can tire users or distract them, diverting their attention from the desired immersive XR experience. However, the inventors recognize and understand that some designs that reduce headset weight also increase headset flexibility, making lightweight headsets more susceptible to changes in sensor position or orientation during use or over time. For example, when a user wears a lightweight headset that includes camera sensors, the relative orientation of these camera sensors may shift. Variations in the spacing of cameras used for stereoscopic imaging can affect the ability of those headsets to obtain accurate stereoscopic information, which relies on cameras having known positional relationships with each other. Thus, a calibration routine, which can be repeated when the headset is worn, can enable a lightweight headset that can use stereoscopic imaging techniques to accurately obtain information about the world around the headset wearer.

[0058] The need to equip XR systems with components to obtain information about objects in the physical world can also limit the usefulness and user enjoyment of these systems. Although the obtained information is used to realistically present computer-generated virtual objects in appropriate locations and with appropriate appearance relative to physical objects, the need to obtain the information imposes limitations on the size, power consumption, and realism of the XR system.

[0059] An XR system may, for example, use sensors worn by a user to obtain information about objects in the physical world around a user, including information about the position of physical world objects in the user's field of view. Challenges arise because physical objects may move relative to a user's field of view either as a result of the object entering or leaving the user's field of view, or as a result of the object moving in the physical world so that the position of the physical object in the user's field of view changes, or as a result of the user changing its orientation relative to the physical world. To present a realistic XR display, models of physical objects in the physical world must be updated frequently enough to capture these changes, processed with sufficiently low latency, and accurately predicted into the future to cover the full latency path, including rendering, so that virtual objects displayed based on that information will have the appropriate position and appearance relative to the physical objects as they are displayed. Otherwise, virtual objects will appear misaligned with the physical objects, and combined scenes including physical and virtual objects will not appear realistic. For example, virtual objects may appear as if they are floating in space or bouncing around relative to the physical objects rather than resting on them. Visual tracking errors are amplified when there is significant evolution in the scene, especially when the user is moving at high speed.

[0060] Such problems can be avoided by sensors that obtain new data at a high rate. However, the power consumed by such sensors can lead to the need for larger batteries, increase the weight of the system, or limit the length of use of such systems. Similarly, the processors required to process data generated at a high rate can drain the battery and add additional weight to a wearable system, further limiting the usefulness or enjoyability of such systems. A known approach is to operate a sensor at a higher frame rate for increased temporal resolution, for example, with a higher resolution to capture sufficient visual detail. An alternative solution could complement this solution with an IR time-of-flight sensor that can directly indicate the position of a physical object relative to the sensor, and simple processing, resulting in low latency, can be performed using this information to display virtual objects. However, such sensors consume substantial amounts of power, especially when they operate in sunlight.

[0061] The inventors have recognized and appreciated that an XR system can address changes in sensor position or orientation during use or over time by repeatedly performing a calibration routine. This calibration routine can determine the current relative separation and orientation of sensors contained within a headset. The wearable XR system can then consider this relative separation and orientation of the headset sensors when calculating stereoscopic depth information. With such calibration capabilities, the XR system can accurately obtain depth information indicating distance to objects in the physical world without using active depth sensors or with only occasional use of active depth sensing. Because active depth sensing can consume substantial power, reducing or eliminating active depth sensing enables the device to draw less power, which can increase the device's operating time without recharging the battery or reduce the size of the device as a result of reducing the battery size.

[0062] The inventors have also recognized and appreciated that with the appropriate combination of image sensors and appropriate techniques for processing image information from those sensors, an XR system may obtain information about physical objects with low latency, even at reduced power consumption, by reducing the number of sensors used, eliminating, disabling, or selectively activating resource-intensive sensors, and / or reducing overall sensor usage. As a specific example, an XR system may include a headset with two world cameras and a color camera. The world cameras may produce grayscale images and may have a global shutter. These grayscale images may be smaller in size than color images of similar resolution. The world cameras may require less power than color cameras of similar resolution. Information from these cameras may be used at different times and in different ways to support the operation of the XR system.

[0063] The techniques described herein may be used with many types of devices or separately and for many types of scenarios. Figure 1 illustrates such a scenario. Figures 2, 3A, and 3B illustrate example AR systems including one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.

[0064] Referring to FIG. 1 , an AR scene 4 is depicted in which a user of the AR system sees a physical-world park-like setting 6 featuring people, trees, a building in the background, and a concrete platform 8. In addition to these physical objects, a user of the AR technology also perceives as "seeing" a virtual object, here illustrated as a robotic figure 10 standing on the physical-world concrete platform 8, and a flying, cartoon-like avatar character 2 that appears to be an anthropomorphic bumblebee, although these elements (e.g., avatar character 2 and robotic figure 10) do not exist within the physical world. Due to the significant complexity of human visual perception and the nervous system, it is difficult to produce an AR system that facilitates a comfortable, natural-feeling, and rich presentation of virtual image elements among other virtual or physical-world image elements.

[0065] Such scenes may be presented to a user by presenting image information representing the actual environment around the user and overlaying it with information representing virtual objects that are not in the actual environment. In an AR system, a user may be able to see objects in the physical world, and the AR system provides information that renders the virtual objects so that they appear in the appropriate locations and with appropriate visual characteristics so that the virtual objects appear to coexist with objects in the physical world. In an AR system, for example, a user may look through a transparent screen so that the user can see objects in the physical world. The AR system may render virtual objects on the screen so that the user can see both the physical world and the virtual objects. In some embodiments, the screen may be worn by the user like a pair of goggles or glasses.

[0066] The scene may be presented to the user via a system including multiple components, including a user interface, that may stimulate one or more of the user's senses, including sight, hearing, and / or touch. In addition, the system may include one or more sensors that may measure parameters of the physical portion of the scene, including the user's position and / or movement within the physical portion of the scene. Furthermore, the system may include one or more computing devices along with associated computer hardware such as memory. These components may be integrated into a single device or even distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated into a wearable device.

[0067] In some embodiments, the AR experience may be provided to a user through a wearable display system. FIG. 2 illustrates an example of a wearable display system 80 (hereinafter referred to as “system 80”). System 80 includes a head-mounted display device 62 (hereinafter referred to as “display device 62”) and various mechanical and electronic modules and systems for supporting the functionality of display device 62. Display device 62 may be coupled to a frame 64, which is wearable by a user or viewer 60 of the display system (hereinafter referred to as “user 60”) and configured to position display device 62 directly in front of the user's eyes 60. According to various embodiments, display device 62 may be a sequential display. Display device 62 may be monocular or binocular.

[0068] In some embodiments, a speaker 66 is coupled to the frame 64 and positioned proximate to the ear canal of the user 60. In some embodiments, another speaker, not shown, is positioned adjacent to another ear canal of the user 60 to provide stereo / adjustable sound control.

[0069] System 80 may include a local data processing module 70. Local data processing module 70 may be operatively coupled to display device 62 through a communications link 68, such as by wired leads or wireless connectivity. Local data processing module 70 may be mounted in a variety of configurations, such as fixedly attached to frame 64, fixedly attached to a helmet or hat worn by user 60, built into headphones, or otherwise removably attached to user 60 (e.g., in a backpack-style configuration, in a belt-coupled configuration). In some embodiments, local data processing module 70 may not be present, as components of local data processing module 70 may be integrated within display device 62 or implemented within a remote server or other component to which display device 62 is coupled, such as through wireless communication over a wide area network.

[0070] The local data processing module 70 may include a processor and digital memory, such as non-volatile memory (e.g., flash memory), both of which may be utilized to assist in processing, caching, and storing data. The data may include a) data captured from sensors (e.g., operably coupled to the frame 64 or otherwise attached to the user 60, such as image capture devices (e.g., cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, and / or gyroscopes) and / or b) data obtained and / or processed using the remote processing module 72 and / or the remote data repository 74, possibly for passage to the display device 62 after processing or retrieval. The local data processing module 70 may be operably coupled to the remote processing module 72 and the remote data repository 74 by communication links 76, 78, such as via wired or wireless communication links, respectively, such that these remote modules 72, 74 are operably coupled to each other and available as resources to the local processing and data module 70.

[0071] In some embodiments, local data processing module 70 may include one or more processors (e.g., a central processing unit and / or one or more graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, remote data repository 74 may include a digital data storage facility, which may be available through the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in local data processing module 70, allowing for fully autonomous use from remote modules.

[0072] In some embodiments, local data processing module 70 is operably coupled to battery 82. In some embodiments, battery 82 is a removable power source, such as a commercially available battery. In other embodiments, battery 82 is a lithium-ion battery. In some embodiments, battery 82 includes both an internal lithium-ion battery that is rechargeable by user 60 during periods of non-operation of system 80, and a removable battery so that user 60 can operate system 80 for longer periods of time without having to plug in a power source and charge the lithium-ion battery or shut off system 80 and replace the battery.

[0073] FIG. 3A illustrates a user 30 wearing an AR display system that renders AR content as the user 30 moves through a physical-world environment 32 (hereinafter referred to as “environment 32”). The user 30 positions the AR display system at a location 34, and the AR display system records ambient information of the passable world (e.g., digital representations of real objects in the physical world that can be stored and updated as the real objects in the physical world change) for the location 34. Each location 34 may further be associated with a “pose” and / or mapped features or directional audio inputs related to the environment 32. A user wearing the AR display system on their head may look in a particular direction and tilt their head to create a head pose of the system relative to the environment. At each location and / or pose within the same location, sensors on the AR display system may capture different information about the environment 32. Thus, information collected at the location 34 may be aggregated into a data input 36 and processed by at least a passable world module 38, which may be implemented, for example, by processing on the remote processing module 72 of FIG. 2 .

[0074] The passable world module 38 determines, at least in part, where and how the AR content 40 can be placed relative to the physical world, as determined from the data input 36. The AR content is "placed" in the physical world by presenting the AR content in a manner that allows the user to see both the AR content and the physical world. Such an interface may be created, for example, using glasses that the user can see through and view the physical world, and that can be controlled so that virtual objects appear at controlled locations within the user's field of view. The AR content is rendered as if interacting with objects in the physical world. The user interface is such that the user's view of objects in the physical world can be obscured, when appropriate, to create the appearance that the AR content obscures the user's view of those objects. For example, the AR content may be placed by appropriately selecting a portion of an element 42 (e.g., a table) in the environment 32 and displaying and displaying the AR content 40 shaped and positioned as if resting on or otherwise interacting with the element 42. AR content may also be placed within structures not already within the field of view 44 or relative to a mapped mesh model 46 of the physical world.

[0075] As depicted, elements 42 are examples of what may be multiple elements in the physical world that may be treated as if fixed and stored within passable world module 38. Once stored within passable world module 38, information about those fixed elements may be used to present information to user 30 so that user 30 may perceive content on fixed elements 42 without the system having to map it to the fixed elements 42 each time user 30 sees it. Fixed elements 42 may thus be mesh models mapped from a previous modeling session or determined from a separate user, but stored on passable world module 38 for future reference by multiple users. Thus, passable world module 38 may recognize environment 32 from a previously mapped environment and display AR content without user 30's device having to map environment 32 first, saving computational processes and cycles and avoiding latency for any rendered AR content.

[0076] Similarly, a mapped mesh model 46 of the physical world can be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying the AR content 40 can be mapped and stored in the passable world module 38 for future retrieval by the user 30 or other users without the need to remap or model. In some embodiments, the data inputs 36 are inputs such as geographic location, user identification, and current activity to indicate to the passable world module 38 which fixed elements 42 of one or more fixed elements are available, which AR content 40 was last placed on the fixed elements 42, and whether that same content should be displayed (such AR content is “persistent” content regardless of whether the user is viewing a particular passable world model).

[0077] Even in embodiments in which objects are considered to be fixed, the passable world module 38 may be updated from time to time to account for possible changes in the physical world. Models of fixed objects may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered fixed. To render the AR scene with a sense of realism, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update fixed objects. To enable accurate tracking of all of the objects in the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.

[0078] FIG. 3B is a schematic diagram of the viewing optics assembly 48 and associated optional components. A specific configuration is described in FIG. 21 below. When oriented toward the user's eye 49, in some embodiments, two eye-tracking cameras 50 detect metrics of the user's eye 49, such as eye shape, eyelid occlusion, pupil direction, and phosphenes on the user's eye 49. In some embodiments, one of the sensors is a depth sensor 51, such as a time-of-flight sensor, that emits signals into the world, detects reflections of those signals from nearby objects, and determines the distance to a given object. The depth sensor can quickly determine, for example, whether objects have entered the user's field of view, either as a result of the objects' motion or a change in the user's posture. However, information about the location of objects within the user's field of view may alternatively or additionally be collected using other sensors. In some embodiments, the world camera 52 records a view larger than the peripheral field of view, maps the environment 32, and detects inputs that may affect the AR content. In some embodiments, world camera 52 and / or camera 53 may be grayscale and / or color image sensors that output grayscale and / or color image frames at fixed time intervals. Camera 53 may also capture physical world images within the user's field of view at specific times. Pixels of a frame-based image sensor may be repeatedly sampled even if their values ​​remain constant. World camera 52, camera 53, and depth sensor 51 each have separate fields of view 54, 55, and 56, and collect and record data from a physical world scene, such as physical world environment 32 depicted in FIG. 3A.

[0079] The inertial measurement unit 57 may determine the movement and / or orientation of the viewing optics assembly 48. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 51 may be operably coupled to the eye tracking camera 50 to ascertain the actual distance of points and / or areas in the physical world that the user's eyes 49 are viewing.

[0080] It should be understood that viewing optics assembly 48 may include several of the components illustrated in FIG. 3B . For example, viewing optics assembly 48 may include a different number of components. In some embodiments, for example, viewing optics assembly 48 may include one world camera 52, two world cameras 52, or more world cameras instead of the four world cameras depicted. Alternatively, or in addition, cameras 52 and 53 need not capture visible light images of their full field of view. Viewing optics assembly 48 may include other types of components. In some embodiments, viewing optics assembly 48 may include one or more dynamic vision sensors (DVS), the pixels of which may respond asynchronously to relative changes in light intensity above a threshold.

[0081] In some embodiments, viewing optics assembly 48 may not include a depth sensor 51 based on time-of-flight information. In some embodiments, for example, viewing optics assembly 48 may include one or more plenoptic cameras, the pixels of which may capture not only light intensity but also the angle of incident light. For example, a plenoptic camera may include an image sensor overlaid with a transmissive diffractive mask (TDM). Alternatively, or in addition, a plenoptic camera may include an image sensor containing angle-sensing pixels and / or phase-detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such sensors may serve as a source of depth information instead of, or in addition to, depth sensor 51.

[0082] 3B is illustrated as an example. Viewing optics assembly 48 may include components in any suitable configuration so that a user can have the maximum field of view for a particular set of components. For example, if viewing optics assembly 48 has one world camera 52, the world camera may be located within a central region of the viewing optics assembly instead of on a side.

[0083] Information from these sensors in the viewing optics assembly 48 may be coupled to one or more processors in the system. The processor may generate data that can be rendered to allow the user to perceive virtual content interacting with objects in the physical world. The rendering may be implemented in any suitable manner, including generating image data depicting both physical and virtual objects. In other embodiments, physical and virtual content may be depicted in a single scene by modulating the opacity of a display device through which the user views the physical world. The opacity may be controlled to create the appearance of virtual objects and also to block the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, the image data may include only the virtual content that can be viewed through a user interface and may be modified to interact realistically with the physical world (e.g., clipping content and accounting for occlusion). Regardless of how the content is presented to the user, a model of the physical world may be used so that properties of the virtual objects that can be affected by physical objects, including their shape, position, motion, and visibility, can be correctly calculated.

[0084] A model of the physical world may be created from data collected from sensors on a user's wearable device. In some embodiments, the model may be created from data collected by multiple users, which may be aggregated in a computing device remote from all of the users (and which may be "in the cloud").

[0085] In some embodiments, at least one of the sensors may be configured to obtain information about physical objects in a scene, particularly non-stationary objects, at high frequency with low latency using compact and low-power components, and the sensor may employ patch tracking to limit the amount of data output.

[0086] 4 depicts an image sensing system 400 according to some embodiments. The image sensing system 400 may include an image sensor 402, which may include an image array 404, which may contain a plurality of pixels, each responsive to light, as in a conventional image sensor. The sensor 402 may further include circuitry for accessing each pixel. Accessing a pixel may involve obtaining information about incident light generated by that pixel. Alternatively, or in addition, accessing a pixel may involve controlling that pixel, such as by configuring it to only provide an output in response to the detection of an event.

[0087] In the illustrated embodiment, the image array 404 is configured as an array with multiple rows and columns of pixels. In such an embodiment, the access circuitry may be implemented as a row address encoder / decoder 406 and a column address encoder / decoder 408. The image sensor 402 may further contain circuitry that generates inputs to the access circuitry and controls the timing and order in which information is read from the pixels in the image array 404. In the illustrated embodiment, that circuitry is a patch tracking engine 410. In contrast to conventional image sensors that may continuously output image information captured by pixels in each row, the image sensor 402 may be controlled to output image information in defined patches. Furthermore, the location of those patches relative to the image array may change over time. In the illustrated embodiment, the patch tracking engine 410 outputs image array access information to control the output of image information from portions of the image array 404 that correspond to the location of the patches, and the access information may change dynamically based on estimates of the movement of objects in the environment and / or the movement of the image sensor relative to those objects.

[0088] In some embodiments, the image sensor 402 may have dynamic vision sensor (DVS) functionality, such that image information is provided by the sensor only when there is a change in an image property (e.g., intensity) for a pixel. For example, the image sensor 402 may apply one or more thresholds that define the on and off states of pixels. The image sensor may detect that a pixel has changed state and selectively provide an output for only those pixels, or only those pixels in a patch, that have changed state. These outputs may occur asynchronously as they are detected, rather than as part of a read of all pixels in the array. The output may be in the form of an address-event representation (AER) 418, which may include, for example, a pixel address (e.g., row and column) and a type of event (on or off). An on event may indicate that a pixel cell at a particular pixel address senses an increase in light intensity, and an off event may indicate that a pixel cell at a particular pixel address senses a decrease in light intensity. The increase or decrease may be relative to an absolute level or may be a change relative to the level in the last output from the pixel. The change may be expressed, for example, as a fixed offset or as a percentage of the value in the final output from the pixel.

[0089] The use of DVS techniques in conjunction with patch tracking may enable an image sensor suitable for use in an XR system. When combined in an image sensor, the amount of data generated may be limited to data from pixel cells within the patch that detect a change that will trigger the output of an event.

[0090] In some scenarios, high-resolution image information is desirable. However, to generate high-resolution image information, large sensors with cells exceeding a million pixels may generate a large amount of image information when DVS techniques are used. The inventors recognize and understand that DVS sensors may produce a large number of events that reflect changes in the image other than those resulting from movement in the background or the movement of the object being tracked. Currently, the resolution of DVS sensors is limited to below 1 MB, e.g., 128 x 128, 240 x 180, and 346 x 260, to limit the number of events generated. Such sensors sacrifice resolution for tracking objects and may not detect, for example, the fine movements of fingers on a hand. Furthermore, if the image sensor outputs image information in other formats, limiting the resolution of the sensor array to output a manageable number of events may also limit the use of the image sensor to generate high-resolution image frames with DVS functionality. Sensors as described herein may have resolutions higher than VGA, including up to 8 or 12 megapixels, in some embodiments. Nevertheless, patch tracking as described herein may be used to limit the number of events output by the image sensor per second. As a result, an image sensor may be enabled that operates in at least two modes. For example, an image sensor with megapixel resolution may operate in a first mode that outputs events within the specific patch being tracked. In a second mode, it may output a high-resolution image frame or a portion of an image frame. Such an image sensor may be controlled within an XR system to operate in these different modes based on the capabilities of the system.

[0091] The imaging array 404 may include a plurality of pixel cells 500 arranged in the array. FIG. 5A depicts an example of a pixel cell 500, which, in this embodiment, is configured for use in an imaging array that implements DVS techniques. The pixel cell 500 may include a photosensitive circuit 502, a difference circuit 506, and a comparator 508. The photosensitive circuit 502 may include a photodiode 504 that converts light striking the photodiode into a measurable electrical signal. In this example, the conversion is to an electrical current I. A transconductance amplifier 510 converts the photocurrent I to a voltage. The conversion may be linear or nonlinear, such as according to a function of log I. Regardless of the specific transfer function, the output of the transconductance amplifier 510 indicates the amount of light detected at the photodiode 504. While a photodiode is illustrated as an example, it should be understood that other light-sensing components that produce a measurable output in response to incident light may be implemented in the photosensitive circuit in place of, or in addition to, the photodiode.

[0092] In the embodiment of FIG. 5A, circuitry is built into the pixel itself to determine whether the pixel's output has changed sufficiently and trigger an output for that pixel cell. In this example, that function is implemented by a difference circuit 506 and a comparator 508. The difference circuit 506 may be configured to reduce DC mismatch between pixel cells, for example, by balancing the difference circuit's output and resetting the level after generating an event. In this example, the difference circuit 506 is configured to produce an output indicative of the change in the output of the photodiode 504 since the last output. The difference circuit may include an amplifier 512 having a gain -A, a capacitor 514, which may be implemented as a single circuit element or one or more capacitors connected in a network, and a reset switch 516.

[0093] In operation, the pixel cell may be reset by momentarily closing switch 516. Such a reset may occur at the beginning of circuit operation and any time thereafter after an event is detected. When pixel 500 is reset, the voltage across capacitor 514, when subtracted from the output of transconductance amplifier 510, will result in zero voltage being present at the input of amplifier 512. When switch 516 is open, the output of transconductance amplifier 510, combined with the voltage drop across capacitor 514, will cause zero voltage to be present at the input of amplifier 512. The output of transconductance amplifier 510 changes as a result of changes in the amount of light striking photodiode 504. As the output of transconductance amplifier 510 increases or decreases, the output of amplifier 512 will swing positive or negative by an amount amplified by the gain of amplifier 512.

[0094] The comparator 508 may determine whether an event is generated and the sign of the event, for example, by comparing the output voltage V of the difference circuit to a predetermined threshold voltage C. In some embodiments, the comparator 508 may include two comparators comprising transistors; one pair may operate when the output of the amplifier 512 indicates a positive change and may detect an increasing change (an ON event), and the other comparator may operate when the output of the amplifier 512 indicates a negative change and may detect a decreasing change (an OFF event). However, it should be understood that the amplifier 512 may have a negative gain. In such an embodiment, an increase in the output of the transconductance amplifier 510 may be detected as a negative voltage change at the output of the amplifier 512. Similarly, it should be understood that the positive and negative voltages may be relative to ground or any suitable reference level. Regardless, the value of the threshold voltage C may be controlled by transistor characteristics (e.g., transistor size, transistor threshold voltage) and / or by the value of a reference voltage that may be applied to the comparator 508.

[0095] FIG. 5B depicts an example of an event output (on, off) over time t of pixel cell 500. In the illustrated example, at time t1, the output of the difference circuit has a value of V1, at time t2, the output of the difference circuit has a value of V2, and at time t3, the output of the difference circuit has a value of V3. Between time t1 and t2, the photodiode senses a slight increase in light intensity, but the pixel cell does not output an event because the change in V does not exceed the value of threshold voltage C. At time t2, the pixel cell outputs an on event because V2 is greater than V1 by the value of threshold voltage C. Between time t2 and t3, the photodiode senses a slight decrease in light intensity, but the change in V does not exceed the value of threshold voltage C, so the pixel cell does not output an event. At time t3, V3 is less than V2 by the value of threshold voltage C, so the pixel cell outputs an off event.

[0096] Each event may trigger an output in the AER 418. The output may include, for example, an indication of whether the event is an on or off event and the identity of the pixel, such as its row and column. Other information may alternatively or additionally be included with the output. For example, a timestamp may be included, which may be useful if the event is queued for later transmission or processing. As another example, the current level at the output of the amplifier 510 may be included. Such information may optionally be included, for example, if further processing is to be performed in addition to detecting object motion.

[0097] It should be understood that the frequency of the event output, and therefore the sensitivity of the pixel cell, may be controlled by the value of the threshold voltage C. For example, the frequency of the event output may be reduced by increasing the value of the threshold voltage C, or increased by decreasing the threshold voltage C. It should also be understood that the threshold voltage C may be different for on-events and off-events, for example, by setting different reference voltages for the comparator for detecting on-events and the comparator for detecting off-events. It should also be understood that the pixel cell may also output a value indicative of the size of the light intensity change instead of, or in addition to, a sign signal indicating the detection of an event.

[0098] The pixel cell 500 of Figures 5A and 5B is illustrated as an example, according to some embodiments. Other designs may also be suitable for the pixel cell. In some embodiments, the pixel cell includes a light-sensing circuit and a difference circuit, but may share a comparator circuit with one or more other pixel cells. In some embodiments, the pixel cell may include circuitry, such as an active pixel sensor at the pixel level, configured to calculate a change value.

[0099] The ability to configure pixels to output only in response to the detection of an event, regardless of the manner in which the event is detected for each pixel cell, may be used to limit the amount of information required to maintain a model of the location of non-stationary (i.e., movable) objects. For example, pixels within a patch may have a threshold voltage C set to trigger when a relatively small change occurs. Other pixels outside the patch may have a larger threshold, such as three or five times larger. In some embodiments, the threshold voltage C for any out-of-patch pixel may be set large so that the pixel is effectively disabled and does not produce any output, regardless of the amount of change. In other embodiments, out-of-patch pixels may be disabled in other ways. In such embodiments, the threshold voltage may be fixed for all pixels, but pixels may be selectively enabled or disabled based on whether they are within the patch.

[0100] In yet other embodiments, the threshold voltage for one or more pixels may be adaptively set to modulate the amount of data output from the image array. For example, an AR system may have the processing capacity to process a certain number of events per second. The threshold for some or all pixels may be increased when the number of events per second being output exceeds an upper limit. Alternatively, or in addition, the threshold may be lowered when the number of events per second drops below a lower limit, enabling more data for more accurate processing. The number of events per second may be 200 to 2,000 events, as a specific example. Such a number of events constitutes a substantial reduction in the number of data to be processed per second compared to processing all of the pixel values ​​scanned out of an image sensor, which may constitute, for example, 30 million or more pixel values ​​per second. The number of events is a further reduction compared to processing only pixels within a patch, which, while smaller, may still generate tens of thousands of pixel values ​​or more per second.

[0101] The control signals for enabling and / or setting the threshold voltages for each of the pixels may be generated in any suitable manner, however, in the illustrated embodiment, the control signals are set by patch tracking engine 410 or based on processing within processing module 72 or other processor.

[0102] Referring back to FIG. 4 , image sensing system 400 may receive input from any suitable component such that patch tracking engine 410 may dynamically select at least one region of image array 404 to be enabled and / or disabled to implement a patch based at least on the received input. Patch tracking engine 410 may be digital processing circuitry having memory that stores one or more parameters of the patch. The parameters may be, for example, the boundary of the patch and may include other information, such as information about a scale factor between movement of the image array and movement within the image array of an image of a movable object associated with the patch. Patch tracking engine 410 may also include circuitry configured to perform calculations on the stored values ​​and other measured values ​​provided as input.

[0103] In the illustrated embodiment, the patch tracking engine 410 receives as input a designation of a current patch. The patch may be designated based on its size and location within the image array 404, such as by defining the range of row and column addresses for the patch. Such a designation may be provided as an output of the processing module 72 ( FIG. 2 ) or other component that processes information about the physical world. The processing module 72 may designate a patch to encompass, for example, the current location of each movable object or a subset of tracked movable objects in the physical world in order to render virtual objects with an appropriate appearance relative to the physical world. For example, if an AR scene is to include, as a virtual object, a toy figure balanced on a physical object such as a moving toy car, a patch may be designated to encompass the toy car. A patch may not be designated for another toy car moving in the background, since there may be little need to have up-to-date information about that object to render a realistic AR scene.

[0104] Regardless of how a patch is selected, information about the patch's current location may be provided to the patch tracking engine 410. In some embodiments, the patch may be rectangular, so that the patch's location may simply be defined as a start and end row and column. In other embodiments, the patch may have other shapes, such as a circle, and the patch may be defined in other ways, such as by a center point and a radius.

[0105] In some embodiments, trajectory information may also be provided for the patch. The trajectory may, for example, define the movement of the patch relative to the coordinates of the image array 404. The processing module 72 may, for example, build a model of the movement of a movable object in the physical world and / or the movement of the image array 404 relative to the physical world. As either or both motions may affect the location in the image array 404 onto which the image of the object is projected, the trajectory of the patch within the image array 404 may be calculated based on either or both. The trajectory may be defined in any suitable manner, such as by parameters of a linear, quadratic, cubic, or other polynomial equation.

[0106] In other embodiments, the patch tracking engine 410 may dynamically calculate the location of the patch based on input from sensors that provide information about the physical world. The information from the sensors may be provided directly from the sensors. Alternatively, or in addition, the sensor information may be processed to extract information about the physical world before being provided to the patch tracking engine 410. The extracted information may include, for example, the motion of the image array 404 relative to the physical world, the distance between the image array 404 and the object whose image falls within the patch, or other information that can be used to dynamically align the patch in the image array 404 with the image of the object in the physical world as the image array 404 and / or the object move.

[0107] Examples of input components may include an image sensor 412 and an inertial sensor 414. Examples of image sensors 412 may include an eye tracking camera 50, a depth sensor 51, a world camera 52, and / or a camera 52. An example of an inertial sensor 414 may include an inertial measurement unit 57. In some embodiments, the input components may be selected to provide data at a relatively high rate. The inertial measurement unit 57 may have an output rate of 200-2,000 measurements per second, such as 800-1,200 measurements per second. Patch locations may also be updated at a similarly high rate. By using the inertial measurement unit 57 as a source of input to the patch tracking engine 410, patch locations may be updated 800-1,200 times per second, as one specific example. In this manner, movable objects may be tracked with high accuracy using relatively small patches, limiting the number of events that need to be processed. Such an approach can lead to very low latency between changes in the relative position of the image sensor and the movable object, with similarly low latency in updating the rendering of the virtual object, providing a desirable user experience.

[0108] In some scenarios, the movable object being tracked using the patch may be a stationary object in the physical world. The AR system may, for example, identify the stationary object by analyzing multiple images captured from the physical world and select features of one or more of the stationary objects as reference points for determining the movement of a wearable device having an image sensor thereon. Frequent, low-latency updates of the locations of these reference points relative to the sensor array may be used to provide frequent, low-latency calculations of the head pose of a user of the wearable device. Frequent, low-latency updates of the head pose improve the user experience of the AR system because the head pose can be used to realistically render virtual objects via a user interface on the wearable. Thus, having input to the patch tracking engine 410, which controls the position of the patch, come solely from sensors with high output rates, such as one or more inertial measurement units, can lead to a desirable user experience of the AR system.

[0109] However, in some embodiments, other information may also be provided to the patch tracking engine 410, enabling it to calculate and / or apply a trajectory to the patch. This other information may include stored information 416, such as the passable world module 38 and / or the mapped mesh model 46. This information may indicate one or more previous positions of the object relative to the physical world, such that consideration of changes in these previous positions and / or changes in the current position relative to the previous position may indicate the trajectory of the object in the physical world, which may then be mapped to the trajectory of the patch across the image array 404. Other information in the model of the physical world may be used instead or in addition. For example, other information regarding the size of a movable object and / or its distance or position relative to the image array 404 may be used to calculate either the location or trajectory of the patch across the image array 404 associated with that object.

[0110] Regardless of the manner in which the trajectory is determined, the patch tracking engine 410 may apply the trajectory to calculate updated locations of patches within the image array 404 at a high rate, such as once per second or faster than 800 times per second. The rate may be limited by processing power, such as less than 2,000 times per second in some embodiments.

[0111] It should be understood that the process for tracking changes in movable objects may involve less than a complete reconstruction of the physical world. However, reconstructions of the physical world may occur at intervals longer than the interval between updates in the positions of movable objects, such as every 30 seconds or every 5 seconds. The locations of objects to be tracked and the locations of patches that will capture information about those objects may be recalculated when reconstructions of the physical world occur.

[0112] 4 illustrates an embodiment in which processing circuitry for both dynamically generating patches and controlling the selective output of image information from within the patches is configured to directly control the image array 404 so that the image information output from the array is limited to the selected information. Such circuitry may be integrated, for example, within the same semiconductor chip that stores the image array 404, or may be integrated into a separate controller chip for the image array 404. However, it should be understood that the circuitry generating the control signals for the image array 404 may be distributed throughout the XR system. For example, some or all of the functions may be performed by programming within the processing module 72 or other processors in the system.

[0113] The image sensing system 400 may output image information for each of a plurality of pixels. Each pixel of the image information may correspond to one of the pixel cells of the image array 404. The image information output from the image sensing system 400 may be image information for one or more patches corresponding to at least one region of the image array 404 selected by the patch tracking engine 410. In some embodiments, such as when each pixel of the image array 404 has a different configuration than that illustrated in FIG. 5A , the pixels in the output image information may identify pixels at which changes in light intensity are detected by the image sensor 400 in one or more patches.

[0114] In some embodiments, the image information output from the image sensing system 400 may be image information for pixels outside each of one or more patches corresponding to at least one region of the image array selected by the patch tracking engine 410. For example, a deer may be running in a physical world with a flowing river. Details of the river's waves may not be of interest but may trigger pixel cells in the image array 402. The patch tracking engine 410 may create a patch that encompasses the river and disable the portion of the image array 402 that corresponds to the patch that encompasses the river.

[0115] Based on the identification of the changed pixels, further processing may be performed. For example, the portion of the world model corresponding to the portion of the physical world imaged by the changed pixels may be updated. These updates may be performed based on information collected using other sensors. In some embodiments, further processing may be adjusted on or triggered by multiple changed pixels within a patch. For example, an update may be performed once 10% or some other threshold amount of pixels within a patch detect a change.

[0116] In some embodiments, image information in other formats may be output from the image sensor and may be used in combination with the change information to provide updates to the world model. In some embodiments, the format of the image information output from the image sensor may change from time to time during operation of the VR system. In some embodiments, for example, pixel cell 500 may be operated at times to produce a differential output such as that produced in comparator 508. The output of amplifier 510 may be switchable at other times to output the magnitude of light incident on photodiode 504. For example, the output of amplifier 510 may be switchably connected to a sense line, which in turn is connected to an A / D converter that may provide a digital indication of the magnitude of the incident light based on the magnitude of the output of amplifier 510.

[0117] In this configuration, the image sensor is operated as part of the AR system and may output differentially, often outputting only events related to pixels where a change above a threshold is detected, or only events related to pixels within a patch where a change above a threshold is detected. Periodically, such as every 5-30 seconds, a full image frame, carrying information about all pixels in the image array, may be output. Low latency and accurate processing may be achieved in this manner; differential information may be used to quickly update selected portions of the world model where changes most likely to affect user perception have occurred, while full images may be used to more frequently update larger portions of the world model. Full updates to the world model occur only at a slower rate, but any delay in updating the model may not meaningfully affect the user's perception of the AR scene.

[0118] The output mode of the image sensor may be changed at any time throughout operation of the image sensor such that the sensor outputs one or more of intensity information for some or all of the pixels and an indication of change for some or all of the pixels in the array.

[0119] It is not a requirement that image information from the patches be selectively output from the image sensor by limiting the information output from the image array. In some embodiments, image information may be output by all pixels in the image array, or only information about specific regions of the array may be output from the image sensor. FIG. 6 depicts an image sensor 600 according to some embodiments. The image sensor 600 may include an image array 602. In this embodiment, the image array 602 may resemble a conventional image array that scans out rows and columns of pixel values. The operation of such an image array may be accommodated by other components. The image sensor 600 may further include a patch tracking engine 604 and / or a comparator 606. The image sensor 600 may provide an output 610 to an image processor 608. The processor 608 may be, for example, part of the processing module 72 (FIG. 2).

[0120] The patch tracking engine 604 may have a structure and functionality similar to the patch tracking engine 410. It may be configured to receive signals defining at least one selected region of the image array 602 and then generate control signals that define a dynamic location of that region within the image array 602 of an image of an object represented by that region based on a calculated trajectory. In some embodiments, the patch tracking engine 604 may receive signals defining at least one selected region of the image array 602, which may include trajectory information for the region or regions. The patch tracking engine 604 may be configured to perform calculations that dynamically identify pixel cells within the at least one selected region based on the trajectory information. Variations in the implementation of the patch tracking engine 604 are also possible. For example, the patch tracking engine may update the location of a patch based on sensors indicating movement of the image array 602 and / or projected movement of an object associated with the patch.

[0121] In the embodiment illustrated in FIG. 6 , the image sensor 600 is configured to output difference information regarding pixels within an identified patch. The comparator 606 may be configured to receive a control signal from the patch tracking engine 604 that identifies a pixel within the patch. The comparator 606 may selectively act on pixels being output from the image array 602 having an address within the patch as indicated by the patch tracking engine 604. The comparator 606 may act on the pixel cells to generate a signal indicative of a change in sensed light detected by at least one region of the image array 602. As one example of implementation, the comparator 606 may contain memory elements that store reset values ​​for pixel cells within the array. As the current values ​​of those pixels are scanned out of the image array 602, circuitry within the comparator 606 may compare the stored values ​​with the current values ​​and output an indication if the difference exceeds a threshold. Digital circuitry may be used, for example, to store values ​​and perform such comparisons. In this embodiment, the output of image sensor 600 may be processed like the output of image sensor 400 .

[0122] In some embodiments, the image array 602, the patch tracking engine 604, and the comparator 606 may be implemented within a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, the image array 602 may be implemented within a single integrated circuit. The patch tracking engine 604 and the comparator 606 may be implemented within a second single integrated circuit, for example, configured as a driver for the image array 602. Alternatively, or in addition, some or all of the functionality of the patch tracking engine and / or the comparator 606 may be distributed to other digital processors within the AR system.

[0123] Other configurations or processing circuitry are also possible. Figure 7 depicts an image sensor 700 according to some embodiments. Image sensor 700 may include an image array 702. In this embodiment, image array 702 may have pixel cells with a differential configuration such as shown for pixel 500 in Figure 5A. However, embodiments herein are not limited to differential pixel cells, and patch tracking may be implemented using an image sensor that outputs intensity information.

[0124] 7, patch tracking engine 704 produces control signals indicating addresses of pixel cells within one or more patches being tracked. Patch tracking engine 704 may be constructed and operate similarly to patch tracking engine 604. Here, patch tracking engine 704 provides the control signals to pixel filter 706, which passes image information from only those pixels within the patch to output 710. As shown, output 710 is coupled to image processor 708, which may further process the image information for the pixels within the patch using techniques as described herein or in other suitable manners.

[0125] A further variation is illustrated in FIG. 8 , which depicts an image sensor 800 according to some embodiments. Image sensor 800 may include an image array 802, which may be a conventional image array that scans out intensity values ​​for pixels. The image array may be adapted to provide difference image information as described herein through the use of a comparator 806. Similar to comparator 606, comparator 806 may calculate difference information based on stored values ​​for pixels. Selected ones of those difference values ​​may be passed by a pixel filter 808 to an output 812. Similar to pixel filter 706, pixel filter 808 may receive a control input from a patch tracking engine 804. Patch tracking engine 804 may be similar to patch tracking engine 704. Output 812 may be coupled to an image processor 810. Some or all of the above-described components of image sensor 800 may be implemented within a single integrated circuit. Alternatively, the components may be distributed across one or more integrated circuits or other components.

[0126] Image sensors as described herein may be operated as part of an augmented reality system to maintain information about movable objects or other information about the physical world that is useful in realistically rendering images of virtual objects in combination with information about the physical environment. Figure 9 depicts a method 900 for image sensing, according to some embodiments.

[0127] At least a portion of method 900 may be implemented to operate an image sensor, including, for example, image sensor 400, 600, 700, or 800. Method 900 may begin with receiving image information from one or more inputs (act 902), including, for example, image sensor 412, inertial sensor 414, and stored information 416. Method 900 may include identifying one or more patches on an image output of the image sensing system based, at least in part, on the received information (act 904). An example of act 904 is illustrated in FIG. 10. In some embodiments, method 900 may include calculating a moving trajectory for the one or more patches (act 906). An example of act 906 is illustrated in FIG. 11.

[0128] Method 900 may also include configuring the image sensing system (act 908) based, at least in part, on the identified one or more patches and / or their estimated moving trajectories. The configuration may be achieved by enabling a portion of pixel cells of the image sensing system based, at least in part, on the identified one or more patches and / or their estimated moving trajectories, for example, through comparator 606, pixel filter 706, etc. In some embodiments, comparator 606 may receive a first reference voltage value for pixel cells corresponding to selected patches on the image and a second reference voltage value for pixel cells that do not correspond to any selected patches on the image. Comparator 606 may set the second reference voltage to be much higher than the first reference voltage such that reasonable light intensity changes sensed by pixel cells having comparator cells with the second reference voltage may not result in an output by the pixel cells. In some embodiments, pixel filter 706 may disable output from pixel cells with addresses (e.g., rows and columns) that do not correspond to any selected patches on the image.

[0129] 10 depicts patch identification 904, according to some embodiments. Patch identification 904 may include segmenting one or more images from one or more inputs based, at least in part, on color, light intensity, angle of arrival, depth, and semantics (act 1002).

[0130] Patch identification 904 may also include recognizing one or more objects in one or more images (act 1004). In some embodiments, object recognition 1004 may be based, at least in part, on predetermined features of the object, including, for example, hands, eyes, or facial features. In some embodiments, object recognition 1004 may be based, at least in part, on one or more virtual objects. For example, a virtual animal character is walking on a physical pencil. Object recognition 1004 may target the virtual animal character as an object. In some embodiments, object recognition 1004 may be based, at least in part, on artificial intelligence (AI) training received by an image sensing system. For example, an image sensing system may be trained by reading images of cats in different types and colors, and thus the learned characteristics of cats, enabling it to identify cats within the physical world.

[0131] Patch identification 904 may include generating a patch based on one or more objects (act 1006). In some embodiments, object patching 1006 may generate the patch by calculating a convex hull or bounding box for one or more objects.

[0132] 11 depicts patch trajectory estimation 906, according to some embodiments. Patch trajectory estimation 906 may include predicting movement over time for one or more patches (act 1102). Movement for one or more patches may occur for a number of reasons, including, for example, a moving object and / or a moving user. Motion prediction 1102 may include deriving a movement velocity for the moving object and / or the moving user based on received images and / or received AI training.

[0133] Patch trajectory estimation 906 may include calculating a trajectory over time for one or more patches (act 1104) based at least in part on the predicted movement. In some embodiments, the trajectory may be calculated by modeling a first-order linear equation, assuming that an object under motion will continue to move at the same speed and in the same direction. In some embodiments, the trajectory may be calculated by curve fitting or using heuristics, including pattern detection.

[0134] 12 and 13 illustrate coefficients that may be applied in calculating a patch trajectory. FIG. 12 depicts an example of a movable object, which in this example is a moving object 1202 (e.g., a hand) that moves relative to a user of the AR system. In this example, the user wears an image sensor as part of a head-mounted display 62. In this example, the user's eye 49 looks straight ahead so that image array 1200 captures the field of view (FOV) for eye 49 relative to one viewpoint 1204. Object 1202 is within the FOV and therefore appears by creating intensity variations in corresponding pixels in array 1200.

[0135] Array 1200 has a number of pixels 1208 arranged in the array. For a system tracking a hand 1202, a patch 1206 in the array that contains object 1202 at time t0 may include a portion of a number of pixels. If object 1202 is moving, the location of the patch capturing the object will change over time. That change can be captured in the patch trajectory from patch 1206 to patches X and Y used at a later time.

[0136] A patch trajectory may be estimated, such as in act 906, by identifying features 1210 associated with an object within the patch, e.g., a fingertip in the illustrated example. A motion vector 1212 may be calculated for the feature. In this example, the trajectory is modeled as a first-order linear equation, and the prediction is based on the assumption that the object 1202 continues on its same patch trajectory 1214 over time, leading to patch locations X and Y at each of two consecutive times.

[0137] As the patch location changes, the image of the moving object 1202 remains within the patch. Although the image information is limited to the information gathered using the pixels within the patch, the image information is sufficient to represent the movement of the moving object 1202. This would be true whether the image information is intensity information or difference information such as produced by a difference circuit. In the case of a difference circuit, for example, events indicating an increase in intensity may occur as the image of the moving object 1202 moves across the pixels. Conversely, as the image of the moving object 1202 passes from a pixel, an event indicating a decrease in intensity may occur. The pattern of pixels with increase and decrease events may be used as a reliable indication of the movement of the moving object 1202, which can be updated quickly with low latency due to the relatively small amount of data indicating the events. As a specific example, such a system may lead to a realistic XR system that tracks a user's hands and alters the rendering of a virtual object, creating the sensation for the user that the user is interacting with the virtual object.

[0138] The position of the patch may change for other reasons as well, any or all of which may be reflected in the trajectory calculation. One such other change is the movement of the user while wearing the image sensor. FIG. 13 depicts an example of a moving user creating a changing viewpoint for the user and the image sensor. In FIG. 13, the user may initially be looking straight ahead at an object with viewpoint 1302. In this configuration, pixel array 1300 of the image array will capture an object directly in front of the user. The object directly in front of the user may be within patch 1312.

[0139] The user may then change their viewpoint, such as by turning their head. The viewpoint may change to viewpoint 1304. An object that was previously directly in front of the user, even if it has not moved, will have a different position in the user's field of view at viewpoint 1304. It will also be at a different point in the field of view of the image sensor worn by the user, and therefore at a different position in image array 1300. The object may be contained within a patch, for example, at location 1314.

[0140] If the user further changes their viewpoint to viewpoint 1306 and the image sensor moves with the user, the location of the object that was previously directly in front of the user will be imaged at a different point in the field of view of the image sensor worn by the user, and therefore at a different position in image array 1300. The object may be contained within a patch at location 1316, for example.

[0141] As can be seen, as the user further changes their viewpoint, the location of the patch in the image array needed to capture the object moves further. The trajectory of this movement from location 1312 to location 1314 to location 1316 may be estimated and used to track the future position of the patch.

[0142] The trajectory may be estimated in other ways. For example, measurements using inertial sensors may indicate the acceleration and velocity of the user's head when the user has the gaze point 1302. This information may be used to predict the trajectory of the patches in the image array based on the user's head movement.

[0143] Based at least in part on these inertial measurements, patch trajectory estimate 906 may predict that at time t1, the user will have viewpoint 1304 and at time t2, viewpoint 1306. Thus, patch trajectory estimate 906 may predict that patch 1308 may move to patch 1310 at time t1 and to patch 1312 at time t2.

[0144] As an example of such an approach, it may be used to provide accurate and low-latency estimation of head pose within an AR system. The patch may be positioned to encompass an image of a stationary object in the user's environment. As a specific example, processing of image information may identify the corner of a picture frame hanging on a wall as a recognizable and stationary object for tracking. The processing may center the patch on the object. As with the moving object 1202 described above in connection with FIG. 12, relative movement between the object and the user's head will produce an event that can be used to calculate relative motion between the user and the tracked object. In this example, because the tracked object is stationary, the relative motion represents the motion of the imaging array worn by the user. That motion, therefore, represents changes in the user's head pose relative to the physical world and can be used to maintain an accurate calculation of the user's head pose, which can be used in realistically rendering virtual objects. Because imaging arrays as described herein can provide fast updates using a relatively small amount of data per update, calculations for rendering virtual objects remain accurate (they can be performed quickly and updated frequently).

[0145] Referring back to FIG. 11 , patch trajectory estimation 906 may include adjusting the size of at least one of the patches (act 1106) based, at least in part, on the calculated patch trajectory. For example, the size of the patch may be set large enough to include pixels onto which an image of the movable object or at least a portion of the object for which image information is to be generated will be projected. The patch may be set slightly larger than the projected size of the image of the portion of the object of interest so that if there is any error in estimating the trajectory of the patch, the patch may still include a relevant portion of the image. As the object moves relative to the image sensor, the size of the object's image in pixels may change based on distance, angle of incidence, object orientation, or other factors. The processor defining the patch associated with the object may set the size of the patch associated with the object, such as by measuring the size of the patch based on other sensor data or by calculating the size of the patch based on a world model. Other parameters of the patch, such as its shape, may be set or updated as well.

[0146] 14 depicts an image sensing system 1400 configured for use in an XR system, according to some embodiments. Like image sensing system 400 (FIG. 4), image sensing system 1400 includes circuitry for selectively outputting values ​​within a patch and may also be configured to output events related to pixels within the patch, as described above. In addition, image sensing system 1400 is configured to selectively output measured intensity values, which may be output for a complete image frame.

[0147] In the illustrated embodiment, separate outputs are shown for events and intensity values ​​generated using the DVS techniques described above. The output generated using the DVS techniques may be output as AER 1418, using the representation as described above in connection with AER 418. The output representing the intensity values ​​may be output through an output, here designated as APS 1420. These intensity outputs may be for a patch or for the entire image frame. The AER and APS outputs may be active simultaneously. However, in the illustrated embodiment, the image sensor 1400 operates in either a mode for outputting events or a mode in which intensity information is output at any given time. When such an image sensor is used, the system may selectively use the event output and / or the intensity information.

[0148] The image sensing system 1400 may include an image sensor 1402, which may include an image array 1404, which may contain a plurality of pixels 1500, each responsive to light. The sensor 1402 may further include circuitry for accessing the pixel cells. The sensor 1402 may further include circuitry that generates input to the access circuitry to control the mode in which information is read from the pixel cells in the image array 1404.

[0149] In the illustrated embodiment, the image array 1404 is configured as an array with multiple rows and columns of pixel cells, both of which are accessible in read mode. In such an embodiment, the access circuitry may include a row address encoder / decoder 1406, a column address encoder / decoder 1408 that controls a column select switch 1422, and / or a register 1424 that may temporarily hold information about incident light sensed by one or more corresponding pixel cells. The patch tracking engine 1410 may generate inputs to the access circuitry to control the pixel cells providing image information at any time.

[0150] In some embodiments, the image sensor 1402 may be configured to operate in a rolling shutter mode, a global shutter mode, or both. For example, the patch tracking engine 1410 may generate inputs to the access circuitry and control the readout mode of the image array 1402.

[0151] When sensor 1402 operates in rolling shutter readout mode, a single column of pixel cells is selected during each system clock, for example, by closing a single column switch 1422 of the plurality of column switches. During that system clock, the selected column of pixel cells is exposed and read out by APS 1420. To generate an image frame in rolling shutter mode, the columns of pixel cells in sensor 1402 are read out column by column and then processed by an image processor to generate the image frame.

[0152] When sensor 1402 operates in global shutter mode, columns of pixel cells are exposed simultaneously, e.g., within a single system clock, so that information captured by pixel cells in multiple columns can be read out simultaneously to APS 1420b, storing the information in register 1424. Such a readout mode allows for direct output of an image frame without requiring further data processing. In the illustrated embodiment, information about incident light sensed by a pixel cell is stored in individual registers 1424. It should be understood that multiple pixel cells may share one register 1424.

[0153] In some embodiments, the sensor 1402 may be implemented within a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, the image array 1404 may be implemented within a single integrated circuit. The patch tracking engine 1410, the row address encoder / decoder 1406, the column address encoder / decoder 1408, the column select switch 1422, and / or the register 1424 may be implemented within a second single integrated circuit, for example, configured as a driver for the image array 1404. Alternatively, or in addition, some or all of the functionality of the patch tracking engine 1410, the row address encoder / decoder 1406, the column address encoder / decoder 1408, the column select switch 1422, and / or the register 1424 may be distributed to other digital processors within the AR system.

[0154] 15 illustrates an example pixel cell 1500. In the illustrated embodiment, each pixel cell may be configured to output either event or intensity information. However, it should be understood that in some embodiments, the image sensor may be configured to output both types of information in parallel.

[0155] Both the event information and the intensity information are based on the output of the photodetector 504, as described above in connection with FIG. 5. The pixel cell 1500 includes circuitry for generating the event information. The circuitry includes a light-sensing circuit 502, a difference circuit 506, and a comparator 508, also as described above. A switch 1520, when in a first state, connects the photodetector 504 to the event-generating circuitry. The switch 1520 or other control circuitry may be controlled by a processor that controls the AR system so that a relatively small amount of image information is provided for a substantial period of time when the AR system is operational.

[0156] The switches 1520 or other control circuitry may also be controlled to configure the pixel cells 1500 to output intensity information. In the illustrated example, the intensity information is provided as a complete image frame, represented continuously as a stream of pixel intensity values ​​for each pixel in the image array. To operate in this mode, the switches 1520 in each pixel cell may be set to a second position, exposing the output of the photodetector 504 so that it can be connected to an output line after passing through the amplifier 510.

[0157] In the illustrated embodiment, the output lines are illustrated as column lines 1510. There may be one such column line per column in the image array. Each pixel cell in a column may be coupled to a column line 1510, but the pixel array may be controlled so that one pixel cell at a time is coupled to the column line 1510. There is one such switch within each pixel cell, switch 1530, which controls when the pixel cell 1500 is connected to its respective column line 1510. Access circuitry, such as row address decoder 410, may close switch 1530, ensuring that only one pixel cell is connected to each column line at a time. Switches 1520 and 1530 may be implemented using one or more transistors or similar components that are part of the image array.

[0158] 15 shows additional components that may be included within each pixel cell, according to some embodiments. A sample and hold circuit (S / H) 1532 may be connected between the photodetector 504 and the column line 1510. When present, the S / H 1532 may enable the image sensor 1402 to operate in a global shutter mode. In global shutter mode, a trigger signal is sent to each pixel cell in the array in parallel. Within each pixel cell, the S / H 1532 captures a value that indicates the intensity at the time of the trigger signal. The S / H 1532 stores that value and generates an output based on that value until the next value is captured.

[0159] As shown in FIG. 15 , a signal representing the value stored by S / H 1532 may be coupled to column line 1510 when switch 1530 is closed. The signal coupled to the column line may be processed to produce the output of the image array. The signal may be buffered and / or amplified in amplifier 1512, for example, upon exiting column line 1510, and then applied to analog-to-digital converter (A / D) 1514. The output of A / D 1514 may be passed through other readout circuitry 1516 to output 1420. Readout circuit 1516 may include, for example, column switch 1422. Other components in readout circuit 1516 may perform other functions, such as serializing the multi-bit output of A / D 1514.

[0160] Those skilled in the art will understand how to implement circuits to perform the functions described herein. S / H 1532 may be implemented, for example, as one or more capacitors and one or more switches. However, it should be understood that S / H 1532 may be implemented using other components or in circuit configurations other than those illustrated in FIG. 15A. It should be understood that other components other than those illustrated may also be implemented. For example, FIG. 15 shows one amplifier and one A / D converter per column. In other embodiments, there may be one A / D converter shared across multiple columns.

[0161] In a pixel array configured for a global shutter, each S / H 1532 may store intensity values ​​reflecting image information at the same moment in time. These values ​​may be stored as the values ​​stored in each pixel are successively read during a readout phase. Successive readout may be achieved, for example, by connecting the S / H 1532 of each pixel cell in a row to its respective column line. The values ​​on the column lines may then be passed to the APS output 1420 one at a time. Such information flow may be controlled by sequencing the opening and closing of column switches 1422, whose operation may be controlled, for example, by the column address decoder 1408. Once the values ​​for each pixel in one row have been read, pixel cells in the next row may be connected to column lines in their place. Their values ​​may be read one column at a time. This process of reading values ​​one row at a time may be repeated until intensity values ​​for all pixels in the image array have been read. In embodiments where intensity values ​​are read for one or more patches, the process will be complete once the values ​​for the pixel cells within the patch have been read.

[0162] The pixel cells may be read in any suitable order. The rows may be interleaved, for example, so that every third row is read in sequence. The AR system may still process the image data as frames of image data by deinterleaving the data.

[0163] In embodiments in which S / H 1532 is not present, values ​​may still be read from each pixel cell sequentially as rows and columns of values ​​are scanned out. However, the value read from each pixel cell may represent the intensity of light detected at the cell's photodetector at the time the value in that cell was captured as part of the readout process, such as when that value is applied to A / D 1514. As a result, with a rolling shutter, the pixels of an image frame may represent images incident on the image array at slightly different times. For an image sensor that outputs complete frames at a rate of 30 Hz, the difference in time between when the first pixel value for a frame is captured and when the last pixel value for the frame is captured may vary by 1 / 30 of a second, which is imperceptible for many applications.

[0164] For some XR functions, such as tracking objects, the XR system may perform calculations on image information collected using an image sensor using a rolling shutter. Such calculations may interpolate between successive image frames and, for each pixel, calculate an interpolated value representing the pixel's estimated value at a point in time between successive frames. The same time may be used for all pixels, such that the interpolated image frame, via the calculation, contains pixels representing the same point in time as might be produced using an image sensor with a global shutter. Alternatively, a global shutter image array may be used for one or more image sensors in a wearable device forming part of the XR system. A global shutter for full or partial image frames may avoid interpolation, other processing that may be performed to compensate for variations in capture time in image information captured using a rolling shutter. Interpolation calculations can therefore be avoided even when image information is used to track a hand or other movable object, or to determine the head pose of a user of a wearable device in an AR system, or even to track the movement of an object, such as may occur when the image information is collected using a camera on the wearable device, which may be moving at the time the image information is collected, for processing to construct an accurate representation of the physical environment.

[0165] Distinguished pixel cells

[0166] In some embodiments, each pixel cell in the sensor array may be identical. Each pixel cell may be responsive to a broad spectrum of visible light, for example. Each photodetector may therefore provide image information indicative of the intensity of the visible light. In this scenario, the output of the image array may be a "grayscale" output indicative of the amount of visible light incident on the image array.

[0167] In other embodiments, pixel cells may be differentiated. For example, different pixel cells in a sensor array may output image information indicative of the intensity of light within a particular portion of the spectrum. A suitable technique for differentiating pixel cells is to position a filter element in the light path leading to the photodetector in the pixel cell. The filter element may be, for example, a bandpass filter that allows visible light of a particular color to pass through. Applying such a color filter over a pixel cell configures that pixel cell to provide image information indicative of the intensity of light of the color corresponding to the filter.

[0168] Filters may be applied across pixel cells regardless of the pixel cell's structure. They may be applied across pixel cells in a sensor array, for example, with a global shutter or a rolling shutter. Similarly, filters may be applied to pixel cells configured to output an intensity or intensity change using DVS techniques.

[0169] In some embodiments, a filter element that selectively passes light of a primary color may be mounted across the photodetectors in each pixel cell in the sensor array. For example, a filter that selectively passes red, green, or blue light may be used. The sensor array may have multiple subarrays, each having one or more pixels configured to sense light of a respective primary color. In this way, the pixel cells in each subarray provide both intensity and color information about the object being imaged by the image sensor.

[0170] The inventors have recognized and appreciated that in an XR system, some functions require color information, while some functions can be implemented with grayscale information. A wearable device equipped with an image sensor and providing image information related to the operation of an XR system may have multiple cameras, some of which may be formed with image sensors that can provide color information. Others of the cameras may be grayscale cameras. The inventors have recognized and appreciated that a grayscale camera may consume less power, be more sensitive in low light conditions, output data more quickly, and / or output less data to represent the same range of the physical world with the same resolution as a camera formed with a comparable image sensor configured to sense color. However, a grayscale camera may output sufficient image information for many functions implemented within the XR system. Thus, an XR system may be configured with both grayscale and color cameras, using primarily grayscale cameras or multiple cameras and selectively using color cameras.

[0171] For example, an XR system may collect and process image information to create a passable world model. That processing may use color information, which may improve the effectiveness of certain functions, such as distinguishing between objects, identifying surfaces associated with the same object, and / or recognizing objects. Such processing may be performed or updated at any time, for example, when a user first turns on the system, moves into a new environment by walking into another room, or when a change in the user's environment is otherwise detected.

[0172] Other functions are not significantly improved through the use of color information. For example, once a passable world model is created, the XR system may use images from one or more cameras to determine the orientation of the wearable device relative to features within the passable world model. Such a function may be performed, for example, as part of head pose tracking. Some or all of the cameras used for such a function may be grayscale. Because head pose tracking is performed frequently as the XR system operates, in some embodiments, persistent use of one or more grayscale cameras for this function may provide significant power savings, reduced computation, or other benefits.

[0173] Similarly, at various times during operation of an XR system, the system may use stereoscopic information from two or more cameras to determine distances to movable objects. Such functionality may require high-rate processing of image information as part of tracking a user's hand or other movable object. Using one or more grayscale cameras for this functionality may provide lower latency or other advantages associated with processing high-resolution image information.

[0174] In some embodiments of the XR system, the XR system may have both a color and at least one grayscale camera and may selectively enable the grayscale and / or color cameras based on the function for which the image information from those cameras is to be used.

[0175] Pixel cells within an image sensor may be distinguished in ways other than based on the spectrum of light to which they are sensitive. In some embodiments, some or all of the pixel cells may produce an output having an intensity indicative of the angle of arrival of light incident on the pixel cell. The angle of arrival information may be processed to calculate the distance to the object being imaged.

[0176] In such an embodiment, the image sensor may obtain depth information passively by placing a component in the light path to a pixel cell in the array such that the pixel cell outputs information indicative of the angle of arrival of light striking that pixel cell. An example of such a component is a transmissive diffractive mask (TDM) filter.

[0177] The angle of arrival information may be converted, through calculations, into distance information indicating the distance to the object from which the light is being reflected. In some embodiments, pixel cells configured to provide angle of arrival information may be interspersed with pixel cells that capture light intensities of one or more colors. As a result, the angle of arrival information, and therefore the distance information, may be combined with other image information about the object.

[0178] In some embodiments, one or more of the sensors may be configured to obtain information about physical objects in a scene at high frequencies with low latency using compact and low-power components. The image sensor may, for example, draw less than 50 mW, allowing the device to be powered by a battery small enough to be used as part of a wearable system. The sensor may be an image sensor configured to passively obtain depth information instead of, or in addition to, image information indicating changes in the intensity and / or intensity of one or more color information. Such sensors may also be configured to provide small amounts of data by using patch tracking or by using DVS techniques to provide a differential output.

[0179] Passive depth information may be obtained by configuring an image array, such as an image array incorporating any one or more of the techniques described herein, with a component that adapts one or more of the pixel cells in the array to output information indicative of the light field emanating from the object being imaged. The information may be based on the angle of arrival of light striking that pixel. In some embodiments, pixel cells, such as those described above, may be configured to output an indication of angle of arrival by placing a plenoptic component in the light path to the pixel cell. An example of a plenoptic component is a transmissive diffractive mask (TDM). The angle of arrival information may be converted, through calculations, into distance information indicative of the distance to the object from which the light is being reflected, forming the captured image. In some embodiments, pixel cells configured to provide angle of arrival information may be interspersed with pixel cells that capture light intensity in grayscale or one or more colors. As a result, the angle of arrival information may also be combined with other image information about the object.

[0180] FIG. 16 illustrates a pixel subarray 100 according to some embodiments. In the illustrated embodiment, the subarray has two pixel cells, but the number of pixel cells in the subarray is not a limitation on the present invention. Here, a first pixel cell 121 and a second pixel cell 122 are shown, one of which is configured to capture angle-of-arrival information (first pixel cell 121), but it should be understood that the number and location within the array of pixel cells configured to measure angle-of-arrival information can be varied. In this example, the other pixel cell (second pixel cell 122) is configured to measure the intensity of one color of light, but other configurations are possible, including pixel cells sensitive to different colors of light, or one or more pixel cells sensitive to a broad spectrum of light, such as in a grayscale camera.

[0181] The first pixel cell 121 of the pixel subarray 100 of FIG. 16 includes an angle-of-arrival / intensity converter 101, a photodetector 105, and differential readout circuitry 107. The second pixel cell 122 of the pixel subarray 100 includes a color filter 102, a photodetector 106, and differential readout circuitry 108. It should be understood that not all of the components illustrated in FIG. 16 need be included in all embodiments. For example, some embodiments may not include differential readout circuitry 107 and / or 108, and some embodiments may not include color filter 102. Furthermore, additional components not shown in FIG. 16 may be included. For example, some embodiments may include a polarizer arranged to allow light of a particular polarization to reach the photodetector. As another example, some embodiments may include scan output circuitry instead of or in addition to the differential readout circuitry 107. As another example, the first pixel cell 121 may also include a color filter such that the first pixel 121 measures both the angle of arrival and the intensity of a particular color of light incident on the first pixel 121.

[0182] The angle-of-arrival-to-intensity converter 101 of the first pixel 121 is an optical component that converts the angle θ of the incident light 111 to an intensity that can be measured by a photodetector. In some embodiments, the angle-of-arrival-to-intensity converter 101 may include refractive optics. For example, one or more lenses may be used to convert the angle of incidence of the light to a position on the image plane, the amount of which is detected by one or more pixel cells. In some embodiments, the angle-of-arrival-to-position-intensity converter 101 may include diffractive optics. For example, one or more diffraction gratings (e.g., transmissive diffractive masks (TDMs)) may convert the angle of incidence of the light to an intensity that can be measured by a photodetector below the TDM.

[0183] The photodetector 105 of the first pixel cell 121 receives the incident light 110 that passes through the angle of arrival to intensity converter 101 and generates an electrical signal based on the intensity of the light incident on the photodetector 105. The photodetector 105 is located at an image plane associated with the angle of arrival to intensity converter 101. In some embodiments, the photodetector 105 may be a single pixel of an image sensor, such as a CMOS image sensor.

[0184] The differential readout circuitry 107 of the first pixel 121 receives a signal from the photodetector 105 and outputs an event only when the amplitude of the electrical signal from the photodetector differs from the amplitude of the previous signal from the photodetector 105, implementing the DVS technique as described above.

[0185] The second pixel cell 122 includes a color filter 102 for filtering the incident light 112 so that only light within a certain wavelength range passes through the color filter 102 and impinges on the photodetector 106. The color filter 102 may be a bandpass filter, for example, allowing one of red, green, or blue light to pass through it and rejecting light of other wavelengths, and / or may limit the IR light reaching the photodetector 106 to only a certain portion of the spectrum.

[0186] In this embodiment, the second pixel cell 122 also includes a photodetector 106 and a differential readout circuitry 108, which may function in a similar manner to the photodetector 105 and differential readout circuitry 107 of the first pixel cell 121.

[0187] As mentioned above, in some embodiments, an image sensor may include an array of pixels, each pixel associated with a photodetector and readout circuitry. A subset of the pixels may be associated with an angle-of-arrival-to-intensity converter used to determine the angle of detected light incident on the pixel. Another subset of the pixels may be associated with color filters used to determine color information about the scene being viewed, or may selectively pass or block light based on other characteristics.

[0188] In some embodiments, the angle of arrival of light may be determined using a single photodetector and diffraction gratings at two different depths. For example, light may be incident on a first TDM, which converts the angle of arrival to position, and a second TDM may be used to selectively pass light incident at a particular angle. Such an arrangement may utilize the Talbot effect, a near-field diffraction effect in which, when a plane wave is incident on a diffraction grating, an image of the diffraction grating is created at a distance from the diffraction grating. If a second diffraction grating is placed at the image plane, where the image of the first diffraction grating is formed, the angle of arrival may be determined from the intensity of the light measured by a single photodetector positioned after the second grating.

[0189] 17A illustrates a first arrangement of pixel cells 140 including a first TDM 141 and a second TDM 143 that are aligned with one another such that the increased refractive index ridges and / or regions for the two gratings are aligned horizontally (Δs=0), where Δs is the horizontal offset between the first TDM 141 and the second TDM 143. Both the first TDM 141 and the second TDM 143 may have the same grating period d, and the two gratings may be separated by a distance / depth z. The depth z, known as the Talbot length, at which the second TDM 143 is located relative to the first TDM 141 may be determined by the grating period d and the wavelength λ of the light being analyzed, and is given by the following equation: [ka]

[0190] As shown in FIG. 17A , incident light 142 with a zero-degree arrival angle θ is diffracted by a first TDM 141. A second TDM 143 is located at a depth equal to the Talbot length such that an image of the first TDM 141 is created, resulting in most of the incident light 142 passing through the second TDM 143. An optional dielectric layer 145 may separate the second TDM 143 from a photodetector 147. As the light passes through the dielectric layer 145, the photodetector 147 detects the light and generates an electrical signal (e.g., voltage or current) proportional to the intensity of the light incident on the photodetector. On the other hand, incident light 144 with a non-zero arrival angle θ is also diffracted by the first TDM 141, but the second TDM 143 prevents at least a portion of the incident light 144 from reaching the photodetector 147. The amount of incident light reaching photodetector 147 depends on the arrival angle θ, with larger angles resulting in less light reaching the photodetector. The dashed line emanating from light 144 illustrates the attenuation of the amount of light reaching photodetector 147. In some cases, light 144 may be completely blocked by diffraction grating 143. Therefore, information about the arrival angle of the incident light may be obtained using a single photodetector 147 using two TDMs.

[0191] In some embodiments, information obtained by neighboring pixel cells without an angle-of-arrival-to-intensity converter may provide an indication of the intensity of the incident light and may be used to determine the portion of the incident light that passes through the angle-of-arrival-to-intensity converter. From this image information, the angle of arrival of the light detected by photodetector 147 may be calculated, as described in more detail below.

[0192] 17B illustrates a second arrangement of pixel cells 150 including a first TDM 151 and a second TDM 153 that are misaligned with one another such that the increased refractive index ridges and / or regions for the two gratings are not horizontally aligned (Δs ≠ 0), where Δs is the horizontal offset between the first TDM 151 and the second TDM 153. Both the first TDM 151 and the second TDM 153 may have the same grating period d, and the two gratings may be separated by a distance / depth z. Unlike the situation discussed in connection with FIG. 17A , where the two TDMs are aligned, the misalignment results in incident light at angles other than zero passing through the second TDM 153.

[0193] As shown in FIG. 17B, incident light 152 with a zero-degree arrival angle is diffracted by the first TDM 151. The second TDM 153 is located at a depth equal to the Talbot length, but due to the horizontal offset of the two gratings, at least a portion of the light 152 is blocked by the second TDM 153. The dashed line emanating from the light 152 illustrates that the amount of light reaching the photodetector 157 is attenuated. In some cases, the light 152 may be completely blocked by the diffraction grating 153. On the other hand, incident light 154 with a non-zero arrival angle θ is diffracted by the first TDM 151 but passes through the second TDM 153. After traversing the optional dielectric layer 155, the photodetector 157 detects the light incident thereon and generates an electrical signal (e.g., voltage or current) proportional to the intensity of the light incident thereon.

[0194] Pixel cells 140 and 150 have different output functions, with different intensities of detected light for different angles of incidence. In either case, however, the relationship may be fixed and based on the pixel cell design, or determined by measurements as part of a calibration process. Regardless of the precise transfer function, the measured intensity may be converted to an angle of arrival, which may in turn be used to determine the distance to the object being imaged.

[0195] In some embodiments, different pixel cells of an image sensor may have different arrangements of TDMs. For example, a first subset of pixel cells may include a first horizontal offset between the grids of two TDMs associated with each pixel, while a second subset of pixel cells may include a second horizontal offset between the grids of two TDMs associated with each pixel cell, the first offset being different from the second offset. Each subset of pixel cells with a different offset may be used to measure a different angle of arrival or a range of different angles of arrival. For example, the first subset of pixels may include an arrangement of TDMs similar to pixel cell 140 of FIG. 17A, while the second subset of pixels may include an arrangement of TDMs similar to pixel cell 150 of FIG. 17B.

[0196] In some embodiments, not all pixel cells of an image sensor include a TDM. For example, a subset of pixel cells may include a color filter, while a different subset of pixel cells may include a TDM for determining angle-of-arrival information. In other embodiments, color filters are not used, such that a first subset of pixel cells simply measures the overall intensity of incident light, and a second subset of pixel cells measures angle-of-arrival information. In some embodiments, information about the intensity of light from neighboring pixel cells without a TDM may be used to determine the angle of arrival for light incident on a pixel cell with one or more TDMs. For example, using two TDMs arranged to exploit the Talbot effect, the intensity of light incident on a photodetector after the second TDM is a sinusoidal function of the angle of arrival of light incident on the first TDM. Thus, if the total intensity of light incident on the first TDM is known, the angle of arrival of the light may be determined from the intensity of light detected by the photodetector.

[0197] In some embodiments, the configuration of pixel cells within a subarray may be selected to provide various types of image information with appropriate resolution. Figures 18A-C illustrate an example arrangement of pixel cells within a pixel subarray of an image sensor. It should be understood that the illustrated example is non-limiting and that alternative pixel arrangements are also contemplated by the inventors. This arrangement may be repeated across the image array, which may contain millions of pixels. The subarray may include one or more pixel cells that provide angle-of-arrival information for the incident light and one or more other pixel cells (with or without color filters) that provide intensity information for the incident light.

[0198] 18A is an example of a pixel sub-array 160 including a first set of pixel cells 161 and a second set of pixel cells 163 that are different from one another and are rectangular rather than square. Pixel cells labeled "R" are pixel cells with a red filter so that incident red light passes through the filter to an associated photodetector, pixel cells labeled "B" are pixel cells with a blue filter so that incident blue light passes through the filter to an associated photodetector, and pixel cells labeled "G" are pixel cells with a green filter so that incident green light passes through the filter to an associated photodetector. In the example sub-array 160, there are more green pixel cells than red or blue pixel cells, illustrating that the various types of pixel cells need not be present in equal proportions.

[0199] The pixel cells labeled A1 and A2 are pixels that provide angle-of-arrival information. For example, pixel cells A1 and A2 may include one or more grids for determining angle-of-arrival information. The pixel cells that provide angle-of-arrival information may be similarly configured or may be differently configured, such as to be sensitive to different ranges of arrival angles or arrival angles relative to different axes. In some embodiments, the pixels labeled A1 and A2 include two TDMs, and the TDMs of pixel cells A1 and A2 may be oriented in different directions, for example, perpendicular to each other. In other embodiments, the TDMs of pixel cells A1 and A2 may be oriented parallel to each other.

[0200] In embodiments using pixel sub-array 160, both color image data and angle-of-arrival information may be obtained. To determine the angle of arrival of light incident on a set of pixel cells 161, the total light intensity incident on set 161 is estimated using electrical signals from the RGB pixel cells. Using the fact that the intensity of light detected by the A1 / A2 pixel varies in a predictable manner as a function of the angle of arrival, the angle of arrival may be determined by comparing the total intensity (estimated from the RGB pixel cells in the group of pixels) with the intensity measured by the A1 and / or A2 pixel cells. For example, the intensity of light incident on the A1 and / or A2 pixel may vary sinusoidally with the angle of arrival of the incident light. The angle of arrival of light incident on set 163 of pixel cells is determined in a similar manner using electrical signals generated by the pixels of set 163.

[0201] 18A shows a specific embodiment of a subarray, and it should be understood that other configurations are possible. In some embodiments, for example, a subarray may be only a set of pixel cells 161 or 163.

[0202] FIG. 18B illustrates an alternative pixel subarray 170 including a first set of pixel cells 171, a second set of pixel cells 172, a third set of pixel cells 173, and a fourth set of pixel cells 174. Each set of pixel cells 171-174 is square and has identical arrangements of pixel cells therein, but may have pixel cells for determining angle-of-arrival information spanning different angular ranges or for different planes (e.g., the TDMs of pixels A1 and A2 may be oriented perpendicular to each other). Each set of pixels 171-174 includes one red pixel cell (R), one blue pixel cell (B), one green pixel cell (G), and one angle-of-arrival pixel cell (A1 or A2). Note that in the exemplary pixel subarray 170, equal numbers of red / green / blue pixel cells are present in each set. It should be further understood that pixel subarrays may be repeated in one or more directions to form larger arrays of pixels.

[0203] In embodiments using pixel sub-array 170, both color image data and angle-of-arrival information may be obtained. To determine the angle of arrival of light incident on a set of pixel cells 171, the total light intensity incident on the set 171 may be estimated using signals from the RGB pixel cells. Using the fact that the intensity of light detected by the angle-of-arrival pixel cells has a sinusoidal or other predictable response to the angle of arrival, the angle of arrival may be determined by comparing the total intensity (estimated from the RGB pixel cells) with the intensity measured by the A1 pixel. The angle of arrival of light incident on a set of pixel cells 172-174 may be determined in a similar manner using the electrical signals generated by the pixel cells of each individual pixel of the set.

[0204] 18C illustrates an alternative pixel sub-array 180 including a first set of pixel cells 181, a second set of pixel cells 182, a third set of pixel cells 183, and a fourth set of pixel cells 184. Each set of pixel cells 181-184 is square and has identical arrangements of pixel cells therein, and no color filters are used. Each set of pixel cells 181-184 includes two "white" pixels (e.g., no color filters so that red, blue, and green light are detected to form a grayscale image), one arriving pixel cell (A1) with a TDM oriented in a first direction, and one arriving pixel cell (A2) with a TDM oriented with a second spacing or in a second direction (e.g., perpendicular) relative to the first direction. Note that no color information is present in the exemplary pixel sub-array 170. The resulting image is grayscale, illustrating that passive depth information can be obtained using techniques as described herein in color or grayscale image arrays. As with the other subarray configurations described herein, the pixel subarray arrangement may be repeated in one or more directions to form a larger array of pixels.

[0205] In embodiments using pixel sub-array 180, both grayscale image data and angle-of-arrival information may be obtained. To determine the angle of arrival of light incident on set 181, the total light intensity incident on set 181 is estimated using electrical signals from two white color pixels. Using the fact that the intensity of light detected by the A1 and A2 pixels has a sinusoidal or other predictable response to angle of arrival, the angle of arrival may be determined by comparing the total intensity (estimated from the white pixels) with the intensity measured by the A1 and / or A2 pixel cells. The angle of arrival of light incident on set 182-184 may be determined in a similar manner using the electrical signals generated by each individual pixel in the set.

[0206] In the above examples, the pixel cells are illustrated as square and arranged in a square grid. The embodiments are not so limited. For example, in some embodiments, the pixel cells may be rectangular. Furthermore, the subarrays may be triangular, arranged diagonally, or have other geometric shapes.

[0207] In some embodiments, the angle of arrival information is obtained using the image processor 708 or a processor associated with the local data processing module 70, which may further determine the distance of the object based on the angle of arrival. For example, the angle of arrival information may be combined with one or more other types of information to obtain the distance of the object. In some embodiments, the object in the mesh model 46 may be associated with the angle of arrival information from the pixel array. The mesh model 46 may include the location of the object, including its distance from the user, which may be updated to a new distance value based on the angle of arrival information.

[0208] Using angle-of-arrival information to determine distance values ​​can be particularly useful in scenarios where objects are close to the user. This is because a change in distance from the image sensor results in a larger change in the angle of arrival of light for a nearby object than a similar magnitude change in distance for an object located farther away from the user. Thus, a processing module utilizing passive distance information based on angle-of-arrival may selectively use that information based on the object's estimated distance and may utilize one or more other techniques to determine distances to objects that, in some embodiments, exceed a threshold distance, such as up to 1 meter, up to 3 meters, or up to 5 meters. As a specific example, the processing module of an AR system may be programmed to use passive distance measurement using angle-of-arrival information for objects within 3 meters of a user of a wearable device, but for objects outside that range, it may use stereoscopic image processing using images captured by two cameras.

[0209] Similarly, pixels configured to detect angle of arrival information may be most sensitive to changes in distance within a certain range of angles from the normal to the image array. The processing module similarly uses distance information derived from angle of arrival measurements within that range of angles, but may be configured to use other sensors and / or other techniques to determine distances outside that range.

[0210] One exemplary application for determining the distance of an object from an image sensor is hand tracking. Hand tracking may be used in an AR system, for example, to provide a gesture-based user interface for system 80 and / or to allow a user to move virtual objects within an environment in an AR experience provided by system 80. The combination of an image sensor providing angle-of-arrival information for accurate depth determination and differential readout circuitry for reducing the amount of data to process for determining the user's hand movement provides an efficient interface by which the user can interact with virtual objects and / or provide input to system 80. A processing module for determining the location of the user's hand may use distance information obtained using different techniques depending on the location of the user's hand within the field of view of the image sensor of the wearable device. Hand tracking, according to some embodiments, may be implemented as a form of patch tracking during the image sensing process.

[0211] Another application for which depth information can be useful is occlusion processing. Occlusion processing uses depth information to determine that certain portions of a model of the physical world do not need to or cannot be updated based on image information captured by one or more image sensors that collect image information about the physical environment around the user. For example, if a first object is determined to be present at a first distance from the sensor, the system 80 may determine not to update the model of the physical world for distances greater than the first distance. For example, even if the model includes a second object at a second distance from the sensor, and the second distance is greater than the first distance, model information for that object may not be updated if it is behind the first object. In some embodiments, the system 80 may generate an occlusion mask based on the location of the first object and update only portions of the model that are not masked by the occlusion mask. In some embodiments, the system 80 may generate more than one occlusion mask for more than one object. Each occlusion mask may be associated with a distinct distance from the sensor. For each occlusion mask, model information associated with objects at a distance from the sensor greater than the distance associated with the respective occlusion mask will not be updated. By limiting the portion of the model that is updated at any given time, the speed at which the AR environment is generated and the amount of computational resources required to generate the AR environment are reduced.

[0212] 18A-C , some embodiments of the image sensor may include pixels with IR filters in addition to or instead of color filters. For example, the IR filter may allow light of a wavelength, such as approximately equal to 940 nm, to pass through and be detected by an associated photodetector. Some embodiments of the wearable may include an IR light source (e.g., an IR LED) that emits light of the same wavelength as that associated with the IR filter (e.g., 940 nm). The IR light source and IR pixels may be used as alternative methods of determining the distance of an object from the sensor. By way of example, and not limitation, the IR light source may be pulsed and time-of-flight measurements may be used to determine the object distance from the sensor.

[0213] In some embodiments, system 80 may be capable of operating in one or more operating modes. A first mode may be a mode in which depth determination is performed using passive depth measurements, for example, based on the angle of arrival of light determined using pixels with angle-of-arrival-to-intensity converters. A second mode may be a mode in which depth determination is performed using active depth measurements, for example, based on the time-of-flight of IR light measured using IR pixels of an image sensor. A third mode may determine the distance of an object using stereoscopic measurements from two separate image sensors. Such stereoscopic measurements may be more accurate than using the angle of arrival of light determined using pixels with angle-of-arrival-to-intensity converters when the object is very far from the sensors. Other suitable methods of determining depth may also be used for one or more additional operating modes for depth determination.

[0214] In some embodiments, it may be preferable to use passive depth determination because such techniques utilize less power. However, the system may determine that it should operate in active mode under certain conditions. For example, if the intensity of visible light being detected by the sensor is below a threshold, it may be too dark to accurately perform passive depth determination. As another example, an object may be too far away for passive depth determination to be accurate. Therefore, the system may be programmed to select to operate in a third mode in which depth is determined based on stereoscopic measurements of the scene using two spatially separated image sensors. As another example, determining the depth of an object based on the angle of arrival of light determined using pixels with an angle-of-arrival / intensity converter may be inaccurate at the periphery of the image sensor. Thus, if an object is detected by pixels near the periphery of the image sensor, the system may select to operate in the second mode using active depth determination.

[0215] While the image sensor embodiments described above used individual pixel cells with stacked TDMs to determine the angle of arrival of light incident on the pixel cell, other embodiments may determine angle-of-arrival information using a group of multiple pixel cells with a single TDM across all pixels of the group. The TDM may project a pattern of light across the sensor array, where the pattern depends on the angle of arrival of the incident light. Multiple photodetectors associated with a single TDM may detect the pattern more accurately because each photodetector of the multiple photodetectors is located at a different position in the image plane (the image plane comprising the photodetector that senses light). The relative intensity sensed by each photodetector may indicate the angle of arrival of the incident light.

[0216] FIG. 19A is a top plan view example of a plurality of photodetectors (in the form of a photodetector array 120, which may be a subarray of pixel cells of an image sensor) associated with a single transmissive diffractive mask (TDM), according to some embodiments. FIG. 19B is a cross-sectional view of the same photodetector array as FIG. 19A taken along line A in FIG. 19A. In the example shown, the photodetector array 120 includes 16 individual photodetectors 121, which may be within pixel cells of the image sensor. The photodetector array 120 includes a TDM 123 disposed above the photodetectors. It should be understood that each group of pixel cells is illustrated with four pixels (e.g., forming a grid of four pixels by four pixels) for clarity and simplicity. Some embodiments may include more than four pixel cells. For example, 16 pixel cells, 64 pixel cells, or any other number of pixels may be included in each group.

[0217] The TDM 123 is located a distance x from the photodetector 121. In some embodiments, the TDM 123 is formed on the top surface of the dielectric layer 125, as illustrated in FIG. 19B. For example, the TDM 123 may be formed from ridges or by valleys etched into the surface of the dielectric layer 125, as illustrated. In other embodiments, the TDM 123 may be formed within the dielectric layer. For example, portions of the dielectric layer may be modified to have a higher or lower refractive index relative to other portions of the dielectric layer, resulting in a holographic phase grating. Light incident on the photodetector array 120 from above is diffracted by the TDM, resulting in the angle of arrival of the incident light being translated into a position in an image plane at a distance x from the TDM 123, where the photodetector 121 is located. The intensity of the incident light measured at each photodetector 121 in the photodetector array may be used to determine the angle of arrival of the incident light.

[0218] FIG. 20A illustrates an example of multiple photodetectors (in the form of a photodetector array 130) associated with multiple TDMs, according to some embodiments. FIG. 20B is a cross-sectional view of the same photodetector array as FIG. 20A taken through line B of FIG. 20A. FIG. 20C is a cross-sectional view of the same photodetector array as FIG. 20A taken through line C of FIG. 20A. The photodetector array 130, in the example shown, includes 16 individual photodetectors, which may reside within pixel cells of an image sensor. As shown, there are four groups 131a, 131b, 131c, and 131d of four pixel cells. The photodetector array 130 includes four individual TDMs 133a, 133b, 133c, and 133d, with each TDM provided above an associated group of pixel cells. It should be understood that each group of pixel cells is illustrated with four pixel cells for clarity and simplicity. Some embodiments may include more than four pixel cells. For example, 16 pixel cells, 64 pixel cells, or any other number of pixel cells may be included in each group.

[0219] Each TDM 133a-d is located a distance x from the photodetectors 131a-d. In some embodiments, the TDMs 133a-d are formed on the top surface of the dielectric layer 135, as illustrated in FIG. 20B. For example, the TDMs 133a-d may be formed from ridges or by valleys etched into the surface of the dielectric layer 135, as illustrated. In other embodiments, the TDMs 133a-d may be formed within the dielectric layer. For example, portions of the dielectric layer may be modified to have a higher or lower refractive index relative to other portions of the dielectric layer, resulting in a holographic phase grating. Light incident on the photodetector array 130 from above is diffracted by the TDMs, resulting in the angle of arrival of the incident light being translated into a position in an image plane at a distance x from the TDMs 133a-d, where the photodetectors 131a-d are located. The intensity of the incident light measured at each photodetector 131a-d in the photodetector array may be used to determine the angle of arrival of the incident light.

[0220] TDMs 133a-d may be oriented in different directions from one another. For example, TDM 133a is perpendicular to TDM 133b. Thus, the intensity of light detected using photodetector group 131a may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133a, and the intensity of light detected using photodetector group 131b may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133b. Similarly, the intensity of light detected using photodetector group 131c may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133c, and the intensity of light detected using photodetector group 131d may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133d.

[0221] Pixel cells configured to passively obtain depth information may be integrated into an image array with features as described herein to support operations useful in cross-reality systems. According to some embodiments, pixel cells configured to obtain depth information may be implemented as part of an image sensor used to implement a camera with a global shutter. Such an arrangement may provide, for example, a complete frame output. A complete frame may include image information for different pixels simultaneously, indicating depth and intensity. Using an image sensor in this arrangement, a processor may obtain depth information for an entire scene at once.

[0222] In other embodiments, pixel cells of the image sensor that provide depth information may be configured to operate according to the DVS techniques described above. In such a scenario, an event may indicate a change in the depth of an object as indicated by the pixel cell. The event output by the image array may indicate the pixel cell for which a change in depth was detected. Alternatively, or in addition, the event may include the value of the depth information for that pixel cell. With an image sensor in this configuration, the processor may obtain depth information updates at a very high rate to provide high temporal resolution.

[0223] In yet other embodiments, the image sensor may be configured to operate in either full-frame or DVS mode. In such embodiments, a processor processing image information from the image sensor may programmatically control the operational mode of the image sensor based on the function being performed by the processor. For example, while performing a function involving tracking objects, the processor may configure the image sensor to output image information as DVS events. On the other hand, while processing to update the world reconstruction, the processor may configure the image sensor to output full-frame depth information.

[0224] Wearable Configuration

[0225] Multiple image sensors may be used within an XR system. The image sensors may be combined with optical components such as lenses and control circuitry to create a camera. The image sensors may obtain imaging information using one or more of the techniques described above, such as grayscale imaging, color imaging, a global shutter, DVS techniques, plenoptic pixel cells, and / or dynamic patches. Regardless of the imaging technique used, the resulting camera may be mounted to a support member to form a headset, which may include or be connected to a processor.

[0226] FIG. 21 is a schematic diagram illustrating a headset 2100 of a wearable display system consistent with disclosed embodiments. As shown in FIG. 21 , headset 2100 may include a display device comprising monoculars 2110a and 2110b, which may be optical eyepieces or displays configured to transmit and / or display visual information to a user's eyes. Headset 2100 may also include a frame 2101, which may be similar to frame 64 described above with respect to FIG. 3B . Headset 2100 may further include three cameras (camera 2120a, camera 2120b, and camera 2140) and additional components such as emitter 2130a, emitter 2130b, inertial measurement unit 2170a (IMU 2170a), and inertial measurement unit 2170b (IMU 2170b).

[0227] Camera 2120a, camera 2120b, and camera 2140 are world cameras that are oriented to image the physical world as seen by a user wearing headset 2100. In some embodiments, these three cameras may be sufficient to obtain image information about the physical world, and these three cameras may be the only world-facing cameras. Headset 2100 may also include additional components, such as an eye-tracking camera, as discussed above with respect to FIG. 3B.

[0228] The monoculars 2110a and 2110b may be mechanically coupled to a support member, such as the frame 2101, using techniques such as adhesives, fasteners, or pressure fitting. Similarly, the three cameras and ancillary components (e.g., emitters, inertial measurement units, eye tracking cameras, etc.) may be mechanically coupled to the frame 2101 using techniques such as adhesives, fasteners, pressure fitting, etc. These mechanical couplings may be direct or indirect. For example, one or more of the cameras and / or one or more of the ancillary components may be directly attached to the frame 2101. As an additional example, one or more of the cameras and / or one or more of the ancillary components may be directly attached to the monocular, which may then be attached to the frame 2101. The mechanism of attachment is not intended to be limiting.

[0229] Alternatively, monocular subassemblies may be formed and then mounted to the frame 2101. Each subassembly may include, for example, a support member to which the monocular 2110a or 2110b is mounted. The IMU and one or more cameras may likewise be mounted to the support member. Mounting both the camera and the IMU on the same support member may allow inertial information about the camera to be obtained based on the output of the IMU. Similarly, mounting the monocular on the same support member as the camera may allow image information about the world to be spatially correlated with information rendered on the monocular.

[0230] The headset 2100 may be lightweight. For example, the headset 2100 may weigh between 30 and 300 grams. The headset 2100 may be made from a material that flexes during use, such as plastic or thin metal components. Such materials may allow for a lightweight and comfortable headset that can be worn by a user for extended periods of time. An XR system with such a lightweight headset may still support high-accuracy stereoscopic image analysis, which requires that the separation between the cameras be understood using a calibration routine that can be repeated as the headset is worn and compensate for any inaccuracies that may result from flexing of the headset during use. In some embodiments, the lightweight headset may include a battery pack. The battery pack may include one or more batteries, which may be rechargeable or non-rechargeable. The battery pack may be built into the lightweight frame or may be removable. The battery pack and the lightweight frame may be formed as a single unit, or the battery pack may be formed as a unit separate from the lightweight frame.

[0231] Camera 2120a and camera 2120b may each include an image sensor and a lens. The image sensor may be configured to produce grayscale images. The image sensor may be configured to acquire images with a size between 1 megapixel and 4 megapixels. For example, the image sensor may be configured to acquire images with a horizontal resolution of 1,016 lines by a vertical resolution of 1,016 lines. The image sensor may be configured to acquire images repeatedly or periodically. For example, the image sensor may be configured to acquire images at a frequency between 30 Hz and 120 Hz, such as 60 Hz. The image sensor may be a CMOS image sensor. The image sensor may be configured with a global shutter. As discussed above with respect to FIGS. 14 and 15, the global shutter may allow each pixel to acquire an intensity measurement simultaneously.

[0232] Camera 2120a and camera 2120b may each be configured to have a wide field of view, consistent with disclosed embodiments. For example, camera 2120a and camera 2120b may include equidistant lenses (e.g., fisheye lenses). Camera 2120a and camera 2120b may each be angled inward on headset 2100. For example, a vertical plane passing through the center of field of view 2121a, the field of view associated with camera 2120a, may intersect and form an angle with a vertical plane passing through the midline of headset 2100. This angle may be between 1 and 40 degrees. In some embodiments, field of view 2121a may have a horizontal field of view and a vertical field of view. The horizontal field of view may range from 90 degrees to 175 degrees, while the vertical field of view may range from 70 to 125 degrees. Similarly, the field of view associated with camera 2120b, a vertical plane passing through the center of field of view 2121b, may intersect and form an angle with a vertical plane passing through the midline of headset 2100. This angle may also be between 1 and 40 degrees. In some embodiments, camera 2120a and camera 2120b may be angled inward by the same amount. Field of view 2121b may have a horizontal field of view and a vertical field of view consistent with disclosed embodiments. This horizontal field of view may range from 90 degrees to 175 degrees, while the vertical field of view may range from 70 degrees to 125 degrees.

[0233] Camera 2120a and camera 2120b may be configured to provide overlapping views of central field of view 2150. The angular range of central field of view 2150 may be 20 to 80 degrees. For example, the angular range of central field of view 2150 may be approximately 40 degrees (e.g., 40±4 degrees). In addition to central field of view 2150, camera 2120a and camera 2120b may be positioned to provide at least two peripheral fields of view. Peripheral field of view 2160a may be associated with camera 2120a and may include that portion of field of view 2121a that does not overlap with field of view 2121b. In some embodiments, the angular range of peripheral field of view 2160b may be in the range of 40 to 80 degrees. For example, the angular range of peripheral field of view 2160a may be approximately 60 degrees (e.g., 60±6 degrees). Peripheral field of view 2160b may be associated with camera 2120b and may include that portion of field of view 2121b that does not overlap with field of view 2121a. In some embodiments, the angular range of peripheral field of view 2160b may be in the range of 40 to 80 degrees. For example, the angular range of peripheral field of view 2160b may be approximately 60 degrees (e.g., 60±6 degrees).

[0234] Emitters 2130a and 2130b may enable imaging and / or active depth sensing by headset 2100 in low-light conditions. Emitters 2130a and 2130b may be configured to emit light at specific wavelengths. This light can be reflected by physical objects in the physical world around the user. Headset 2100 may be configured with sensors for detecting this reflected light, including image sensors as described herein. In some embodiments, these sensors may be incorporated into at least one of cameras 2120a, 2120b, or 2140. For example, as described above with respect to FIGS. 18A-18C , these cameras may be configured with detectors corresponding to emitters 2130a and / or 2130b. For example, these cameras may include pixels configured to detect light emitted by emitters 2130a and / or 2130b.

[0235] Emitter 2130a and emitter 2130b may be configured to emit IR light consistent with disclosed embodiments. The IR light may have a wavelength between 900 nanometers and 1 micrometer. The IR light may be a 940 nm light source, for example, with the emitted light energy centered around 940 nm. Emitters emitting light of other wavelengths may alternatively or additionally be used. For systems intended for indoor use only, for example, an emitter emitting light centered around 850 nm may be used. At least one of camera 2120a, camera 2120b, or camera 2140a may include one or more IR filters positioned across at least a subset of pixels in the camera's image sensor. The filters may pass light at wavelengths emitted by emitter 2130a and / or emitter 2130b while attenuating light at other wavelengths. For example, the IR filters may be notch filters that pass IR light with wavelengths matching those of the emitters. The notch filter may substantially attenuate other IR light. In some embodiments, the notch filter may be an IR notch filter that blocks IR light and allows light from the emitter to pass. The IR notch filter may also allow light outside the IR band to pass. Such a notch filter may allow the image sensor to receive both visible light and light from the emitter reflected from objects within the field of view of the image sensor. In this manner, a subset of pixels may act as detectors for IR light emitted by emitter 2130a and / or emitter 2130b.

[0236] In some embodiments, a processor of an XR system may selectively enable emitters, such as to enable imaging under low-light conditions. The processor may process image information generated by one or more image sensors and detect whether images output by those image sensors provide adequate information about objects in the physical world without the emitters enabled. The processor may enable an emitter in response to detecting that the images do not provide adequate image information as a result of low ambient light conditions. For example, an emitter may be turned on when stereoscopic information is being used to track an object and a lack of ambient light results in an image with insufficient contrast between features of the tracked object to accurately determine distance using stereoscopic imaging techniques.

[0237] Alternatively, or in addition, emitter 2130a and / or emitter 2130b may be configured for use in performing active depth measurements, such as by emitting light in short pulses. The wearable display system may be configured to perform time-of-flight measurements by detecting reflections of such pulses from objects within emitter 2130a's illumination field 2131a and / or emitter 2130b's illumination field 2131b. These time-of-flight measurements may provide additional depth information for tracking objects or updating a pathable world model. In other embodiments, one or more of the emitters may be configured to emit patterned light, and the XR system may be configured to process images of objects illuminated by the patterned light. Such processing may detect variations in the pattern, which may reveal the distance to the object.

[0238] In some embodiments, the extent of the illumination field associated with the emitter may be sufficient to at least illuminate the field of view of the camera and obtain image information about the object. For example, the emitters may collectively illuminate central field of view 2150. In the illustrated embodiment, emitter 2130a and emitter 2130b may be positioned to collectively illuminate illumination fields 2131a and 2131b, which span the range over which active lighting may be provided. In this exemplary embodiment, two emitters are shown, but it should be understood that more or fewer emitters may be used to span the desired range.

[0239] In some embodiments, emitters such as emitters 2130a and 2130b may be turned off by default but may be enabled when additional illumination is desired to obtain more information than can be obtained using passive imaging. The wearable display system may be configured to enable emitter 2130a and / or emitter 2130b when additional depth information is required. For example, if the wearable display system detects that it cannot obtain adequate depth information for tracking hand or head pose using stereoscopic image information, the wearable display system may be configured to enable emitter 2130a and / or emitter 2130b. Emitter 2130a and / or emitter 2130b may be disabled when additional depth information is not required, thereby reducing power consumption and improving battery life.

[0240] Furthermore, even if the headset is configured with an image sensor configured to detect IR light, it is not a requirement that the IR emitter be mounted on or solely on the headset 2100. In some embodiments, the IR emitter may be an external device disposed within a space, such as a room, in which the headset 2100 may be used. Such an emitter may project IR light, such as in an ArUco pattern, at 940 nm, which is invisible to the human eye. Light with such a pattern may facilitate “instrumentation / assisted tracking,” in which the headset 2100 does not need to provide power to provide the IR pattern but may still provide IR image information as a result of the pattern being presented so that processing performed on that image information can determine the distance to or location of an object in the space. A system with an external illumination source may also allow more devices to operate in the space. When multiple headsets, each moving about the space without a fixed positional relationship, are operating in the same space, there is a risk that light emitted by one headset will be projected onto the image sensor of another headset and thus interfere with its operation. The risk of such interference between headsets may limit the number of headsets that can operate in a space to, for example, 3 or 4. By using one or more IR emitters in the space that illuminate objects that can be imaged by image sensors on the headsets, in some embodiments many more than 10 headsets can operate in the same space without interference.

[0241] As disclosed above with respect to FIG. 3B , camera 2140 may be configured to capture images of the physical world within field of view 2141. Camera 2140 may include an image sensor and a lens. The image sensor may be configured to produce color images. The image sensor may be configured to acquire images with a size between 6 megapixels and 24 megapixels. For example, the image sensor may acquire images with a 4K x 2K resolution (e.g., a horizontal resolution of 3,840 or 4,096 lines and a vertical resolution of 1,716 or 2,160 lines). The image sensor may be configured to acquire images repeatedly or periodically. For example, the image sensor may be configured to acquire images at a frequency between 30 Hz and 120 Hz, such as 60 Hz.

[0242] The image sensor may be a CMOS image sensor. The image sensor may be configured with a rolling shutter. As discussed above with respect to FIGS. 14 and 15 , the rolling shutter may repeatedly read subsets of pixels in the image sensor such that pixels in different subsets reflect light intensity data collected at different times. For example, the image sensor may be configured to read a first row of pixels in the image sensor at a first time and a second row of pixels in the image sensor at a later time. In some embodiments, the sensor may be a CMOS sensor.

[0243] The field of view 2141 of the camera 2140 may include a horizontal field of view and a vertical field of view. The horizontal field of view may extend from 75 to 125 degrees, while the vertical field of view may extend from 60 to 125 degrees.

[0244] IMU 2170a and / or IMU 2170b may be configured to provide acceleration and / or velocity and / or tilt information to the wearable display system. For example, as a user wearing headset 2100 moves, IMU 2170a and / or IMU 2170b may provide information describing the acceleration and / or velocity of the user's head.

[0245] The wearable display system may be coupled to a processor, which may be configured to process image information obtained using the camera, extract information from images captured using the camera as described herein, and / or render virtual objects on a display device. The processor may be mechanically coupled to the frame 2101. Alternatively, the processor may be mechanically coupled to a display device, such as a display device including the monocular 2110a or monocular 2110b. As a further alternative, the processor may be operably coupled to the headset 2100 and / or the display device through a communications link. For example, the XR system may include a local data processing module. This local data processing module may include a processor and may be connected to the headset 2100 or the display device through a physical connection (e.g., a wire or cable) or a wireless (e.g., Bluetooth, Wi-Fi, Zigbee, or equivalent) connection.

[0246] The processor may be configured to perform world reconstruction, head pose tracking, and object tracking operations. For example, the processor may be configured to create a passable world model using camera 2120a, camera 2120b, and camera 2140. In creating the passable world model, the processor may be configured to stereoscopically determine depth information using multiple images of the same physical object acquired by camera 2120a and camera 2120b. As an additional example, the processor may be configured to update an existing passable world model using camera 2120a and camera 2120b, but without camera 2140. As mentioned above, camera 2120a and camera 2120b may be grayscale cameras with relatively lower resolution than color camera 2140. As a result, updating the passable world model using images acquired by camera 2120a and camera 2120b rather than camera 2140 may be performed quickly with reduced power consumption and improved battery life. In some embodiments, the processor may be configured to occasionally or periodically update the passable world model using camera 2120a, camera 2120b, and camera 2140. For example, the processor may be configured to determine that a passable world quality criterion is no longer met, and / or a predetermined time interval has elapsed since the last acquisition and / or use of an image acquired by camera 2140, and / or changes have occurred to objects within the portion of the physical world currently within the field of view of camera 2120a, camera 2120b, and camera 2140.

[0247] According to some embodiments, the XR system may include a hardware accelerator. The hardware accelerator may be implemented as an application-specific integrated circuit (ASIC) or other semiconductor device and may be integrated within or otherwise coupled to the headset 2100 to receive image information from the camera 2120a and the camera 2120b. This hardware accelerator may assist in the stereoscopic determination of depth information using images acquired by the two world cameras 2120a and 2120b. These images may be grayscale images. Using hardware acceleration may accelerate the determination of depth information and reduce power consumption, thus increasing battery life.

[0248] The processor may be configured to perform object tracking within the central field of view 2150 using images from camera 2120a and camera 2120b. In some embodiments, the processor may perform object tracking using depth information determined stereoscopically from a first image acquired by the two cameras. As a non-limiting example, the tracked object may be a hand of a user of the wearable display system. The processor may be configured to perform object tracking within the peripheral field of view using images acquired from one of camera 2120a and camera 2120b.

[0249] Exemplary Calibration Process

[0250] FIG. 22 depicts a simplified flowchart of a calibration routine (method 2200) according to some embodiments. The processor may be configured to perform the calibration routine while the wearable display system is being worn. The calibration routine may account for distortion resulting from the lightweight structure of the headset 2100. In some embodiments, the calibration routine may account for distortion in the frame 2101 that occurs during use due to temperature changes or mechanical distortion of the frame 2101. For example, the processor may perform the calibration routine repeatedly so that the calibration routine compensates for distortion in the frame 2101 during use of the wearable display system. The compensation routine may be performed automatically or in response to manual input (e.g., a user request to perform the calibration routine). The calibration routine may include determining the relative position and orientation of the cameras 2120a and 2120b. The processor may be configured to perform the calibration routine using images acquired by the cameras 2120a and 2120b. In some embodiments, the processor may be further configured to use the outputs of IMU 2170a and IMU 2170b.

[0251] After starting at block 2201, method 2200 may proceed to block 2210. In block 2210, the processor may identify corresponding features in images obtained from camera 2120a and camera 2120b. The corresponding features may be part of an object in the physical world. In some embodiments, the object may have an easily identifiable feature in the image that may be placed in the central field of view 2150 by the user for calibration purposes and may have a predetermined relative position. However, the calibration techniques described herein may be implemented to allow calibration to be repeated during use of headset 2100 based on features on objects present in the central field of view 2150 at the time of calibration. In various embodiments, the processor may be configured to automatically select features detected in both field of view 2121a and field of view 2121b. In some embodiments, the processor can be configured to use the estimated locations of features in field of view 2121a and field of view 2121b to determine correspondence between the features. Such estimation may be based on a passable world model constructed for the object containing these features or other information about the features.

[0252] The method 2200 may proceed to block 2230. At block 2230, the processor may receive inertial measurement data. The inertial measurement data may be received from the IMU 2170a and / or IMU 2170b. The inertial measurement data may include tilt and / or acceleration and / or velocity measurements. In some embodiments, the IMUs 2170a and 2170b may be mechanically coupled, directly or indirectly, to the cameras 2120a and 2120b, respectively. In such embodiments, differences in the inertial measurements, such as tilt, made by the IMUs 2170a and 2170b may indicate differences in the positions and / or orientations of the cameras 2120a and 2120b. Thus, the outputs of the IMUs 2170a and 2170b may provide a basis for making an initial estimate of the relative positions of the cameras 2120a and 2120b.

[0253] After block 2230, method 2200 may proceed to block 2250. In block 2250, the processor may calculate an initial estimate of the relative position and orientation of camera 2120a and camera 2120b. This initial estimate may be calculated using measurements received from IMU 2170a and / or IMU 2170b. In some embodiments, for example, the headset may be designed with a nominal relative position and orientation of camera 2120a and camera 2120b. The processor may be configured to attribute differences in the received measurements between IMU 2170a and IMU 2170b to distortions in frame 2101, which may alter the position and / or orientation of camera 2120a and camera 2120b. For example, IMU 2170a and IMU 2170b may be mechanically coupled, directly or indirectly, to frame 2101 such that tilt and / or acceleration and / or velocity measurements by these sensors have a predetermined relationship. This relationship may be affected when the frame 2101 becomes distorted. As a non-limiting example, IMU 2170a and IMU 2170b may be mechanically coupled to the frame 2101 such that, when there is no distortion of the frame 2101, these sensors measure similar tilt, acceleration, or velocity vectors during movement of the headset. In this non-limiting example, a twist or bend that rotates IMU 2170a relative to IMU 2170b may result in a corresponding rotation of the tilt, acceleration, or velocity vector measurement relative to IMU 2170a relative to the corresponding vector measurement relative to IMU 2170b. The processor may therefore adjust the nominal relative positions and orientations for cameras 2120a and 2120b to match the measured relationship between IMU 2170a and IMU 2170b, because IMU 2170a and IMU 2170b are mechanically coupled to cameras 2120a and 2120b, respectively.

[0254] Other techniques may alternatively or additionally be used to make the initial estimate. In embodiments in which the calibration method 2200 is performed repeatedly during operation of the XR system, the initial estimate may be, for example, the most recently calculated estimate.

[0255] After block 2250, a subprocess begins in which further estimates of the relative position and orientation of cameras 2120a and 2120b are made. One of the estimates is selected as the relative position and orientation of cameras 2120a and 2120b for calculating stereoscopic depth information from images acquired by cameras 2120a and 2120b. The subprocess may be performed iteratively, with further estimates made in each iteration, until an acceptable estimate is identified. In the example of FIG. 22, the subprocess includes blocks 2270, 2272, 2274, and 2290.

[0256] In block 2270, the processor may calculate an error related to the estimated relative orientation of the cameras and the features being compared. In calculating this error, the processor may be configured to estimate how an identified feature should appear or where an identified feature should be located in corresponding images acquired using cameras 2120a and 2120b based on the estimated relative orientations of cameras 2120a and 2120b and the estimated location feature being used for calibration. In some embodiments, this estimate may be compared to the appearance or apparent location of the corresponding feature in images acquired using each of the two cameras to generate an error for each estimated relative orientation. Such an error may be calculated using linear algebraic techniques. For example, the mean squared deviation between the calculated location and the actual location of each of multiple features in the image may be used as a metric for the error.

[0257] After block 2270, method 2200 may proceed to block 2272, where a check may be made as to whether the error meets acceptance criteria. The criteria may be, for example, the overall magnitude of the error or the change in error between iterations. If the error meets the acceptance criteria, method 2200 proceeds to block 2290.

[0258] At block 2290, the processor may select one of the estimated relative orientations based on the error calculated at block 2272. The selected estimated relative position and orientation may be the estimated relative position and orientation with the lowest error. In some embodiments, the processor may be configured to select the estimated relative position and orientation associated with this lowest error as the current relative position and orientation of camera 2120a and camera 2120b. After block 2290, method 2200 may proceed to block 2299. Method 2200 ends at block 2299, and the selected positions and orientations of camera 2120a and camera 2120b may be used to calculate stereoscopic image information based on images formed using the cameras.

[0259] If the error does not meet the acceptance criteria at block 2272, method 2200 may proceed to block 2274. At block 2274, the estimates used in calculating the error at block 2270 may be updated. These updates may be to the estimated relative positions and / or orientations of camera 2120a and camera 2120b. In embodiments in which the relative positions of a set of features being used for calibration are estimated, the updated estimates selected at block 2274 may alternatively or additionally include updates to the positions of feature locations within the set. Such updates may be made according to linear algebra techniques used to solve sets of equations involving multiple variables. As a specific example, one or more of the estimated positions or orientations may be increased or decreased. If the change reduces the calculated error in one iteration of the subprocess, then in a subsequent iteration, the same estimated positions or orientations may be further changed in the same direction. Conversely, if the change increases the error, then in a subsequent iteration, the estimated positions or orientations may be changed in the opposite direction. The estimated positions and orientations of the cameras and features used in the calibration process can be varied in this manner, either sequentially or in combination.

[0260] Once an updated estimate has been calculated, the subprocess returns to block 2270. A further iteration of the subprocess then begins, with calculation of the error for the estimated relative position. In this manner, the estimated position and orientation are updated until an updated relative position and orientation that provides an acceptable error is selected. However, it should be understood that processing at block 2272 may apply other criteria for terminating the iterative subprocess, such as completing a certain number of iterations without finding an acceptable error.

[0261] Although method 2200 is described with reference to camera 2120a and camera 2120b, similar calibration may be performed for any pair of cameras used for stereoscopic imaging, or for any set of cameras for which relative positions and orientations are desired. For example, the position and orientation of camera 2140 may be calculated for one or both of camera 2120a and camera 2120b.

[0262] Example Camera Configurations

[0263] According to some embodiments, components are incorporated into headset 2100 to provide fields of view and lighting to support multiple functions of the XR system. Figures 23A-23C are exemplary diagrams of the fields of view or lighting associated with headset 2100 of Figure 21, according to some embodiments. The exemplary diagrams each depict the fields of view or lighting from different orientations and distances from the headset. Figure 23A depicts the fields of view or lighting from an elevated, off-axis line of sight, at a distance of one meter from the headset. Figure 23A depicts the overlap between the fields of view for cameras 2120a, 2120b, and 2140, particularly how cameras 2120a and 2120b are angled so that fields of view 2121a and 2121b intersect the midline of headset 2100. As depicted, the lighting fields for emitter 2130a and emitter 2130b largely overlap. In this manner, emitter 2130a and emitter 2130b may be configured to support imaging or depth measurement for objects within central field of view 2150 under low ambient light conditions. Figure 23B depicts the field of view or illumination from an up-down eye line at a distance of 0.3 meters from the headset. Figure 23B depicts that overlap between field of view 2121a, field of view 2121b, and field of view 2141 exists at 0.3 meters from the headset. Figure 23C depicts the field of view or illumination from a front-on eye line at a distance of 0.25 meters from the headset. Figure 23C depicts that overlap between field of view 2121a, field of view 2121b, and field of view 2141 exists at 0.25 meters from the headset.

[0264] As can be seen from Figures 23A-23C, the overlap of field of view 2121a, field of view 2121b, and field of view 2141 creates a central field of view in which stereoscopic imaging techniques can be employed using grayscale images acquired with cameras 2120a and 2120b, with or without IR illumination from emitters 2130a and 2130b. In this central field of view, color information from camera 2140 may be combined with the grayscale image information. In addition, there is a peripheral field of view in which there is no overlap, but monocular grayscale image information is available from one of cameras 2120a or 2120b. Different operations may be performed on the image information acquired for the central and peripheral fields, as described herein.

[0265] World Model Generation

[0266] In some embodiments, image information obtained within the central field of view may be used to build or update a world model. FIG. 24 is a simplified flowchart 2400 of a method for creating or updating a passable world model, according to some embodiments. As disclosed above with respect to FIG. 21 , the wearable display system may be configured to determine and update the passable world model using a processor. In some embodiments, the processor may be configured to perform this determination and update based on the output of cameras 2120a and 2120b and 2140 without using emitters 2130a and 2130b. However, in some embodiments, the passable world model may be incomplete. For example, the processor may incompletely determine depth relative to walls or other flat surfaces. As an additional example, the passable world model may incompletely represent objects with many corners, curved surfaces, transparent surfaces, or large surfaces, such as windows, doors, balls, tables, and the like. The processor 2400 may be configured to identify such incomplete information, obtain additional information, and use the additional depth information to update the world model.

[0267] In some embodiments, emitters 2130a and 2130b may be selectively enabled to collect additional image information from which to build or update a passable world model. In some scenarios, the processor may be configured to perform object recognition within the acquired images, select templates for the recognized objects, and add information to the passable world model based on the templates. In this manner, the wearable display system may improve the passable world model with little or no utilization of power-intensive components such as emitters 2130a and 2130b, thereby extending battery life.

[0268] Method 2400 may be initiated one or more times during operation of the wearable display system. The processor may be configured to create a passable world model when the user first moves to a new environment by turning the system on, walking into another room, etc., or generally when the processor detects a change in the user's physical environment. Alternatively, or in addition, method 2400 may be performed periodically during operation of the wearable display system, or when a significant change in the physical world is detected, or in response to a user input, such as an input indicating that the world model is out of sync with the physical world.

[0269] In some embodiments, all or a portion of the passable world model may be stored, provided by other users of the XR system, or otherwise obtained. Thus, while creation of a world model is described, it should be understood that method 2400 may be used for portions of the world model, along with other portions of the world model that are derived from other sources.

[0270] After starting at block 2401, method 2400 may proceed to block 2410. At block 2410, a passable world model may be created. In the illustrated embodiment, the processor may create the passable world model using cameras 2120a and 2120b, along with camera 2140. As described above, in generating the passable world model, the processor may be configured to stereoscopically determine depth information about objects in the physical world using grayscale images obtained from cameras 2120a and 2120b when building the passable world model. In some embodiments, the processor may receive color information from camera 2140. This color information may be used to distinguish objects or to identify surfaces associated with the same object. The color information may also be used to recognize objects.

[0271] After creating the passable world model at block 2410, method 2400 may proceed to block 2430. At block 2430, the processor may identify surfaces and / or objects to use to update the passable world model. The processor may identify such surfaces or objects using grayscale images obtained from camera 2120a and camera 2120b. In some embodiments, the processor may use these grayscale images to stereoscopically determine depth information about objects in the physical world. The depth information may be used when updating the passable world model. Alternatively, or in addition, monocular information may be used to update the world model.

[0272] For example, once a world model showing a surface at a particular location within the passable world is created in block 2410, grayscale images obtained from camera 2120a and / or camera 2120b may be used to detect a surface with approximately identical characteristics and determine that the passable world model should be updated by updating the position of that surface within the passable world model. A surface at approximately the same location with approximately the same shape as a surface in the passable world model may, for example, be considered comparable to that surface in the passable world model, and the passable world model may be updated accordingly. As another example, the position of an object represented in the passable world model may be updated based on grayscale images obtained from camera 2120a and / or camera 2120b.

[0273] The grayscale information may be stereoscopic information or monocular information. In some embodiments, the update may be based on stereoscopic information about an object in the central field of view, where stereoscopic information is available, and on monocular information in the peripheral field of view of one of camera 2120a or camera 2120b, where only monocular information is available. In some embodiments, the update process may be performed differently based on whether the object is in the central or peripheral field of view. For example, the update may be performed for detected surfaces in the central field of view. In the peripheral field of view, the update may be performed only for objects whose models the processor has such that the processor can verify that any updates to the passable world model are consistent with that object. Alternatively, or in addition, new objects or surfaces may be recognized based on processing on the grayscale image. Even if such processing leads to a less accurate representation of the object or surface than the processing in block 2410, trading off accuracy for faster and lower power processing may lead to a better overall system in some scenarios. Additionally, lower accuracy information may be periodically replaced with higher accuracy information by periodically repeating method 2400 to create portions of the world model generated using only a grayscale camera and portions with information generated through the use of a color camera in combination with a grayscale camera.

[0274] After updating the passable world model at block 2430, method 2400 may proceed to block 2450. At block 2450, the processor may identify whether the passable world model includes incomplete depth information. Incomplete depth information can arise in any of several ways. For example, some objects do not produce detectable structures in the image. For example, areas in the physical world that are very dark cannot be imaged with sufficient resolution to extract depth information from images obtained with ambient lighting. As another example, the surface of a window or glass table may not appear or be recognized by computerized processing in the visible image. As yet another example, large uniform surfaces, such as the surface of a table or a wall, may lack sufficient features that can be correlated in two stereoscopic images to enable stereoscopic image processing. As a result, the processor may be unable to determine the location of such objects using stereoscopic processing. In these scenarios, a "hole" exists in the world model, and a process using a passable world model to determine the distance to the surface in a particular direction passing through the "hole" will not be able to obtain any depth information.

[0275] When the passable world model does not include incomplete depth information, the method 2400 may return to updating the passable world model using grayscale images acquired from the cameras 2120a and 2120b.

[0276] Following identification of incomplete depth information, the processor controlling method 2400 may take one or more actions to obtain additional depth information. Method 2400 may proceed to block 2471 and / or block 2473. In block 2471, the processor may enable emitter 2130a and / or emitter 2130b. As disclosed above, one or more of camera 2120a, camera 2120b, or camera 2140 may be configured to detect light emitted by emitter 2130a and / or emitter 2130b. The processor may then obtain depth information by causing emitter 2130a and / or 2130b to emit light that may enhance the obtained images of objects in the physical world. For example, when camera 2120a and camera 2120b are sensitive to the emitted light, images obtained using camera 2120a and camera 2120b may be processed to extract stereoscopic information. Other analysis techniques may alternatively or additionally be used to obtain depth information when emitter 2130a and / or emitter 2130b are enabled. Time-of-flight measurements and / or structured light techniques may alternatively or additionally be used in some embodiments.

[0277] In block 2473, the processor may determine additional depth information from previously obtained depth information. In some embodiments, for example, the processor may be configured to identify objects in images formed using camera 2120a and / or camera 2120b and fill any holes in the passable world model based on models of the identified objects. For example, the process may detect planar surfaces in the physical world. The planar surfaces may be detected using existing depth information obtained using camera 2120a and / or camera 2120b or depth information stored in the passable world model. The planar surfaces may be detected in response to determining that a portion of the world model includes incomplete depth information. The processor may be configured to estimate additional depth information based on the detected planar surfaces. For example, the processor may be configured to extend the identified planar surfaces through areas of incomplete depth information. In some embodiments, the processor may be configured to interpolate missing depth information based on surrounding portions of the passable world model when extending the planar surfaces.

[0278] In some embodiments, as an additional example, the processor may be configured to detect an object within a portion of the world model that includes incomplete depth information. In some embodiments, this detection may involve recognizing the object using a neural network or other machine learning tool. In some embodiments, the processor may be configured to access a database of stored templates and select an object template that corresponds to the identified object. For example, when the identified object is a window, the processor may be configured to access the database of stored templates and select a corresponding window template. As a non-limiting example, the template may be a three-dimensional model that represents a class of object, such as a type of window, door, bowl, or the like. The processor may construct an instance of the object template based on an image of the object in the updated world model. For example, the processor may scale, rotate, and translate the template to match the detected location of the object in the updated world model. Additional depth information may then be estimated based on the boundary of the constructed template, which represents the surface of the recognized object.

[0279] After block 2471 and / or block 2473, method 2400 may proceed to block 2490. In block 2490, the processor may update the passable world model using the additional depth information obtained in block 2471 and / or block 2473. For example, the processor may be configured to blend additional depth information obtained from measurements made using active IR illumination into the existing passable world model. As additional examples, the processor may be configured to blend interpolated depth information obtained by extending a detected planar surface into the existing passable world model, or to blend additional depth information estimated from the boundary of a configured template into the existing passable world model. Information may be blended in one or more ways depending on the nature of the additional depth information and / or the information in the passable world model. Blending may be performed, for example, by adding to the passable world model additional depth information collected about locations where holes exist in the passable world model. Alternatively, the additional depth information may overwrite information at a corresponding location in the passable world model. As yet another alternative, blending may involve selecting between information already in the passable world model and the additional depth information. Such a selection may be based, for example, on selecting depth information either already in the passable world model or in the additional depth information that represents a surface closest to the camera being used to collect the additional depth information.

[0280] In some embodiments, a passable world model may be represented by a mesh of connected points. Updating the world model may be done by computing a mesh representation of the object or surface to be added to the world model and then combining that mesh representation with the mesh representation of the world model. The inventors have recognized and appreciated that performing the processing in this order may require less processing than adding the object or surface to the world model and then computing a mesh for the updated model.

[0281] 24 shows the world model updated at both blocks 2430 and 2490. The processing at each block may be performed in the same manner or in different manners, for example, by generating a mesh representation of the object or surface to be added to the world model and combining the generated mesh with the mesh of the world model. In some embodiments, this merging operation may be performed once for both the object or surface identified at block 2430 and block 2490. Such combined processing may be performed, for example, as described in connection with block 2490.

[0282] In some embodiments, method 2400 may return to block 2430 to repeat the process of updating the world model based on information obtained using one or more grayscale cameras. The processing at block 2430 may be performed on fewer and smaller images than the processing at block 2410 and may therefore be repeated at a higher rate. This processing may be performed at a rate less than 10 times per second, such as 3-7 times per second.

[0283] Method 2400 may be repeated in this manner until an exit condition is detected. For example, method 2400 may be repeated for a predetermined period of time, until user input is received, or until a particular type or magnitude of change in the portion of the physical world model within the field of view of the camera of headset 2100 is detected. Method 2400 may then end at block 2499. Method 2400 may be restarted so that new information for the world model is captured at block 2410, including that obtained using a higher resolution color camera. Method 2400 may be terminated and restarted to repeat the process at block 2401 to create a portion of the world model using the color camera at a slower average rate than the rate at which the world model is updated based solely on grayscale image information. The process using the color camera may be repeated, for example, once per second or at a slower average rate.

[0284] Object Tracking

[0285] As described above, the processor of an XR system may track objects in the physical world to support realistic rendering of virtual objects relative to the physical objects. Tracking has been described with reference to movable objects, such as the hands of a user of an XR system. Rapidly updating the position of movable objects enables realistic rendering of virtual objects because rendering can reflect occlusion of physical objects by virtual objects or vice versa or interactions between virtual objects within physical objects. Tracking of fixed objects in the physical world has also been described as part of head pose tracking. Determining the user's head pose allows information in the passable world model to be transformed into the frame of reference of the user's wearable display device so that the information in the passable world model can be used in rendering objects on the wearable display device.

[0286] According to some embodiments, grayscale stereoscopic information may be used to track objects in the central field of view 2150. Such objects may not be tracked when they are in the peripheral fields of view 2160a or 2160b. Alternatively, or in addition, a different tracking approach that uses only monocular image information may be used for tracking in these peripheral fields of view.

[0287] Hand tracking may be used as an example of object tracking. FIG. 25 is a simplified flowchart of a method of hand tracking (method 2500) according to some embodiments. In various embodiments, the processor may be configured to perform hand tracking using camera 2120a and camera 2120b when the user's hand is in the central field of view 2150, using only camera 2120a when the user's hand is in the peripheral field of view 2120b, and using only camera 2120b when the user's hand is in the peripheral field of view 2120b. In this configuration, the wearable display system may be configured to provide adequate hand tracking using a reduced number of available cameras, allowing for reduced power consumption and increased battery life.

[0288] Method 2500 may be implemented under control of a processor of an XR system. The method may be initiated in response to the detection of an object to be tracked, such as a hand, as a result of analysis of images acquired using any one of the cameras on headset 2100. The analysis may involve recognizing the object as a hand based on regions of the image having photometric properties characteristic of a hand. Alternatively, or in addition, depth information acquired based on stereoscopic image analysis may be used to detect the hand. As a specific example, the depth information may indicate the presence of an object having a shape that matches a 3D model of a hand. Detecting the presence of a hand in this manner may also involve setting parameters of the hand model to match the orientation of the hand. In some embodiments, such a model may also be used for high-speed hand tracking by using photometric information from one or more grayscale cameras to determine how the hand has moved from its original position.

[0289] Other trigger conditions may initiate method 2500, such as the XR system performing an action that involves tracking an object, such as rendering a virtual button that the user is likely to attempt to press, such that the user's hand is expected to enter the field of view of one or more cameras. Method 2500 may be repeated at a relatively high rate, such as 30-100 times per second, e.g., 40-60 times per second. As a result, updated position information about the tracked object may be made available with low latency for processing to render virtual objects that interact with the physical object.

[0290] After starting at block 2501, method 2500 may proceed to block 2510. In block 2510, the processor may determine potential hand locations. In some embodiments, the potential hand locations may be locations of detected objects within the acquired image. In embodiments where hands are detected based on matching depth information to a 3D model of a hand, that same information may be used as the initial hand position in block 2510.

[0291] After block 2510, method 2500 may proceed to block 2520. In block 2520, the processor may determine whether the object is in the central or peripheral field of view. When the object is in the central field of view, method 2500 may proceed to block 2531. In block 2531, the processor may obtain depth information about the object. The depth information may be obtained based on stereoscopic image analysis, from which a distance between a camera collecting image information from the camera and the tracked object may be calculated. The processor may, for example, select a feature in the central field of view and determine depth information about the selected feature. The processor may stereoscopically determine depth information about the feature using images obtained by cameras 2120a and 2120b.

[0292] In some embodiments, the selected features may represent different segments of a human hand as defined by bones and joints. Feature selection may be based on matching image information to a model of a human hand. Such matching may be performed heuristically, for example. A human hand may be represented by a finite number of segments, such as 16, and points in an image of the hand may be mapped to one of those segments so that features on each segment can be selected. Alternatively, or in addition, such matching may use a deep neural net or a classification / decision forest to apply a series of yes / no decisions in the analysis to identify different parts of the hand and select features that represent different parts of the hand. The matching may, for example, identify whether a particular point in the image belongs to the palm area, back of the hand, non-thumb fingers, thumb, fingertips, and / or knuckles. Any suitable classifier may be used for this analysis stage. For example, a deep learning module or neural network mechanism may be used instead of or in addition to a classification forest. Additionally, a regression forest (e.g., using a Hough transform, etc.) may be used in addition to a classification forest.

[0293] Regardless of the specific number of features selected and the technique used to select those features, after block 2531, method 2500 may proceed to block 2550. In block 2550, the processor may then construct a hand model based on the depth information. In some embodiments, the hand model may reflect structural information about a human hand, for example, representing each of the bones in the hand as segments within the hand, and each joint defining a range of likely angles between adjacent segments. By assigning a location to each of the segments in the hand model based on the depth information of the selected features, information about the position of the hand may be provided for subsequent processing by the XR system.

[0294] In some embodiments, the processing at blocks 2531 and 2550 may be performed iteratively, whereby feature selection for which depth information is collected is refined based on the configuration of the hand model. The hand model may include shape and motion constraints that the processor may be configured to use to refine the selection of features representing the hand components. For example, when a feature selected to represent a hand segment indicates a position or motion of that section that violates the constraints of the hand model, a different feature may be selected to represent that segment.

[0295] In some embodiments, successive iterations of the hand tracking process may be performed using photometric image information instead of or in addition to depth information. At each iteration, the 3D model of the hand may be updated to reflect potential movement of the hand. Potential movement may be determined from depth information, photometric information, a projection of the hand trajectory, or the like. If depth information is used, it may relate to a more limited set of features than those used to establish the initial configuration of the hand model in order to speed up processing.

[0296] Regardless of how the 3D hand model is updated, the updated model may be refined based on photometric image information. The model may be used, for example, to render a virtual image of the hand, representing how the image of the hand is expected to appear. The expected image may be compared to photometric image information obtained using an image sensor. The 3D model may then be adjusted to reduce the error between the expected and obtained photometric information. The adjusted 3D model then provides an indication of the position of the hand. As this update process is repeated, the 3D model provides an indication of the position of the hand as it moves.

[0297] In contrast, if processing at block 2520 determines that the tracked object is within peripheral vision, depth information from stereoscopic image analysis techniques may not be available. Nevertheless, in processing at block 2535, the processor may select features in the image that represent the structure of a human hand. Such features may be identified heuristically or using AI techniques, as described above, for processing at block 2531.

[0298] In block 2555, the processor may then attempt to match the selected features and their movement per image to a hand model without the benefit of depth information. This match may result in less robust information or may be less accurate than that generated in block 2550. Nevertheless, the information identified based on the monocular information may provide useful information regarding the operation of the XR system. The determined information may include, for example, recognition of hand movements corresponding to hand gestures. This gesture recognition may be performed using the hand tracking method described in U.S. Patent Publication No. 2016 / 0026253 (which teaches related to hand tracking and the use of information about hands obtained from image information in an XR system and is incorporated herein by reference in its entirety).

[0299] After matching the image portion to the portion of the hand model at block 2550, method 2500 may end at block 2599. However, it should be understood that object tracking may occur continuously during operation of the XR system, or may occur during an interval during which the object is within the field of view of one or more cameras. Thus, once one iteration of method 2500 is completed, another iteration may be performed, and the process may be performed over the interval during which object tracking is performed. In some embodiments, information used in one iteration may be used in subsequent iterations. In various embodiments, for example, the processor may be configured to estimate an updated location of the user's hand based on a previously detected hand location. For example, the processor may estimate the next likely location of the user's hand based on the previous location and the velocity of the user's hand. Such information may be used to narrow down the amount of image information processed to detect the object location, as described above in connection with patch tracking techniques.

[0300] Having thus described several aspects of several embodiments, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art.

[0301] As an example, embodiments are described in relation to an augmented reality (AR) environment, although it should be understood that some or all of the techniques described herein may also be applied in mixed reality (MR) environments, or more generally, other XR environments.

[0302] Also described are image array embodiments in which one patch is applied to the image array to control the selective output of image information for one movable object. It should be understood that there may be more than one movable object in a physical embodiment. Furthermore, in some embodiments, it may be desirable to selectively obtain frequent updates of image information in areas other than where the movable object is located. For example, a patch may be configured to selectively obtain image information for an area of ​​the physical world where a virtual object is to be rendered. Thus, some image sensors may be capable of selectively providing information about two or more patches, with or without circuitry for tracking the trajectories of those patches.

[0303] As a still further example, the image array is described as outputting information related to the magnitude of incident light. The magnitude may be a representation of power across a spectrum of light frequencies. The spectrum may be relatively broad, capturing energy at frequencies corresponding to any color of visible light, such as in a black-and-white camera. Alternatively, the spectrum may be narrow, corresponding to a single color of visible light. Filters to limit the light incident on the image array to light of a particular color may be used for this purpose. When pixels are limited to receive light of a specific color, different pixels may be limited to different colors. In such an embodiment, the outputs of pixels sensitive to the same color may be processed together.

[0304] A process for setting a patch in an image array and then updating the patch for an object of interest has been described. This process may be performed for each movable object, for example, as it enters the field of view of the image sensor. The patch may be released when the object of interest exits the field of view so that the patch is no longer tracked or image information is not output for the patch. It should be understood that from time to time, a patch may be updated, such as by determining the location of the object associated with the patch and setting the patch's position to correspond to that location. Similar adjustments can be made to the calculated trajectory of the patch. A motion vector for the object and / or a motion vector of the image sensor may be calculated from other sensor information and used to reset values ​​programmed into the image sensor or other components for patch tracking.

[0305] For example, the location, motion, and other characteristics of an object may be determined by analyzing the output of a wide-angle video camera or a pair of video cameras with stereoscopic information. Data from these other sensors may be used to update the world model. In connection with the updates, patch position and / or trajectory information may be updated. Such updates may occur at a slower rate than the patch positions are updated by the patch tracking engine. The patch tracking engine may calculate new patch positions, for example, at a rate of about 1 to 30 times per second. Patch position updates based on other information may occur at a slower rate, such as once per second to about once per 30-second interval.

[0306] As yet a further example of a variation, FIG. 2 shows a system with a head-mounted display separate from the remote processing module. Image sensors as described herein can lead to a compact design of the system. Such sensors generate less data, which in turn leads to lower processing requirements and less power consumption. The need for less processing and power allows for size reduction, such as by reducing the size of the battery. Thus, in some embodiments, the entire augmented reality system may be integrated within the head-mounted display without a remote processing module. The head-mounted display may be configured as a pair of goggles or may be similar in size and shape to a pair of eyeglasses, as shown in FIG. 2.

[0307] Additionally, embodiments are described in which the image sensor is responsive to visible light. It should be understood that the techniques described herein are not limited to operation with visible light. They may alternatively or additionally respond to "light" in other parts of the spectrum, such as IR light or UV. Furthermore, image sensors as described herein respond to naturally occurring light. Alternatively, or additionally, the sensor may be used in a system with an illumination source. In some embodiments, the sensitivity of the image sensor may be tuned to the portion of the spectrum in which the illumination source emits light.

[0308] As another example, a selected region of an image array where changes should be output from the image sensor is described as being defined by defining a "patch" on which image analysis is to be performed. However, it should be understood that the patch and selected region may be of different sizes. The selected region may be larger than the patch, for example, to account for movement of an object in the image being tracked that deviates from a predicted trajectory and / or to allow processing around the edges of the patch.

[0309] Such alterations, modifications, and improvements are intended to be part of this disclosure, and are intended to be within the spirit and scope of this disclosure.

[0310] For example, in some embodiments, the color filter 102 of a pixel of the image sensor may not be a separate component, but instead may be incorporated into one of the other components of the pixel sub-array 100. For example, in an embodiment including a single pixel with both an angle-of-arrival / position-intensity converter and a color filter, the angle-of-arrival / intensity converter may be a transmissive optical component formed from a material that filters certain wavelengths.

[0311] According to some embodiments, a wearable display system is provided, the wearable display system may include: two first cameras mechanically coupled to the frame to provide a central field of view associated with both the frame and the cameras and a first peripheral field of view associated with a first of the two first cameras; a color camera mechanically coupled to the frame to provide a color field of view overlapping the central field of view; and a processor operably coupled to the two first cameras and the color camera and configured to track hand movement within the central field of view using stereoscopically determined depth information from first images acquired by the two first cameras, track hand movement within the first peripheral field of view using one or more second images acquired by a first of the two first cameras, create a world model using the two first cameras and the color camera, and update the world model using the two first cameras.

[0312] In some embodiments, the two first cameras may be configured with a global shutter. In some embodiments, the wearable display system may further include a hardware accelerator for stereoscopically determining depth information using the first grayscale images acquired by the two first cameras. In some embodiments, the two first cameras may have equidistant lenses. In some embodiments, the two first cameras may each have a horizontal field of view between 90 degrees and 175 degrees. In some embodiments, the central field of view may have an angular range between 40 and 80 degrees. In some embodiments, the processor may further be configured to perform a calibration routine to determine a relative orientation of the two first cameras.

[0313] In some embodiments, the calibration routine may include the steps of identifying corresponding features in images acquired using each of the two first cameras; calculating an error for each of a plurality of estimated relative orientations of the two first cameras, the error indicating the difference between the corresponding feature as it appears in the images acquired using each of the two first cameras and an estimate of the identified feature calculated based on the estimated relative orientations of the two first cameras; and selecting a relative orientation from the plurality of estimated relative orientations as the determined relative orientation based on the calculated error.

[0314] In some embodiments, the wearable display system may further include a first inertial measurement unit mechanically coupled to a first of the two first cameras and a second inertial measurement unit mechanically coupled to a second of the two first cameras, and the calibration routine may further include selecting at least one of a plurality of estimated relative orientations based in part on outputs of the first inertial measurement unit and the second inertial measurement unit.

[0315] In some embodiments, the processor may be further configured to repeatedly perform the calibration routine while the wearable display system is worn, such that the calibration routine compensates for distortion in the frame during use of the wearable display system. In some embodiments, the calibration routine may compensate for distortion in the frame caused by changes in temperature. In some embodiments, the calibration routine may compensate for distortion in the frame caused by mechanical distortion. In some embodiments, the two first cameras may be configured to obtain grayscale images. In some embodiments, the wearable display system may have one color camera. In some embodiments, the processor may be mechanically coupled to the frame. In some embodiments, the display device may be mechanically coupled to the frame, and the display device comprises the processor. In some embodiments, the local data processing module may comprise the processor, and the local data processing module may be operably coupled to the display device through a communications link, and the display device may be mechanically coupled to the frame. In some embodiments, the first image may comprise one or more second images.

[0316] According to some embodiments, a wearable display system is provided, which may include a frame, two first cameras mechanically coupled to the frame, a color camera mechanically coupled to the frame, and a processor operably coupled to the two first cameras and the color camera and configured to create a world model using a first grayscale image acquired by the two first cameras and one or more color images acquired by the color camera, update the world model using a second grayscale image acquired by the two first cameras, determine portions of the world model that include incomplete depth information, and update the world model with additional depth information for portions of the world model that include incomplete depth information.

[0317] In some embodiments, the wearable display system may further include one or more emitters, and additional depth information may be obtained using the one or more emitters, and the processor may be further configured to enable the one or more emitters in response to determining that the portion of the world model includes incomplete depth information. In some embodiments, the one or more emitters may include an infrared emitter, and the two first cameras may include filters configured to pass infrared light. In some embodiments, the infrared emitter may emit light having a wavelength between 900 nanometers and 1 micrometer.

[0318] In some embodiments, the processor may be further configured to detect a planar surface in the physical world in response to determining that the portion of the world model includes incomplete depth information, and may estimate additional depth information based on the detected planar surface. In some embodiments, the processor may be further configured to detect an object in the portion of the world model including incomplete depth information, identify an object template corresponding to the detected object, construct an instance of the object template based on an image of the object in the updated world model, and estimate additional depth information based on the constructed instance of the object template. In some embodiments, the two first cameras may be mechanically coupled to the frame to provide a central field of view associated with both of the first cameras and a peripheral field of view associated with a first of the two first cameras, and the processor may be further configured to track hand movements within the central field of view using depth information determined from images obtained by the two first cameras.

[0319] In some embodiments, tracking hand movement within the central visual field using depth information may include selecting a point within the central visual field, stereoscopically determining depth information for the selected point using images acquired by the two first cameras, generating a depth map using the stereoscopically determined depth information, and matching portions of the depth map to corresponding portions of a hand model, including both shape and movement constraints. In some embodiments, the processor may be further configured to track hand movement within the peripheral visual field using one or more images acquired from a first of the two first cameras by matching portions of the images to corresponding portions of a hand model, including both shape and movement constraints. In some embodiments, the processor may be mechanically coupled to the frame. In some embodiments, a display device mechanically coupled to the frame may comprise the processor. In some embodiments, a local data processing module may comprise the processor, and the local data processing module may be operably coupled to the display device through a communications link, and the display device may be mechanically coupled to the frame.

[0320] According to some embodiments, a wearable display system is provided, the wearable display system may include a frame; two grayscale cameras mechanically coupled to the frame, the two grayscale cameras being a first grayscale camera having a first field of view and a second grayscale camera having a second field of view, the first grayscale camera and the second grayscale camera being positioned to provide a central field of view where the first field of view overlaps the second field of view, and a first peripheral field of view that is within the first field of view and outside the second field of view; and a color camera mechanically coupled to the frame to provide a color field of view that overlaps the central field of view.

[0321] In some embodiments, the two grayscale cameras may have a global shutter. In some embodiments, the color camera may have a rolling shutter. In some embodiments, the wearable display system may include two inertial measurement units mechanically coupled to the frame and one or more emitters mechanically coupled to the frame. In some embodiments, the one or more emitters may be infrared emitters, and the two grayscale cameras may include a filter configured to pass infrared light. In some embodiments, the infrared emitter may emit light having a wavelength between 900 nanometers and 1 micrometer. In some embodiments, the illumination field of the one or more emitters may overlap with the central field of view. In some embodiments, a first inertial measurement unit of the two inertial measurement units may be mechanically coupled to a first grayscale camera, and a second inertial measurement unit of the two inertial measurement units may be mechanically coupled to a second grayscale camera of the two grayscale cameras. In some embodiments, the two inertial measurement units may be configured to measure tilt, acceleration, velocity, or any combination thereof. In some embodiments, the two grayscale cameras may have a horizontal field of view between 90 degrees and 175 degrees. In some embodiments, the central field of view may have an angular range between 40 and 80 degrees.

[0322] Furthermore, while advantages of the present disclosure are illustrated, it should be understood that not all embodiments of the present disclosure include all described advantages. Some embodiments may not implement any feature described herein as advantageous. Thus, the foregoing description and drawings are by way of example only.

[0323] The foregoing embodiments of the present disclosure can be implemented in any of numerous ways. For example, embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. Such a processor may be implemented as an integrated circuit, with one or more processors in an integrated circuit component, including commercially available integrated circuit components known in the art, such as a CPU chip, a GPU chip, a microprocessor, a microcontroller, or a coprocessor, to name a few. In some embodiments, the processor may be implemented in a custom circuit, such as an ASIC, or in a semi-custom circuit resulting from configuring a programmable logic device. As a further alternative, the processor may be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores, such that one or a subset of those cores may constitute a processor. However, a processor may be implemented using circuitry in any suitable format.

[0324] Further, it should be understood that a computer may be embodied in any of several forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer, etc. Additionally, a computer may be embodied in devices not generally considered computers but with suitable processing capabilities, including a personal digital assistant (PDA), a smartphone, or any suitable portable or fixed electronic device.

[0325] A computer may also have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or display screen for visual presentation of output, or a speaker or other sound-generating device for audible presentation of output. Examples of input devices that can be used for a user interface include a keyboard and pointing devices such as a mouse, touchpad, and digitizing tablet. As another example, a computer may receive input information through speech recognition or in other audible formats. In the illustrated embodiment, the input / output devices are illustrated as physically separate from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated within the same unit as the processor or other elements of the computing device. For example, a keyboard may be implemented as a soft keyboard on a touchscreen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated through a wireless connection.

[0326] Such computers may be interconnected by one or more networks in any suitable form, including as a local area network or a wide area network, such as an enterprise network or the Internet. Such networks may be based on any suitable technology and operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.

[0327] The various methods and processes outlined herein may also be coded as software that is executable on one or more processors employing any one of a variety of operating systems or platforms. In addition, such software may be written using any of a number of suitable programming languages ​​and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code that runs on a framework or virtual machine.

[0328] In this aspect, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., computer memory, one or more diskettes, compact disks (CDs), optical disks, digital video disks (DVDs), magnetic tape, flash memory, circuitry in a field programmable gate array or other semiconductor device, or other tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement various embodiments of the present disclosure discussed above. As is evident from the foregoing examples, a computer-readable storage medium may retain information for a time sufficient to provide computer-executable instructions in a non-transitory form. Such a computer-readable storage medium or media may be transportable, as described above, such that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the present disclosure. As used herein, the term "computer-readable storage medium" encompasses only computer-readable media that can be considered a manufacture (i.e., an article of manufacture) or machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagated signal.

[0329] The terms "program" or "software" are used herein in a generic sense to refer to any type of computer code or set of computer-executable instructions that can be employed to program a computer or other processor to implement various aspects of the present disclosure, as described above. Additionally, according to one aspect of the present embodiments, it should be understood that one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but can be distributed in a modular manner among several different computers or processors to implement various aspects of the present disclosure.

[0330] Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.

[0331] Also, data structures may be stored in a computer-readable medium in any suitable form. For ease of illustration, data structures may be shown to have fields that are related through their locations within the data structure. Such relationships may also be achieved by allocating storage for the fields with locations within the computer-readable medium that convey the relationship between the fields. However, any suitable mechanism may be used to establish relationships between information within fields of a data structure, including through the use of pointers, tags, or other mechanisms that establish relationships between data elements.

[0332] Various aspects of the present disclosure may be used alone, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and therefore, its application is not limited to the details and arrangements of components set forth in the foregoing description or illustrated in the drawings. For example, aspects described in one embodiment may be combined in any manner with aspects described in other embodiments.

[0333] The present disclosure may also be embodied as a method, examples of which are provided. The acts performed as part of the method may be ordered in any suitable manner. Thus, while the illustrative embodiments are shown as sequential acts, embodiments may be constructed in which acts are performed in an order different from that shown, which may include performing some acts simultaneously.

[0334] The use of ordinal terms such as "first," "second," "third," etc. in the claims to modify claim elements does not, by itself, imply any priority, precedence, or ordering of one claim element relative to another element, or the chronological order in which acts of a method are performed, but rather the ordinal terms are used solely as markers to distinguish one claim element having a certain name from another element having the same name (due to the use of the ordinal terms) to distinguish between claim elements.

[0335] Also, the phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use herein of "including," "comprising," "having," "containing," "with," and variations thereof, is meant to encompass the items listed thereafter and equivalents and additional items.

Claims

1. 1. A wearable display system, comprising: The frame and a first camera and a second camera, the first camera and the second camera comprising: a central field of view associated with both cameras; a peripheral field of view associated with each of the first camera and the second camera; a first camera and a second camera mechanically coupled to the frame to provide a processor operably coupled to the first camera and the second camera, the processor comprising: Determining a potential location of the object; when the potential location of the object is determined to be in the central field of view, tracking object movement within the central field of view using stereoscopically determined depth information from images acquired by the first and second cameras; when the potential location of the object is determined to be in the peripheral field of view of only one of the first camera or the second camera, tracking the movement of the object within the peripheral field of view using at least one image acquired by only the first camera or the second camera; a processor configured to execute A wearable display system comprising:

2. The processor: determining whether the object is within the central or peripheral visual field; selecting a feature in the at least one image that represents a structure of the object when the object is determined to be within the peripheral vision; The wearable display system of claim 1 , further configured to:

3. The wearable display system of claim 1 , further comprising a color camera mechanically coupled to the frame to provide a color field of view overlapping the central field of view.

4. The processor: creating a world model representing objects in the physical world, including the object, using stereoscopically determined depth information from images acquired by the first camera and the second camera and color information obtained from the color camera; updating the world model using the depth information determined stereoscopically from images acquired by the first camera and the second camera; and The wearable display system of claim 3 , further configured to perform:

5. The wearable display system of claim 4 , wherein updating the world model includes updating the world model 30 to 100 times per second.

6. The wearable display system of claim 1 , wherein the processor is further configured to recognize an object as a human body part based on a region of at least one image having photometric properties that are characteristic of the human body part.

7. The wearable display system of claim 1 , wherein tracking the movement of the object within the central visual field further comprises creating a model representing the object.

8. Tracking the movement of the object within the central visual field includes: The wearable display system of claim 7 , further comprising: using depth information of selected features of the object to assign a location to each respective segment of a plurality of segments in the model.

9. Tracking the movement of the object within the peripheral vision includes: identifying selected features of the object in a first image acquired by the first camera or the second camera; identifying selected features of the object in a second image acquired by the same camera that acquired the first image; The wearable display system of claim 1 , further comprising:

10. 10. The wearable display system of claim 9, wherein tracking the movement of the object within the peripheral vision further comprises matching the selected feature of the object in the first image to the selected feature of the object in the second image.

11. The wearable display system of claim 10 , wherein the matching can include recognizing a motion of the object.

12. 1. A wearable display system, comprising: The frame and two grayscale cameras mechanically coupled to the frame, the two grayscale cameras comprising a first grayscale camera having a first field of view and a second grayscale camera having a second field of view, the first grayscale camera and the second grayscale camera comprising: a central field of view where the first field of view overlaps with the second field of view; a first peripheral field of view within the first field of view and outside the second field of view; two grayscale cameras positioned to provide a processor operatively coupled to the two grayscale cameras, the processor comprising: Determining a potential location of the object; when the potential location of the object is determined to be in the central field of view, tracking object movement within the central field of view using stereoscopically determined depth information from images acquired by the two grayscale cameras; when the potential location of the object is determined to be in the peripheral field of view of the first grayscale camera, tracking the movement of the object within the peripheral field of view using at least one image acquired by the first grayscale camera; a processor configured to execute A wearable display system comprising:

13. a first emitter having a third field of view mechanically coupled to the frame; a second emitter mechanically coupled to the frame, the second emitter having a fourth field of view; The wearable display system of claim 12 further comprising:

14. The wearable display system of claim 13 , wherein the third field of view overlaps with the fourth field of view.

15. The wearable display system of claim 13 , wherein the third field of view overlaps with the central field of view and the fourth field of view overlaps with the central field of view.