Lightweight cross-display device with passive depth extraction

Through the lightweight head-mounted device design and sensor combination, combined with grayscale cameras and all-optical cameras, the weight and power consumption problems of wearable XR display systems are solved, low latency and high accuracy information acquisition is achieved, and user experience and system practicality are improved.

CN113711587BActive Publication Date: 2025-08-08MAGIC LEAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080026110.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-02-07
Filing Date
2020-02-07
Publication Date
2025-08-08
Estimated Expiration
2040-02-07

AI Technical Summary

Technical Problem

The existing wearable XR display system is too heavy, causing users to be tired and distracted. At the same time, high-power sensors limit the practicality and user experience of the system, making it difficult to accurately obtain information about the physical world under low latency and low power consumption.

Method used

Designed with a lightweight head-mounted device, combined with a grayscale camera and an all-optical camera, reduce the number of sensors by repeating the appropriate combination of calibration routines and image sensors, use low-power sensors to perform compensation and size reduction routines, reduce power consumption and improve the accuracy of information acquisition.

Benefits of technology

A lightweight wearable XR display system is realized, reducing power consumption, extending battery life, and accurately obtaining information from the physical world at low latency, improving the immersion of the user experience and the practicality of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113711587B_ABST
    Figure CN113711587B_ABST
Patent Text Reader

Abstract

A wearable display system including multiple cameras and a processor is disclosed. A grayscale camera and a color camera can be arranged to provide a central field of view associated with the two cameras and a peripheral field of view associated with one of the two cameras. One or more of the two cameras can be plenoptic cameras. The wearable display system can acquire light field information using at least one plenoptic camera and create a world model using the first light field information and first depth information stereoscopically determined from images acquired by the grayscale camera and the color camera. The wearable display system can track head pose using the at least one plenoptic camera and the world model. When the object meets a depth criterion, the wearable display system can track the object in the central field of view and the peripheral field of view using one or both plenoptic cameras.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application generally relates to a wearable cross-reality display system including at least one plenoptic camera. Background Art

[0002] A computer can control a human user interface to create an X Reality (XR or cross-reality) environment in which some or all of the XR environments perceived by the user are generated by the computer. These XR environments can be virtual reality (VR), augmented reality (AR), or mixed reality (MR) environments, some or all of which can be generated by a computer in part using data describing the environment. For example, the data can describe a virtual object that can be rendered in a way that the user feels or perceives it as part of the physical world, so that the user can interact with the virtual object. Because the data is rendered and presented by a user interface device (such as a head-mounted display device, for example), the user can experience these virtual objects. The data can be displayed to the user, or can control audio played to the user, or can control a tactile (or haptic) interface, allowing the user to experience the touch sensation that the user feels or perceives as feeling the virtual object.

[0003] XR systems can be used for a wide range of applications across scientific visualization, medical training, engineering design and prototyping, telemanipulation and telepresence, and personal entertainment. Compared to VR, AR and MR involve one or more virtual objects that are associated with real objects in the physical world. The experience of virtual objects interacting with real objects significantly enhances the user's enjoyment of using XR systems and also opens the door to a variety of applications that present realistic and easily understood information about how to modify the physical world. Summary of the Invention

[0004] Aspects of the present application relate to a wearable cross-reality display system comprising at least one plenoptic camera.The techniques described herein can be used together, separately, or in any suitable combination.

[0005] According to some embodiments, a wearable display system is provided, comprising: a head-mounted device including a first camera with a global shutter and a second camera with a rolling shutter, the first camera and the second camera positioned to provide overlapping views of a central field of view; and a processor operably coupled to the first camera and the second camera and configured to: execute a compensation routine to adjust an image acquired using the second camera for rolling shutter image distortion; and create a world model in part using depth information stereoscopically determined based on the image acquired using the first camera and the adjusted image.

[0006] In some embodiments, the first camera and the second camera may be angled inward asymmetrically.

[0007] In some embodiments, the field of view of the first camera may be larger than the field of view of the second camera.

[0008] In some embodiments, the first camera may be angled inward between 20 and 40 degrees, and the second camera may be angled inward between 1 and 20 degrees.

[0009] In some embodiments, the first camera may have an angular pixel resolution between 1 arc minute and 5 arc minutes per pixel.

[0010] In some embodiments, the processor may be further configured to perform a downsizing routine to resize images acquired using the second camera.

[0011] In some embodiments, the downsizing routine may include generating a downsized image by binning pixels in an image acquired by the second camera.

[0012] In some embodiments, the compensation routine may include: comparing a first image acquired using the first camera with a second image acquired using the second camera to detect a skew of at least a portion of the second image; and adjusting the at least a portion of the second image to compensate for the detected skew.

[0013] In some embodiments, comparing a first image acquired using the first camera with a second image acquired using the second camera may include performing a line-by-line comparison between the first image acquired by the first camera and the second image acquired by the second camera.

[0014] In some embodiments, the processor may be further configured to disable the second camera or modulate a frame rate of the second camera based on at least one of a power conservation criterion or a world model integrity criterion.

[0015] In some embodiments, at least one of the first camera or the second camera may include a plenoptic camera; and the processor may be further configured to create a world model using, in part, light field information acquired by at least one of the first camera or the second camera.

[0016] In some embodiments, the first camera may include a plenoptic camera; and the processor may be further configured to perform a world model update routine using depth information acquired using the plenoptic camera.

[0017] In some embodiments, the processor may be mechanically coupled to the head-mounted device.

[0018] In some embodiments, the head mounted device may include a display device mechanically coupled to the processor.

[0019] In some embodiments, a local data processing module may include the processor, the local data processing module being operably coupled to a display device via a communication link, and wherein the head mounted device includes the display device.

[0020] According to some embodiments, a method for creating a world model using a wearable display system may be provided, the wearable display system comprising: a head-mounted device comprising a first camera having a global shutter and a second camera having a rolling shutter, the first camera and the second camera being positioned to provide overlapping views of a central field of view; and a processor operably coupled to the first camera and the second camera; wherein the method comprises, using the processor, performing a compensation routine to adjust an image acquired using the second camera for rolling shutter image distortion; and creating the world model in part using depth information stereoscopically determined based on the image acquired using the first camera and the adjusted image.

[0021] According to some embodiments, a wearable display system may be provided, comprising: a head-mounted device having a grayscale camera and a color camera positioned to provide overlapping views of a central field of view; and a processor operably coupled to the grayscale camera and the color camera and configured to: create a world model using first depth information stereoscopically determined based on images acquired by the grayscale camera and the color camera; and track head pose using the grayscale camera and the world model.

[0022] According to some embodiments, a method for tracking head pose using a wearable display system may be provided, the wearable display system comprising: a head-mounted device having a grayscale camera and a color camera positioned to provide overlapping views of a central field of view; and a processor operably coupled to the grayscale camera and the color camera; wherein the method comprises using the processor to: create a world model using first depth information stereoscopically determined based on images acquired by the grayscale camera and the color camera; and tracking the head pose using the grayscale camera and the world model.

[0023] According to some embodiments, a wearable display system may be provided, comprising: a frame; a first camera mechanically coupled to the frame and a second camera mechanically coupled to the frame, wherein the first camera and the second camera are positioned to provide a central field of view associated with both the first camera and the second camera, and wherein at least one of the first camera and the second camera comprises a plenoptic camera; and a processor operably coupled to the first camera and the second camera and configured to: determine whether an object is within the central field of view; determine whether the object meets a depth criterion when the object is within the central field of view; track the object using depth information stereoscopically determined based on images acquired by the first camera and the second camera when the tracked object is within the central field of view and does not meet the depth criterion; and track the object using depth information determined based on light field information acquired by one of the first camera or the second camera when the tracked object is within the central field of view and meets the depth criterion.

[0024] In accordance with some embodiments, a method for tracking an object using a wearable display system is provided, the wearable display system comprising: a frame; a first camera mechanically coupled to the frame and a second camera mechanically coupled to the frame, wherein the first camera and the second camera are positioned to provide a central field of view associated with both the first camera and the second camera, and wherein at least one of the first camera and the second camera comprises a plenoptic camera; and a processor operably coupled to the first camera and the second camera; wherein the method comprises using the processor to: determine whether an object is within the central field of view; determine whether the object meets a depth criterion when the object is within the central field of view; when the tracked object is within the central field of view and does not meet the depth criterion, track the object using depth information determined stereoscopically based on images acquired by the first camera and the second camera; and when the tracked object is within the central field of view and meets the depth criterion, track the object using depth information determined based on light field information acquired by one of the first camera or the second camera.

[0025] According to some embodiments, a wearable display system may be provided, comprising: a frame; two cameras mechanically coupled to the frame, wherein the two cameras comprise: a first camera with a global shutter having a first field of view; and a second camera with a rolling shutter having a second field of view; wherein the first camera and the second camera are positioned to provide: a central field of view in which the first field of view overlaps with the second field of view; and a peripheral field of view outside the central field of view; and a processor operably coupled to the first camera and the second camera.

[0026] The foregoing summary has been provided by way of illustration and is not intended to be limiting. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component that is illustrated in various figures is represented by a like numeral. For clarity, not every component may be labeled in every figure. In the drawings:

[0028] Figure 1 is a diagram illustrating an example of a simplified augmented reality (AR) scene according to some embodiments.

[0029] Figure 2 is a schematic diagram illustrating an example of an AR display system according to some embodiments.

[0030] Figure 3A is a diagram illustrating an AR display system rendering AR content as a user moves through a physical world environment while the user is wearing the system, according to some embodiments.

[0031] Figure 3B is a schematic diagram illustrating a viewing optics assembly and accompanying components according to some embodiments.

[0032] Figure 4 is a schematic diagram illustrating an image sensing system according to some embodiments.

[0033] Figure 5A is a diagram showing a method according to some embodiments Figure 4 Schematic diagram of a pixel unit in .

[0034] Figure 5B is a diagram showing a method according to some embodiments Figure 5A Schematic diagram of the output events of a pixel unit.

[0035] Figure 6 is a schematic diagram illustrating an image sensor according to some embodiments.

[0036] Figure 7 is a schematic diagram illustrating an image sensor according to some embodiments.

[0037] Figure 8 is a schematic diagram illustrating an image sensor according to some embodiments.

[0038] Figure 9 is a simplified flow chart of a method for image sensing according to some embodiments.

[0039] Figure 10 According to some embodiments Figure 9 A simplified flowchart of the actions of patch recognition.

[0040] Figure 11 According to some embodiments Figure 9 Simplified flowchart of the actions for block-wise trajectory estimation.

[0041] Figure 12 is a diagram showing a viewpoint according to some embodiments. Figure 11 Schematic diagram of block trajectory estimation.

[0042] Figure 13 is a diagram showing a change in viewpoint according to some embodiments. Figure 11 Schematic diagram of block trajectory estimation.

[0043] Figure 14 is a schematic diagram illustrating an image sensing system according to some embodiments.

[0044] Figure 15 is a diagram showing a method according to some embodiments Figure 14 Schematic diagram of a pixel unit in .

[0045] Figure 16 is a schematic diagram of a pixel subarray according to some embodiments.

[0046] Figure 17A is a cross-sectional view of a plenoptic device with an arrival angle-to-intensity converter in the form of two aligned stacked transmissive diffraction masks (TDMs) according to some embodiments.

[0047] Figure 17B is a cross-sectional view of an all-optical device with an arrival angle-to-intensity converter in the form of two non-aligned stacked TDMs, according to some embodiments.

[0048] Figure 18A is a pixel subarray having color pixel units and angle-of-arrival pixel units according to some embodiments.

[0049] Figure 18B is a pixel subarray having color pixel units and angle-of-arrival pixel units according to some embodiments.

[0050] Figure 18C is a pixel subarray having white pixel cells and corner-of-arrival pixel cells according to some embodiments.

[0051] Figure 19A is a top view of a photodetector array with a single TDM according to some embodiments.

[0052] Figure 19B is a side view of a photodetector array with a single TDM according to some embodiments.

[0053] Figure 20A is a top view of a photodetector array with multiple arrival angle-to-intensity converters in TDM format, according to some embodiments.

[0054] Figure 20B is a side view of a photodetector array with multiple TDMs according to some embodiments.

[0055] Figure 20C is a side view of a photodetector array with multiple TDMs according to some embodiments.

[0056] Figure 21 is a schematic diagram illustrating a headset including two cameras and accompanying components according to some embodiments.

[0057] Figure 22 is a simplified flow chart of a calibration routine according to some embodiments.

[0058] Figures 23A to 23C shows the Figure 21 An example field of view diagram associated with a head-mounted device.

[0059] Figure 24 is a simplified flow chart of a method for creating and updating a traversable world model according to some embodiments.

[0060] Figure 25 is a simplified flow chart of a method for head pose tracking according to some embodiments.

[0061] Figure 26 is a simplified flow chart of a method for object tracking according to some embodiments.

[0062] Figure 27 is a simplified flow chart of the process of hand tracking according to some embodiments. DETAILED DESCRIPTION

[0063] The inventors have recognized and understood design and operating techniques for wearable XR display systems that enhance the enjoyment and usefulness of such systems. These design and / or operating techniques can enable the acquisition of information to perform a variety of functions, including hand tracking, head pose tracking, and world reconstruction using a limited number of cameras, which can be used to realistically render virtual objects so that they appear to interact realistically with physical objects. The wearable cross-reality display system can be lightweight and can consume low power in operation. The system can use a specifically configured set of sensors to acquire image information about physical objects in the physical world with low latency. The system can perform various routines to improve the accuracy and / or realism of the displayed XR environment. Such routines can include calibration routines that improve the accuracy of stereo depth measurements even if the lightweight frame deforms during use, and routines that detect and resolve incomplete depth information in the model of the physical world around the user.

[0064] The weight of known XR system headsets can limit user enjoyment. Such XR headsets can weigh over 340 grams (sometimes even over 700 grams). In comparison, glasses may weigh less than 50 grams. Wearing such relatively heavy headsets for extended periods can fatigue or distract users, thereby detracting from the desired immersive XR experience. However, the inventors have recognized and appreciated that some designs that reduce headset weight also increase headset flexibility, making lightweight headsets susceptible to changes in sensor position or orientation during use or over time. For example, when a user wears a lightweight headset that includes camera sensors, the relative orientation of these camera sensors may shift. Changes in the spacing between cameras used for stereoscopic imaging can affect the ability of these headsets to acquire accurate stereo information, which relies on cameras having a known positional relationship relative to each other. Therefore, a calibration routine that can be repeated while the headset is worn can enable lightweight headsets to accurately acquire information about the world around the headset wearer using stereo imaging technology.

[0065] The need to equip XR systems with components to obtain information about objects in the physical world also limits the usefulness and user enjoyment of these systems. While the information obtained is used to realistically render computer-generated virtual objects in the proper location and appearance relative to physical objects, the need to obtain this information imposes limitations on the size, power consumption, and realism of XR systems.

[0066] For example, XR systems can use sensors worn by the user to obtain information about objects in the physical world around the user, including information about their positions in the user's field of view. Challenges arise because objects may move relative to the user's field of view due to movement in the physical world or changes in the user's posture relative to the physical world, causing physical objects to enter or exit the user's field of view, or because the position of physical objects within the user's field of view changes. To render realistic XR displays, models of physical objects in the physical world must be updated frequently enough to capture these changes, processed with sufficiently low latency, and accurately predicted into the future to cover the full latency path, including rendering. This ensures that when virtual objects are displayed, they will have the appropriate position and appearance relative to the physical objects based on this information. Otherwise, virtual objects will be misaligned with physical objects, and the combined scene including physical and virtual objects will appear unrealistic. For example, virtual objects may appear to float in space rather than resting on physical objects, or may appear to bounce around relative to physical objects. Errors in visual tracking are particularly amplified when the user is moving at high speeds and there is noticeable motion in the scene.

[0067] These problems could be avoided by using sensors that acquire new data at a high rate. However, the power consumed by such sensors can lead to the need for larger batteries, increase the weight of the system, or limit the usability of such systems. Similarly, the processors required to process the data generated at a high rate can drain batteries and increase the weight of wearable systems, further limiting the practicality or enjoyment of such systems. For example, one known approach is to operate at a higher resolution to capture sufficient visual detail and operate the sensor at a higher frame rate to increase temporal resolution. An alternative solution might supplement this solution with an IR time-of-flight sensor, which might directly indicate the position of a physical object relative to the sensor, and simple processing might be performed when using this information to display virtual objects, resulting in low latency. However, such sensors consume a lot of power, especially if they operate in sunlight.

[0068] The inventors have recognized and appreciated that an XR system can account for changes in sensor position or orientation during use or over time by repeatedly performing a calibration routine. This calibration routine can determine the current relative separation and orientation of sensors included in a head-mounted device. The wearable XR system can then take into account the current relative separation and orientation of the head-mounted device sensors when calculating stereo depth information. With such a calibration capability, the XR system can accurately obtain depth information to indicate the distance to objects in the physical world without the need for active depth sensors or only occasionally using active depth sensing. Because active depth sensing can consume a lot of power, reducing or eliminating active depth sensing can enable the device to consume less power, which can increase the device's operating time without charging the battery or reduce the size of the device due to a reduced battery size.

[0069] The inventors have also recognized and understood that, with an appropriate combination of image sensors, and appropriate techniques for processing image information from these sensors, an XR system can acquire information about physical objects with low latency and even with reduced power consumption by: reducing the number of sensors used; eliminating, disabling, or selectively activating resource-intensive sensors; and / or reducing overall sensor usage. As a specific example, an XR system may include a head-mounted device with two cameras. A first of the cameras may generate grayscale images and may have a global shutter. These grayscale images may be smaller than color images of similar resolution, in some cases being represented with fewer than one-third the number of bits. Such a grayscale camera may require less power than a color camera of similar resolution. A second of the cameras may be an RGB camera. The wearable cross-reality display system may be configured to selectively use this camera, reducing power consumption and extending battery life without impacting the user's XR experience.

[0070] The techniques described herein may be used in various types of scenarios, with various types of devices or alone. Figure 1 Such a scenario is shown. Figure 2 、 3A 3B illustrate an exemplary AR system including one or more processors, memory, sensors, and a user interface that can operate according to the techniques described herein.

[0071] refer to Figure 1, shows an AR scene 4 in which a user of the AR system sees a physical-world park-like setting 6, which features people, trees, buildings in the background, and a concrete platform 8. In addition to these physical objects, the user of the AR technology also perceives that they "see" virtual objects, illustrated here as a robotic statue 10 standing on the physical-world concrete platform 8, and a flying cartoon-like avatar character 2 that appears to be the head of a bumblebee, even though these elements (e.g., avatar character 2 and robotic statue 10) do not exist in the physical world. Due to the extreme complexity of human visual perception and the nervous system, it is challenging to create an AR system that promotes a comfortable, natural-feeling, rich presentation of virtual image elements among other virtual or physical-world image elements.

[0072] Such a scene can be presented to the user by presenting image information representing the actual environment around the user and overlaying information representing virtual objects that are not in the actual environment. In an AR system, the user may be able to see objects in the physical world, and the AR system provides information for rendering virtual objects so that they appear in the appropriate location and have appropriate visual characteristics so that the virtual objects appear to coexist with objects in the physical world. For example, in an AR system, the user can look through a transparent screen so that the user can see objects in the physical world. The AR system can render virtual objects on the screen so that the user can see the physical world and the virtual objects at the same time. In some embodiments, the screen can be worn by the user, such as a pair of goggles or glasses.

[0073] The scene can be presented to the user via a system including multiple components, including a user interface that can stimulate one or more user senses (including vision, sound and / or touch). In addition, the system can include one or more sensors that can measure parameters of the physical part of the scene, including the position and / or movement of the user within the physical part of the scene. In addition, the system can include one or more computing devices, and associated computer hardware, such as memory. These components can be integrated into a single device, or more distributed across multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.

[0074] In some embodiments, an AR experience may be provided to a user via a wearable display system. Figure 2An example of a wearable display system 80 (hereinafter referred to as "system 80") is shown. System 80 includes a head-mounted display device 62 (hereinafter referred to as "display device 62"), and various mechanical and electronic modules and systems that support the functionality of display device 62. Display device 62 can be coupled to a frame 64 that can be worn by a display system user or viewer 60 (hereinafter referred to as "user 60") and is configured to position display device 62 in front of the eyes of user 60. According to various embodiments, display device 62 can be sequentially displayed. Display device 62 can be monocular or binocular.

[0075] In some embodiments, speaker 66 is coupled to frame 64 and positioned near the ear canal of user 60. In some embodiments, another speaker, not shown, is positioned near the other ear canal of user 60 to provide stereo / plastic sound control.

[0076] The system 80 may include a local data processing module 70. The local data processing module 70 may be operably coupled to the display device 62 via a communication link 68 (such as via a wired conductor or a wireless connection). The local data processing module 70 may be mounted in a variety of configurations, such as fixedly attached to the frame 64, fixedly attached to a helmet or hat worn by the user 60, embedded in headphones, or otherwise removably attached to the user 60 (e.g., in a backpack-style configuration, in a belt-coupled configuration). In some embodiments, the local data processing module 70 may not be present, as components of the local data processing module 70 may be integrated into the display device 62 or implemented in a remote server or other component to which the display device 62 is coupled, such as by wireless communication over a wide area network.

[0077] The local data processing module 70 may include a processor and digital storage, such as non-volatile memory (e.g., flash memory), both of which may be used to facilitate processing, caching, and storage of data. The data may include: a) data captured from sensors (e.g., which may be operably coupled to the frame 64) or otherwise attached to the user 60, such as an image capture device (e.g., a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio, and / or a gyroscope; and / or b) data acquired and / or processed using the remote processing module 72 and / or remote data repository 74, and possibly transferred to the display device 62 after such processing or acquisition. The local data processing module 70 may be operably coupled to the remote processing module 72 and the remote data repository 74 via communication links 76, 78, respectively (e.g., via wired or wireless communication links), such that these remote modules 72, 74 are operably coupled to each other and can serve as resources for the local processing and data module 70.

[0078] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a central processing unit and / or one or more graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, the remote data repository 74 may include a digital data storage facility that may be available via the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 70, allowing for fully autonomous use from a remote module.

[0079] In some embodiments, local data processing module 70 is operably coupled to battery 82. In some embodiments, battery 82 is a removable power source, such as a counter battery. In other embodiments, battery 82 is a lithium-ion battery. In some embodiments, battery 82 includes both an internal lithium-ion battery that can be recharged by user 60 during non-operating times of system 80, and a removable battery, allowing user 60 to operate system 80 for extended periods without having to connect to a power source to recharge the lithium-ion battery or shut down system 80 to replace the battery.

[0080] Figure 3A A user 30 wearing an AR display system that renders AR content is shown as the user 30 moves through a physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 positions the AR display system at a location 34, and the AR display system records environmental information of the traversable world relative to the location 34 (e.g., digital representations of real objects in the physical world, which can be stored and updated as changes are made to the real objects in the physical world). Each location 34 can also be associated with a "pose" and / or mapped features or directional audio input related to the environment 32. A user wearing the AR display system on their head might look in a particular direction and tilt their head, thereby creating a head pose of the system relative to the environment. At each location and / or pose within the same location, the sensors on the AR display system can capture different information about the environment 32. Therefore, the information collected at the location 34 can be aggregated to a data input 36 and processed by at least a traversable world module 38, which can, for example, be processed by Figure 2 This is achieved by processing on the remote processing module 72.

[0081] Navigable world module 38 determines where and how AR content 40 can be placed relative to the physical world, based at least in part on data input 36. AR content is "placed" within the physical world by presenting it in a way that allows the user to see both the AR content and the physical world simultaneously. For example, such an interface can be created using glasses through which the user can view the physical world, and the glasses can be controlled to cause virtual objects to appear at controlled locations within the user's field of view. AR content is rendered as if interacting with objects in the physical world. The user interface allows the user's view of objects in the physical world to be obscured to create locations where AR content appears to obscure the user's view of those objects when appropriate. For example, AR content can be placed by appropriately selecting a portion of an element 42 (e.g., a table) in environment 32 to be displayed and displaying the shape and position of AR content 40 as if it were resting on or otherwise interacting with that element 42. AR content can also be placed within structures that are not already within field of view 44 or within a mapped mesh model 46 relative to the physical world.

[0082] As shown, element 42 is an example of a plurality of elements in the physical world that can be considered fixed and stored in traversable world module 38. Once stored in traversable world module 38, information about these fixed elements can be used to present information to the user, so that the user 30 can perceive the content on the fixed element 42 without the system having to map to the fixed element 42 each time the user 30 sees the fixed element 42. Thus, the fixed element 42 can be a mapped mesh model from a previous modeling session, or can be determined by an individual user but still stored in traversable world module 38 for future reference by multiple users. Thus, traversable world module 38 can identify environment 32 from a previously mapped environment and display AR content without requiring the user 30's device to first map the environment 32, thereby saving computational processes and cycles and avoiding any latency in rendering the AR content.

[0083] Similarly, a mapped mesh model 46 of the physical world can be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying AR content 40 can be mapped and stored in the traversable world module 38 for future retrieval by the user 30 or other users without the need for remapping or modeling. In some embodiments, data input 36 is input such as geographic location, user identification, and current activity to indicate to the traversable world module 38 which of one or more fixed elements 42 is available, which AR content 40 was last placed on the fixed element 42, and whether to display the same content (such AR content is "persistent" regardless of how the user views the particular traversable world model).

[0084] Even in embodiments where objects are considered stationary, the traversable world module 38 may be updated from time to time to account for the possibility of changes in the physical world. Models of stationary objects may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered stationary. To render realistic AR scenes, the AR system may update the positions of these non-stationary objects at a much higher frequency than that used to update fixed objects. To be able to accurately track all objects in the physical world, the AR system may obtain information from multiple sensors, including one or more image sensors.

[0085] Figure 3B is a schematic diagram of the viewing optical assembly 48 and accompanying optional components. Figure 21 Specific configurations are described in [ ]. Directed toward the user's eyes 49, in some embodiments, two eye-tracking cameras 50 detect metrics of the user's eyes 49, such as eye shape, eyelid occlusion, pupil orientation, and glint. In some embodiments, one of the sensors may be a depth sensor 51, such as a time-of-flight sensor, which transmits signals into the world and detects reflections of those signals from nearby objects to determine the distance to a given object. A depth sensor can quickly determine whether an object has entered the user's field of view due to, for example, the motion of those objects or a change in the user's posture. However, information about the position of an object in the user's field of view may alternatively or additionally be collected by other sensors. In some embodiments, world camera 52 records a view larger than the periphery to map environment 32 and detect inputs that can affect AR content. In some embodiments, world camera 52 and / or camera 53 may be grayscale and / or color image sensors that can output grayscale and / or color image frames at fixed time intervals. Camera 53 may further capture images of the physical world within the user's field of view at specific times. Pixels of a frame-based image sensor may be repeatedly sampled even if their values do not change. Each of the world camera 52, camera 53 and depth sensor 51 has a corresponding field of view 54, 55 and 56 to obtain a Figure 3A The data is collected in the physical world scene of the physical world environment 32 shown in FIG and the physical world scene is recorded.

[0086] Inertial measurement unit 57 can determine the motion and / or orientation of viewing optics assembly 48. In some embodiments, each component is operably coupled to at least one other component. For example, depth sensor 51 can be operably coupled to eye-tracking camera 50 to determine the actual distance of a point and / or area in the physical world that the user's eye 49 is looking at.

[0087] It should be understood that the viewing optics assembly 48 may include Figure 3BComponents shown. For example, viewing optics assembly 48 may include a different number of components. In some embodiments, for example, viewing optics assembly 48 may include one world camera 52, two world cameras 52, or more world cameras instead of the four world cameras shown. Alternatively or additionally, cameras 52 and 53 need not capture visible light images of their entire fields of view. Viewing optics assembly 48 may include other types of components. In some embodiments, viewing optics assembly 48 may include one or more dynamic vision sensors (DVS) whose pixels may asynchronously respond to relative changes in light intensity exceeding a threshold.

[0088] In some embodiments, based on the time-of-flight information, the viewing optical assembly 48 may not include the depth sensor 51. For example, in some embodiments, the viewing optical assembly 48 may include one or more plenoptic cameras whose pixels can capture not only the light intensity, but also the angle of the incident light. For example, the plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or in addition, the plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such a sensor may be used as a source of depth information instead of or in addition to the depth sensor 51.

[0089] It should also be understood that Figure 3B The configuration of the components in FIG. 4 is shown as an example. The viewing optics assembly 48 may include components having any suitable configuration to allow a user to have the maximum field of view for a particular set of components. For example, if the viewing optics assembly 48 includes a world camera 52, the world camera may be placed in the center area of the viewing optics assembly rather than on the side.

[0090] Information from these sensors in the viewing optical assembly 48 can be coupled to one or more processors in the system. The processor can generate data that can be rendered so that the user perceives virtual content interacting with objects in the physical world. This rendering can be achieved in any suitable manner, including generating image data showing both physical and virtual objects. In other embodiments, physical and virtual content can be shown in a scene by modulating the opacity of a display device that the user is viewing in the physical world. The opacity can be controlled to create the appearance location of virtual objects and also prevent the user from seeing objects in the physical world that are obscured by virtual objects. In some embodiments, the image data may only include virtual content that can be modified to interact with the physical world in a realistic manner (e.g., clipping content to resolve obstructions), which can be viewed through the user interface. Regardless of how the content is presented to the user, a model of the physical world can be used so that the characteristics of virtual objects that can be affected by physical objects can be correctly calculated, including the shape, position, movement, and visibility of virtual objects.

[0091] A model of the physical world can be created based on data collected from sensors on a user's wearable device. In some embodiments, a model can be created from data collected from multiple users, which can be aggregated in a computing device remote from all users (and this data can be in the "cloud").

[0092] In some embodiments, at least one of the sensors can be configured to acquire information about physical objects in a scene, particularly non-stationary objects, at high frequency and low latency using compact and low-power components. The sensor can employ block tracking to limit the amount of data output.

[0093] Figure 4 An image sensing system 400 is shown in accordance with some embodiments. Image sensing system 400 may include an image sensor 402, which may include an image array 404. Image array 404 may include a plurality of pixels, each pixel responsive to light as in a conventional image sensor. Sensor 402 may also include circuitry for accessing each pixel. Accessing a pixel may require obtaining information about incident light generated by the pixel. Alternatively or additionally, accessing a pixel may require controlling the pixel, for example by configuring it to provide an output only upon detecting a certain event.

[0094] In the illustrated embodiment, image array 404 is configured as an array having multiple rows and columns of pixels. In such an embodiment, the access circuitry may be implemented as a row address encoder / decoder 406 and a column address encoder / decoder 408. Image sensor 402 may also include circuitry that generates inputs to the access circuitry to control the timing and order in which information is read out from the pixels in image array 404. In the illustrated embodiment, this circuitry is a tile tracking engine 410. Compared to conventional image sensors that can continuously output image information captured by pixels in each row, image sensor 402 can be controlled to output image information in specified tiles. Furthermore, the positions of these tiles relative to the image array can change over time. In the illustrated embodiment, tile tracking engine 410 may output image array access information to control the output of image information from portions of image array 404 corresponding to the positions of the tiles, and the access information may be dynamically changed based on an estimate of the movement of objects in the environment and / or an estimate of the movement of the image sensor relative to these objects.

[0095] In some embodiments, the image sensor 402 may include a dynamic vision sensor (DVS) function, such that image information is provided by the sensor only when a pixel's image property (e.g., intensity) changes. For example, the image sensor 402 may apply one or more thresholds that define the on (ON) and off (OFF) states of a pixel. The image sensor may detect when a pixel changes state and selectively provide outputs only for those pixels, or those pixels in a block, that changed state. These outputs may be asynchronous upon detection, rather than as part of a readout of all pixels in the array. For example, the output may be in the form of an address event representation (AER) 418, which may include a pixel address (e.g., row and column) and an event type (ON or OFF). An ON event may indicate that the pixel cell at the corresponding pixel address sensed an increase in light intensity; and an OFF event may indicate that the pixel cell at the corresponding pixel address sensed a decrease in light intensity. The increase or decrease may be relative to an absolute level, or may be a change in the level of the pixel's last output. For example, the change may be expressed as a fixed offset or as a percentage of the value of the pixel's last output.

[0096] Using DVS technology in conjunction with tile tracking can make image sensors suitable for XR systems. When combined in an image sensor, the amount of data generated may be limited to the pixel cells within the tile and the data of the pixel cells that detect the change that will trigger the event output.

[0097] In some scenarios, high-resolution image information is required. However, when using DVS technology, large sensors with over one million pixels for generating high-resolution image information can generate a large amount of image information. The inventors have recognized and understood that DVS sensors can generate a large number of events reflecting background motion or image changes, rather than the motion of the tracked object. Currently, DVS sensor resolutions are limited to less than 1MB, such as 128x128, 240x180, or 346x260, to limit the number of events generated. Such sensors sacrifice resolution for tracking the object and may not be able to detect, for example, fine finger movements in a hand. Furthermore, if the image sensor outputs image information in other formats, limiting the resolution of the sensor array to output a manageable number of events can also limit the image sensor's use in conjunction with DVS functionality to generate high-resolution image frames. In some embodiments, sensors as described herein can have resolutions higher than VGA, including up to 8 or 12 megapixels. Nevertheless, the block tracking described herein can be used to limit the number of events output per second by the image sensor. Consequently, image sensors can be enabled to operate in at least two modes. For example, an image sensor with a megapixel resolution can operate in a first mode, in which it outputs events in a specific tile being tracked. In a second mode, it can output a high-resolution image frame or portion of an image frame. Such an image sensor can be controlled in an XR system to operate in these different modes based on the system's capabilities.

[0098] Image array 404 may include a plurality of pixel cells 500 arranged in an array. Figure 5A An example of a pixel cell 500 is shown, which in this embodiment is configured for use in an imaging array implementing DVS technology. Pixel cell 500 may include a photoreceptor circuit 502, a differential circuit 506, and a comparator 508. Photoreceptor circuit 502 may include a photodiode 504 that converts light striking the photodiode into a measurable electrical signal. In this example, the conversion is to a current I. A transconductance amplifier 510 converts the photocurrent I into a voltage. The conversion may be linear or nonlinear, for example, according to a function logI. Regardless of the specific transfer function, the output of transconductance amplifier 510 indicates the amount of light detected at photodiode 504. While a photodiode is shown as an example, it should be understood that other photosensitive components that generate a measurable output in response to incident light may be implemented in the photoreceptor circuit in place of or in addition to the photodiode.

[0099] exist Figure 5AIn an embodiment of the present invention, the circuitry for determining whether the output of a pixel has changed sufficiently to trigger an output for that pixel cell is incorporated into the pixel itself. In this example, this functionality is implemented by a differential circuit 506 and a comparator 508. The differential circuit 506 can be configured to reduce DC mismatches between pixel cells by, for example, balancing the output of the differential circuit to reset the level after an event is generated. In this example, the differential circuit 506 is configured to generate an output that shows the change in the output of the photodiode 504 since the last output. The differential circuit can include an amplifier 512 with a gain of A, a capacitor 514 that can be implemented as a single circuit element or as one or more capacitors connected in a network, and a reset switch 516.

[0100] In operation, the pixel cell is reset by temporarily closing switch 516. Such a reset can occur at the beginning of the circuit's operation and at any time after an event is detected. When the pixel 500 is reset, the voltage across capacitor 514 is 0.01 V. When the voltage across capacitor 514 is subtracted from the output of transconductance amplifier 510, a zero voltage is generated at the input of amplifier 512. When switch 516 is open, the output of transconductance amplifier 510 is combined with the voltage drop across capacitor 514, resulting in a zero voltage at the input of amplifier 512. The output of transconductance amplifier 510 changes due to changes in the amount of light striking photodiode 504. When the output of transconductance amplifier 510 increases or decreases, the output of amplifier 512 will swing positively or negatively by an amount that is amplified by the gain of amplifier 512.

[0101] Comparator 508 can determine whether an event has occurred and the sign of the event by, for example, comparing the output voltage V of the differential circuit to a predetermined threshold voltage C. In some embodiments, comparator 508 can include two comparators having transistors. When the output of amplifier 512 shows a positive change, one pair of comparators can operate and detect an increasing change (ON event); when the output of amplifier 512 shows a negative change, the other comparator can operate and detect a decreasing change (OFF event). However, it should be understood that amplifier 512 can have a negative gain. In such an embodiment, an increase in the output of transconductance amplifier 510 can be detected as a negative voltage change at the output of amplifier 512. Similarly, it should be understood that the positive and negative voltages can be relative to ground or any suitable reference level. In any case, the value of threshold voltage C can be controlled by the characteristics of the transistors (e.g., transistor size, transistor threshold voltage) and / or the value of a reference voltage that can be applied to comparator 508.

[0102] Figure 5BAn example of event output (ON, OFF) of the pixel unit 500 over time t is shown. In the example shown, at time t1, the output value of the differential circuit is V1; at time t2, the output value of the differential circuit is V2; and at time t3, the output value of the differential circuit is V3. Between time t1 and time t2, although the photodiode senses an increase in light intensity, the pixel unit does not output any event because the change in V does not exceed the value of the threshold voltage C. At time t2, because V2 is greater than V1 by the value of the threshold voltage C, the pixel unit outputs an ON event. Between time t2 and time t3, although the photodiode senses a decrease in light intensity, the pixel unit does not output any event because the change in V does not exceed the value of the threshold voltage C. At time t3, because V3 is less than V2 by the value of the threshold voltage C, the pixel unit outputs an OFF event.

[0103] Each event can trigger an output at AER 418. The output can include, for example, an indication of whether the event is an ON or OFF event and the identity of the pixel, such as the row and column of the pixel. Other information can alternatively or additionally be included in the output. For example, a timestamp might be included, which can be useful if the event is queued for later transmission or processing. As another example, the current level at the output of amplifier 510 can be included. Such information can optionally be included, for example, if further processing is to be performed in addition to detecting the movement of the object.

[0104] It should be understood that the frequency of the event output and the sensitivity of the pixel unit can be controlled by the value of the threshold voltage C. For example, the frequency of the event output can be reduced by increasing the value of the threshold voltage C, or the frequency of the event output can be increased by decreasing the value of the threshold voltage C. It should also be understood that the threshold voltage C can be different for ON events and OFF events, for example, by setting different reference voltages for the comparator used to detect ON events and the comparator used to detect OFF events. It should also be understood that the pixel unit can also output a value indicating the magnitude of the light intensity change instead of the sign signal indicating event detection, or output a value indicating the magnitude of the light intensity change in addition to the sign signal indicating event detection.

[0105] Figure 5A and Figure 5B Pixel cell 500 is shown as an example according to some embodiments. Other designs may also be applicable to pixel cells. In some embodiments, a pixel cell may include a photoreceptor circuit and a differential circuit, but share a comparator circuit with one or more other pixel cells. In some embodiments, a pixel cell may include circuitry configured to calculate a change value, for example, an active pixel sensor at the pixel level.

[0106] Regardless of how events are detected for each pixel unit, the ability to configure pixels to output only when an event is detected can be used to limit the amount of information required to maintain a model of the position of a non-fixed (i.e., movable) object. For example, pixels within a block can be set using a threshold voltage C that is triggered when a relatively small change occurs. Other pixels outside the block may have a larger threshold, such as three or five times the threshold. In some embodiments, the threshold voltage C of pixels outside any block can be set large enough so that the pixel is effectively disabled and does not produce any output, regardless of the amount of change. In other embodiments, pixels outside the block can be disabled in other ways. In such an embodiment, the threshold voltage can be fixed for all pixels, but the pixel can be selectively enabled or disabled based on whether the pixel is within a block.

[0107] In other embodiments, the threshold voltage of one or more pixels can be adaptively set in a manner that modulates the amount of data output from the image array. For example, the AR system can have a processing capability to process multiple events per second. When the number of events output per second exceeds an upper limit, the threshold of some or all pixels can be increased. Alternatively or additionally, when the number of events per second drops below a lower limit, the threshold can be lowered, thereby enabling more data to be used for more accurate processing. As a specific example, the number of events per second can be between 200 and 2000 events. Compared to, for example, processing all pixel values scanned from an image sensor (constituting 30 million or more pixel values per second), this number of events constitutes a significant reduction in the number of data blocks to be processed per second. This number of events is even reduced compared to processing only the pixels within the block (which may be fewer but still may be tens of thousands or more pixel values per second).

[0108] The control signals used to enable and / or set the threshold voltage for each of the plurality of pixels may be generated in any suitable manner. However, in the illustrated embodiment, those control signals are set by the tile tracking engine 410 or based on processing within the processing module 72 or other processor.

[0109] Return to reference Figure 4, image sensing system 400 can receive input from any suitable component such that tile tracking engine 410 can dynamically select at least one region of image array 404 to enable and / or disable in order to implement tile based at least on the received input. Tile tracking engine 410 can be a digital processing circuit having a memory storing one or more parameters for the tile. The parameters can be, for example, the boundaries of the tile and can include other information such as information related to a scaling factor between movement of the image array and movement within the image array of an image of a movable object associated with the tile. Tile tracking engine 410 can also include circuitry configured to perform calculations on stored values and other measurements provided as input.

[0110] In the illustrated embodiment, the tile tracking engine 410 receives as input a designation of a current tile. A tile may be designated based on its size and location within the image array 404, for example, by specifying a range of row and column addresses for the tile. This designation may be used as a proxy for the processing module 72 ( Figure 2 ) or other components that process information about the physical world. For example, the processing module 72 may designate a tile to contain the current position of each movable object in the physical world, or the current position of a subset of movable objects being tracked, so that virtual objects are rendered with appropriate appearance positions relative to the physical world. For example, if an AR scene will include a toy doll balanced on a physical object (e.g., a moving toy car) as a virtual object, a tile containing the toy car may be designated. A tile may not be designated for another toy car moving in the background because up-to-date information about that object may not be needed to render a realistic AR scene.

[0111] Regardless of how the tiles are selected, information about the current location of the tiles can be provided to the tile tracking engine 410. In some embodiments, the tiles can be rectangular, so that the location of the tiles can be simply specified as a starting row and column and an ending row and column. In other embodiments, the tiles can have other shapes, such as circles, and the tiles can be specified in other ways, such as by a center point and a radius.

[0112] In some embodiments, trajectory information about the tiles may also be provided. For example, the trajectory may specify the movement of the tiles relative to the coordinates of image array 404. For example, processing module 72 may model the movement of movable objects within the physical world and / or the movement of image array 404 relative to the physical world. Because the movement of either or both of these may affect the location within image array 404 where the image of the object is projected, the trajectory of the tiles within image array 404 may be calculated based on either or both of these. The trajectory may be specified in any suitable manner, such as parameters of a linear, quadratic, cubic, or other polynomial equation.

[0113] In other embodiments, the tile tracking engine 410 can dynamically calculate the position of the tiles based on input from sensors that provide information about the physical world. The information from the sensors can be provided directly from the sensors. Alternatively or additionally, the sensor information can be processed to extract information about the physical world before being provided to the tile tracking engine 410. For example, the extracted information can include the movement of the image array 404 relative to the physical world, the distance between the image array 404 and the object whose image falls within the tile, and other information that can be used to dynamically align the tiles in the image array 404 with the image of the object in the physical world as the image array 404 and / or the object moves.

[0114] Examples of input components may include image sensors 412 and inertial sensors 414. Examples of image sensors 412 may include eye-tracking camera 50, depth sensor 51, world camera 52, and / or camera 52. Examples of inertial sensors 414 may include inertial measurement unit 57. In some embodiments, the input component may be selected to provide data at a relatively high rate. For example, inertial measurement unit 57 may have an output rate of between 200 and 2000 measurements per second, such as between 800 and 1200 measurements per second. Tile positions may be updated at a similarly high rate. As a specific example, by using inertial measurement unit 57 as an input source for tile tracking engine 410, tile positions may be updated 800 to 1200 times per second. In this way, movable objects can be tracked with high accuracy using relatively small tiles, limiting the number of events that need to be processed. This approach can result in very low latency between changes in the relative position of the image sensor and the movable object, and similarly low latency in updating virtual object rendering, thereby providing an ideal user experience.

[0115] In some scenarios, the movable objects tracked using tiles can be stationary objects in the physical world. For example, an AR system can identify stationary objects by analyzing multiple images of the physical world taken, and select one or more features of the stationary objects as reference points to determine the movement of a wearable device having an image sensor thereon. Frequent and low-latency updates of the positions of these reference points relative to the sensor array can be used to provide frequent and low-latency calculations of the head pose of the wearable device user. Since the head pose can be used to realistically render virtual objects through a user interface on the wearable device, the frequent and low-latency updates of the head pose improve the user experience of the AR system. Therefore, having the input of the tile tracking engine 410 that controls the tile positions come only from sensors with a high output rate, such as one or more inertial measurement units, can produce a satisfactory user experience of the AR system.

[0116] However, in some embodiments, other information may be provided to the tile tracking engine 410 to enable it to calculate trajectories and / or apply trajectories to the tiles. This other information may include stored information 416, such as the traversable world module 38 and / or the mapping grid model 46. This information may indicate one or more previous positions of the object relative to the physical world, such that changes in these previous positions and / or changes in the current position relative to the previous positions may indicate a trajectory of the object in the physical world, which may then be mapped to a trajectory across the tiles of the image array 404. Other information in the physical world model may be used alternatively or additionally. For example, the size and / or distance of a movable object or other information regarding its position relative to the image array 404 may be used to calculate the position or trajectory across the tiles of the image array 404 associated with the object.

[0117] Regardless of how the trajectory is determined, tile tracking engine 410 can use the trajectory to calculate updated positions of tiles within image array 404 at a high rate, for example, faster than once per second or more than 800 times per second. In some embodiments, this rate may be limited by processing power to less than 2000 times per second.

[0118] It should be understood that tracking changes in movable objects is not sufficient to reconstruct the complete physical world. However, the interval between reconstructing the physical world may be longer than the interval between position updates of movable objects, such as every 30 seconds or every 5 seconds. When there is a reconstruction of the physical world, the positions of the objects to be tracked and the positions of the tiles that will capture information about these objects can be recalculated.

[0119] Figure 4 An embodiment is shown in which the processing circuitry for dynamically generating tiles and controlling the selective output of image information from within the tiles is configured to directly control image array 404 so that the image information output from the array is limited to selected information. For example, such circuitry can be integrated into the same semiconductor chip that houses image array 404, or can be integrated into a separate controller chip for image array 404. However, it should be understood that the circuitry for generating control signals for image array 404 can be distributed throughout the XR system. For example, some or all of the functionality can be performed by programming in processing module 72 or other processors within the system.

[0120] Image sensing system 400 may output image information for each of a plurality of pixels. Each pixel of the image information may correspond to one of the pixel cells of image array 404. The output image information from image sensing system 400 may be image information for each of one or more tiles selected by tile tracking engine 410 corresponding to at least one region of image array 404. In some embodiments, for example, when each pixel of image array 404 has a different Figure 5A When configured as shown in , pixels in the output image information may identify pixels within one or more tiles where image sensor 400 detected a change in light intensity.

[0121] In some embodiments, the output image information from image sensing system 400 may be image information of pixels outside each of one or more tiles corresponding to at least one region of an image array, selected by tile tracking engine 410. For example, a deer may be running in a physical world with a flowing river. The details of the river's waves may not be of interest, but may trigger pixel cells in image array 402. Tile tracking engine 410 may create a tile surrounding the river and disable a portion of image array 402 corresponding to the tile surrounding the river.

[0122] Based on the identification of changed pixels, further processing can be performed. For example, the portion of the world model corresponding to the portion of the physical world imaged by the changed pixels can be updated. These updates can be performed based on information collected using other sensors. In some embodiments, further processing can be conditioned on or triggered by a number of changed pixels in a tile. For example, once it is detected that 10% or some other threshold number of pixels in a tile have changed, an update can be performed.

[0123] In some embodiments, image information in other formats may be output from the image sensor and may be used in conjunction with the change information to update the world model. In some embodiments, the format of the image information output from the image sensor may change at any time during operation of the VR system. For example, in some embodiments, the pixel unit 500 may be operated to generate a differential output at certain times, such as generated in the comparator 508. The output of the amplifier 510 may be switchable so as to output the magnitude of light incident on the photodiode 504 at other times. For example, the output of the amplifier 510 may be switchably connected to a sense line, which in turn is connected to an A / D converter that may provide a digital indication of the magnitude of the incident light based on the magnitude of the output of the amplifier 510.

[0124] An image sensor in this configuration can be operated as part of an AR system to perform differential output most of the time, outputting events only for pixels where a change exceeding a threshold is detected, or outputting events only for pixels within a tile where a change exceeding a threshold is detected. A full image frame with amplitude information for all pixels in the image array can be output periodically (e.g., every 5 to 30 seconds). In this way, low latency and accurate processing can be achieved, with differential information used to quickly update selected parts of the world model that are most likely to change and affect user perception, while the full image can be used to update larger parts of the world model more frequently. Although full updates to the world model occur only at a slower rate, any delay in updating the model may not have a meaningful impact on the user's perception of the AR scene.

[0125] The output mode of the image sensor can be changed at any time throughout the operation of the image sensor, so that the sensor outputs one or more intensity information for some or all pixels, as well as an indication of the change for some or all pixels in the array.

[0126] It is not essential that image information from the tiles be selectively output from the image sensor by limiting the information output from the image array. In some embodiments, image information may be output by all pixels in the image array, and only information about a specific region of the array may be output from the image sensor. Figure 6 An image sensor 600 is shown according to some embodiments. The image sensor 600 may include an image array 602. In this embodiment, the image array 602 may be similar to a conventional image array that scans out rows and columns of pixel values. The operation of such an image array may be coordinated by other components. The image sensor 600 may also include a block tracking engine 604 and / or a comparator 606. The image sensor 600 may provide an output 610 to an image processor 608. For example, the processor 608 may be the processing module 72 ( Figure 2 ) part.

[0127] The block tracking engine 604 can have a structure and function similar to that of the block tracking engine 410. It can be configured to receive a signal specifying at least one selected area of the image array 602 and then generate a control signal specifying the dynamic position of the area based on a calculated trajectory within the image array 602 of the image of the object represented by the area. In some embodiments, the block tracking engine 604 can receive a signal specifying at least one selected area of the image array 602, which signal can include trajectory information for one or more areas. The block tracking engine 604 can be configured to perform calculations to dynamically identify pixel cells within the at least one selected area based on the trajectory information. Variations in the implementation of the block tracking engine 604 are possible. For example, the block tracking engine can update the position of the block based on a sensor that indicates movement of the image array 602 and / or projected movement of an object associated with the block.

[0128] exist Figure 6 In the illustrated embodiment, image sensor 600 is configured to output differential information for pixels within identified blocks. Comparator 606 can be configured to receive control signals from block tracking engine 604 that identify pixels within the blocks. Comparator 606 can selectively operate on pixels output from image array 602, which have addresses within the blocks as indicated by block tracking engine 604. Comparator 606 can operate on pixel cells to generate a signal indicating a change in sensed light detected by at least one region of image array 602. As one example embodiment, comparator 606 can include a memory element that stores reset values for pixel cells within the array. When the current values of these pixels are scanned from image array 602, circuitry within comparator 606 can compare the stored value with the current value and output an indication if the difference exceeds a threshold. For example, digital circuitry can be used to store the values and perform this comparison. In this example, the output of image sensor 600 can be processed similarly to the output of image sensor 400.

[0129] In some embodiments, the image array 602, the tile tracking engine 604, and the comparator 606 can be implemented in a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, the image array 602 can be implemented in a single integrated circuit. The tile tracking engine 604 and the comparator 606 can be implemented in a second single integrated circuit, which can be configured as a driver for the image array 602, for example. Alternatively or additionally, some or all of the functionality of the tile tracking engine and / or the comparator 606 can be distributed to other digital processors within the AR system.

[0130] Other configurations or processing circuits are possible. Figure 7An image sensor 700 is shown according to some embodiments. The image sensor 700 may include an image array 702. In this embodiment, the image array 702 may have pixel cells with a differential configuration, e.g. Figure 5A 500. However, the embodiments herein are not limited to differential pixel units, as block tracking can be implemented using an image sensor that outputs intensity information.

[0131] exist Figure 7 In the illustrated embodiment, a block tracking engine 704 generates control signals indicating the addresses of pixel locations within one or more blocks to be tracked. The block tracking engine 704 can be constructed and operated similarly to the block tracking engine 604. Here, the block tracking engine 704 provides control signals to a pixel filter 706, which passes image information from only those pixels within the block to an output 710. As shown, the output 710 is coupled to an image processor 708, which can further process the image information of the pixels within the block using techniques as described herein or in other suitable manners.

[0132] Figure 8 , which illustrates an image sensor 800 according to some embodiments. Image sensor 800 may include an image array 802, which may be a conventional image array that scans out pixel intensity values. The image array may be used to provide differential image information as described herein using a comparator 806. Similar to comparator 606, comparator 806 may calculate differential information based on stored pixel values. Pixel filter 808 may pass selected values from those differential values to output 812. Like pixel filter 706, pixel filter 808 may receive control inputs from a block tracking engine 804. Block tracking engine 804 may be similar to block tracking engine 704. Output 812 may be coupled to an image processor 810. Some or all of the above-described components of image sensor 800 may be implemented in a single integrated circuit. Alternatively, the components may be distributed across one or more integrated circuits or other components.

[0133] Image sensors as described herein may operate as part of an augmented reality system to maintain information about movable objects or other information about the physical world that is useful in realistically rendering images of virtual objects in conjunction with information about the physical environment. Figure 9 A method 900 for image sensing is shown in accordance with some embodiments.

[0134] At least a portion of method 900 may be performed to operate an image sensor including, for example, image sensor 400, 600, 700, or 800. Method 900 may begin by receiving (act 902) imaging information from one or more inputs including, for example, image sensor 412, inertial sensor 414, and stored information 416. Method 900 may include identifying (act 904) one or more patches on an image output of the image sensing system based at least in part on the received information. An example of act 904 is shown in FIG. Figure 10 In some embodiments, method 900 may include calculating (act 906) movement trajectories of one or more segments. An example of act 906 is shown in Figure 11 Shown in.

[0135] The method 900 may also include setting (action 908) the image sensing system based at least in part on the identified one or more blocks and / or their estimated movement trajectories. This setting can be achieved by enabling a portion of the pixel cells of the image sensing system through, for example, a comparator 606, a pixel filter 706, etc., based at least in part on the identified one or more blocks and / or their estimated movement trajectories. In some embodiments, the comparator 606 may receive a first reference voltage value for a pixel cell corresponding to a selected block on the image, and a second reference voltage value for a pixel cell that does not correspond to any selected block on the image. The comparator 606 may set the second reference voltage to be much higher than the first reference voltage so that an unreasonable light intensity change sensed by a pixel cell of a comparator cell having the second reference voltage may result in an output of the pixel cell. In some embodiments, the pixel filter 706 may disable the output of a pixel cell having an address (e.g., row and column) that does not correspond to any selected block on the image.

[0136] Figure 10 Shown is block identification 904 in accordance with some embodiments. Block identification 904 may include segmenting (act 1002) one or more images from one or more inputs based at least in part on color, light intensity, angle of arrival, depth, and semantics.

[0137] Block recognition 904 may also include recognizing (action 1004) one or more objects in one or more images. In some embodiments, object recognition 1004 may be based at least in part on predetermined features of the object, including, for example, hands, eyes, and facial features. In some embodiments, object recognition 1004 may be based on one or more virtual objects. For example, a virtual animal character is walking on a physical pencil. Object recognition 1004 may use a virtual animal character target as an object. In some embodiments, object recognition 1004 may be based at least in part on artificial intelligence (AI) training received by the image sensing system. For example, the image sensing system may be trained by reading images of cats of different types and colors, thereby learning the characteristics of cats and being able to recognize cats in the physical world.

[0138] Segment identification 904 may include generating (act 1006) segments based on one or more objects. In some embodiments, object segmentation 1006 may generate segments by computing a convex hull or bounding box of one or more objects.

[0139] Figure 11 Block trajectory estimation 906 is shown in accordance with some embodiments. Block trajectory estimation 906 may include predicting (act 1102) the motion of one or more blocks over time. The motion of one or more blocks may be caused by a variety of reasons, including, for example, moving objects and / or moving users. Motion prediction 1102 may include deriving the speed of movement of the moving objects and / or moving users based on received images and / or received AI training.

[0140] Block trajectory estimation 906 may include calculating (act 1104) trajectories of one or more blocks over time based at least in part on the predicted motion. In some embodiments, the trajectories may be calculated by modeling with a first-order linear equation, assuming that the moving object will continue to move in the same direction at the same speed. In some embodiments, the trajectories may be calculated by curve fitting or using heuristics, including pattern detection.

[0141] Figure 12 and Figure 13 Shows factors that can be applied in the calculation of the block trajectory. Figure 12 An example of a movable object is shown, in which the movable object is a moving object 1202 (e.g., a hand) that is moving relative to a user of an AR system. In this example, the user is wearing an image sensor as part of a head-mounted display 62. In this example, the user's eye 49 is looking straight ahead, so that the image array 1200 captures the field of view (FOV) of the eye 49 relative to one viewpoint 1204. The object 1202 is within the FOV and therefore appears in the corresponding pixel in the array 1200 by creating intensity variations.

[0142] Array 1200 has a plurality of pixels 1208 arranged in an array. For the system to track hand 1202, a block 1206 in the array containing object 1202 at time t0 may include a portion of the plurality of pixels. If object 1202 is moving, the position of the block capturing it will change over time. This change can be captured in a block trajectory, from block 1206 to blocks X and Y used later.

[0143] The segment trajectory can be estimated, for example, in action 906, by identifying a feature 1210 of an object in the segment, such as a fingertip in the example shown. A motion vector 1212 can be calculated for the feature. In this example, the trajectory is modeled as a first-order linear equation, and the prediction is based on the assumption that the object 1202 will continue on the same segment trajectory 1214 over time, resulting in segment positions X and Y at each of two consecutive times.

[0144] As the position of the tiles changes, the image of the moving object 1202 remains within the tiles. Even though the image information is limited to information collected using pixels within the tiles, the image information is sufficient to represent the movement of the moving object 1202. This is true regardless of whether the image information is intensity information or differential information, such as that generated by a differential circuit. For example, in the case of a differential circuit, an event indicating an increase in intensity can occur when the image of the moving object 1202 moves over a pixel. Conversely, an event indicating a decrease in intensity can occur when the image of the moving object 1202 passes over a pixel. A pattern of pixels with increasing and decreasing events can serve as a reliable indicator of the movement of the moving object 1202, and because the amount of data indicating the events is relatively small, it can be updated quickly with low latency. As a specific example, such a system can result in a realistic XR system that tracks a user's hand and changes the rendering of a virtual object, thereby creating the feeling for the user that the user is interacting with the virtual object.

[0145] The positions of the tiles may change for other reasons, and any or all of these reasons may be reflected in the trajectory calculation. One such other change is the user's movement while wearing the image sensor. Figure 13 An example of a moving user is shown, which creates a changing viewpoint for the user as well as the image sensor. Figure 13 13, the user may initially be looking directly ahead at an object at viewpoint 1302. In this configuration, pixel array 1300 of the image array will capture the object in front of the user. The object in front of the user may be in segment 1312.

[0146] The user can then change their viewpoint, for example by turning their head. The viewpoint can be changed to viewpoint 1304. Even if the object that was previously directly in front of the user has not moved, it will be at a different location within the field of view at the user's viewpoint 1304. It will also be at a different point within the field of view of the image sensor worn by the user, and therefore at a different location within image array 1300. For example, the object may be contained within the tile at location 1314.

[0147] If the user further changes their viewpoint to viewpoint 1306, and the image sensor moves with the user, the location of the object that was previously directly in front of the user will be imaged at a different point within the field of view of the image sensor worn by the user, and therefore at a different location within image array 1300. For example, the object may be contained within the tile at location 1316.

[0148] It can be seen that as the user further changes their viewpoint, the position of the tile in the image array required to capture the object moves further. The trajectory of this movement from position 1312 to position 1314 to position 1316 can be estimated and used to track the future position of the tile.

[0149] Trajectories can be estimated in other ways. For example, measurements from inertial sensors can indicate the acceleration and velocity of the user's head when the user has viewpoint 1302. This information can be used to predict the trajectory of the tiles within the image array based on the movement of the user's head.

[0150] Tile trajectory estimation 906 can predict, based at least in part on these inertial measurements, that the user will have viewpoint 1304 at time t1 and viewpoint 1306 at time t2. Thus, tile trajectory estimation 906 can predict that tile 1308 can move to tile 1310 at time t1 and to tile 1312 at time t2.

[0151] As an example of this approach, it can be used to provide accurate and low-latency head pose estimation in an AR system. The tiles can be localized to images containing stationary objects within the user's environment. As a specific example, processing of image information can identify the corner of a picture frame hanging on a wall as a recognizable and stationary object to be tracked. The processing can focus the tiles on that object. Combined with the above Figure 12As with the moving object 1202 described herein, relative motion between the object and the user's head will generate events that can be used to calculate relative motion between the user and the tracked object. In this example, since the tracked object is stationary, the relative motion indicates movement of the imaging array being worn by the user. Therefore, this motion indicates changes in the user's head pose relative to the physical world and can be used to maintain accurate calculations of the user's head pose, which can be used to realistically render virtual objects. Because the imaging arrays described herein can provide rapid updates, with relatively small amounts of data per update, the calculations for rendering virtual objects remain accurate (they can be performed quickly and updated frequently).

[0152] Return to reference Figure 11 , the block trajectory estimation 906 may include adjusting (action 1106) the size of at least one block based at least in part on the calculated block trajectory. For example, the size of the block may be set to be large enough so that it includes a number of pixels in which the image of the movable object or at least a portion of the object for which image information is to be generated will be projected. The block may be set to be slightly larger than the projected size of the image of the portion of the object of interest so that if there are any errors in estimating the trajectory of the block, the block can still include the relevant portion of the image. When the object moves relative to the image sensor, the image size (in pixels) of the object can change based on the distance, the angle of incidence, the orientation of the object, or other factors. The processor that defines the blocks associated with the object can set the size of the blocks, for example, by measuring based on other sensor data or calculating the size of the blocks associated with the object based on a world model. Other parameters of the blocks, such as their shape, can be similarly set or updated.

[0153] Figure 14 An image sensing system 1400 is shown configured for use in an XR system according to some embodiments. Figure 4 ), image sensing system 1400 includes circuitry to selectively output values within a block and can be configured to output events for pixels within the block, as also described above. Additionally, image sensing system 1400 is configured to selectively output measured intensity values, which can be output for a complete image frame.

[0154] In the illustrated embodiment, separate outputs are shown for events and intensity values generated using the DVS technique described above. Outputs generated using the DVS technique can be represented as AER 1418 using the representation described above in conjunction with AER 418. Outputs representing intensity values can be output via outputs designated herein as APS 1420. Those intensity outputs can be for blocks or for the entire image frame. The AER and APS outputs can be activated simultaneously. However, in the illustrated embodiment, image sensor 1400 operates in either a mode outputting events or a mode outputting intensity information at any given time. Systems utilizing such image sensors can selectively utilize event outputs and / or intensity information.

[0155] Image sensing system 1400 may include image sensor 1402, which may include image array 1404. Image array 1404 may include a plurality of pixels 1500, each pixel responsive to light. Sensor 1402 may also include circuitry for accessing the pixel cells. Sensor 1402 may also include circuitry for generating inputs to the access circuitry to control a mode for reading information from the pixel cells in image array 1404.

[0156] In the illustrated embodiment, image array 1404 is configured as an array having multiple rows and columns of pixel cells that are accessible in both readout modes. In such an embodiment, access circuitry may include a row address encoder / decoder 1406, a column address encoder / decoder 1408 that controls column select switches 1422, and / or registers 1424 that may temporarily store information about incident light sensed by one or more corresponding pixel cells. A tile tracking engine 1410 may generate inputs to the access circuitry to control which pixel cells provide image information at any given time.

[0157] In some embodiments, image sensor 1402 may be configured to operate in rolling shutter mode, global shutter mode, or both.For example, tile tracking engine 1410 may generate inputs to access circuitry to control a readout mode of image array 1402.

[0158] When the sensor 1402 is operated in a rolling shutter readout mode, a single column of pixel cells is selected during each system clock by, for example, closing a single column switch 1422 from among the plurality of column switches. During the system clock, the selected column of pixel cells is exposed and read out to the APS 1420. To generate an image frame using the rolling shutter mode, the columns of pixel cells in the sensor 1402 can be read out column by column and then processed by the image processor to generate an image frame.

[0159] When sensor 1402 is operated in global shutter mode, for example, within a single system clock, columns of pixel cells are exposed simultaneously and the information is stored in registers 1424, allowing information captured by pixel cells in multiple columns to be read out simultaneously to APS 1420b. This readout mode allows image frames to be directly output without further data processing. In the example shown, information about the incident light sensed by a pixel cell is stored in a corresponding register 1424. It should be understood that multiple pixel cells can share a single register 1424.

[0160] In some embodiments, sensor 1402 can be implemented in a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, image array 1404 can be implemented in a single integrated circuit. Tile tracking engine 1410, row address encoder / decoder 1406, column address encoder / decoder 1408, column select switch 1422, and / or register 1424 can be implemented in a second single integrated circuit, such as a driver for image array 1404. Alternatively or additionally, some or all of the functionality of tile tracking engine 1410, row address encoder / decoder 1406, column address encoder / decoder 1408, column select switch 1422, and / or register 1424 can be distributed to other digital processors in the AR system.

[0161] Figure 15 An exemplary pixel cell 1500 is shown. In the embodiment shown, each pixel cell can be configured to output either event or intensity information. However, it should be understood that in some embodiments, the image sensor can be configured to output both types of information simultaneously.

[0162] Both event information and intensity information are based on the output of the photodetector 504, as described above in conjunction with Figure 5A Pixel unit 1500 includes circuitry for generating event information. This circuitry includes photoreceptor circuitry 502, differential circuitry 506, and comparator 508, also described above. When in a first state, switch 1520 connects photodetector 504 to the event generation circuitry. Switch 1520 or other control circuitry can be controlled by a processor controlling the AR system, thereby providing a relatively small amount of image information over a relatively long period of time while the AR system is operating.

[0163] Switch 1520 or other control circuitry can also be controlled to configure pixel cell 1500 to output intensity information. In the illustrated information, the intensity information is provided as a complete image frame, represented continuously as a stream of pixel intensity values for each pixel in the image array. To operate in this mode, switch 1520 in each pixel cell can be set to a second position that exposes the output of photodetector 504 after passing through amplifier 510 so that it can be connected to an output line.

[0164] In the illustrated embodiment, the output lines are illustrated as column lines 1510. There can be one such column line for each column in the image array. Each pixel cell in the column can be coupled to a column line 1510, but the pixel array can be controlled so that one pixel cell is coupled to a column line 1510 at a time. Switch 1530, one such switch in each pixel cell, controls when the pixel cell 1500 is connected to its corresponding column line 1510. Access circuitry, such as row address decoder 410, can close switch 1530 to ensure that only one pixel cell is connected to each column line at a time. Switches 1520 and 1530 can be implemented using one or more transistors as part of an image array or the like.

[0165] Figure 15 15. Additional components that may be included in each pixel cell according to some embodiments are shown. A sample and hold circuit (S / H) 1532 may be connected between the photodetector 504 and the column line 1510. When present, the S / H 1532 may enable the image sensor 1402 to operate in global shutter mode. In global shutter mode, a trigger signal is sent simultaneously to each pixel cell in the array. Within each pixel cell, the S / H 1532 captures a value indicating intensity upon the trigger signal. The S / H 1532 stores the value and generates an output based on the value until the next value is captured.

[0166] like Figure 15 As shown, when switch 1530 is closed, a signal representing the value stored by S / H 1532 can be coupled to column line 1510. The signal coupled to the column line can be processed to produce the output of the image array. For example, the signal can be buffered and / or amplified in amplifier 1512 at the end of column line 1510 and then applied to analog-to-digital converter (A / D) 1514. The output of A / D 1514 can be passed to output 1420 through other readout circuitry 1516. Readout circuitry 1516 can include, for example, column switch 1422. Other components within readout circuitry 1516 can perform other functions, such as serializing the multi-bit output of A / D 1514.

[0167] Those skilled in the art will understand how to implement circuits to perform the functions described herein. S / H 1532 can be implemented as, for example, one or more capacitors and one or more switches. However, it should be understood that S / H 1532 can use other components or be implemented in a manner different from that described herein. Figure 15 It should be understood that other components can be implemented in addition to those shown. For example, Figure 15 One amplifier and one A / D converter per column is shown. In other embodiments, there may be one A / D converter shared across multiple columns.

[0168] In a pixel array configured for a global shutter, each S / H 1532 can simultaneously store intensity values reflecting image information. These values can be stored during the readout phase, as the values stored in each pixel are read out continuously. For example, continuous readout can be achieved by connecting the S / H 1532 of each pixel cell in a row to its corresponding column line. The values on the column lines can then be passed one at a time to the APS output 1420. This information flow can be controlled by sequencing the opening and closing of column switches 1422. For example, this operation can be controlled by column address decoder 1408. Once the values of each pixel in a row have been read out, the pixel cells in the next row can be connected to the column lines in their place. These values can be read out one column at a time. This process of reading out values for one row at a time can be repeated until the intensity values of all pixels in the image array have been read out. In embodiments where intensity values are read out in one or more blocks, this process is completed when the values of the pixel cells within the block have been read out.

[0169] The pixel cells can be read out in any suitable order. For example, the rows can be interleaved so that every other row is read out in sequence. Nevertheless, the AR system can still process the image data into a frame of image data by deinterleaving the data.

[0170] In embodiments where S / H 1532 is not present, values may still be read sequentially from each pixel cell as rows and columns of values are scanned out. However, the value read from each pixel cell may represent the value in that cell at the time it was captured as part of the readout process, such as the light intensity detected at the cell's photodetector when the value was applied to, for example, A / D 1514. Thus, in a rolling shutter, the pixels of an image frame may represent images incident on the image array at slightly different times. For an image sensor outputting a full frame at a rate of 30 Hz, the time difference between the first pixel value of a captured frame and the last pixel value of the frame may differ by as much as 1 / 30 of a second, which is imperceptible for many applications.

[0171] For some XR functions, such as object tracking, the XR system can perform computations on image information collected using an image sensor using a rolling shutter. Such computations can interpolate between consecutive image frames to calculate an interpolated value for each pixel, representing an estimated value of the pixel at a point in time between the consecutive frames. The same time can be used for all pixels so that, through computation, the interpolated image frames contain pixels representing the same point in time, such as can be generated using an image sensor with a global shutter. Alternatively, a global shutter image array can be used for one or more image sensors in a wearable device forming part of the XR system. Using a global shutter for full or partial image frames can avoid interpolation of other processing that may be performed to compensate for variations in capture time in image information captured using a rolling shutter. Thus, even if the image information is used to track the movement of an object, such as can occur for processing such as tracking hands or other movable objects, determining the head pose of a user of a wearable device in AR, or even using a camera on a wearable device to construct an accurate representation of the physical environment, where the device may move as the image information is collected, interpolation computations can be avoided.

[0172] Differentiated pixel units

[0173] In some embodiments, each pixel cell in a sensor array can be identical. For example, each pixel cell can respond to a broad spectrum of visible light. Thus, each photodetector can provide image information indicating the intensity of visible light. In this scenario, the output of the image array can be a "grayscale" output representing the amount of visible light incident on the image array.

[0174] In other embodiments, the pixel cells can be differentiated. For example, different pixel cells in a sensor array can output image information indicating the intensity of light in specific parts of the spectrum. A suitable technique for differentiating pixel cells is to place a filter element in the light path leading to the photodetector in the pixel cell. The filter element can be bandpass, for example, allowing visible light of a specific color to pass. Applying such a filter to a pixel cell configures the pixel cell to provide image information indicating the intensity of light of the color corresponding to the filter.

[0175] Filters can be applied to pixel cells regardless of their structure. For example, they can be applied to pixel cells in a sensor array with a global shutter or a rolling shutter. Similarly, filters can be applied to pixel cells configured to output intensity or intensity variations using DVS technology.

[0176] In some embodiments, a filter element that selectively passes primary colors of light can be mounted above the photodetector in each pixel cell in the sensor array. For example, a filter that selectively passes red, green, or blue light can be used. The sensor array can have multiple subarrays, each with one or more pixels configured to sense light of each primary color. In this way, the pixel cells in each subarray provide both intensity and color information related to the object being imaged by the image sensor.

[0177] The inventors have recognized and understood that in an XR system, some functions require color information, while some functions can be performed using grayscale information. A wearable device equipped with an image sensor to provide image information for operation of the XR system may have multiple cameras, some of which may be formed from image sensors that can provide color information. Other cameras may be grayscale cameras. The inventors have recognized and understood that a grayscale camera may consume less power, be more sensitive in low light conditions, output data faster, and / or output less data than a camera formed from a comparable image sensor configured to sense color, to represent the same range of the physical world at the same resolution. However, the image information output by the grayscale camera is sufficient for many functions performed in the XR system. Therefore, the XR system may be configured with grayscale cameras and color cameras, primarily using one or more grayscale cameras, and selectively using color cameras.

[0178] For example, an XR system may collect and process image information to create a navigable world model. This processing may use color information, which may enhance the effectiveness of certain functions, such as distinguishing objects, identifying surfaces associated with the same object, and / or recognizing objects. Such processing may be performed or updated from time to time, such as when a user first turns on the system, moves to a new environment (e.g., walks into another room), or otherwise detects a change in the user's environment.

[0179] Other functions are not significantly improved by using color information. For example, once a navigable world model is created, the XR system can use images from one or more cameras to determine the orientation of the wearable device relative to features in the navigable world model. For example, such a function can be performed as part of head pose tracking. Some or all of the cameras used for such functions can be grayscale. Because head pose tracking is performed frequently, and in some embodiments continuously, while the XR system is running, using one or more grayscale cameras for this function can provide considerable power savings, reduced computation, or other benefits.

[0180] Similarly, at various times during XR system operation, the system may use stereoscopic information from two or more cameras to determine the distance to a movable object. Such functionality may require high-speed processing of image information as part of tracking a user's hands or other movable objects. Using one or more grayscale cameras for this functionality may provide lower latency or other benefits associated with processing high-resolution image information.

[0181] In some embodiments of an XR system, the XR system may have both a color camera and at least one grayscale camera, and may selectively enable the grayscale and / or color cameras based on the functions for which image information from those cameras will be used.

[0182] Pixel cells in an image sensor can be differentiated based on a spectrum other than the spectrum to which the pixel cells are sensitive. In some embodiments, some or all pixel cells can generate an output having an intensity that indicates the angle of arrival of light incident on the pixel cell. The angle of arrival information can be processed to calculate the distance to the imaged object.

[0183] In such an embodiment, the image sensor can passively acquire depth information. Passive depth information can be obtained by placing a component in the optical path to a pixel cell in the array, causing the pixel cell to output information indicating the angle of arrival of the light illuminating the pixel cell. An example of such a component is a transmissive diffraction mask (TDM) filter.

[0184] The angle of arrival information can be converted into distance information through calculation, indicating the distance to the object reflecting the light. In some embodiments, the pixel cells configured to provide the angle of arrival information can be interspersed with pixel cells that capture the intensity of light of one or more colors. As a result, the angle of arrival information, and therefore the distance information, can be combined with other image information about the object.

[0185] In some embodiments, one or more sensors may be configured to acquire information about physical objects in a scene at high frequency and low latency using compact and low-power components. For example, the power consumption of an image sensor may be less than 50 milliwatts, enabling the device to be powered by a battery that is small enough to be used as part of a wearable system. The sensor may be an image sensor configured to passively acquire depth information as a supplement to or alternative to image information indicating the intensity of one or more colors and / or changes in intensity information. Such a sensor may also be configured to provide a small amount of data to provide a differential output by using block tracking or by using DVS techniques.

[0186] Passive depth information can be obtained by configuring an image array, such as an image array incorporating any one or more of the techniques described herein, with components that adapt one or more pixel cells in the array to output information indicative of a light field emanating from an object being imaged. This information can be based on the angle of arrival of light illuminating the pixel. In some embodiments, pixel cells such as those described above can be configured to output an indication of the angle of arrival by placing a plenoptic component in the light path to the pixel cell. An example of a plenoptic component is a transmissive diffraction mask (TDM). The angle of arrival information can be computationally converted into distance information indicating the distance to the object from which light was reflected to form the image being captured. In some embodiments, the pixel cells configured to provide angle of arrival information can be interspersed with pixel cells that capture light intensities in grayscale or one or more colors. As a result, the angle of arrival information can also be combined with other image information about the object.

[0187] Figure 16 1. A pixel subarray 100 is shown in accordance with some embodiments. In the embodiment shown, the subarray has two pixel cells, but the number of pixel cells in the subarray is not limiting of the present invention. Here, a first pixel cell 121 and a second pixel cell 122 are shown, one of which is configured to capture angle of arrival information (first pixel cell 121), but it should be understood that the number and position of pixel cells within the array that are configured to measure angle of arrival information can vary. In this example, the other pixel cell (second pixel cell 122) is configured to measure the intensity of one color of light, but other configurations are possible, including pixel cells that are sensitive to different colors of light or one or more pixel cells that are sensitive to a broad spectrum of light, such as in a grayscale camera.

[0188] Figure 16 The first pixel unit 121 of the pixel subarray 100 includes an arrival angle to intensity converter 101, a photodetector 105, and a differential readout circuit 107. The second pixel unit 122 of the pixel subarray 100 includes a filter 102, a photodetector 106, and a differential readout circuit 108. It should be understood that it is not Figure 16 All components shown in need to be included in each embodiment. For example, some embodiments may not include differential readout circuits 107 and / or 108, and some embodiments may not include filter 102. In addition, may include Figure 16Other components not shown in the figure. For example, some embodiments may include a polarizer arranged to allow light of a specific polarization to reach the photodetector. As another example, some embodiments may include a scan output circuit instead of the differential readout circuit 107, or include a scan output circuit in addition to the differential readout circuit 107. As another example, the first pixel unit 121 may further include a filter so that the first pixel 121 measures the arrival angle and intensity of light of a specific color incident on the first pixel 121.

[0189] The arrival angle to intensity converter 101 of the first pixel 121 is an optical component that converts the angle θ of the incident light 111 into an intensity that can be measured by a photodetector. In some embodiments, the arrival angle to intensity converter 101 may include a refractive optical device. For example, one or more lenses can be used to convert the incident angle of light into a position on the image plane, that is, the amount of incident light detected by one or more pixel units. In some embodiments, the arrival angle to position intensity converter 101 may include a diffraction optical device. For example, one or more diffraction gratings (e.g., a transmissive diffraction mask (TDM)) can convert the incident angle of light into an intensity that can be measured by a photodetector below the TDM.

[0190] The photodetector 105 of the first pixel unit 121 receives the incident light 110 that passes through the arrival angle-to-intensity converter 101 and generates an electrical signal based on the intensity of the light incident on the photodetector 105. The photodetector 105 is located at an image plane associated with the arrival angle-to-intensity converter 101. In some embodiments, the photodetector 105 may be a single pixel of an image sensor, such as a CMOS image sensor.

[0191] The differential readout circuit 107 of the first pixel 121 receives a signal from the photodetector 105 and outputs an event only when the amplitude of the electrical signal from the photodetector is different from the amplitude of the previous signal from the photodetector 105, implementing the DVS technique as described above.

[0192] The second pixel unit 122 includes an optical filter 102 for filtering the incident light 112 so that only light within a specific wavelength range passes through the optical filter 102 and is incident on the photodetector 106. The optical filter 102 can be, for example, a bandpass filter that allows one of red, green, or blue light to pass and rejects light of other wavelengths and / or can limit the IR light reaching the photodetector 106 to only a specific portion of the spectrum.

[0193] In this example, the second pixel cell 122 also includes a photodetector 106 and a differential readout circuit 108 , which may operate similarly to the photodetector 105 and the differential readout circuit 107 of the first pixel cell 121 .

[0194] As described above, in some embodiments, an image sensor may include an array of pixels, each pixel being associated with a photodetector and readout circuitry. A subset of the pixels may be associated with an angle-of-arrival-to-intensity converter for determining the angle of detection light incident on the pixel. Other subsets of the pixels may be associated with filters for determining color information about the scene being viewed or that may selectively pass or block light based on other characteristics.

[0195] In some embodiments, a single photodetector and two diffraction gratings of different depths can be used to determine the angle of arrival of light. For example, light can be incident on a first TDM, the angle of arrival is converted to position, and a second TDM can be used to selectively pass light incident at a specific angle. This arrangement can take advantage of the Talbot effect, a near-field diffraction effect in which when a plane wave is incident on a diffraction grating, an image of the diffraction grating is produced at a certain distance from the diffraction grating. If a second diffraction grating is placed at the image plane that forms the image of the first diffraction grating, the angle of arrival can be determined based on the light intensity measured by a single photodetector located after the second grating.

[0196] Figure 17A A first arrangement of pixel cells 140 is shown, comprising a first TDM 141 and a second TDM 143 aligned with one another such that the ridges and / or regions of increased refractive index of the two gratings are aligned in the horizontal direction (Δs=0), where Δs is the horizontal offset between the first TDM 141 and the second TDM 143. The first TDM 141 and the second TDM 143 can both have the same grating period d, and the two gratings can be separated by a distance / depth z. The depth z at which the second TDM 143 is located relative to the first TDM 141 is known as the Talbot length and can be determined by the grating period d and the wavelength λ of the light being analyzed, and is given by:

[0197]

[0198] like Figure 17AAs shown, incident light 142 with a zero-degree arrival angle is diffracted by first TDM 141. Second TDM 143 is located at a depth equal to the Talbot length, creating an image of first TDM 141, causing most of incident light 142 to pass through second TDM 143. An optional dielectric layer 145 can separate second TDM 143 from photodetector 147. When light passes through dielectric layer 145, photodetector 147 detects the light and generates an electrical signal whose characteristics (e.g., voltage or current) are proportional to the intensity of the light incident on the photodetector. On the other hand, when incident light 144 with a non-zero arrival angle θ is also diffracted by first TDM 141, second TDM 143 blocks at least a portion of incident light 144 from reaching photodetector 147. The amount of incident light reaching photodetector 147 depends on the arrival angle θ, with larger angles resulting in less light reaching the photodetector. The dashed line generated by light 144 indicates that the amount of light reaching photodetector 147 is attenuated. In some cases, light 144 may be completely blocked by diffraction grating 143. Thus, two TDMs may be used, using a single photodetector 147 to obtain information about the angle of arrival of the incident light.

[0199] In some embodiments, information obtained from neighboring pixel cells that do not have an arrival angle to intensity converter can provide an indication of the intensity of the incident light and can be used to determine the portion of the incident light that passes through the arrival angle to intensity converter. Based on this image information, the arrival angle of the light detected by the photodetector 147 can be calculated, as described in more detail below.

[0200] Figure 17B A second arrangement of pixel cells 150 is shown, comprising a first TDM 151 and a second TDM 153 that are misaligned with each other such that the ridges and / or regions of increased refractive index of the two gratings are misaligned in the horizontal direction (Δs≠0), where Δs is the horizontal offset between the first TDM 151 and the second TDM 153. The first TDM 151 and the second TDM 153 may have the same grating period d, and the two gratings may be separated by a distance / depth z. In combination with Figure 17A Unlike the discussed case where two TDMs are aligned, the misalignment results in incident light having an angle different from zero passing through the second TDM 153 .

[0201] like Figure 17BAs shown, incident light 152 with a zero-degree arrival angle is diffracted by first TDM 151. Second TDM 153 is located at a depth equal to the Talbot length, but due to the horizontal offset of the two gratings, at least a portion of light 152 is blocked by second TDM 153. The dashed line generated by light 152 shows that the amount of light reaching photodetector 157 is attenuated. In some cases, light 152 can be completely blocked by diffraction grating 153. On the other hand, incident light 154 with a non-zero arrival angle θ is diffracted by first TDM 151 but passes through second TDM 153. After passing through optional dielectric layer 155, photodetector 157 detects the light incident on photodetector 157 and generates an electrical signal whose characteristic (e.g., voltage or current) is proportional to the intensity of the light incident on the photodetector.

[0202] Pixel cells 140 and 150 have different output functions, where different light intensities are detected for different angles of incidence. However, in each case, the relationship is fixed and can be determined based on the design of the pixel cell or through measurements as part of a calibration process. Regardless of the exact transfer function, the measured intensity can be converted to an angle of arrival, which in turn can be used to determine the distance to the imaged object.

[0203] In some embodiments, different pixel cells of an image sensor may have different TDM arrangements. For example, a first subset of pixel cells may include a first horizontal offset between the gratings of the two TDMs associated with each pixel cell, while a second subset of pixel cells may include a second horizontal offset between the gratings of the two TDMs associated with each pixel cell, wherein the first offset is different from the second offset. Each subset of pixel cells with different offsets may be used to measure a different angle of arrival or a different range of angles of arrival. For example, a first subset of pixel cells may include something like Figure 17A The second pixel subset may include a TDM arrangement of pixel cells 140 similar to Figure 17B TDM arrangement of pixel cells 150.

[0204] In some embodiments, not all pixel cells of an image sensor include a TDM. For example, a subset of pixel cells may include a filter, while a different subset of pixel cells may include a TDM for determining angle of arrival information. In other embodiments, no filter is used, such that a first subset of pixel cells simply measures the total intensity of incident light, while a second subset of pixel cells measures angle of arrival information. In some embodiments, information related to the light intensity from nearby pixel cells without a TDM can be used to determine the angle of arrival of light incident on a pixel cell with one or more TDMs. For example, using two TDMs arranged to exploit the Talbot effect, the intensity of light incident on a photodetector following the second TDM is a sinusoidal function of the angle of arrival of the light incident on the first TDM. Therefore, if the total intensity of light incident on the first TDM is known, the angle of arrival of the light can be determined based on the intensity of the light detected by the photodetector.

[0205] In some embodiments, the configuration of pixel elements in a subarray can be selected to provide various types of image information with appropriate resolution. 18A to 18C An example arrangement of pixel cells in a pixel subarray of an image sensor is shown. The example shown is a non-limiting arrangement, as it should be understood that the inventors have devised alternative pixel arrangements. This arrangement can be repeated across an image array, which can contain millions of pixels. A subarray can include one or more pixel cells that provide information about the angle of arrival of incident light and one or more other pixel cells (with or without filters) that provide information about the intensity of the incident light.

[0206] Figure 18A is an example of a pixel subarray 160 that includes a first group of pixel cells 161 and a second group of pixel cells 163 that are different from each other and are rectangular rather than square. Pixel cells labeled "R" are pixel cells with red filters, allowing red incident light to pass through the filters to their associated photodetectors; pixel cells labeled "B" are pixel cells with blue filters, allowing blue incident light to pass through the filters to their associated photodetectors; and pixel cells labeled "G" are pixel cells with green filters, allowing green incident light to pass through the filters to their associated photodetectors. In example subarray 160, there are more green pixel cells than red or blue pixel cells, illustrating that the various types of pixel cells do not need to be present in equal proportions.

[0207] The pixel cells labeled A1 and A2 are pixels that provide angle of arrival information. For example, pixel cells A1 and A2 may include one or more gratings for determining angle of arrival information. The pixel cells that provide angle of arrival information may be configured similarly or may be configured differently, such as being sensitive to different ranges of angles of arrival or angles of arrival relative to different axes. In some embodiments, the pixels labeled A1 and A2 include two TDMs, and the TDMs of pixel cells A1 and A2 may be oriented in different directions, such as perpendicular to each other. In other embodiments, the TDMs of pixel cells A1 and A2 may be oriented parallel to each other.

[0208] In an embodiment using pixel subarray 160, color image data and angle of arrival information can be obtained. In order to determine the angle of arrival of light incident on pixel cell group 161, the electrical signal from the RGB pixel cell is used to estimate the total light intensity incident on pixel cell group 161. Using the fact that the light intensity detected by the A1 / A2 pixels varies with the angle of arrival in a predictable manner, the angle of arrival can be determined by comparing the total intensity (estimated based on the RGB pixel cells within the pixel group) with the intensity measured by the A1 and / or A2 pixel cells. For example, the intensity of light incident on the A1 and / or A2 pixels can vary sinusoidally with respect to the angle of arrival of the incident light. The angle of arrival of light incident on pixel cell group 163 is determined in a similar manner using the electrical signal generated by pixel cell group 163.

[0209] It should be understood that Figure 18A A particular embodiment of a sub-array is shown, and other configurations are possible. In some embodiments, for example, a sub-array may be only a group of pixel cells 161 or 163 .

[0210] Figure 18B 1 is an alternative pixel subarray 170 including a first group of pixel cells 171, a second group of pixel cells 172, a third group of pixel cells 173, and a fourth group of pixel cells 174. Each group of pixel cells 171-174 is square and has the same pixel cell arrangement therein, but in order to make it possible to have pixel cells for determining arrival angle information within different angular ranges or relative to different planes (for example, the TDMs of pixels A1 and A2 can be oriented perpendicular to each other). Each group of pixels 171-174 includes a red pixel cell (R), a blue pixel cell (B), a green pixel cell (G), and an arrival angle pixel cell (A1 or A2). Note that in the example pixel subarray 170, there are the same number of red / green / blue pixel cells in each group. In addition, it should be understood that the pixel subarray can be repeated in one or more directions to form a larger pixel array.

[0211] In an embodiment using pixel subarray 170, color image data and angle of arrival information may be obtained. To determine the angle of arrival of light incident on pixel cell group 171, the signals from the RGB pixel cells may be used to estimate the total light intensity incident on pixel cell group 171. Using the fact that the light intensity detected by the angle of arrival pixel cells has a sinusoidal or other predictable response with respect to the angle of arrival, the angle of arrival may be determined by comparing the total intensity (estimated from the RGB pixel cells) with the intensity measured by the A1 pixel. The angle of arrival of light incident on pixel cell groups 172-174 may be determined in a similar manner using the electrical signals generated by the pixel cells of each respective pixel group.

[0212] Figure 18C 1 . An alternative pixel subarray 180 includes a first group of pixel cells 181, a second group of pixel cells 182, a third group of pixel cells 183, and a fourth group of pixel cells 184. Each group of pixel cells 181-184 is square and has the same pixel cell arrangement, wherein no filters are used. Each group of pixel cells 181-184 includes: two "white" pixels (e.g., no filters so that red, blue, and green light are detected to form a grayscale image); an arrival angle pixel cell (A1) in which the TDM is oriented in a first direction; and an arrival angle pixel cell (A2) in which the TDM is oriented at a second spacing or in a second direction relative to the first direction (e.g., perpendicular). Note that there is no color information in the example pixel subarray 170. The resulting image is grayscale, illustrating that passive depth information can be obtained in a color or grayscale image array using the techniques described herein. As with the other subarray arrangements described herein, the pixel subarray arrangement can be repeated in one or more directions to form a larger pixel array.

[0213] In an embodiment using pixel subarray 180, grayscale image data and angle of arrival information can be obtained. In order to determine the angle of arrival of light incident on pixel cell group 181, the total light intensity incident on pixel cell group 181 is estimated using the electrical signals from the two white pixels. Using the fact that the light intensity detected by the A1 and A2 pixels has a sinusoidal or other predictable response with respect to the angle of arrival, the angle of arrival can be determined by comparing the total intensity (estimated from the white pixels) with the intensity measured by the A1 and / or A2 pixel cells. The angle of arrival of light incident on pixel cell groups 182-184 can be determined in a similar manner using the electrical signals generated by the pixels of each corresponding pixel group.

[0214] In the above examples, the pixel cells are illustrated as squares and arranged in a square grid. Embodiments are not limited thereto. For example, in some embodiments, the shape of the pixel cells can be rectangular. Furthermore, the subarrays can be triangular or arranged on a diagonal or have other geometric shapes.

[0215] In some embodiments, the angle of arrival information is obtained using image processor 708 or a processor associated with local data processing module 70, and the processor can further determine the distance of the object based on the angle of arrival. For example, the angle of arrival information can be combined with one or more other types of information to obtain the distance of the object. In some embodiments, objects in grid model 46 can be associated with the angle of arrival information from the pixel array. Grid model 46 can include the location of the object, including the distance from the user, which can be updated to a new distance value based on the angle of arrival information.

[0216] Using angle of arrival information to determine distance values can be particularly useful in scenarios where objects are close to the user. This is because changes in distance from the image sensor cause changes in the angle of arrival of light for nearby objects to be larger than similar magnitude distance changes for objects positioned further away from the user. Therefore, a processing module that utilizes passive distance information based on angle of arrival can selectively use this information based on an estimated object distance, and can utilize one or more other techniques to determine the distance to objects that exceed a threshold distance, such as, in some embodiments, up to 1 meter, up to 3 meters, or up to 5 meters. As a specific example, a processing module of an AR system can be programmed to use passive distance measurement that uses angle of arrival information for objects within 3 meters of the wearable device user, but for objects outside of this range, stereo image processing can be used that uses images captured by two cameras.

[0217] Similarly, pixels configured to detect angle of arrival information may be most sensitive to distance variations within an angular range from the normal to the image array. The processing module may similarly be configured to use distance information derived from angle of arrival measurements within this angular range, but use other sensors and / or other techniques to determine distances outside of this range.

[0218] One example application of determining the distance of an object from an image sensor is hand tracking. Hand tracking may be used in an AR system, for example, to provide a gesture-based user interface for system 80 and / or to allow a user to move virtual objects within an environment in an AR experience provided by system 80. The combination of an image sensor that provides angle of arrival information for accurate depth determination and a differential readout circuit for reducing the amount of processed data for determining the user's hand movement provides an efficient interface through which a user may interact with virtual objects and / or provide input to system 80. A processing module that determines the position of a user's hand may use distance information that is acquired using different techniques depending on the position of the user's hand in the field of view of the image sensor of the wearable device. According to some embodiments, hand tracking may be implemented as a form of tile tracking during the image sensing process.

[0219] Another application where depth information can be useful is in occlusion processing. Occlusion processing uses depth information to determine that certain portions of the physical world model do not need to or cannot be updated based on image information captured by one or more image sensors that collect image information about the user's surrounding physical environment. For example, if a first object is determined to be at a first distance from the sensor, system 80 may determine not to update the physical world model for distances greater than the first distance. For example, even if the model includes a second object at a second distance from the sensor, where the second distance is greater than the first distance, if the second object is behind the first object, the model information for that object may not be updated. In some embodiments, system 80 may generate an occlusion mask based on the position of the first object and only update the portions of the model that are not obscured by the occlusion mask. In some embodiments, system 80 may generate more than one occlusion mask for more than one object. Each occlusion mask may be associated with a corresponding distance from the sensor. For each occlusion mask, model information associated with objects whose distance from the sensor is greater than the distance associated with the corresponding occlusion mask will not be updated. By limiting the portion of the model that is updated at any given time, the speed of generating the AR environment and the amount of computing resources required to generate the AR environment are reduced.

[0220] Although not in 18A to 18C As shown in , some embodiments of the image sensor may include pixels with IR filters in addition to or instead of filters. For example, the IR filter may allow light of a wavelength, such as approximately 940 nm, to pass through and be detected by an associated photodetector. Some embodiments of the wearable device may include an IR light source (e.g., an IR LED) that emits light of the same wavelength as the wavelength associated with the IR filter (e.g., 940 nm). The IR light source and IR pixels may be used as an alternative way to determine the distance of an object from the sensor. By way of example and not limitation, the IR light source may be pulsed and time-of-flight measurements may be used to determine the distance of an object from the sensor.

[0221] In some embodiments, the system 80 can operate in one or more operating modes. A first mode can be a mode in which depth is determined using passive depth measurement, for example, based on the angle of arrival of light determined using pixels with an angle of arrival to intensity converters. A second mode can be a mode in which depth is determined using active depth measurement, for example, based on the time of flight of IR light measured using IR pixels of an image sensor. A third mode can use stereo measurements from two independent image sensors to determine the distance to an object. When the object is far from the sensor, such stereo measurements can be more accurate than the angle of arrival of light determined using pixels with an angle of arrival to intensity converters. Other appropriate methods of determining depth can be used for one or more additional depth determination operating modes.

[0222] In some embodiments, it may be preferable to use passive depth determination because such techniques use less power. However, the system may determine that it should operate in active mode under certain conditions. For example, if the visible light intensity detected by the sensor is below a threshold, it may be too dark to accurately perform passive depth determination. As another example, the object may be too far away for passive depth determination to be inaccurate. Therefore, the system can be programmed to elect to operate in a third mode, in which depth is determined based on stereo measurements of the scene using two spatially separated image sensors. As another example, determining the depth of an object based on the angle of arrival of light determined using pixels with angle of arrival to intensity converters may be inaccurate at the periphery of the image sensor. Therefore, if the object is being detected by pixels near the periphery of the image sensor, the system may elect to operate in a second mode, using active depth determination.

[0223] While the above-described image sensor embodiments use individual pixel cells with stacked TDMs to determine the angle of arrival of light incident on the pixel cells, other embodiments may use multiple pixel cells with a single TDM on all pixels in a group to determine the angle of arrival information. The TDM can project a light pattern on the sensor array that depends on the angle of arrival of the incident light. Multiple photodetectors associated with one TDM can more accurately detect the pattern because each of the multiple photodetectors is located at a different position in the image plane (the image plane includes the photodetectors that sense the light). The relative intensity sensed by each photodetector can indicate the angle of arrival of the incident light.

[0224] Figure 19A is a top view example of multiple photodetectors (in the form of a photodetector array 120, which may be a subarray of pixel cells of an image sensor) associated with a single transmissive diffraction mask (TDM) in accordance with some embodiments. Figure 19B is with Figure 19A The same photodetector array along Figure 19A . In the example shown, the photodetector array 120 includes 16 individual photodetectors 121, which may be within a pixel cell of the image sensor. The photodetector array 120 includes a TDM 123 disposed above the photodetectors. It will be appreciated that for clarity and simplicity, each group of pixel cells is shown with four pixels (e.g., forming a four-pixel by four-pixel grid). Some embodiments may include more than four pixel cells. For example, each group may include 16 pixel cells, 64 pixel cells, or any other number of pixels.

[0225] TDM 123 is located at a distance x from photodetector 121. In some embodiments, TDM 123 is formed on the top surface of dielectric layer 125, such as Figure 19B123. For example, as shown, the TDM 123 can be formed by ridges, or by valleys etched into the surface of the dielectric layer 125. In other embodiments, the TDM 123 can be formed within the dielectric layer. For example, portions of the dielectric layer can be modified to have a higher or lower refractive index relative to other portions of the dielectric layer, thereby producing a holographic phase grating. Light incident on the photodetector array 120 from above is diffracted by the TDMs, resulting in the angle of arrival of the incident light being converted to a position in the image plane at a distance x from the TDM 123, where the photodetector 121 is located. The incident light intensity measured at each photodetector 121 of the photodetector array can be used to determine the angle of arrival of the incident light.

[0226] Figure 20A An example of multiple photodetectors (in the form of a photodetector array 130 ) associated with multiple TDMs is shown in accordance with some embodiments. Figure 20B is with Figure 20A The same photodetector array is passed Figure 20A Cross-sectional view along line B. Figure 20C is with Figure 20A The same photodetector array is passed Figure 20A 1 . A cross-sectional view of line C of FIG. 1 . In the example shown, the photodetector array 130 includes 16 individual photodetectors, which may be within pixel cells of an image sensor. Four groups 131a, 131b, 131c, 131d of four pixel cells are shown. The photodetector array 130 includes four individual TDMs 133a, 133b, 133c, 133d, each TDM being disposed above an associated group of pixel cells. It will be appreciated that for clarity and simplicity, each group of pixel cells is illustrated using four pixel cells. Some embodiments may include more than four pixel cells. For example, each group may include 16 pixel cells, 64 pixel cells, or any other number of pixel cells.

[0227] Each TDM 133a-d is located at a distance x from the photodetectors 131a-d. In some embodiments, the TDMs 133a-d are formed on the top surface of the dielectric layer 135, such as Figure 20B133a-d. The holographic phase grating is shown in FIG. 135. For example, the TDMs 123a-d can be formed by ridges, as shown, or by valleys etched into the surface of the dielectric layer 135. In other embodiments, the TDMs 133a-d can be formed within the dielectric layer. For example, portions of the dielectric layer can be modified to have a higher or lower refractive index relative to other portions of the dielectric layer, thereby producing a holographic phase grating. Light incident on the photodetector array 130 from above is diffracted by the TDMs, resulting in the angle of arrival of the incident light being converted to a position in the image plane at a distance x from the TDMs 133a-d, at which position the photodetectors 131a-d are located. The incident light intensity measured at each photodetector 131a-d of the photodetector array can be used to determine the angle of arrival of the incident light.

[0228] TDMs 133a-d can be oriented in different directions from one another. For example, TDM 133a is perpendicular to TDM 133b. Thus, the light intensity detected by photodetector group 131a can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133a, and the light intensity detected by photodetector group 131b can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133b. Similarly, the light intensity detected by photodetector group 131c can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133c, and the light intensity detected by photodetector group 131d can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133d.

[0229] Pixel cells configured to passively acquire depth information can be integrated into an image array having the features described herein to support operations useful in an X-reality system. According to some embodiments, pixel cells configured to acquire depth information can be implemented as part of an image sensor for implementing a camera with a global shutter. For example, such a configuration can provide a full-frame output. A full frame can include image information from different pixels that simultaneously indicates depth and intensity. Using an image sensor configured in this manner, a processor can acquire depth information for an entire scene at once.

[0230] In other embodiments, the pixel cells of the image sensor providing depth information can be configured to operate according to the DVS technique described above. In such a scenario, an event can indicate a change in the depth of an object, as indicated by the pixel cell. The event output by the image array can indicate the pixel cell that detected the depth change. Alternatively or additionally, the event can include the value of the depth information of the pixel cell. Using an image sensor configured in this way, the processor can obtain depth information updates at a very high rate, thereby providing high temporal resolution.

[0231] In yet other embodiments, the image sensor can be configured to operate in either full-frame or DVS mode. In such embodiments, the processor processing image information from the image sensor can programmatically control the image sensor's operating mode based on the function being performed by the processor. For example, when performing a function involving tracking an object, the processor can configure the image sensor to output image information as DVS events. On the other hand, when processing an update world reconstruction, the processor can configure the image sensor to output full-frame depth information.

[0232] Wearable configuration

[0233] Multiple image sensors can be used in an XR system. Image sensors can be combined with optical components (e.g., lenses) and control circuitry to create a camera. Those image sensors can use one or more of the above-mentioned technologies to acquire imaging information, such as grayscale imaging, color imaging, global shutter, DVS technology, plenoptic pixel cells, and / or dynamic segmentation. Regardless of the imaging technology used, the resulting camera can be mounted on a support member to form a head-mounted device, which can include or be connected to a processor.

[0234] Figure 21 FIG is a schematic diagram of a head mounted device 2100 of a wearable display system consistent with the disclosed embodiments. Figure 21 As shown, the head-mounted device 2100 may include a display device, which includes a monocular 2110a and a monocular 2110b. The monocular 2110a and the monocular 2110b may be optical eyepieces or displays configured to transmit and / or display visual information to the user's eyes. The head-mounted device 2100 may also include a frame 2101, which may be similar to the frame 2101 described above. Figure 3B The described frame 64. The head mounted device 2100 may also include two cameras (camera 2120 and camera 2140) and additional components, such as a transmitter 2130a, a transmitter 2130b, an inertial measurement unit 2170a (IMU 2170a), and an inertial measurement unit 2170b (IMU 2170b).

[0235] Camera 2120 and camera 2140 are world cameras because they are oriented to image the physical world as seen by the user wearing the head mounted device 2100. In some embodiments, these two cameras may be sufficient to acquire image information about the physical world, and these two cameras may be the only world-facing cameras. The head mounted device 2100 may also include additional components, such as an eye tracking camera, as described above with respect to Figure 3B discussed.

[0236] Monocular 2110a and monocular 2110b can be mechanically coupled to a support member such as frame 2101 using techniques such as adhesives, fasteners, or press fits. Similarly, two cameras and accompanying components (e.g., a transmitter, an inertial measurement unit, an eye-tracking camera, etc.) can be mechanically coupled to frame 2101 using techniques such as adhesives, fasteners, press fits, etc. These mechanical couplings can be direct or indirect. For example, one or more cameras and / or one or more accompanying components can be directly attached to frame 2101. As an additional example, one or more cameras and / or one or more accompanying components can be directly attached to a monocular, which can then be attached to frame 2101. The mechanism of attachment is not intended to be limiting.

[0237] Alternatively, a monocular mirror assembly can be formed and then attached to the frame 2101. Each subassembly can include, for example, a support member to which the monocular 2110a or 2110b is attached. The IMU and one or more cameras can be similarly attached to the support member. Attaching the camera and IMU to the same support member can obtain inertial information about the camera based on the output of the IMU. Similarly, attaching the monocular to the same support member as the camera can spatially correlate image information about the world with the information rendered on the monocular.

[0238] The headset 2100 can be lightweight. For example, the headset 2100 can weigh between 30 and 300 grams. The headset 2100 can be made of materials that flex during use, such as plastic or thin metal components. Such materials can achieve a lightweight and comfortable headset that can be worn by a user for extended periods of time. Nevertheless, an XR system with such a lightweight headset can support high-precision stereo image analysis, which requires knowledge of the separation between cameras, using a calibration procedure that can be repeated while the headset is worn to compensate for any inaccuracies caused by flexing of the headset during use. In some embodiments, the lightweight headset can include a battery pack. The battery pack can include one or more batteries, which can be rechargeable or non-rechargeable. The battery pack can be built into the lightweight frame or removable. The battery pack and the lightweight frame can be formed as a single unit, or the battery pack can be formed as a separate unit from the lightweight frame.

[0239] The camera 2120 may include an image sensor and a lens. The image sensor may be configured to generate a grayscale image. The image sensor may be configured to acquire an image having a size between 1 million pixels and 4 million pixels. For example, the image sensor may be configured to acquire an image with a horizontal resolution of 1016 lines and a vertical resolution of 1016 lines. The image sensor may be configured to acquire an image repeatedly or periodically. For example, the image sensor may be configured to acquire an image at a frequency between 30 Hz and 120 Hz (e.g., 60 Hz). The image sensor may be a CMOS image sensor. The image sensor may be configured with a global shutter. As described above, with respect to Figure 14 and Figure 15 , a global shutter can enable each pixel to obtain intensity measurements simultaneously. In some embodiments, the camera 2120 can be configured as a plenoptic camera. For example, as described above with respect to Figure 3B and Figures 15 to 20C As discussed, components can be placed in the optical path to one or more pixel cells of an image sensor so that these pixel cells generate an output having an intensity that indicates the angle of arrival of light incident on the pixel cells. In such embodiments, the image sensor can passively acquire depth information. An example of a component suitable for placement in the optical path is a TDM filter. The processor can be configured to use the angle of arrival information to calculate the distance to the imaged object. For example, the angle of arrival information can be converted into distance information indicating the distance to the object from which the light was reflected. In some embodiments, the pixel cells configured to provide angle of arrival information can be interspersed with pixel cells that capture light intensity of one or more colors. As a result, the angle of arrival information, and therefore the distance information, can be combined with other image information about the object. In some embodiments, the camera 2120 can be configured to provide tile tracking functionality. The processor of the head-mounted device 2100 can be configured to provide instructions to the camera 2120 to limit image capture to a subset of pixels. In some embodiments, the image sensor can be an IMX418 image sensor or an equivalent image sensor.

[0240] Camera 2120 can be configured to have a wide field of view, consistent with the disclosed embodiments. For example, camera 2120 can include an equidistant lens (e.g., a fisheye lens). Camera 2120 can be angled inward on head-mounted device 2100. For example, a vertical plane through the center of field of view 2121, i.e., the field of view associated with camera 2120, can intersect and form an angle with a vertical plane through the centerline of head-mounted device 2100. This angle can be between 1 and 40 degrees. In some embodiments, field of view 2121 can have a horizontal field of view and a vertical field of view. The horizontal field of view can range from 90 degrees to 175 degrees, while the vertical field of view can range from 70 degrees to 125 degrees. In some embodiments, camera 2120 can be configured to have an angular pixel resolution between 1 and 5 arc minutes per pixel.

[0241] Emitters 2130a and 2130b may be configured to emit light of a specific wavelength in low light conditions and / or when active depth sensing is being performed by the head mounted device 2100. Emitters 2130a and 2130b may be configured to emit light of a specific wavelength. This light may be reflected by physical objects in the physical world around the user. The head mounted device 2100 may be configured with sensors to detect this reflected light, including image sensors as described herein. In some embodiments, these sensors may be incorporated into at least one of the cameras 2120 or 2140. For example, as described above with respect to 18A to 18C As described above, the cameras may be configured with detectors corresponding to emitter 2130a and / or emitter 2130b. For example, the cameras may include pixels configured to detect light emitted by emitter 2130a and emitter 2130b.

[0242] Emitters 2130a and 2130b can be configured to emit IR light, consistent with the disclosed embodiments. The wavelength of the IR light can be between 900 nanometers and 1 micron. The IR light can be, for example, a 940 nm light source, with the emitted light energy concentrated near 940 nm. Emitters emitting light at other wavelengths can alternatively or additionally be used. For example, for a system intended for indoor use only, an emitter emitting light concentrated near 850 nm can be used. At least one of cameras 2120 or 2140 can include one or more IR filters disposed over at least a subset of pixels in the camera's image sensor. The filters can pass light of the wavelength emitted by emitters 2130a and / or 2130b while attenuating light of other wavelengths. For example, the IR filters can be notch filters that pass IR light with a wavelength matching the wavelength of the emitters. Notch filters can significantly attenuate other IR light. In some embodiments, the notch filters can be IR notch filters that block IR light while allowing light from the emitters to pass. The IR notch filter can also allow light outside the IR band to pass through. This notch filter can enable the image sensor to receive visible light and light from the emitter that has been reflected from objects in the image sensor's field of view. In this way, a subset of pixels can function as a detector for IR light emitted by emitter 2130a and / or emitter 2130b.

[0243] In some embodiments, the processor of the XR system can selectively enable the emitters, for example to enable imaging in low light conditions. The processor can process image information generated by one or more image sensors and can detect whether images output by those image sensors without the emitters being activated provide sufficient information about objects in the physical world. The processor can enable the emitters in response to detecting that the images do not provide sufficient image information due to low ambient light conditions. For example, the emitters can be turned on when stereo information is being used to track an object and the lack of ambient light causes insufficient contrast between features of the tracked object to accurately determine distance using stereo imaging techniques.

[0244] Alternatively or in addition, emitter 2130a and / or emitter 2130b can be configured to perform active depth measurements, for example by emitting light in short pulses. The wearable display system can be configured to perform time-of-flight measurements by detecting reflections of such pulses from objects in the illumination field 2131a of emitter 2130a and / or the illumination field 2131b of emitter 2130b. These time-of-flight measurements can provide additional depth information for tracking objects or updating a traversable world model. In other embodiments, one or more emitters can be configured to emit patterned light, and the XR system can be configured to process images of objects illuminated by the patterned light. Such processing can detect changes in the pattern, which can reveal the distance to the object.

[0245] In some embodiments, the range of the illumination field associated with the emitters may be sufficient to illuminate at least the field of view of the camera to obtain image information about the object. For example, the emitters may collectively illuminate a central field of view 2150. In the illustrated embodiment, emitters 2130a and 2130b may be positioned to illuminate illumination fields 2131a and 2131b, which together span the range over which active illumination may be provided. In this example embodiment, two emitters are shown, but it should be understood that more or fewer emitters may be used to span the desired range.

[0246] In some embodiments, emitters, such as emitters 2130a and 2130b, can be turned off by default, but can be enabled when additional illumination is needed to obtain more information than passive imaging can obtain. The wearable display system can be configured to enable emitter 2130a and / or emitter 2130b when additional depth information is needed. For example, when the wearable display system detects that sufficient depth information for tracking hand or head gestures cannot be obtained using stereo image information, the wearable display system can be configured to enable emitter 2130a and / or emitter 2130b. The wearable display system can be configured to disable emitter 2130a and / or emitter 2130b when additional depth information is not needed, thereby reducing power consumption and improving battery life.

[0247] Furthermore, even if the headset is equipped with an image sensor configured to detect IR light, it is not necessary to mount an IR emitter on or solely on the headset 2100. In some embodiments, the IR emitter can be an external device installed in the space, such as an interior room, where the headset 2100 is used. Such an emitter can emit IR light, for example, in a 940nm ArUco pattern, which is invisible to the human eye. Light with such a pattern can enable "instrumented / assisted tracking," in which the headset 2100 does not need to be powered to provide the IR pattern, but can still provide IR image information as a result of the pattern, allowing the distance or position of objects within the space to be determined by processing this image information. Systems with external illumination sources can also enable more devices to operate within the space. If multiple headsets are operated in the same space, each headset moves within the space without a fixed positional relationship, there is a risk that light emitted by one headset could project onto the image sensor of another headset, thereby interfering with its operation. For example, the risk of such interference between headsets may limit the number of headsets that can operate in a space to 3 or 4. By utilizing one or more IR emitters in the space that illuminate objects that can be imaged by image sensors on the headsets, more headsets (in some embodiments, more than 10) can operate in the same space without interference.

[0248] As mentioned above about Figure 3B As disclosed, camera 2140 can be configured to capture images of the physical world within field of view 2141. Camera 2140 can include an image sensor and a lens. The image sensor can be configured to generate a color image. The image sensor can be configured to acquire images having a size between 4 megapixels and 16 megapixels. For example, the image sensor can acquire an image of 12 megapixels. The image sensor can be configured to repeatedly or periodically acquire images. For example, the image sensor can be configured to acquire images at a frequency between 30 Hz and 120 Hz (e.g., 60 Hz).

[0249] The image sensor may be a CMOS image sensor. The image sensor may be configured with a rolling shutter. Figure 14 and Figure 15 As discussed, the rolling shutter may iteratively read subsets of pixels in the image sensor such that pixels in different subsets reflect light intensity data collected at different times. For example, the image sensor may be configured to read a first row of pixels in the image sensor at a first time and a second row of pixels in the image sensor at a later time. In some embodiments, the camera 2140 may be configured as a plenoptic camera. For example, as described above with respect to Figure 3B and Figures 15 to 20C As discussed, components can be placed in the optical path to one or more pixel cells of an image sensor so that these pixel cells produce an output having an intensity that indicates the angle of arrival of light incident on the pixel cells. In such embodiments, the image sensor can passively acquire depth information. An example of a component suitable for placement in the optical path is a transmissive diffraction mask (TDM) filter. The processor can be configured to use this angle of arrival information to calculate the distance to the imaged object. For example, the angle of arrival information can be converted into distance information indicating the distance to the object from which the light was reflected. In some embodiments, the pixel cells configured to provide angle of arrival information can be interspersed with pixel cells that capture the intensity of light of one or more colors. As a result, the angle of arrival information, and therefore the distance information, can be combined with other image information about the object. In some embodiments, the camera 2140 can be configured to provide a tiled tracking function. The processor of the head-mounted device 2100 can be configured to provide instructions to the camera 2140 to limit image capture to a subset of pixels. In some embodiments, the sensor can be a CMOS sensor. In some embodiments, the image sensor can be an IMX380 image sensor or an equivalent image sensor.

[0250] The camera 2140 may be positioned on a side of the head mounted device 2100 opposite the camera 2120. For example, Figure 21 As shown, when camera 2140 and monocular 2110a are located on the same side of the head-mounted device 2100, camera 2120 can be located on the same side of the head-mounted device as monocular 2110b. Camera 2140 can be angled inward on the head-mounted device 2100. For example, a vertical plane passing through the center of field of view 2141, i.e., the field of view associated with camera 2140, can intersect and form an angle with a vertical plane passing through the centerline of the head-mounted device 2100. This angle can be between 1 and 20 degrees. Field of view 2141 of camera 2140 can have a horizontal field of view and a vertical field of view. The horizontal field of view can range between 75 and 125 degrees, while the vertical field of view can range between 60 and 125 degrees.

[0251] Camera 2140 and camera 2120 can be angled asymmetrically inward toward the midline of headset 2100. The angle of camera 2140 can be between 1 and 20 degrees inward toward the midline of headset 2100. The angle of camera 2120 can be between 20 and 40 degrees inward toward the midline of headset 2100 and can be different from the angle of camera 2140. The angular range of field of view 2120 can exceed the angular range of field of view 2141.

[0252] Camera 2120 and camera 2140 can be configured to provide overlapping views of a central field of view 2150. The angular range of central field of view 2150 can be between 40 degrees and 120 degrees. For example, the angular range of central field of view 2150 can be approximately 60 degrees (e.g., 60±6 degrees). Central field of view 2150 can be asymmetric. For example, central field of view 2150 can extend further toward a side of head mounted device 2100 including camera 2140, such as Figure 21 As shown. In addition to the central field of view 2150, the camera 2120 and the camera 2140 can be positioned to provide at least two peripheral fields of view. The peripheral field of view 2160a can be associated with the camera 2120 and can include a portion of the field of view 2121 that does not overlap with the field of view 2141. In some embodiments, the horizontal angular range of the peripheral field of view 2160a can be within a range between 20 degrees and 80 degrees. For example, the angular range of the peripheral field of view 2160a can be approximately 40 degrees (e.g., 40±4 degrees). The peripheral field of view 2160b ( Figure 21 2121 ). In some embodiments, the peripheral field of view 2160b may have a horizontal angular range of about 10 to 40 degrees. For example, the angular range of the peripheral field of view 2160a may be about 20 degrees (e.g., 20±2 degrees). The location of the peripheral field of view may vary. For example, for a particular configuration of the head-mounted device 2100, the peripheral field of view 2160b may not extend within 0.25 meters of the head-mounted device 2100 because the field of view 2141 may fall entirely within the field of view 2120 at that distance. Conversely, the peripheral field of view 2160a may extend within 0.25 meters of the head-mounted device 2100. In such a configuration, the wider field of view and greater inward angle of the camera 2120 may ensure that, even within 0.25 meters of the head-mounted device 2100, the field of view 2120 at least partially falls outside the field of view 2140.

[0253] IMU 2170a and / or IMU 2170b can be configured to provide acceleration and / or velocity and / or inclination information to the wearable display system. For example, when a user wearing head mounted device 2100 moves, IMU 2170a and / or IMU 2170b can provide information describing the acceleration and / or velocity of the user's head.

[0254] The wearable display system can be coupled to a processor that can be configured to process image information acquired using the camera in order to extract information from the images captured using the camera described herein and / or render virtual objects on a display device. The processor can be mechanically coupled to the frame 2101. Alternatively, the processor can be mechanically coupled to a display device, such as a display device including a monocular 2110a or a monocular 2110b. As a further alternative, the processor can be operably coupled to the head-mounted device 2100 and / or the display device via a communication link. For example, the XR system can include a local data processing module. The local data processing module can include a processor and can be connected to the head-mounted device 2100 and / or the display device via a physical connection (e.g., a wire or cable) or a wireless (e.g., Bluetooth, Wi-Fi, Zigbee, etc.) connection.

[0255] The processor may be configured to perform world reconstruction, head pose tracking, and object tracking operations. For example, the processor may be configured to create a traversable world model using camera 2120 and camera 2140. When creating the traversable world model, the processor may be configured to stereoscopically determine depth information using multiple images of the same physical object acquired by camera 2120 and camera 2140. As an additional example, the processor may be configured to update an existing traversable world model using camera 2120 instead of camera 2140. As described above, camera 2120 may be a grayscale camera with a lower resolution than color camera 2140. Therefore, using images acquired by camera 2120 instead of camera 2140 to quickly update the traversable world model can reduce power consumption and extend battery life. In some embodiments, the processor may be configured to occasionally or periodically update the traversable world model using camera 2120 and camera 2140. For example, the processor may be configured to determine that the traversable world quality criteria are no longer met, and / or a predetermined time interval has elapsed since the last acquisition and / or use of images acquired by camera 2140, and / or that objects in a portion of the physical world currently in the field of view of cameras 2120 and 2140 have changed.

[0256] The processor may be configured to acquire light field information, such as angle of arrival information of light incident on an image sensor, using a plenoptic camera. In some embodiments, the plenoptic camera may be at least one of camera 2120 and camera 2140. Consistent with the disclosed embodiments, when depth information is described herein or can be enhanced for processing, such depth information may be determined from or supplemented by light field information acquired by the plenoptic camera. For example, when camera 2140 includes a TDM filter, the processor may be configured to create a traversable world model using images acquired from cameras 2120 and 2140 and light field information acquired from camera 2140. Alternatively or additionally, when camera 2120 includes a TDM filter, the processor may use camera light field information acquired from camera 2120. After creating the traversable world model, the processor may be configured to track head pose using the at least one plenoptic camera and the traversable world model.

[0257] The processor can be configured to perform a size reduction routine to adjust the image acquired using camera 2140. Camera 2140 can generate larger images than camera 2120. For example, camera 2140 can generate an image with 12 million pixels, while camera 2120 can generate an image with 1 million pixels. The image generated by camera 2140 may include more information than required for performing a traversable world model creation, head tracking, or object tracking operation. Processing this additional information may require additional power, thereby shortening battery life or increasing latency. Therefore, the processor can be configured to discard or combine pixels in the image acquired by camera 2140. For example, the processor can be configured to output an image with one-sixteenth the number of pixels as the acquired image. Each pixel in the output image can have a value based on the corresponding 4x4 pixel group in the acquired image (e.g., the average value of the values of these 16 pixels).

[0258] The processor may be configured to perform object tracking in the central field of view 2150 using images from camera 2120 and camera 2140. In some embodiments, the processor may perform object tracking using depth information determined stereoscopically from the first images acquired by the two cameras. As a non-limiting example, the tracked object may be a hand of a user of the wearable display system. The processor may be configured to perform object tracking in the peripheral field of view using images acquired from one of cameras 2120 and camera 2140.

[0259] According to some embodiments, the XR system may include a hardware accelerator. The hardware accelerator may be implemented as an application-specific integrated circuit (ASIC) or other semiconductor device and may be integrated within the head-mounted device 2100 or otherwise coupled to the head-mounted device 2100 so that it receives image information from the camera 2120 and the camera 2140. The hardware accelerator may use the images acquired by the two world cameras to assist in stereoscopically determining depth information. The image from the camera 2120 may be a grayscale image and the image from the camera 2140 may be a color image. Using hardware acceleration may speed up the determination of depth information and reduce power consumption, thereby extending battery life.

[0260] Example Calibration Process

[0261] Figure 22 A simplified flowchart of a calibration routine (method 2200) according to some embodiments is shown. The processor can be configured to perform the calibration routine while the wearable display system is being worn. The calibration routine can account for deformation caused by the lightweight structure of the head-mounted device 2100. In some embodiments, the calibration routine can account for deformation of the frame 2101 due to temperature changes or mechanical strain on the frame 2101 during use. For example, the processor can repeatedly perform the calibration routine so that the calibration routine compensates for deformation of the frame 2101 during use of the wearable display system. The compensation routine can be performed automatically or in response to manual input (e.g., a user request to perform the calibration routine). The calibration routine can include determining the relative position and orientation of camera 2120 and camera 2140. The processor can be configured to use images acquired by camera 2120 and camera 2140 to perform the calibration routine. In some embodiments, the processor can also be configured to use the output of IMU 2170a and IMU 2170b.

[0262] After starting in block 2201, method 2200 may proceed to block 2210. In block 2210, the processor may identify corresponding features in images acquired from cameras 2120 and 2140. The corresponding features may be parts of objects in the physical world. In some embodiments, for calibration purposes, an object may be placed by a user within central field of view 2150 and may have features that are easily identifiable in the image, which may have predetermined relative positions. However, calibration techniques as described herein may be performed based on features on an object that appear in central field of view 2150 during calibration, enabling repeatable calibration during use of the head-mounted device 2100. In various embodiments, the processor may be configured to automatically select features detected within fields of view 2121 and 2141. In some embodiments, the processor may be configured to determine correspondences between features using the estimated positions of the features within fields of view 2121 and 2141. This estimation may be based on a traversable world model constructed for the object that contains these features or other information about these features.

[0263] Method 2200 may proceed to block 2230. In block 2230, the processor may receive inertial measurement data. The inertial measurement data may be received from IMU 2170a and / or IMU 2170b. The inertial measurement data may include tilt and / or acceleration and / or velocity measurements. In some embodiments, IMUs 2170a and 2170b may be mechanically coupled directly or indirectly to camera 2140 and camera 2120, respectively. In such embodiments, a difference in inertial measurements (e.g., tilt) taken by IMUs 2170a and 2170b may indicate a difference in the position and / or orientation of camera 2140 and camera 2120. Thus, the outputs of IMUs 2170a and 2170b may provide a basis for an initial estimate of the relative position of camera 2140 and camera 2120.

[0264] After block 2230, method 2200 may proceed to block 2250. In block 2250, the processor may calculate an initial estimated relative position and orientation of cameras 2120 and 2140. This initial estimate may be calculated using measurements received from IMU 2170b and / or IMU 2170a. In some embodiments, for example, the head-mounted device may be designed with nominal relative positions and orientations of cameras 2120 and 2140. The processor may be configured to attribute differences in measurements received between IMU 2170a and IMU 2170b to deformation of frame 2101, which may change the position and / or orientation of cameras 2120 and 2140. For example, IMU 2170a and IMU 2170b may be mechanically coupled directly or indirectly to frame 2101 such that the tilt and / or acceleration and / or velocity measurements of these sensors have a predetermined relationship. This relationship may be affected when frame 2101 deforms. As a non-limiting example, IMU 2170a and IMU 2170b can be mechanically coupled to frame 2101 so that these sensors measure similar tilt, acceleration, or velocity vectors during head-mounted device motion when there is no deformation of frame 2101. In this non-limiting example, a twist or bend that rotates IMU 2170a relative to IMU 2170b can result in a corresponding rotation of the tilt, acceleration, or velocity vector measurement of IMU 2170a relative to the corresponding vector measurement of IMU 2170b. Because IMU 2170a and IMU 2170b are mechanically coupled to camera 2140 and camera 2120, respectively, the processor can adjust the nominal relative position and orientation of camera 2120 and camera 2140 consistent with the measurement relationship between IMU 2170a and IMU 2170b.

[0265] Other techniques may alternatively or additionally be used to make the initial estimate.In embodiments where calibration method 2200 is performed repeatedly during operation of the XR system, the initial estimate may be, for example, the most recently calculated estimate.

[0266] After block 2250, a sub-process is initiated in which an estimate of the relative position and orientation of camera 2120 and camera 2140 is also made. One of the estimates is selected as the relative position and orientation of camera 2120 and camera 2140 to calculate stereo depth information based on the images acquired by camera 2120 and camera 2140. This sub-process may be performed iteratively, making additional estimates in each iteration until an acceptable estimate is determined. Figure 22 In the example of , the subprocess includes boxes 2270, 2272, 2274 and 2290.

[0267] In block 2270, the processor may calculate an error between the estimated relative orientations of the cameras and the compared features. In calculating the error, the processor may be configured to estimate how the identified features should appear or where the identified features should be located within corresponding images acquired using cameras 2120 and 2140 based on the estimated relative orientations of cameras 2120 and 2140 and the estimated position features used for calibration. In some embodiments, this estimate may be compared to the appearance or apparent position of the corresponding features in the images acquired using each of the two cameras to generate an error for each of the estimated relative orientations. Linear algebra techniques may be used to calculate such errors. For example, the mean squared deviation between the calculated position and the actual position of each of the plurality of features within the image may be used as a measure of error.

[0268] After block 2270, method 2200 may proceed to block 2272, where it may be checked whether the error meets an acceptance criterion. For example, the criterion may be the overall magnitude of the error, or the change in error between iterations. If the error meets the acceptance criterion, method 2200 may proceed to block 2290.

[0269] In block 2290, the processor may select one of the estimated relative orientations based on the error calculated in block 2272. The selected estimated relative position and orientation may be the estimated relative position and orientation with the lowest error. In some embodiments, the processor may be configured to select the estimated relative position and orientation associated with the lowest error as the current relative position and orientation of camera 2120 and camera 2140. After block 2290, method 2200 may proceed to block 2299. Method 2200 may end in block 2299, where the selected positions and orientations of camera 2120 and camera 2140 are used to calculate stereoscopic image information based on images formed by those cameras.

[0270] If the error does not meet the acceptance criteria at block 2272, method 2200 may proceed to block 2274. At block 2274, the estimates used to calculate the error at block 2270 may be updated. These updates may be to the estimated relative positions and / or orientations of cameras 2120 and 2140. In embodiments where the relative positions of a set of features used for calibration are estimated, the updated estimates selected at block 2274 may alternatively or additionally include updates to the positions or orientations of the features in the set. Such updates may be performed based on linear algebraic techniques for solving systems of equations with multiple variables. As a specific example, one or more of the estimated positions or orientations may be increased or decreased. If the change reduces the calculated error in one iteration of the sub-process, the same estimated position or orientation may be further changed in the same direction in subsequent iterations. Conversely, if the change increases the error, the estimated positions or orientations may be changed in the opposite direction in subsequent iterations. The estimated positions and orientations of the cameras and features used in the calibration process may be changed sequentially or in combination in this manner.

[0271] Once the updated estimate is calculated, the sub-process returns to block 2270. Here, further iterations of the sub-process begin, calculating the error in the estimated relative position. In this manner, the estimated position and orientation are updated until an updated relative position and orientation that provides an acceptable error is selected. However, it should be understood that the processing at block 2272 can apply other criteria to terminate the iterative sub-process, such as completing multiple iterations without finding an acceptable error.

[0272] Although method 2200 is described in conjunction with camera 2120 and camera 2140, similar calibration may be performed for any pair of cameras used for stereoscopic imaging or for any camera group for which relative position and orientation are desired.

[0273] Example Camera Configurations

[0274] According to some embodiments, components are incorporated into the head-mounted device 2100 to provide a field of view and a field of illumination to support various functions of the XR system. Figures 23A to 23C According to some embodiments Figure 21 21. Example diagrams of fields of view or illumination associated with a head-mounted device 2100. Each example diagram shows the field of view or illumination from a different orientation and distance relative to the head-mounted device. Figure 23A The field of view or illumination is shown from elevated off-axis angles at a distance of 1 meter from the head-mounted device. Figure 23AThe overlap between the fields of view of camera 2120 and camera 2140 is shown, specifically how camera 2120 and camera 2140 are angled so that field of view 2121 and field of view 2141 pass through the centerline of head-mounted device 2100. In the illustrated configuration, field of view 2141 extends beyond field of view 2121 to form peripheral field of view 2160b. As shown, the illumination fields of emitter 2130a and emitter 2130b overlap to a large extent. In this way, emitter 2130a and emitter 2130b can be configured to support imaging or depth measurement of objects in central field of view 2150 under low ambient light conditions. Figure 23B The field of view or illumination is shown from a top-down perspective at a distance of 0.3 meters from the head-mounted device. Figure 23B The overlap of field of view 2121 and field of view 2141 is shown at a distance of 0.3 meters from the head mounted device. However, in the illustrated configuration, field of view 2141 does not extend very far beyond field of view 2121, limiting the extent of peripheral field of view 2160b and exhibiting an asymmetry between peripheral field of view 2160a and peripheral field of view 2160b. Figure 23C The field of view or illumination is shown from a front-facing perspective at a distance of 0.25 meters from the head-mounted device. Figure 23C The overlap of field of view 2121 and field of view 2141 is shown as existing at 0.25 meters from the head mounted device. However, field of view 2141 is completely contained within field of view 2121, so the peripheral field of view 2160b does not exist at this distance from the head mounted device 2100 in the illustrated configuration.

[0275] from Figures 23A to 23C As will be appreciated, the overlap of fields of view 2121 and 2141 creates a central field of view in which stereoscopic imaging techniques can be employed using images acquired by cameras 2120 and 2140, with or without IR illumination from emitters 2130a and 2130b. Within this central field of view, color information from camera 2140 can be combined with grayscale image information from camera 2120. Additionally, there are non-overlapping peripheral fields of view, but monocular grayscale image information or color image information is available from camera 2120 or camera 2140, respectively. As described herein, different operations can be performed on the image information acquired for the central and peripheral fields of view.

[0276] World model generation

[0277] In some embodiments, image information acquired in the central field of view may be used to construct or update a world model. Figure 24 is a simplified flow chart of a method 2400 for creating or updating a traversable world model according to some embodiments. Figure 21As disclosed, the wearable display system can be configured to use a processor to determine and update a traversable world model. In some embodiments, the processor can determine and update the traversable world model based on the outputs of camera 2120 and camera 2140. However, camera 2120 may be configured with a global shutter, while camera 2140 may be configured with a rolling shutter. The processor can therefore perform a compensation routine to compensate for rolling shutter distortion in images acquired by camera 2140. In various embodiments, the processor can determine and update the traversable world model without using emitters 2130a and 2130b. However, in some embodiments, the traversable world model may be incomplete. For example, the processor may incompletely determine the depth of a wall or other flat surface. As an additional example, the traversable world model may incompletely represent objects with many corners, curved surfaces, transparent surfaces, or large surfaces, such as windows, doors, balls, tables, etc. The processor can be configured to identify such incomplete information, obtain additional information, and update the world model using the additional depth information. In some embodiments, transmitters 2130a and 2130b can be selectively enabled to collect additional image information from which to construct or update the traversable world model. In some scenarios, the processor can be configured to perform object recognition in the acquired imagery, select a template for the recognized object, and add information to the traversable world model based on the template. In this way, the wearable display system can improve the traversable world model while using little or no power-intensive components such as transmitters 2130a and 2130b, thereby extending battery life.

[0278] Method 2400 can be initiated one or more times during operation of the wearable display system. The processor can be configured to create a navigable world model when the user first turns on the system, moves to a new environment (e.g., walks into another room), or generally when the processor detects a change in the user's physical environment. Alternatively or additionally, method 2400 can be executed periodically during operation of the wearable display system, or upon detecting a significant change in the physical world, or in response to user input (e.g., input indicating that the world model is out of sync with the physical world).

[0279] In some embodiments, all or part of a traversable world model may be stored, provided by other users of the XR system or otherwise obtained. Thus, while the creation of a world model is described, it should be understood that method 2400 may be used for a portion of a world model, while other portions of the world model are derived from other sources.

[0280] In block 2405, the processor may execute a compensation routine to compensate for rolling shutter distortion in images acquired by camera 2140. As described above, images acquired by an image sensor with a global shutter (e.g., the image sensor in camera 2120) include pixel values acquired simultaneously. In contrast, images acquired by an image sensor with a rolling shutter include pixel values acquired at different times. During image acquisition by camera 2140, relative motion of the head-mounted device and the environment may introduce spatial distortions into the images. These spatial distortions may affect the accuracy of methods that rely on comparing images acquired by camera 2140 with images acquired by camera 2120.

[0281] Performing the compensation routine may include using a processor to compare an image acquired using camera 2120 with an image acquired using camera 2140. The processor performs this comparison to identify any distortion in the image acquired using camera 2140. This distortion may include skew in at least a portion of the image. For example, if the image sensor in camera 2140 is acquiring pixel values row by row from the top of the image sensor to the bottom of the image sensor, and the headset 2100 is being translated sideways, the position where an object or portion of the object appears may be offset in consecutive pixel rows by an amount that depends on the speed of the translation and the time difference between acquiring each row of values. Similar distortions may occur when rotating the headset. These distortions may result in an overall skew in the position and / or orientation of the object or portion of the object in the image. The processor may be configured to perform a row-by-row comparison between the image acquired by camera 2120 and the image acquired by camera 2140 to determine the amount of skew. The image acquired by camera 2140 may then be transformed to remove the distortion (e.g., to remove the detected skew).

[0282] In block 2410, a traversable world model may be created. In the illustrated embodiment, the processor may use camera 2120 and camera 2140 to create the traversable world model. As described above, when generating the traversable world model, the processor may be configured to use images acquired from camera 2120 and camera 2140 to stereoscopically determine depth information of objects in the physical world when constructing the traversable world model. In some embodiments, the processor may receive color information from camera 2140. This color information may be used to distinguish objects or identify surfaces associated with the same object. Color information may also be used to identify objects. As described above with respect to Figure 3AAs disclosed, the processor can create a navigable world model by associating information about the physical world with information about the position and orientation of the head-mounted device 2100. As a non-limiting example, the processor can be configured to determine the distance from the head-mounted device 2100 to a feature in a view (e.g., field of view 2120 and / or field of view 2140). The processor can be configured to estimate the current position and orientation of the view. The processor can be configured to accumulate such distance and position and orientation information. By triangulating the distance to the feature obtained from multiple positions and orientations, the position and orientation of the feature in the environment can be determined. In various embodiments, the navigable world model can be a combination of raster images, point and descriptor clouds, and polygon / geometric definitions that describe the position and orientation of such features in the environment. In some embodiments, the distance from the head-mounted device 2100 to features in the central field of view 2150 can be determined stereoscopically using images obtained from camera 2120 and compensated images obtained from camera 2140. In various embodiments, light field information can be used to supplement or refine this determination. For example, through calculation, the angle of arrival information can be converted into distance information indicating the distance to the object from which the light was reflected.

[0283] In block 2415, the processor may disable camera 2140 or reduce the frame rate of camera 2140. For example, the frame rate of camera 2140 may be reduced from 30 Hz to 1 Hz. As disclosed above, color camera 2140 may consume more power than grayscale camera 2120. By disabling or reducing the frame rate of camera 2140, the processor may reduce power consumption and extend the battery life of the wearable display system. Thus, the processor may disable camera 2140 or reduce the frame rate of camera 2140 to save power. This lower power state may be maintained until a condition is detected indicating that an update in the world model may be needed. Such a condition may be detected based on the passage of time or input, such as from a sensor collecting information about the user's surroundings or from the user.

[0284] Alternatively or additionally, once the traversable world model is sufficiently complete, it may be sufficient to use images from camera 2120 to determine the location and orientation of features in the physical environment. As non-limiting examples, the traversable world model may be identified as sufficiently complete based on the percentage of space surrounding the user's location represented in the model or based on the amount of new image information that matches the traversable world model. For the latter approach, newly acquired images may be associated with locations in the traversable world. If features in these images have features that match features identified as landmarks in the traversable world model, the world model may be considered complete. Coverage or matching need not be 100% complete. Instead, appropriate thresholds may be applied for each criterion, such as greater than 95% coverage or greater than 90% feature matching previously identified landmarks. Regardless of how the traversable world model is determined to be complete, once complete, the processor may use the existing traversable world information to refine its estimate of the location and orientation of features in the physical world. This process may reflect the assumption that features in the physical world are changing position and / or orientation slowly (if at all) compared to the rate at which the processor processes images acquired by camera 2120.

[0285] In block 2420, after creating the traversable world model, the processor may identify surfaces and / or objects for updating the traversable world model. In some embodiments, the processor may use a grayscale image acquired from camera 2120 to identify such surfaces or objects. For example, once a world model has been created at block 2410 indicating a surface at a particular location within the traversable world, the grayscale image acquired from camera 2120 may be used to detect surfaces having substantially the same features and determine that the traversable world model should be updated by updating the position of that surface within the traversable world model. For example, a surface having substantially the same shape as a surface in the traversable world model at substantially the same location may be equated to that surface in the traversable world model, and the traversable world model may be updated accordingly. As another example, the position of an object represented in the traversable world model may be updated based on the grayscale image acquired from camera 2120.

[0286] In some embodiments, the processor can use light field information obtained from the camera 2120 to determine depth information for objects in the physical world. The depth information can be determined based on angle-of-arrival information, which may be more accurate for objects in the physical world that are closer to the headset 2100. Therefore, in some embodiments, the processor can be configured to use the light field information to update only the portion of the traversable world model that meets a depth criterion. The depth criterion can be based on a maximum distinguishable distance. For example, the processor may not be able to distinguish objects at different distances from the headset 2100 when those distances exceed a threshold distance. The depth criterion can be based on a maximum error threshold. For example, the error in the estimated distance may increase with distance, with a particular distance corresponding to a maximum error threshold. In some embodiments, the depth criterion can be based on a minimum distance. For example, the processor may not be able to accurately determine distance information for objects from the headset 2100 within the minimum distance. Therefore, portions of the world model that are 15 cm or more from the headset may meet the depth criterion. In some embodiments, the traversable world model can be composed of three-dimensional voxel "bricks." In such embodiments, updating the traversable world model can include identifying voxel blocks for updating. In some embodiments, the processor may be configured to determine a viewing frustum. The viewing frustum may have a maximum depth, such as 1.5 meters. The processor may be configured to identify blocks within the viewing frustum. The processor may then update the traversable world information for the voxels within the identified blocks. In some embodiments, the processor may be configured to update the traversable world information for the voxels using the light field information acquired in step 2450, as described herein.

[0287] In some embodiments, the update process can be performed differently based on whether the object is in the central field of view or the peripheral field of view. For example, updates can be performed for surfaces detected in the central field of view. In the peripheral field of view, for example, updates can be performed only for objects for which the processor has a model, so that the processor can confirm that any updates to the traversable world model are consistent with that object. Alternatively or additionally, new objects or surfaces can be identified based on processing of grayscale images. Even if such processing results in a less accurate representation of the object or surface than the processing at block 2410, in some scenarios, balancing accuracy with faster and less power-intensive processing can result in a better overall system. Furthermore, by periodically repeating method 2400, less accurate information can be periodically replaced with more accurate information, so that portions of the world model generated using only monocular grayscale images are replaced with portions generated stereoscopically using color camera 2140 in combination with grayscale camera 2120.

[0288] In some embodiments, the processor can be configured to determine whether the updated world model meets a quality standard. When the world model meets the quality standard, the processor can continue to update the world model with camera 2140 disabled or at a reduced frame rate. When the updated world model does not meet the quality standard, method 2400 can enable camera 2140 or increase the frame rate of camera 2140. Method 2400 can also return to step 2410 and recreate a traversable world model.

[0289] In block 2425, after updating the traversable world model, the processor may identify whether the traversable world model includes incomplete depth information. Incomplete depth information may occur in any of a number of ways. For example, some objects will not produce detectable structure in the image. For example, very dark areas in the physical world may not be imaged with sufficient resolution to extract depth information from an image acquired using ambient lighting. As another example, a window or glass tabletop may not appear in a visible image or be recognized by computer processing. As yet another example, a large uniform surface, such as a tabletop or a wall, may lack sufficient features that can be correlated in two stereo images to enable stereo image processing. As a result, the processor may not be able to determine the position of such an object using stereo processing. In these scenarios, there will be a "hole" in the world model because a process that seeks to use the traversable world model to determine the distance to the surface in a particular direction through the "hole" will not be able to obtain any depth information.

[0290] When the traversable world model does not include incomplete depth information, method 2400 may return to updating the traversable world model using the grayscale image obtained from camera 2120 .

[0291] After identifying incomplete depth information, the processor controlling method 2400 may take one or more actions to obtain additional depth information. Method 2400 may proceed to blocks 2431, 2433, and / or 2435. In block 2431, the processor may enable emitter 2130a and / or emitter 2130b. As described above, one or more of camera 2120 and camera 2140 may be configured to detect light emitted by emitter 2130a and / or emitter 2130b. The processor may then obtain depth information by causing emitter 2130a and / or 2130b to emit light that may enhance images of objects captured in the physical world. For example, when cameras 2120 and 2140 are sensitive to the emitted light, images captured using cameras 2120 and 2140 may be processed to extract stereo information. When emitter 2130a and / or emitter 2130b are enabled, other analysis techniques may alternatively or additionally be used to obtain depth information. In some embodiments, time-of-flight measurements and / or structured light techniques may be used alternatively or additionally.

[0292] In block 2433, the processor may determine additional depth information from previously acquired depth information. In some embodiments, for example, the processor may be configured to identify objects in an image formed using camera 2120 and / or camera 2140 and fill any holes in the traversable world model based on a model of the identified object. For example, the process may detect flat surfaces in the physical world. The flat surface may be detected using existing depth information acquired using camera 2120 and / or camera 2140 or depth information stored in the traversable world model. The flat surface may be detected in response to determining that a portion of the world model includes incomplete depth information. The processor may be configured to estimate additional depth information based on the detected flat surface. For example, the processor may be configured to extend the identified flat surface through an area of incomplete depth information. In some embodiments, the processor may be configured to interpolate the missing depth information based on the surrounding portion of the traversable world model when extending the flat surface.

[0293] In some embodiments, as an additional example, the processor may be configured to detect an object in a portion of the world model that includes incomplete depth information. In some embodiments, the detection may involve using a neural network or other machine learning tools to identify the object. In some embodiments, the processor may be configured to access a database storing templates and select an object template corresponding to the identified object. For example, when the identified object is a window, the processor may be configured to access a database storing templates and select a corresponding window template. As a non-limiting example, the template may be a three-dimensional model representing a class of objects, such as a window, door, ball, etc. The processor may configure an instance of the object template based on an image of the object in the updated world model. For example, the processor may scale, rotate, and translate the template to match the detected position of the object in the updated world model. Additional depth information may then be estimated based on the boundaries of the configured template, representing the surface of the identified object.

[0294] In block 2435, the processor may acquire light field information. In some embodiments, the light field information may be acquired along with the image and may include angle of arrival information. In some embodiments, the camera 2120 may be configured as a plenoptic camera to acquire the light field information.

[0295] After blocks 2431, 2433, and / or 2435, method 2400 may proceed to block 2440. In block 2490, the processor may update the traversable world model using the additional depth information obtained in blocks 2431 and / or 2473. For example, the processor may be configured to blend additional depth information obtained from measurements made using active IR illumination into the existing traversable world model. Similarly, additional depth information determined from light field information, such as using triangulation based on angle of arrival information, may be blended into the existing traversable world model. As additional examples, the processor may be configured to blend interpolated depth information obtained by extending detected flat surfaces into the existing traversable world model, or to blend additional depth information estimated based on the boundaries of a configuration template into the existing traversable world model.

[0296] The information can be blended in one or more ways, depending on the nature of the additional depth information and / or the information in the traversable world model. For example, blending can be performed by adding to the traversable world model additional depth information collected for locations in the traversable world model where holes exist. Alternatively, the additional depth information can overwrite information at corresponding locations in the traversable world model. As yet another alternative, blending can involve selecting between information already in the traversable world model and the additional depth information. Such a selection can, for example, be based on selecting depth information already in the traversable world model or depth information in the additional depth information that represents the surface closest to the camera used to collect the additional depth information.

[0297] In some embodiments, a traversable world model can be represented by a mesh of connected points. Updating the world model can be accomplished by computing a mesh representation of the object or surface to be added to the world model, and then combining that mesh representation with the mesh representation of the world model. The inventors have recognized and appreciated that performing the processes in this order may require less processing than adding the object or surface to the world model and then computing the mesh for the updated model.

[0298] Figure 24 It is shown that the world model can be updated at blocks 2420 and 2440. The processing of each block can be performed in the same manner, such as by generating a mesh representation of the object or surface to be added to the world model and combining the generated mesh with the mesh of the world model, or in a different manner. In some embodiments, this merging operation can be performed once for both the objects or surfaces identified at blocks 2420 and 2440. For example, this combining process can be performed as described in conjunction with block 2440.

[0299] In some embodiments, method 2400 may loop back to block 2420 to repeat the process of updating the world model based on information acquired with one or more grayscale cameras. Because the process of block 2420 may be performed on fewer and smaller images than the process of block 2410, it may be repeated at a higher rate. The process may be performed at a rate of less than 10 times per second, for example, between 3 and 7 times per second.

[0300] Method 2400 may repeat in this manner until an end condition is detected. For example, method 2400 may repeat for a predetermined time period, until user input is received, or until a change of a particular type or magnitude in a portion of the physical world model is detected in the field of view of a camera of the head mounted device 2100. Method 2400 may then terminate at block 2499. Method 2400 may be initiated again to capture new information for the world model at block 2405, including information acquired with a higher resolution color camera. Method 2400 may terminate and restart to repeat the processing at block 2405 using the color camera, thereby creating a portion of the world model at an average rate that is slower than the rate at which the world model is updated based solely on grayscale image information. For example, the processing using the color camera may be repeated at an average rate of once per second or slower.

[0301] Head pose tracking

[0302] The XR system can track the position and orientation of the head of a user wearing the XR display system. Determining the user's head pose enables information in a traversable world model to be converted to the reference frame of the user's wearable display device so that the information in the traversable world model can be used to render objects on the wearable display device. Because head pose is frequently updated, performing head pose tracking using only camera 2120 can provide energy savings, reduced computation, or other benefits. The XR system can therefore be configured to disable or reduce the frame rate of color camera 2140 as needed to balance head tracking accuracy with power consumption and computational requirements.

[0303] Figure 25 2 is a simplified flow chart of a method 2500 for head pose tracking according to some embodiments. The method 2500 may include creating a world model, tracking a head pose, and determining whether head pose tracking criteria are met. When the head pose tracking criteria are not met, the method 2500 may further include enabling the camera 2140 and tracking the head pose using stereoscopically determined depth information.

[0304] In block 2510, the processor may create a traversable world model. In some embodiments, the processor may be configured to create the traversable world model as described above with respect to blocks 2405-2415 of method 2400. For example, the processor may be configured to acquire images from camera 2120 and camera 2140. In some embodiments, the processor may compensate for rolling shutter distortion in camera 2140. The processor may then use the images from 2120 and the compensated images from 2140 to determine the depths of features in the physical world. Using these depths, the processor may create the traversable world model. In some embodiments, after creating the traversable world model, the processor may be configured to disable or reduce the frame rate of camera 2140. By disabling or reducing the frame rate of camera 2140 after generating the traversable world model, the XR system may reduce power consumption and computational requirements.

[0305] After creating a traversable world model in block 2510, method 2500 may proceed to block 2520. In block 2520, the processor may track head pose. In some embodiments, the processor may be configured to calculate the user's head pose in real time or near real time based on information acquired by camera 2120. The information may be a grayscale image acquired by camera 2120. Additionally or alternatively, the information may be light field information, such as angle of arrival information.

[0306] In block 2530, method 2500 may determine whether the head pose meets tracking quality criteria. Tracking criteria may depend on the stability of the estimated head pose, the noise in the estimated head pose, the consistency of the estimated head pose with a world model, or similar factors. As a specific example, the calculated head pose may be compared with other information that may indicate inaccuracies, such as the output of an inertial measurement unit or a model of the range of human head motion, to identify errors in the head pose. The specific tracking criteria may vary depending on the tracking method used. For example, in methods using event-based information, correspondence between the positions of features may be used, as indicated by comparing the event-based output with the positions of corresponding features in the full-frame image. Alternatively or additionally, the visual distinctiveness of a feature relative to its surroundings may be used as a tracking criterion. For example, when the field of view is filled with one or more objects, making it difficult to discern the motion of a particular feature, the tracking criteria of an event-based method may indicate poor tracking. The percentage of the field of view that is occluded is an example of a criterion that may be used. For example, a threshold value greater than 40% may be used as an indication to switch from using an image-based method for head pose tracking. As a further example, reprojection error can be used as a measure of head pose tracking quality. Such a criterion can be calculated by matching features in the acquired image to a previously determined traversable world model. The positions of the features in the image can be related to positions in the traversable world model using a geometric transformation calculated based on the head pose. The deviation between the calculated position and the feature in the traversable world model, expressed as a mean squared error, can thus indicate an error in the head pose, allowing the deviation to be used as a tracking criterion.

[0307] The difficulty of tracking head poses can depend on the position and orientation of the user's head and the content of the navigable world model. Thus, in some cases, the processor may be unable or become unable to track head poses using only light field information and / or grayscale images acquired by camera 2120. When the tracking quality criteria are met, method 2500 may return to block 2520 and continue tracking head poses.

[0308] When the tracking quality criteria are not met, method 2500 may proceed to block 2540. In block 2540, the processor may be configured to enable camera 2140. After enabling camera 2140, method 2500 may proceed to block 2550. In block 2550, the processor may stereoscopically determine depth information using images acquired by camera 2140 and camera 2120. In some embodiments, after the processor determines the depth information, method 2500 may return to block 2520 and the processor may resume tracking the head pose. In some embodiments, the processor may be configured to disable camera 2140 when resuming tracking the head pose. In various embodiments, the processor may be configured to continue tracking the head pose using images acquired by camera 2140 and camera 2120 for a duration or a time, or until the quality criteria are met.

[0309] When the attempt to stereoscopically determine depth information is unsuccessful, method 2500 may proceed to block 2599, as shown in FIG. Figure 25 When the tracking quality criteria are not met even after enabling the camera 2140 , the method 2500 may instead proceed to block 2500 .

[0310] Object Tracking

[0311] As described above, the processor of the XR system can track objects in the physical world to support realistic rendering of virtual objects relative to the physical objects. For example, tracking is described in conjunction with movable objects, such as the hands of a user of the XR system. For example, the XR system can track objects in the central field of view 2150, the peripheral field of view 2160a, and / or the peripheral field of view 2160b. Rapidly updating the position of the movable objects enables realistic rendering of virtual objects, because such rendering can reflect occlusion of physical objects by virtual objects, vice versa, or interactions between virtual objects and physical objects. In some embodiments, for example, updates to the position of the physical objects can be calculated at an average rate of at least 10 times per second, and in some embodiments, at least 20 times per second, such as approximately 30 or 60 times per second. When the tracked object is a user's hand, tracking can enable gesture control by the user. For example, a particular gesture can correspond to a command for the XR system.

[0312] In some embodiments, the XR system can be configured to track objects having features that provide high contrast when imaged with an image sensor that is sensitive to IR light. In some embodiments, objects with high contrast features can be created by adding markers to the object. For example, a physical object can be equipped with one or more markers that appear as high contrast areas when imaged with IR light. The markers can be passive markers that are highly reflective or highly absorptive of IR light. In some embodiments, at least 25% of light in the frequency range of interest can be absorbed or reflected. Alternatively or additionally, the markers can be active markers that emit IR light, such as IR LEDs. By tracking these features, for example using a DVS camera, information that accurately represents the location of the physical object can be quickly determined.

[0313] As with head pose tracking, the positions of tracked objects are frequently updated, so performing object tracking using only camera 2120 may provide power savings, reduced computation, or other benefits. The XR system can therefore be configured to disable or reduce the frame rate of color camera 2140 as needed to balance object tracking accuracy against power consumption and computational requirements.

[0314] Figure 26 FIG26 is a simplified flow chart of a method 2600 for object tracking according to some embodiments. According to the method 2600, the processor performs object tracking differently depending on whether the tracked object is in a peripheral field of view (peripheral field of view 2160a or peripheral field of view 2160b) or in the central field of view 2150. The processor may further perform object tracking differently in the central field of view 2150 depending on whether the tracked object meets a depth criterion. The processor may alternatively or additionally apply other criteria to dynamically select an object tracking method, such as available battery power or the operations of the XR system being performed and the need for those operations to track the position of the object or to track the position of the object with high accuracy.

[0315] Method 2600 can begin in block 2601. In some embodiments, camera 2140 can be disabled or have a reduced frame rate. The processor can have disabled camera 2140 or reduced the frame rate of camera 2140 to reduce power consumption and extend battery life. In various embodiments, the processor can track an object in the physical world (e.g., a user's hand). The processor can be configured to predict the next position of an object or the trajectory of an object based on one or more previous positions of the object.

[0316] After starting in block 2601, method 2600 may proceed to block 2610. In block 2610, the processor may be configured to determine whether an object is within central view 2150. In some embodiments, the processor may make this determination based on the object's current location (e.g., whether the object is currently within central field of view 2150). In various embodiments, the processor may make this determination based on an estimate of the object's location. For example, the processor may determine that an object that is leaving central field of view 2150 may enter peripheral field of view 2160a or peripheral field of view 2160b. When the object is not within central view 2150, the processor may be configured to proceed to block 2620.

[0317] In block 2620, the processor may enable the camera 2140 or increase the frame rate of the camera 2140. The processor may be configured to enable the camera 2140 when the processor estimates that the object will soon leave or has left the central field of view 2150 and entered the peripheral field of view 2160b. The processor may be configured to increase the frame rate of the camera 2140 when the processor determines that the object is in the peripheral field of view 2160b based on one or more images received from the camera 2140. The processor may be configured to track the object using the camera 2140 if the object is within the peripheral field of view 2160b.

[0318] The processor may determine whether the object meets the depth criteria in steps 2630a and 2630b. In some embodiments, the depth criteria may be the same or similar to the depth criteria described above with respect to block 2460 of method 2400. For example, the depth criteria may relate to a maximum error rate or a maximum distance beyond which the processor cannot distinguish between different distances between objects.

[0319] In some embodiments, camera 2140 may be configured as a plenoptic camera. In such embodiments, when an object is within peripheral field of view 2160b, the processor may determine whether a distance criterion is met in step 2630a. If the distance criterion is met, the processor may be configured to track the object using light field information in block 2640a. In some embodiments, camera 2120 may be configured as a plenoptic camera. In such embodiments, when an object is within central field of view 2150, the processor may determine whether a distance criterion is met in block 2630b. If the distance criterion is met, the processor may be configured to track the object using light field information in block 2640b. If the distance criterion is not met, method 2600 may proceed to block 2645.

[0320] In block 2645, the processor may be configured to enable camera 2140 or increase the frame rate of camera 2140. After the processor enables camera 2140 or increases the frame rate of camera 2140, method 2600 may proceed to block 2650. In block 2650, the processor may be configured to track the object using stereoscopically determined depth information. The processor may be configured to determine depth information based on images acquired by camera 2120 and camera 2140. After steps 2640a, 2640b, or 2650, method 2600 proceeds to block 2699. In step 2699, method 2600 may end.

[0321] Hand tracking can include further steps beyond general object tracking. Figure 27 2 is a simplified flow chart of a hand tracking method 2700 according to some embodiments. The object tracked in method 2700 can be a user's hand. In various embodiments, the processor can be configured to perform hand tracking using both camera 2120 and camera 2140 when the user's hand is in central field of view 2150, to use only camera 2120 when the user's hand is in peripheral field of view 2160a, and to use only camera 2140 when the user's hand is in peripheral field of view 2160b. If necessary, the XR system can enable camera 2140 or increase the frame rate of camera 2140 to enable tracking in field of view 2160b. In this way, the wearable display system can be configured to provide adequate hand tracking using a reduced number of available cameras in this configuration, thereby reducing power consumption and increasing battery life.

[0322] Method 2700 can be executed under the control of a processor of the XR system. The method can be initiated when an object to be tracked (e.g., a hand) is detected as a result of analyzing an image acquired using any one of the cameras on the head-mounted device 2100. The analysis may require identifying the object as a hand based on an image area having photometric features that are characteristic of a hand. Alternatively or additionally, depth information acquired based on stereo image analysis can be used to detect the hand. As a specific example, the depth information may indicate the presence of an object having a shape that matches a 3D model of a hand. Detecting the presence of a hand in this manner may also require setting the parameters of the hand model to match the orientation of the hand. In some embodiments, such a model can also be used for fast hand tracking by using photometric information from one or more grayscale cameras to determine how the hand has moved from its original position.

[0323] Other trigger conditions can initiate method 2700, such as the XR system performing an operation involving a tracked object, such as rendering a virtual button that a user may attempt to press with a hand, thereby anticipating the user's hand entering the field of view of one or more cameras. Method 2700 can be repeated at a relatively high rate, such as between 30 and 100 times per second, such as between 40 and 60 times per second. As a result, updated position information of the tracked object can be made available with low latency for use in rendering processes for virtual objects interacting with physical objects.

[0324] After starting in block 2701, method 2700 may proceed to block 2710. In block 2710, the processor may determine a potential hand position. In some embodiments, the potential hand position may be the position of an object detected in the acquired image. In embodiments where the hand is detected based on matching depth information to a 3D model of the hand, the same information may be used as the initial hand position at block 2710.

[0325] After block 2710, method 2700 may proceed to block 2720. In block 2720, the processor may determine whether object tracking robustness or fine object tracking detail is desired. In this particular case, method 2700 may proceed to block 2730. In block 2730, the processor may obtain depth information for the object. This depth information may be obtained based on stereo image analysis, from which the distance between the camera collecting the image information and the tracked object may be calculated. For example, the processor may select a feature in the center field of view and determine the depth information for the selected feature. The processor may determine the depth information for the feature stereoscopically using images acquired by cameras 2120 and 2140.

[0326] In some embodiments, the selected features can represent different segments of a human hand defined by bones and joints. Feature selection can be based on matching image information with a human hand model. For example, this matching can be performed heuristically. For example, a human hand can be represented by a finite number of segments, such as 16 segments, and points in the image of the hand can be mapped to one of those segments, so that features on each segment can be selected. Alternatively or additionally, this matching can use a deep neural network or classification / decision forest to apply a series of yes / no decisions in the analysis to identify different parts of the hand and select features representing different parts of the hand. For example, matching can identify whether a specific point in the image belongs to the palm, the back of the hand, a non-thumb finger, the thumb, a fingertip, and / or a knuckle. Any suitable classifier can be used for this analysis stage. For example, a deep learning module or a neural network mechanism can be used instead of or in addition to a classification forest. In addition, in addition to the classification forest, a regression forest (e.g., using the Hough transform, etc.) can also be used.

[0327] Regardless of the specific number of features selected and the technique used to select those features, after block 2730, method 2700 can proceed to block 2740. In block 2740, the processor can then configure a hand model based on the depth information. In some embodiments, the hand model can reflect structural information about the human hand, for example, representing each bone in the hand as a segment in the hand, and each joint defining the possible angle range between adjacent segments. By assigning a position to each segment in the hand model based on the depth information of the selected features, information about the position of the hand can be provided for subsequent processing by the XR system.

[0328] In some embodiments, the processing at blocks 2730 and 2740 may be performed iteratively, where the selection of features for which depth information is collected is refined based on the configuration of the hand model. The hand model may include shape constraints and movement constraints, which the processor may be configured to use to refine the selection of features representing portions of the hand. For example, when a feature selected to represent a segment of the hand indicates that the position or movement of the segment violates a hand model constraint, a different feature may be selected to represent the segment.

[0329] In some embodiments, photometric image information can be used instead of or in addition to depth information to perform successive iterations of the hand tracking process. At each iteration, the 3D model of the hand can be updated to reflect the potential movement of the hand. Potential movement can be determined, for example, based on depth information, photometric information, or a projection of the hand trajectory. If depth information is used, it can apply a more limited set of features than those used to set the initial configuration of the hand model, thereby speeding up the process.

[0330] Regardless of how the 3D hand model is updated, the updated model can be refined based on the photometric image information. For example, the model can be used to render a virtual image of the hand, representing how the hand image is expected to appear. This expected image can be compared with the photometric image information acquired using the image sensor. The 3D model can be adjusted to reduce the error between the expected and acquired photometric information. The adjusted 3D model then provides an indication of the hand's position. As this updating process is repeated, the 3D model provides an indication of the hand's position as the hand moves.

[0331] The XR system may alternatively determine that robustness or fine detail is not required, or that the tracked object may be outside the central view 2150. In this case, in block 2750, the XR system may be configured to acquire grayscale image information from camera 2120. The processor may then select features in the image that represent the structure of a person's hand. Such features may be identified heuristically or using AI techniques, as described above with respect to the processing at block 2730. For example, as described above, features may be heuristically selected by representing a person's hand in a finite number of segments and mapping points in the image to corresponding segments in those segments, such that features on each segment may be selected. Alternatively or additionally, this matching may utilize a deep neural network or classification / decision forest to apply a series of yes / no decisions in the analysis to identify different parts of the hand and select features representing the different parts of the hand. Any suitable classifier may be used in this analysis phase. For example, a deep learning module or neural network mechanism may be used in place of or in addition to a classification forest. Furthermore, a regression forest (e.g., using a Hough transform, etc.) may also be used in addition to a classification forest.

[0332] At block 2760, the processor may then attempt to match the selected features and the movement of those selected features between images to the hand model without the aid of depth information. This matching may result in less robust or potentially less accurate information than that generated in block 2740. Nonetheless, the information identified based on the monocular information may provide useful information for the operation of the XR system.

[0333] After matching the image portion with the portion of the hand model in block 2740 or 2760, the processor may use the determined hand model information to recognize a hand gesture in block 270. This gesture recognition may be performed using the hand tracking method described in U.S. Patent Publication No. 2016 / 0026253, which is incorporated herein by reference in its entirety, which teaches the use of information about the hand obtained from image information in conjunction with hand tracking in an XR system.

[0334] Following block 2770, method 2700 may end in block 2799. However, it should be understood that object tracking may occur continuously during operation of the XR system or may occur during intervals when an object is in the field of view of one or more cameras. Thus, once one iteration of method 2700 is completed, another iteration may be performed, and the process may be performed during the intervals in which object tracking is being performed. In some embodiments, information used in one iteration may be used in subsequent iterations. In various embodiments, for example, the processor may be configured to estimate an updated position of a user's hand based on previously detected hand positions. For example, the processor may estimate where the user's hand will be next based on the previous position and velocity of the user's hand. Such information may be used to reduce the amount of image information processed to detect the position of an object, as described above in conjunction with the block tracking techniques.

[0335] Having thus described several aspects of some embodiments, it is to be understood that various alterations, modifications, and improvements will readily occur to those skilled in the art.

[0336] As an example, embodiments are described in conjunction with an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein may be applied to an MR environment or more generally to other XR environments.

[0337] Furthermore, embodiments of image arrays are described in which a partition is applied to the image array to control the selective output of image information about a movable object. It will be appreciated that there may be more than one movable object in a physical environment. Furthermore, in some embodiments, it may be desirable to selectively obtain frequent updates of image information in areas other than the area where the movable object is located. For example, partitions may be configured to selectively obtain image information about an area of the physical world where a virtual object is to be rendered. Thus, some image sensors may be capable of selectively providing information from two or more partitions, with or without circuitry for tracking the trajectories of these partitions.

[0338] As another example, an image array is described as outputting information related to the amplitude of incident light. The amplitude can be a representation of power across a spectrum of light frequencies. This spectrum can be relatively broad, capturing energy at frequencies corresponding to any color of visible light, such as in a black and white camera. Alternatively, the spectrum can be narrow, corresponding to a single color of visible light. To this end, a filter can be used that limits the light incident on the image array to that of a particular color. Where pixels are restricted to receiving light of a particular color, different pixels can be restricted to different colors. In such an embodiment, the outputs of pixels sensitive to the same color can be processed together.

[0339] A process is described for setting up tiles in an image array and then updating the tiles for objects of interest. For example, the process can be performed for each movable object as it enters the field of view of the image sensor. When the object of interest leaves the field of view, the tile can be cleared so that the tile is no longer tracked or image information is no longer output for the tile. It will be understood that the tiles can be updated from time to time, for example by determining the position of an object associated with the tile and setting the position of the tile to correspond to that position. Similar adjustments can be made to the calculated trajectory of the tile. The motion vector of the object and / or the motion vector of the image sensor can be calculated based on other sensor information and used to reset values programmed into the image sensor or other components for tile tracking.

[0340] For example, the position, movement, and other characteristics of an object can be determined by analyzing the output of a wide-angle camera or a pair of cameras with stereo information. Data from these other sensors can be used to update the world model. In conjunction with the updates, the tile positions and / or trajectory information can be updated. Such updates can occur at a lower rate than the tile tracking engine updates the tile positions. For example, the tile tracking engine can calculate new tile positions at a rate of approximately 1 to 30 times per second. Updates to tile positions based on other information can occur at a slower rate, such as once per second to approximately once every 30 seconds.

[0341] As another example of a variation, Figure 2 A system is shown with a head mounted display that is separate from the remote processing module. An image sensor as described herein can result in a compact design of the system. Such a sensor generates less data, which in turn results in lower processing requirements and lower power consumption. The reduced processing and power requirements enable size reduction by reducing the size of the battery. Thus, in some embodiments, the entire augmented reality system can be integrated into a head mounted display without a remote processing module. The head mounted display can be configured as a pair of goggles, or as Figure 2 As shown, it can be similar in size and shape to a pair of glasses.

[0342] Furthermore, embodiments are described in which the image sensor responds to visible light. It should be understood that the techniques described herein are not limited to operation with visible light. They may alternatively or additionally respond to IR light or "light" in other parts of the spectrum, such as UV. Furthermore, the image sensors described herein respond to naturally occurring light. Alternatively or additionally, the sensors may be used in systems with an illumination source. In some embodiments, the sensitivity of the image sensor may be tuned to the portion of the spectrum where the illumination source emits light.

[0343] As another example, the image sensor is described as outputting changes in a selected region of an image array by specifying a "tile" on which image analysis is to be performed. However, it should be understood that the tile and the selected region can be of different sizes. For example, the selected region can be larger than the tile to account for deviations from the predicted trajectory of a tracked object in the image and / or to enable processing around the edges of the tile.

[0344] In addition, multiple processes are described, such as navigable world model generation, object tracking, head pose tracking, and hand tracking. These and other processes in some embodiments can be performed by the same or different processors. The processor can be operated so that these processes can operate concurrently. However, each process can be executed at a different rate. In the case where different processes request data from image sensors or other sensors at different rates, the acquisition of sensor data can be managed, for example by another process, so that data is provided to each process at a rate appropriate for its operation.

[0345] Such changes, modifications, and improvements are intended to be part of the present disclosure and are intended to be within the spirit and scope of the present disclosure. For example, in some embodiments, the filter 102 of a pixel of an image sensor may not be a separate component, but rather incorporated into one of the other components of the pixel subarray 100. For example, in an embodiment including a single pixel having an arrival angle to position intensity converter and an optical filter, the arrival angle to intensity converter may be a transmissive optical component formed of a material that filters a specific wavelength.

[0346] According to some embodiments, a wearable display system may be provided, comprising: a head-mounted device having a grayscale camera and a color camera positioned to provide overlapping views of a central field of view; and a processor operably coupled to the grayscale camera and the color camera and configured to: create a world model using first depth information stereoscopically determined from images acquired by the grayscale camera and the color camera; and track a head pose using the grayscale camera and the world model.

[0347] In some embodiments, the grayscale camera may include a plenoptic camera, and the processor may be further configured to acquire light field information using the plenoptic camera, and perform a world model update routine on portions of the world model that meet the depth criteria using the light field information.

[0348] In some embodiments, portions of the world model that are less than 1.5 meters from the head-mounted device may meet the depth criteria.

[0349] In some embodiments, the processor may also be configured to disable the color camera or reduce the frame rate of the color camera.

[0350] In some embodiments, the processor may be further configured to evaluate a head pose tracking quality criterion and, based on the evaluation, enable the color camera and track the head pose using depth information stereoscopically determined from images acquired by the grayscale camera and the color camera.

[0351] In some embodiments, the grayscale camera may include a plenoptic camera and

[0352] Head pose tracking can use light field information acquired by a grayscale camera.

[0353] In some embodiments, the plenoptic camera may include a transmissive diffraction mask.

[0354] In some embodiments, the horizontal field of view of the grayscale camera may range between 90 degrees and 175 degrees and the central field of view may range between 40 degrees and 120 degrees.

[0355] In some embodiments, the world model may include three-dimensional blocks, the three-dimensional blocks including voxels, and the world model update routine may include identifying the blocks within a viewing frustum, the viewing frustum having a maximum depth, and updating the voxels within the identified blocks using depth information acquired using the plenoptic camera.

[0356] In some embodiments, the processor may further be configured to: determine whether the object is within the central field of view; and based on the determination, track the object using depth information stereoscopically determined from images acquired by the grayscale camera and the color camera when the object is within the central view; and track the object using one or more images acquired by the color camera when the object is within the peripheral field of view outside the central field of view of the color camera.

[0357] In some embodiments, the head mounted device may weigh between 30 and 300 grams, and the processor may be further configured to perform a calibration routine to determine the relative orientation of the grayscale camera and the color camera.

[0358] In some embodiments, the calibration routine may include identifying corresponding features in images acquired with each of the two cameras; calculating an error for each of a plurality of estimated relative orientations of the two cameras, wherein the error indicates a difference between the corresponding feature appearing in the images acquired with each of the two cameras and an estimate of the identified feature calculated based on the estimated relative orientations of the two cameras; and selecting a relative orientation from the plurality of estimated relative orientations as the determined relative orientation based on the calculated error.

[0359] In some embodiments, the processor may be further configured to repeatedly perform the calibration routine while the wearable display system is worn, such that the calibration routine compensates for distortions in the frame during use of the wearable display system.

[0360] In some embodiments, the calibration routine may compensate for distortions in the frame caused by temperature changes.

[0361] In some embodiments, the calibration procedure can compensate for distortions in the frame caused by mechanical strain.

[0362] In some embodiments, the processor may be mechanically coupled to the head-mounted device.

[0363] In some embodiments, the head mounted device may include a display device mechanically coupled to the processor.

[0364] In some embodiments, the local data processing module may include a processor, the local data processing module is operably coupled to the display device via a communication link, and wherein the head-mounted device may include the display device.

[0365] According to some embodiments, a wearable display system is provided, comprising: a frame; a first camera mechanically coupled to the frame and a second camera mechanically coupled to the frame, wherein the first camera and the second camera are positioned to provide a central field of view associated with the first camera and the second camera, and wherein at least one of the first camera and the second camera comprises a plenoptic camera; a processor operably coupled to the first camera and the second camera and configured to: determine whether an object is within the central field of view; when the object is within the central field of view, determine whether the object meets a depth criterion; when the tracked object is within the central field of view and does not meet the depth criterion, track the object using depth information determined stereoscopically from images acquired by the first camera and the second camera; when the tracked object is within the central field of view and meets the depth criterion, track the object using depth information determined from light field information acquired by one of the first camera and the second camera.

[0366] In some embodiments, the object may be the hand of the wearer of the head mounted device.

[0367] In some embodiments, the plenoptic camera may include a transmissive diffraction mask.

[0368] In some embodiments, the first camera may include a plenoptic camera, and the processor may be further configured to: create a world model using images acquired by the first camera and the second camera; update the world model using light field information acquired by the first camera; determine that the updated world model does not meet a quality standard; and based on the determination, enable the second camera.

[0369] In some embodiments, the horizontal field of view of the first camera may range between 90 degrees and 175 degrees, and the central field of view may range between 40 degrees and 80 degrees.

[0370] In some embodiments, the processor may be further configured to provide instructions to at least one of the first camera or the second camera to limit image capture to a subset of pixels.

[0371] In some embodiments, the second camera may be positioned to provide a peripheral field of view, and the processor may be further configured to determine whether the object is within the peripheral field of view and track the object using the second camera when the object to be tracked is within the peripheral field of view.

[0372] In some embodiments, the second camera may include a plenoptic camera, and the processor may be further configured to track the object using light field information obtained from the second camera when the tracked object is within the peripheral field of view and meets the depth criteria.

[0373] In some embodiments, the processor may be further configured to determine whether the object has moved into the peripheral field of view, and in response to the determination, enable the second camera or increase the frame rate of the second camera.

[0374] In some embodiments, tracking an object using depth information when the object is within the central field of view may include selecting a point in the central field of view; determining depth information for the selected point using at least one of depth information determined stereoscopically based on images acquired by the first camera and the second camera or depth information determined based on light field information acquired by one of the first camera or the second camera; generating a depth map using the determined depth information; and matching portions of the depth map with corresponding portions of the hand model including shape constraints and motion constraints.

[0375] In some embodiments, the processor may be further configured to track hand motion in the peripheral field of view using one or more images acquired from the second camera by matching portions of the images to corresponding portions of the hand model including shape constraints and motion constraints.

[0376] In some embodiments, the processor may be mechanically coupled to the frame.

[0377] In some embodiments, a display device mechanically coupled to the frame may include a processor.

[0378] In some embodiments, the local data processing module may include a processor, the local data processing module operatively coupled to the display device via a communication link, the display device mechanically coupled to the frame.

[0379] According to some embodiments, a wearable display system may be provided, comprising: a frame; two cameras mechanically coupled to the frame, wherein the two cameras comprise: a first camera with a global shutter having a first field of view; and a second camera with a rolling shutter having a second field of view; wherein the first camera and the second camera are positioned to provide: a central field of view, wherein the first field of view overlaps with the second field of view; a peripheral field of view outside the central field of view; and a processor operably coupled to the first camera and the second camera.

[0380] In some embodiments, the first camera and the second camera may be angled inward asymmetrically.

[0381] In some embodiments, the field of view of the first camera may be larger than the field of view of the second camera.

[0382] In some embodiments, the first camera may be angled inward between 20 and 40 degrees and the second camera may be angled inward between 1 and 20 degrees.

[0383] In some embodiments, the first camera may be a color camera and the second camera may be a grayscale camera.

[0384] In some embodiments, at least one of the two cameras may be a plenoptic camera.

[0385] In some embodiments, the plenoptic camera may include a transmissive diffraction mask.

[0386] In some embodiments, the processor may be mechanically coupled to the frame.

[0387] In some embodiments, a display device mechanically coupled to the frame may include a processor.

[0388] In some embodiments, the horizontal field of view of the first camera may range between 90 degrees and 175 degrees, and the central field of view may range between 40 degrees and 80 degrees.

[0389] Furthermore, although the advantages of the present disclosure have been pointed out, it should be understood that not every embodiment of the present disclosure will include every described advantage. Some embodiments may not implement any of the features described herein as advantageous. Therefore, the foregoing description and accompanying drawings are intended only as examples.

[0390] The above-mentioned embodiments of the present disclosure can be implemented in any of a variety of ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or set of processors, whether provided in a single computer or distributed in multiple computers. Such a processor can be implemented as an integrated circuit, having one or more processors in an integrated circuit component, including commercially available integrated circuit components known in the art, with names such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors. In some embodiments, the processor can be implemented in a custom circuit (such as an ASIC) or in a semi-custom circuit generated by configuring a programmable logic device. As another alternative, the processor can be part of a larger circuit or semiconductor device, whether commercially available, semi-custom, or custom. As a specific example, some commercially available microprocessors have multiple cores, so that one or a subset of these cores can constitute a processor. However, the processor can be implemented using circuits of any appropriate format.

[0391] Furthermore, it should be understood that a computer may be embodied in any of a variety of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer may be embedded in a device not generally considered a computer but having suitable processing capabilities, including a personal digital assistant (PDA), a smart phone, or any other suitable portable or fixed electronic device.

[0392] In addition, the computer may have one or more input and output devices. These devices may be used, in particular, to present a user interface. Examples of output devices that may be used to provide a user interface include printers or display screens for visual presentation output, and speakers or other sound generating devices for auditory presentation output. Examples of input devices that may be used for a user interface include keyboards and pointing devices, such as mice, touchpads, and digital tablet computers. As another example, a computer may receive input information through voice recognition or other audible formats. In the illustrated embodiment, the input / output devices are shown as being physically separated from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated into the same unit as the processor or other elements of the computing device. For example, a keyboard may be implemented as a soft keyboard on a touch screen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated via a wireless connection.

[0393] Such computers may be interconnected by one or more networks of any suitable form, including as a local area network or a wide area network such as an enterprise network or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.

[0394] Furthermore, the various methods or processes outlined herein may be encoded as software executable on one or more processors employing any of a variety of operating systems or platforms. Furthermore, such software may be written using any of a variety of suitable programming languages and / or programming or scripting tools and further compiled into executable machine language code or intermediate code that is executed on a framework or virtual machine.

[0395] In this regard, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., a computer memory, one or more floppy disks, compact discs (CDs), optical disks, digital video discs (DVDs), magnetic tapes, flash memory, field programmable gate arrays or other semiconductor devices or circuitry in other tangible computer storage media) encoded with one or more programs that, when executed on one or more computers or other processors, will perform methods that implement the various embodiments of the present disclosure discussed above. As will be apparent from the foregoing examples, the computer-readable storage medium may retain information for a sufficient time to provide computer-executable instructions in a non-transient form. Such one or more computer-readable storage media may be removable so that one or more programs stored thereon may be loaded onto one or more different computers or other processors to implement various aspects of the present disclosure as described above. As used herein, the term "computer-readable storage medium" encompasses only computer-readable media that can be considered an article of manufacture (i.e., an article of manufacture) or a machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagating signal.

[0396] The terms "program" or "software" are used herein in a general sense to refer to computer code or a set of computer-executable instructions that can be used to program a computer or other processor to implement various aspects of the present disclosure as described above. In addition, it should be understood that according to one aspect of this embodiment, one or more computer programs that, when executed, perform the methods of the present disclosure need not reside on a single computer or processor, but can be distributed in a modular manner among multiple different computers or processors to implement various aspects of the present disclosure.

[0397] Computer-executable instructions can take many forms, such as program modules, which are executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In general, in various embodiments, the functionality of program modules can be combined or distributed as needed.

[0398] In addition, the data structure can be stored in a computer-readable medium in any suitable form. To simplify the description, the data structure can be shown to have fields that are related by position in the data structure. Similarly, such relationships can be achieved by allocating storage for the fields by conveying the positions in the computer-readable medium that convey the relationships between the fields. However, any suitable mechanism can be used to establish the relationships between the information in the fields of the data structure, including by using pointers, tags, or other mechanisms that establish relationships between data elements.

[0399] The various aspects of the present disclosure may be used alone, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and therefore, are not limited in their application to the details and arrangements of components set forth in the foregoing description or shown in the accompanying drawings. For example, aspects described in one embodiment may be combined in any manner with aspects described in other embodiments.

[0400] Furthermore, the present disclosure may be embodied as a method, an example of which has been provided. The actions performed as part of the method may be ordered in any suitable manner. Thus, embodiments may be constructed in which actions are performed in an order different from that shown, and even if shown as sequential actions in an illustrative embodiment, the actions may include performing some actions simultaneously.

[0401] The use of ordinal terms such as “first,” “second,” “third,” etc. in the claims to modify claim elements does not, in itself, indicate any priority, precedence, or order of one claim element with respect to another sequential or temporal order of performing the method actions, but serves merely as a marker to distinguish one claim element having a certain name from another element having the same name (but used in ordinal numbers) to distinguish the claim elements.

[0402] In addition, the words and terms used herein are for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," or "having," "containing," "involving," and variations thereof herein is intended to encompass the items listed thereafter and their equivalents as well as other items.

Claims

1. A wearable display system, comprising: a head-mounted device comprising a first camera having a global shutter and a second camera having a rolling shutter, the first camera and the second camera being positioned to provide overlapping views of a central field of view; as well as a processor operatively coupled to the first camera and the second camera and configured to: performing a compensation routine to adjust a second image acquired using the second camera for rolling shutter image distortion, wherein the compensation routine comprises: detecting a skew of at least a portion of the second image; and adjusting the at least a portion of the second image to compensate for the detected skew; and A world model is created using, in part, depth information stereoscopically determined from images acquired using the first camera and the adjusted images.

2. The wearable display system according to claim 1, wherein: The first camera and the second camera are asymmetrically angled inward.

3. The wearable display system according to claim 2, wherein: The field of view of the first camera is larger than the field of view of the second camera.

4. The wearable display system according to claim 2, wherein: The first camera is angled inwardly between 20 and 40 degrees, and the second camera is angled inwardly between 1 and 20 degrees.

5. The wearable display system according to claim 1, wherein: The first camera has an angular pixel resolution between 1 arc minute and 5 arc minutes per pixel.

6. The wearable display system according to claim 1, wherein: The processor is further configured to perform a downsizing routine to resize images acquired using the second camera.

7. The wearable display system according to claim 6, wherein: The downsizing routine includes generating a downsized image by binning pixels in an image acquired by the second camera.

8. The wearable display system according to claim 1, wherein: The compensation routine also includes: A first image acquired using the first camera is compared with a second image acquired using the second camera to detect the skew.

9. The wearable display system according to claim 8, wherein: Comparing a first image acquired using the first camera with a second image acquired using the second camera includes performing a line-by-line comparison between the first image acquired by the first camera and the second image acquired by the second camera.

10. The wearable display system according to claim 1, wherein: The processor is further configured to: The second camera is disabled or a frame rate of the second camera is modulated based on at least one of a power conservation criterion or a world model integrity criterion.

11. The wearable display system according to claim 1, wherein: At least one of the first camera or the second camera comprises a plenoptic camera; and The processor is further configured to: A world model is created using, in part, light field information acquired by the at least one of the first camera or the second camera.

12. The wearable display system according to claim 1, wherein: The first camera comprises a plenoptic camera; and The processor is further configured to: A world model update routine is performed using depth information acquired using the plenoptic camera.

13. The wearable display system according to claim 1, wherein: The processor is mechanically coupled to the head-mounted device.

14. The wearable display system according to claim 1, wherein: The head mounted device includes a display device mechanically coupled to the processor.

15. The wearable display system according to claim 1, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head-mounted device includes the display device.

16. A method for creating a world model using a wearable display system, the wearable display system comprising: a head-mounted device comprising a first camera having a global shutter and a second camera having a rolling shutter, the first camera and the second camera being positioned to provide overlapping views of a central field of view; as well as a processor operatively coupled to the first camera and the second camera; wherein the method comprises using the processor to: performing a compensation routine to adjust a second image acquired using the second camera for rolling shutter image distortion, wherein the compensation routine comprises: detecting a skew of at least a portion of the second image; and adjusting the at least a portion of the second image to compensate for the detected skew; and The world model is created in part using depth information determined stereoscopically from images acquired using the first camera and the adjusted images.

17. The method according to claim 16, wherein The first camera and the second camera are asymmetrically angled inward.

18. The method according to claim 17, wherein The field of view of the first camera is larger than the field of view of the second camera.

19. The method according to claim 17, wherein The first camera is angled inwardly between 20 and 40 degrees, and the second camera is angled inwardly between 1 and 20 degrees.

20. The method according to claim 16, wherein The first camera has an angular pixel resolution between 1 arc minute and 5 arc minutes per pixel.

21. The method according to claim 16, wherein The method includes using the processor to: perform a downsizing routine to resize an image acquired using the second camera.

22. The method according to claim 21, wherein The downsizing routine includes generating a downsized image by binning pixels in an image acquired by the second camera.

23. The method of claim 16, wherein: The compensation routine also includes: A first image acquired using the first camera is compared with a second image acquired using the second camera to detect the skew.

24. The method of claim 23, wherein: Comparing a first image acquired using the first camera with a second image acquired using the second camera includes performing a line-by-line comparison between the first image acquired by the first camera and the second image acquired by the second camera.

25. The method of claim 16, wherein: The method includes using the processor to: The second camera is disabled or a frame rate of the second camera is modulated based on at least one of a power conservation criterion or a world model integrity criterion.

26. The method of claim 16, wherein: At least one of the first camera or the second camera comprises a plenoptic camera; and The method includes using the processor to: The world model is created using, in part, light field information acquired by the at least one of the first camera or the second camera.

27. The method of claim 16, wherein: The first camera comprises a plenoptic camera; and The method includes using the processor to: A world model update routine is performed using depth information acquired using the plenoptic camera.

28. The method according to claim 16, wherein The processor is mechanically coupled to the head-mounted device.

29. The method according to claim 16, wherein The head mounted device includes a display device mechanically coupled to the processor.

30. The method of claim 16, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head-mounted device includes the display device.

31. A wearable display system, comprising: frame; two cameras mechanically coupled to the frame, wherein the two cameras comprise: a first camera with a global shutter having a first field of view; and a second camera with a rolling shutter having a second field of view; wherein the first camera and the second camera are positioned to provide: a central field of view in which the first field of view overlaps with the second field of view; and a peripheral field of view outside the central field of view; and a processor operatively coupled to the first camera and the second camera and configured to: performing a compensation routine to adjust a second image acquired using the second camera for rolling shutter image distortion, wherein the compensation routine comprises: detecting a skew of at least a portion of the second image; and The at least a portion of the second image is adjusted to compensate for the detected skew.

32. The wearable display system of claim 31 , further comprising a head-mounted device, the head-mounted device comprising the frame, the first camera, and the second camera, wherein The processor is further configured to: A world model is created using, in part, depth information stereoscopically determined from images acquired using the first camera and the adjusted images.

33. The wearable display system according to claim 31, wherein: The first camera and the second camera are asymmetrically angled inward.

34. The wearable display system according to claim 33, wherein: The field of view of the first camera is larger than the field of view of the second camera.

35. The wearable display system according to claim 33, wherein: The first camera is angled inwardly between 20 and 40 degrees, and the second camera is angled inwardly between 1 and 20 degrees.

36. The wearable display system according to claim 31, wherein: The first camera has an angular pixel resolution between 1 arc minute and 5 arc minutes per pixel.

37. The wearable display system according to claim 32, wherein: The processor is further configured to perform a downsizing routine to resize images acquired using the second camera.

38. The wearable display system according to claim 37, wherein: The downsizing routine includes generating a downsized image by binning pixels in an image acquired by the second camera.

39. The wearable display system of claim 31 , wherein: The compensation routine also includes: A first image acquired using the first camera is compared with a second image acquired using the second camera to detect the skew.

40. The wearable display system of claim 39, wherein: Comparing a first image acquired using the first camera with a second image acquired using the second camera includes performing a line-by-line comparison between the first image acquired by the first camera and the second image acquired by the second camera.

41. The wearable display system of claim 32, wherein: The processor is further configured to: The second camera is disabled or a frame rate of the second camera is modulated based on at least one of a power conservation criterion or a world model integrity criterion.

42. The wearable display system of claim 32, wherein: At least one of the first camera or the second camera comprises a plenoptic camera; and The processor is further configured to: A world model is created using, in part, light field information acquired by the at least one of the first camera or the second camera.

43. The wearable display system of claim 32, wherein: The first camera comprises a plenoptic camera; and The processor is further configured to: A world model update routine is performed using depth information acquired using the plenoptic camera.

44. The wearable display system of claim 32, wherein: The processor is mechanically coupled to the head-mounted device.

45. The wearable display system of claim 32, wherein: The head mounted device includes a display device mechanically coupled to the processor.

46. The wearable display system of claim 32, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head-mounted device includes the display device.

47. The wearable display system of claim 31 , wherein: The wearable display system is configured to perform object tracking in the peripheral field of view using images acquired by the first camera having the global shutter.

Citation Information

Patent Citations

  • Methods and systems for creating virtual and augmented reality

    US20160026253A1

  • Information processing device, position and / or orientation estimating method, and computer program

    CN108139204A