Lightweight and low-power cross-reality device with high temporal resolution
By using DVS cameras and processors in wearable XR systems, creating world models and tracking head postures, the weight, power consumption and delay problems in the prior art are solved, and efficient and accurate acquisition of physical world information is achieved.
Patent Information
- Application Number
- CN202080026109.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-07
- Filing Date
- 2020-02-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-02-07
AI Technical Summary
Existing XR systems face weight, power consumption and delay problems when obtaining information about physical world objects, affecting user experience and system practicality.
A wearable head-mounted device equipped with a DVS camera, which includes two cameras and a processor, determines depth information through stereoscopic image data, creates a world model, and uses the model to track head posture.
It realizes accurate acquisition of information about physical world objects under low latency and low power consumption, reduces system weight, improves user experience and system practicality.
Smart Images

Figure CN113678435B_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to a wearable cross-reality display system (XR system) including a dynamic vision sensor (DVS) camera. Background Art
[0002] A computer can control a human user interface to create an X Reality (XR or cross reality) environment in which some or all of the XR environments perceived by the user are generated by the computer. These XR environments can be virtual reality (VR), augmented reality (AR), or mixed reality (MR) environments, some or all of which can be generated by a computer using data describing the environment in part. For example, the data can describe a virtual object, which can be rendered in a way that the user feels or perceives it as part of the physical world so that the user can interact with the virtual object. Since the data is rendered and presented by a user interface device (such as, for example, a head-mounted display device), the user can experience these virtual objects. The data can be displayed to the user, or the audio played to the user can be controlled, or a tactile (or tactile) interface can be controlled, so that the user can experience the touch sensation that the user feels or perceives as feeling the virtual object.
[0003] XR systems can be used for many applications across scientific visualization, medical training, engineering design and prototyping, telemanipulation and telepresence, and personal entertainment. Compared to VR, AR and MR include one or more virtual objects related to real objects in the physical world. The experience of virtual objects interacting with real objects significantly enhances the user's enjoyment of using XR systems, and also opens the door to a variety of applications that present realistic and easy-to-understand information about how to change the physical world. Summary of the invention
[0004] Aspects of the present application relate to a wearable cross-reality display system configured with a DVS camera.The techniques described herein can be used together, separately, or in any suitable combination.
[0005] According to some embodiments, a wearable display system may be provided, comprising: a head-mounted device, comprising: a first camera, configured to output image frames or image data that meet an intensity variation criterion; and a second camera; wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and a processor, operably coupled to the first camera and the second camera and configured to: create a world model using depth information stereoscopically determined based on images output by the first camera and the second camera; and track head posture using the world model and the image data output by the first camera.
[0006] In some embodiments, the intensity variation criteria may include absolute or relative intensity variation criteria.
[0007] In some embodiments, the first camera may be configured to output the image data asynchronously.
[0008] In some embodiments, the processor may be further configured to track head pose asynchronously.
[0009] In some embodiments, the processor may be further configured to perform a tracking routine to restrict image data acquisition to points of interest within the world model.
[0010] In some embodiments, the first camera may be configured to limit image acquisition to one or more portions of a field of view of the first camera; and the tracking routine may include: identifying a point of interest within the world model; determining one or more first portions of the field of view of the first camera corresponding to the point of interest; and providing instructions to the first camera to limit image acquisition to the one or more first portions of the field of view.
[0011] In some embodiments, the tracking routine may also include: estimating one or more second portions of the field of view of the first camera corresponding to the point of interest based on movement of the point of interest relative to the world model or movement of the head-mounted device relative to the point of interest; and providing instructions to the first camera to limit image acquisition to the one or more second portions of the field of view.
[0012] In some embodiments, the head-mounted device may also include an inertial measurement unit; and performing the tracking routine may include: estimating an updated relative position of the object based at least in part on an output of the inertial measurement unit.
[0013] In some embodiments, the tracking routine may include: repeatedly calculating the position of the point of interest within the world model; and the repeated calculation may be performed at a temporal resolution exceeding 60 Hz.
[0014] In some embodiments, the intervals between the repeated calculations may be between 1 millisecond and 15 milliseconds in duration.
[0015] In some embodiments, the processor may be further configured to: determine whether the head pose tracking meets a quality standard; and when the head pose tracking does not meet the quality standard, enable the second camera or modulate the frame rate of the second camera.
[0016] In some embodiments, the processor may be mechanically coupled to the head mounted device.
[0017] In some embodiments, the head mounted device may include a display device mechanically coupled to the processor.
[0018] In some embodiments, a local data processing module may include the processor, the local data processing module may be operably coupled to a display device via a communication link, and wherein the head mounted device may include the display device.
[0019] In some embodiments, the head mounted device may also include an IR transmitter.
[0020] In some embodiments, the processor may be configured to selectively enable the IR emitter to enable head gesture tracking in low light conditions.
[0021] According to some embodiments, a method for tracking head posture using a wearable display system can be provided, the wearable display system comprising: a head-mounted device, comprising: a first camera, which can be configured to output image frames or image data that meet an intensity change criterion; and a second camera; wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and a processor, which is operably coupled to the first camera and the second camera; wherein the method includes using the processor to: create a world model using depth information determined stereoscopically based on images output by the first camera and the second camera; and track head posture using the world model and the image data output by the first camera.
[0022] According to some embodiments, a wearable display system may be provided, comprising: a frame; a first camera mechanically coupled to the frame, wherein the first camera is capable of being configured to output image data satisfying an intensity variation criterion in a first field of view of the first camera; and a processor operably coupled to the first camera and configured to: determine whether an object is within the first field of view; and track movement of the object using image data received from the first camera for one or more portions of the first field of view.
[0023] According to some embodiments, a method for tracking the movement of an object using a wearable display system may be provided, the wearable display system comprising: a frame; a first camera mechanically coupled to the frame, wherein the first camera is capable of being configured to output image data that satisfies an intensity variation criterion in a first field of view of the first camera; and a processor operably coupled to the first camera; wherein the method comprises using the processor to: determine whether the object is within the first field of view; and, for one or more portions of the first field of view, track the movement of the object using image data received from the first camera.
[0024] According to some embodiments, a wearable display system may be provided, comprising: a frame; two cameras mechanically coupled to the frame, wherein the two cameras comprise: a first camera configured to output image data satisfying an intensity variation criterion; and a second camera, wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and a processor operably coupled to the first camera and the second camera.
[0025] The foregoing summary is provided by way of illustration and is not intended to be limiting. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in various figures is represented by a like numeral. For clarity, not every component may be labeled in every figure. In the drawings:
[0027] Figure 1 is a schematic diagram illustrating an example of a simplified augmented reality (AR) scene according to some embodiments.
[0028] Figure 2 is a schematic diagram illustrating an example of an AR display system according to some embodiments.
[0029] Figure 3A is a schematic diagram showing an AR display system rendering AR content as the user moves through a physical world environment when the user is wearing the AR display system, according to some embodiments.
[0030] Figure 3B is a schematic diagram illustrating a viewing optics assembly and accompanying components according to some embodiments.
[0031] Figure 4 is a schematic diagram illustrating an image sensing system according to some embodiments.
[0032] Figure 5A is a diagram showing a method according to some embodiments Figure 4 Schematic diagram of a pixel unit in .
[0033] Figure 5B is a diagram showing a method according to some embodiments Figure 5A Schematic diagram of the output events of a pixel unit.
[0034] Figure 6 is a schematic diagram illustrating an image sensor according to some embodiments.
[0035] Figure 7 is a schematic diagram illustrating an image sensor according to some embodiments.
[0036] Figure 8 is a schematic diagram illustrating an image sensor according to some embodiments.
[0037] Fig. 9 is a simplified flow chart of a method for image sensing according to some embodiments.
[0038] Fig.10 According to some embodiments Fig. 9 A simplified flowchart of the actions of patch recognition.
[0039] Fig.11 According to some embodiments Fig. 9 Simplified flowchart of the actions for block-wise trajectory estimation.
[0040] Fig.12 is a diagram showing a viewpoint according to some embodiments Fig.11 Schematic diagram of block trajectory estimation.
[0041] Fig.13 is a diagram showing a change in viewpoint according to some embodiments Fig.11 Schematic diagram of block trajectory estimation.
[0042] Fig.14 is a schematic diagram illustrating an image sensing system according to some embodiments.
[0043] Fig.15 is a diagram showing a method according to some embodiments Fig.14 Schematic diagram of a pixel unit in .
[0044] Fig.16 is a schematic diagram of a pixel subarray according to some embodiments.
[0045] Fig.17A is a cross-sectional view of a plenoptic device with an arrival angle to intensity converter in the form of two aligned stacked transmissive diffraction masks (TDMs) according to some embodiments.
[0046] Fig. 17B is a cross-sectional view of an all-optical device with an arrival angle-to-intensity converter in the form of two non-aligned stacked TDMs according to some embodiments.
[0047] Fig.18A is a pixel subarray having color pixel units and angle-of-arrival pixel units according to some embodiments.
[0048] Fig.18B is a pixel subarray having color pixel units and angle-of-arrival pixel units according to some embodiments.
[0049] Fig.18C is a pixel subarray having white pixel cells and corner-of-arrival pixel cells according to some embodiments.
[0050] Fig.19A is a top view of a photodetector array with a single TDM according to some embodiments.
[0051] Fig.19B is a side view of a photodetector array with a single TDM according to some embodiments.
[0052] Fig. 20A is a top view of a photodetector array with multiple angle-of-arrival-to-intensity converters in TDM format, according to some embodiments.
[0053] Fig. 20B is a side view of a photodetector array with multiple TDMs according to some embodiments.
[0054] Fig. 20C is a side view of a photodetector array with multiple TDMs according to some embodiments.
[0055] Fig.21 is a schematic diagram of a headset including two cameras and associated components according to some embodiments.
[0056] Fig. 22 is a flow chart of a calibration routine according to some embodiments.
[0057] FIG. 23A to FIG. 23C According to some embodiments, Fig.21 An exemplary field of view diagram associated with a head-mounted device.
[0058] Fig.24 is a flow chart of a method for creating and updating a traversable world model according to some embodiments.
[0059] Fig.25 is a flow chart of a method for head pose tracking according to some embodiments.
[0060] Fig.26 is a flow chart of a method for object tracking according to some embodiments.
[0061] Fig. 27 is a flow chart of a method for hand tracking according to some embodiments. DETAILED DESCRIPTION
[0062] The inventors have recognized and understood design and operating techniques for wearable XR display systems that enhance the enjoyment and usefulness of such systems. These design and / or operating techniques can enable acquisition of information to perform a variety of functions, including hand tracking, head pose tracking, and world reconstruction using a limited number of cameras, which can be used to realistically render virtual objects so that they appear to realistically interact with physical objects. The wearable cross-reality display system can be lightweight and can consume low power in operation. The system can use sensors in a specific configuration to acquire image information about physical objects in the physical world with low latency. The system can perform various routines to improve the accuracy and / or realism of the displayed XR environment. Such routines can include calibration routines that improve the accuracy of stereo depth measurements even if the lightweight frame deforms during use, and routines that detect and resolve incomplete depth information in the model of the physical world around the user.
[0063] The weight of known XR system headsets can limit user enjoyment. Such XR headsets can weigh more than 340 grams (sometimes even more than 700 grams). In contrast, glasses may weigh less than 50 grams. Wearing such relatively heavy headsets for long periods of time can cause fatigue to users or distract their attention, thereby detracting from the ideal immersive XR experience. However, the inventors have recognized and understood that some designs that reduce the weight of the headset also increase the flexibility of the headset, making the lightweight headset susceptible to changes in sensor position or orientation during use or over time. For example, when a user wears a lightweight headset that includes camera sensors, the relative orientation of these camera sensors may shift. Changes in the camera spacing used for stereoscopic imaging may affect the ability of these headsets to obtain accurate stereo information, which depends on cameras having a known positional relationship relative to each other. Therefore, a calibration routine that can be repeated while wearing the headset can enable the lightweight headset to accurately obtain information about the world around the wearer of the headset using stereo imaging technology.
[0064] The need to equip XR systems with components to obtain information about objects in the physical world can also limit the usefulness and user enjoyment of these systems. While the information obtained is used to realistically render computer-generated virtual objects in the proper location and with the proper appearance relative to the physical objects, the need to obtain the information imposes limitations on the size, power consumption, and realism of XR systems.
[0065] For example, an XR system may use sensors worn by a user to obtain information about objects in the physical world around the user, including information about the location of physical world objects in the user's field of view. Challenges arise because objects may move relative to the user's field of view as the object moves in the physical world or the user changes his or her posture relative to the physical world so that the physical object enters or leaves the user's field of view, or the location of the physical object within the user's field of view changes. In order to present a realistic XR display, the model of the physical objects in the physical world must be updated frequently enough to capture these changes, processed with low enough latency, and accurately predicted into the future to cover the full latency path including rendering, so that when the virtual object is displayed, the virtual object displayed based on this information will have the appropriate location and appearance relative to the physical object. Otherwise, the virtual object will not be aligned with the physical object, and the combined scene including the physical object and the virtual object will appear unreal. For example, the virtual object may appear to float in space instead of resting on the physical object, or may appear to bounce around relative to the physical object. Errors in visual tracking are particularly amplified when the user is moving at high speed and there is significant movement in the scene.
[0066] These problems can be avoided by sensors that acquire new data at a high rate. However, the power consumed by such sensors can result in the need for larger batteries, increase the weight of the system, or limit the time such systems can be used. Similarly, processors that need to process data generated at a high rate can drain batteries and increase the weight of wearable systems, further limiting the usefulness or enjoyment of such systems. For example, one known approach is to operate at a higher resolution to capture sufficient visual detail and operate higher frame rate sensors to increase temporal resolution. An alternative solution might supplement this solution with an IR time-of-flight sensor, which might directly indicate the position of a physical object relative to the sensor, and simple processing might be performed when using this information to display virtual objects, resulting in low latency. However, such sensors consume a lot of power, especially if they operate in sunlight.
[0067] The inventors have recognized and appreciated that an XR system can account for changes in sensor position or orientation during use or over time by repeatedly performing a calibration routine. The calibration routine can determine the current relative separation and orientation of sensors included in a head-mounted device. The wearable XR system can then take into account the current relative separation and orientation of the head-mounted device sensors when calculating stereo depth information. With such a calibration capability, the XR system can accurately obtain depth information to indicate the distance to objects in the physical world without the need for active depth sensors or only occasionally using active depth sensing. Because active depth sensing can consume a lot of power, reducing or eliminating active depth sensing can cause the device to consume less power, which can increase the operating time of the device without charging the battery, or reduce the size of the device due to reduced battery size.
[0068] The inventors also recognize and understand that, through the appropriate combination of image sensors, and the appropriate technology for processing image information from these sensors, the XR system can obtain information about physical objects with low latency and even with reduced power consumption by: reducing the number of sensors used; eliminating, disabling or selectively activating resource-intensive sensors; and / or reducing the overall use of sensors. As a specific example, the XR system may include a head-mounted device with two world-facing cameras. The first camera in the camera can generate grayscale images and can have a global shutter. These grayscale images can be smaller than color images of similar resolution, in some cases, represented by less than one-third the number of bits. Such a grayscale camera can require less power than a color camera of similar resolution. The grayscale camera can be configured to use event-based data acquisition and / or block tracking to limit the amount of data output. The second camera in the camera can be an RGB camera. The wearable cross-reality display system can be configured to selectively use the camera to reduce power consumption and extend battery life without affecting the user's XR experience. For example, an RGB camera can be used with grayscale to build a world model of the user's surroundings, but then the grayscale camera can be used to track the user's head posture based on the world model. Alternatively or additionally, a grayscale camera may be used primarily to track objects, including the user's hands. Based on detected conditions indicating poor tracking quality, image information from one or more other sensors may be employed. For example, an RGB camera may acquire color information that helps distinguish an object from the background, or output from a light field sensor that passively provides depth information may be used.
[0069] The techniques described herein may be used in various types of scenarios, either together with various types of devices or alone. Figure 1 Such a scene is shown. Figure 2 , 3A3B illustrate an exemplary AR system including one or more processors, memory, sensors, and user interfaces that may operate according to the techniques described herein.
[0070] refer to Figure 1 , shows an AR scene 4 in which the user of the AR system sees a park-like setting 6 of the physical world, which features people, trees, buildings in the background, and a concrete platform 8. In addition to these physical objects, the user of the AR technology also perceives that they "see" virtual objects, here illustrated as a robot statue 10 standing on the physical world concrete platform 8, and a flying cartoon-like avatar character 2 that appears to be the head of a bumblebee, even though these elements (e.g., avatar character 2 and robot statue 10) do not exist in the physical world. Due to the extreme complexity of human visual perception and nervous system, it is challenging to produce an AR system that promotes comfortable, natural feeling, rich presentation of virtual image elements among other virtual or physical world image elements.
[0071] Such a scene can be presented to the user by presenting image information representing the actual environment around the user and overlaying information representing virtual objects that are not in the actual environment. In an AR system, the user may be able to see objects in the physical world, and the AR system provides information for rendering virtual objects so that they appear in the appropriate location and have appropriate visual features for the virtual objects to appear to coexist with objects in the physical world. For example, in an AR system, the user can look through a transparent screen so that the user can see objects in the physical world. The AR system can render virtual objects on the screen so that the user sees both the physical world and the virtual objects at the same time. In some embodiments, the screen can be worn by the user, such as a pair of goggles or glasses.
[0072] The scene can be presented to the user via a system including multiple components, including a user interface that can stimulate one or more user senses (including vision, sound and / or touch). In addition, the system can include one or more sensors that can measure parameters of the physical part of the scene, including the position and / or movement of the user within the physical part of the scene. In addition, the system can include one or more computing devices, and associated computer hardware, such as memory. These components can be integrated into a single device, or more distributed across multiple interconnected devices. In some embodiments, some or all of these components can be integrated into a wearable device.
[0073] In some embodiments, an AR experience may be provided to a user via a wearable display system. Figure 2An example of a wearable display system 80 (hereinafter "system 80") is shown. System 80 includes a head mounted display device 62 (hereinafter "display device 62"), and various mechanical and electronic modules and systems that support the functionality of display device 62. Display device 62 can be coupled to a frame 64 that can be worn by a display system user or viewer 60 (hereinafter "user 60") and is configured to position display device 62 in front of the eyes of user 60. According to various embodiments, display device 62 can be displayed sequentially. Display device 62 can be monocular or binocular.
[0074] In some embodiments, speaker 66 is coupled to frame 64 and positioned near the ear canal of user 60. In some embodiments, another speaker, not shown, is positioned near the other ear canal of user 60 to provide stereo / plastic sound control.
[0075] The system 80 may include a local data processing module 70. The local data processing module 70 may be operably coupled to the display device 62 via a communication link 68 (such as via a wired conductor or a wireless connection). The local data processing module 70 may be mounted in a variety of configurations, such as fixedly attached to the frame 64, fixedly attached to a helmet or hat worn by the user 60, embedded in headphones, or otherwise removably attached to the user 60 (e.g., in a backpack configuration, in a belt-coupled configuration). In some embodiments, the local data processing module 70 may not be present, as components of the local data processing module 70 may be integrated into the display device 62 or implemented in a remote server or other component to which the display device 62 is coupled, such as by wireless communication over a wide area network.
[0076] The local data processing module 70 may include a processor and digital storage such as non-volatile memory (e.g., flash memory), both of which may be used to assist in the processing, caching, and storage of data. The data may include: a) data captured from a sensor (e.g., which may be operably coupled to the frame 64) or otherwise attached to the user 60, such as an image capture device (such as a camera), a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a radio device, and / or a gyroscope; and / or b) data acquired and / or processed using the remote processing module 72 and / or the remote data repository 74, which may be passed to the display device 62 after such processing or acquisition. The local data processing module 70 may be operably coupled to the remote processing module 72 and the remote data repository 74 by communication links 76, 78 (such as via a wired or wireless communication link), respectively, so that these remote modules 72, 74 are operably coupled to each other and can be used as resources for the local processing and data module 70.
[0077] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a central processing unit and / or one or more graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, the remote data repository 74 may include a digital data storage facility that may be available via the Internet or other networking configuration in a "cloud" resource configuration. In some embodiments, all data is stored and all computations are performed in the local data processing module 70, allowing for fully autonomous use from a remote module.
[0078] In some embodiments, the local data processing module 70 is operably coupled to a battery 82. In some embodiments, the battery 82 is a removable power source, such as above a counter battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 includes both an internal lithium-ion battery that can be charged by the user 60 during the non-operating time of the system 80, and a removable battery so that the user 60 can operate the system 80 for a longer period of time without having to connect to a power source to charge the lithium-ion battery, or without having to shut down the system 80 to replace the battery.
[0079] Figure 3A A user 30 wearing an AR display system that renders AR content is shown as the user 30 moves through a physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 positions the AR display system at a location 34, and the AR display system records environmental information of the navigable world relative to the location 34 (e.g., digital representations of real objects in the physical world, which can be stored and updated as changes are made to the real objects in the physical world). Each location 34 can also be associated with a "pose" and / or mapped features or directional audio input related to the environment 32. A user wearing the AR display system on the user's head may look in a particular direction and tilt his head, thereby creating a head pose of the system relative to the environment. At each location and / or pose within the same location, the sensors on the AR display system can capture different information about the environment 32. Therefore, the information collected at the location 34 can be aggregated into a data input 36 and processed by at least a navigable world module 38, which can, for example, be processed by Figure 2 This is achieved by processing on the remote processing module 72.
[0080] The traversable world module 38 determines where and how the AR content 40 can be placed relative to the physical world, at least in part based on the data input 36. The AR content is "placed" in the physical world by presenting the AR content in a manner that the user can see the AR content and the physical world at the same time. For example, such an interface can be created using glasses through which the user can see the physical world, and the glasses can be controlled so that virtual objects appear at controlled locations within the user's field of view. The AR content is rendered as if interacting with objects in the physical world. The user interface allows the user's view of objects in the physical world to be blocked to create an appearance location where the AR content blocks the user's view of these objects when appropriate. For example, the AR content can be placed by appropriately selecting a portion of an element 42 (e.g., a table) in the environment 32 to be displayed, and displaying the shape and position of the AR content 40 as if it rests on the element 42 or otherwise interacts with the element 42. The AR content can also be placed within a structure that is not yet within the field of view 44 or within a map grid model 46 relative to the physical world.
[0081] As shown, element 42 is an example of multiple elements in the physical world that can be considered fixed and stored in the traversable world module 38. Once stored in the traversable world module 38, information about these fixed elements can be used to present information to the user so that the user 30 can perceive the content on the fixed element 42 without the system having to map to the fixed element 42 every time the user 30 sees the fixed element 42. Therefore, the fixed element 42 can be a mapped mesh model from a previous modeling session, or can be determined by a separate user but still stored on the traversable world module 38 for future reference by multiple users. Therefore, the traversable world module 38 can identify the environment 32 from the previously mapped environment and display AR content without the user 30's device first mapping the environment 32, thereby saving computing processes and cycles and avoiding any latency of rendered AR content.
[0082] Similarly, a mapped mesh model 46 of the physical world can be created by the AR display system, and appropriate surfaces and metrics for interacting with and displaying AR content 40 can be mapped and stored in the traversable world module 38 for future retrieval by the user 30 or other users without the need for remapping or modeling. In some embodiments, the data input 36 is an input such as a geographic location, user identification, and current activity to indicate to the traversable world module 38 which fixed element 42 of one or more fixed elements is available, which AR content 40 was last placed on the fixed element 42, and whether to display the same content (such AR content is "persistent" content regardless of how the user views the particular traversable world model).
[0083] Even in embodiments where objects are considered fixed, the traversable world module 38 may be updated from time to time to account for the possibility of changes in the physical world. Models of fixed objects may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not considered fixed. In order to render a realistic AR scene, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update fixed objects. In order to be able to accurately track all objects in the physical world, the AR system may obtain information from multiple sensors (including one or more image sensors).
[0084] Figure 3B is a schematic diagram of viewing optical assembly 48 and accompanying optional components. Fig.21 Specific configurations are described in . Directed to the user's eye 49, in some embodiments, two eye tracking cameras 50 detect metrics of the user's eye 49, such as eye shape, eyelid occlusion, pupil direction, and flicker on the user's eye 49. In some embodiments, one of the sensors may be a depth sensor 51, such as a time-of-flight sensor, which transmits signals to the world and detects reflections of those signals from nearby objects to determine the distance to a given object. The depth sensor can quickly determine whether an object has entered the user's field of view, for example, due to the movement of those objects or changes in the user's posture. However, information about the position of the object in the user's field of view may be collected alternatively or additionally by other sensors. In some embodiments, the world camera 52 records a view greater than the periphery to map the environment 32 and detect inputs that may affect AR content. In some embodiments, the world camera 52 and / or the camera 53 may be a grayscale and / or color image sensor that may output grayscale and / or color image frames at fixed time intervals. The camera 53 may further capture an image of the physical world within the user's field of view at a specific time. Even if the values of the pixels of the frame-based image sensor do not change, the pixels thereof may be repeatedly sampled. Each of the world camera 52, camera 53 and depth sensor 51 has a corresponding field of view 54, 55 and 56 to obtain a Figure 3A Collect data in the physical world scene of the physical world environment 32 shown in and record the physical world scene.
[0085] Inertial measurement unit 57 can determine the motion and / or orientation of viewing optical assembly 48. In some embodiments, each component is operably coupled to at least one other component. For example, depth sensor 51 can be operably coupled to eye tracking camera 50 to confirm the actual distance of a point and / or area in the physical world that the user's eye 49 is looking at.
[0086] It should be understood that the viewing optics assembly 48 may include Figure 3BSome of the components shown. For example, viewing optical assembly 48 may include a different number of components. In some embodiments, for example, viewing optical assembly 48 may include one world camera 52, two world cameras 52, or more world cameras instead of the four world cameras shown. Alternatively or additionally, cameras 52 and 53 do not need to capture visible light images of their entire fields of view. Viewing optical assembly 48 may include other types of components. In some embodiments, viewing optical assembly 48 may include one or more dynamic vision sensors whose pixels may asynchronously respond to relative changes in light intensity that exceed a threshold.
[0087] In some embodiments, based on the time-of-flight information, the viewing optical assembly 48 may not include the depth sensor 51. For example, in some embodiments, the viewing optical assembly 48 may include one or more plenoptic cameras whose pixels may capture not only the light intensity, but also the angle of the incident light. For example, the plenoptic camera may include an image sensor covered with a transmissive diffraction mask (TDM). Alternatively or in addition, the plenoptic camera may include an image sensor that includes angle-sensitive pixels and / or phase detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such a sensor may be used as a source of depth information instead of or in addition to the depth sensor 51.
[0088] It should also be understood that Figure 3B The configuration of the components in is shown as an example. The viewing optical assembly 48 may include components having any suitable configuration so that a user can have a maximum field of view for a particular set of components. For example, if the viewing optical assembly 48 has a world camera 52, the world camera may be placed in the central area of the viewing optical assembly rather than on the side.
[0089] The information from these sensors in the viewing optical assembly 48 can be coupled to one or more processors in the system. The processor can generate data that can be rendered so that the user perceives virtual content that interacts with objects in the physical world. The rendering can be implemented in any suitable manner, including generating image data showing both physical and virtual objects. In other embodiments, physical and virtual content can be shown in a scene by modulating the opacity of the display device that the user browses in the physical world. The opacity can be controlled to create the location where the virtual object appears, and also prevent the user from seeing the object blocked by the virtual object in the physical world. In some embodiments, the image data can only include virtual content that can be modified to interact with the physical world (for example, clip content to solve occlusion), which can be viewed through the user interface. No matter how the content is presented to the user, a model of the physical world can be used so that the characteristics of the virtual object that can be affected by the physical object can be correctly calculated, including the shape, position, motion and visibility of the virtual object.
[0090] A model of the physical world can be created from data collected from sensors on a user's wearable device. In some embodiments, a model can be created from data collected from multiple users, which can be aggregated in a computing device remote from all users (and the data can be in the "cloud").
[0091] In some embodiments, at least one of the sensors can be configured to acquire information about physical objects in the scene, particularly non-stationary objects, at high frequency and low latency using compact and low-power components. The sensor can employ block tracking to limit the amount of data output.
[0092] Figure 4 An image sensing system 400 is shown in accordance with some embodiments. Image sensing system 400 may include image sensor 402, which may include image array 404, which may include a plurality of pixels, each pixel responding to light as in a conventional image sensor. Sensor 402 may also include circuitry to access each pixel. Accessing a pixel may require obtaining information about incident light generated by the pixel. Alternatively or additionally, accessing a pixel may require controlling the pixel, such as by configuring it to provide an output only when a certain event is detected.
[0093] In the illustrated embodiment, the image array 404 is configured as an array having multiple rows and columns of pixels. In such an embodiment, the access circuit can be implemented as a row address encoder / decoder 406 and a column address encoder / decoder 408. The image sensor 402 may also include a circuit that generates an input to the access circuit to control the timing and order of reading information from the pixels in the image array 404. In the illustrated embodiment, the circuit is a block tracking engine 410. Compared to conventional image sensors that can continuously output image information captured by pixels in each row, the image sensor 402 can be controlled to output image information in a specified block. In addition, the position of these blocks relative to the image array can change over time. In the illustrated embodiment, the block tracking engine 410 can output image array access information to control the output of image information from the portion of the image array 404 corresponding to the position of the block, and the access information can be dynamically changed based on an estimate of the movement of objects in the environment and / or an estimate of the movement of the image sensor relative to these objects.
[0094] In some embodiments, the image sensor 402 may have the functionality of a dynamic vision sensor (DVS), so that image information is provided by the sensor only when the image attribute (e.g., intensity) of the pixel changes. For example, the image sensor 402 may apply one or more thresholds that define the on (ON) and off (OFF) states of the pixel. The image sensor may detect that the pixel has changed state, and selectively provide outputs only for those pixels or those pixels in the block that have changed state. These outputs may be asynchronous when detected, rather than as part of all pixel readouts in the array. For example, the output may be in the form of an address event representation (AER) 418, which may include a pixel address (e.g., row and column) and an event type (ON or OFF). An ON event may indicate that a pixel unit at the corresponding pixel address senses an increase in light intensity; and an OFF event may indicate that a pixel unit at the corresponding pixel address senses a decrease in light intensity. The increase or decrease may be relative to an absolute level, or may be a change in the level relative to the last output of the pixel. For example, the change may be expressed as a fixed offset or a percentage of the value of the last output of the pixel.
[0095] Using DVS technology in conjunction with tile tracking can make image sensors suitable for XR systems. When combined in an image sensor, the amount of data generated may be limited to the data of the pixel cells within the tile and detecting a change that will trigger an event output.
[0096] In some scenarios, high-resolution image information is required. However, when using DVS technology, a large sensor with more than one million pixel units for generating high-resolution image information may generate a large amount of image information. The inventors have recognized and understood that a DVS sensor may generate a large number of events reflecting motion or image changes in the background, rather than caused by the motion of the tracked object. Currently, the resolution of DVS sensors is limited to less than 1MB, such as 128x128, 240x180, 346x260, to limit the number of events generated. Such sensors sacrifice resolution for tracking objects and may not be able to detect, for example, fine finger movements of the hand. In addition, if the image sensor outputs image information in other formats, limiting the resolution of the sensor array to output a manageable number of events may also limit the image sensor to be used with the DVS function to generate high-resolution image frames. In some embodiments, the sensor as described herein may have a resolution higher than VGA, including up to 8 million pixels or 12 million pixels. Nevertheless, the block tracking as described herein can be used to limit the number of events output by the image sensor per second. Therefore, an image sensor operating in at least two modes can be enabled. For example, an image sensor with a megapixel resolution may be operated in a first mode in which it outputs events in a particular tile being tracked. In a second mode, it may output a high-resolution image frame or portion of an image frame. Such an image sensor may be controlled in an XR system to operate in these different modes based on the functionality of the system.
[0097] Image array 404 may include a plurality of pixel cells 500 arranged in an array. Figure 5A An example of a pixel unit 500 is shown, which in this embodiment is configured for use in an imaging array implementing DVS technology. The pixel unit 500 may include a photoreceptor circuit 502, a differential circuit 506, and a comparator 508. The photoreceptor circuit 502 may include a photodiode 504, which converts light that illuminates the photodiode into a measurable electrical signal. In this example, the conversion is to a current I. A transconductance amplifier 510 converts the photocurrent I into a voltage. The conversion may be linear or nonlinear, for example according to a function logI. Regardless of the specific transfer function, the output of the transconductance amplifier 510 indicates the amount of light detected at the photodiode 504. Although a photodiode is shown as an example, it should be understood that other photosensitive components that generate a measurable output in response to incident light may be implemented in the photoreceptor circuit in place of or in addition to the photodiode.
[0098] exist Figure 5AIn an embodiment of the invention, the circuit for determining whether the output of a pixel has changed enough to trigger the output of the pixel unit is incorporated into the pixel itself. In this example, this function is implemented by a differential circuit 506 and a comparator 508. The differential circuit 506 can be configured to reduce the DC mismatch between pixel units by, for example, balancing the output of the differential circuit to reset the level after an event is generated. In this example, the differential circuit 506 is configured to generate an output that shows the change in the output of the photodiode 504 since the last output. The differential circuit can include an amplifier 512 with a gain A, a capacitor 514 that can be implemented as a single circuit element or one or more capacitors connected in a network, and a reset switch 516.
[0099] In operation, the pixel cell will be reset by temporarily closing switch 516. Such a reset can occur at the beginning of the operation of the circuit and at any time after the event is detected. When the pixel 500 is reset, the voltage across capacitor 514 is When the voltage across capacitor 514 is subtracted from the output of transconductance amplifier 510, zero voltage is generated at the input of amplifier 512. When switch 516 is opened, the output of transconductance amplifier 510 will be combined with the voltage drop across capacitor 514, and there will be zero voltage at the input of amplifier 512. The output of transconductance amplifier 510 changes due to changes in the amount of light illuminating photodiode 504. When the output of transconductance amplifier 510 increases or decreases, the output of amplifier 512 will swing positively or negatively by a change amount, which is amplified by the gain of amplifier 512.
[0100] The comparator 508 can determine whether an event is generated and the sign of the event by, for example, comparing the output voltage V of the differential circuit with a predetermined threshold voltage C. In some embodiments, the comparator 508 may include two comparators with transistors, and when the output of the amplifier 512 shows a positive change, a pair of comparators can operate and an increased change (ON event) can be detected; when the output of the amplifier 512 shows a negative change, another comparator can operate and a reduced change (OFF event) can be detected. However, it should be understood that the amplifier 512 can have a negative gain. In such an embodiment, an increase in the output of the transconductance amplifier 510 can be detected as a negative voltage change at the output of the amplifier 512. Similarly, it should be understood that the positive voltage and the negative voltage can be relative to ground or any appropriate reference level. In any case, the value of the threshold voltage C can be controlled by the characteristics of the transistor (e.g., transistor size, transistor threshold voltage) and / or the value of the reference voltage that can be applied to the comparator 508.
[0101] Figure 5BAn example of event output (ON, OFF) of the pixel unit 500 over time t is shown. In the example shown, at time t1, the output value of the differential circuit is V1; at time t2, the output value of the differential circuit is V2; and at time t3, the output value of the differential circuit is V3. Between time t1 and time t2, although the photodiode senses an increase in light intensity, the pixel unit does not output any event because the change in V does not exceed the value of the threshold voltage C. At time t2, because V2 is greater than V1 by the value of the threshold voltage C, the pixel unit outputs an ON event. Between time t2 and time t3, although the photodiode senses a decrease in light intensity, the pixel unit does not output any event because the change in V does not exceed the value of the threshold voltage C. At time t3, because V3 is less than V2 by the value of the threshold voltage C, the pixel unit outputs an OFF event.
[0102] Each event may trigger an output at AER 418. The output may include, for example, an indication of whether the event is an ON event or an OFF event and an identification of the pixel, such as the row and column of the pixel. Other information may alternatively or additionally be included in the output. For example, a timestamp may be included, which may be useful if the event is queued for later transmission or processing. As another example, the current level at the output of amplifier 510 may be included. Such information may optionally be included, for example, if further processing is to be performed in addition to detecting movement of an object.
[0103] It should be understood that the frequency of the event output and the sensitivity of the pixel unit can be controlled by the value of the threshold voltage C. For example, the frequency of the event output can be reduced by increasing the value of the threshold voltage C, or the frequency of the event output can be increased by reducing the value of the threshold voltage C. It should also be understood that the threshold voltage C can be different for ON events and OFF events, for example, by setting different reference voltages for the comparator used to detect the ON event and the comparator used to detect the OFF event. It should also be understood that the pixel unit can also output a value indicating the magnitude of the light intensity change instead of the sign signal indicating the event detection, or output a value indicating the magnitude of the light intensity change in addition to the sign signal indicating the event detection.
[0104] Figure 5A and Figure 5B Pixel cell 500 is shown as an example according to some embodiments. Other designs may also be applicable to pixel cells. In some embodiments, a pixel cell may include a photoreceptor circuit and a differential circuit, but share a comparator circuit with one or more other pixel cells. In some embodiments, a pixel cell may include a circuit configured to calculate a change value, for example, an active pixel sensor at the pixel level.
[0105] Regardless of how events are detected for each pixel unit, the ability to configure pixels to output only when events are detected can be used to limit the amount of information required to maintain a model of the position of a non-fixed (i.e., movable) object. For example, a threshold voltage C that is triggered when a relatively small change occurs can be used to set the pixels within the block. Other pixels outside the block may have a larger threshold, such as a threshold of three or five times. In some embodiments, the threshold voltage C of the pixels outside any block can be set large enough so that the pixel is effectively disabled and does not produce any output, regardless of the amount of change. In other embodiments, pixels outside the block can be disabled in other ways. In such an embodiment, the threshold voltage can be fixed for all pixels, but the pixel can be selectively enabled or disabled based on whether the pixel is within the block.
[0106] In other embodiments, the threshold voltage of one or more pixels can be adaptively set to modulate the amount of data output from the image array. For example, the AR system may have the processing capability to process multiple events per second. When the number of events output per second exceeds an upper limit, the threshold of some or all pixels can be increased. Alternatively or additionally, when the number of events per second drops below a lower limit, the threshold can be lowered so that more data can be used for more accurate processing. As a specific example, the number of events per second may be between 200 and 2000 events. Compared to, for example, processing all pixel values scanned from an image sensor (constituting 30 million or more pixel values per second), this number of events constitutes a significant reduction in the number of data blocks to be processed per second. Compared to processing only pixels within a block (which may be less but still may be tens of thousands or more pixel values per second), this number of events is even reduced.
[0107] The control signals used to enable and / or set the threshold voltage for each of the plurality of pixels may be generated in any suitable manner. However, in the illustrated embodiment, those control signals are set by the tile tracking engine 410 or based on processing within the processing module 72 or other processor.
[0108] Return to reference Figure 4, the image sensing system 400 may receive input from any suitable component such that the tile tracking engine 410 may dynamically select at least one region of the image array 404 to be enabled and / or disabled in order to implement the tile based at least on the received input. The tile tracking engine 410 may be a digital processing circuit having a memory storing one or more parameters for the tile. The parameters may be, for example, the boundaries of the tile, and may include other information, such as information related to a scaling factor between movement of the image array and movement within the image array of an image of a movable object associated with the tile. The tile tracking engine 410 may also include circuitry configured to perform calculations on stored values and other measurements provided as input.
[0109] In the illustrated embodiment, the block tracking engine 410 receives as input a designation of a current block. The block may be designated based on its size and location within the image array 404, for example by specifying a range of row and column addresses for the block. This designation may be used as a parameter in the processing module 72 ( Figure 2 ) or the output of other components that process information about the physical world. For example, the processing module 72 can specify a tile to contain the current position of each movable object in the physical world or the current position of a subset of movable objects being tracked, so as to render virtual objects with appropriate appearance positions relative to the physical world. For example, if the AR scene will include a toy doll balanced on a physical object (such as a moving toy car) as a virtual object, a tile containing the toy car can be specified. A tile may not be specified for another toy car moving in the background because the latest information about the object may not be needed to render a realistic AR scene.
[0110] Regardless of how the tiles are selected, information about the current position of the tiles can be provided to the tile tracking engine 410. In some embodiments, the tiles can be rectangular, so that the location of the tiles can be simply specified as a starting row and column and an ending row and column. In other embodiments, the tiles can have other shapes, such as circular, and the tiles can be specified in other ways, such as by a center point and a radius.
[0111] In some embodiments, trajectory information about the tiles may also be provided. For example, the trajectory may specify the movement of the tiles relative to the coordinates of the image array 404. For example, the processing module 72 may construct a model of the movement of movable objects in the physical world and / or the movement of the image array 404 relative to the physical world. Since the movement of either or both of the above may affect the position within the image array 404 to which the image of the object is projected, the trajectory of the tiles within the image array 404 may be calculated based on either or both. The trajectory may be specified in any suitable manner, such as parameters of a linear, quadratic, cubic, or other polynomial equation.
[0112] In other embodiments, the block tracking engine 410 can dynamically calculate the position of the block based on input from sensors that provide information about the physical world. The information from the sensor can be provided directly from the sensor. Alternatively or additionally, the sensor information can be processed to extract information about the physical world before being provided to the block tracking engine 410. For example, the extracted information can include the movement of the image array 404 relative to the physical world, the distance between the image array 404 and the object whose image falls within the block, and other information that can be used to dynamically align the block in the image array 404 with the image of the object in the physical world when the image array 404 and / or the object moves.
[0113] Examples of input components may include image sensors 412 and inertial sensors 414. Examples of image sensors 412 may include eye tracking cameras 50, depth sensors 51, world cameras 52, and / or cameras 52. Examples of inertial sensors 414 may include inertial measurement units 57. In some embodiments, input components may be selected to provide data at a relatively high rate. For example, the inertial measurement unit 57 may have an output rate between 200 and 2000 measurements per second, such as between 800 and 1200 measurements per second. The position of the tiles may be updated at a similarly high rate. As a specific example, by using the inertial measurement unit 57 as an input source to the tile tracking engine 410, the position of the tiles may be updated 800 to 1200 times per second. In this way, a movable object may be tracked with high accuracy using relatively small tiles that limit the number of events that need to be processed. This approach may result in very low latency between changes in the relative position of the image sensor and the movable object, and equally low latency in updating virtual object rendering, thereby providing an ideal user experience.
[0114] In some scenarios, the movable objects tracked using tiles can be stationary objects in the physical world. For example, the AR system can identify stationary objects by analyzing multiple images of the physical world taken, and select one or more features of the stationary objects as reference points to determine the movement of the wearable device having an image sensor thereon. Frequent and low-latency updates of the positions of these reference points relative to the sensor array can be used to provide frequent and low-latency calculations of the head pose of the wearable device user. Since the head pose can be used to realistically render virtual objects through the user interface on the wearable device, the frequent and low-latency updates of the head pose improve the user experience of the AR system. Therefore, making the input of the tile tracking engine 410 that controls the tile position only come from sensors with high output rates, such as one or more inertial measurement units, can produce a satisfactory user experience of the AR system.
[0115] However, in some embodiments, other information may be provided to the tile tracking engine 410 to enable it to calculate trajectories and / or apply trajectories to the tiles. This other information may include stored information 416, such as the traversable world module 38 and / or the mapping grid model 46. This information may indicate one or more previous positions of the object relative to the physical world, such that taking into account changes in these previous positions and / or changes in the current position relative to the previous positions may indicate a trajectory of the object in the physical world, which may then be mapped to a trajectory across the tiles of the image array 404. Other information in the physical world model may be used alternatively or additionally. For example, the size and / or distance of a movable object or other information about the position relative to the image array 404 may be used to calculate the position or trajectory across the tiles of the image array 404 associated with the object.
[0116] Regardless of the manner in which the trajectory is determined, the tile tracking engine 410 can use the trajectory to calculate the updated position of the tiles within the image array 404 at a high rate, such as faster than once per second or more than 800 times per second. In some embodiments, the rate may be limited by processing power to less than 2000 times per second.
[0117] It should be understood that the process of tracking changes in movable objects is not sufficient to reconstruct the complete physical world. However, the interval for reconstructing the physical world may be longer than the interval between position updates of movable objects, such as every 30 seconds or every 5 seconds. When there is a reconstruction of the physical world, the positions of the objects to be tracked and the positions of the tiles that will capture information about these objects can be recalculated.
[0118] Figure 4 An embodiment is shown in which the processing circuitry for dynamically generating tiles and controlling the selective output of image information from within the tiles is configured to directly control the image array 404 so that the image information output from the array is limited to the selected information. For example, such circuitry may be integrated into the same semiconductor chip that houses the image array 404, or may be integrated into a separate controller chip for the image array 404. However, it should be understood that the circuitry for generating control signals for the image array 404 may be distributed throughout the XR system. For example, some or all of the functions may be performed by programming in the processing module 72 or other processor within the system.
[0119] Image sensing system 400 may output image information for each of the plurality of pixels. Each pixel of the image information may correspond to one of the pixel cells of image array 404. The output image information from image sensing system 400 may be image information for each of one or more blocks selected by block tracking engine 410 corresponding to at least one region of image array 404. In some embodiments, for example, when each pixel of image array 404 has a different pixel size than the pixel size of image array 404, the image information may be output to the image sensing system 400. Figure 5A When configured as shown in , pixels in the output image information may identify pixels within one or more tiles where image sensor 400 detected a change in light intensity.
[0120] In some embodiments, the output image information from the image sensing system 400 may be image information of pixels outside each of the one or more tiles corresponding to at least one region of the image array selected by the tile tracking engine 410. For example, a deer may be running in a physical world with a flowing river. The details of the river's waves may not be of interest but may trigger pixel cells of the image array 402. The tile tracking engine 410 may create a tile surrounding the river and disable a portion of the image array 402 corresponding to the tile surrounding the river.
[0121] Based on the identification of changed pixels, further processing can be performed. For example, the portion of the world model corresponding to the portion of the physical world imaged by the changed pixels can be updated. These updates can be performed based on information collected using other sensors. In some embodiments, further processing can be conditioned on or triggered by a number of changed pixels in a tile. For example, once it is detected that 10% or some other threshold amount of pixels in a tile have changed, an update can be performed.
[0122] In some embodiments, image information in other formats may be output from the image sensor and may be used in conjunction with the change information to update the world model. In some embodiments, the format of the image information output from the image sensor may change at any time during operation of the VR system. For example, in some embodiments, the pixel unit 500 may be operated to produce a differential output at certain times, such as in the comparator 508. The output of the amplifier 510 may be switchable so as to output the magnitude of light incident on the photodiode 504 at other times. For example, the output of the amplifier 510 may be switchably connected to a sense line, which in turn is connected to an A / D converter, which may provide a digital indication of the magnitude of the incident light based on the magnitude of the output of the amplifier 510.
[0123] An image sensor in this configuration can be operated as part of an AR system to perform differential output most of the time, output events only for pixels where a change exceeding a threshold is detected, or output events only for pixels within a block where a change exceeding a threshold is detected. A full image frame with amplitude information for all pixels in the image array can be output periodically (e.g., every 5 to 30 seconds). In this way, low latency and accurate processing can be achieved, with differential information used to quickly update selected parts of the world model that are most likely to affect user perception of the change, while the full image can be used to update larger parts of the world model more. Although full updates to the world model occur only at a slower rate, any delay in updating the model may not have a meaningful impact on the user's perception of the AR scene.
[0124] The output mode of the image sensor can be changed at any time throughout the operation of the image sensor, so that the sensor outputs one or more intensity information for some or all pixels, as well as an indication of changes in some or all pixels in the array.
[0125] It is not a requirement that image information from the tiles be selectively output from the image sensor by limiting the information output from the image array. In some embodiments, image information may be output by all pixels in the image array, and only information about a particular area of the array may be output from the image sensor. Figure 6 An image sensor 600 is shown according to some embodiments. The image sensor 600 may include an image array 602. In this embodiment, the image array 602 may be similar to a conventional image array that scans out rows and columns of pixel values. The operation of such an image array may be coordinated by other components. The image sensor 600 may also include a block tracking engine 604 and / or a comparator 606. The image sensor 600 may provide an output 610 to an image processor 608. For example, the processor 608 may be a processing module 72 ( Figure 2 ) part.
[0126] The block tracking engine 604 may have a structure and function similar to the block tracking engine 410. It may be configured to receive a signal specifying at least one selected area of the image array 602, and then generate a control signal specifying the dynamic position of the area based on a calculated trajectory within the image array 602 of the image of the object represented by the area. In some embodiments, the block tracking engine 604 may receive a signal specifying at least one selected area of the image array 602, which may include trajectory information for one or more areas. The block tracking engine 604 may be configured to perform calculations to dynamically identify pixel cells within at least one selected area based on the trajectory information. Changes in the implementation of the block tracking engine 604 are possible. For example, the block tracking engine may update the position of the block based on a sensor indicating movement of the image array 602 and / or projected movement of an object associated with the block.
[0127] exist Figure 6 In the illustrated embodiment, the image sensor 600 is configured to output differential information of pixels within the identified blocks. The comparator 606 may be configured to receive a control signal from the block tracking engine 604 to identify the pixels within the blocks. The comparator 606 may selectively operate on pixels output from the image array 602, which have addresses within the blocks as indicated by the block tracking engine 604. The comparator 606 may operate on the pixel cells to generate a signal indicating a change in the sensed light detected by at least one area of the image array 602. As an example of an implementation, the comparator 606 may include a memory element storing a reset value of the pixel cells within the array. When the current values of these pixels are scanned out from the image array 602, the circuit within the comparator 606 may compare the stored value with the current value and output an indication when the difference exceeds a threshold. For example, a digital circuit may be used to store values and perform such a comparison. In this example, the output of the image sensor 600 may be processed like the output of the image sensor 400.
[0128] In some embodiments, the image array 602, the block tracking engine 604, and the comparator 606 can be implemented in a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, the image array 602 can be implemented in a single integrated circuit. The block tracking engine 604 and the comparator 606 can be implemented in a second single integrated circuit, which is configured as a driver for the image array 602, for example. Alternatively or additionally, some or all of the functionality of the block tracking engine and / or the comparator 606 can be distributed to other digital processors within the AR system.
[0129] Other configurations or processing circuits are possible. Figure 7An image sensor 700 according to some embodiments is shown. The image sensor 700 may include an image array 702. In this embodiment, the image array 702 may have pixel cells with a differential configuration, e.g. Figure 5A 500 in FIG. However, the embodiments herein are not limited to differential pixel units, since block tracking can be implemented using an image sensor that outputs intensity information.
[0130] exist Figure 7 In the illustrated embodiment, a block tracking engine 704 generates control signals indicating addresses of pixel locations within one or more blocks being tracked. The block tracking engine 704 can be constructed and operated like the block tracking engine 604. Here, the block tracking engine 704 provides control signals to a pixel filter 706, which passes image information from only those pixels within the block to an output 710. As shown, the output 710 is coupled to an image processor 708, which can further process the image information of the pixels within the block using techniques as described herein or in other suitable manners.
[0131] Figure 8 , which shows an image sensor 800 according to some embodiments. Image sensor 800 may include an image array 802, which may be a conventional image array that scans out pixel intensity values. The image array may be used to provide differential image information as described herein by using a comparator 806. Similar to comparator 606, comparator 806 may calculate differential information based on stored pixel values. Pixel filter 808 may pass selected values from those differential values to output 812. As with pixel filter 706, pixel filter 808 may receive control inputs from a block tracking engine 804. Block tracking engine 804 may be similar to block tracking engine 704. Output 812 may be coupled to an image processor 810. Some or all of the above components of image sensor 800 may be implemented in a single integrated circuit. Alternatively, the components may be distributed on one or more integrated circuits or other components.
[0132] Image sensors as described herein may operate as part of an augmented reality system to maintain information about movable objects or other information about the physical world that is useful in realistically rendering images of virtual objects in conjunction with information about the physical environment. Fig. 9 A method 900 for image sensing is shown in accordance with some embodiments.
[0133] At least a portion of method 900 may be performed to operate an image sensor including, for example, image sensor 400, 600, 700, or 800. Method 900 may begin by receiving (act 902) imaging information from one or more inputs including, for example, image sensor 412, inertial sensor 414, and stored information 416. Method 900 may include identifying (act 904) one or more patches on an image output of the image sensing system based at least in part on the received information. An example of act 904 is shown in FIG. Fig.10 In some embodiments, method 900 may include calculating (act 906) movement trajectories of one or more blocks. An example of act 906 is shown in Fig.11 Shown in.
[0134] The method 900 may also include setting (action 908) the image sensing system based at least in part on the identified one or more blocks and / or their estimated movement trajectories. The setting may be achieved by enabling a portion of the pixel units of the image sensing system through, for example, a comparator 606, a pixel filter 706, etc., based at least in part on the identified one or more blocks and / or their estimated movement trajectories. In some embodiments, the comparator 606 may receive a first reference voltage value for a pixel unit corresponding to a selected block on the image, and a second reference voltage value for a pixel unit that does not correspond to any selected block on the image. The comparator 606 may set the second reference voltage to be much higher than the first reference voltage, so that an unreasonable light intensity change sensed by a pixel unit of a comparator unit having the second reference voltage may result in an output of the pixel unit. In some embodiments, the pixel filter 706 may disable the output of a pixel unit having an address (e.g., row and column) that does not correspond to any selected block on the image.
[0135] Fig.10 Block identification 904 is shown in accordance with some embodiments. Block identification 904 may include segmenting (act 1002) one or more images from one or more inputs based at least in part on color, light intensity, angle of arrival, depth, and semantics.
[0136] Block recognition 904 may also include recognizing (action 1004) one or more objects in one or more images. In some embodiments, object recognition 1004 may be based at least in part on predetermined features of an object, including, for example, hands, eyes, facial features. In some embodiments, object recognition 1004 may be based on one or more virtual objects. For example, a virtual animal character is walking on a physical pencil. Object recognition 1004 may use a virtual animal character target as an object. In some embodiments, object recognition 1004 may be based at least in part on artificial intelligence (AI) training received by an image sensing system. For example, an image sensing system may be trained by reading images of cats of different types and colors, thereby learning the features of a cat and being able to recognize a cat in the physical world.
[0137] Segment identification 904 may include generating (act 1006) segments based on one or more objects. In some embodiments, object segmentation 1006 may generate segments by computing a convex hull or bounding box of one or more objects.
[0138] Fig.11 Block trajectory estimation 906 is shown according to some embodiments. Block trajectory estimation 906 may include predicting (action 1102) the movement of one or more blocks over time. The movement of one or more blocks may be caused by a variety of reasons, including, for example, moving objects and / or moving users. Movement prediction 1102 may include deriving the movement speed of the moving objects and / or moving users based on received images and / or received AI training.
[0139] The block trajectory estimation 906 may include calculating (action 1104) a trajectory of one or more blocks over time based at least in part on the predicted motion. In some embodiments, the trajectory may be calculated by modeling with a first order linear equation, assuming that the moving object will continue to move in the same direction at the same speed. In some embodiments, the trajectory may be calculated by curve fitting or using heuristics, including pattern detection.
[0140] Fig.12 and Fig.13 Shows factors that can be applied in the calculation of the block trajectory. Fig.12 An example of a movable object is shown, in which the movable object is a moving object 1202 (e.g., a hand) that is moving relative to a user of the AR system. In this example, the user wears an image sensor as part of a head mounted display 62. In this example, the user's eye 49 is looking straight ahead so that the image array 1200 captures the field of view (FOV) of the eye 49 relative to one viewpoint 1204. The object 1202 is in the FOV and therefore appears in the corresponding pixel in the array 1200 by creating intensity variations.
[0141] Array 1200 has a plurality of pixels 1208 arranged in an array. For the system to track hand 1202, a block 1206 in the array containing object 1202 at time t0 may include a portion of the plurality of pixels. If object 1202 is moving, the position of the block capturing the object will change over time. This change may be captured in a block trajectory, from block 1206 to blocks X and Y used later.
[0142] The tile trajectory can be estimated, for example, in action 906, by identifying a feature 1210 of an object in the tile, such as a fingertip in the illustrated example. A motion vector 1212 can be calculated for the feature. In this example, the trajectory is modeled as a first order linear equation, and the prediction is based on the assumption that the object 1202 will continue on the same tile trajectory 1214 over time, producing a tile position X and Y at each of two consecutive times.
[0143] As the position of the block changes, the image of the moving object 1202 remains within the block. Even if the image information is limited to information collected using pixels within the block, the image information is sufficient to represent the movement of the moving object 1202. This is true whether the image information is intensity information or differential information such as that produced by a differential circuit. For example, in the case of a differential circuit, an event indicating an increase in intensity may occur when the image of the moving object 1202 moves over a pixel. Conversely, an event indicating a decrease in intensity may occur when the image of the moving object 1202 passes over a pixel. A pixel pattern with increasing and decreasing events can be used as a reliable indication of the movement of the moving object 1202, and since the amount of data indicating the event is relatively small, it can be updated quickly with low latency. As a specific example, such a system can result in a realistic XR system that tracks a user's hands and changes the rendering of virtual objects, thereby creating a sense for the user that the user is interacting with the virtual object.
[0144] The positions of the tiles may change for other reasons, and any or all of these reasons may be reflected in the trajectory calculation. One such other change is the movement of the user while the user is wearing the image sensor. Fig.13 An example of a moving user is shown, which creates a changing viewpoint for the user as well as the image sensor. Fig.13 13, the user may initially be looking straight ahead at an object at viewpoint 1302. In this configuration, pixel array 1300 of the image array will capture the object in front of the user. The object in front of the user may be in block 1312.
[0145] The user may then change the viewpoint, such as by turning their head. The viewpoint may change to viewpoint 1304. Even though the object that was previously directly in front of the user has not moved, it will have a different location within the field of view at the user's viewpoint 1304. It will also be at a different point within the field of view of the image sensor worn by the user, and therefore at a different location within image array 1300. For example, the object may be contained within the tile at location 1314.
[0146] If the user further changes their viewpoint to viewpoint 1306, and the image sensor moves with the user, the location of the object that was previously directly in front of the user will be imaged at a different point within the field of view of the image sensor worn by the user, and therefore at a different location within image array 1300. For example, the object may be contained within the tile at location 1316.
[0147] It can be seen that as the user further changes their viewpoint, the position of the tile in the image array required to capture the object moves further. The trajectory of this movement from position 1312 to position 1314 to position 1316 can be estimated and used to track the future position of the tile.
[0148] The trajectory may be estimated in other ways. For example, measurements from an inertial sensor may indicate the acceleration and velocity of the user's head when the user has viewpoint 1302. This information may be used to predict the trajectory of a tile within the image array based on the movement of the user's head.
[0149] Tile trajectory estimation 906 can predict, based at least in part on these inertial measurements, that the user will have viewpoint 1304 at time t1 and viewpoint 1306 at time t2. Thus, tile trajectory estimation 906 can predict that tile 1308 can move to tile 1310 at time t1 and to tile 1312 at time t2.
[0150] As an example of this approach, it can be used to provide accurate and low latency head pose estimation in an AR system. The tile can be localized to contain images of stationary objects within the user's environment. As a specific example, processing of image information can identify the corner of a picture frame hanging on a wall as a recognizable and stationary object to be tracked. The processing can focus the tile on that object. Combined with the above Fig.12As is the case with the moving object 1202 described herein, the relative motion between the object and the user's head will generate events that can be used to calculate the relative motion between the user and the tracked object. In this example, since the tracked object is stationary, the relative motion indicates the movement of the imaging array that the user is wearing. Therefore, the movement indicates the change in the user's head posture relative to the physical world and can be used to maintain an accurate calculation of the user's head posture, which can be used to realistically render virtual objects. Because the imaging array as described herein can provide fast updates, the amount of data for each update is relatively small, so the calculations for rendering virtual objects remain accurate (they can be performed quickly and updated frequently).
[0151] Return to reference Fig.11 , the block trajectory estimation 906 may include adjusting (action 1106) the size of at least one block based at least in part on the calculated block trajectory. For example, the size of the block may be set to be large enough so that it includes a number of pixels in which an image of a movable object or at least a portion of an object for which image information is to be generated will be projected. The block may be set to be slightly larger than the projected size of the image of the portion of the object of interest so that if there is any error in estimating the trajectory of the block, the block can still include the relevant portion of the image. When an object moves relative to the image sensor, the image size (in pixels) of the object may change based on distance, angle of incidence, orientation of the object, or other factors. A processor that defines blocks associated with an object may set the size of the block, for example by measuring based on other sensor data or calculating the size of the block associated with the object based on a world model. Other parameters of the block, such as its shape, may be similarly set or updated.
[0152] Fig.14 An image sensing system 1400 configured for use in an XR system according to some embodiments is shown. Figure 4 ), the image sensing system 1400 includes circuitry that selectively outputs values within a block and can be configured to output events for pixels within the block, as also described above. In addition, the image sensing system 1400 is configured to selectively output measured intensity values, which can be output for a complete image frame.
[0153] In the illustrated embodiment, separate outputs are shown for events and intensity values generated using the DVS technique as described above. The output generated using the DVS technique can be represented as AER1418 using the representation output described above in conjunction with AER 418. The output representing the intensity value can be output through the output terminal designated here as APS1420. Those intensity outputs can be for blocks or for the entire image frame. The AER and APS outputs can be activated simultaneously. However, in the illustrated embodiment, the image sensor 1400 operates in a mode of outputting events or a mode of outputting intensity information at any given time. A system using such an image sensor can selectively use the event output and / or the intensity information.
[0154] Image sensing system 1400 may include image sensor 1402, which may include image array 1404, which may include multiple pixels 1500, each pixel responsive to light. Sensor 1402 may also include circuitry to access pixel cells. Sensor 1402 may also include circuitry to generate inputs to the access circuitry to control a mode of reading out information from pixel cells in image array 1404.
[0155] In the illustrated embodiment, image array 1404 is configured as an array having multiple rows and columns of pixel cells that are accessible in both readout modes. In such an embodiment, access circuitry may include a row address encoder / decoder 1406, a column address encoder / decoder 1408 that controls a column select switch 1422, and / or a register 1424 that may temporarily hold information about incident light sensed by one or more corresponding pixel cells. A tile tracking engine 1410 may generate inputs to the access circuitry to control which pixel cells provide image information at any time.
[0156] In some embodiments, image sensor 1402 may be configured to operate in a rolling shutter mode, a global shutter mode, or both.For example, tile tracking engine 1410 may generate inputs to access circuits to control a readout mode of image array 1402 .
[0157] When the sensor 1402 is operated in a rolling shutter readout mode, a single column of pixel cells is selected during each system clock by, for example, closing a single column switch 1422 of the plurality of column switches. During the system clock, the selected column of pixel cells is exposed and read out to the APS 1420. To generate an image frame using the rolling shutter mode, the columns of pixel cells in the sensor 1402 may be read out column by column and then processed by an image processor to generate an image frame.
[0158] When the sensor 1402 is operated in a global shutter mode, for example, in a single system clock, columns of pixel cells are exposed simultaneously and the information is stored in registers 1424, so that information captured by pixel cells in multiple columns can be read out to APS 1420b simultaneously. This readout mode allows image frames to be output directly without further data processing. In the example shown, information about the incident light sensed by the pixel cells is stored in corresponding registers 1424. It should be understood that multiple pixel cells can share one register 1424.
[0159] In some embodiments, the sensor 1402 may be implemented in a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, the image array 1404 may be implemented in a single integrated circuit. The block tracking engine 1410, the row address encoder / decoder 1406, the column address encoder / decoder 1408, the column select switch 1422, and / or the register 1424 may be implemented in a second single integrated circuit, such as a driver configured for the image array 1404. Alternatively or additionally, some or all of the functions of the block tracking engine 1410, the row address encoder / decoder 1406, the column address encoder / decoder 1408, the column select switch 1422, and / or the register 1424 may be distributed to other digital processors in the AR system.
[0160] Fig.15 An exemplary pixel cell 1500 is shown. In the embodiment shown, each pixel cell can be configured to output either event or intensity information. However, it should be understood that in some embodiments, the image sensor can be configured to output both types of information simultaneously.
[0161] Both the event information and the intensity information are based on the output of the photodetector 504, as described above in conjunction with Figure 5A The pixel unit 1500 includes a circuit for generating event information. The circuit includes a photoreceptor circuit 502, a differential circuit 506, and a comparator 508, also as described above. When in a first state, a switch 1520 connects the photodetector 504 to the event generating circuit. The switch 1520 or other control circuit can be controlled by a processor controlling the AR system, thereby providing a relatively small amount of image information over a relatively long period of time when the AR system is running.
[0162] The switch 1520 or other control circuitry may also be controlled to configure the pixel cell 1500 to output intensity information. In the illustrated information, the intensity information is provided as a complete image frame, continuously represented as a stream of pixel intensity values for each pixel in the image array. To operate in this mode, the switch 1520 in each pixel cell may be set to a second position that exposes the output of the photodetector 504 after passing through the amplifier 510 so that it can be connected to an output line.
[0163] In the illustrated embodiment, the output lines are illustrated as column lines 1510. There may be one such column line for each column in the image array. Each pixel cell in the column may be coupled to a column line 1510, but the pixel array may be controlled so that one pixel cell is coupled to a column line 1510 at a time. Switch 1530, which may have one such switch in each pixel cell, controls when the pixel cell 1500 is connected to its corresponding column line 1510. Access circuitry such as row address decoder 410 may close switch 1530 to ensure that only one pixel cell is connected to each column line at a time. Switches 1520 and 1530 may be implemented using one or more transistors as part of an image array or similar component.
[0164] Fig.15 Additional components that may be included in each pixel cell according to some embodiments are shown. A sample and hold circuit (S / H) 1532 may be connected between the photodetector 504 and the column line 1510. When present, the S / H 1532 may enable the image sensor 1402 to operate in a global shutter mode. In global shutter mode, a trigger signal is sent to each pixel cell in the array simultaneously. Within each pixel cell, the S / H 1532 captures a value indicating intensity at the time of the trigger signal. The S / H 1532 stores the value and generates an output based on the value until the next value is captured.
[0165] like Fig.15 As shown, when switch 1530 is closed, a signal representing the value stored by S / H 1532 can be coupled to column line 1510. The signal coupled to the column line can be processed to produce the output of the image array. For example, the signal can be buffered and / or amplified in amplifier 1512 at the end of column line 1510 and then applied to analog-to-digital converter (A / D) 1514. The output of A / D 1514 can be passed to output 1420 through other readout circuit 1516. Readout circuit 1516 can include, for example, column switch 1422. Other components within readout circuit 1516 can perform other functions, such as serializing the multi-bit output of A / D 1514.
[0166] Those skilled in the art will understand how to implement circuits to perform the functions described herein. S / H 1532 may be implemented as, for example, one or more capacitors and one or more switches. However, it should be understood that S / H 1532 may use other components or be implemented in a manner different from that described herein. Fig.15 It should be understood that other components may be implemented in addition to those shown. For example, Fig.15 One amplifier and one A / D converter per column is shown. In other embodiments, there may be one A / D converter shared across multiple columns.
[0167] In a pixel array configured for a global shutter, each S / H 1532 can store intensity values reflecting image information at the same time. These values can be stored during the readout phase because the values stored in each pixel are read out continuously. For example, continuous readout can be achieved by connecting the S / H 1532 of each pixel unit in a row to its corresponding column line. The values on the column line can then be passed to the APS output 1420 one at a time. This information flow can be controlled by sequencing the opening and closing of the column switch 1422. For example, the operation can be controlled by the column address decoder 1408. Once the value of each pixel of a row is read out, the pixel units in the next row can be connected to the column line at their location. These values can be read out one column at a time. The process of reading out the value of a row can be repeated until the intensity values of all pixels in the image array are read out. In an embodiment where the intensity values are read out for one or more blocks, the process will be completed when the values of the pixel units within the blocks are read out.
[0168] The pixel cells may be read out in any suitable order. For example, the rows may be interlaced so that every other row is read out in sequence. Nevertheless, the AR system may still process the image data into a frame of image data by de-interlacing the data.
[0169] In embodiments where S / H 1532 is not present, values may still be read sequentially from each pixel cell as rows and columns of values are scanned out. However, the value read from each pixel cell may represent the value in that cell as it was captured as part of the readout process, such as the intensity of light detected at the photodetector of that cell when the value was applied to, for example, A / D 1514. Thus, in a rolling shutter, the pixels of an image frame may represent images incident on the image array at slightly different times. For an image sensor that outputs a full frame at a 30 Hz rate, the time difference between the first pixel value of a captured frame and the last pixel value of the frame may differ by as much as 1 / 30 of a second, which is imperceptible for many applications.
[0170] For some XR functions, such as tracking an object, the XR system may perform calculations on image information collected using an image sensor using a rolling shutter. Such calculations may interpolate between consecutive image frames to calculate an interpolated value for each pixel, which represents an estimated value of the pixel at a point in time between consecutive frames. The same time may be used for all pixels so that, by calculation, the interpolated image frame contains pixels representing the same point in time, such as may be produced using an image sensor with a global shutter. Alternatively, a global shutter image array may be used for one or more image sensors in a wearable device that forms part of the XR system. A global shutter for a full or partial image frame may avoid interpolation of other processing that may be performed to compensate for variations in capture time in image information captured using a rolling shutter. Thus, even if the image information is used to track the movement of an object, such as may occur for processing such as tracking a hand or other movable object, or determining the head pose of a user of a wearable device in AR, or even using a camera on a wearable device to build an accurate representation of the physical environment, the device may move as the image information is collected, and interpolation calculations may be avoided.
[0171] Differentiated pixel unit
[0172] In some embodiments, each pixel cell in the sensor array can be identical. For example, each pixel cell can respond to a broad spectrum of visible light. Thus, each photodetector can provide image information indicating the intensity of visible light. In this scenario, the output of the image array can be a "grayscale" output, representing the amount of visible light incident on the image array.
[0173] In other embodiments, the pixel cells can be differentiated. For example, different pixel cells in the sensor array can output image information indicating the intensity of light in a specific part of the spectrum. A suitable technique for differentiating pixel cells is to place a filter element in the light path leading to the photodetector in the pixel cell. The filter element can be bandpass, for example, allowing visible light of a specific color to pass. Applying such a filter on a pixel cell configures the pixel cell to provide image information indicating the intensity of light of the color corresponding to the filter.
[0174] Filters can be applied to pixel cells regardless of their structure. For example, they can be applied to pixel cells in a sensor array with a global shutter or a rolling shutter. Similarly, filters can be applied to pixel cells that are configured to output intensity or intensity changes using DVS technology.
[0175] In some embodiments, a filter element that selectively passes primary color light can be mounted above a photodetector in each pixel cell in the sensor array. For example, a filter that selectively passes red light, green light, or blue light can be used. The sensor array can have multiple subarrays, each subarray having one or more pixels configured to sense light of each primary color. In this way, the pixel cells in each subarray provide intensity and color information related to the object imaged by the image sensor.
[0176] The inventors have recognized and understood that in an XR system, some functions require color information, while some functions can be performed using grayscale information. A wearable device equipped with an image sensor to provide image information for operation of the XR system may have multiple cameras, some of which may be formed by image sensors that can provide color information. Other cameras may be grayscale cameras. The inventors have recognized and understood that a grayscale camera may consume less power, be more sensitive in low light conditions, output data faster and / or output less data than a camera formed by a comparable image sensor configured to sense color, to represent the same range of the physical world at the same resolution. However, the image information output by the grayscale camera is sufficient for many functions performed in the XR system. Therefore, the XR system may be configured with a grayscale camera and a color camera, primarily using one or more grayscale cameras, and selectively using a color camera.
[0177] For example, an XR system may collect and process image information to create a navigable world model. This processing may use color information, which may enhance the effectiveness of certain functions, such as distinguishing objects, identifying surfaces associated with the same object, and / or identifying objects. Such processing may be performed or updated from time to time, such as when a user first turns on the system, moves to a new environment (e.g., walks into another room), or otherwise detects a change in the user's environment.
[0178] Other functions are not significantly improved by using color information. For example, once a traversable world model is created, the XR system can use images from one or more cameras to determine the orientation of the wearable device relative to features in the traversable world model. For example, such functionality can be accomplished as part of head pose tracking. Some or all of the cameras used for such functionality can be grayscale. Because head pose tracking is performed frequently, and in some embodiments continuously, while the XR system is running, using one or more grayscale cameras for this functionality can provide considerable power savings, reduced computation, or other benefits.
[0179] Similarly, at various times during operation of the XR system, the system may use stereo information from two or more cameras to determine the distance to a movable object. Such functionality may require high-speed processing of image information as part of tracking a user's hands or other movable objects. Using one or more grayscale cameras for this functionality may provide lower latency or other benefits associated with processing high-resolution image information.
[0180] In some embodiments of an XR system, the XR system may have both a color camera and at least one grayscale camera, and may selectively enable the grayscale and / or color cameras based on the functions for which image information from those cameras will be used.
[0181] Pixel cells in an image sensor may be distinguished in a manner other than based on the spectrum to which the pixel cells are sensitive. In some embodiments, some or all pixel cells may produce an output having an intensity indicative of the angle of arrival of light incident on the pixel cells. The angle of arrival information may be processed to calculate the distance to the imaged object.
[0182] In such an embodiment, the image sensor can passively acquire depth information. Passive depth information can be obtained by placing a component in the optical path to a pixel cell in the array so that the pixel cell outputs information indicating the angle of arrival of light illuminating the pixel cell. An example of such a component is a transmissive diffraction mask (TDM) filter.
[0183] The angle of arrival information can be converted into distance information by calculation, indicating the distance to the object reflecting the light. In some embodiments, the pixel cells configured to provide the angle of arrival information can be interspersed with pixel cells that capture the light intensity of one or more colors. As a result, the angle of arrival information and therefore the distance information can be combined with other image information about the object.
[0184] In some embodiments, one or more sensors may be configured to acquire information about physical objects in a scene at high frequency and low latency using compact and low power components. For example, the power consumption of an image sensor may be less than 50 milliwatts, enabling the device to be powered by a battery that is small enough to be used as part of a wearable system. The sensor may be an image sensor configured to passively acquire depth information as a supplement or alternative to image information indicating the intensity of one or more colors and / or changes in intensity information. Such a sensor may also be configured to provide a small amount of data using block tracking or by using DVS techniques to provide a differential output.
[0185] Passive depth information can be obtained by configuring an image array, such as an image array in combination with any one or more of the techniques described herein, wherein components in the image array adapt one or more pixel cells in the array to output information indicating a light field emitted from an imaged object. The information may be based on the angle of arrival of the light illuminating the pixel. In some embodiments, pixel cells such as those described above may be configured to output an indication of the angle of arrival by placing a plenoptic component in the light path to the pixel cell. An example of a plenoptic component is a transmissive diffraction mask (TDM). The angle of arrival information may be converted by calculation into distance information indicating the distance to the object from which the light is reflected to form the image being captured. In some embodiments, the pixel cells configured to provide angle of arrival information may be interspersed with pixel cells that capture light intensities in grayscale or one or more colors. As a result, the angle of arrival information may also be combined with other image information about the object.
[0186] Fig.16 A pixel subarray 100 is shown according to some embodiments. In the illustrated embodiment, the subarray has two pixel cells, but the number of pixel cells in the subarray does not limit the present invention. Here, a first pixel cell 121 and a second pixel cell 122 are shown, one of which is configured to capture angle of arrival information (first pixel cell 121), but it should be understood that the number and position within the array of pixel cells configured to measure angle of arrival information may vary. In this example, the other pixel cell (second pixel cell 122) is configured to measure the intensity of one color of light, but other configurations are possible, including pixel cells that are sensitive to different colors of light or one or more pixel cells that are sensitive to a wide spectrum of light, such as in a grayscale camera.
[0187] Fig.16 The first pixel unit 121 of the pixel subarray 100 includes an arrival angle to intensity converter 101, a photodetector 105, and a differential readout circuit 107. The second pixel unit 122 of the pixel subarray 100 includes a filter 102, a photodetector 106, and a differential readout circuit 108. It should be understood that it is not Fig.16 All components shown in need to be included in each embodiment. For example, some embodiments may not include differential readout circuits 107 and / or 108, and some embodiments may not include filter 102. In addition, Fig.16107. Other components not shown in the drawings. For example, some embodiments may include a polarizer arranged to allow light of a specific polarization to reach the photodetector. As another example, some embodiments may include a scan output circuit instead of the differential readout circuit 107, or include a scan output circuit in addition to the differential readout circuit 107. As another example, the first pixel unit 121 may also include a filter so that the first pixel 121 measures the arrival angle and intensity of light of a specific color incident on the first pixel 121.
[0188] The arrival angle to intensity converter 101 of the first pixel 121 is an optical component that converts the angle θ of the incident light 111 into an intensity that can be measured by a photodetector. In some embodiments, the arrival angle to intensity converter 101 may include a refractive optical device. For example, one or more lenses can be used to convert the incident angle of light into a position on the image plane, that is, the amount of incident light detected by one or more pixel units. In some embodiments, the arrival angle to position intensity converter 101 may include a diffractive optical device. For example, one or more diffraction gratings (e.g., TDM) can convert the incident angle of light into an intensity that can be measured by a photodetector below the TDM.
[0189] The photodetector 105 of the first pixel unit 121 receives incident light 110 passing through the arrival angle to intensity converter 101 and generates an electrical signal based on the intensity of the light incident on the photodetector 105. The photodetector 105 is located at an image plane associated with the arrival angle to intensity converter 101. In some embodiments, the photodetector 105 may be a single pixel of an image sensor, such as a CMOS image sensor.
[0190] The differential readout circuit 107 of the first pixel 121 receives the signal from the photodetector 105 and outputs an event only when the amplitude of the electrical signal from the photodetector is different from the amplitude of the previous signal from the photodetector 105, implementing the DVS technique as described above.
[0191] The second pixel unit 122 includes a filter 102 for filtering the incident light 112 so that only light within a specific wavelength range passes through the filter 102 and is incident on the photodetector 106. The filter 102 may be, for example, a bandpass filter that allows one of red, green or blue light to pass and rejects light of other wavelengths and / or may limit the IR light reaching the photodetector 106 to only a specific portion of the spectrum.
[0192] In this example, the second pixel cell 122 also includes a photodetector 106 and a differential readout circuit 108 , which may operate similarly to the photodetector 105 and the differential readout circuit 107 of the first pixel cell 121 .
[0193] As described above, in some embodiments, the image sensor may include an array of pixels, each pixel being associated with a photodetector and readout circuitry. A subset of the pixels may be associated with an angle of arrival to intensity converter for determining the angle of detection light incident on the pixel. Other subsets of the pixels may be associated with filters for determining color information about the scene being viewed or that may selectively pass or block light based on other characteristics.
[0194] In some embodiments, a single photodetector and two diffraction gratings of different depths can be used to determine the angle of arrival of light. For example, light can be incident on a first TDM, the angle of arrival is converted to position, and a second TDM can be used to selectively pass light incident at a specific angle. This arrangement can take advantage of the Talbot effect, which is a near-field diffraction effect in which an image of the diffraction grating is produced at a certain distance from the diffraction grating when a plane wave is incident on the diffraction grating. If a second diffraction grating is placed at an image plane that forms an image of the first diffraction grating, the angle of arrival can be determined based on the light intensity measured by a single photodetector located after the second grating.
[0195] Fig.17A A first arrangement of pixel cells 140 is shown, which includes a first TDM 141 and a second TDM 143 aligned to each other such that the ridges and / or regions of increased refractive index of the two gratings are aligned in the horizontal direction (Δs=0), where Δs is the horizontal offset between the first TDM 141 and the second TDM 143. The first TDM 141 and the second TDM 143 may both have the same grating period d, and the two gratings may be separated by a distance / depth z. The depth z at which the second TDM 143 is located relative to the first TDM 141, known as the Talbot length, may be determined by the grating period d and the wavelength λ of the light being analyzed, and is given by:
[0196]
[0197] like Fig.17AAs shown, incident light 142 having a zero degree arrival angle is diffracted by the first TDM 141. The second TDM 143 is located at a depth equal to the Talbot length, so that an image of the first TDM 141 is created, resulting in most of the incident light 142 passing through the second TDM 143. The optional dielectric layer 145 can separate the second TDM 143 from the photodetector 147. When the light passes through the dielectric layer 145, the photodetector 147 detects the light and generates an electrical signal, the characteristics of which (e.g., voltage or current) are proportional to the intensity of the light incident on the photodetector. On the other hand, when the incident light 144 having a non-zero arrival angle θ is also diffracted by the first TDM 141, the second TDM 143 blocks at least a portion of the incident light 144 from reaching the photodetector 147. The amount of incident light reaching the photodetector 147 depends on the arrival angle θ, and the larger the angle, the less light reaches the photodetector. The dashed line generated by the light 144 indicates that the amount of light reaching the photodetector 147 is attenuated. In some cases, light 144 may be completely blocked by diffraction grating 143. Thus, two TDMs may be used, using a single photodetector 147 to obtain information about the angle of arrival of the incident light.
[0198] In some embodiments, information obtained from neighboring pixel cells that do not have an arrival angle to intensity converter can provide an indication of the intensity of the incident light and can be used to determine the portion of the incident light that passes through the arrival angle to intensity converter. Based on this image information, the arrival angle of the light detected by the photodetector 147 can be calculated, as described in more detail below.
[0199] Fig. 17B A second arrangement of pixel cells 150 is shown, which includes a first TDM 151 and a second TDM 153 that are misaligned with each other, such that the ridges and / or regions of increased refractive index of the two gratings are misaligned in the horizontal direction (Δs≠0), where Δs is the horizontal offset between the first TDM 151 and the second TDM 153. The first TDM 151 and the second TDM 153 may have the same grating period d, and the two gratings may be separated by a distance / depth z. In combination with Fig.17A Unlike the discussed case where two TDMs are aligned, the misalignment results in incident light having an angle different from zero passing through the second TDM 153 .
[0200] like Fig. 17BAs shown, incident light 152 having a zero degree arrival angle is diffracted by the first TDM 151. The second TDM 153 is located at a depth equal to the Talbot length, but due to the horizontal offset of the two gratings, at least a portion of the light 152 is blocked by the second TDM 153. The dashed line generated by the light 152 shows that the amount of light reaching the photodetector 157 is attenuated. In some cases, the light 152 can be completely blocked by the diffraction grating 153. On the other hand, incident light 154 having a non-zero arrival angle θ is diffracted by the first TDM 151, but passes through the second TDM 153. After passing through the optional dielectric layer 155, the photodetector 157 detects the light incident on the photodetector 157 and generates an electrical signal whose characteristic (e.g., voltage or current) is proportional to the intensity of the light incident on the photodetector.
[0201] Pixel units 140 and 150 have different output functions, where different light intensities are detected for different angles of incidence. However, in each case, the relationship is fixed and can be determined based on the design of the pixel unit or by measurements as part of a calibration process. Regardless of the exact transfer function, the measured intensity can be converted to an angle of arrival, which in turn can be used to determine the distance to the imaged object.
[0202] In some embodiments, different pixel cells of the image sensor can have different TDM arrangements. For example, a first subset of pixel cells can include a first horizontal offset between the gratings of the two TDMs associated with each pixel cell, and a second subset of pixel cells can include a second horizontal offset between the gratings of the two TDMs associated with each pixel cell, where the first offset is different from the second offset. Each subset of pixel cells with different offsets can be used to measure a different angle of arrival or a different range of angles of arrival. For example, a first subset of pixel cells can include something like Fig.17A The second pixel subset may include a TDM arrangement of pixel units 140 similar to Fig. 17B TDM arrangement of pixel unit 150.
[0203] In some embodiments, not all pixel cells of an image sensor include a TDM. For example, a subset of pixel cells may include a filter, while a different subset of pixel cells may include a TDM for determining angle of arrival information. In other embodiments, no filter is used, so that a first subset of pixel cells simply measures the total intensity of incident light, while a second subset of pixel cells measures angle of arrival information. In some embodiments, information related to the light intensity from nearby pixel cells without TDMs can be used to determine the angle of arrival of light incident on a pixel cell with one or more TDMs. For example, using two TDMs arranged to utilize the Talbot effect, the light intensity incident on the photodetector after the second TDM is a sinusoidal function of the angle of arrival of the light incident on the first TDM. Therefore, if the total intensity of light incident on the first TDM is known, the angle of arrival of the light can be determined based on the intensity of light detected by the photodetector.
[0204] In some embodiments, the configuration of pixel elements in a subarray may be selected to provide various types of image information with appropriate resolution. 18A to 18C An example arrangement of pixel cells in a pixel subarray of an image sensor is shown. The example shown is a non-limiting arrangement, as it should be understood that the inventors have designed alternative pixel arrangements. This arrangement can be repeated across an image array, which can contain millions of pixels. A subarray can include one or more pixel cells that provide information about the angle of arrival of incident light and one or more other pixel cells (with or without filters) that provide information about the intensity of the incident light.
[0205] Fig.18A 1 is an example of a pixel subarray 160, which includes a first group of pixel cells 161 and a second group of pixel cells 163 that are different from each other and are rectangular rather than square. The pixel cells marked as "R" are pixel cells with red filters, so that red incident light passes through the filter to reach the associated photodetector; the pixel cells marked as "B" are pixel cells with blue filters, so that blue incident light passes through the filter to reach the associated photodetector; the pixel cells marked as "G" are pixels with green filters, so that green incident light passes through the filter to reach the associated photodetector. In the example subarray 160, there are more green pixel cells than red or blue pixel cells, indicating that the various types of pixel cells do not need to exist in the same proportion.
[0206] The pixel units labeled A1 and A2 are pixels that provide angle of arrival information. For example, pixel units A1 and A2 may include one or more gratings for determining angle of arrival information. The pixel units that provide angle of arrival information may be configured similarly or may be configured differently, such as being sensitive to different ranges of angles of arrival or angles of arrival relative to different axes. In some embodiments, the pixels labeled A1 and A2 include two TDMs, and the TDMs of pixel units A1 and A2 may be oriented in different directions, such as perpendicular to each other. In other embodiments, the TDMs of pixel units A1 and A2 may be oriented parallel to each other.
[0207] In an embodiment using pixel subarray 160, color image data and angle of arrival information can be obtained. In order to determine the angle of arrival of light incident on pixel unit group 161, the total light intensity incident on pixel unit group 161 is estimated using electrical signals from RGB pixel units. Using the fact that the light intensity detected by A1 / A2 pixels varies with the angle of arrival in a predictable manner, the angle of arrival can be determined by comparing the total intensity (estimated based on the RGB pixel units within the pixel group) with the intensity measured by A1 and / or A2 pixel units. For example, the intensity of light incident on A1 and / or A2 pixels can vary sinusoidally with respect to the angle of arrival of the incident light. The angle of arrival of light incident on pixel unit group 163 is determined in a similar manner using electrical signals generated by pixel unit group 163.
[0208] It should be understood that Fig.18A A particular embodiment of a sub-array is shown, and other configurations are possible. In some embodiments, for example, a sub-array may be only a pixel unit group 161 or 163 .
[0209] Fig.18B 170 is an alternative pixel subarray 170 including a first group of pixel cells 171, a second group of pixel cells 172, a third group of pixel cells 173, and a fourth group of pixel cells 174. Each group of pixel cells 171-174 is square and has the same pixel cell arrangement therein, but in order to have pixel cells for determining the angle of arrival information within different angle ranges or relative to different planes (for example, the TDM of pixels A1 and A2 can be oriented perpendicular to each other). Each group of pixels 171-174 includes a red pixel cell (R), a blue pixel cell (B), a green pixel cell (G), and an angle of arrival pixel cell (A1 or A2). Note that in the example pixel subarray 170, there are the same number of red / green / blue pixel cells in each group. In addition, it should be understood that the pixel subarray can be repeated in one or more directions to form a larger pixel array.
[0210] In an embodiment using pixel subarray 170, color image data and angle of arrival information may be obtained. In order to determine the angle of arrival of light incident on pixel cell group 171, the signals from the RGB pixel cells may be used to estimate the total light intensity incident on pixel cell group 171. Using the fact that the light intensity detected by the angle of arrival pixel cells has a sinusoidal or other predictable response with respect to the angle of arrival, the angle of arrival may be determined by comparing the total intensity (estimated from the RGB pixel cells) with the intensity measured by the A1 pixel. The angle of arrival of light incident on pixel cell groups 172-174 may be determined in a similar manner using electrical signals generated by the pixel cells of each corresponding pixel group.
[0211] Fig.18C 1 is an alternative pixel subarray 180 including a first group of pixel cells 181, a second group of pixel cells 182, a third group of pixel cells 183, and a fourth group of pixel cells 184. Each group of pixel cells 181-184 is square and has the same pixel cell arrangement, in which no filters are used. Each group of pixel cells 181-184 includes: two "white" pixels (e.g., no filters so that red, blue, and green light are detected to form a grayscale image); an arrival angle pixel cell (A1) in which the TDM is oriented in a first direction; and an arrival angle pixel cell (A2) in which the TDM is oriented at a second pitch or in a second direction relative to the first direction (e.g., vertically). Note that there is no color information in the example pixel subarray 170. The resulting image is grayscale, illustrating that passive depth information can be obtained in a color or grayscale image array using the techniques described herein. As with other subarray arrangements described herein, the pixel subarray arrangement can be repeated in one or more directions to form a larger pixel array.
[0212] In an embodiment using pixel subarray 180, grayscale image data and angle of arrival information can be obtained. In order to determine the angle of arrival of light incident on pixel unit group 181, the total light intensity incident on pixel unit group 181 is estimated using electrical signals from two white pixels. Using the fact that the light intensity detected by A1 and A2 pixels has a sinusoidal or other predictable response relative to the angle of arrival, the angle of arrival can be determined by comparing the total intensity (estimated based on the white pixels) with the intensity measured by the A1 and / or A2 pixel units. The angle of arrival of light incident on pixel unit groups 182-184 can be determined in a similar manner using electrical signals generated by the pixels of each corresponding pixel group.
[0213] In the above examples, the pixel units are illustrated as squares and arranged in a square grid. Embodiments are not limited thereto. For example, in some embodiments, the shape of the pixel unit may be rectangular. In addition, the subarrays may be triangular or arranged on a diagonal or have other geometric shapes.
[0214] In some embodiments, the angle of arrival information is obtained using the image processor 708 or a processor associated with the local data processing module 70, and the processor can further determine the distance of the object based on the angle of arrival. For example, the angle of arrival information can be combined with one or more other types of information to obtain the distance of the object. In some embodiments, the objects of the grid model 46 can be associated with the angle of arrival information from the pixel array. The grid model 46 can include the location of the object, including the distance from the user, which can be updated to a new distance value based on the angle of arrival information.
[0215] Using angle of arrival information to determine distance values can be particularly useful in scenarios where objects are close to the user. This is because changes in distance from the image sensor cause changes in the angle of light arrival for nearby objects to be greater than similar magnitude distance changes for objects positioned farther away from the user. Therefore, a processing module that utilizes passive distance information based on the angle of arrival can selectively use this information based on an estimated object distance, and can utilize one or more other techniques to determine the distance to objects that exceed a threshold distance, in some embodiments, such as up to 1 meter, up to 3 meters, or up to 5 meters. As a specific example, a processing module of an AR system can be programmed to use passive distance measurement that uses angle of arrival information for objects within 3 meters of the wearable device user, but for objects outside of this range, stereo image processing can be used that uses images captured by two cameras.
[0216] Similarly, pixels configured to detect angle of arrival information may be most sensitive to distance changes within an angular range from normal to the image array. The processing module may similarly be configured to use distance information derived from angle of arrival measurements within this angular range, but use other sensors and / or other techniques to determine distances outside of this range.
[0217] An example application of determining the distance of an object from an image sensor is hand tracking. Hand tracking can be used in AR systems, for example, to provide a gesture-based user interface for system 80 and / or to allow a user to move virtual objects within an environment in an AR experience provided by system 80. The combination of an image sensor that provides angle of arrival information for accurate depth determination and a differential readout circuit for reducing the amount of processed data for determining user hand movement provides an effective interface through which a user can interact with virtual objects and / or provide input to system 80. A processing module that determines the position of a user's hand can use distance information that is acquired using different techniques, depending on the position of the user's hand in the field of view of the image sensor of the wearable device. In accordance with some embodiments, hand tracking can be implemented as a form of block tracking during the image sensing process.
[0218] Another application where depth information can be useful is occlusion processing. Occlusion processing uses depth information to determine that certain parts of the physical world model do not need to or cannot be updated based on image information captured by one or more image sensors that collect image information about the physical environment around the user. For example, if it is determined that there is a first object at a first distance from the sensor, the system 80 can determine not to update the model of the physical world for distances greater than the first distance. For example, even if the model includes a second object at a second distance from the sensor, the second distance is greater than the first distance, if the second object is behind the first object, the model information of the object may not be updated. In some embodiments, the system 80 can generate an occlusion mask based on the position of the first object, and only update the parts of the model that are not occluded by the occlusion mask. In some embodiments, the system 80 can generate more than one occlusion mask for more than one object. Each occlusion mask can be associated with a corresponding distance from the sensor. For each occlusion mask, the model information associated with the following objects will not be updated: the distance of the object from the sensor is greater than the distance associated with the corresponding occlusion mask. By limiting the part of the model that is updated at any given time, the speed of generating the AR environment and the amount of computing resources required to generate the AR environment are reduced.
[0219] Although not in 18A to 18C As shown in , some embodiments of the image sensor may include pixels with IR filters in addition to or instead of filters. For example, the IR filter may allow light of a wavelength such as approximately equal to 940nm to pass through and be detected by an associated photodetector. Some embodiments of the wearable device may include an IR light source (e.g., an IR LED) that emits light of the same wavelength as the wavelength associated with the IR filter (e.g., 940nm). The IR light source and IR pixels may be used as an alternative way to determine the distance of an object from the sensor. By way of example and not limitation, the IR light source may be pulsed and a time-of-flight measurement may be used to determine the distance of an object from the sensor.
[0220] In some embodiments, the system 80 can operate in one or more operating modes. The first mode can be a mode in which depth is determined using passive depth measurement, for example, based on the angle of arrival of light determined using pixels with an angle of arrival to intensity converter. The second mode can be a mode in which depth is determined using active depth measurement, for example, based on the time of flight of IR light measured using IR pixels of an image sensor. The third mode can use stereo measurements from two independent image sensors to determine the distance of an object. When the object is far away from the sensor, such stereo measurements can be more accurate than the angle of arrival of light determined using pixels with an angle of arrival to intensity converter. Other appropriate methods of determining depth can be used for one or more additional depth determination operating modes.
[0221] In some embodiments, it may be preferred to use passive depth determination because such a technique uses less power. However, the system may determine that it should operate in active mode under certain conditions. For example, if the visible light intensity detected by the sensor is below a threshold, it may be too dark to accurately perform passive depth determination. In some embodiments, IR illumination can be used to offset the effects of low lighting. IR illumination can be selectively enabled in response to detected low lighting conditions that prevent the acquisition of images sufficient for the task being performed by the device (e.g., head tracking or hand tracking). As another example, the object may be too far away for passive depth determination to be inaccurate. Therefore, the system can be programmed to choose to operate in a third mode, in which the depth is determined based on stereo measurements of a scene using two spatially separated image sensors. As another example, determining the depth of an object based on the angle of arrival of light determined using pixels with an angle of arrival to intensity converters may be inaccurate in the periphery of the image sensor. Therefore, if an object is being detected by pixels near the periphery of the image sensor, the system can choose to operate in a second mode, using active depth determination.
[0222] While the above-described image sensor embodiments use individual pixel cells with stacked TDMs to determine the angle of arrival of light incident on the pixel cells, other embodiments may use multiple pixel cells with a single TDM on all pixels in a group to determine the angle of arrival information. The TDM may project a light pattern on the sensor array that depends on the angle of arrival of the incident light. Multiple photodetectors associated with one TDM may more accurately detect the pattern because each of the multiple photodetectors is located at a different position in the image plane (the image plane includes the photodetectors that sense the light). The relative intensity sensed by each photodetector may indicate the angle of arrival of the incident light.
[0223] Fig.19A is an example of a top view of multiple photodetectors (in the form of a photodetector array 120, which may be a subarray of pixel cells of an image sensor) associated with a single transmissive diffraction mask (TDM) in accordance with some embodiments. Fig.19B is with Fig.19A The same photodetector array along Fig.19A 1. A cross-sectional view of line A of FIG. 1. In the example shown, the photodetector array 120 includes 16 individual photodetectors 121, which may be within a pixel unit of the image sensor. The photodetector array 120 includes a TDM 123 disposed above the photodetector. It should be understood that for clarity and simplicity, each group of pixel units is shown with four pixels (e.g., forming a four-pixel by four-pixel grid). Some embodiments may include more than four pixel units. For example, 16 pixel units, 64 pixel units, or any other number of pixels may be included in each group.
[0224] TDM 123 is located at a distance x from photodetector 121. In some embodiments, TDM 123 is formed on the top surface of dielectric layer 125, such as Fig.19B As shown. For example, as shown, the TDM 123 can be formed by ridges, or by valleys etched into the surface of the dielectric layer 125. In other embodiments, the TDM 123 can be formed within the dielectric layer. For example, portions of the dielectric layer can be modified to have a higher or lower refractive index relative to other portions of the dielectric layer, thereby producing a holographic phase grating. Light incident on the photodetector array 120 from above is diffracted by the TDM, causing the angle of arrival of the incident light to be converted into a position in the image plane at a distance x from the TDM 123, where the photodetector 121 is located. The incident light intensity measured at each photodetector 121 of the photodetector array can be used to determine the angle of arrival of the incident light.
[0225] Fig. 20A An example of multiple photodetectors (in the form of a photodetector array 130) associated with multiple TDMs is shown in accordance with some embodiments. Fig. 20B is with Fig. 20A The same photodetector array is Fig. 20A A cross-sectional view along line B of FIG. Fig. 20C is with Fig. 20A The same photodetector array is Fig. 20A . A cross-sectional view of line C of FIG. 1 . In the example shown, the photodetector array 130 includes 16 individual photodetectors, which may be within a pixel cell of an image sensor. Four groups 131a, 131b, 131c, 131d of four pixel cells are shown. The photodetector array 130 includes four individual TDMs 133a, 133b, 133c, and 133d, each TDM being disposed above an associated group of pixel cells. It should be understood that for clarity and simplicity, each group of pixel cells is illustrated with four pixel cells. Some embodiments may include more than four pixel cells. For example, each group may include 16 pixel cells, 64 pixel cells, or any other number of pixel cells.
[0226] Each TDM 133a-d is located at a distance x from the photodetectors 131a-d. In some embodiments, the TDMs 133a-d are formed on the top surface of the dielectric layer 135, such as Fig. 20BAs shown. For example, the TDMs 123a-d can be formed by ridges, as shown, or by valleys etched into the surface of the dielectric layer 135. In other embodiments, the TDMs 133a-d can be formed within the dielectric layer. For example, portions of the dielectric layer can be modified to have a higher or lower refractive index relative to other portions of the dielectric layer, thereby producing a holographic phase grating. Light incident on the photodetector array 130 from above is diffracted by the TDMs, causing the angle of arrival of the incident light to be converted into a position in the image plane at a distance x from the TDMs 133a-d, where the photodetectors 131a-d are located. The incident light intensity measured at each photodetector 131a-d of the photodetector array can be used to determine the angle of arrival of the incident light.
[0227] TDM 133a-d can be oriented in different directions from each other. For example, TDM 133a is perpendicular to TDM 133b. Therefore, the light intensity detected using photodetector group 131a can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133a, and the light intensity detected using photodetector group 131b can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133b. Similarly, the light intensity detected using photodetector group 131c can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133c, and the light intensity detected using photodetector group 131d can be used to determine the angle of arrival of incident light in a plane perpendicular to TDM 133d.
[0228] Pixel units configured to passively acquire depth information may be integrated into an image array having features as described herein to support useful operations in an X-reality system. According to some embodiments, a pixel unit configured to acquire depth information may be implemented as part of an image sensor for implementing a camera with a global shutter. For example, such a configuration may provide a full-frame output. A full frame may include image information of different pixels that simultaneously indicate depth and intensity. Using an image sensor of such a configuration, a processor may acquire depth information of a full scene at once.
[0229] In other embodiments, the pixel cells of the image sensor providing depth information may be configured to operate according to the DVS technique as described above. In such a scenario, an event may indicate a change in the depth of an object, as indicated by the pixel cell. The event output by the image array may indicate the pixel cell that detected the depth change. Alternatively or additionally, the event may include the value of the depth information of the pixel cell. Using an image sensor configured in this way, the processor can obtain depth information updates at a very high rate, thereby providing high temporal resolution. In some embodiments, high temporal resolution may involve updating depth information more frequently than at 1 Hz or 5 Hz, and may alternatively involve updating depth information hundreds or thousands of times per second, for example every millisecond.
[0230] In yet other embodiments, the image sensor may be configured to operate in full frame or DVS mode. In such embodiments, a processor that processes image information from the image sensor may programmatically control the operating mode of the image sensor based on the functions performed by the processor. For example, when performing functions involving tracking objects, the processor may configure the image sensor to output image information as DVS events. On the other hand, when processing update world reconstruction, the processor may configure the image sensor to output full frame depth information. In some embodiments, in full frame and / or DVS mode, active illumination may be used to illuminate the imaged scene. For example, IR illumination may be provided when the light intensity detected by the sensor is below a threshold, in which case it may be too dim to accurately perform passive depth determination.
[0231] Wearable Configuration
[0232] Multiple image sensors may be used in an XR system. Image sensors may be combined with optical components (e.g., lenses) and control circuitry to create a camera. Those image sensors may use one or more of the above-mentioned techniques to acquire imaging information, such as grayscale imaging, color imaging, global shutter, DVS techniques, plenoptic pixel cells, and / or dynamic blocking. Regardless of the imaging technology used, the resulting camera may be mounted to a support member to form a head-mounted device, which may include or be connected to a processor.
[0233] Fig.21 2 is a schematic diagram of a head mounted device 2100 of a wearable display system consistent with the disclosed embodiments. Fig.21 As shown, the head mounted device 2100 may include a display device, the display device including a monocular 2110a and a monocular 2110b, the monocular 2110a and the monocular 2110b may be optical eyepieces or displays configured to transmit and / or display visual information to the user's eyes. The head mounted device 2100 may also include a frame 2101, which may be similar to the frame 2101 described above. Figure 3B The described frame 64. The head mounted device 2100 may also include two cameras (DVS camera 2120 and camera 2140) and additional components, such as a transmitter 2130a, a transmitter 2130b, an inertial measurement unit 2170a (IMU 2170a), and an inertial measurement unit 2170b (IMU 2170b).
[0234] Camera 2120 and camera 2140 are world cameras because they are oriented to image the physical world as seen by a user wearing head mounted device 2100. In some embodiments, these two cameras may be sufficient to acquire image information about the physical world, and these two cameras may be the only cameras facing the world. Head mounted device 2100 may also include additional components, such as an eye tracking camera, as described above with respect to Figure 3B discussed.
[0235] Monocular 2110a and monocular 2110b can be mechanically coupled to a support member such as frame 2101 using techniques such as adhesives, fasteners, or press fits. Similarly, two cameras and accompanying components (e.g., a transmitter, an inertial measurement unit, an eye tracking camera, etc.) can be mechanically coupled to frame 2101 using techniques such as adhesives, fasteners, press fits, etc. These mechanical couplings can be direct or indirect. For example, one or more cameras and / or one or more accompanying components can be directly attached to frame 2101. As an additional example, one or more cameras and / or one or more accompanying components can be directly attached to a monocular, which can then be attached to frame 2101. The mechanism of attachment is not intended to be limiting.
[0236] Alternatively, a monocular mirror assembly can be formed and then attached to the frame 2101. Each subassembly can include, for example, a support member to which the monocular 2110a or 2110b is attached. The IMU and one or more cameras can be similarly attached to a support member. Attaching the camera and the IMU to the same support member can obtain inertial information about the camera based on the output of the IMU. Similarly, attaching the monocular to the same support member as the camera can spatially correlate image information about the world with information rendered on the monocular.
[0237] The head-mounted device 2100 can be lightweight. For example, the weight of the head-mounted device 2100 can be between 30 and 300 grams. The head-mounted device 2100 can be made of a material that bends during use, such as plastic or thin metal. Such materials can achieve a lightweight and comfortable head-mounted device that can be worn by a user for a long time. Nevertheless, an XR system with such a lightweight head-mounted device can support high-precision stereo image analysis (which requires knowing the separation between the cameras), using a calibration procedure that can be repeated while the head-mounted device is worn to compensate for any inaccuracies caused by the bending of the head-mounted device during use. In some embodiments, the lightweight head-mounted device may include a battery pack. The battery pack may include one or more batteries, which may be rechargeable or non-rechargeable. The battery pack may be built into the lightweight frame or may be removable. The battery pack and the lightweight frame may be formed as a single unit, or the battery pack may be formed as a unit separate from the lightweight frame.
[0238] The DVS camera 2120 may include an image sensor and a lens. The image sensor may be configured to generate a grayscale image. The image sensor may be configured to output an image having a size between 1 megapixel and 4 megapixels. For example, the image sensor may be configured to output an image having a horizontal resolution of 1016 lines and a vertical resolution of 1016 lines. In some aspects, the image sensor may be a CMOS image sensor.
[0239] DVS camera 2120 can support the above Figure 4 , Figure 5A and Figure 5B Dynamic visual sensing disclosed. In an operating mode, the DVS camera 2120 can be configured to output image information in response to a detected change in an image characteristic (e.g., a pixel-level change in light intensity). In some embodiments, the detected change can meet an intensity change criterion. For example, the image sensor can be configured to have a threshold for an incremental increase in light intensity and another threshold for an incremental decrease in light intensity. When such a change is detected, image information can be provided asynchronously. In another operating mode, the DVS camera 120 can be configured to repeatedly or periodically output image frames without responding to a detected change in image characteristics. For example, the DVS camera 2120 can be configured to output images at a frequency between 30 Hz and 120 Hz (e.g., at 60 Hz).
[0240] DVS camera 2120 can support the above Figures 6 to 15For example, in various aspects, the DVS camera 2120 can be configured to provide image information for a subset of pixels in the image sensor (e.g., a block of pixels in the image sensor). The DVS camera 2120 can be configured to combine dynamic visual sensing with block tracking to provide image information for those pixels in the block that experience changes in image characteristics. In various aspects, the image sensor can be configured with a global shutter. As described above, with respect to Fig.14 and Fig.15 , a global shutter enables each pixel to acquire intensity measurements simultaneously.
[0241] DVS camera 2120 may be as described above with respect to Figure 3B and Figures 15 to 20C The plenoptic camera discussed. For example, a component may be placed in the optical path of one or more pixel cells of the image sensor of the DVS camera 2120 so that the pixel cells produce an output having an intensity indicating the angle of arrival of light incident on the pixel cells. In such an embodiment, the image sensor may passively acquire depth information. An example of a component suitable for placement in the optical path is a TDM filter. The processor may be configured to use the angle of arrival information to calculate the distance to the imaged object. For example, the angle of arrival information may be converted into distance information indicating the distance to the object from which the light is reflected. In some embodiments, the pixel cells configured to provide the angle of arrival information may be interspersed with pixel cells that capture the light intensity of one or more colors. As a result, the angle of arrival information, and therefore the distance information, may be combined with other image information about the object.
[0242] The DVS camera 2120 can be configured to have a wide field of view, consistent with the disclosed embodiments. For example, the DVS camera 2120 can include an equidistant lens (e.g., a fisheye lens). The DVS camera 2120 can be tilted toward the camera 2140, thereby creating an area imaged by the two cameras directly in front of the user of the head-mounted device 2100. For example, a vertical plane through the center of the field of view 2121, i.e., the field of view associated with the DVS camera 2120, can intersect and form an angle with a vertical plane through the centerline of the head-mounted device 2100. In some embodiments, the field of view 2121 can have a horizontal field of view and a vertical field of view. The horizontal field of view can range between 90 degrees and 175 degrees, while the vertical field of view can range between 70 degrees and 125 degrees. In some embodiments, the DVS camera 2120 can be configured to have an angular pixel resolution between 1 and 5 arc minutes per pixel.
[0243] Emitters 2130a and 2130b may be configured to emit light of a particular wavelength in low light conditions and / or when active depth sensing is performed by the head mounted device 2100. Emitters 2130a and 2130b may be configured to emit light of a particular wavelength. This light may be reflected by physical objects in the physical world around the user. The head mounted device 2100 may be configured with sensors to detect this reflected light, including image sensors as described herein. In some embodiments, these sensors may be incorporated into at least one of the cameras 2120 or 2140. For example, as described above with respect to 18A to 18C As described above, the cameras may be configured with detectors corresponding to emitter 2130a and / or emitter 2130b. For example, the cameras may include pixels configured to detect light emitted by emitter 2130a and emitter 2130b.
[0244] Emitter 2130a and emitter 2130b may be configured to emit IR light, consistent with the disclosed embodiments. The wavelength of the IR light may be between 900 nanometers and 1 micron. The IR light may be, for example, a 940nm light source, the emitted light energy being concentrated near 940nm. Emitters emitting light of other wavelengths may be used alternatively or additionally. For example, for a system intended for indoor use only, an emitter emitting light concentrated near 850nm may be used. At least one of the DVS cameras 2120 or 2140 may include one or more IR filters disposed above at least one subset of pixels in the image sensor of the camera. The filter may pass light of the wavelength emitted by emitter 2130a and / or emitter 2130b while attenuating some other wavelengths of light. For example, the IR filter may be a notch filter that passes IR light having a wavelength matching that of the emitter. The notch filter may significantly attenuate other IR light. In some embodiments, the notch filter may be an IR notch filter that blocks IR light and allows light from the emitter to pass. The IR notch filter may also allow light outside the IR band to pass. Such a notch filter may enable the image sensor to receive visible light and light from the emitter that has been reflected from an object in the field of view of the image sensor. In this way, a subset of pixels may be used as a detector of IR light emitted by emitter 2130a and / or emitter 2130b.
[0245] In some embodiments, the processor of the XR system can selectively enable the emitters, for example to enable imaging in low light conditions. The processor can process image information generated by one or more image sensors and can detect whether images output by those image sensors without enabling the emitters provide sufficient information about objects in the physical world. The processor can enable the emitters in response to detecting that the images do not provide sufficient image information due to low ambient light conditions. For example, the emitters can be turned on when stereo information is used to track an object and the lack of ambient light causes insufficient contrast between features of the tracked object to accurately determine distance using stereo image technology.
[0246] Alternatively or additionally, emitter 2130a and / or emitter 2130b may be configured to perform active depth measurements, for example by emitting light in short pulses. The wearable display system may be configured to perform time-of-flight measurements by detecting reflections of such pulses from objects in the illumination field 2131a of emitter 2130a and / or the illumination field 2131b of emitter 2130b. These time-of-flight measurements may provide additional depth information for tracking objects or updating a navigable world model. In other embodiments, one or more emitters may be configured to emit patterned light, and the XR system may be configured to process images of objects illuminated by the patterned light. Such processing may detect changes in the pattern, which may reveal the distance to the object.
[0247] In some embodiments, the range of the illumination field associated with the emitters may be sufficient to illuminate at least the field of view of a camera used to acquire image information about an object. For example, the emitters may collectively illuminate a central field of view 2150. In the illustrated embodiment, emitters 2130a and emitters 2130b may be positioned to illuminate illumination fields 2131a and 2131b, which collectively span a range over which active illumination may be provided. In this example embodiment, two emitters are shown, but it should be understood that more or fewer emitters may be used to span the desired range.
[0248] In some embodiments, a transmitter, such as transmitters 2130a and 2130b, may be turned off by default, but may be enabled when additional illumination is needed to obtain more information than passive imaging can obtain. The wearable display system may be configured to enable transmitter 2130a and / or transmitter 2130b when additional depth information is needed. For example, when the wearable display system detects that sufficient depth information for tracking hand or head gestures cannot be obtained using stereoscopic image information, the wearable display system may be configured to enable transmitter 2130a and / or transmitter 2130b. The wearable display system may be configured to disable transmitter 2130a and / or transmitter 2130b when additional depth information is not needed, thereby reducing power consumption and improving battery life.
[0249] Furthermore, even if the head-mounted device is configured with an image sensor configured to detect IR light, it is not a necessary condition to mount the IR emitter on the head-mounted device 2100 or only on the head-mounted device 2100. In some embodiments, the IR emitter may be an external device mounted in a space, such as an interior room, in which the head-mounted device 2100 may be used. Such an emitter may emit IR light, for example, in an ArUco pattern at 940 nm, which is invisible to the human eye. Light with such a pattern may enable "instrumented / assisted tracking", in which the head-mounted device 2100 does not have to supply power to provide the IR pattern, but may still provide IR image information as a result of the presence of the pattern, so that the distance or position of an object in the space may be determined by processing the image information. Systems with external illumination sources may also enable more devices to operate in the space. If multiple head-mounted devices are operated in the same space, each head-mounted device moves in the space without a fixed positional relationship, there is a risk that light emitted by one head-mounted device may be projected onto the image sensor of another head-mounted device, thereby interfering with its operation. For example, the risk of such interference between head mounted devices may limit the number of head mounted devices that can operate in a space to 3 or 4. With one or more IR emitters in the space illuminating objects that can be imaged by image sensors on the head mounted devices, more head mounted devices (more than 10 in some embodiments) can operate in the same space without interference.
[0250] As mentioned above about Figure 3BAs disclosed, the camera 2140 can be configured to capture images of the physical world within the field of view 2141. The camera 2140 can include an image sensor and a lens. The image sensor can be configured to generate a color image. The image sensor can be configured to output an image with a size between 4 megapixels and 16 megapixels. For example, the image sensor can output an image of 12 megapixels. The image sensor can be configured to output images repeatedly or periodically. For example, the image sensor can be configured to output images at a frequency between 30 Hz and 120 Hz (e.g., 60 Hz) when enabled. In some embodiments, the image sensor can be configured to selectively output images based on the task being performed. The image sensor can be configured with a rolling shutter. As described above with respect to Fig.14 and Fig.15 As discussed, a rolling shutter may iteratively read subsets of pixels in an image sensor such that pixels in different subsets reflect light intensity data collected at different times. For example, an image sensor may be configured to read a first row of pixels in the image sensor at a first time and a second row of pixels in the image sensor at a later time.
[0251] In some embodiments, camera 2140 can be configured as a plenoptic camera. Figure 3B and Figures 15 to 20C As discussed, components may be placed in the optical path of one or more pixel cells leading to the image sensor so that the pixel cells produce an output having an intensity indicating the angle of arrival of light incident on the pixel cells. In such an embodiment, the image sensor may passively acquire depth information. An example of a component suitable for placement in the optical path is a transmissive diffraction mask (TDM) filter. The processor may be configured to use the angle of arrival information to calculate the distance to the imaged object. For example, the angle of arrival information may be converted into distance information indicating the distance from the object from which the light is reflected. In some embodiments, the pixel cells configured to provide angle of arrival information may be interspersed with pixel cells that capture the light intensity of one or more colors. As a result, the angle of arrival information and therefore the distance information may be combined with other image information about the object. In some embodiments, similar to the DVS camera 2120, the camera 2140 may be configured to provide event detection and block tracking functionality. The processor of the head mounted device 2100 may be configured to provide instructions to the camera 2140 to limit image capture to a subset of pixels. In some embodiments, the sensor may be a CMOS sensor.
[0252] The camera 2140 may be positioned on a side of the head mounted device 2100 opposite to the DVS camera 2120. For example, Fig.21As shown, when camera 2140 is on the same side of head mounted device 2100 as monocular 2110a, DVS camera 2120 can be on the same side of head mounted device 2100 as monocular 2110b. Camera 2140 can be angled inward on head mounted device 2100. For example, a vertical plane through the center of field of view 2141, i.e., the field of view associated with camera 2140, can intersect and form an angle with a vertical plane through the centerline of head mounted device 2100. Field of view 2141 of camera 2140 can have a horizontal field of view and a vertical field of view. The horizontal field of view can range between 75 degrees and 125 degrees, while the vertical field of view can range between 60 degrees and 125 degrees.
[0253] The camera 2140 and the DVS camera 2120 may be angled asymmetrically inwardly toward the midline of the head mounted device 2100. The angle of the camera 2140 may be between 1 degree and 20 degrees inwardly toward the midline of the head mounted device 2100. The angle of the DVS camera 2120 may be between 1 degree and 40 degrees inwardly toward the midline of the head mounted device 2100, and may be different from the angle of the camera 2140. The angular range of the field of view 2121 may exceed the angular range of the field of view 2141.
[0254] DVS camera 2120 and camera 2140 may be configured to provide overlapping views of a central field of view 2150. The angular range of central field of view 2150 may be between 40 degrees and 120 degrees. For example, the angular range of central field of view 2150 may be approximately 70 degrees (e.g., 70±7 degrees). Central field of view 2150 may be asymmetric. For example, central field of view 2150 may further extend toward a side of head mounted device 2100 including camera 2140, such as Fig.21 As shown. In addition to the central field of view 2150, the DVS camera 2120 and the camera 2140 can be positioned to provide at least two peripheral fields of view. The peripheral field of view 2160a can be associated with the DVS camera 2120 and can include a portion of the field of view 2121 that does not overlap with the field of view 2141. In some embodiments, the horizontal angle range of the peripheral field of view 2160a can be within a range between 20 degrees and 80 degrees. For example, the angle range of the peripheral field of view 2160a can be approximately 40 degrees (e.g., 40±4 degrees). The peripheral field of view 2160b ( Fig.212121). In some embodiments, the horizontal angle range of the peripheral field of view 2160b can be in the range of 10 degrees to 40 degrees. For example, the angle range of the peripheral field of view 2160a can be approximately 20 degrees (e.g., 20±2 degrees). The location of the peripheral field of view can be different. For example, for a particular configuration of the head-mounted device 2100, the peripheral field of view 2160b may not extend within 0.25 meters of the head-mounted device 2100 because the field of view 2141 can fall completely within the field of view 2121 within this distance. In contrast, the peripheral field of view 2160a can extend within 0.25 meters of the head-mounted device 2100. In such a configuration, the wider field of view and larger inward angle of the DVS camera 2120 can ensure that, even within 0.25 meters of the head-mounted device 2100, the field of view 2121 at least partially falls outside the field of view 2141.
[0255] IMU 2170a and / or IMU 2170b can be configured to provide acceleration and / or velocity and / or inclination information to the wearable display system. For example, when a user wearing the head mounted device 2100 moves, IMU 2170a and / or IMU 2170b can provide information describing the acceleration and / or velocity of the user's head.
[0256] The XR system can be coupled to a processor that can be configured to process image data output by the camera and / or render virtual objects on a display device. The processor can be mechanically coupled to the frame 2101. Alternatively, the processor can be mechanically coupled to a display device, such as a display device including a monocular 2110a or a monocular 2110b. As a further alternative, the processor can be operably coupled to the head-mounted device 2100 and / or the display device via a communication link. For example, the XR system may include a local data processing module. The local data processing module may include a processor and may be connected to the head-mounted device 2100 and / or the display device via a physical connection (e.g., a wire or cable) or a wireless (e.g., Bluetooth, Wi-Fi, Zigbee, etc.) connection.
[0257] The processor may be configured to perform world reconstruction, head pose tracking, and object tracking operations. The processor may be configured to create a passable world model using the DVS camera 2120 and the camera 2140. When creating the passable world model, the processor may be configured to stereoscopically determine depth information using multiple images of the same physical object acquired by the DVS camera 2120 and the camera 2140. The processor may be configured to update an existing passable world model using the DVS camera 2120 instead of the camera 2140. As described above, the DVS camera 2120 may be a grayscale camera with a lower resolution than the color camera 2140. In addition, the DVS camera 2120 may output image information asynchronously (e.g., in response to a detected event), so that the processor can asynchronously update the passable world model, head pose, and / or object position only when a change is detected. In some embodiments, the processor may provide instructions to the DVS camera 2120 to limit image data acquisition to one or more pixel blocks (patch of pixel) in the image sensor of the DVS camera 2120. The DVS camera 2120 can then restrict image data acquisition to these pixel blocks. Thus, the traversable world model is updated quickly using image information output by the DVS camera 2120 rather than the camera 2140, which can reduce power consumption and extend battery life. In some embodiments, the processor can be configured to use the image information output by the DVS camera 2120 to occasionally or periodically update the traversable world model.
[0258] While the processor may preferentially use DVS camera 2120 to update the traversable world model, in some embodiments, the processor may occasionally or periodically update the traversable world model using DVS camera 2120 and camera 2140. For example, the processor may be configured to determine that the traversable world quality criteria are no longer met, a predetermined time interval has passed since the last time an image output by camera 2140 was acquired and / or used, and / or an object in a portion of the physical world currently in the field of view of DVS camera 2120 and camera 2140 has changed.
[0259] After creating a traversable world model, the processor may preferentially use the image information output by the DVS camera 2120 to track the head, object, and / or hand position. As described above, the DVS camera 2120 may be configured to acquire images based on events. The acquired image information may be specific to one or more patches in the image sensor of the DVS camera 2120. Alternatively, the processor may use monocular images output by the DVS camera 2120 or the camera 2140 to track the head, object, and / or hand position. These monocular images may be full-frame images and may be repeatedly or periodically output by the DVS camera 2120. For example, the DVS camera may be configured to output images at a frequency between 30 Hz and 120 Hz (e.g., at 60 Hz), while changes may be output at a higher effective rate, such as hundreds or thousands of times per second. In some embodiments, the processor may use stereoscopic images output by both the DVS camera 2120 and the camera 2140 to track the head, object, and / or hand position. These stereoscopic images may be output repeatedly or periodically.
[0260] The processor may be configured to acquire light field information, such as angle of arrival information of light incident on the image sensor, using a plenoptic camera. In some embodiments, the plenoptic camera may be at least one of the DVS camera 2120 or the camera 2140. Consistent with the disclosed embodiments, when depth information is described herein or may enhance processing, such depth information may be determined from or supplemented by light field information obtained by the plenoptic camera. For example, when the camera 2140 includes a TDM filter, the processor may be configured to create a traversable world model using images obtained from the DVS cameras 2120 and 2140 and light field information obtained from the camera 2140. Alternatively or additionally, when the camera DVS 2120 includes a TDM filter, the processor may use light field information obtained from the DVS camera 2120.
[0261] The processor may be configured to detect conditions where one or more types of image information are unavailable or insufficient to provide resolution for a particular function (e.g., world reconstruction, object tracking, or head posture determination). The processor may be configured to select additional or alternative sources of image information to provide sufficient image information for the function. These sources may be selected in an order that results in obtaining suitable image information with a low processing burden in each case. For example, the DVS camera 2120 may be configured to output image data that meets the intensity variation criteria, but when an object with uniform visual characteristics fills the field of view 2121, the pixel intensity may not change sufficiently to trigger image data acquisition. This may be the case, for example, when the user's hand fills the field of view even when the user's hand moves. In some embodiments, even partial filling of the field of view 2121 may inhibit or prevent the acquisition of appropriate image data. For example, when the blocks tracked by the processor are filled with images of objects with uniform appearance positions, the image data for these blocks may not be sufficient to track the object. Similarly, if a close object fills the camera's field of view, or at least fills a patch used to track points of interest, such as points of interest tracked for tracking head pose, then image information from the camera may not be sufficient for head tracking functionality.
[0262] However, given Fig.21 2100, the camera 2140 can be positioned to output images of objects that fill the relevant portion of the field of view. Alternatively or additionally, these images can be used to determine the depth of the object. The camera 2140 can also be positioned to image the points of interest that are tracked within the tiles used for head pose tracking.
[0263] Thus, the processor may be configured to determine whether the object meets the filling criteria of the field of view 2121. Based on the determination, the processor may be configured to enable the camera 2140 or increase the frame rate of the camera 2140. The processor may then receive at least image data from the camera 2140. In some embodiments, the image data may include light field information. For object tracking, the processor may be configured to determine depth information using the received image data. The processor may use the determined depth information to track occluding objects. For head pose tracking, the processor may be configured to determine changes in head pose using the received image data. When the DVS camera 2120 is configured for block tracking of points of interest, the processor may be configured to use the camera 2120 to track points of interest. The processor may be configured to disable or reduce the frame rate of the camera 2140 based on a determination that the filling criteria are no longer met, resuming the use of data that can be acquired faster and / or processed with lower latency, but still sufficient for the function being performed (e.g., object tracking or head pose tracking). This may allow the DVS camera 2120 to track images with an ultra-high temporal resolution: for example, the DVS camera 2120 may output information indicating the current position of a moving object at a frequency higher than 60 Hz, such as thousands or hundreds of times per second.
[0264] It should be understood that the processor can enable or disable the camera so as to dynamically provide different image information sources for one or more functions in any one or more ways based on the detected operating conditions. The processor can send a control signal to the underlying image sensor hardware to change the operation of the hardware. Alternatively or additionally, the processor can enable the camera by reading the image information generated by the camera, or disable the camera by not accessing or using the image information generated by the camera. These techniques can be used to enable or suppress image information in whole or in part. For example, the processor can be configured to perform a size reduction routine to adjust the image output using the camera 2140. The camera 2140 can generate an image larger than the DVS camera 2120. For example, the camera 2140 can generate an image of 12 megapixels and the DVS camera 2120 can generate a 1 megapixel image. The image generated by the camera 2140 may include more information than is required for performing a traversable world model creation, head tracking, or object tracking operation. Processing this additional information may require additional power, thereby shortening battery life or increasing latency. Therefore, the processor can be configured to discard or combine pixels output by the camera 2140 in the image. For example, the processor may be configured to output an image having one sixteenth the number of pixels as the original generated image. Each pixel in the output image may have a value based on a corresponding 4x4 pixel group in the original generated image (e.g., the average value of the values of the 16 pixels).
[0265] According to some embodiments, the XR system may include a hardware accelerator. The hardware accelerator may be implemented as an application specific integrated circuit (ASIC) or other semiconductor device and may be integrated within the head mounted device 2100 or otherwise coupled to the head mounted device 2100 so that it receives image information from the camera 2120 and the camera 2140. The hardware accelerator may use the images output by the two world cameras to assist in stereoscopically determining depth information. The image from the camera 2120 may be a grayscale image and the image from the camera 2140 may be a color image. Using hardware acceleration may speed up the determination of depth information and reduce power consumption, thereby extending battery life.
[0266] Example Calibration Process
[0267] Fig. 22 A simplified flow chart of a calibration routine (method 2200) according to some embodiments is shown. The processor may be configured to perform the calibration routine while wearing the wearable display system. The calibration routine may account for deformations caused by the lightweight structure of the head mounted device 2100. For example, the processor may repeatedly perform the calibration routine so that the calibration routine compensates for deformations of the frame 2101 during use of the wearable display system. The compensation routine may be performed automatically or in response to manual input (e.g., a user request to perform the calibration routine). The calibration routine may include determining the relative position and orientation of the camera 2120 and the camera 2140. The processor may be configured to perform the calibration routine using images output by the DVS camera 2120 and the camera 2140. In some embodiments, the processor may be configured to also use the outputs of the IMU 2170a and the IMU 2170b.
[0268] After starting in box 2201, method 2200 may proceed to box 2210. In box 2210, the processor may identify corresponding features in the images output from the DVS camera 2120 and the camera 2140. The corresponding features may be parts of objects in the physical world. In some embodiments, for the purpose of calibration, the object may be placed by the user within the central field of view 2150 and may have features that are easily identifiable in the image, which may have predetermined relative positions. However, the calibration techniques described herein may be performed based on features on the object that appear in the central field of view 2150 at the time of calibration, so that calibration can be repeated during use of the head mounted device 2100. In various embodiments, the processor may be configured to automatically select features detected within the field of view 2121 and the field of view 2141. In some embodiments, the processor may be configured to determine the correspondence between the features using the estimated positions of the features within the field of view 2121 and the field of view 2141. The estimation may be based on a traversable world model constructed for an object containing these features or other information about these features.
[0269] The method 2200 may proceed to block 2230. In block 2230, the processor may receive inertial measurement data. The inertial measurement data may be received from the IMU 2170a and / or the IMU 2170b. The inertial measurement data may include inclination and / or acceleration and / or velocity measurements. In some embodiments, the IMUs 2170a and 2170b may be mechanically coupled directly or indirectly to the camera 2140 and the DVS camera 2120, respectively. In such embodiments, the difference in inertial measurements (e.g., inclination) made by the IMUs 2170a and 2170b may indicate a difference in the position and / or orientation of the camera 2140 and the camera 2120. Therefore, the outputs of the IMUs 2170a and 2170b may provide a basis for an initial estimate of the relative position of the camera 2140 and the DVS camera 2120.
[0270] After block 2230, method 2200 may proceed to block 2250. In block 2250, the processor may calculate an initial estimated relative position and orientation of the DVS camera 2120 and the camera 2140. The initial estimate may be calculated using measurements received from the IMU 2170b and / or the IMU 2170a. In some embodiments, for example, the head mounted device may be designed to have a nominal relative position and orientation of the DVS camera 2120 and the camera 2140. The processor may be configured to attribute differences in measurements received between the IMU 2170a and the IMU 2170b to deformations of the frame 2101, which may change the position and / or orientation of the DVS camera 2120 and the camera 2140. For example, the IMU 2170a and the IMU 2170b may be mechanically coupled directly or indirectly to the frame 2101 so that the tilt and / or acceleration and / or velocity measurements of these sensors have a predetermined relationship. This relationship may be affected when the frame 2101 is deformed. As a non-limiting example, IMU 2170a and IMU 2170b can be mechanically coupled to frame 2101 so that these sensors measure similar inclination, acceleration, or velocity vectors during movement of the head mounted device when there is no deformation of frame 2101. In this non-limiting example, a twist or bend that rotates IMU 2170a relative to IMU 2170b can result in a corresponding rotation of the inclination, acceleration, or velocity vector measurement of IMU 2170a relative to the corresponding vector measurement of IMU 2170b. Since IMU 2170a and IMU 2170b are mechanically coupled to DVS camera 2140 and camera 2120, respectively, the processor can adjust the nominal relative position and orientation of DVS camera 2120 and camera 2140 consistent with the measurement relationship between IMU 2170a and IMU 2170b.
[0271] Other techniques may alternatively or additionally be used to make the initial estimate.In embodiments where calibration method 2200 is performed repeatedly during operation of the XR system, the initial estimate may be, for example, the most recently calculated estimate.
[0272] After block 2250, a subprocess is initiated in which an estimate of the relative position and orientation of camera 2120 and camera 2140 is also made. One of the estimates is selected as the relative position and orientation of camera 2120 and camera 2140 to calculate stereo depth information based on the images output by camera 2120 and camera 2140. This subprocess may be performed iteratively, making additional estimates in each iteration until an acceptable estimate is determined. Fig. 22 In the example of , the subprocess includes boxes 2270, 2272, 2274 and 2290.
[0273] In block 2270, the processor may calculate the error of the estimated relative orientation of the camera and the compared feature. In calculating the error, the processor may be configured to estimate how the identified feature should appear or where the identified feature should be located within the corresponding image output by camera 2120 and camera 2140 based on the estimated relative orientation of DVS camera 2120 and camera 2140 and the estimated position feature used for calibration. In some embodiments, the estimate may be compared with the appearance position or apparent position of the corresponding feature in the image output by each of the two cameras to generate an error for each of the estimated relative orientations. Such errors may be calculated using linear algebra techniques. For example, the mean square deviation between the calculated position and the actual position of each of the plurality of features within the image may be used as a measure of the error.
[0274] After block 2270, method 2200 may proceed to block 2272, where it may be checked whether the error meets acceptance criteria. For example, the criterion may be the overall magnitude of the error, or may be the error variation between iterations. If the error meets the acceptance criteria, method 2200 proceeds to block 2290.
[0275] In block 2290, the processor may select one of the estimated relative orientations based on the error calculated in block 2272. The selected estimated relative position and orientation may be the estimated relative position and orientation with the lowest error. In some embodiments, the processor may be configured to select the estimated relative position and orientation associated with the lowest error as the current relative position and orientation of the DVS camera 2120 and the camera 2140. After block 2290, the method 2200 may proceed to block 2299. The method 2200 may end in block 2299, where the selected positions and orientations of the DVS camera 2120 and the camera 2140 are used to calculate stereoscopic image information based on the images formed by those cameras.
[0276] If the error does not meet the acceptance criteria at box 2272, the method 2200 can proceed to box 2274. At box 2274, the estimates used to calculate the error at box 2270 can be updated. These updates can be the estimated relative positions and / or orientations of camera 2120 and camera 2140. In an embodiment where the relative positions of a set of features used for calibration are estimated, the updated estimates selected at box 2274 can alternatively or additionally include updates to the positions or locations of the features in the set. Such updates can be made based on linear algebra techniques for solving systems of equations with multiple variables. As a specific example, one or more of the estimated positions or orientations can be increased or decreased. If the change reduces the calculated error in one iteration of the subprocess, the same estimated position or orientation can be further changed in the same direction in subsequent iterations. Conversely, if the change increases the error, those estimated positions or orientations can be changed in the opposite direction in subsequent iterations. The estimated positions and orientations of the cameras and features used in the calibration process can be changed sequentially or in combination in this manner.
[0277] Once the updated estimate is calculated, the subprocess returns to block 2270. Here, further iterations of the subprocess begin, calculating the error of the estimated relative position. In this way, the estimated position and orientation are updated until an updated relative position and orientation that provides an acceptable error is selected. However, it should be understood that the processing at block 2272 can apply other criteria to end the iterative subprocess, such as completing multiple iterations without finding an acceptable error.
[0278] Although method 2200 is described in conjunction with DVS camera 2120 and camera 2140, similar calibration may be performed for any pair of cameras used for stereoscopic imaging or for any camera group for which relative position and orientation are desired.
[0279] Example Camera Configuration
[0280] The head mounted device 2100 incorporates components that provide a field of view and field of illumination to support the various functions of the XR system. FIG. 23A to FIG. 23C According to some embodiments Fig.21 21. Example diagrams of fields of view or illumination associated with a head mounted device 2100. Each example diagram shows the field of view or illumination from a different orientation and distance relative to the head mounted device. Fig.23A The field of view or illumination is shown from elevated off-axis angles at a distance of 1 meter from the head-mounted device. Fig.23AThe overlap between the fields of view of the DVS camera 2120 and the camera 2140 is shown, and in particular how the DVS camera 2120 and the camera 2140 are angled so that the fields of view 2121 and the field of view 2141 pass through the centerline of the head mounted device 2100. In the illustrated configuration, the field of view 2141 extends beyond the field of view 2121 to form a peripheral field of view 2160b. As shown, the illumination fields of the emitters 2130a and 2130b overlap to a large extent. In this way, the emitters 2130a and 2130b can be configured to support imaging or depth measurement of objects in the central field of view 2150 under conditions of low ambient light. Fig. 23B The field of view or illumination is shown from a top-down perspective at a distance of 0.3 meters from the head-mounted device. Fig. 23B The overlap of field of view 2121 and field of view 2141 is shown at 0.3 meters from the head mounted device. However, in the illustrated configuration, field of view 2141 does not extend very far beyond field of view 2121, limiting the extent of peripheral field of view 2160b and exhibiting an asymmetry between peripheral field of view 2160a and peripheral field of view 2160b. Fig.23C The field of view or illumination is shown from a front-view perspective at a distance of 0.25 meters from the head-mounted device. Fig.23C The overlap of field of view 2121 and field of view 2141 that exist at 0.25 meters from the head mounted device is shown. However, field of view 2141 is completely contained within field of view 2121, so the peripheral field of view 2160b does not exist at this distance from the head mounted device 2100 in the illustrated configuration.
[0281] from FIG. 23A to FIG. 23C As can be appreciated, the overlap of fields of view 2121 and 2141 creates a central field of view in which stereoscopic imaging techniques may be employed using images output by cameras 2120 and 2140, with or without IR illumination from emitters 2130a and 2130b. In the central field of view, color information from camera 2140 may be combined with grayscale image information from camera 2120. Additionally, there are non-overlapping peripheral fields of view, but monocular grayscale image information or color image information is available from camera 2120 or camera 2140, respectively. As described herein, different operations may be performed on the image information acquired for the central field of view and the peripheral fields of view.
[0282] World Model Generation
[0283] In some embodiments, image data output by DVS camera 2120 and / or camera 2140 may be used to construct or update a world model. Fig.24 is a simplified flow chart of a method 2400 for creating or updating a traversable world model according to some embodiments. Fig.21As disclosed, the XR system may use a processor to determine and update a traversable world model. In some embodiments, the DVS camera 2120 may be configured to operate to output image information representing the intensity level detected at each of a plurality of pixels or to output image information indicating pixels where an intensity change exceeding a threshold has been detected. In some embodiments, the processor may determine and update the traversable world model based on the output of the DVS camera 2120 representing the detected intensity, which may be combined with the image information output from the camera 2140 to stereoscopically determine the position of an object in the traversable world. In some embodiments, the output of the DVS camera 2120 representing the detected intensity change may be used to identify areas of the world model to be updated based on changes in image information from those locations.
[0284] In some embodiments, the DVS camera 2120 may be configured to output image information reflecting intensity using a global shutter, while the camera 2140 may be configured to have a rolling shutter. The processor may therefore perform a compensation routine to compensate for rolling shutter distortion in the image output by the camera 2140. In various embodiments, the processor may determine and update the traversable world model without using the transmitters 2130a and 2130b. However, in some embodiments, the traversable world model may be incomplete. For example, the processor may incompletely determine the depth of a wall or other flat surface. As an additional example, the traversable world model may incompletely represent objects with many corners, curved surfaces, transparent surfaces, or large surfaces, such as windows, doors, balls, tables, etc. The processor may be configured to identify such incomplete information, obtain additional information, and update the world model using additional depth information. In some embodiments, the transmitters 2130a and 2130b may be selectively enabled to collect additional image information from which to construct or update the traversable world model. In some scenarios, the processor can be configured to perform object recognition in the acquired image, select a template for the recognized object, and add information to the traversable world model based on the template. In this way, the wearable display system can improve the traversable world model while using little or no power-intensive components such as transmitters 2130a and 2130b, thereby extending battery life.
[0285] Method 2400 may be initiated one or more times during operation of the wearable display system. The processor may be configured to create a navigable world model when the user first turns on the system, moves to a new environment (e.g., walks into another room), or generally when the processor detects a change in the user's physical environment. Alternatively or additionally, method 2400 may be performed periodically during operation of the wearable display system or when a significant change in the physical world is detected or in response to user input (e.g., input indicating that the world model is out of sync with the physical world).
[0286] In some embodiments, all or part of a traversable world model may be stored that is provided by other users of the XR system or otherwise obtained. Thus, while the creation of a world model is described, it should be understood that method 2400 may be used for a portion of a world model while other portions of the world model are derived from other sources.
[0287] In block 2405, the processor may perform a compensation routine to compensate for rolling shutter distortion in the image output by camera 2140. As described above, an image acquired by an image sensor with a global shutter (e.g., an image sensor in DVS camera 2120) includes pixel values acquired at the same time. In contrast, an image acquired by an image sensor with a rolling shutter includes pixel values acquired at different times. During the acquisition of the image by camera 2140, the relative motion of the head mounted device and the environment may introduce spatial distortions into the image. These spatial distortions may affect the accuracy of methods that rely on comparing images acquired by camera 2140 with images acquired by DVS camera 2120.
[0288] Performing the compensation routine may include using a processor to compare an image output by the DVS camera 2120 with an image output by the camera 2140. The processor performs the comparison to identify any distortion in the image output by the camera 2140. Such distortion may include skew in at least a portion of the image. For example, if the image sensor in the camera 2140 is acquiring pixel values row by row from the top of the image sensor to the bottom of the image sensor, and the head mounted device 2100 is being translated sideways, the position where an object or a portion of an object appears may be offset by a certain amount in consecutive pixel rows, the offset depending on the speed of the translation and the time difference between acquiring each row of values. Similar distortions may occur when rotating the head mounted device. These distortions may result in an overall skew in the position and / or orientation of an object or a portion of an object in the image. The processor may be configured to perform a row by row comparison between the image output by the camera 2120 and the image output by the camera 2140 to determine the amount of skew. The image output by the camera 2140 may then be transformed to remove the distortion (e.g., to remove the detected skew).
[0289] In block 2410, a traversable world model may be created. In the illustrated embodiment, the processor may use the DVS camera 2120 and the camera 2140 to create the traversable world model. As described above, when generating the traversable world model, the processor may be configured to use the images output by the DVS camera 2120 and the camera 2140 to stereoscopically determine depth information of objects in the physical world when constructing the traversable world model. In some embodiments, the processor may receive color information from the camera 2140. The color information may be used to distinguish objects or identify surfaces associated with the same object. Color information may also be used to identify objects. As described above with respect to Figure 3A As disclosed, the processor may create a traversable world model by associating information about the physical world with information about the position and orientation of the head mounted device 2100. As a non-limiting example, the processor may be configured to determine the distance from the head mounted device 2100 to a feature in a view (e.g., field of view 2121 and / or field of view 2141). The processor may be configured to estimate the current position and orientation of the view. The processor may be configured to accumulate such distance and position and orientation information. By triangulating the distance to the feature acquired from multiple positions and orientations, the position and orientation of the feature in the environment may be determined. In various embodiments, the traversable world model may be a combination of a raster image, a cloud of points and descriptors, and a polygon / geometric definition that describes the position and orientation of such a feature in the environment. In some embodiments, the distance from the head mounted device 2100 to the feature in the central field of view 2150 may be determined stereoscopically using image data output by the DVS camera 2120 and a compensated image generated using image data output by the camera 2140. In various embodiments, light field information may be used to supplement or refine this determination. For example, through calculation, the angle of arrival information can be converted into distance information indicating the distance to the object from which the light was reflected.
[0290] In block 2415, the processor may disable the camera 2140 or reduce the frame rate of the camera 2140. For example, the frame rate of the camera 2140 may be reduced from 30 Hz to 1 Hz. As disclosed above, the color camera 2140 may consume more power than the DVS camera 2120 (grayscale camera). By disabling or reducing the frame rate of the camera 2140, the processor may reduce power consumption and extend the battery life of the wearable display system. Thus, the processor may disable the camera 2140 or reduce the frame rate of the camera 2140 to save power. This lower power state may be maintained until a condition is detected indicating that an update in the world model may be required. Such a condition may be detected based on the passage of time or input, such as from a sensor collecting information about the user's surroundings or from the user.
[0291] Alternatively or additionally, once the traversable world model is sufficiently complete, it is sufficient to be able to use images or image patches acquired using camera 2120 to determine the location and orientation of features in the physical environment. As non-limiting examples, the traversable world model can be identified as sufficiently complete based on the percentage of space around the user's location represented in the model or based on the amount of new image information that matches the traversable world model. For the latter approach, newly acquired images can be associated with locations in the traversable world. If the features in these images have features that match features identified as landmarks in the traversable world model, the world model can be considered complete. Coverage or matching does not have to be 100% complete. Instead, for each criterion, an appropriate threshold can be applied, such as greater than 95% coverage or greater than 90% of features matching previously identified landmarks. Regardless of how the traversable world model is determined to be complete, once completed, the processor can use the existing traversable world information to refine the estimate of the location and orientation of features in the physical world. This process may reflect the assumption that features in the physical world are changing position and / or orientation slowly, if at all, compared to the rate at which the processor processes images output by the DVS camera 2120 .
[0292] In block 2420, after creating the traversable world model, the processor may identify surfaces and / or objects for updating the traversable world model. In some embodiments, the processor may use a grayscale image or image block output by the DVS camera 2120 to identify such surfaces or objects. For example, once a world model is created at block 2410 that indicates a surface at a particular location within the traversable world, the grayscale image or image block output by the DVS camera 2120 may be used to detect surfaces having substantially the same features, and determine that the traversable world model should be updated by updating the location of the surface within the traversable world model. For example, a surface having substantially the same shape as a surface in the traversable world model at substantially the same location may be equivalent to the surface in the traversable world model, and the traversable world model may be updated accordingly. As another example, the location of an object represented in the traversable world model may be updated based on a grayscale image or image block output by the DVS camera 2120. As described herein, the DVS camera 2120 may detect events associated with a grayscale image or more image blocks. Upon detecting such an event, the DVS camera 2120 may be configured to update the traversable world model using the grayscale image or image patches.
[0293] In some embodiments, the processor may use light field information obtained from the camera 2120 to determine depth information of objects in the physical world. For example, angle of arrival information may be used to determine depth information. This determination may be more accurate for objects in the physical world that are closer to the head mounted device 2100. Therefore, in some embodiments, the processor may be configured to use light field information to update only the portion of the traversable world model that meets the depth criteria. The depth criteria may be based on a maximum distinguishable distance. For example, the processor may not be able to distinguish objects at different distances from the head mounted device 2100 when these distances exceed a threshold distance. The depth criteria may be based on a maximum error threshold. For example, the error in estimating the distance may increase as the distance increases, and a specific distance corresponds to a maximum error threshold. In some embodiments, the depth criteria may be based on a minimum distance. For example, the processor may not be able to accurately determine the distance information of an object within a minimum distance from the head mounted device 2100 (such as 15 cm). Therefore, portions of the world model that are greater than 16 cm from the head mounted device may meet the depth criteria. In some embodiments, the traversable world model may be composed of three-dimensional voxel "bricks". In such an embodiment, updating the traversable world model may include identifying a block of voxels for updating. In some embodiments, the processor may be configured to determine a viewing cone. The viewing cone may have a maximum depth, such as 1.5 m. The processor may be configured to identify a block within the viewing cone. The processor may then update the traversable world information for the voxels within the identified block. In some embodiments, the processor may be configured to update the traversable world information for the voxels using the light field information acquired in step 2450, as described herein.
[0294] The process for updating the world model can be different based on whether the object is in the central field of view or the peripheral field of view. For example, updates can be performed on surfaces detected in the central field of view. In the peripheral field of view, for example, updates can be performed only for objects for which the processor has a model, so that the processor can confirm that any updates to the traversable world model are consistent with the object. Alternatively or additionally, new objects or surfaces can be identified based on processing of grayscale images. Even if such processing results in a less accurate representation of the object or surface compared to the processing at box 2410, in some scenarios, a better overall system can be produced by balancing the accuracy of faster and less power-consuming processing. In addition, by periodically repeating method 2400, information of lower accuracy can be periodically replaced by information of higher accuracy so as to replace portions of the world model generated using only a monocular grayscale image with portions generated stereoscopically using the color camera 2140 in combination with the DVS camera 2120.
[0295] In some embodiments, the processor may be configured to determine whether the updated world model meets the quality criteria. When the world model meets the quality criteria, the processor may continue to update the world model with the camera 2140 disabled or with a reduced frame rate. When the updated world model does not meet the quality criteria, the method 2400 may enable the camera 2140 or increase the frame rate of the camera 2140. The method 2400 may also return to step 2410 and recreate a traversable world model.
[0296] In box 2425, after updating the traversable world model, the processor can identify whether the traversable world model includes incomplete depth information. Incomplete depth information can appear in any of a variety of ways. For example, some objects do not produce detectable structures in the image. For example, very dark areas in the physical world may not be imaged with sufficient resolution to extract depth information from images acquired using ambient lighting. As another example, a window or glass tabletop may not appear in a visible image or be recognized by computer processing. As another example, a large uniform surface, such as a tabletop or a wall, may lack sufficient features that can be associated in two stereo images to achieve stereo image processing. As a result, the processor may not be able to use stereo processing to determine the location of such an object. In these scenarios, there will be "holes" in the world model because the process of seeking to use a traversable world model to determine the distance to the surface in a specific direction through the "hole" will not be able to obtain any depth information.
[0297] When the traversable world model does not include incomplete depth information, the method 2400 may return to updating the traversable world model using grayscale images or image patches obtained from the DVS camera 2120 .
[0298] After identifying the incomplete depth information, the processor controlling the method 2400 may take one or more actions to obtain additional depth information. The method 2400 may proceed to box 2431, box 2433, and / or box 2435. In box 2431, the processor may enable the emitter 2130a and / or the emitter 2130b. As described above, one or more of the camera 2120 and the camera 2140 may be configured to detect light emitted by the emitter 2130a and / or the emitter 2130b. The processor may then acquire depth information by causing the emitter 2130a and / or 2130b to emit light that can enhance the image of the object acquired in the physical world. For example, when the DVS camera 2120 and the camera 2140 are sensitive to the emitted light, the images output by the DVS camera 2120 and the camera 2140 may be processed to extract stereo information. When the emitter 2130a and / or the emitter 2130b are enabled, other analysis techniques may be used alternatively or additionally to obtain depth information. In some embodiments, time-of-flight measurements and / or structured light techniques may alternatively or additionally be used.
[0299] In box 2433, the processor may determine additional depth information from previously acquired depth information. In some embodiments, for example, the processor may be configured to identify objects in an image formed using the DVS camera 2120 and / or the camera 2140, and fill any holes in the traversable world model based on the model of the identified object. For example, the process may detect a flat surface in the physical world. The flat surface may be detected using existing depth information acquired using the DVS camera 2120 and / or the camera 2140 or depth information stored in the traversable world model. The flat surface may be detected in response to determining that a portion of the world model includes incomplete depth information. The processor may be configured to estimate additional depth information based on the detected flat surface. For example, the processor may be configured to extend the identified flat surface through an area of incomplete depth information. In some embodiments, the processor may be configured to interpolate lost depth information based on the surrounding portion of the traversable world model when extending the flat surface.
[0300] In some embodiments, as an additional example, the processor may be configured to detect an object in a portion of the world model that includes incomplete depth information. In some embodiments, the detection may involve using a neural network or other machine learning tool to identify the object. In some embodiments, the processor may be configured to access a database storing templates and select an object template corresponding to the identified object. For example, when the identified object is a window, the processor may be configured to access a database storing templates and select a corresponding window template. As a non-limiting example, the template may be a three-dimensional model representing a class of objects, such as types of windows, doors, balls, etc. The processor may configure an instance of the object template based on an image of the object in the updated world model. For example, the processor may scale, rotate, and translate the template to match the detected position of the object in the updated world model. Additional depth information may then be estimated based on the boundaries of the configured template, representing the surface of the identified object.
[0301] In block 2435, the processor may acquire light field information. In some embodiments, the light field information may be acquired with the image and may include angle of arrival information. In some embodiments, the camera 2120 may be configured as a plenoptic camera to acquire the light field information.
[0302] After box 2431, box 2433 and / or box 2435, method 2400 can proceed to box 2440. In box 2490, the processor can use the additional depth information obtained in box 2431 and / or box 2473 to update the traversable world model. For example, the processor can be configured to blend additional depth information obtained from measurements made using active IR illumination into the existing traversable world model. Similarly, additional depth information determined from light field information, such as using triangulation based on angle of arrival information, can be blended into the existing traversable world model. As additional examples, the processor can be configured to blend interpolated depth information obtained by extending a detected flat surface into the existing traversable world model, or to blend additional depth information estimated based on the boundaries of a configuration template into the existing traversable world model.
[0303] The information may be blended in one or more ways, depending on the nature of the additional depth information and / or the information in the traversable world model. For example, blending may be performed by adding to the traversable world model additional depth information collected for locations in the traversable world model where holes exist. Alternatively, the additional depth information may overwrite information at corresponding locations in the traversable world model. As yet another alternative, blending may involve selecting between information already in the traversable world model and the additional depth information. Such a selection may, for example, be based on selecting depth information already in the traversable world model, or depth information in the additional depth information that represents the surface closest to the camera used to collect the additional depth information.
[0304] In some embodiments, a traversable world model may be represented by a mesh of connected points. Updating the world model may be accomplished by computing a mesh representation of an object or surface to be added to the world model, and then combining that mesh representation with the mesh representation of the world model. The inventors have recognized and appreciated that performing the processing in this order may require less processing than adding an object or surface to the world model and then computing a mesh to update the model.
[0305] Fig.24 It is shown that the world model can be updated at blocks 2420 and 2440. The processing of each block can be performed in the same manner, such as by generating a mesh representation of the object or surface to be added to the world model and combining the generated mesh with the mesh of the world model, or in a different manner. In some embodiments, this merging operation can be performed once for both the objects or surfaces identified at blocks 2420 and 2440. For example, this combining process can be performed as described in conjunction with block 2440.
[0306] In some embodiments, the method 2400 may loop back to block 2420 to repeat the process of updating the world model based on information acquired using the DVS camera 2120. Since the process of block 2420 may be performed on fewer and smaller images than the process of block 2410, it may be repeated at a higher rate. The process may be performed at a rate of less than 10 times per second, for example, between 3 and 7 times per second.
[0307] Method 2400 may be repeated in this manner until an end condition is detected. For example, method 2400 may be repeated for a predetermined time period until user input is received, or until a change of a particular type or magnitude in a portion of the physical world model is detected in the field of view of a camera of the head mounted device 2100. Method 2400 may then terminate in box 2499. Method 2400 may be initiated again to capture new information for the world model at box 2405, including information acquired with a higher resolution color camera. Method 2400 may be terminated and restarted to repeat the processing at box 2405 using a color camera, thereby creating a portion of the world model at an average rate that is slower than the rate at which the world model is updated based solely on grayscale image information. For example, the processing using the color camera may be repeated at an average rate of once per second or slower.
[0308] Head pose tracking
[0309] The XR system can track the position and orientation of the head of a user wearing the XR display system. Determining the user's head pose enables information in a traversable world model to be converted to the reference frame of the user's wearable display device so that the information in the traversable world model can be used to render objects on the wearable display device. Because the head pose is frequently updated, performing head pose tracking using only the DVS camera 2120 can provide energy savings, reduced computation, or other benefits. This tracking can be event-based and can obtain full image and / or image tile data, as described above with respect to Figures 4 to 16 The XR system can therefore be configured to disable or reduce the frame rate of the color camera 2140 as needed to balance head tracking accuracy against power consumption and computational requirements.
[0310] Fig.25 is a simplified flow chart of a method 2500 for head pose tracking according to some embodiments. The method 2500 may include creating a world model, selecting a tracking method, tracking head pose using the selected method, and evaluating tracking quality. According to the method 2500, the processor may preferentially use image block data acquired based on events by the DVS camera 2120 to track the head pose. If the preferred method proves insufficient, the processor may use full frame images periodically output by the DVS camera 2120 to track the head pose. If the secondary method proves insufficient, the processor may stereoscopically track the head pose using images output by the DVS camera 2120 and the camera 2140.
[0311] In box 2510, the processor may create a traversable world model. In some embodiments, the processor may be configured to create a traversable world model as described above with respect to boxes 2405-2415 of method 2400. For example, the processor may be configured to acquire images from the DVS camera 2120 and the camera 2140. In some embodiments, the processor may compensate for rolling shutter distortion in the camera 2140. The processor may then use the images from the DVS camera 2120 and the compensated images from the camera 2140 to stereoscopically determine the depth of features in the physical world. Using these depths, the processor may create a traversable world model. In some embodiments, after creating the traversable world model, the processor may be configured to disable or reduce the frame rate of the camera 2140. By disabling or reducing the frame rate of the camera 2140 after generating the traversable world model, the XR system may reduce power consumption and computing requirements.
[0312] For example, the processor may select from world model features corresponding to stationary features, as described above in conjunction with Fig.13 Image information indicating the position of a fixed feature relative to a camera mounted on a device worn on a user's head may be used to compute a change in position of the user's head relative to a world model. According to some embodiments, a processor may select a tracking method to track the relative position of the fixed feature that meets quality criteria while requiring low computational effort relative to other tracking methods.
[0313] In block 2520, the processor may select a tracking method. The processor may use events detected by the DVS camera 2120 to prioritize head pose tracking. The processor may continue to use the preferred method while the head pose tracking meets the tracking quality criteria. For example, the difficulty of tracking head pose may depend on the position and orientation of the user's head and the content of the navigable world model. Therefore, in some cases, the processor may or may become unable to track head pose using only the asynchronously acquired image data output by the DVS camera 2120.
[0314] If the processor determines that the head pose tracking provided by the preferred method does not meet the tracking quality criteria, the processor may select a secondary method to track head pose. For example, the processor may select to track head pose using an indication of an event output by the DVS camera 2120 in combination with color information obtained using the camera 2140. If the processor determines that the secondary method does not meet the tracking quality criteria, the processor may select a tertiary method to track head pose. For example, the processor may select to track head pose stereoscopically using images output by the DVS camera 2120 and the camera 2140. When the head pose tracking meets the tracking quality criteria, the processor may continue to use the selected method. Alternatively, the processor may revert to the preferred method after a predetermined duration or number or quantity of head pose updates; or in response to meeting the criteria.
[0315] In each approach, image information can be acquired for the entire field of view of each camera being used. Fig.13 As described, image information may be collected only for image patches corresponding to portions containing the tracked features.
[0316] Fig.25 A first tracking method performed in blocks 2530a and 2540a is shown. In block 2530a, the processor may enable the DVS function of the DVS camera 2120 when the function is not already enabled. The function may be enabled by setting a threshold change in intensity associated with the movement of fixed features selected from the world model. In some embodiments, a block containing those features may also be set. In block 2540a, the processor may track head pose using block data acquired in response to events detected by the DVS camera 2120. In some embodiments, the processor may be configured to calculate real-time or near real-time user head pose based on the block data.
[0317] A secondary tracking method is shown in boxes 2530b and 2540b. In this example, color image information can be used in conjunction with event information to track the relative position of features. In box 2530b, the processor can enable camera 2140 when camera 2140 is not already enabled. Camera 2140 can be enabled to provide images at a rate that is the same as or slower than the rate at which head pose updates are provided. For example, head pose updates can be provided at an average rate between 30 and 60 Hz by using asynchronous event data. Camera 2140 can be enabled to provide frames at a rate less than 30 Hz (e.g., between 5 and 10 Hz).
[0318] Color information can be used to improve the accuracy of tracking fixed features. For example, color information can be used to calculate an updated position of a feature being tracked, which can be determined more accurately than using grayscale events alone. As a result, the position of the tile being tracked may be updated or changed to include other features. Alternatively or additionally, information from camera 2140 can be used to identify alternative features to track. As yet another alternative, color information can allow the relative position of the camera to be calculated based on analyzing surfaces, edges, or larger features than tracked with DVS camera 2120.
[0319] The third level tracking method is shown in boxes 2530c and 2540c. In this example, the third level tracking can be based on stereo information. In box 2530b, the processor can disable the DVS function of the DVS camera 2120 so that the camera 2120 outputs intensity information instead of event information representing intensity changes. In box 2540b, the processor can use images periodically output by the DVS camera 2120 to track head postures, which images can be full-frame images or can be images within a specific block being tracked. In box 2540c, the processor can use stereo image data obtained from images output by the DVS camera 2120 and the camera 2140 to track head postures. For example, the processor can be configured to determine depth information based on the stereo image data. In some embodiments, the processor can be configured to calculate real-time or near real-time user head postures based on these images.
[0320] In block 2550, the processor may evaluate the tracking quality according to a tracking criterion. The tracking criterion may depend on the stability of the estimated head pose, the noise of the estimated head pose, the consistency of the estimated head pose with the world model, or similar factors. As a specific example, the calculated head pose may be compared with other information that may indicate inaccuracy, such as the output of an inertial measurement unit or a model of the range of human head movement, so that errors in the head pose can be identified. The specific tracking criterion may vary depending on the tracking method used. For example, in a method using event-based information, the correspondence between the positions of features may be used, as indicated by the comparison of the event-based output with the position of the corresponding feature in the full-frame image. Alternatively or additionally, the visual uniqueness of the feature relative to its surrounding environment may be used as a tracking criterion. For example, when the field of view is filled with one or more objects, making it difficult to identify the movement of a particular feature, the tracking criterion of the event-based method may show poor tracking. The percentage of the field of view that is blocked is an example of a criterion that can be used. For example, a threshold value greater than 40% can be used as an indication to switch from using an image-based method for head pose tracking. As a further example, the reprojection error can be used as a measure of the quality of head pose tracking. Such a criterion may be calculated by matching features in the acquired image to a previously determined traversable world model. The position of the feature in the image may be related to the position in the traversable world model using a geometric transformation calculated based on the head pose. The deviation between the calculated position and the feature in the traversable world model, expressed as a mean squared error, may thus indicate an error in the head pose, such that the deviation may be used as a tracking criterion.
[0321] In some embodiments, the processor may be configured to calculate an error in the estimated head pose based on the world model. In calculating this error, the processor may be configured to estimate how the world model (or multiple features in the world model) should appear based on the estimated head pose. In some embodiments, the estimate may be compared to the world model (or features in the world model) to generate an error in the estimated head pose. Such an error may be calculated using linear algebra techniques. For example, the mean square error between the calculated position and the actual position of each feature in a plurality of features within an image may be used as a measure of the error. This measure, in turn, may be used as a measure of the quality of head pose tracking.
[0322] After evaluating the head pose tracking quality, method 2500 may return to block 2520 where the processor may use the measured head pose tracking quality to select a tracking method. In scenarios where the tracking quality is low, such as below a threshold, an alternative tracking method may be selected.
[0323] The method 2500 may repeat in this manner until an end condition is detected. For example, the method 2500 may repeat for a predetermined period of time, until a user input is received or until a change of a particular type or magnitude in a portion of the physical world model is detected in the field of view of a camera of the head mounted device 2100. The method 2500 may then terminate in block 2599.
[0324] The method 2500 may be initiated again to capture new information for the world model at block 2510, including information acquired with the higher resolution color camera. The method 2500 may terminate and be restarted to repeat the processing of block 2510, using the color camera to create a portion of the world model at an average rate that is slower than the rate at which head pose tracking is performed in blocks 2520-2550. For example, the processing using the color camera may be repeated at an average rate of once per second or slower.
[0325] Other tracking methods may be used instead of or in addition to the tracking methods described above as examples. In some embodiments, the processor may be configured to compute real-time or near real-time user head pose based on image information, which may include grayscale images and / or light field information, such as angle-of-arrival information. Alternatively or additionally, in scenarios where image-based head pose tracking methods have unacceptable quality metrics, a "dead reckoning" method may be selected, in which the head pose may be computed using the movement of the user's head measured by an inertial measurement unit.
[0326] Object Tracking
[0327] As described above, the processor of the XR system can track objects in the physical world to support realistic rendering of virtual objects relative to physical objects. For example, tracking is described in conjunction with a movable object, such as a hand of a user of the XR system. For example, the XR system can track objects in the central field of view 2150, the peripheral field of view 2160a, and / or the peripheral field of view 2160b. Rapidly updating the position of the movable object enables realistic rendering of the virtual object, because such rendering can reflect the occlusion of the virtual object by the physical object, and vice versa, or the interaction between the virtual object and the physical object. In some embodiments, for example, the update of the position of the physical object can be calculated at an average rate of at least 10 times per second, and, in some embodiments, at least 20 times per second, such as approximately 30 times per second. When the tracked object is a user's hand, tracking can enable gesture control of the user. For example, a particular gesture can correspond to a command of the XR system.
[0328] In some embodiments, the XR system can be configured to track objects having features that provide high contrast when imaged with an image sensor that is sensitive to IR light. In some embodiments, objects with high contrast features can be created by adding markers to the object. For example, a physical object can be equipped with one or more markers that appear as high contrast areas when imaged with IR light. The markers can be passive markers that are highly reflective or highly absorptive to IR light. In some embodiments, at least 25% of light in the frequency range of interest can be absorbed or reflected. Alternatively or additionally, the marker can be an active marker that emits IR light, such as an IR LED. By tracking these features, for example using a DVS camera, information that accurately represents the location of the physical object can be quickly determined.
[0329] As with head pose tracking, the tracked object positions are frequently updated, so performing object tracking using only the DVS camera 2120 may provide power savings, reduced computational requirements, or provide other benefits. The XR system may therefore be configured to disable or reduce the frame rate of the color camera 2140 as needed to balance object tracking accuracy against power consumption and computational requirements. In addition, the XR system may be configured to track objects asynchronously in response to events generated by the DVS camera 2120.
[0330] Fig.26 is a simplified flow chart of a method 2600 for object tracking in accordance with some embodiments. According to the method 2600, the processor may perform object tracking differently depending on the field of view containing the object and the value of the tracking quality criterion. In addition, the processor may or may not use light field information, depending on whether the tracked object meets the depth criterion. The processor may alternatively or additionally apply other criteria to dynamically select the object tracking method, such as available battery power or the operation of the XR system being performed and the need for those operations to track the object position or to track the object position with high accuracy.
[0331] Method 2600 may begin in block 2601. In some embodiments, camera 2140 may be disabled or have a reduced frame rate. The processor may have disabled camera 2140 or reduced the frame rate of camera 2140 to reduce power consumption and extend battery life. In various embodiments, the processor may track an object in the physical world (e.g., a user's hand). The processor may be configured to predict the next position of an object or the trajectory of an object based on one or more previous positions of the object.
[0332] After starting in block 2601, method 2600 may proceed to block 2610. In block 2610, the processor may determine a field of view (e.g., field of view 2121, field of view 2141, peripheral field of view 2160a, peripheral field of view 2160b, or central field of view 2150) that surrounds the object. In some embodiments, the processor may make this determination based on the current position of the object (e.g., whether the object is currently in the central field of view 2150). In various embodiments, the processor may make this determination based on an estimate of the position of the object. For example, the processor may determine that an object that is leaving the central field of view 2150 may enter the peripheral field of view 2160a or the peripheral field of view 2160b.
[0333] In block 2620, the processor may select an object tracking method. According to method 2600, when an object is within field of view 2121 (e.g., within peripheral field of view 2160a or central field of view 2150), the processor may prioritize object tracking using DVS camera 2120. Additionally, the processor may prioritize event-based asynchronous object tracking.
[0334] In some embodiments, block tracking as described above may be used, wherein one or more blocks are established as described above to contain features of the object being tracked. Block tracking may be used for some or all object tracking methods and some or all cameras. A block may be selected to contain an estimated position of the tracked object in the field of view.
[0335] As above combined Fig.25 As described above with respect to head pose tracking, a processor can dynamically select an appropriate tracking method for object tracking. A method can be selected to provide appropriate tracking quality with low processing overhead compared to other methods. Thus, if event-based asynchronous object tracking does not meet the object tracking criteria, the processor can use another method. Fig.26 In the example of FIG. 2600 , four methods are shown in boxes 2640a, 2640b, 2640c, and 2640d. The methods are ranked and method 2600 will select the first method in the order that meets the tracking quality criteria. For example, the methods can be ranked to reflect the trade-offs in accuracy, latency, and power consumption. For example, the method with the lowest latency can be ranked first, and methods with more latency or more power consumption are lower in the ranking.
[0336] In block 2630a, the processor may enable the DVS function of the DVS camera 2120 when the function is not already enabled. In block 2640a, the processor may use the events detected by the DVS camera 2120 to track the object. In some embodiments, the DVS camera 2120 may be configured to limit image acquisition to a tile in the image sensor that contains the location of the object in the field of view 2121. Changes in image data within a tile (e.g., caused by movement of the object) may trigger an event. In response to the event, the DVS camera 2120 may acquire image data for the tile and update the location and / or orientation of the object based on the acquired tile image data.
[0337] In some embodiments, an event may indicate a change in intensity. As described above, for example, in conjunction with Fig.12 and Fig.13 , changes in intensity can be tracked to track the movement of an object. In some embodiments, the DVS camera 2120 may be or may be configured to operate as a plenoptic camera. When an object meets a depth criterion, the processor may be configured to additionally or alternatively acquire angle of arrival information. The depth criterion may be the same or similar to the depth criterion described above with respect to box 2420 of method 2400. For example, the depth criterion may involve a maximum error rate or a maximum distance beyond which the processor cannot distinguish between different distances between objects. Depth information may therefore be used to determine the position of an object and / or changes in the position of an object. Such use of plenoptic image information may be part of the methods described herein or may be other methods that may be used in conjunction with other methods.
[0338] The second method is shown in blocks 2630b and 2640b. In block 2630b, the processor may enable camera 2140 when it is not already enabled. As described above in conjunction with block 2530b, the color camera may be operated to acquire color images at a relatively low average rate. Color information may be used primarily in block 2640b using events output by DVS camera 2120, where information from the color image better identifies the features to be tracked or their locations.
[0339] The third level method is shown in boxes 2630c and 2640c. In box 2630c, when camera 2140 has not been enabled, the processor can enable camera 2140. Camera 2140 is described as being able to output color image information. For the third level method, color image information can be used, but in some embodiments or some scenes, intensity information can be obtained only from camera 2140 or only intensity information can be processed. The processor can also increase the frame rate of camera 2140 to a rate sufficient for object tracking (e.g., a frame rate between 40Hz and 120Hz). In some embodiments, the rate can match the sampling frequency of DVS camera 2120. When the DVS function of DVS camera 2120 has not been disabled, the processor can also disable the function. In this configuration, the output of DVS camera 2120 can represent intensity information. In box 2640c, the processor can track an object using stereoscopic image data obtained from images output by DVS camera 2120 and camera 2140.
[0340] The fourth method is shown in boxes 2630d and 2640d. In box 2630d, the processor can enable the camera when the camera is not already enabled. Other cameras can be disabled. The enabled camera can be a DVS camera 2120 or a camera 2140. If the DVS camera 2120 is used, it can be configured to output image intensity information. When the camera 2140 is enabled, it can be enabled to output color information or only grayscale intensity information. The processor can also increase the frame rate of the enabled camera to a rate sufficient for object tracking (e.g., a frame rate between 40Hz and 120Hz). In box 2640d, the processor can use the image output by the enabled camera to track the object.
[0341] In block 2650, the processor may evaluate the tracking quality. One or more of the metrics described above in connection with block 2550 may be used in block 2650. However, in block 2650, these metrics will be applied to features on the tracked object rather than fixed features in the tracking environment.
[0342] After evaluating the object tracking quality, method 2600 may return to blocks 2610 and 2620, where the processor may again determine the field of view that contains the object (block 2610), and then select an object tracking method using the determined field of view and the measured object tracking quality (block 2620). In the illustrated embodiment, the first method is selected in the order that provides a quality that exceeds a threshold associated with adequate performance.
[0343] Method 2600 may repeat in this manner until an end condition is detected. For example, method 2600 may repeat for a predetermined time period until user input is received, or until the tracked object leaves the field of view of the XR device. Method 2600 may then terminate at block 2699.
[0344] Method 2600 can be used to track any object, including a user's hand. However, in some embodiments, different or additional actions can be performed when tracking a user's hand. Fig. 27 is a simplified flow chart of a hand tracking method 2700 according to some embodiments. The object tracked in method 2700 can be a user's hand. In various embodiments, the XR system can use images obtained from camera 2140 and / or images or image data obtained from DVS camera 2120 to perform hand tracking. These cameras can be configured to operate in one of a variety of modes. For example, one or both can be configured to acquire image segment data. Alternatively or additionally, DVS camera 2140 can be configured to output events, as described above with respect to Figures 4 to 16 Alternatively or additionally, the camera 2140 may be configured to output color information, or only grayscale intensity information. Furthermore, alternatively or additionally, the camera 2140 may be configured to output full light information, as described above in conjunction with Figure 16 to Figure 2 0. In some embodiments, the DVS camera 2120 can be similarly configured to output plenoptic information. Any combination of these cameras and functions can be selected to generate information for hand tracking.
[0345] Combined with Fig.26 As with object tracking described above, the processor can be configured to select a hand tracking method based on the determined hand position and an assessment of the hand tracking quality provided by the selected method. For example, stereo depth information may be obtained when the hand is within the center field of view 2150. As another example, plenoptic information may have sufficient resolution for hand tracking only within an angular range relative to the center of the plenoptic camera's field of view (e.g., ±20 degrees). Thus, techniques that rely on stereo image information or plenoptic image information may be used only when the hand is detected to be within the appropriate field of view.
[0346] If necessary, the XR system can enable camera 2140 or increase the frame rate of camera 2140 to enable tracking in field of view 2160b. In this way, the wearable display system can be configured to provide adequate hand tracking using a reduced number of available cameras in this configuration, thereby allowing for reduced power consumption and increased battery life.
[0347] Method 2700 can be performed under the control of a processor of the XR system. The method can be initiated when an object to be tracked (e.g., a hand) is detected as a result of analyzing an image acquired using any one of the cameras on the head-mounted device 2100. The analysis may require identifying the object as a hand based on an image area having photometric features that are characteristic of a hand. Alternatively or additionally, depth information acquired based on stereo image analysis can be used to detect the hand. As a specific example, the depth information may indicate the presence of an object having a shape that matches a 3D model of a hand. Detecting the presence of a hand in this manner may also require setting parameters of the hand model to match the orientation of the hand. In some embodiments, such a model can also be used for fast hand tracking by using photometric information from one or more grayscale cameras to determine how the hand has moved from an original position.
[0348] Other trigger conditions may initiate method 2700, such as the XR system performing an operation involving a tracked object, such as rendering a virtual button that a user may attempt to press with a hand, in anticipation of the user's hand entering the field of view of one or more cameras. Method 2700 may be repeated at a relatively high rate, such as between 30 and 100 times per second, such as between 40 and 60 times per second. As a result, updated position information of the tracked object may be made available with low latency for use in rendering a process for virtual object interaction with a physical object.
[0349] After starting in block 2701, method 2700 may proceed to block 2710. In block 2710, the processor may determine a potential hand position. In some embodiments, the potential hand position may be the position of an object detected in the acquired image. In embodiments where the hand is detected based on matching depth information to a 3D model of the hand, the same information may be used as the initial position of the hand at block 2710.
[0350] In some embodiments, image information for configuring the hand model may be dynamically selected. After box 2710, method 2700 may proceed to box 2720. In box 2720, the processor may select a processing method for configuring the hand model. The selection may depend on the potential position of the hand and whether the hand model has already been configured. In some embodiments, the selection may further depend on the desired accuracy and / or quality of the hand model and / or the available processing computer power in the case of other tasks being performed by one or more processors. As an example, when initially configuring the hand model upon detection of a hand, more extensive but more accurate image information may be used, but additional processing may be required. For example, stereo information may be used to initially configure the model. Thereafter, the model itself provides information about the position of the hand because there are limitations on the way a human hand can move. Therefore, when the XR system is running, the processing of reconfiguring the hand model to account for hand movement may optionally be performed using less comprehensive but faster processing image information, or when more comprehensive image information is not available. For example, when the object is in the central field of view 2150 and robust hand tracking or fine hand details are required, or if other methods prove insufficient, the processor can use the images output by the DVS camera 2120 and the camera 2140 to stereoscopically perform hand tracking. Alternatively, when the object is not in the central field of view 2150, or robust hand tracking or fine hand details are not required, the processor can be configured to preferentially use other image information to configure the hand model, such as a monocular color image or a grayscale image. When the object is within the peripheral field 2160a, the processor can be configured to use the image acquired from the DVS camera 2120 to configure the hand model. When the object is within the peripheral field 2160b, the processor can be configured to use the image acquired from the camera 2140 to configure the hand model. In some embodiments of some embodiments, for example, the processor can be configured to use the lowest power consumption drain method that still meets the quality standard.
[0351] Fig. 27 Four methods are shown for collecting information to configure a hand model. These four methods are shown in the flow chart as four parallel paths, including paths through boxes 2740a, 2740b, 2740c, and 2740d.
[0352] In the first path, at box 2730a, when a potential hand is in the center field of view 2150 and hand tracking robustness or fine hand details are required, the processor can enable the camera 2140 when it is not already enabled. This path can also be selected for the initial configuration of the hand model at box 2720. At box 2730a, the processor can also increase the frame rate of the camera 2140 to a rate sufficient for hand tracking (e.g., a frame rate between 40 Hz and 120 Hz). In some embodiments, the rate can match the sampling frequency of the DVS camera 2120. The processor can also disable the DVS function of the DVS camera 2120 when it has not been disabled. As a result, both cameras can provide image information representing intensity. The image information can be grayscale, or for cameras that support color imaging, can alternatively include color information.
[0353] In block 2740a, the processor may obtain depth information for the potential hand. The depth information may be obtained based on stereo image analysis, from which the distance between the camera collecting the image information and the segments or features of the potential hand may be calculated. For example, the processor may select a feature in the center field of view and determine the depth information for the selected feature.
[0354] In some embodiments, the selected features can represent different segments of a human hand defined by bones and joints. Feature selection can be based on matching image information with a human hand model. For example, this matching can be performed heuristically. For example, a human hand can be represented by a limited number of segments, such as 16 segments, and points in the image of the hand can be mapped to one of those segments, so that features on each segment can be selected. Alternatively or additionally, this matching can use a deep neural network or a classification / decision forest to apply a series of yes / no decisions in the analysis to identify different parts of the hand and select features representing different parts of the hand. For example, matching can identify whether a specific point in the image belongs to the palm part, the back of the hand, a non-thumb finger, a thumb, a fingertip and / or a knuckle. Any suitable classifier can be used for this analysis stage. For example, a deep learning module or a neural network mechanism can be used instead of a classification forest or as a supplement to a classification forest. In addition, in addition to the classification forest, a regression forest (e.g., using Hough transform, etc.) can also be used.
[0355] Regardless of the specific number of features selected and the techniques used to select those features, after box 2740a, method 2700 can proceed to box 2750a. In box 2750a, the processor can configure a hand model based on the depth information. In some embodiments, the hand model can reflect structural information about a person's hand, such as representing each bone in the hand as a segment in the hand, and each joint defines a possible angle range between adjacent segments. By assigning a position to each segment in the hand model based on the depth information of the selected features, information about the position of the hand can be provided for subsequent processing by the XR system.
[0356] Regardless of how the 3D hand model is updated, the updated model can be refined based on the photometric image information. For example, the model can be used to generate a projection of the hand that represents how an image of the hand is expected to appear. This expected image can be compared to photometric image information acquired with an image sensor. The 3D model can be adjusted to reduce the error between the expected and acquired photometric information. The adjusted 3D model then provides an indication of the position of the hand. As this updating process is repeated, the 3D model provides an indication of the position of the hand as the hand moves.
[0357] In some embodiments, the processing at blocks 2740a and 2750a may be performed iteratively, where the selection of features for which depth information is collected is refined based on the configuration of the hand model. The hand model may include shape constraints and movement constraints, which the processor may be configured to use to refine the selection of features representing portions of the hand. For example, when a feature selected to represent a segment of the hand indicates that the position or movement of the segment violates a hand model constraint, a different feature may be selected to represent the segment.
[0358] In box 2750a, the processor can select features in the blocks or images output by the DVS camera 2120 or the camera 2140 that represent the structure of the human hand. Such features can be identified in a heuristic manner or using AI techniques. For example, features can be heuristically selected by representing the human hand with a limited number of segments and mapping points in the image to corresponding segments in those segments, so that features on each segment can be selected. Alternatively or additionally, this matching can use a deep neural network or a classification / decision forest to apply a series of yes / no decisions in the analysis to identify different parts of the hand and select features representing different parts of the hand. Any suitable classifier can be used for this analysis stage. For example, a deep learning module or a neural network mechanism can be used instead of a classification forest or as a supplement to a classification forest. In addition, in addition to the classification forest, a regression forest (e.g., using a Hough transform, etc.) can also be used. The processor can try to match the selected features and the movement of those selected features between images with the hand model without the help of depth information. This matching may result in less robust information or less accurate information than the information generated in box 2750a. Nonetheless, information based on monocular information recognition can provide useful information for the operation of XR systems.
[0359] In box 2750a, the processor may also evaluate the hand model. The evaluation may depend on the completeness of the match between the selected features and the hand model, the stability of the match between the selected features and the hand model, the noise of the estimated hand position, the consistency between the position and orientation of the detected features and the constraints imposed by the hand model, or similar factors. In some embodiments, the processor may determine where the selected features should appear based on the hand model. As a specific example, the processor may parameterize a general hand model and then check the photometric consistency of the model edges and compare these edges with the edges detected in the acquired hand image. In some embodiments, the estimate may be compared with the estimated position and orientation of the selected features to generate an error of the hand model. Such errors may be calculated using linear algebra techniques. For example, the mean square deviation between the calculated position and the actual position of each of a plurality of selected features within the image may be used as a measure of error. This measure, in turn, may be used as a measure of the quality of the hand model.
[0360] After matching the image portion with the portion of the hand model in box 2750a, the model can be used for one or more operations performed by the XR system. As an example, the model can be used directly as an indication of the location of the object in the physical world for rendering the virtual object. In such an embodiment, the processing at box 2760 can be optionally omitted. Alternatively or additionally, the model can be used to determine whether the user has made a gesture with their hand, such as a gesture indicating a command or interaction with a virtual object. Therefore, in box 2760, the processor can use the determined hand model information to recognize the hand gesture. The gesture recognition can be performed using the hand tracking method described in U.S. Patent Publication No. 2016 / 0026253, which is incorporated herein by reference, and is generally taught in conjunction with hand tracking and hand information obtained using image information from the XR system.
[0361] In some embodiments, gestures can be recognized without stereo depth information. For example, gestures can be recognized based on a hand model configured based on monocular image information. Thus, in some embodiments, gesture recognition can be performed on hands even in peripheral fields of view 2160a and 2160b where stereo depth information is not available. Alternatively or additionally, monocular image information can be used when performing hand tracking for gesture recognition, rather than other functions that may require more precise determination of hand position (e.g., rendering virtual objects to appear to interact realistically with the user's hands). In such embodiments, at boxes 2730a and 2740b, different cameras can be enabled and / or used to collect image information.
[0362] In some embodiments, continuous iterations of the hand tracking process may be performed while the system is operating. For example, a gesture may be identified by continuous determination of the hand position. Alternatively or additionally, continuous iterations may be performed such that a configuration model of the user's hand matches the actual hand position as the hand position changes. The same or different imaging techniques may be used in each iteration.
[0363] In some embodiments, a source of image information may be selected for an iteration at block 2720 based on one or more factors, including the hand model to be used and / or the quality of tracking performed using a particular tracking method. For example, the image information may be selected as described above in conjunction with block 2650 ( Fig.26 ) described herein to determine tracking quality, and as also described above, the selection of a technique may be made that requires the least amount of processing, power consumption, and other computing resources to achieve the desired quality metric.
[0364] therefore, Fig. 27It is shown that once an iteration is performed, method 2700 loops back to block 2720, where a processing method can be selected for further iterations. These alternative methods can use image information instead of stereo depth information or as a supplement to stereo depth information. At each iteration, the 3D model of the hand can be updated to reflect the movement of the hand.
[0365] A method that requires low processing volume is to track the hand position based on event information. This method can be selected at box 2720, for example, in a scene where a hand model of appropriate accuracy has been calculated, for example, by using stereo image information in a previous iteration. This method can be implemented by a branch to box 2730b. In box 2730b, when the DVS function of the DVS camera 2120 is not yet enabled, the processor can enable the function. In box 2740b, the processor can acquire image data in response to an event detected by the DVS camera 2120. In some embodiments, the DVS camera 2120 can be configured to acquire information of a block in the image sensor, which contains a potential hand position in the field of view 2121. Changes in image data within the block (e.g., caused by the movement of the potential hand) can trigger an event. In response to the event, the DVS camera 2120 can obtain image data. This information about the movement of the user's hand can then be used to update the hand model at box 2750b. The update can be performed using the technology described above for box 2750a. This update may take into account other information, including the previously calculated position of the hand indicated by the model and constraints on the motion of the human hand.
[0366] Once updated, the hand model can be used by the system in the same manner as the initial hand model calculated at box 2750a, including determining interaction between virtual objects and the user's hand and / or recognizing gestures at box 2760.
[0367] In some scenarios, event-based image information may not be sufficient to update the hand model. For example, such a scenario may occur if the user's hand fills the field of view of the DVS camera 2120. In such a scenario, the hand model may be updated based on color information rather than event information, such as using color information from the camera 2140. This approach may be implemented by branching to box 2730d when processing at box 2720 detects that the image within the field of view of the DVS camera 2120 has an intensity change below a predetermined threshold or other characteristics indicating a lack of sufficiently distinct features for event-based tracking. At box 2730d, the processor may disable the DVS function of the DVS camera 2120 if it has not already been disabled. The camera 2140 may be enabled to obtain color information.
[0368] In block 2740d, the processor may use the color image information to identify the motion of the user's hand. As described above, this information about the motion of the user's hand may then be used to update the hand model at block 2750b.
[0369] In some scenarios, color information may not be available or required for tracking hand movements. For example, when the hand is in the peripheral field of view 2160a, color information may not be available, but intensity image information obtained using the DVS camera 2120 operating in a mode in which the DVS functionality is disabled may be available and may be appropriate. Alternatively, in some scenarios where event-based tracking produces a quality metric below a threshold, tracking using grayscale image information may produce a quality metric above a suitability threshold. This approach may be implemented by branching to box 2730c. At box 2730c, the processor may enable the camera 2140 when it is not already enabled. The processor may also increase the frame rate of the camera 2140 to a rate sufficient for hand tracking (e.g., a frame rate between 40 Hz and 120 Hz)
[0370] In block 2740d, the processor may acquire image data using camera 2140. In other scenarios, DVS camera 2120 may be configured to collect intensity information and may collect monocular image information in lieu of camera 2140. Regardless of which camera is used, full frame information may be used, or tile tracking as described above may be used to reduce the amount of image information processed.
[0371] Fig. 27 A method is shown in which processing at box 2720 selects between three alternative methods for tracking a hand. The methods may be ranked according to the degree to which one or more criteria are met, such as low processing volume or low power consumption. Processing at box 2720 may select a processing method by selecting a first method in an order that is operable in a detected scene (e.g., hand position) and produces a quality metric that satisfies a threshold. Different or additional processing techniques may be included. For example, for an XR system with a plenoptic camera that provides image information indicating depth, the method may be based on using that depth information alone or in combination with any other data source. As another variation, as described above in conjunction with box 2640b ( Fig.26 ) as described in , color-assisted DVS tracking can be used to track hand features.
[0372] Regardless of the selected method or the set of methods from which such selection is made, the hand model can be updated and used for XR functions, such as rendering virtual objects that interact with the user's hands or detecting gestures at box 2760. After box 2760, method 2700 can end in box 2799. However, it should be understood that hand tracking can occur continuously during the operation of the XR system or can occur during intervals when the hand is in the field of view of one or more cameras. Therefore, once one iteration of method 2700 is completed, another iteration can be performed, and the process can be performed during the interval in which hand tracking is being performed. In some embodiments, information used in one iteration can be used in subsequent iterations. In various embodiments, for example, the processor can be configured to estimate an updated position of the user's hand based on a previously detected hand position. For example, the processor can estimate where the user's hand will be next based on the previous position and velocity of the user's hand. Such information can be used to reduce the amount of image information processed to detect the position of the object, as described above in conjunction with the block tracking technique.
[0373] Having thus described several aspects of some embodiments, it is to be appreciated various alterations, modifications, and improvements will readily occur to those skilled in the art.
[0374] As an example, embodiments are described in conjunction with an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein may be applied to an MR environment or more generally to other XR environments.
[0375] In addition, embodiments of image arrays are described in which a partition is applied to the image array to control the selective output of image information about a movable object. It should be understood that there can be more than one movable object in a physical environment. In addition, in some embodiments, it may be necessary to selectively obtain frequent updates of image information in areas other than the area where the movable object is located. For example, partitions can be set to selectively obtain image information about areas of the physical world where virtual objects are to be rendered. Therefore, some image sensors may be able to selectively provide information from two or more partitions, with or without circuitry for tracking the trajectories of these partitions.
[0376] As yet another example, an image array is described as outputting information related to the amplitude of incident light. The amplitude may be a representation of power across a spectrum of light frequencies. The spectrum may be relatively broad, capturing energy at frequencies corresponding to visible light of any color, such as in a black and white camera. Alternatively, the spectrum may be narrow, corresponding to a single color of visible light. To this end, filters may be used that limit the light incident on the image array to light of a particular color. In the case where pixels are restricted to receiving light of a particular color, different pixels may be restricted to different colors. In such an embodiment, the outputs of pixels sensitive to the same color may be processed together.
[0377] A process is described for setting up tiles in an image array and then updating the tiles for an object of interest. For example, the process can be performed for each movable object as it enters the field of view of an image sensor. When the object of interest leaves the field of view, the tile can be cleared so that the tile is no longer tracked or image information is no longer output for the tile. It should be understood that the tiles can be updated from time to time, such as by determining the position of an object associated with the tile and setting the position of the tile to correspond to that position. Similar adjustments can be made to the calculated trajectory of the tile. The motion vector of the object and / or the motion vector of the image sensor can be calculated based on other sensor information and used to reset values programmed into the image sensor or other components for tile tracking.
[0378] For example, the position, movement, and other features of an object can be determined by analyzing the output of a wide-angle camera or a pair of cameras with stereo information. Data from these other sensors can be used to update the world model. In conjunction with the update, the block position and / or trajectory information can be updated. Such updates can occur at a lower rate than the block tracking engine updates the block position. For example, the block tracking engine can calculate new block positions at a rate between approximately 1 and 30 times per second. Updates to the block position based on other information can occur at a slower rate, such as once per second to approximately once every 30 seconds.
[0379] As yet another example of a variation, Figure 2 A system with a head mounted display separate from a remote processing module is shown. An image sensor as described herein can result in a compact design of the system. Such a sensor generates less data, which in turn results in lower processing requirements and lower power consumption. The reduced demand for processing and power enables size reduction by reducing the size of the battery. Thus, in some embodiments, the entire augmented reality system can be integrated into a head mounted display without a remote processing module. The head mounted display can be configured as a pair of goggles, or as Figure 2 As shown, it can be similar in size and shape to a pair of glasses.
[0380] In addition, embodiments are described where the image sensor is responsive to visible light. It should be understood that the techniques described herein are not limited to operation with visible light. They may alternatively or additionally respond to IR light or "light" in other parts of the spectrum, such as UV. Furthermore, the image sensors as described herein are responsive to naturally occurring light. Alternatively or additionally, the sensor may be used in a system having an illumination source. In some embodiments, the sensitivity of the image sensor may be tuned to the portion of the spectrum where the illumination source emits light.
[0381] As another example, it is described that a selected area of an image array is specified by specifying a "block" on which image analysis is to be performed, and the image sensor should output changes in the selected area. However, it should be understood that the block and the selected area can have different sizes. For example, the selected area can be larger than the block in order to account for the movement of the tracked object in the image deviating from the predicted trajectory and / or to implement processing around the edges of the block.
[0382] In addition, multiple processes are described, such as navigable world model generation, object tracking, head posture tracking, and hand tracking. These and other processes in some embodiments can be performed by the same or different processors. The processor can be operated so that these processes can operate concurrently. However, each process can be executed at a different rate. In the case where different processes request data from an image sensor or other sensor at different rates, the acquisition of sensor data can be managed, for example, by another process, so as to provide data to each process at a rate suitable for its operation.
[0383] Such changes, modifications, and improvements are intended to be part of the present disclosure and are intended to be within the spirit and scope of the present disclosure. For example, in some embodiments, the filter 102 of a pixel of an image sensor may not be a separate component, but may be incorporated into one of the other components of the pixel subarray 100. For example, in an embodiment including a single pixel having an arrival angle to position intensity converter and an optical filter, the arrival angle to intensity converter may be a transmissive optical component formed of a material that filters a specific wavelength.
[0384] According to some embodiments, a wearable display system may be provided, comprising: a frame; a first camera mechanically coupled to the frame, wherein the first camera is configurable to output image data satisfying an intensity variation criterion in a first field of view of the first camera; and a processor operably coupled to the first camera, wherein the processor is configured to determine whether an object is within the first field of view and track movement of the object for one or more portions of the first field of view using image data received from the first camera.
[0385] In some embodiments, the object may be a hand, and tracking the movement of the object may include updating a corresponding portion of the hand model including shape constraints and / or motion constraints based on image data from the first camera that satisfies the intensity change criterion.
[0386] In some embodiments, the processor may be further configured to provide instructions to the first camera to restrict image data acquisition to one or more tiles of the first field of view corresponding to the object.
[0387] In some embodiments, a second camera can be mechanically coupled to the frame to provide a second field of view that at least partially overlaps the first field of view, and the processor can be further configured to: determine whether the object meets occlusion criteria of the first field of view; enable the second camera or increase the frame rate of the second camera; determine depth information of the object using the second camera; and track the object using the determined depth information.
[0388] In some embodiments, depth information may be determined stereoscopically using images output by the first camera and the second camera.
[0389] In some embodiments, depth information may be determined using light field information output by the second camera.
[0390] In some embodiments, the object can be a hand, and tracking the movement of the hand can include using the determined depth information in the following manner: selecting a point in the first field of view; associating the selected point with the depth information; generating a depth map using the selected point; matching portions of the depth map with corresponding portions of the hand model including shape constraints and motion constraints.
[0391] In some embodiments, tracking the motion of the object may include updating the position of the object in the world model, and the duration of the intervals between updates may be between 1 millisecond and 15 milliseconds.
[0392] In some embodiments, a second camera can be mechanically coupled to the frame to provide a second field of view that at least partially overlaps the first field of view; and the processor can be operably coupled to the second camera and further configured to: create a world model using images output by the first camera and the second camera; and update the world model using light field information output by the second camera.
[0393] In some embodiments, the processor may be mechanically coupled to the frame.
[0394] In some embodiments, a display device mechanically coupled to the frame may include a processor.
[0395] In some embodiments, the local data processing module may include a processor, the local data processing module is operably coupled to the display device via a communication link, and the display device is mechanically coupled to the frame.
[0396] According to some embodiments, a wearable display system may be provided, comprising: a frame; two cameras mechanically coupled to the frame, wherein the two cameras include a first camera and a second camera that are configurable to output image data that meets an intensity variation criterion, wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and a processor operably coupled to the first camera and the second camera.
[0397] In some embodiments, the processor can also be configured to determine whether the object is within the central field of view; track the object using image data output by the first camera; determine whether the object tracking meets the quality standard; when the object tracking does not meet the quality standard, enable the second camera or increase the frame rate of the second camera and track the object using the first camera and the second camera.
[0398] In some embodiments, the first camera may be configured to selectively output image frames or image data that meet the intensity change criteria; the processor may also be configured to provide instructions to the first camera to limit image data acquisition to one or more portions of the central field of view corresponding to the object; and the image data output by the first camera may be used for one or more portions of the central field of view.
[0399] In some embodiments, the first camera may be configured to selectively output image frames or image data that meet the intensity change criteria; the second camera may include a plenoptic camera; the processor may also be configured to: determine whether the object meets the depth criteria; tracking the object using the first camera and the second camera may include using light field information obtained from the plenoptic camera to track the object when the depth criteria is met; when the depth criteria is not met, using depth information determined by image stereoscopy output from the first camera and the second camera to track the object.
[0400] In some embodiments, the plenoptic camera may include a transmissive diffraction mask.
[0401] In some embodiments, the plenoptic camera may have a horizontal field of view between 90 and 140 degrees and a central field of view extending between 40 and 80 degrees.
[0402] In some embodiments, the first camera may provide a grayscale image and the second camera may provide a color image.
[0403] In some embodiments, the first camera may be configured to selectively output image frames or image data that meet an intensity variation criterion; the first camera includes a global shutter; the second camera includes a rolling shutter; the processor may also be configured to: compare a first image acquired using the first camera with a second image acquired using the second camera to detect a skew in at least a portion of the second image; and adjust at least a portion of the second image to compensate for the detected skew.
[0404] In some embodiments, comparing a first image acquired using a first camera to a second image acquired using a second camera may include performing a line-by-line comparison between the first image and the second image.
[0405] In some embodiments, the processor may be mechanically coupled to the frame.
[0406] In some embodiments, the frame includes a display device mechanically coupled to the processor.
[0407] In some embodiments, the local data processing module may include a processor, the local data processing module may be operably coupled to the display device via a communication link, and wherein the framework may include the display device.
[0408] In some embodiments, the processor may be further configured to determine that an occlusion criterion for the first camera is satisfied; enable or increase a frame rate of the second camera; and track the object using image data output by the second camera.
[0409] In some embodiments, the occlusion criteria may be met when the object occupies more than a threshold amount of the field of view of the first camera.
[0410] In some embodiments, the object may be a stationary object within the environment.
[0411] In some embodiments, the object may be a hand of a user of the wearable display system.
[0412] In some embodiments, the wearable display system also includes an IR transmitter mechanically coupled to the frame.
[0413] In some embodiments, the IR emitter is configured to be selectively activated to provide IR illumination.
[0414] In addition, although the advantages of the present disclosure are pointed out, it should be understood that not every embodiment of the present disclosure will include every described advantage. Some embodiments may not implement any feature described as advantageous herein. Therefore, the foregoing description and drawings are only intended as examples.
[0415] The above-mentioned embodiments of the present disclosure can be implemented in any of a variety of ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or processor set, whether provided in a single computer or distributed in multiple computers. Such a processor can be implemented as an integrated circuit, with one or more processors in an integrated circuit component, including commercially available integrated circuit components known in the art, with names such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors. In some embodiments, the processor can be implemented in a custom circuit (such as an ASIC) or in a semi-custom circuit generated by configuring a programmable logic device. As another alternative, the processor can be part of a larger circuit or semiconductor device, whether commercially available, semi-custom or custom. As a specific example, some commercially available microprocessors have multiple cores, so that one or a subset of these cores can constitute a processor. However, a circuit of any appropriate format can be used to implement the processor.
[0416] Furthermore, it should be understood that a computer may be embodied in any of a variety of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer may be embedded in a device that is not generally considered a computer but has suitable processing capabilities, including a personal digital assistant (PDA), a smart phone, or any other suitable portable or fixed electronic device.
[0417] In addition, the computer may have one or more input and output devices. These devices may be used in particular for presenting a user interface. Examples of output devices that may be used to provide a user interface include a printer or display screen for visual presentation output, and a speaker or other sound generating device for auditory presentation output. Examples of input devices that may be used for a user interface include a keyboard and a pointing device, such as a mouse, a touch pad, and a digitized tablet computer. As another example, a computer may receive input information by voice recognition or other audible formats. In the illustrated embodiment, the input / output device is shown as being physically separated from the computing device. However, in some embodiments, the input and / or output device may be physically integrated into the same unit as other elements of a processor or computing device. For example, a keyboard may be implemented as a soft keyboard on a touch screen. In some embodiments, the input / output device may be completely disconnected from the computing device and functionally integrated by a wireless connection.
[0418] Such computers may be interconnected by one or more networks of any suitable form, including as a local area network or a wide area network such as an enterprise network or the Internet. Such a network may be based on any suitable technology and may operate according to any suitable protocol, and may include a wireless network, a wired network, or a fiber optic network.
[0419] In addition, the various methods or processes outlined herein may be encoded as software that can be executed on one or more processors using any of a variety of operating systems or platforms. In addition, such software may be written using any of a variety of suitable programming languages and / or programming or scripting tools, and may also be compiled into executable machine language code or intermediate code that is executed on a framework or virtual machine.
[0420] In this regard, the present disclosure may be embodied as a computer-readable storage medium (or multiple computer-readable media) (e.g., a computer memory, one or more floppy disks, compact discs (CDs), optical disks, digital video discs (DVDs), tapes, flash memory, field programmable gate arrays or other semiconductor devices or other tangible computer storage media) encoded with one or more programs, which when executed on one or more computers or other processors will perform methods for implementing the various embodiments of the present disclosure discussed above. It is apparent from the foregoing examples that the computer-readable storage medium can retain information for a sufficient time to provide computer-executable instructions in a non-transient form. Such one or more computer-readable storage media may be removable so that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement various aspects of the present disclosure as described above. As used herein, the term "computer-readable storage medium" covers only computer-readable media that can be considered as an article of manufacture (i.e., an article of manufacture) or a machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium, such as a propagating signal.
[0421] The term "program" or "software" is used herein in a general sense to refer to computer code or a computer executable instruction set that can be used to program a computer or other processor to implement various aspects of the present disclosure as described above. In addition, it should be understood that according to one aspect of this embodiment, one or more computer programs that perform the methods of the present disclosure when executed do not need to reside on a single computer or processor, but can be distributed in a modular manner between multiple different computers or processors to implement various aspects of the present disclosure.
[0422] Computer executable instructions can have many forms, such as program modules that are executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. Typically, in various embodiments, the functionality of program modules can be combined or distributed as needed.
[0423] In addition, the data structure can be stored in a computer-readable medium in any suitable form. To simplify the description, the data structure can be shown to have fields that are related by position in the data structure. Similarly, such relationships can be achieved by allocating storage for the fields by conveying the positions in the computer-readable medium that communicate the relationships between the fields. However, any suitable mechanism can be used to establish the relationship between the information in the fields of the data structure, including by using pointers, tags, or other mechanisms that establish relationships between data elements.
[0424] Various aspects of the present disclosure may be used alone, in combination, or in various arrangements not specifically discussed in the foregoing embodiments, and therefore, are not limited in their application to the details and arrangements of components set forth in the foregoing description or shown in the accompanying drawings. For example, aspects described in one embodiment may be combined in any manner with aspects described in other embodiments.
[0425] In addition, the present disclosure may be embodied as a method, an example of which has been provided. The actions performed as part of the method may be ordered in any suitable manner. Thus, embodiments may be constructed in which the actions are performed in an order different from that shown, and even if shown as sequential actions in an illustrative embodiment, the actions may include performing some actions simultaneously.
[0426] The use of ordinal terms such as "first", "second", "third" and the like in the claims to modify claim elements does not in itself indicate any priority, precedence or order of one claim element relative to another sequential or temporal order of performing the method actions, but serves merely as a marker to distinguish one claim element having a certain name from another element having the same name (but used in ordinal numbers) to distinguish the claim elements.
[0427] In addition, the words and terms used herein are for the purpose of description and should not be regarded as limiting. The use of "including," "comprising," or "having," "containing," "involving," and variations thereof herein is intended to encompass the items listed thereafter and their equivalents as well as other items.
Claims
1. A wearable display system, comprising: A head mounted device comprising: a first camera configured to operate in a first mode to output image frames and in a second mode to output image data satisfying an intensity variation criterion; a second camera; and wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and a processor operably coupled to the first camera and the second camera and configured to: creating a world model using depth information determined stereoscopically from images output by the first camera and the second camera; and Head pose is tracked using the world model and the image data output by the first camera.
2. The wearable display system according to claim 1, wherein: The intensity change standard includes an absolute or relative intensity change standard.
3. The wearable display system according to claim 1, wherein: The first camera is configured to output the image data asynchronously.
4. The wearable display system according to claim 3, wherein: The processor is also configured to asynchronously track head pose.
5. The wearable display system according to claim 1, wherein: The processor is further configured to execute a tracking routine to restrict image data acquisition to points of interest within the world model.
6. The wearable display system according to claim 5, wherein: the first camera being configured to restrict image acquisition to one or more portions of a field of view of the first camera; as well as The tracking routine includes: identifying points of interest within the world model; determining one or more first portions of the first camera's field of view corresponding to the points of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more first portions of the field of view.
7. The wearable display system according to claim 6, wherein: The tracking routine also includes: estimating one or more second portions of the first camera's field of view corresponding to the point of interest based on movement of the point of interest relative to the world model or movement of the head mounted device relative to the point of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more second portions of the field of view.
8. The wearable display system according to claim 5, wherein: The head mounted device also includes an inertial measurement unit; as well as Executing the tracking routine includes estimating an updated relative position of an object based at least in part on an output of the inertial measurement unit.
9. The wearable display system according to claim 5, wherein: The tracking routine includes: repeatedly calculating the position of the point of interest within the world model; and The calculations were repeated with a temporal resolution of more than 60 Hz.
10. The wearable display system according to claim 9, wherein: The intervals between the repeated calculations are between 1 millisecond and 15 milliseconds in duration.
11. The wearable display system according to claim 1, wherein: The processor is further configured to: determining whether head pose tracking meets quality standards; and When the head pose tracking does not meet the quality criterion, enabling the second camera or modulating a frame rate of the second camera.
12. The wearable display system according to claim 1, wherein: The processor is mechanically coupled to the head mounted device.
13. The wearable display system according to claim 1, wherein: The head mounted device includes a display device mechanically coupled to the processor.
14. The wearable display system according to claim 1, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head mounted device includes the display device.
15. The wearable display system according to claim 1, wherein: The head mounted device also includes an IR transmitter.
16. The wearable display system according to claim 15, wherein: The processor is configured to selectively enable the IR emitter to enable head gesture tracking in low light conditions.
17. A method for tracking head posture using a wearable display system, the wearable display system comprising: A head mounted device comprising: a first camera configured to operate in a first mode to output image frames and in a second mode to output image data satisfying an intensity variation criterion; a second camera; and wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and a processor operably coupled to the first camera and the second camera; The method comprises using the processor to: creating a world model using depth information determined stereoscopically from images output by the first camera and the second camera; and Head pose is tracked using the world model and image data output by the first camera.
18. The method according to claim 17, wherein: The intensity change standard includes an absolute or relative intensity change standard.
19. The method according to claim 17, wherein: The first camera is configured to output the image data asynchronously.
20. The method according to claim 19, wherein: The method includes using the processor to: asynchronously track head pose.
21. The method of claim 17, wherein: The method comprises using the processor to: A tracking routine is performed to restrict image data acquisition to points of interest within the world model.
22. The method of claim 21, wherein: the first camera being configured to restrict image acquisition to one or more portions of a field of view of the first camera; as well as The tracking routine includes: identifying points of interest within the world model; determining one or more first portions of the first camera's field of view corresponding to the points of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more first portions of the field of view.
23. The method of claim 22, wherein: The tracking routine also includes: estimating one or more second portions of the first camera's field of view corresponding to the point of interest based on movement of the point of interest relative to the world model or movement of the head mounted device relative to the point of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more second portions of the field of view.
24. The method of claim 21, wherein: The head mounted device also includes an inertial measurement unit; as well as Executing the tracking routine includes estimating an updated relative position of an object based at least in part on an output of the inertial measurement unit.
25. The method of claim 21, wherein: The tracking routine includes: repeatedly calculating the position of the point of interest within the world model; and The calculations were repeated with a temporal resolution of more than 60 Hz.
26. The method according to claim 25, wherein: The intervals between the repeated calculations are between 1 millisecond and 15 milliseconds in duration.
27. The method of claim 17, wherein: The method comprises using the processor to: determining whether head pose tracking meets quality standards; and When the head pose tracking does not meet the quality criterion, enabling the second camera or modulating a frame rate of the second camera.
28. The method of claim 17, wherein: The processor is mechanically coupled to the head mounted device.
29. The method of claim 17, wherein: The head mounted device includes a display device mechanically coupled to the processor.
30. The method of claim 17, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head mounted device includes the display device.
31. The method of claim 17, wherein: The head mounted device also includes an IR transmitter.
32. The method according to claim 31, wherein: The method includes using the processor to: selectively enable the IR transmitter to enable head gesture tracking in low light conditions.
33. A wearable display system, comprising: frame; a first camera mechanically coupled to the frame, wherein the first camera is configured to operate in a first mode to output image frames in a first field of view of the first camera and in a second mode to output image data satisfying an intensity variation criterion; and a processor operably coupled to the first camera and configured to: determining whether an object is within the first field of view; and The motion of the object is tracked using image data received from the first camera for one or more portions of the first field of view.
34. The wearable display system of claim 33, further comprising a head mounted device, the head mounted device comprising the frame, the first camera, and a second camera, the second camera being mechanically coupled to the frame, wherein The first camera and the second camera are positioned to provide overlapping views of a central field of view; and the processor is operably coupled to the first camera and the second camera and is configured to: creating a world model using depth information determined stereoscopically from images output by the first camera and the second camera; as well as Head pose is tracked using the world model and the image data output by the first camera.
35. The wearable display system of claim 33, wherein: The intensity change standard includes an absolute or relative intensity change standard.
36. The wearable display system of claim 33, wherein: The first camera is configured to output the image data asynchronously.
37. The wearable display system of claim 36, wherein: The processor is also configured to asynchronously track head pose.
38. The wearable display system of claim 34, wherein: The processor is further configured to execute a tracking routine to restrict image data acquisition to points of interest within the world model.
39. A wearable display system according to claim 38, wherein: the first camera being configured to restrict image acquisition to one or more portions of a field of view of the first camera; as well as The tracking routine includes: identifying points of interest within the world model; determining one or more first portions of the first camera's field of view corresponding to the points of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more first portions of the field of view.
40. The wearable display system of claim 39, wherein: The tracking routine also includes: estimating one or more second portions of the first camera's field of view corresponding to the point of interest based on movement of the point of interest relative to the world model or movement of the head mounted device relative to the point of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more second portions of the field of view.
41. The wearable display system of claim 38, wherein: The head mounted device also includes an inertial measurement unit; as well as Executing the tracking routine includes estimating an updated relative position of an object based at least in part on an output of the inertial measurement unit.
42. The wearable display system of claim 38, wherein: The tracking routine includes: repeatedly calculating the position of the point of interest within the world model; and The calculations were repeated with a temporal resolution of more than 60 Hz.
43. The wearable display system of claim 42, wherein: The intervals between the repeated calculations are between 1 millisecond and 15 milliseconds in duration.
44. The wearable display system of claim 34, wherein: The processor is further configured to: determining whether head pose tracking meets quality standards; and When the head pose tracking does not meet the quality criterion, enabling the second camera or modulating a frame rate of the second camera.
45. The wearable display system of claim 34, wherein: The processor is mechanically coupled to the head mounted device.
46. The wearable display system of claim 34, wherein: The head mounted device includes a display device mechanically coupled to the processor.
47. The wearable display system of claim 34, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head mounted device includes the display device.
48. The wearable display system of claim 34, wherein: The head mounted device also includes an IR transmitter.
49. The wearable display system of claim 48, wherein: The processor is configured to selectively enable the IR emitter to enable head gesture tracking in low light conditions.
50. A method of tracking motion of an object using a wearable display system, the wearable display system comprising: frame; a first camera mechanically coupled to the frame, wherein the first camera is configured to operate in a first mode to output image frames in a first field of view of the first camera and in a second mode to output image data satisfying an intensity variation criterion; and a processor operably coupled to the first camera; The method comprises using the processor to: determining whether the object is within the first field of view; and The motion of the object is tracked using image data received from the first camera for one or more portions of the first field of view.
51. The method of claim 50, wherein: The wearable display system also includes a head mounted device, the head mounted device including the frame, the first camera and a second camera, the second camera being mechanically coupled to the frame, wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and the processor is operably coupled to the first camera and the second camera, wherein the method includes using the processor to: creating a world model using depth information determined stereoscopically from images output by the first camera and the second camera; and Head pose is tracked using the world model and the image data output by the first camera.
52. The method of claim 50, wherein: The intensity change standard includes an absolute or relative intensity change standard.
53. The method of claim 50, wherein: The first camera is configured to output the image data asynchronously.
54. The method of claim 53, wherein: The method includes using the processor to: asynchronously track head pose.
55. The method of claim 51, wherein: The method includes using the processor to: execute a tracking routine to restrict image data acquisition to points of interest within the world model.
56. The method of claim 55, wherein: the first camera being configured to restrict image acquisition to one or more portions of a field of view of the first camera; as well as The tracking routine includes: identifying points of interest within the world model; determining one or more first portions of the first camera's field of view corresponding to the points of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more first portions of the field of view.
57. The method of claim 56, wherein: The tracking routine also includes: estimating one or more second portions of the first camera's field of view corresponding to the point of interest based on movement of the point of interest relative to the world model or movement of the head mounted device relative to the point of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more second portions of the field of view.
58. The method of claim 55, wherein: The head mounted device also includes an inertial measurement unit; as well as Executing the tracking routine includes estimating an updated relative position of an object based at least in part on an output of the inertial measurement unit.
59. The method of claim 55, wherein: The tracking routine includes: repeatedly calculating the position of the point of interest within the world model; and The calculations were repeated with a temporal resolution of more than 60 Hz.
60. The method of claim 59, wherein: The intervals between the repeated calculations are between 1 millisecond and 15 milliseconds in duration.
61. The method of claim 51, wherein: The method comprises using the processor to: determining whether head pose tracking meets quality standards; and When the head pose tracking does not meet the quality criterion, enabling the second camera or modulating a frame rate of the second camera.
62. The method of claim 51, wherein: The processor is mechanically coupled to the head mounted device.
63. The method of claim 51, wherein: The head mounted device includes a display device mechanically coupled to the processor.
64. The method of claim 51, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head mounted device includes the display device.
65. The method of claim 51, wherein: The head mounted device also includes an IR transmitter.
66. The method of claim 65, wherein: The method includes using the processor to: selectively enable the IR transmitter to enable head gesture tracking in low light conditions.
67. A wearable display system, comprising: frame; two cameras mechanically coupled to the frame, wherein the two cameras include: a first camera configured to operate in a first mode to output image frames and in a second mode to output image data satisfying an intensity variation criterion; and a second camera, wherein the first camera and the second camera are positioned to provide overlapping views of a central field of view; and A processor is operably coupled to the first camera and the second camera.
68. The wearable display system of claim 67, wherein: The wearable display system further includes a head mounted device, the head mounted device including the frame, the first camera, and the second camera, wherein the processor is configured to: creating a world model using depth information determined stereoscopically from images output by the first camera and the second camera; and Head pose is tracked using the world model and the image data output by the first camera.
69. The wearable display system of claim 67, wherein: The intensity change standard includes an absolute or relative intensity change standard.
70. The wearable display system of claim 67, wherein: The first camera is configured to output the image data asynchronously.
71. The wearable display system of claim 70, wherein: The processor is also configured to asynchronously track head pose.
72. The wearable display system of claim 68, wherein: The processor is further configured to execute a tracking routine to restrict image data acquisition to points of interest within the world model.
73. The wearable display system of claim 72, wherein: the first camera being configured to restrict image acquisition to one or more portions of a field of view of the first camera; as well as The tracking routine includes: identifying points of interest within the world model; determining one or more first portions of the first camera's field of view corresponding to the points of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more first portions of the field of view.
74. A wearable display system according to claim 73, wherein: The tracking routine also includes: estimating one or more second portions of the first camera's field of view corresponding to the point of interest based on movement of the point of interest relative to the world model or movement of the head mounted device relative to the point of interest; and Instructions are provided to the first camera to restrict image acquisition to the one or more second portions of the field of view.
75. The wearable display system of claim 72, wherein: The head mounted device also includes an inertial measurement unit; as well as Executing the tracking routine includes estimating an updated relative position of an object based at least in part on an output of the inertial measurement unit.
76. The wearable display system of claim 72, wherein: The tracking routine includes: repeatedly calculating the position of the point of interest within the world model; and The calculations were repeated with a temporal resolution of more than 60 Hz.
77. A wearable display system according to claim 76, wherein: The intervals between the repeated calculations are between 1 millisecond and 15 milliseconds in duration.
78. The wearable display system of claim 68, wherein: The processor is further configured to: determining whether head pose tracking meets quality standards; and When the head pose tracking does not meet the quality criterion, enabling the second camera or modulating a frame rate of the second camera.
79. The wearable display system of claim 68, wherein: The processor is mechanically coupled to the head mounted device.
80. The wearable display system of claim 68, wherein: The head mounted device includes a display device mechanically coupled to the processor.
81. The wearable display system of claim 68, wherein: A local data processing module includes the processor, the local data processing module is operably coupled to a display device via a communication link, and wherein the head mounted device includes the display device.
82. The wearable display system of claim 68, wherein: The head mounted device also includes an IR transmitter.
83. The wearable display system of claim 82, wherein: The processor is configured to selectively enable the IR emitter to enable head gesture tracking in low light conditions.
Citation Information
Patent Citations
Methods and systems for creating virtual and augmented reality
US20160026253A1
Multi-baseline camera array system architectures for depth augmentation in VR / ar applications
US20160309134A1
Systems and Methods for Synthesizing Images from Image Data Captured by an Array Camera Using Restricted Depth of Field Depth Maps in which Depth Estimation Precision Varies
US20170094243A1
Device for displaying an image sequence and system for displaying a scene
US20170111619A1
Using dynamic vision sensors for motion detection in head mounted displays
US20180295337A1