Lightweight, low-power cross-reality device with high transient resolution
The described XR system addresses weight and power issues by using a combination of grayscale and RGB cameras with selective sensor activation, ensuring accurate and realistic XR experiences with reduced power consumption.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- MAGIC LEAP INC
- Filing Date
- 2024-08-09
- Publication Date
- 2026-04-30
AI Technical Summary
Wearable XR systems face challenges with weight, power consumption, and accuracy in tracking physical objects due to sensor shifts and high power requirements, leading to compromised user experience and realism.
A lightweight XR system using a combination of grayscale and RGB cameras, inertial measurement units, and selective sensor activation to create a world model for accurate head and object tracking with reduced power consumption.
Enables high-accuracy, low-power, and lightweight XR experiences with improved realism by reducing sensor reliance and implementing calibration routines for accurate stereoscopic depth measurement.
Smart Images

Figure 0007854014000002 
Figure 0007854014000003 
Figure 0007854014000004
Abstract
Description
Technical Field
[0001] This application generally relates to a wearable cross-reality display system (XR system) including a dynamic vision sensor (DVS) camera.
Background Art
[0002] A computer can create an X reality (XR or cross-reality) environment that controls a human user interface and in which some or all of the XR environment is generated by the computer as it is perceived by the user. These XR environments can be virtual reality (VR), augmented reality (AR), or mixed reality (MR) environments in which some or all of the XR environment can be generated by the computer using data that describes the environment in part. This data can describe virtual objects that can be rendered in a way that allows the user to perceive or sense them as part of the physical world so that the user can interact with the virtual objects. The user can experience these virtual objects as a result of data being rendered and presented through a user interface device such as a head-mounted display device. The data can be displayed to be visible to the user, or reproduced to be audible to the user, can control audio, or can control a haptic (or tactile) interface to enable the user to experience a touch sensation as the user senses or perceives a virtual object.
[0003] XR systems can be useful for a wide range of applications, including scientific visualization, medical training, engineering design and prototyping, remote operation and telepresence, and personal entertainment. AR and MR, in contrast to VR, involve one or more virtual objects in relation to real objects in the physical world. The experience of virtual objects interacting with real objects generally enhances the user experience when using XR systems and expands the possibilities for various applications, presenting realistic and easily understandable information about how the physical world can be altered. [Overview of the project] [Means for solving the problem]
[0004] Aspects of this application relate to a wearable cross-reality display system configured with a DVS camera. The techniques described herein may be used together, separately, or in any preferred combination.
[0005] According to some embodiments, a wearable display system may be provided, which includes a headset comprising a first camera and a second camera, which are configured to output image frames or image data satisfying an intensity change criterion, the first and second cameras being positioned to provide an overlapping view of the central field of view, and a processor operably coupled to the first and second cameras and configured to create a world model using depth information stereoscopically determined from images output by the first and second cameras, and to track head pose using the world model and image data output by the first camera.
[0006] In some embodiments, the intensity change criterion may include absolute or relative intensity change criteria.
[0007] In some embodiments, the first camera may be configured to output image data asynchronously.
[0008] In some embodiments, the processor may be further configured to track head posture asynchronously.
[0009] In some embodiments, the processor may be further configured to perform tracking routines and restrict image data acquisition to points of interest within the world model.
[0010] In some embodiments, the first camera may be configured to restrict image acquisition to one or more portions of the first camera's field of view, and the tracking routine may include the steps of identifying a point of interest in the world model, determining one or more first portions of the first camera's field of view corresponding to the point of interest, and providing the first camera with an instruction to restrict image acquisition to one or more first portions of its field of view.
[0011] In some embodiments, the tracking routine may further include the steps of: estimating one or more second portions of the field of view of a first camera corresponding to a point of interest, based on the movement of the point of interest relative to a world model or the movement of the headset relative to the point of interest; and providing instructions to the first camera to limit image acquisition to one or more second portions of the field of view.
[0012] In some embodiments, the headset may further include an inertial measurement unit, and the step of performing the tracking routine may at least in part include the step of estimating the updated relative position of the object based on the output of the inertial measurement unit.
[0013] In some embodiments, the tracking routine may include a step of iteratively calculating the position of the point of interest within the world model, and the iterative calculation may be performed with a transient resolution greater than 60 Hz.
[0014] In some embodiments, the interval between repeated calculations may have a duration of 1 ms to 15 ms.
[0015] In some embodiments, the processor may further determine whether head pose tracking meets quality criteria, and if head pose tracking does not meet quality criteria, enable a second camera or modulate the frame rate of the second camera.
[0016] In some embodiments, the processor may be mechanically coupled to the headset.
[0017] In some embodiments, the headset may include a display device that is mechanically coupled to the processor.
[0018] In some embodiments, the local data processing module may include a processor, and the local data processing module may be operably coupled to a display device via a communication link, and the headset may include a display device.
[0019] In some embodiments, the headset may further include an IR emitter.
[0020] In some embodiments, the processor may be configured to selectively enable the IR emitter to enable head attitude tracking under low light conditions.
[0021] According to some embodiments, a method for tracking head posture using a wearable display system may be provided, comprising a headset comprising a first camera and a second camera, which are configurable to output image frames or image data satisfying intensity change criteria, the first and second cameras being positioned to provide overlapping views of the central field of view, and a processor operably coupled to the first and second cameras, the method comprising the steps of creating a world model using depth information stereoscopically determined from images output by the first and second cameras using the processor, and tracking head posture using the world model and image data output by the first camera.
[0022] According to some embodiments, a wearable display system may be provided, which comprises a frame, a first camera mechanically coupled to the frame and configured to output image data that satisfies a first field of view intensity change criterion with respect to the first camera, and a processor operably coupled to the first camera and configured to determine whether an object is in the first field of view and to track the motion of an object using image data received from the first camera with respect to one or more portions of the first field of view.
[0023] According to some embodiments, a method of tracking the movement of an object using a wearable display system is provided. The wearable display system includes a frame and a first camera mechanically coupled to the frame, the first camera being configured to output image data that satisfies an intensity change criterion within a first field of view of the first camera, and a processor operably coupled to the first camera. The method includes determining, using the processor, whether the object is within the first field of view, and tracking the movement of the object using the image data received from the first camera with respect to one or more portions of the first field of view.
[0024] According to some embodiments, a wearable display system is provided. The wearable display system includes a frame and two cameras mechanically coupled to the frame, one first camera and one second camera, the first camera and the second camera being configured to output image data that satisfies an intensity change criterion, and the first camera and the second camera being positioned to provide overlapping views of a central field of view, and a processor operably coupled to the first camera and the second camera.
[0025] The foregoing description is provided by way of illustration and not by way of limitation. The present invention provides, for example, the following. (Item 1) A wearable display system, wherein the wearable display system is a headset, and the headset includes one first camera, the one first camera being configured to output an image frame or image data that satisfies an intensity change criterion, and one second camera and is positioned such that the first camera and the second camera provide overlapping views of a central field of view. a headset A processor, the processor being operably coupled to the first camera and the second camera, the processor using depth information stereoscopically determined from images output by the first camera and the second camera to create a world model, using the world model and image data output by the first camera to track a head pose and configured to perform, a processor A wearable display system comprising. (Item 2) The wearable display system according to item 1, wherein the intensity change criterion includes an absolute or relative intensity change criterion. (Item 3) The wearable display system according to item 1, wherein the first camera is configured to output the image data asynchronously. (Item 4) The wearable display system according to item 3, wherein the processor is further configured to track the head pose asynchronously. (Item 5) The wearable display system according to item 1, wherein the processor is further configured to implement a tracking routine and limit image data acquisition to a point of interest within the world model. (Item 6) The first camera is configurable to limit image acquisition to one or more portions of the field of view of the first camera, The tracking routine identifying a point of interest within the world model, determining one or more first portions of the field of view of the first camera corresponding to the point of interest, and providing an instruction to the first camera to limit image acquisition to one or more first portions of the field of view. The wearable display system according to item 5, including. (Item 7) The tracking routine further Based on the movement of the point of interest relative to the world model or the movement of the headset relative to the point of interest, one or more second portions of the field of view of the first camera corresponding to the point of interest are estimated. The first camera is instructed to limit image acquisition to one or more second portions of the field of view. Wearable display systems as described in item 6, including the one described in item 6. (Item 8) The headset further includes an inertial measurement unit, Implementing the tracking routine includes, at least in part, estimating the updated relative position of the object based on the output of the inertial measurement unit. Wearable display systems as described in item 5. (Item 9) The tracking routine includes repeatedly calculating the location of the point of interest within the world model, The repeated calculations are performed with a transient resolution exceeding 60Hz. Wearable display systems as described in item 5. (Item 10) The wearable display system described in item 9, wherein the interval between the repeated calculations has a duration of 1 ms to 15 ms. (Item 11) The aforementioned processor further, To determine whether the head posture tracking meets the quality standards, When the head posture tracking fails to meet the quality criteria, the second camera is enabled, or the frame rate of the second camera is modulated. A wearable display system as described in item 1, configured to perform the following: (Item 12) The wearable display system according to item 1, wherein the processor is mechanically coupled to the headset. (Item 13) The wearable display system according to item 1, wherein the headset comprises a display device mechanically coupled to the processor. (Item 14) The wearable display system according to item 1, wherein a local data processing module comprises the processor, the local data processing module is operably coupled to a display device via a communication link, and the headset comprises the display device. (Item 15) The headset further includes an IR emitter, as described in item 1, as part of the wearable display system. (Item 16) The wearable display system according to item 15, wherein the processor is configured to selectively enable the IR emitter to enable head pose tracking under low light conditions. (Item 17) A method for tracking head posture using a wearable display system, wherein the wearable display system is It is a headset, A first camera, wherein the first camera can be configured to output an image frame or image data that satisfies an intensity change criterion, One second camera and Includes, The first camera and the second camera are positioned to provide an overlapping view of the central field of view, and the headset is configured to do so. A processor operably coupled to the first camera and the second camera Equipped with, The above method uses the processor, Using depth information stereoscopically determined from the images output by the first camera and the second camera, a world model is created. Using the world model and the image data output by the first camera, the head posture is tracked. Methods that include... (Item 18) A wearable display system, wherein the wearable display system is Frame and, A first camera mechanically coupled to the frame, wherein the first camera can be configured to output image data that satisfies a first intensity change criterion within the field of view relating to the first camera, A processor, wherein the processor is operably coupled to the first camera, Determining whether the object is within the first field of view, Tracking the motion of the object using image data received from the first camera with respect to one or more portions of the first field of view. A processor and A wearable display system equipped with [features / equipment]. (Item 19) A method for tracking the movement of an object using a wearable display system, wherein the wearable display system is Frame and, A first camera mechanically coupled to the frame, wherein the first camera can be configured to output image data that satisfies a first intensity change criterion within the field of view relating to the first camera, A processor operably coupled to the first camera and Equipped with, The above method uses the processor, Determining whether the object is within the first field of view, Tracking the motion of the object using image data received from the first camera with respect to one or more portions of the first field of view. Methods that include... (Item 20) A wearable display system, wherein the wearable display system is Frame and, Two cameras mechanically coupled to the frame, wherein the two cameras are A first camera that can be configured to output image data that satisfies the intensity change criteria, One second camera and Equipped with, The first camera and the second camera are positioned to provide an overlapping view of the central field of view, and the two cameras are: A processor operably coupled to the first camera and the second camera A wearable display system equipped with [features / equipment]. [Brief explanation of the drawing]
[0026] The attached drawings are not intended to be drawn to scale. In the drawings, each identical or nearly identical component shown in various figures is represented by similar numbers. For clarity purposes, not all components are labeled in every drawing.
[0027] [Figure 1] Figure 1 is a sketch illustrating some examples of simplified augmented reality (AR) scenes according to several embodiments.
[0028] [Figure 2] Figure 2 is a schematic diagram illustrating some embodiments of an AR display system.
[0029] [Figure 3A] Figure 3A is a schematic diagram illustrating a user wearing an AR display system that renders AR content as the user moves through the physical world environment, according to several embodiments.
[0030] [Figure 3B] Figure 3B is a schematic diagram illustrating a viewing optical system assembly and associated components according to several embodiments.
[0031] [Figure 4] Figure 4 is a schematic diagram illustrating an image sensing system according to several embodiments.
[0032] [Figure 5A] Figure 5A is a schematic diagram illustrating the pixel cells in Figure 4 according to several embodiments.
[0033] [Figure 5B] Figure 5B is a schematic diagram illustrating the output events of the pixel cells in Figure 5A according to several embodiments.
[0034] [Figure 6] Figure 6 is a schematic diagram illustrating an image sensor according to several embodiments.
[0035] [Figure 7] Figure 7 is a schematic diagram illustrating an image sensor according to several embodiments.
[0036] [Figure 8] Figure 8 is a schematic diagram illustrating an image sensor according to several embodiments.
[0037] [Figure 9] Figure 9 is a simplified flowchart of a method for image detection according to several embodiments.
[0038] [Figure 10] Figure 10 is a simplified flowchart of the patch identification process shown in Figure 9, according to several embodiments.
[0039] [Figure 11] Figure 11 is a simplified flowchart of the patch trajectory estimation process shown in Figure 9, according to several embodiments.
[0040] [Figure 12] Figure 12 is a schematic diagram illustrating the patch trajectory estimation of Figure 11 for one viewpoint, according to several embodiments.
[0041] [Figure 13] Figure 13 is a schematic diagram illustrating the estimation of the patch trajectory in Figure 11 with respect to changes in viewpoint, according to several embodiments.
[0042] [Figure 14] Figure 14 is a schematic diagram illustrating an image sensing system according to several embodiments.
[0043] [Figure 15] Figure 15 is a schematic diagram illustrating the pixel cells in Figure 14 according to several embodiments.
[0044] [Figure 16] Figure 16 is a schematic diagram of a pixel subarray according to several embodiments.
[0045] [Figure 17A] Figure 17A is a cross-sectional view of a prenoptic device with an angle-of-arrival / intensity converter in the form of two aligned, stacked transmission diffraction masks (TDMs), according to several embodiments.
[0046] [Figure 17B] Figure 17B is a cross-sectional view of a pre-optic device with an angle-of-arrival / intensity converter in the form of two misaligned, stacked TDMs, according to several embodiments.
[0047] [Figure 18A] Figure 18A shows a pixel subarray, comprising color pixel cells and arrival angle pixel cells, according to several embodiments.
[0048] [Figure 18B]Figure 18B shows a pixel subarray, comprising color pixel cells and arrival angle pixel cells, according to several embodiments.
[0049] [Figure 18C] Figure 18C shows a pixel subarray with white pixel cells and arrival angle pixel cells, according to several embodiments.
[0050] [Figure 19A] Figure 19A is a top view of a photodetector array with a single TDM according to several embodiments.
[0051] [Figure 19B] Figure 19B is a side view of a photodetector array with a single TDM according to several embodiments.
[0052] [Figure 20A] Figure 20A is a top view of a photodetector array with multiple arrival angle / intensity converters in a TDM configuration, according to several embodiments.
[0053] [Figure 20B] Figure 20B is a side view of a photodetector array with multiple TDMs according to several embodiments.
[0054] [Figure 20C] Figure 20C is a side view of a photodetector array with multiple TDMs according to several embodiments.
[0055] [Figure 21] Figure 21 is a schematic diagram of a headset including two cameras and ancillary components, according to several embodiments.
[0056] [Figure 22] Figure 22 shows flowcharts of calibration routines in several embodiments.
[0057] [Figure 23A] Figures 23A–23C depict exemplary field-of-view diagrams associated with the headset of Figure 21, according to several embodiments. [Figure 23B] Figures 23A–23C depict exemplary field-of-view diagrams associated with the headset of Figure 21, according to several embodiments. [Figure 23C] Figures 23A–23C depict exemplary field-of-view diagrams associated with the headset of Figure 21, according to several embodiments.
[0058] [Figure 24] Figure 24 is a flowchart of methods for creating and updating passable world models according to several embodiments.
[0059] [Figure 25] Figure 25 is a flowchart of a method for head posture tracking according to several embodiments.
[0060] [Figure 26] Figure 26 is a flowchart of a method for object tracking according to several embodiments.
[0061] [Figure 27] Figure 27 is a flowchart of a method for manual tracking according to several embodiments. [Modes for carrying out the invention]
[0062] The inventors have recognized and acknowledged the true value of design and operational techniques for wearable XR display systems that enhance the enjoyability and usefulness of such systems. These design and / or operational techniques may enable the acquisition of information to perform multiple functions, including hand tracking, head pose tracking, and world reconstruction, using a limited number of cameras, which may be used to realistically render virtual objects so that they appear to interact realistically with physical objects. Wearable cross-reality display systems may be lightweight and consume low power during operation. The system may acquire image information about physical objects in the physical world with short latency using sensors in a specific configuration. The system may implement various routines to improve the accuracy and / or realism of the displayed XR environment. Such routines may include calibration routines to improve the accuracy of stereoscopic depth measurement, even if the lightweight frame distorts during use, and routines to detect and address incomplete depth information in the model of the physical world around the user.
[0063] The weight of known XR system headsets can limit user enjoyment. Such XR headsets can weigh over 340 grams (sometimes even over 700 grams). Glasses, in contrast, can weigh less than 50 grams. Wearing such relatively heavy headsets for extended periods can fatigue users or distract them, diverting them from the desired immersive XR experience. However, the inventors recognize and understand that some designs that reduce headset weight also increase headset flexibility, making lightweight headsets more susceptible to changes in sensor position or orientation during or over time. For example, when a user wears a lightweight headset that includes camera sensors, the relative orientation of these camera sensors can shift. Variations in the spacing of cameras used for stereoscopic imaging can affect the ability of those headsets to obtain accurate stereoscopic information, which depends on cameras having known positional relationships to each other. Therefore, a calibration routine that may be repeated when the headset is worn could enable a lightweight headset that uses stereoscopic imaging techniques to accurately obtain information about the world around the wearer.
[0064] The need to equip XR systems with components for obtaining information about objects in the physical world can also limit the usefulness and user enjoyment of these systems. While the obtained information is used to realistically present computer-generated virtual objects in their appropriate locations with a realistic appearance relative to physical objects, the need to obtain this information imposes limitations on the size, power consumption, and realism of XR systems.
[0065] XR systems can, for example, use sensors worn by the user to acquire information about objects in the physical world around the user, including information about the position of physical world objects within the user's field of view. Challenges arise because physical objects can move relative to the user's field of view as a result of either the object moving within the physical world so that the physical object's position within the user's field of view changes, or the user changing their orientation to the physical world. To present a realistic XR display, models of physical objects in the physical world must be updated frequently enough to capture these changes, processed with sufficiently low latency, and accurately predicted in the future to cover a full latency path, including rendering, so that virtual objects displayed based on that information will have the appropriate position and appearance relative to the physical object as the virtual object is displayed. Otherwise, virtual objects will appear out of sync with the physical object, and combined scenes involving physical and virtual objects will not appear realistically. For example, virtual objects may appear to be floating in space rather than resting on a physical object, or they may appear to bounce around the physical object. Errors in visual tracking are amplified, especially when the user is moving at high speed or when there is significant movement within the scene.
[0066] Such problems can be avoided by sensors that acquire new data at a high rate. However, the power consumed by such sensors can increase the weight of the system, leading to the need for larger batteries, or limiting the length of time such systems can be used. Similarly, the processor required to process the data generated at a high rate can deplete the battery, add extra weight to the wearable system, and further limit the usefulness or enjoyment of such systems. A known approach is to operate higher frame rate sensors for increased transient resolution, for example, at a higher resolution to capture sufficient visual detail. An alternative solution is to complement this solution with an IR time-of-flight sensor that can directly indicate the position of a physical object relative to the sensor, resulting in short latency, and simple processing can be performed when displaying a virtual object using this information. However, such sensors consume a substantial amount of power, especially when they operate in sunlight.
[0067] The inventors recognized and appreciated the true value of an XR system being able to address changes in sensor position or orientation during use or over time by repeatedly performing a calibration routine. This calibration routine can determine the current relative separation and orientation of the sensors contained within the headset. The wearable XR system can then take this relative separation and orientation of the headset sensors into account when calculating stereoscopic depth information. Using such calibration capability, the XR system can accurately obtain depth information indicating the distance to objects in the physical world without using an active depth sensor, or with only occasional use of active depth sensing. Since active depth sensing can consume substantial power, reducing or eliminating active depth sensing allows for a device that draws less power, which can reduce the size of the device as a result of increasing the operating time of the device without recharging the battery or reducing the battery size.
[0068] The inventors also recognized and acknowledged the true value of XR systems, which, through a suitable combination of image sensors and appropriate techniques for processing image information from those sensors, can obtain information about physical objects with low latency, even with reduced power consumption, by reducing the number of sensors used, eliminating, disabling, or selectively activating resource-intensive sensors, and / or reducing the overall sensor usage. In a specific embodiment, an XR system may include a headset with two world-facing cameras. The first camera may produce grayscale images and may have a global shutter. These grayscale images may be smaller in size than color images of similar resolution, represented in some instances with less than one-third the number of bits. This grayscale camera may require less power than a color camera of similar resolution. The grayscale camera may employ event-based data acquisition and / or patch tracking and be configured to limit the amount of data output. The second camera may be an RGB camera. A wearable cross-reality display system may selectively use this camera to reduce power consumption and extend battery life without compromising the user's XR experience. The RGB camera may, for example, be used in conjunction with a grayscale camera to build a world model of the user's surrounding environment, and the grayscale camera may then be used to track the user's head pose based on this world model. Alternatively, or in addition, the grayscale camera may be used primarily to track objects, including the user's hands. Image information from one or more other sensors may be employed based on detected conditions indicating poor tracking quality. The RGB camera may be enabled to obtain color information to assist in distinguishing objects from the background, or the output from a light field sensor may be used to passively provide depth information.
[0069] The techniques described herein may be used with or separately from many types of devices for many types of scenarios. Figure 1 illustrates such a scenario. Figures 2, 3A, and 3B illustrate exemplary AR systems including one or more processors, memory, sensors, and a user interface that may operate according to the techniques described herein.
[0070] Referring to Figure 1, AR Scene 4 is depicted, and the user of the AR system sees a park-like setting 6 in the physical world, featuring people, trees, buildings in the background, and a concrete platform 8. In addition to these physical objects, the user of the AR technology also perceives "seeing" a virtual object, here illustrated as a robot figure 10 standing on the concrete platform 8 in the physical world, and a flying cartoon-like avatar character 2 that looks like an anthropomorphic bumblebee, although these elements (e.g., avatar character 2 and robot figure 10) do not exist in the physical world. Due to the significant complexity of human visual perception and the nervous system, it is difficult to produce an AR system that facilitates a comfortable, natural, and rich presentation of virtual image elements among other image elements of the virtual or physical world.
[0071] Such scenarios can be presented to the user by displaying image information representing the actual environment around the user and overlaying it with information representing virtual objects that are not in the actual environment. In an AR system, the user may be able to see objects in the physical world, and the AR system provides information that renders virtual objects so that they appear in the appropriate places with appropriate visual characteristics so that the virtual objects appear to coexist with objects in the physical world. In an AR system, for example, the user may view objects in the physical world through a transparent screen so that the user can see them. The AR system may render virtual objects on its screen so that the user can see both the physical world and the virtual objects. In some embodiments, the screen may be worn by the user, such as a pair of goggles or glasses.
[0072] A scene may be presented to a user via a system comprising multiple components, including a user interface that can stimulate one or more of the user's senses, including sight, hearing, and / or touch. In addition, the system may include one or more sensors that can measure parameters of the physical parts of the scene, including the user's position and / or movement within the physical parts of the scene. Furthermore, the system may include one or more computing devices, along with associated computer hardware such as memory. These components may be integrated into a single device or further distributed across multiple interconnected devices. In some embodiments, some or all of these components may be integrated into a wearable device.
[0073] In some embodiments, an AR experience may be provided to the user through a wearable display system. Figure 2 illustrates an embodiment of a wearable display system 80 (hereinafter referred to as "System 80"). System 80 includes a head-mounted display device 62 (hereinafter referred to as "Display Device 62") and various mechanical and electronic modules and systems to support the functions of the Display Device 62. The Display Device 62 may be coupled to a frame 64, which is wearable by a user or viewer 60 (hereinafter referred to as "User 60") of the display system and is configured to position the Display Device 62 in front of the user's eyes 60. According to various embodiments, the Display Device 62 may be a sequential display. The Display Device 62 may be monocular or binocular.
[0074] In some embodiments, speaker 66 is coupled to frame 64 and positioned close to the user's ear canal. In some embodiments, another speaker (not shown) is positioned adjacent to another ear canal of the user's ear to provide stereo / adjustable sound control.
[0075] The system 80 may include a local data processing module 70. The local data processing module 70 may be operably coupled to the display device 62 via a communication link 68 by wired conductors or wireless connectivity, etc. The local data processing module 70 may be mounted in various configurations, such as being fixedly attached to the frame 64, fixedly attached to a helmet or hat worn by the user 60, built into headphones, or otherwise detachably attached to the user 60 (e.g., in a backpack configuration, in a belt-mounted configuration). In some embodiments, the local data processing module 70 may not be present because its components may be implemented in a remote server or other component, either integrated within the display device 62 or coupled to the display device 62 via wireless communication over a wide area network, etc.
[0076] The local data processing module 70 may include a processor and digital memory such as non-volatile memory (e.g., flash memory), both of which may be used to assist in data processing, caching, and storage. The data may include a) data captured from sensors (e.g., operably coupled to frame 64, or otherwise attached to user 60 such as image acquisition devices (camera, etc.), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, and / or gyroscopes), and / or b) data potentially obtained and / or processed using the remote processing module 72 and / or remote data repository 74 for passage to the display device 62 after processing or reading. The local data processing module 70 may be operably coupled to the remote processing module 72 and the remote data repository 74 by communication links 76, 78 via wired or wireless communication links, etc., so that these remote modules 72, 74 are operably coupled to each other and available as resources to the local processing and data module 70.
[0077] In some embodiments, the local data processing module 70 may include one or more processors (e.g., a central processing unit and / or one or more graphics processing units (GPUs)) configured to analyze and process data and / or image information. In some embodiments, the remote data repository 74 may include a digital data storage facility, which may be available through the internet or other networking configurations in a “cloud” resource configuration. In some embodiments, all data is stored and all calculations are performed in the local data processing module 70, enabling fully autonomous use from the remote module.
[0078] In some embodiments, the local data processing module 70 is operably coupled to a battery 82. In some embodiments, the battery 82 is a removable power source such as a commercially available battery. In other embodiments, the battery 82 is a lithium-ion battery. In some embodiments, the battery 82 includes both an internal lithium-ion battery that can be charged by the user 60 during the non-operating time of the system 80, and a removable battery that allows the user 60 to operate the system 80 over longer time cycles without needing to connect it to a power source and charge the lithium-ion battery, or to shut down the system 80 and replace the battery.
[0079] Figure 3A illustrates a user 30 wearing an AR display system that renders AR content as the user 30 moves through a physical world environment 32 (hereinafter referred to as "environment 32"). The user 30 positions the AR display system at a location 34, and the AR display system records information about the surroundings of the passable world (for example, a digital representation of real objects in the physical world that can be stored and updated in accordance with changes in real objects in the physical world) for location 34. Each location 34 may further be associated with an "attitude" related to environment 32 and / or mapped features or directional audio input. The user wearing the AR display system on their head may turn their eyes in a particular direction and tilt their head to create a head attitude of the system relative to the environment. At each location and / or attitude within the same location, sensors on the AR display system may capture different information about environment 32. Thus, the information collected at location 34 is aggregated into a data input 36 and can be processed by a passable world module 38, which can be implemented, for example, by processing on the teleprocessing module 72 in Figure 2.
[0080] The passable world module 38 determines, at least in part, where and how the AR content 40 may be placed in relation to the physical world, as determined from the data input 36. The AR content is “placed” in the physical world by presenting it in a manner that allows the user to see both the AR content and the physical world. Such an interface may be created, for example, using glasses that are transparent to the user, allowing them to see the physical world, and which can be controlled so that virtual objects appear in controlled locations within the user’s field of view. The AR content is rendered as if it were interacting with objects in the physical world. The user interface may be such that the user’s view of objects in the physical world is obscured, so that the AR content can, when appropriate, obscure the user’s view of those objects. For example, the AR content may be placed by appropriately selecting a portion of an element 42 (e.g., a table) in the environment 32 and displaying and showing the AR content 40 which is placed on that element 42, or otherwise shaped and positioned as if it were interacting with it. AR content may also be placed within structures not yet within the field of view 44, or relative to a mapped mesh model 46 of the physical world.
[0081] As described, element 42 is an embodiment of what could be multiple elements in the physical world, which are treated as if they were fixed and can be stored within the passable world module 38. Once stored within the passable world module 38, information about those fixed elements may be used to present information to the user so that the user 30 can perceive content on the fixed element 42 without the system needing to map it to the fixed element 42 each time the user 30 sees it. The fixed element 42 is therefore a mesh model mapped from a previous modeling session, or determined by a separate user, but can be stored in the passable world module 38 for future reference by multiple users. Thus, the passable world module 38 can recognize the environment 32 from a previously mapped environment and, initially, display AR content without the user 30's device mapping the environment 32, saving computation processes and cycles and avoiding latency for any rendered AR content.
[0082] Similarly, a mapped mesh model 46 of the physical world can be created by an AR display system, and the appropriate surfaces and metrics for interacting with and displaying AR content 40 can be mapped and stored in the passable world module 38 for future readout by user 30 or other users without needing to be remapped or modeled. In some embodiments, the data input 36 is input such as geographical location, user identification, and current activity to indicate to the passable world module 38 which of one or more fixed elements 42 is available, which AR content 40 was last placed on fixed element 42, and whether that same content should be displayed (such AR content is “persistent” content regardless of whether the user is viewing a particular passable world model).
[0083] Even in embodiments where objects are considered to be fixed, the passable world module 38 may be updated from time to time to account for the possibility of changes in the physical world. The model of the fixed object may be updated at a very low frequency. Other objects in the physical world may be moving or otherwise not be considered fixed. To render the AR scene with a sense of realism, the AR system may update the positions of these non-fixed objects at a much higher frequency than that used to update the fixed object. To enable accurate tracking of all objects in the physical world, the AR system may draw information from multiple sensors, including one or more image sensors.
[0084] Figure 3B is a schematic diagram of the visibility optical system assembly 48 and its associated optional components. The specific configuration is described in Figure 21 below. When oriented toward the user's eye 49, in some embodiments, the two eye-tracking cameras 50 detect metrics of the user's eye 49, such as eye shape, eyelid occlusion, pupil direction, and flash on the user's eye 49. In some embodiments, one of the sensors may be a depth sensor 51, such as a time-of-flight sensor, which emits a signal into the world, detects reflections of those signals from nearby objects, and determines the distance to a given object. The depth sensor can quickly determine, for example, whether an object has entered the user's field of view as a result of either the movement of those objects or a change in the user's posture. However, information about the position of objects within the user's field of view may be collected using other sensors, either alternatively or in addition. In some embodiments, a world camera 52 records a view larger than the peripheral field of view, maps the environment 32, and detects inputs that may affect the AR content. In some embodiments, the world camera 52 and / or camera 53 may be grayscale and / or color image sensors that output grayscale and / or color image frames at fixed time intervals. Camera 53 may further capture images of the physical world within the user's field of view at specific times. Pixels of the frame-based image sensor may be sampled iteratively, even if their values remain constant. The world camera 52, camera 53, and depth sensor 51 each have separate fields of view 54, 55, and 56, respectively, which collect and record data from physical world scenes such as the physical world environment 32 depicted in Figure 3A.
[0085] The inertial measurement unit 57 may determine the movement and / or orientation of the viewing optical system assembly 48. In some embodiments, each component is operably coupled to at least one other component. For example, the depth sensor 51 may be operably coupled to an eye-tracking camera 50 to determine the actual distance to a point and / or area in the physical world that the user's eye 49 is looking at.
[0086] It should be understood that the visibility optics assembly 48 may include some of the components illustrated in Figure 3B. For example, the visibility optics assembly 48 may include a different number of components. In some embodiments, for example, the visibility optics assembly 48 may include one world camera 52, two world cameras 52, or more world cameras instead of the four world cameras depicted. Alternatively, or in addition, cameras 52 and 53 do not need to capture visible light images of their entire field of view. The visibility optics assembly 48 may include other types of components. In some embodiments, the visibility optics assembly 48 may include one or more dynamic vision sensors whose pixels may respond asynchronously to relative changes in light intensity above a threshold.
[0087] In some embodiments, the viewing optical system assembly 48 may not include a depth sensor 51 based on time-of-flight information. In some embodiments, for example, the viewing optical system assembly 48 may include one or more plenoptic cameras whose pixels may capture not only light intensity but also the angle of incident light. For example, the plenoptic camera may include an image sensor overlaid with a transmissive diffraction mask (TDM). Alternatively, or in addition, the plenoptic camera may include an image sensor containing angle-sensing pixels and / or phase-detection autofocus pixels (PDAF) and / or a microlens array (MLA). Such a sensor may serve as a source of depth information instead of, or in addition to, the depth sensor 51.
[0088] Furthermore, it should be understood that the component configuration in Figure 3B is illustrated as an embodiment. The viewing optics assembly 48 may include components in any preferred configuration so that the user can have the maximum field of view for a particular set of components. For example, if the viewing optics assembly 48 has one world camera 52, the world camera may be located within the central region of the viewing optics assembly instead of on the side.
[0089] Information from these sensors within the visibility optics assembly 48 may be coupled to one or more processors in the system. The processors may generate data that can be rendered to allow the user to perceive virtual content interacting with objects in the physical world. The rendering may be implemented in any preferred way, including generating image data that depicts both physical and virtual objects. In other embodiments, the physical and virtual content may be depicted in a single scene by modulating the opacity of a display device through which the user views the physical world. The opacity may be controlled to create the appearance of the virtual objects and also to block the user from seeing objects in the physical world that are occluded by the virtual objects. In some embodiments, the image data may include only virtual content that can be viewed through a user interface and modified to realistically interact with the physical world (e.g., by clipping the content and considering occlusion). Regardless of how the content is presented to the user, a model of the physical world may be used so that the properties of the virtual objects, which may be influenced by physical objects, including the shape, position, motion, and visibility of the virtual objects, can be correctly calculated.
[0090] A model of the physical world may be constructed from data collected from sensors on a user's wearable device. In some embodiments, the model may be constructed from data collected by multiple users, which can be aggregated in a computing device that is remote (and may be "in the cloud") from all of the users.
[0091] In some embodiments, at least one of the sensors may be configured to use a compact, low-power component to acquire information about physical objects in the scene, particularly non-fixed objects, at high frequencies with short latency. The sensor may employ patch tracking to limit the amount of data output.
[0092] Figure 4 illustrates an image sensing system 400 according to several embodiments. The image sensing system 400 may include an image sensor 402, which may include an image array 404, which may contain a plurality of pixels, each responding to light as in a conventional image sensor. The sensor 402 may further include a network for accessing each pixel. Accessing a pixel may involve obtaining information about the incident light generated by that pixel. Alternatively, or in addition, accessing a pixel may involve controlling that pixel, for example, by configuring it to provide only an output in response to the detection of a certain event.
[0093] In the illustrated embodiment, the image array 404 is configured as an array with multiple rows and columns of pixels. In such an embodiment, the access network may be implemented as a row-address encoder / decoder 406 and a column-address encoder / decoder 408. The image sensor 402 may further include a network that generates inputs to the access network and controls the timing and order in which information is read from pixels in the image array 404. In the illustrated embodiment, the network is a patch tracking engine 410. In contrast to conventional image sensors that can continuously output image information captured by pixels in each row, the image sensor 402 may be controlled to output image information in defined patches. Furthermore, the location of those patches relative to the image array may change over time. In the illustrated embodiment, the patch tracking engine 410 may output image array access information and control the output of image information from portions of the image array 404 corresponding to the patch locations, and the access information may change dynamically based on estimates of the motion of objects in the environment and / or the motion of the image sensor relative to those objects.
[0094] In some embodiments, the image sensor 402 may have the functionality of a dynamic visual sensor (DVS) such that image information is provided by the sensor only when there is a change in image properties (e.g., intensity) relating to a pixel. For example, the image sensor 402 may apply one or more thresholds that define the on and off states of pixels. The image sensor may detect when pixels have changed state and selectively provide an output for only those pixels that have changed state, or for only those pixels in a patch. These outputs may be asynchronous as they are detected, rather than as part of reading all pixels in the array. The outputs may be in the form of an address-event representation (AER) 418, which may include, for example, a pixel address (e.g., row and column) and the type of event (on or off). An on event may indicate that a pixel cell at a particular pixel address has detected an increase in light intensity, and an off event may indicate that a pixel cell at a particular pixel address has detected a decrease in light intensity. The increase or decrease may be relative to an absolute level, or a change in level at the last output from the pixel. The change may be expressed, for example, as a fixed offset, or as a percentage of the value in the final output from the pixel.
[0095] The use of DVS techniques related to patch tracking can enable image sensors suitable for use in XR systems. When combined within an image sensor, the amount of data generated can be limited to data from pixel cells that detect changes within the patch that would trigger the output of an event.
[0096] In some scenarios, high-resolution image information is desirable. However, the use of large sensors with more than one million pixels of cells to generate high-resolution image information can generate a large amount of image information when DVS techniques are used. We recognize and understand that DVS sensors can produce a large number of events that reflect image changes other than those resulting from the movement of moving or tracked objects in the background. Currently, the resolution of DVS sensors is limited to less than 1 MB, for example, 128×128, 240×180, and 346×260, to limit the number of events generated. Such sensors may sacrifice resolution for tracking objects and may not detect, for example, the subtle movement of fingers on a hand. Furthermore, if the image sensor outputs image information in other formats, limiting the resolution of the sensor array to output a manageable number of events may also limit the use of the image sensor for generating high-resolution image frames, along with DVS functionality. Sensors such as those described herein may have resolutions higher than VGA, including up to 8 megapixels or 12 megapixels, in some embodiments. Nevertheless, patch tracking as described herein may be used to limit the number of events output by the image sensor per second. As a result, an image sensor operating in at least two modes may be enabled. For example, an image sensor with megapixel resolution may operate in a first mode, outputting events within a specific patch being tracked. In a second mode, it may output high-resolution image frames or portions of image frames. Such an image sensor may be controlled within the XR system to operate in these different modes based on the system's capabilities.
[0097] The image array 404 may include a plurality of pixel cells 500 arranged within the array. Figure 5A depicts an embodiment of the pixel cell 500, which in this embodiment is configured for use in an imaging array implementing the DVS technique. The pixel cell 500 may include a photosensitive circuit 502, a difference circuit 506, and a comparator 508. The photosensitive circuit 502 may include a photodiode 504 that converts light striking the photodiode into a measurable electrical signal. In this embodiment, the conversion is performed on an electrical current I. A transconductance amplifier 510 converts the photocurrent I into a voltage. The conversion may be linear or nonlinear according to a function of logI, etc. Regardless of the specific transfer function, the output of the transconductance amplifier 510 indicates the amount of light detected in the photodiode 504. Although a photodiode is illustrated as an embodiment, it should be understood that other photosensing components that produce a measurable output in response to incident light may be implemented in place of or in addition to the photodiode within the photosensitive circuit.
[0098] In the embodiment shown in Figure 5A, a network for determining whether the output of a pixel has changed sufficiently and for triggering an output for that pixel cell is incorporated within the pixel itself. In this embodiment, this function is implemented by a difference circuit 506 and a comparator 508. The difference circuit 506 may be configured, for example, to reduce DC mismatch between pixel cells by balancing the output of the difference circuit and resetting the level after the generation of an event. In this embodiment, the difference circuit 506 is configured to produce an output indicating the change in the output of the photodiode 504 since the last output. The difference circuit may include an amplifier 512 having a gain-A, a capacitor 514 which may be implemented as a single circuit element or one or more capacitors connected in the network, and a reset switch 516.
[0099] During operation, the pixel cell will be reset by momentarily closing switch 516. Such resets may occur at the start of circuit operation and at any time thereafter after an event is detected. When pixel 500 is reset, the voltage across capacitor 514, when subtracted from the output of transconductance amplifier 510, will result in a zero voltage at the input of amplifier 512. When switch 516 is opened, the output of transconductance amplifier 510, combined with the voltage drop across capacitor 514, will result in a zero voltage at the input of amplifier 512. The output of transconductance amplifier 510 changes as a result of changes in the amount of light striking photodiode 504. As the output of transconductance amplifier 510 increases or decreases, the output of amplifier 512 will fluctuate positively or negatively by an amount amplified by the gain of amplifier 512.
[0100] Comparator 508 may determine whether an event is generated and the sign of the event, for example, by comparing the output voltage V of the difference circuit with a predetermined threshold voltage C. In some embodiments, comparator 508 may include two comparators, each comprising a transistor, one pair of which may operate when the output of amplifier 512 shows a positive change and detect an increasing change (on-event), and the other comparator may operate when the output of amplifier 512 shows a negative change and detect a decreasing change (off-event). However, it should be understood that amplifier 512 may have negative gain. In such embodiments, an increase in the output of transconductance amplifier 510 may be detected as a negative voltage change at the output of amplifier 512. Similarly, it should be understood that the positive and negative voltages may be relative to ground or at any preferred reference level. Nevertheless, the value of the threshold voltage C may be controlled by the characteristics of the transistor (e.g., transistor size, transistor threshold voltage) and / or by the value of a reference voltage that may be applied to comparator 508.
[0101] Figure 5B illustrates an example of event output (on, off) of pixel cell 500 over time t. In the illustrated example, at time t1, the output of the difference circuit has the value V1; at time t2, the output of the difference circuit has the value V2; and at time t3, the output of the difference circuit has the value V3. Between time t1 and time t2, the photodiode senses a certain increase in light intensity, but the pixel cell does not output an event because the change in V does not exceed the threshold voltage C. At time t2, the pixel cell outputs an on event because V2 is greater than V1 by the threshold voltage C. Between time t2 and time t3, the photodiode senses a certain decrease in light intensity, but the pixel cell does not output an event because the change in V does not exceed the threshold voltage C. At time t3, the pixel cell outputs an off event because V3 is less than V2 by the threshold voltage C.
[0102] Each event may trigger an output in AER418. The output may include, for example, an indication of whether the event is an on or off event, and identification of the pixel, such as its row and column. Other information may be included with the output, either as an alternative or in addition. For example, a timestamp may be included, which may be useful if the event is queued for later transmission or processing. In another embodiment, the current level at the output of amplifier 510 may be included. Such information may be included optionally, for example, if further processing is to be performed in addition to detecting the motion of an object.
[0103] It should be understood that the frequency of the event output, and therefore the sensitivity of the pixel cell, may be controlled by the value of the threshold voltage C. For example, the frequency of the event output may be reduced by increasing the value of the threshold voltage C, or increased by decreasing the threshold voltage C. It should also be understood that the threshold voltage C may differ with respect to on-events and off-events, for example by setting different reference voltages for comparators for detecting on-events and comparators for detecting off-events. It should also be understood that the pixel cell may also output a value indicating the size of the change in light intensity, instead of, or in addition to, a code signal indicating the detection of an event.
[0104] The pixel cells 500 in Figures 5A and 5B are illustrated as examples according to several embodiments. Other designs may also be suitable for the pixel cells. In some embodiments, the pixel cell includes a photosensitive circuit and a difference circuit, but the comparator circuit may be shared with one or more other pixel cells. In some embodiments, the pixel cell may include a network of circuits configured to calculate the value of change, such as an active pixel sensor at the pixel level.
[0105] Regardless of how events are detected per pixel cell, the ability to configure pixels to output only in response to event detection may be used to limit the amount of information required to maintain a model of the position of non-fixed (i.e., movable) objects. For example, pixels within a patch may be configured to trigger a threshold voltage C when relatively small changes occur. Other pixels outside the patch may have a larger threshold, such as 3 or 5 times. In some embodiments, the threshold voltage C for any outside-patch pixel may be set high enough so that the pixel is effectively disabled and does not produce any output regardless of the amount of change. In other embodiments, outside-patch pixels may be disabled in other ways. In such embodiments, the threshold voltage may be fixed for all pixels, but pixels may be selectively enabled or disabled based on whether they are within a patch or not.
[0106] In yet another embodiment, threshold voltages for one or more pixels may be adaptively set to modulate the amount of data output from the image array. For example, an AR system may have processing capacity to handle a certain number of events per second. Thresholds for some or all pixels may be increased when the number of events output per second exceeds an upper limit. Alternatively, or in addition, thresholds may be decreased when the number of events per second falls below a lower limit, enabling more data for more accurate processing. The number of events per second may be 200 to 2,000 in a specific embodiment. Such a number of events constitutes a substantial reduction in the amount of data to be processed per second compared to processing all pixel values scanned and output from an image sensor, which would constitute, for example, 30 million or more pixel values per second. That number of events is a further reduction compared to processing only pixels in a patch, which may be fewer, but still tens of thousands or more pixel values per second.
[0107] Control signals for enabling and / or setting threshold voltages for each of multiple pixels may be generated in any preferred manner. However, in the illustrated embodiment, these control signals are set by the patch tracking engine 410 or based on processing in the processing module 72 or other processors.
[0108] Referring back to Figure 4, the image sensing system 400 may receive input from any suitable component so that the patch tracking engine 410 can dynamically select at least one region of the image array 404 to be enabled and / or disabled based on the received input so as to implement a patch. The patch tracking engine 410 may also include a digital processing network having memory for storing one or more parameters of the patch. The parameters may be, for example, the boundaries of the patch, and may also include other information such as information about a scale factor between the motion of the image array and the motion of an image of a movable object associated with the patch within the image array. The patch tracking engine 410 may also include a network configured to perform calculations on stored values and other measured values supplied as inputs.
[0109] In the illustrated embodiment, the patch tracking engine 410 receives a specification of the current patch as input. The patch may be specified based on its size and location within the image array 404, for example, by defining a range of row and column addresses for the patch. Such a specification may be provided as output of a processing module 72 (Figure 2) or other component that processes information about the physical world. The processing module 72 may specify the patch to encompass the current location of each movable object in the physical world or a subset of tracked movable objects in order to render the virtual object with an appropriate appearance to the physical world. For example, if the AR scene should include a toy doll as a virtual object, balanced on a physical object such as a moving toy car, the patch may be specified to encompass that toy car. The patch may not specify another toy car moving in the background, as there is little need to have up-to-date information about that object in order to render a realistic AR scene.
[0110] Regardless of how the patches are selected, information about the current location of the patches may be supplied to the patch tracking engine 410. In some embodiments, the patches may be rectangular, such that the location of the patch can be defined simply as a start and end row and column. In other embodiments, the patches may have other shapes, such as circles, and the patches may be defined in other ways, such as by a center point and radius.
[0111] In some embodiments, trajectory information may also be supplied for the patch. The trajectory may, for example, define the motion of the patch relative to the coordinates of the image array 404. The processing module 72 may, for example, construct a model of the motion of a movable object in the physical world and / or the motion of the image array 404 relative to the physical world. Since one or both of these motions may affect the location in the image array 404 on which the image of the object is projected, the trajectory of the patch in the image array 404 may be calculated based on one or both. The trajectory may be defined in any preferred way, such as a linear, quadratic, cubic, or other polynomial equation parameter.
[0112] In other embodiments, the patch tracking engine 410 may dynamically calculate the location of patches based on input from sensors providing information about the physical world. The information from the sensors may be supplied directly from the sensors. Alternatively, or in addition, the sensor information may be processed to extract information about the physical world before being supplied to the patch tracking engine 410. The extracted information may include, for example, the motion of the image array 404 relative to the physical world, the distance between the image array 404 and the objects whose images correspond to those within the patches, or other information that can be used to dynamically align the images of patches in the image array 404 with those of objects in the physical world as the image array 404 and / or objects move.
[0113] Embodiments of the input components may include an image sensor 412 and an inertial sensor 414. Embodiments of the image sensor 412 may include an eye-tracking camera 50, a depth sensor 51, a world camera 52, and / or camera 52. Embodiments of the inertial sensor 414 may include an inertial measurement unit 57. In some embodiments, the input components may be selected to provide data at a relatively high rate. The inertial measurement unit 57 may have an output rate of 200 to 2,000 measurements / second, for example, 800 to 1,200 measurements / second. Patch locations may also be updated at a high rate. By using the inertial measurement unit 57 as a source of input to the patch tracking engine 410, patch locations can be updated 800 to 1,200 times / second, in one specific embodiment. In this way, movable objects can be tracked with high accuracy using relatively small patches, which limit the number of events that need to be processed. Such an approach can lead to very short latency between changes in the relative position of the image sensor and the movable object, and similarly, short latency for updating the rendering of the virtual object to provide a desirable user experience.
[0114] In some scenarios, the movable objects being tracked using patches may be stationary objects in the physical world. The AR system may, for example, identify stationary objects by analyzing multiple images taken from the physical world and select features of one or more of these stationary objects as reference points for determining the motion of a wearable device having image sensors on it. Frequent and short-latency updates of the locations of these reference points relative to the sensor array may be used to provide frequent and short-latency calculations of the wearable device user's head posture. Frequent and short-latency updates of the head posture improve the user experience of the AR system, as the head posture may be used to realistically render virtual objects via a user interface on the wearable. Therefore, having inputs to the patch tracking engine 410 that control the position of the patches may result from only sensors with high power rates, such as one or more inertial measurement units, which can lead to a desirable user experience for the AR system.
[0115] However, in some embodiments, other information may also be supplied to the patch tracking engine 410, enabling it to calculate and / or apply trajectories to patches. This other information may include stored information 416 such as the passable world module 38 and / or mapped mesh model 46. This information may indicate one or more previous positions of an object relative to the physical world, such that a consideration of changes in these previous positions and / or changes in the current position relative to the previous positions may indicate the trajectory of the object in the physical world, which can then be mapped to the trajectory of a patch across the image array 404. Other information in the physical world model may be used as an alternative or in addition. For example, other information regarding the size of a movable object and / or its distance or position relative to the image array 404 may be used to calculate either the location or trajectory of a patch across the image array 404 associated with that object.
[0116] Regardless of the manner in which the trajectory is determined, the patch tracking engine 410 may apply the trajectory and calculate the updated location of the patch in the image array 404 at a high rate, such as faster than 1 time per second or more than 800 times per second. The rate may be limited by processing power, such as less than 2,000 times per second in some embodiments.
[0117] It should be understood that the process of tracking changes in movable objects may involve reconstructing a less-than-perfect physical world. However, reconstruction of the physical world may occur at intervals longer than the interval between updates to the location of movable objects, such as every 30 seconds or every 5 seconds. The locations of the objects to be tracked and the locations of the patches that will capture information about those objects may be recalculated when reconstruction of the physical world occurs.
[0118] Figure 4 illustrates an embodiment in which a processing network for both dynamically generating patches and controlling the selective output of image information from those patches is configured to directly control the image array 404 so that the image information output from the array is limited to selected information. Such a network may be integrated, for example, within the same semiconductor chip housing the image array 404, or it may be integrated into a separate controller chip for the image array 404. However, it should be understood that the network for generating control signals for the image array 404 may be distributed throughout the entire XR system. For example, some or all of the functions may be performed by programming within the processing module 72 or by other processors in the system.
[0119] The image sensing system 400 may output image information for each of several pixels. Each pixel in the image information may correspond to one of the pixel cells of the image array 404. The image information output from the image sensing system 400 may be image information for one or more patches, corresponding to at least one region of the image array 404 selected by the patch tracking engine 410. In some embodiments, when each pixel of the image array 404 has a configuration different from that shown in Figure 5A, the pixels in the output image information may identify pixels in which a change in light intensity has been detected by the image sensor 400 in one or more patches.
[0120] In some embodiments, the image information output from the image sensing system 400 may be image information relating to pixels outside of one or more patches corresponding to at least one region of the image array selected by the patch tracking engine 410. For example, a deer may be running in the physical world with a flowing river. Details of the river waves may not be of interest, but they may trigger pixel cells in the image array 402. The patch tracking engine 410 may create a patch surrounding the river and disable a portion of the image array 402 corresponding to the patch surrounding the river.
[0121] Further processing may be performed based on the identification of the changed pixels. For example, a portion of the world model corresponding to the portion of the physical world imaged by the changed pixels may be updated. These updates may be performed based on information collected using other sensors. In some embodiments, further processing may be coordinated on or triggered on multiple changed pixels within a patch. For example, an update may be performed once 10% or some other threshold amount of pixels in the patch have detected a change.
[0122] In some embodiments, image information in other formats may be output from the image sensor and used in combination with change information to update the world model. In some embodiments, the format of the image information output from the image sensor may change from time to time during the operation of the VR system. In some embodiments, for example, the pixel cell 500 may at times operate to produce a differential output, such as that produced in the comparator 508. The output of the amplifier 510 may be switchable to at other times output the magnitude of the light incident on the photodiode 504. For example, the output of the amplifier 510 may be switchably connected to a sensing line, which in turn is connected to an A / D converter that can provide digital indication of the magnitude of the incident light based on the magnitude of the output of the amplifier 510.
[0123] In this configuration, the image sensor operates as part of the AR system and, in most cases, outputs differentially, outputting only events for pixels where a change exceeding a threshold is detected, or only events for pixels within a patch where a change exceeding a threshold is detected. Periodically, such as every 5 to 30 seconds, a full image frame with information size for all pixels in the image array may be output. Short latency and accurate processing can thus be achieved, with differential information being used to quickly update selected parts of the world model where the most likely changes affecting user perception have occurred, while full images may be used to update larger parts of the world model. Full updates to the world model occur only at slower rates, but any delay in updating the model cannot significantly affect the user's perception of the AR scene.
[0124] The output mode of the image sensor may change from time to time throughout the operation of the image sensor so that the sensor outputs one or more of the following: intensity information for some or all of the pixels and indications of change for some or all of the pixels in the array.
[0125] It is not a requirement that image information from a patch be selectively output from the image sensor by limiting the information output from the image array. In some embodiments, image information may be output by all pixels in the image array, or only information about a specific region of the array may be output from the image sensor. Figure 6 depicts an image sensor 600 according to some embodiments. The image sensor 600 may include an image array 602. In this embodiment, the image array 602 may be similar to a conventional image array that scans and outputs rows and columns of pixel values. The operation of such an image array may be adapted by other components. The image sensor 600 may further include a patch tracking engine 604 and / or a comparator 606. The image sensor 600 may provide an output 610 to an image processor 608. The processor 608 may be, for example, part of a processing module 72 (Figure 2).
[0126] The patch tracking engine 604 may have a structure and function similar to that of the patch tracking engine 410. It may be configured to receive a signal defining at least one selected region of the image array 602, and then generate a control signal that defines the dynamic location of that region within the image array 602 based on a calculated trajectory of an image of an object represented by that region. In some embodiments, the patch tracking engine 604 may receive a signal defining at least one selected region of the image array 602, which may include trajectory information about one or more regions. The patch tracking engine 604 may be configured to perform a calculation that dynamically identifies pixel cells within the at least one selected region based on the trajectory information. Variations in the implementation of the patch tracking engine 604 are also possible. For example, the patch tracking engine may update the location of a patch based on a sensor indicating the motion of the image array 602 and / or the projected motion of an object associated with the patch.
[0127] In the embodiment illustrated in Figure 6, the image sensor 600 is configured to output difference information about pixels in an identified patch. The comparator 606 may be configured to receive control signals from the patch tracking engine 604 that identify pixels in the patch. The comparator 606 may selectively act on pixels output from the image array 602 that have addresses in the patch as indicated by the patch tracking engine 604. The comparator 606 may act on pixel cells to generate signals that indicate changes in sensed light detected by at least one region of the image array 602. In one embodiment of implementation, the comparator 606 may include a memory element that stores reset values of pixel cells in the array. As the current values of those pixels are scanned and output from the image array 602, the network in the comparator 606 may compare the stored values with the current values and output an indication if the difference exceeds a threshold. A digital network may be used, for example, to store the values and perform such comparisons. In this embodiment, the output of the image sensor 600 may be processed in the same way as the output of the image sensor 400.
[0128] In some embodiments, the image array 602, the patch tracking engine 604, and the comparator 606 may be implemented within a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, the image array 602 may be implemented within a single integrated circuit. The patch tracking engine 604 and the comparator 606 may be implemented within a second single integrated circuit, configured, for example, as a driver for the image array 602. Alternatively, or in addition, some or all of the functions of the patch tracking engine and / or comparator 606 may be distributed to other digital processors in the AR system.
[0129] Other configurations or processing networks are also possible. Figure 7 depicts an image sensor 700 according to several embodiments. The image sensor 700 may include an image array 702. In this embodiment, the image array 702 may have pixel cells with a differential configuration as shown with respect to pixel 500 in Figure 5A. However, the embodiments herein are not limited to differential pixel cells, and patch tracking may be implemented using an image sensor that outputs intensity information.
[0130] In the embodiment illustrated in Figure 7, the patch tracking engine 704 produces a control signal indicating the addresses of pixel cells in one or more patches being tracked. The patch tracking engine 704 may be constructed and operated in the same way as the patch tracking engine 604. Here, the patch tracking engine 704 provides the control signal to a pixel filter 706, which passes image information from only those pixels in the patch to the output 710. As shown, the output 710 is coupled to an image processor 708, which may further process the image information about the pixels in the patch using techniques such as those described herein or in other preferred methods.
[0131] Further modifications are illustrated in Figure 8, which depicts the image sensor 800 according to several embodiments. The image sensor 800 may include an image array 802, which may be a conventional image array that scans and outputs intensity values for pixels. The image array may be adapted to provide difference image information as described herein through the use of a comparator 806. The comparator 806, like the comparator 606, may calculate difference information based on stored values for pixels. Selected of these difference values may be passed to output 812 by a pixel filter 808. Similar to the pixel filter 706, the pixel filter 808 may receive a control input from a patch tracking engine 804. The patch tracking engine 804 may be similar to the patch tracking engine 704. Output 812 may be coupled to an image processor 810. Some or all of the components of the image sensor 800 described above may be implemented in a single integrated circuit. Alternatively, the components may be distributed across one or more integrated circuits or other components.
[0132] Image sensors, such as those described herein, operate as part of an augmented reality system and may maintain other information about the physical world that is useful in realistically rendering images of virtual objects in combination with information about movable objects or information about the physical environment. Figure 9 illustrates a method 900 for image sensing according to several embodiments.
[0133] At least part of Method 900 may be performed to operate an image sensor, including, for example, an image sensor 400, 600, 700, or 800. Method 900 may begin with the step of receiving image information from one or more inputs, including, for example, an image sensor 412, an inertial sensor 414, and stored information 416 (Action 902). Method 900 may at least partially include the step of identifying one or more patches on the image output of an image sensing system based on the received information (Action 904). An embodiment of Action 904 is illustrated in Figure 10. In some embodiments, Method 900 may include the step of calculating a moving trajectory for one or more patches (Action 906). An embodiment of Action 906 is illustrated in Figure 11.
[0134] Method 900 may also include the step (act 908) of configuring the image sensing system based at least in part on one or more identified patches and / or their estimated moving trajectories. Configuration may be achieved, for example, by enabling a portion of the pixel cells of the image sensing system based at least in part on one or more identified patches and / or their estimated moving trajectories, through a comparator 606, a pixel filter 706, etc. In some embodiments, the comparator 606 may receive a first reference voltage value for pixel cells corresponding to selected patches on the image and a second reference voltage value for pixel cells not corresponding to any selected patch on the image. The comparator 606 may have a comparator cell with the second reference voltage, and may set the second reference voltage to be much higher than the first reference voltage so that a reasonable change in light intensity sensed by a pixel cell does not result in an output from the pixel cell. In some embodiments, the pixel filter 706 may disable outputs from pixel cells with addresses (e.g., rows and columns) not corresponding to any selected patch on the image.
[0135] Figure 10 illustrates patch identification 904 according to several embodiments. Patch identification 904 may include, at least in part, the step (act 1002) of segmenting one or more images from one or more inputs based on color, light intensity, angle of arrival, depth, and semantics.
[0136] Patch recognition 904 may also include the step (action 1004) of recognizing one or more objects in one or more images. In some embodiments, object recognition 1004 may be based at least in part on predetermined features of the object, including, for example, hands, eyes, and facial features. In some embodiments, object recognition 1004 may be based on one or more virtual objects, for example, a virtual animal character walking on a physical pencil. Object recognition 1004 may target the virtual animal character as an object. In some embodiments, object recognition 1004 may be based at least in part on artificial intelligence (AI) training received by an image sensing system. For example, the image sensing system may be trained by reading images of cats in different types and colors, and thus learned characteristics of cats, and be able to identify cats in the physical world.
[0137] Patch identification 904 may include the step (action 1006) of generating a patch based on one or more objects. In some embodiments, object patching 1006 may generate a patch by calculating a convex hull or bounding box relating to one or more objects.
[0138] Figure 11 illustrates patch trajectory estimation 906 in several embodiments. Patch trajectory estimation 906 may include the step (action 1102) of predicting movement over time for one or more patches. Movement for one or more patches may occur for several reasons, including, for example, moving objects and / or moving users. Motion prediction 1102 may include the step of deriving movement velocities for moving objects and / or moving users based on received images and / or received AI training.
[0139] Patch trajectory estimation 906 may include, at least in part, the step (act 1104) of calculating the trajectory over time for one or more patches based on the predicted movement. In some embodiments, the trajectory may be calculated by modeling a linear equation, assuming that the objects under motion will continue to move in the same direction at the same velocity. In some embodiments, the trajectory may be calculated by curve fitting or by using heuristics, including pattern detection.
[0140] Figures 12 and 13 illustrate coefficients that may be applied in the calculation of patch trajectories. Figure 12 depicts an embodiment of a movable object, which in this embodiment is a moving object 1202 (e.g., a hand) that moves relative to the user of the AR system. In this embodiment, the user is wearing an image sensor as part of a head-mounted display 62. In this embodiment, the user's eye 49 is looking straight ahead so that the image array 1200 captures a field of view (FOV) related to the eye 49 for one viewpoint 1204. The object 1202 is within the FOV and therefore appears by creating an intensity variation in the corresponding pixel in the array 1200.
[0141] Array 1200 has multiple pixels 1208 arranged within the array. With respect to a system tracking hand 1202, patch 1206 in its array containing object 1202 at time t0 may contain a portion of multiple pixels. If object 1202 is moving, the location of the patch capturing that object will change over time. This change may be captured within the patch trajectory from patch 1206 to patches X and Y used at later times.
[0142] The patch trajectory may be estimated in action 906, etc., by identifying an object within the patch, for example, a feature 1210 relating to a fingertip in the illustrated embodiment. A motion vector 1212 may be calculated with respect to the feature. In this embodiment, the trajectory is modeled as a linear equation, and the prediction is based on the assumption that the object 1202 continues over time on its same patch trajectory 1214, leading to patch locations X and Y at each of two consecutive time periods.
[0143] As the patch location changes, the image of the moving object 1202 remains within the patch. The image information is limited to the information gathered using the pixels within the patch, but this image information is sufficient to represent the motion of the moving object 1202. This would be true whether the image information is intensity information or difference information, such as that produced by a difference circuit. In the case of a difference circuit, for example, an event indicating an increase in intensity may occur as the image of the moving object 1202 moves across pixels. Conversely, an event indicating a decrease in intensity may occur as the image of the moving object 1202 passes a certain pixel. The pixel pattern, accompanied by increase and decrease events, may be used as a reliable indication of the motion of the moving object 1202, which can be rapidly updated with short latency due to a relatively small amount of data indicating the events. In concrete examples, such a system could lead to realistic XR systems that track a user's hand, modify the rendering of virtual objects, and create a sense for the user that the user is interacting with virtual objects.
[0144] The position of the patch may change for other reasons, and any or all of these may be reflected in the trajectory calculation. One such other change is the user's movement while wearing the image sensor. Figure 13 depicts an embodiment of a moving user that creates a changing viewpoint with respect to the user and the image sensor. In Figure 13, the user may initially be looking straight at an object with viewpoint 1302. In this configuration, the pixel array 1300 of the image array will capture an object in front of the user. The object in front of the user may be within patch 1312.
[0145] The user can then change their viewpoint by turning their head, for example. The viewpoint may change to viewpoint 1304. An object that was previously directly in front of the user will have a different position within the user's field of view at viewpoint 1304, even if it does not move. It will also be at a different point within the field of view of the image sensor worn by the user, and therefore at a different position within the image array 1300. The object may be contained within a patch at location 1314, for example.
[0146] If the user further changes their viewpoint to viewpoint 1306, and the image sensor moves with the user, the location of the object that was previously directly in front of the user will be imaged at a different point within the field of view of the image sensor worn by the user, and therefore at a different position within the image array 1300. That object may be contained within a patch at location 1316, for example.
[0147] As can be seen, as the user further changes their viewpoint, the position of the patch in the image array, which is required to capture the object, moves further. The trajectory of this movement from location 1312 to location 1314 and location 1316 may be estimated and used to track the future position of the patch.
[0148] The trajectory may be estimated by other means. For example, when the user has viewpoint 1302, measurements using an inertial sensor may indicate the acceleration and velocity of the user's head. This information may be used to predict the trajectory of a patch in an image array based on the user's head movement.
[0149] The patch trajectory estimate 906 can, at least partially, predict, based on these inertial measurements, that the user will have viewpoint 1304 at time t1 and viewpoint 1306 at time t2. Therefore, the patch trajectory estimate 906 can predict that patch 1308 may move to patch 1310 at time t1 and to patch 1312 at time t2.
[0150] As an embodiment of such an approach, it may be used to provide accurate and low-latency estimation of head posture within an AR system. The patch may be positioned to encompass an image of a stationary object in the user's environment. In a specific embodiment, the processing of image information may identify the corner of a photographic frame suspended on a wall as a recognizable and stationary object for tracking. The processing may center the patch on that object. As with the case of the moving object 1202 described above in relation to Figure 12, the relative movement between the object and the user's head will produce an event that can be used to calculate the relative motion between the user and the tracked object. In this embodiment, since the tracked object is stationary, the relative motion represents the motion of an imaging array worn by the user. This motion, therefore, represents a change in the user's head posture relative to the physical world and can be used to maintain an accurate calculation of the user's head posture, which can be used when rendering virtual objects realistically. Because imaging arrays such as those described herein can provide high-speed updates using relatively small amounts of data per update, calculations for rendering virtual objects remain accurate (they can be performed quickly and updated frequently).
[0151] Referring back to Figure 11, the patch trajectory estimation 906 may include, at least in part, a step (action 1106) of adjusting the size of at least one of the patches based on the calculated patch trajectory. For example, the patch size may be set to be large enough to include pixels onto which an image of a movable object or at least a portion of an object from which image information should be generated should be projected. The patch may be set to be slightly larger than the projected size of the image of the portion of the object of interest so that, if any errors exist in estimating the patch trajectory, the patch can still include the relevant portion of the image. As the object moves relative to the image sensor, the size of the image of that object in pixels may change based on distance, angle of incidence, object orientation, or other factors. The processor defining the patch associated with the object may set the patch size by measuring the size of the patch associated with the object based on other sensor data, or by calculating it based on a world model, etc. Other parameters of the patch, such as its shape, may also be set or updated.
[0152] Figure 14 depicts an image sensing system 1400 configured for use in an XR system according to several embodiments. As in the image sensing system 400 (Figure 4), the image sensing system 1400 includes a network for selectively outputting values within a patch and may also be configured to output events relating to pixels within a patch, as similarly described above. In addition, the image sensing system 1400 may be configured to selectively output measured intensity values, which may be output for a complete image frame.
[0153] In the illustrated embodiment, separate outputs are shown for events and intensity values, generated using the DVS technique described above. The output generated using the DVS technique may be output as AER1418 using the expression described above in relation to AER418. The output representing the intensity value may be output through an output designated here as APS1420. These intensity outputs may relate to a patch or to an entire image frame. The AER and APS outputs may be active simultaneously. However, in the illustrated embodiment, the image sensor 1400 operates in a mode for outputting events or in a mode for outputting intensity information at any given time. Using such an image sensor, the system may selectively use event outputs and / or intensity information.
[0154] The image sensing system 1400 may include an image sensor 1402, which may include an image array 1404, which may contain a plurality of pixels 1500, each responding to light. The sensor 1402 may further include a network for accessing pixel cells. The sensor 1402 may further include a network that generates inputs to the access network to control the mode in which information is read from pixel cells in the image array 1404.
[0155] In the illustrated embodiment, the image array 1404 is configured as an array comprising multiple rows and columns of pixel cells, both accessible in read mode. In such an embodiment, the access network may include a row-address encoder / decoder 1406, a column-address encoder / decoder 1408 that controls a column selection switch 1422, and / or registers 1424 that can temporarily hold information about incident light sensed by one or more corresponding pixel cells. The patch tracking engine 1410 may generate inputs to the access network from time to time to control the pixel cells providing image information.
[0156] In some embodiments, the image sensor 1402 may be configured to operate in roll shutter mode, global shutter mode, or both. For example, the patch tracking engine 1410 may generate inputs to the access network and control the reading mode of the image array 1402.
[0157] When sensor 1402 operates in roll shutter reading mode, a single column of pixel cells is selected between each system clock, for example, by closing a single-column switch 1422 of a multi-column switch. During that system clock, the selected column of pixel cells is exposed and read by APS 1420. To generate an image frame by roll shutter mode, the columns of pixel cells in sensor 1402 are read one column at a time and then processed by the image processor to generate an image frame.
[0158] When sensor 1402 operates in global shutter mode, rows of pixel cells are exposed simultaneously, for example, within a single system clock, so that information captured by pixel cells in multiple rows can be read simultaneously by APS1420b, and the information is stored in register 1424. Such a reading mode allows for direct output of image frames without requiring further data processing. In the illustrated embodiment, information about incident light sensed by pixel cells is stored in separate registers 1424. It should be understood that multiple pixel cells may share a single register 1424.
[0159] In some embodiments, the sensor 1402 may be implemented within a single integrated circuit, such as a CMOS integrated circuit. In some embodiments, the image array 1404 may be implemented within a single integrated circuit. The patch tracking engine 1410, row address encoder / decoder 1406, column address encoder / decoder 1408, column selection switch 1422, and / or register 1424 may be implemented within a second single integrated circuit, configured, for example, as a driver for the image array 1404. Alternatively, or in addition, some or all of the functions of the patch tracking engine 1410, row address encoder / decoder 1406, column address encoder / decoder 1408, column selection switch 1422, and / or register 1424 may be distributed to other digital processors in the AR system.
[0160] Figure 15 illustrates an exemplary pixel cell 1500. In the illustrated embodiment, each pixel cell may be configured to output either event or intensity information. However, it should be understood that in some embodiments, the image sensor may be configured to output both types of information in parallel.
[0161] Both event information and intensity information are based on the output of the photodetector 504, as described above in relation to Figure 5A. The pixel cell 1500 includes a network for generating event information. This network includes a photosensitive circuit 502, a difference circuit 506, and a comparator 508, as also described above. When the switch 1520 is in a first state, it connects the photodetector 504 to the event generation network. The switch 1520 or other control network may be controlled by a processor that controls the AR system so that a relatively small amount of image information is provided for the substantial time period during which the AR system is operating.
[0162] Switch 1520 or other control networks may also be controlled to configure the pixel cell 1500 to output intensity information. In the illustrated information, the intensity information is provided as a complete image frame, represented sequentially as a stream of pixel intensity values for each pixel in the image array. To operate in this mode, the switch 1520 in each pixel cell may be set to a second position, exposing the output of the photodetector 504 so that it can be connected to an output line after passing through the amplifier 510.
[0163] In the illustrated embodiment, the output line is illustrated as a column line 1510. There may be one such column line for each column in the image array. Each pixel cell in a column may be coupled to a column line 1510, but the pixel array may be controlled so that only one pixel cell is coupled to a column line 1510 at a time. One such switch exists in each pixel cell; switch 1530 controls when a pixel cell 1500 is coupled to its individual column line 1510. An access network, such as a row address decoder 410, may close switch 1530 to ensure that only one pixel cell is coupled to each column line at a time. Switches 1520 and 1530 may be implemented using one or more transistors or similar components that are part of the image array.
[0164] Figure 15 shows further components that may be included within each pixel cell according to several embodiments. A sample-and-hold circuit (S / H) 1532 may be connected between the photodetector 504 and the column line 1510. When present, the S / H 1532 may enable the image sensor 1402 to operate in global shutter mode. In global shutter mode, trigger signals are sent in parallel to each pixel cell in the array. Within each pixel cell, the S / H 1532 captures a value indicating the intensity of the trigger signal at a given time. The S / H 1532 stores this value and generates an output based on it until the next value is captured.
[0165] As shown in Figure 15, the signal representing the value stored by S / H1532 may be coupled to column line 1510 when switch 1530 is closed. The signal coupled to the column line may be processed to produce the output of the image array. This signal may be buffered and / or amplified in amplifier 1512, for example, at the end of column line 1510, and then applied to analog-to-digital converter (A / D) 1514. The output of A / D 1514 may pass through another reading circuit 1516 to output 1420. The reading circuit 1516 may include, for example, column switch 1422. Other components in the reading circuit 1516 may perform other functions, such as serializing the multibit output of A / D 1514.
[0166] Those skilled in the art will understand how to implement a circuit for performing the functions described herein. The S / H1532 may be implemented, for example, as one or more capacitors and one or more switches. However, it should be understood that the S / H1532 may be implemented using other components or in circuit configurations other than those illustrated in Figure 15A. It should also be understood that other components other than those illustrated may also be implemented. For example, Figure 15 shows one amplifier and one A / D converter per column. In other embodiments, there may be one A / D converter shared across multiple columns.
[0167] In a pixel array configured for a global shutter, each S / H1532 may store intensity values that reflect image information at the same moment. These values may be stored as the values stored within each pixel are read sequentially during the reading phase. Sequential reading may be achieved, for example, by connecting the S / H1532 of each pixel cell in a row to its individual column line. The values on the column line may then pass to the APS output 1420, one at a time. Such information flow may be controlled by sequencing the opening and closing of the column switch 1422. This operation may be controlled, for example, by the column address decoder 1408. Once the values for each pixel in one row have been read, the pixel cells in the next row may be connected to the column line in their location. Their values may be read one column at a time. The process of reading values for one row at a time may be repeated until intensity values for all pixels in the image array have been read. In embodiments where intensity values are read for one or more patches, the process would be completed once values for pixel cells within a patch have been read.
[0168] Pixel cells may be read in any preferred order. Rows may be interleaved, for example, so that every other row is read sequentially. The AR system can still process the image data as frames of image data by deinterleaving the data.
[0169] In embodiments where S / H1532 is absent, values may still be read sequentially from each pixel cell as the rows and columns of values are scanned and output. However, the value read from each pixel cell may represent the intensity of light detected by the cell's photodetector at the time the value in that cell was captured as part of the reading process, for example, when the value is applied to A / D1514. As a result, in a roll shutter, pixels in an image frame may represent images incident on the image array at slightly different times. For an image sensor outputting a complete frame at a rate of 30 Hz, the time difference between when the first pixel value for a frame is captured and when the last pixel value for a frame is captured can be as little as 1 / 30 of a second, which is imperceptible for many applications.
[0170] For some XR functions, such as object tracking, the XR system may use a roll shutter to perform calculations on image information acquired using an image sensor. Such calculations may interpolate between consecutive image frames, and for each pixel, an interpolated value may be calculated that represents the estimated value of the pixel at a point in time between consecutive frames. The same time may be used for all pixels so that, through the calculation, the interpolated image frame contains pixels representing the same point in time that could be produced using an image sensor with a global shutter. Alternatively, a global shutter image array may be used for one or more image sensors in a wearable device that form part of the XR system. A global shutter for a complete or partial image frame may avoid interpolation of other processes that can be performed to compensate for variations in capture time in the image information captured using a roll shutter. Interpolation calculations can therefore be avoided even when image information is used to track the movement of an object, such as a hand or other movable object, or to determine the head posture of a user of a wearable device in an AR system, or further, to process the movement of an object using a camera on a wearable device, which may be moving when the image information is collected, in order to construct an accurate representation of the physical environment.
[0171] Distinguished pixel cells
[0172] In some embodiments, the pixel cells in the sensor array may be identical. Each pixel cell may respond to, for example, a broad spectrum of visible light. Each photodetector can therefore provide image information indicating the intensity of visible light. In this scenario, the output of the image array may be a "grayscale" output indicating the amount of visible light incident on the image array.
[0173] In other embodiments, pixel cells may be distinguished. For example, different pixel cells in a sensor array may output image information indicating the intensity of light in a particular part of the spectrum. A preferred technique for distinguishing pixel cells is to position a filter element in the optical path leading to a photodetector within the pixel cell. The filter element may be, for example, a bandpass filter that allows visible light of a particular color to pass through. Applying such a color filter across a pixel cell configures that pixel cell to provide image information indicating the intensity of light of the color corresponding to the filter.
[0174] Filters may be applied across pixel cells regardless of the structure of the pixel cells. They may be applied across pixel cells in a sensor array, for example, with a global shutter or a roll shutter. Similarly, filters may be applied to pixel cells configured to output intensity or intensity changes using DVS techniques.
[0175] In some embodiments, filter elements that selectively allow primary color light to pass through may be mounted across the photodetectors in each pixel cell of the sensor array. For example, filters that selectively allow red, green, or blue light to pass through may be used. The sensor array may have multiple subarrays, each subarray having one or more pixels configured to sense each of the primary colors of light. Thus, the pixel cells in each subarray provide both intensity and color information about the object being imaged by the image sensor.
[0176] The inventors recognized and appreciated the value of grayscale cameras in XR systems, recognizing that while some functions require color information, others can be performed using grayscale information. A wearable device equipped with image sensors to provide image information relating to the operation of an XR system may have multiple cameras, some of which may be formed with image sensors capable of providing color information. Other cameras may be grayscale cameras. The inventors recognized and appreciated the value of grayscale cameras, recognizing that they can consume less power, be more sensitive in low-light conditions, output data faster, and / or output less data to represent the same range of the physical world with the same resolution as a camera formed with a comparable image sensor configured to sense color. However, grayscale cameras can output sufficient image information for many functions performed within an XR system. Therefore, an XR system may be configured primarily with grayscale cameras or with multiple cameras, or with both grayscale and color cameras, or with the selective use of color cameras.
[0177] For example, an XR system may collect and process image information to create a passable world model. This processing may use color information, which can improve the effectiveness of several functions, such as distinguishing objects, identifying surfaces associated with the same object, and / or recognizing objects. Such processing may be performed or updated from time to time, for example, when the user moves to a new environment, such as by first turning on the system or walking into another room, or when a change in the user's environment is detected in a different way.
[0178] Other functions are not significantly improved through the use of color information. For example, once a passable world model is created, the XR system may use images from one or more cameras to determine the orientation of the wearable device to features in the passable world model. Such a function may be performed, for example, as part of head pose tracking. Some or all of the cameras used for such a function may be grayscale. Since head pose tracking is performed frequently as the XR system operates, in some embodiments, the continuous use of one or more grayscale cameras for this function may provide considerable power savings, reduced computation, or other advantages.
[0179] Similarly, during multiple XR system operations, the system may use stereoscopic information from two or more cameras to determine the distance to a movable object. Such functionality may require high-rate processing of image information as part of tracking the user's hand or other movable objects. Using one or more grayscale cameras for this functionality may provide lower latency or other advantages associated with processing high-resolution image information.
[0180] In some embodiments of the XR system, the XR system may have both a color camera and at least one grayscale camera, and the grayscale and / or color camera may be selectively enabled based on the function for which the image information from those cameras should be used.
[0181] Pixel cells within an image sensor may be distinguished in ways other than based on the spectrum of light to which the pixel cell is sensitive. In some embodiments, some or all of the pixel cells may produce an output having an intensity that indicates the angle of arrival of light incident on the pixel cell. The angle of arrival information may be processed to calculate the distance to the object being imaged.
[0182] In such embodiments, the image sensor may passively acquire depth information. Passive depth information may be acquired by positioning a component in the optical path to the pixel cells in the array such that the pixel cell outputs information indicating the angle of arrival of light striking the pixel cell. An embodiment of such a component is a transmissive diffraction mask (TDM) filter.
[0183] The angle of arrival information may be converted through calculation into distance information, which indicates the distance to the object from which the light is reflected. In some embodiments, the pixel cells configured to provide the angle of arrival information may be scattered with pixel cells that capture the light intensity of one or more colors. As a result, the angle of arrival information, and therefore the distance information, may be combined with other image information about the object.
[0184] In some embodiments, one or more of the sensors may be configured to acquire information about physical objects in a scene at high frequencies with short latency, using compact, low-power components. The image sensor may draw, for example, less than 50 mW, allowing the device to be powered by a battery small enough to be used as part of a wearable system. The sensor may also be an image sensor configured to passively acquire depth information in addition to, or instead of, image information, indicating the intensity and / or changes in intensity information of one or more colors. Such a sensor may also be configured to provide small amounts of data by using patch tracking or by using DVS techniques to provide differential output.
[0185] Passive depth information may be obtained by configuring an image array, such as an image array, which incorporates any one or more of the techniques described herein, along with a component that adapts one or more of the pixel cells in the array to output information indicating the light field emanating from the imaged object. The information may be based on the angle of arrival of light striking the pixel. In some embodiments, a pixel cell, such as those described above, may be configured to output an indication of the angle of arrival by placing a prenoptic component in the light path to the pixel cell. An embodiment of the prenoptic component is a transmissive diffraction mask (TDM). The angle of arrival information may be converted through computation into distance information indicating the distance from there to the object from which the light is reflected, forming a captured image. In some embodiments, the pixel cells configured to provide the angle of arrival information may be scattered with pixel cells that capture light intensity in grayscale or in one or more colors. As a result, the angle of arrival information may also be combined with other image information about the object.
[0186] Figure 16 illustrates a pixel subarray 100 according to several embodiments. In the illustrated embodiments, the subarray has two pixel cells, but the number of pixel cells in the subarray is not a limitation of the present invention. Here, a first pixel cell 121 and a second pixel cell 122 are shown, one of which is configured to capture arrival angle information (first pixel cell 121), but it should be understood that the number and location of pixel cells in the array can vary. In this embodiment, the other pixel cell (second pixel cell 122) is configured to measure the intensity of one color of light, but other configurations are also possible and may include pixel cells that are sensitive to different colors of light, or one or more pixel cells that are sensitive to a wide spectrum of light, such as in a grayscale camera.
[0187] The first pixel cell 121 of the pixel subarray 100 in Figure 16 includes an arrival angle / intensity converter 101, a photodetector 105, and a difference reading network 107. The second pixel cell 122 of the pixel subarray 100 includes a color filter 102, a photodetector 106, and a difference reading network 108. It should be understood that not all components illustrated in Figure 16 are included in all embodiments. For example, some embodiments may not include the difference reading network 107 and / or 108, and some embodiments may not include the color filter 102. Furthermore, additional components not shown in Figure 16 may be included. For example, some embodiments may include a polarizer arranged to allow light of a specific polarization to reach the photodetector. In another embodiment, some embodiments may include a scanning output network instead of, or in addition to, the difference reading network 107. In another embodiment, the first pixel cell 121 may also include a color filter such that the first pixel 121 measures both the arrival angle and intensity of a particular color of light incident on the first pixel 121.
[0188] The arrival angle / intensity converter 101 of the first pixel 121 is an optical component that converts the angle θ of incident light 111 into an intensity that can be measured by a photodetector. In some embodiments, the arrival angle / intensity converter 101 may include a refractive optical system. For example, one or more lenses may be used to convert the angle of incidence of light into a position on the image plane, the amount of which is detected by one or more pixel cells. In some embodiments, the arrival angle / position intensity converter 101 may include a diffraction optical system. For example, one or more diffraction gratings (e.g., TDMs) may convert the angle of incidence of light into an intensity that can be measured by a photodetector below the TDM.
[0189] The photodetector 105 of the first pixel cell 121 receives incident light 110 passing through the arrival angle / intensity converter 101 and generates an electrical signal based on the intensity of the light incident on the photodetector 105. The photodetector 105 is located in the image plane and is associated with the arrival angle / intensity converter 101. In some embodiments, the photodetector 105 may be a single pixel of an image sensor, such as a CMOS image sensor.
[0190] The first pixel 121 difference reading network 107 receives a signal from the photodetector 105 and outputs an event only when the amplitude of the electrical signal from the photodetector is different from the amplitude of the previous signal from the photodetector 105, implementing the DVS technique as described above.
[0191] The second pixel cell 122 includes a color filter 102 for filtering the incident light 112 such that only light within a specific wavelength range passes through the color filter 102 and is incident on the photodetector 106. The color filter 102 may be a band-pass filter that allows, for example, one of red, green, or blue light to pass through it and rejects light of other wavelengths, and / or restricts the IR light reaching the photodetector 106 to only a specific portion of the spectrum.
[0192] In this embodiment, the second pixel cell 122 also includes a photodetector 106 and a difference reading network 108, which may function similarly to the photodetector 105 and difference reading network 107 of the first pixel cell 121.
[0193] As described above, in some embodiments, the image sensor may include an array of pixels, each pixel being associated with a photodetector and a reading circuit. A subset of pixels may be associated with an angle-of-arrival / intensity converter, used to determine the angle of detected light incident on the pixel. Other subsets of pixels may be associated with color filters, used to determine color information about the scene being observed, or may selectively pass or block light based on other properties.
[0194] In some embodiments, the angle of arrival of light may be determined using a single photodetector and diffraction gratings at two different depths. For example, light may be incident on a first TDM that converts the angle of arrival into position, and a second TDM may be used to selectively pass light incident at a particular angle. Such an arrangement may utilize the Talbot effect, a near-field diffraction effect, in which, when a plane wave is incident on a diffraction grating, an image of the diffraction grating is created at a distance from the grating. If the second diffraction grating is positioned in the image plane where the image of the first diffraction grating is formed, the angle of arrival may be determined from the intensity of light measured by a single photodetector positioned behind the second grating.
[0195] Figure 17A illustrates a first array of pixel cells 140, including a first TDM141 and a second TDM143, which are aligned with each other such that the increased refractive index bumps and / or regions with respect to the two gratings are aligned horizontally (Δs=0) (where Δs is the horizontal offset between the first TDM141 and the second TDM143). Both the first TDM141 and the second TDM143 may have the same grating period d, and the two gratings may be separated by distance / depth z. The depth z, known as the Talbot length, to which the second TDM143 is located relative to the first TDM141, can be determined by the grating period d and wavelength λ of the light being analyzed, and is given by the following equation: [ka]
[0196] As shown in Figure 17A, incident light 142 with an arrival angle of zero degrees is diffracted by the first TDM 141. The second TDM 143 is positioned at a depth equal to the Talbot length such that an image of the first TDM 141 is created, resulting in most of the incident light 142 passing through the second TDM 143. An optional dielectric layer 145 may separate the second TDM 143 from the photodetector 147. As light passes through the dielectric layer 145, the photodetector 147 detects the light and generates an electrical signal with properties (e.g., voltage or current) that are proportional to the intensity of the light incident on the photodetector. On the other hand, incident light 144 with a non-zero arrival angle θ is also diffracted by the first TDM 141, but the second TDM 143 prevents at least a portion of the incident light 144 from reaching the photodetector 147. The amount of incident light reaching the photodetector 147 depends on the arrival angle θ; at larger angles, less light reaches the photodetector. The dashed line originating from light 144 illustrates the attenuation of the amount of light reaching the photodetector 147. In some cases, light 144 may be completely blocked by the diffraction grating 143. Therefore, information about the arrival angle of the incident light may be obtained using a single photodetector 147 with two TDMs.
[0197] In some embodiments, information acquired by adjacent pixel cells without an arrival angle / intensity converter may provide an indication of the intensity of the incident light and may be used to determine the portion of the incident light that passes through the arrival angle / intensity converter. From this image information, the arrival angle of the light detected by the photodetector 147 may be calculated as described in more detail below.
[0198] Figure 17B illustrates a second array of pixel cells 150, including a first TDM151 and a second TDM153, in which the increased refractive index bulges and / or regions with respect to the two gratings are misaligned with each other such that they are not aligned horizontally (Δs ≠ 0) (where Δs is the horizontal offset between the first TDM151 and the second TDM153). Both the first TDM151 and the second TDM153 may have the same grating period d, and the two gratings may be separated by distance / depth z. Unlike the situation discussed in relation to Figure 17A, where the two TDMs are aligned, the misalignment results in incident light passing through the second TDM153 at angles other than zero.
[0199] As shown in Figure 17B, incident light 152 with an arrival angle of zero degrees is diffracted by the first TDM 151. The second TDM 153 is located at a depth equal to the Talbot length, but due to the horizontal offset of the two gratings, at least a portion of the light 152 is blocked by the second TDM 153. The dashed line resulting from the light 152 illustrates the attenuation of the amount of light reaching the photodetector 157. In some cases, the light 152 may be completely blocked by the diffraction grating 153. On the other hand, incident light 154 with a non-zero arrival angle θ is diffracted by the first TDM 151 but passes through the second TDM 153. After crossing the optional dielectric layer 155, the photodetector 157 detects the light incident on the photodetector 157 and generates an electrical signal with properties (e.g., voltage or current) that are proportional to the intensity of the light incident on the photodetector.
[0200] Pixel cells 140 and 150 have different output functions, with different intensities of detected light for different angles of incidence. However, in either case, the relationship may be fixed and determined based on the pixel cell design or by measurement as part of a calibration process. Regardless of the precise transfer function, the measured intensity may be converted to an arrival angle, which may then be used to determine the distance to the imaged object.
[0201] In some embodiments, different pixel cells of an image sensor may have different TDM arrays. For example, a first subset of pixel cells may include a first horizontal offset between the grids of two TDMs associated with each pixel, while a second subset of pixel cells may include a second horizontal offset between the grids of two TDMs associated with each pixel cell, where the first offset is different from the second offset. Each subset of pixel cells with different offsets may be used to measure different arrival angles or ranges of different arrival angles. For example, a first subset of pixels may include an array of TDMs similar to pixel cell 140 in Figure 17A, and a second subset of pixels may include an array of TDMs similar to pixel cell 150 in Figure 17B.
[0202] In some embodiments, not all pixel cells of an image sensor include a TDM. For example, a subset of pixel cells may include a color filter, while a different subset of pixel cells may include a TDM for determining the angle of arrival information. In other embodiments, the color filter is not used, such that a first subset of pixel cells simply measures the total intensity of incident light, and a second subset of pixel cells measures the angle of arrival information. In some embodiments, information about the intensity of light from neighboring pixel cells without a TDM may be used to determine the angle of arrival for light incident on a pixel cell with one or more TDMs. For example, using two TDMs arranged to utilize the Talbot effect, the intensity of light incident on the photodetector after the second TDM is the synusoid function of the angle of arrival of light incident on the first TDM. Therefore, if the total intensity of light incident on the first TDM is known, the angle of arrival of the light may be determined from the intensity of light detected by the photodetector.
[0203] In some embodiments, the configuration of pixel cells within a subarray may be selected to provide various types of image information with appropriate resolution. Figures 18A–C illustrate exemplary arrangements of pixel cells within a pixel subarray of an image sensor. It should be understood that the illustrated embodiments are non-limiting and alternative pixel arrangements may also be considered by the inventors. This arrangement may be repeated across the image array and may contain millions of pixels. The subarray may include one or more pixel cells that provide arrival angle information for incident light, and one or more other pixel cells (with or without color filters) that provide intensity information for incident light.
[0204] Figure 18A shows an embodiment of a pixel subarray 160, comprising a first set of pixel cells 161 and a second set of pixel cells 163, which are different from each other and are rectangular rather than square. A pixel cell labeled "R" is a pixel cell with a red filter so that red incident light passes through the filter to the associated photodetector; a pixel cell labeled "B" is a pixel cell with a blue filter so that blue incident light passes through the filter to the associated photodetector; and a pixel cell labeled "G" is a pixel with a green filter so that green incident light passes through the filter to the associated photodetector. The exemplary subarray 160 illustrates that there are more green pixel cells than red or blue pixel cells, and that the different types of pixel cells do not need to be present in equal proportions.
[0205] Pixel cells labeled A1 and A2 are pixels that provide arrival angle information. For example, pixel cells A1 and A2 may include one or more grids for determining the arrival angle information. Pixel cells providing arrival angle information may be configured similarly or differently, such as being sensitive to different ranges of arrival angles or arrival angles relative to different axes. In some embodiments, pixels labeled A1 and A2 include two TDMs, and the TDMs of pixel cells A1 and A2 may be oriented in different directions, for example, perpendicular to each other. In other embodiments, the TDMs of pixel cells A1 and A2 may be oriented parallel to each other.
[0206] In embodiments using a pixel subarray 160, both color image data and angle of arrival information may be acquired. To determine the angle of arrival of light incident on a set of pixel cells 161, the total light intensity incident on set 161 is estimated using electrical signals from the RGB pixel cells. Using the fact that the light intensity detected by A1 / A2 pixels fluctuates in a predictable manner as a function of the angle of arrival, the angle of arrival may be determined by comparing the total intensity (estimated from the RGB pixel cells in the group of pixels) with the intensity measured by the A1 and / or A2 pixel cells. For example, the light intensity incident on A1 and / or A2 pixels may fluctuate sinusoidally with respect to the angle of arrival of the incident light. The angle of arrival of light incident on a set of pixel cells 163 is determined by a similar method using electrical signals generated by the pixels of set 163.
[0207] Figure 18A shows a specific embodiment of the subarray, and it should be understood that other configurations are also possible. In some embodiments, for example, the subarray may consist of only a set of pixel cells 161 or 163.
[0208] Figure 18B shows an alternative pixel subarray 170, comprising a first set of pixel cells 171, a second set of pixel cells 172, a third set of pixel cells 173, and a fourth set of pixel cells 174. Each set of pixel cells 171-174 is square and contains identical arrangements of pixel cells, but may have pixel cells for determining arrival angle information spanning different ranges of angles or relative to different planes (for example, the TDMs of pixels A1 and A2 may be oriented perpendicular to each other). Each set of pixels 171-174 contains one red pixel cell (R), one blue pixel cell (B), one green pixel cell (G), and one arrival angle pixel cell (A1 or A2). Note that in the exemplary pixel subarray 170, an equal number of red / green / blue pixel cells are present in each set. Furthermore, understand that pixel subarrays may be repeated in one or more directions to form larger arrays of pixels.
[0209] In embodiments using the pixel subarray 170, both color image data and angle of arrival information may be acquired. To determine the angle of arrival of light incident on a set of pixel cells 171, the total light intensity incident on set 171 may be estimated using signals from RGB pixel cells. The angle of arrival may be determined by comparing the total intensity (estimated from RGB pixel cells) with the intensity measured by the A1 pixel, using the fact that the light intensity detected by the angle of arrival pixel cell has a sinusoidal or other predictable response to the angle of arrival. The angles of arrival of light incident on sets of pixel cells 172-174 may be determined in a similar manner using electrical signals generated by the pixel cells of each individual set of pixels.
[0210] Figure 18C shows an alternative pixel subarray 180, which includes a first set of pixel cells 181, a second set of pixel cells 182, a third set of pixel cells 183, and a fourth set of pixel cells 184. Each set of pixel cells 181-184 is square and contains identically arranged pixel cells, and no color filters are used. Each set of pixel cells 181-184 includes two "white" pixels (e.g., without color filters, so that red, blue, and green light are detected to form a grayscale image), one arrival angle pixel cell (A1) with TDM oriented in a first direction, and one arrival angle pixel cell (A2) with TDM oriented with a second spacing or in a second direction (e.g., perpendicular to the first direction). Note that in the exemplary pixel subarray 170, no color information is present. The resulting image is grayscale, and passive depth information can be obtained within a color or grayscale image array using techniques as described herein. Similar to other subarray configurations described herein, the pixel subarray array may be repeated in one or more directions to form a larger array of pixels.
[0211] In embodiments using the pixel subarray 180, both grayscale image data and angle of arrival information may be acquired. To determine the angle of arrival of light incident on a set of pixel cells 181, the total light intensity incident on set 181 is estimated using electrical signals from two white color pixels. The angle of arrival may be determined by comparing the total intensity (estimated from the white pixels) with the intensity measured by the A1 and / or A2 pixel cells, using the fact that the light intensity detected by pixels A1 and A2 has a sinusoidal or other predictable response to the angle of arrival. The angle of arrival of light incident on sets of pixel cells 182-184 may be determined in a similar manner using electrical signals generated by the pixels of each individual set of pixels.
[0212] In the above embodiment, the pixel cells are illustrated as squares and are arranged in a square grid. The embodiment is not limited thereto. For example, in some embodiments, the pixel cells may be rectangular in shape. Furthermore, the subarrays may be triangular, or arranged diagonally, or have other geometric shapes.
[0213] In some embodiments, arrival angle information is obtained using an image processor 708 or a processor associated with a local data processing module 70, which may further determine the distance of an object based on the arrival angle. For example, the arrival angle information may be combined with one or more other types of information to obtain the distance of an object. In some embodiments, objects in the mesh model 46 may be associated with arrival angle information from a pixel array. The mesh model 46 may include the location of an object, including its distance from the user, which may be updated with a new distance value based on the arrival angle information.
[0214] Using angle of arrival information to determine distance values can be particularly useful in scenarios where the object is close to the user. This is because a change in distance from the image sensor results in a change in the angle of arrival of light for nearby objects that is greater than a similar change in distance for objects located farther from the user. Therefore, a processing module that utilizes passive distance information based on the angle of arrival may selectively use that information based on the estimated distance of the object, and may use one or more other techniques to determine the distance to an object that exceeds a threshold distance such as up to 1 meter, up to 3 meters, or up to 5 meters in some embodiments. As a specific example, a processing module of an AR system may be programmed to use passive distance measurement using angle of arrival information for objects within 3 meters of the user of the wearable device, but for objects outside that range, stereoscopic image processing may be used using images captured by two cameras.
[0215] Similarly, a pixel configured to detect arrival angle information may be most sensitive to changes in distance within a certain angular range from the normal to the image array. The processing module may also use distance information derived from arrival angle measurements within that angular range, but may be configured to use other sensors and / or other techniques to determine distances outside that range.
[0216] One exemplary application of determining the distance to an object from an image sensor is hand tracking. Hand tracking may be used in an AR system, for example, to provide a gesture-based user interface for system 80 and / or to allow the user to move virtual objects in the environment in the AR experience provided by system 80. A combination of an image sensor that provides angle of arrival information for accurate depth determination and a differential reading network to reduce the amount of data to the process of determining the movement of the user's hand provides an efficient interface, thereby allowing the user to interact with virtual objects and / or provide input to system 80. The processing module that determines the location of the user's hand may use distance information obtained using different techniques depending on the location of the user's hand within the field of view of the image sensor of the wearable device. Hand tracking may be implemented in the form of patch tracking during the image sensing process, according to some embodiments.
[0217] Another application where depth information may be useful is occlusion processing. Occlusion processing uses depth information to determine that a portion of a model of the physical world does not need to be updated, or cannot be updated, based on image information captured by one or more image sensors that collect image information about the physical environment around the user. For example, if it is determined that a first object is located at a first distance from the sensor, the system 80 may decide not to update the model of the physical world over distances beyond the first distance. For example, even if the model includes a second object at a second distance from the sensor, and the second distance is greater than the first distance, the model information about that object does not need to be updated if it is behind the first object. In some embodiments, the system 80 may generate an occlusion mask based on the location of the first object and update only the portions of the model that are not masked by the occlusion mask. In some embodiments, the system 80 may generate one or more occlusion masks for one or more objects. Each occlusion mask may be associated with a separate distance from the sensor. For each occlusion mask, model information associated with objects that are at a distance from the sensor exceeding the distance associated with that individual occlusion mask will not be updated. By limiting the parts of the model that are updated at any given time, the speed at which the AR environment is generated and the amount of computational resources required to generate the AR environment are reduced.
[0218] Although not shown in Figures 18A-C, some embodiments of the image sensor may include pixels with an IR filter in addition to, or instead of, a color filter. For example, an IR filter may allow light of a wavelength such as approximately equal to 940 nm to pass through and be detected by an associated photodetector. Some embodiments of the wearable may include an IR light source (e.g., an IR LED) that emits light of the same wavelength as the associated IR filter (e.g., 940 nm). The IR light source and IR pixels may be used as an alternative method for determining the distance of an object from the sensor. As an example, but not limited to, the IR light source may be pulsed, and time-of-flight measurement may be used to determine the distance of an object from the sensor.
[0219] In some embodiments, the system 80 may be capable of operating in one or more operating modes. A first mode may be one in which depth determination is performed using passive depth measurement, for example, based on the angle of arrival of light determined using a pixel with an angle of arrival / intensity converter. A second mode may be one in which depth determination is performed using active depth measurement, for example, based on the time of flight of IR light measured using an IR pixel of an image sensor. A third mode may be one in which the distance of an object is determined using stereoscopic measurements from two separate image sensors. Such stereoscopic measurements may be more accurate than using the angle of arrival of light determined using a pixel with an angle of arrival / intensity converter when the object is very far from the sensor. Other preferred methods for determining depth may also be used for one or more additional operating modes for depth determination.
[0220] In some embodiments, passive depth determination may be preferred because such techniques utilize less power. However, the system may decide that it should operate in active mode under certain conditions. For example, if the intensity of visible light detected by the sensor falls below a threshold, it may be too dark to accurately perform passive depth determination. In some embodiments, IR illumination may be used to compensate for the effects of low light. IR illumination may be selectively enabled in response to detected low light conditions that prevent the device, such as head tracking or hand tracking, from obtaining sufficient images for the task it is performing. In another embodiment, the object may be too far away for accurate passive depth determination. Therefore, the system may be programmed to choose to operate in a third mode in which depth is determined based on a stereoscopic measurement of the scene using two spatially separated image sensors. In yet another embodiment, determining the depth of an object based on the angle of arrival of light determined using pixels with an angle of arrival / intensity converter may be inaccurate at the edges of the image sensor. Therefore, if an object is detected by pixels near the periphery of the image sensor, the system may choose to operate in a second mode using active depth determination.
[0221] The image sensor embodiments described above use individual pixel cells with stacked TDMs to determine the angle of arrival of light incident on the pixel cells, but other embodiments may use groups of multiple pixel cells with a single TDM across all pixels in the group to determine the angle of arrival information. The TDM may project a pattern of light across the sensor array, and this pattern depends on the angle of arrival of the incident light. Multiple photodetectors associated with one TDM can more accurately detect the pattern because each photodetector is located at a different position within the image plane (the image plane containing light-sensing photodetectors). The relative intensity sensed by each photodetector may indicate the angle of arrival of the incident light.
[0222] Figure 19A is a top plan view embodiment of multiple photodetectors (which may be subarrays of the pixel cells of an image sensor, in the form of a photodetector array 120) associated with a single transmission diffraction mask (TDM), according to several embodiments. Figure 19B is a cross-sectional view of the same photodetector array as in Figure 19A, along line A in Figure 19A. In the embodiments shown, the photodetector array 120 includes 16 separate photodetectors 121, which may be located within the pixel cells of an image sensor. The photodetector array 120 includes a TDM 123 positioned above the photodetectors. It should be understood that each group of pixel cells is illustrated with 4 pixels for clarity and simplicity (e.g., forming a 4x4 pixel grid). Some embodiments may include more than 4 pixel cells. For example, 16 pixel cells, 64 pixel cells, or any other number of pixels may be included within each group.
[0223] The TDM123 is located at a distance x from the photodetector 121. In some embodiments, the TDM123 is formed on the upper surface of the dielectric layer 125, as shown in Figure 19B. For example, the TDM123 may be formed from ridges or by valleys etched into the surface of the dielectric layer 125, as shown. In other embodiments, the TDM123 may be formed within the dielectric layer. For example, a portion of the dielectric layer may be modified to have a higher or lower refractive index than other portions of the dielectric layer, resulting in a holographic phase grating. Light incident on the photodetector array 120 from above is diffracted by the TDM, resulting in an arrival angle of the incident light such that it is converted into a position in the image plane at a distance x from the TDM123, where the photodetector 121 is located. The intensity of the incident light measured at each photodetector 121 of the photodetector array may be used to determine the arrival angle of the incident light.
[0224] Figure 20A illustrates an embodiment of multiple photodetectors (in the form of a photodetector array 130) associated with multiple TDMs, according to several embodiments. Figure 20B is a cross-sectional view of the same photodetector array as in Figure 20A, through line B in Figure 20A. Figure 20C is a cross-sectional view of the same photodetector array as in Figure 20A, through line C in Figure 20A. In the embodiments shown, the photodetector array 130 includes 16 distinct photodetectors, which may be located within the pixel cells of an image sensor. As shown, there are four groups of four pixel cells 131a, 131b, 131c, and 131d. The photodetector array 130 includes four distinct TDMs 133a, 133b, 133c, and 133d, each TDM provided above the associated group of pixel cells. It should be understood that each group of pixel cells is illustrated together with four pixel cells for clarity and simplicity. Some embodiments may include more than four pixel cells. For example, each group may contain 16 pixel cells, 64 pixel cells, or any other number of pixel cells.
[0225] Each TDM133a-d is located at a distance x from the photodetectors 131a-d. In some embodiments, the TDM133a-d are formed on the upper surface of the dielectric layer 135, as shown in Figure 20B. For example, the TDM123a-d may be formed from ridges or by valleys etched into the surface of the dielectric layer 135, as shown. In other embodiments, the TDM133a-d may be formed within the dielectric layer. For example, a portion of the dielectric layer may be modified to have a higher or lower refractive index than other portions of the dielectric layer, resulting in a holographic phase grating. Light incident on the photodetector array 130 from above is diffracted by the TDM, resulting in an arrival angle of the incident light such that it is converted into a position in the image plane at a distance x from the TDM133a-d where the photodetectors 131a-d are located. The intensity of the incident light measured at each photodetector 131a-d in the array of photodetectors may be used to determine the arrival angle of the incident light.
[0226] TDM133a-d may be oriented in different directions from each other. For example, TDM133a is perpendicular to TDM133b. Therefore, the intensity of light detected using photodetector group 131a may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM133a, and the intensity of light detected using photodetector group 131b may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM133b. Similarly, the intensity of light detected using photodetector group 131c may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM133c, and the intensity of light detected using photodetector group 131d may be used to determine the angle of arrival of incident light in a plane perpendicular to TDM133d.
[0227] Pixel cells configured to passively acquire depth information may be integrated into an image array with features such as those described herein to support useful operation in a cross-reality system. According to some embodiments, the pixel cells configured to acquire depth information may be implemented as part of an image sensor used to implement a camera with a global shutter. Such a configuration may, for example, provide a full frame output. A full frame may simultaneously contain image information about different pixels indicating depth and intensity. Using an image sensor with this configuration, a processor can acquire depth information about the entire scene at once.
[0228] In other embodiments, the pixel cells of the image sensor providing depth information may be configured to operate according to the DVS technique described above. In such scenarios, events may indicate a change in the depth of an object, as indicated by the pixel cells. Events output by the image array may indicate the pixel cell in which the change in depth was detected. Alternatively, or in addition, events may include a value of depth information for that pixel cell. Using the image sensor in this configuration, the processor can obtain depth information updates at a very high rate to provide high transient resolution. In some embodiments, high transient resolution may involve steps to update depth information more frequently than 1 Hz or 5 Hz, or alternatively, steps to update depth information hundreds or thousands of times per second, such as every millisecond.
[0229] In yet another embodiment, the image sensor may be configured to operate in either full-frame or DVS mode. In such embodiments, the processor, which processes the image information from the image sensor, may programmatically control the operating mode of the image sensor based on the function being performed by the processor. For example, while performing a function involving object tracking, the processor may configure the image sensor to output image information as DVS events. On the other hand, while processing to update world reconstruction, the processor may configure the image sensor to output full-frame depth information. In some embodiments, in full-frame and / or DVS mode, active illumination may be used to illuminate the imaged scene. For example, IR illumination may be provided in instances where the intensity of light detected by the sensor is below a threshold, which may be too dark to accurately perform passive depth determination.
[0230] Wearable configuration
[0231] Multiple image sensors may be used within the XR system. The image sensors may be combined with optical components such as lenses and control circuits to create a camera. These image sensors may obtain imaging information using one or more of the techniques described above, such as grayscale imaging, color imaging, global shutter, DVS techniques, plenooptic pixel cells, and / or dynamic patching. Regardless of the imaging technique used, the resulting camera may be mounted on a support member to form a headset, which may include or be connected to a processor.
[0232] Figure 21 is a schematic diagram illustrating a headset 2100 of a wearable display system consistent with the disclosed embodiments. As shown in Figure 21, the headset 2100 may include a display device comprising monoculars 2110a and 2110b, which may be optical eyepieces or displays configured to transmit and / or display visual information to the user's eyes. The headset 2100 may also include a frame 2101, which may be similar to the frame 64 described above with respect to Figure 3B. The headset 2100 may further include two cameras (DVS cameras 2120 and 2140) and additional components such as emitters 2130a, 2130b, an inertial measurement unit 2170a (IMU 2170a), and an inertial measurement unit 2170b (IMU 2170b).
[0233] Cameras 2120 and 2140 are world cameras, which are oriented to image the physical world as seen by the user wearing the headset 2100. In some embodiments, these two cameras may be sufficient to obtain image information about the physical world, and these two cameras may be world-facing cameras only. The headset 2100 may also include additional components such as eye-tracking cameras, as discussed above with respect to Figure 3B.
[0234] Monoculars 2110a and 2110b may be mechanically coupled to a support member such as frame 2101 using techniques such as adhesive, fasteners, or pressure fitting. Similarly, two cameras and ancillary components (e.g., emitters, inertial measurement units, eye-tracking cameras, etc.) may be mechanically coupled to frame 2101 using techniques such as adhesive, fasteners, or pressure fitting. These mechanical couplings may be direct or indirect. For example, one or more of the cameras and / or ancillary components may be directly attached to frame 2101. In an additional embodiment, one or more of the cameras and / or ancillary components may be directly attached to a monocular, which may then be attached to frame 2101. The mounting mechanism is not intended to be limiting.
[0235] Alternatively, a monocular subassembly may be formed and then attached to the frame 2101. Each subassembly may include a support member to which, for example, a monocular 2110a or 2110b is attached. The IMU and one or more cameras may similarly be attached to the support member. Attaching both the camera and the IMU to the same support member may allow inertial information about the camera to be acquired based on the IMU's output. Similarly, attaching the monocular to the same support member as the camera may allow image information about the world to be spatially correlated with information rendered on the monocular.
[0236] The headset 2100 may be lightweight. For example, the headset 2100 may have a weight of 30 to 300 grams. The headset 2100 may be made of a material that flexes during use, such as plastic or thin metal. Such materials can enable a lightweight and comfortable headset that can be worn by the user for extended periods of time. An XR system with such a lightweight headset can still support high-accuracy stereoscopic image analysis (which requires the separation between cameras to be known) using a calibration routine that can compensate for any inaccuracies that would occur from the flexing of the headset during use as it is repeatedly worn. In some embodiments, the lightweight headset may include a battery pack. The battery pack may include one or more batteries, which may be rechargeable or non-rechargeable. The battery pack may be constructed within the lightweight frame or may be removable. The battery pack and the lightweight frame may be formed as a single unit or the battery pack may be formed as a separate unit from the lightweight frame.
[0237] The DVS camera 2120 may include an image sensor and a lens. The image sensor may be configured to produce grayscale images. The image sensor may be configured to output an image with a size of 1 megapixel to 4 megapixels. For example, the image sensor may be configured to output an image with a horizontal resolution of 1,016 lines × a vertical resolution of 1,016 lines. In some aspects, the image sensor can be a CMOS image sensor.
[0238] The DVS camera 2120 may support dynamic vision perception as disclosed above with respect to FIGS. 4, 5A, and 5B. In one operating mode, the DVS camera 2120 may be configured to output image information in response to detected changes in image properties such as pixel-level changes in light intensity. In some embodiments, the detected changes may satisfy an intensity change criterion. For example, the image sensor may be configured to have one threshold for an incremental increase in light intensity and another threshold for an incremental decrease in light intensity. The image information can be provided asynchronously as such changes are detected. In another operating mode, the DVS camera 120 may be configured to output image frames periodically or repeatedly without responding to detected changes in image properties. For example, the DVS camera 2120 may be configured to output images at a frequency of 30 Hz to 120 Hz, such as 60 Hz.
[0239] The DVS camera 2120 may support patch tracking as disclosed above with respect to FIGS. 6-15. For example, the DVS camera 2120 may be configured to provide image information regarding a subset of pixels (e.g., a patch of pixels within the image sensor) within the image sensor in various aspects. The DVS camera 2120 may be configured to combine dynamic vision perception and patch tracking to provide image information regarding those pixels within the patch that are subject to changes in image properties. In various aspects, the image sensor can be configured with a global shutter. As discussed above with respect to FIGS. 14 and 15, the global shutter may enable each pixel to obtain intensity measurements simultaneously.
[0240] The DVS camera 2120 may be a plenooptic camera, as disclosed above with respect to Figures 3B and 15-20C. For example, a component may be installed in the optical path to one or more pixel cells of the image sensor of the DVS camera 2120, such that these pixel cells produce an output having an intensity indicating the angle of arrival of light incident on the pixel cells. In such embodiments, the image sensor can passively obtain depth information. A suitable embodiment of a component for installation in the optical path is a TDM filter. A processor may be configured to use the angle of arrival information to calculate the distance to the imaged object. For example, the angle of arrival information can be converted into distance information indicating the distance to the object from which light is reflected. In some embodiments, the pixel cells configured to provide the angle of arrival information may be scattered with pixel cells capturing the light intensity of one or more colors. As a result, the angle of arrival information, and therefore the distance information, may be combined with other image information about the object.
[0241] The DVS camera 2120 can be configured to have a wide field of view, consistent with the disclosed embodiments. For example, the DVS camera 2120 may include an equidistant lens (e.g., a fisheye lens). The DVS camera 2120 may be angled toward camera 2140 to create an area directly in front of the user of the headset 2100, which is imaged by both cameras. For example, a vertical plane passing through the center of the field of view 2121, which is the field of view associated with the DVS camera 2120, may intersect a vertical plane passing through the midline of the headset 2100 and form an angle with it. In some embodiments, the field of view 2121 may have a horizontal field of view and a vertical field of view. The range of the horizontal field of view may be 90 to 175 degrees, while the range of the vertical field of view may be 70 to 125 degrees. In some embodiments, the DVS camera 2120 can be configured to have an angular pixel resolution of 1 to 5 minutes per pixel.
[0242] Emitters 2130a and 2130b may enable imaging and / or active depth sensing by the headset 2100 under low light conditions. Emitters 2130a and 2130b may be configured to emit light at specific wavelengths. This light can be reflected by physical objects in the physical world surrounding the user. The headset 2100 may be configured with sensors for detecting this reflected light, including image sensors as described herein. In some embodiments, these sensors may be incorporated into at least one of cameras 2120 or 2140. For example, as described above with respect to Figures 18A-18C, these cameras may be configured with detectors corresponding to emitters 2130a and / or emitters 2130b. For example, these cameras may include pixels configured to detect the light emitted by emitters 2130a and / or emitters 2130b.
[0243] Emitters 2130a and 2130b may be configured to emit IR light, consistent with the disclosed embodiments. The IR light may have wavelengths between 900 nanometers and 1 micrometer. The IR light may be a 940 nm light source, for example, with the emitted light energy concentrated at about 940 nm. Emitters emitting light of other wavelengths may also be used, either as an alternative or in addition. For systems intended for indoor use only, for example, an emitter emitting light concentrated at about 850 nm may be used. At least one of the DVS cameras 2120 or 2140 may include one or more IR filters positioned over at least a subset of pixels in the camera's image sensor. The filters may allow light at wavelengths emitted by emitters 2130a and / or emitter 2130b to pass through, while attenuating light at other wavelengths. For example, the IR filter may be a notch filter that allows IR light with wavelengths matching those of the emitters to pass through. A notch filter can substantially attenuate other IR light. In some embodiments, the notch filter may be an IR notch filter that blocks IR light while allowing light from the emitter to pass through. An IR notch filter may also allow light outside the IR band to pass through. Such a notch filter may allow the image sensor to receive both visible light and light from the emitter reflected from objects within the image sensor's field of view. Thus, a subset of pixels may act as detectors for the IR light emitted by emitters 2130a and / or 2130b.
[0244] In some embodiments, the XR system's processor may selectively enable emitters, for example, to enable imaging under low-light conditions. The processor may process image information generated by one or more image sensors and detect whether the images output by those image sensors provide adequate information about objects in the physical world when the emitters are not enabled. The processor may enable emitters in response to detecting that the images do not provide adequate image information as a result of low ambient light conditions. For example, emitters may be turned on when stereoscopic information is used to track objects, and the lack of ambient light results in an image with insufficient contrast between features of the tracked objects to accurately determine distances using stereoscopic imaging techniques.
[0245] Alternatively, or in addition, emitters 2130a and / or emitters 2130b may be configured for use in performing active depth measurements, such as by emitting light in short pulses. The wearable display system may be configured to perform time-of-flight measurements by detecting reflections of such pulses from objects in the illumination field 2131a of emitter 2130a and / or the illumination field 2131b of emitter 2130b. These time-of-flight measurements may provide additional depth information for tracking objects or updating passable world models. In other embodiments, one or more emitters may be configured to emit patterned light, and the XR system may be configured to process images of objects illuminated by that patterned light. Such processing may detect variations in the pattern, which may reveal the distance to the object.
[0246] In some embodiments, the range of the illumination field associated with the emitter may be sufficient to illuminate at least the field of view of a camera used to obtain image information about an object. For example, the emitters may collectively illuminate the central field of view 2150. In the illustrated embodiment, emitters 2130a and 2130b may be positioned collectively to illuminate illumination fields 2131a and 2131b, which extend to a range where active illumination can be provided. In this exemplary embodiment, two emitters are shown, but it should be understood that more or fewer emitters may be used to extend to a desired range.
[0247] In some embodiments, emitters such as emitters 2130a and 2130b may be turned off by default, but may be enabled when additional illumination is desirable to obtain more information than can be obtained using passive imaging. The wearable display system may be configured to enable emitters 2130a and / or emitters 2130b when additional depth information is required. For example, if the wearable display system detects that it cannot obtain sufficient depth information to track hand or head posture using stereoscopic image information, the wearable display system may be configured to enable emitters 2130a and / or emitters 2130b. The wearable display system may be configured to disable emitters 2130a and / or emitters 2130b when additional depth information is not required, thereby reducing power consumption and improving battery life.
[0248] Furthermore, even when the headset is configured with an image sensor configured to detect IR light, it is not a requirement that the IR emitter be mounted on or only on the headset 2100. In some embodiments, the IR emitter may be an external device located within the space where the headset 2100 may be used, such as a room. Such an emitter may project IR light at 940 nm, invisible to the human eye, in the form of an ArUco pattern, etc. Light with such a pattern does not require the headset 2100 to supply power to provide the IR pattern, but can still facilitate "instrumentation / auxiliary tracking," as the pattern is presented, providing IR image information so that processing performed on the image information can determine the distance or location to an object in the space. Systems with an external light source can also allow more devices to operate in that space. When multiple headsets, each moving around in space without a fixed positional relationship, are operating in the same space, there is a risk that light emitted by one headset will be projected onto the image sensor of another headset, and thus interfere with its operation. The risk of such interference between headsets may limit the number of headsets that can operate in space to, for example, three or four. By using one or more IR emitters in space to illuminate objects that can be imaged by image sensors on the headsets, in some embodiments, more than 10 headsets can operate in the same space without interference.
[0249] As disclosed above with respect to Figure 3B, the camera 2140 may be configured to capture an image of the physical world within the field of view 2141. The camera 2140 may include an image sensor and a lens. The image sensor may be configured to produce a color image. The image sensor may be configured to output an image of size from 4 megapixels to 16 megapixels. For example, the image sensor may output a 12-megapixel image. The image sensor may be configured to output images repeatedly or periodically. For example, when enabled, the image sensor may be configured to output images at a frequency of 30Hz to 120Hz, such as 60Hz. In some embodiments, the image sensor may be configured to selectively output images based on the task being performed. The image sensor may be configured with a roll shutter. As discussed above with respect to Figures 14 and 15, the roll shutter may iteratively read a subset of pixels in the image sensor so that pixels in different subsets reflect light intensity data collected at different times. For example, an image sensor may be configured to read a first row of pixels in the image sensor at a first time and a second row of pixels in the image sensor at a later time.
[0250] In some embodiments, camera 2140 can be configured as a plenooptic camera. For example, as discussed above with respect to Figures 3B and 15-20C, a component may be installed in the optical path to one or more pixel cells of the image sensor so that these pixel cells produce an output having an intensity indicating the angle of arrival of light incident on the pixel cells. In such embodiments, the image sensor can passively acquire depth information. An example of a component suitable for installation in the optical path is a transmissive diffraction mask (TDM) filter. A processor may be configured to use this angle of arrival information to calculate the distance to the imaged object. For example, the angle of arrival information can be converted into distance information indicating the distance to the object from which light is reflected. In some embodiments, the pixel cells configured to provide the angle of arrival information may be scattered with pixel cells capturing the light intensity of one or more colors. As a result, the angle of arrival information, and therefore the distance information, may be combined with other image information about the object. In some embodiments, similar to the DVS camera 2120, camera 2140 may be configured to provide event detection and patch tracking functionality. The processor of the headset 2100 can be configured to provide instructions to the camera 2140 to limit image capture to a subset of pixels. In some embodiments, the sensor may be a CMOS sensor.
[0251] Camera 2140 can be positioned on the opposite side of the headset 2100 from the DVS camera 2120. For example, as shown in Figure 21, when camera 2140 is on the same side of the headset 2100 as monocular 2110a, DVS camera 2120 can be on the same side of the headset as monocular 2110b. Camera 2140 can be angled inward on the headset 2100. For example, a vertical plane passing through the center of the field of view 2141, which is the field of view associated with camera 2140, can intersect a vertical plane passing through the midline of the headset 2100 and form an angle with it. The field of view 2141 of camera 2140 may have a horizontal field of view and a vertical field of view. The range of the horizontal field of view may be 75 to 125 degrees, while the range of the vertical field of view may be 60 to 125 degrees.
[0252] Camera 2140 and DVS camera 2120 can be angled asymmetrically inward toward the midline of the headset 2100. The angle of camera 2140 can be 1 to 20 degrees inward toward the midline of the headset 2100. The angle of DVS camera 2120 can be 1 to 40 degrees inward toward the midline of the headset 2100 and can differ from the angle of camera 2140. The angular range of the field of view 2121 may exceed the angular range of the field of view 2141.
[0253] The DVS cameras 2120 and 2140 may be configured to provide an overlapping view of the central field of view 2150. The angular range of the central field of view 2150 may be 40 to 120 degrees. For example, the angular range of the central field of view 2150 may be about 70 degrees (e.g., 70 ± 7 degrees). The central field of view 2150 may be asymmetrical. For example, the central field of view 2150 may extend further toward the side of the headset 2100, including camera 2140, as shown in Figure 21. In addition to the central field of view 2150, the DVS cameras 2120 and 2140 may be positioned to provide at least two peripheral fields of view. The peripheral field of view 2160a may be associated with the DVS camera 2120 and may include that portion of the field of view 2121 that does not overlap with the field of view 2141. In some embodiments, the horizontal angular range of the peripheral field of view 2160a may be in the range of 20 to 80 degrees. For example, the angular range of the peripheral field of view 2160a may be approximately 40 degrees (e.g., 40 ± 4 degrees). The peripheral field of view 2160b (not depicted in Figure 21) may be associated with the camera 2140 and may include that portion of the field of view 2141 that does not overlap with the field of view 2121. In some embodiments, the horizontal angular range of the peripheral field of view 2160b may be in the range of 10 to 40 degrees. For example, the angular range of the peripheral field of view 2160a may be approximately 20 degrees (e.g., 20 ± 2 degrees). The location of the peripheral field of view may vary. For example, with respect to a certain configuration of the headset 2100, the peripheral field of view 2160b may not extend within 0.25 meters of the headset 2100 because, within its distance, the field of view 2141 may be entirely within the field of view 2121. In contrast, the peripheral field of view 2160a may extend within 0.25 meters of the headset 2100. In such a configuration, the wider field of view and larger inward angle of the DVS camera 2120 can ensure that the field of view 2121 falls outside the field of view 2141, at least partially, even within 0.25 meters of the headset 2100, as shown in Figure 21.
[0254] IMU2170a and / or IMU2170b may be configured to provide acceleration and / or velocity and / or tilt information to the wearable display system. For example, as a user wearing the headset 2100 moves, IMU2170a and / or IMU2170b may provide information describing the acceleration and / or velocity of the user's head.
[0255] The XR system may be coupled to a processor, which may be configured to process image data output using a camera and / or render virtual objects on a display device. The processor may be mechanically coupled to frame 2101. Alternatively, the processor may be mechanically coupled to a display device, such as a display device including a monocular 2110a or monocular 2110b. As a further alternative, the processor may be operably coupled to the headset 2100 and / or the display device via a communication link. For example, the XR system may include a local data processing module. This local data processing module may include a processor and may be connected to the headset 2100 or the display device via a physical connection (e.g., wire or cable) or a wireless connection (e.g., Bluetooth®, Wi-Fi, Zigbee®, or equivalent).
[0256] The processor may be configured to perform world reconstruction, head pose tracking, and object tracking operations. The processor may be configured to create a passable world model using the DVS cameras 2120 and 2140. When creating the passable world model, the processor may be configured to stereoscopically determine depth information using multiple images of the same physical object obtained by the DVS cameras 2120 and 2140. The processor may be configured to update an existing passable world model using the DVS camera 2120 but without using the camera 2140. As mentioned above, the DVS camera 2120 may be a grayscale camera with a relatively lower resolution than the color camera 2140. Furthermore, the DVS camera 2120 may output image information asynchronously (e.g., in response to detected events), allowing the processor to asynchronously update the passable world model, head pose, and / or object locations only when changes are detected. In some embodiments, the processor may instruct the DVS camera 2120 to limit image data acquisition to one or more patches of pixels in the image sensor of the DVS camera 2120. The DVS camera 2120 may then limit image data acquisition to these patches of pixels. As a result, updating the passable world model using image information output by the DVS camera 2120, rather than camera 2140, can be performed quickly with reduced power consumption and improved battery life. In some embodiments, the processor may be configured to update the passable world model using image information output by the DVS camera 2120, either from time to time or periodically.
[0257] The processor may preferentially update the passable world model using the DVS camera 2120, but in some embodiments, the processor may update the passable world model using both the DVS camera 2120 and camera 2140, either from time to time or periodically. For example, the processor may be configured to determine that the passable world quality criteria are no longer met, that a predetermined time interval has elapsed since the last acquisition and / or use of an image output by camera 2140, and / or that a change has now occurred to an object in part of the physical world within the field of view of both the DVS camera 2120 and camera 2140.
[0258] After creating a passable world model, the processor may preferentially track the positions of heads, objects, and / or hands using image information output by the DVS camera 2120. As described above, the DVS camera 2120 may be configurable for event-based image acquisition. The acquired image information may be specific to one or more patches within the image sensor of the DVS camera 2120. Alternatively, the processor may track the positions of heads, objects, and / or hands using monocular images output by the DVS camera 2120 or camera 2140. These monocular images may be full-frame images and may be output repeatedly or periodically by the DVS camera 2120. For example, the DVS camera may be configured to output images at frequencies between 30Hz and 120Hz, such as 60Hz, while changes may be output at much higher effective rates, such as hundreds or thousands of times per second. In some embodiments, the processor may track the positions of heads, objects, and / or hands using stereoscopic images output by both the DVS camera 2120 and camera 2140. These stereoscopic projection images may be output repeatedly or periodically.
[0259] The processor may be configured to obtain light field information, such as arrival angle information regarding light incident on the image sensor, using a plenoptic camera. In some embodiments, the plenoptic camera can be at least one of the DVS camera 2120 or the camera 2140. Consistent with the disclosed embodiments, when depth information is described herein or can improve processing, such depth information may be determined from or complemented by the light field information obtained by the plenoptic camera. For example, when the camera 2140 includes a TDM filter, the processor may be configured to use the images obtained from the DVS cameras 2120 and 2140, together with the light field information obtained from the camera 2140, to create a passable world model. Alternatively, or in addition, when the DVS camera 2120 includes a TDM filter, the processor may use the light field information obtained from the DVS camera 2120.
[0260] The processor may be configured to detect conditions in which one or more types of image information are unavailable or insufficient to provide resolution for a particular function, such as world reconstruction, object tracking, or head pose determination. The processor may be configured to select additional or alternative sources of image information to provide sufficient image information for that function. These sources may be selected in order to yield suitable image information available with a low processing burden in each situation. For example, the DVS camera 2120 may be configured to output image data that satisfies an intensity change criterion, but the pixel intensity may not change sufficiently to trigger image data acquisition when objects with uniform visual characteristics fill the field of view 2121. Such an example may be, for example, when the user's hand moves, but the user's hand fills the field of view. In some embodiments, even partial filling of the field of view 2121 may hinder or interfere with the acquisition of suitable image data. For example, when patches tracked by the processor are filled with images of objects with a uniform appearance, the image data for those patches may be unsuitable for tracking objects. Similarly, if a nearby object fills a patch used to track a point of focus, such as a point of focus tracked to track the camera's field of view or at least head posture, the image information from that camera may be unsuitable for head tracking functionality.
[0261] However, given the geometric shape of the headset 2100 shown in Figure 21, the camera 2140 may be positioned to output images of objects filling the relevant portion of the field of view. These images may, alternatively or in addition, be used to determine the depth of the objects. The camera 2140 may also be positioned to image a point of interest tracked within a patch for head pose tracking.
[0262] Therefore, the processor can be configured to determine whether an object satisfies the fill criteria with respect to the field of view 2121. Based on this determination, the processor can be configured to enable camera 2140 or increase the frame rate of camera 2140. The processor can then receive image data from camera 2140, at least. In some embodiments, this image data may include light field information. For object tracking, the processor may be configured to use the received image data to determine depth information. The processor can use the determined depth information to track occluded objects. For head pose tracking, the processor may be configured to use its received image data to determine changes in head pose. When DVS camera 2120 is configured for patch tracking of points of interest, the processor can be configured to use camera 2120 to track points of interest. The processor may be configured to disable camera 2140 or reduce its frame rate based on determining that the fill criteria are no longer met, and to return to using the data, which may be more readily available and / or processed with shorter latency but still suitable for the function being performed, such as object tracking or head pose tracking. This may enable DVS camera 2120 to track images at ultra-high transient resolution. For example, DVS camera 2120 may output information indicating the current position of a moving object at frequencies higher than 60 Hz, such as hundreds or thousands of times per second.
[0263] It should be understood that the processor may enable or disable cameras to dynamically provide image information from different sources for one or more functions based on operating conditions detected in any one or more ways. The processor may send control signals to the underlying image sensor hardware to modify the operation of the hardware. Alternatively, or in addition, the processor may enable a camera by reading the image information it generates, or disable a camera by not accessing or not using the image information it generates. These techniques may be used to enable or suppress image information, in whole or in part. For example, the processor may be configured to perform a size reduction routine and adjust the image output using camera 2140. Camera 2140 may produce a larger image than DVS camera 2120. For example, camera 2140 may produce a 12-megapixel image, while DVS camera 2120 may produce a 1-megapixel image. The image produced by camera 2140 may contain more information than is necessary to perform passable world creation, head tracking, or object tracking operations. Processing this additional information may require additional power, reduce battery life, or increase latency. Therefore, the processor can be configured to discard or combine pixels in the image output by the camera 2140. For example, the processor can be configured to output an image with 1 / 16th the number of pixels of the original generated image. Each pixel in this output image may have a value based on the corresponding 4x4 set of pixels in the original generated image (e.g., the average of the values of these 16 pixels).
[0264] The XR system may include a hardware accelerator according to some embodiments. The hardware accelerator may be implemented as an application-specific integrated circuit (ASIC) or other semiconductor device and may be integrated into the headset 2100, or otherwise coupled to it, to receive image information from cameras 2120 and 2140. The hardware accelerator can use the images output by these two world cameras to assist in the stereoscopic determination of depth information. The image from camera 2120 may be a grayscale image, and the image from camera 2140 may be a color image. Using hardware acceleration can speed up the determination of depth information, reduce power consumption, and therefore increase battery life.
[0265] Exemplary Calibration Process
[0266] Figure 22 illustrates a simplified flowchart of a calibration routine (Method 2200) according to several embodiments. The processor may be configured to perform the calibration routine while the wearable display system is being worn. The calibration routine may address distortion resulting from the lightweight construction of the headset 2100. For example, the processor may repeatedly perform the calibration routine so that it compensates for distortion in the frame 2101 while the wearable display system is in use. The compensation routine may be performed automatically or in response to manual input (e.g., a user request to perform the calibration routine). The calibration routine may include a step of determining the relative positions and orientations of cameras 2120 and 2140. The processor may be configured to perform the calibration routine using images output by the DVS cameras 2120 and 2140. In some embodiments, the processor may further be configured to use the outputs of IMUs 2170a and 2170b.
[0267] Starting from block 2201, method 2200 may proceed to block 2210. In block 2210, the processor may identify corresponding features in the images output from DVS cameras 2120 and 2140. The corresponding features may be parts of objects in the physical world. In some embodiments, the objects may be placed within the central field of view 2150 by the user for calibration purposes and may have features that are easily identifiable in the image and may have a predetermined relative position. However, the calibration techniques described herein may be implemented in such a way that the calibration can be repeated during the use of the headset 2100, based on features on objects present in the central field of view 2150 at the time of calibration. In various embodiments, the processor may be configured to automatically select features detected in both the field of view 2121 and the field of view 2141. In some embodiments, the processor may be configured to determine correspondences between features using estimated locations of features in the field of view 2121 and the field of view 2141. Such estimations may be based on a passable world model or other information about features constructed for the objects containing these features.
[0268] Method 2200 may proceed to block 2230. In block 2230, the processor may receive inertial measurement data. The inertial measurement data may be received from IMU 2170a and / or IMU 2170b. The inertial measurement data may include tilt and / or acceleration and / or velocity measurements. In some embodiments, IMU 2170a and 2170b may be mechanically coupled directly or indirectly to camera 2140 and DVS camera 2120, respectively. In such embodiments, differences in inertial measurements such as tilt performed by IMU 2170a and 2170b may indicate differences in the position and / or orientation of camera 2140 and camera 2120. Thus, the outputs of IMU 2170a and 2170b may provide a basis for making an initial estimation of the relative positions of camera 2140 and DVS camera 2120.
[0269] After block 2230, method 2200 may proceed to block 2250. In block 2250, the processor may calculate initial estimates of the relative position and orientation of the DVS cameras 2120 and 2140. These initial estimates may be calculated using measurements received from IMU 2170b and / or IMU 2170a. In some embodiments, for example, the headset may be designed using the nominal relative position and orientation of the DVS cameras 2120 and 2140. The processor may be configured to consider any difference in the received measurements between IMU 2170a and IMU 2170b as being due to distortion in frame 2101 that can alter the position and / or orientation of the DVS cameras 2120 and 2140. For example, IMU 2170a and IMU 2170b may be mechanically coupled to frame 2101, directly or indirectly, such that their tilt and / or acceleration and / or velocity measurements have a predetermined relationship. This relationship may be affected if frame 2101 is distorted. In a non-limiting embodiment, IMUs 2170a and 2170b may be mechanically coupled to frame 2101 so that when there is no distortion of frame 2101, these sensors measure similar tilt, acceleration, or velocity vectors during headset movement. In this non-limiting embodiment, twisting or bending that rotates IMU 2170a relative to IMU 2170b may result in a corresponding rotation of the tilt, acceleration, or velocity vector measurement with respect to IMU 2170a relative to the corresponding vector measurement with respect to IMU 2170b. The processor may therefore adjust the nominal relative position and orientation for DVS cameras 2120 and 2140 to match the measured relationship between IMU 2170a and IMU 2170b, since IMUs 2170a and 2170b are mechanically coupled to DVS cameras 2140 and 2120, respectively.
[0270] Other techniques may also be used, either as an alternative or in addition, to make initial estimates. In embodiments where calibration method 2200 is performed repeatedly during the operation of the XR system, the initial estimate may be, for example, the most recently calculated estimate.
[0271] Following block 2250, a subprocess is initiated to make further estimates of the relative positions and orientations of cameras 2120 and 2140. One of these estimates is selected as the relative positions and orientations of cameras 2120 and 2140 in order to calculate stereoscopic depth information from the images output by cameras 2120 and 2140. The subprocess may be performed iteratively, with further estimations made in each iteration until an acceptable estimate is identified. In the embodiment of Figure 22, the subprocess includes blocks 2270, 2272, 2274, and 2290.
[0272] In block 2270, the processor may calculate an error relating to the estimated relative orientation of the cameras and the features being compared. In calculating this error, the processor may be configured to estimate how the identified features should appear or where they should be located in the corresponding images output using the DVS cameras 2120 and 2140, based on the estimated relative orientation of the estimated location features used for calibration. In some embodiments, this estimate may be compared to the appearance or apparent location of the corresponding features in the images output using each of the two cameras to generate an error for each estimated relative orientation. Such errors may be calculated using linear algebra techniques. For example, the mean squared deviation between the estimated location and the actual location of each of several features in the image may be used as a metric for the error.
[0273] After block 2270, method 2200 may proceed to block 2272, where a check may be performed to determine whether the error meets the approval criteria. The criteria may be, for example, the overall magnitude of the error or the change in the error between iterations. If the error meets the approval criteria, method 2200 proceeds to block 2290.
[0274] In block 2290, the processor may select one of the estimated relative orientations based on the error calculated in block 2272. The selected estimated relative position and orientation may be the estimated relative position and orientation having the lowest error. In some embodiments, the processor may be configured to select the estimated relative position and orientation associated with this lowest error as the current relative position and orientation of the DVS cameras 2120 and 2140. After block 2290, method 2200 may proceed to block 2299. Method 2200 may terminate in block 2299, and the selected positions and orientations of the DVS cameras 2120 and 2140 may be used to calculate stereoscopic image information based on the images formed using those cameras.
[0275] If the error in block 2272 does not meet the approval criteria, method 2200 may proceed to block 2274. In block 2274, the estimates used in calculating the error in block 2270 may be updated. These updates may be for the estimated relative positions and / or orientations of cameras 2120 and 2140. In embodiments where the relative positions of a set of features used for calibration are estimated, the updated estimates selected in block 2274 may, alternatively or in addition, include updates to the positions of the locations of features within the set. Such updates may be performed according to linear algebraic techniques used to solve a set of equations with multiple variables. In specific embodiments, one or more of the estimated positions or orientations may be increased or decreased. If the change reduces the calculated error in one iteration of the subprocess, the same estimated positions or orientations may be further changed in the same direction in subsequent iterations. Conversely, if the change increases the error, those estimated positions or orientations may be changed in the opposite direction in subsequent iterations. The estimated positions and orientations of the cameras and features used in the calibration process may vary sequentially or in combination.
[0276] Once the updated estimate is calculated, the subprocess returns to block 2270. There, further iterations of the subprocess begin, along with the calculation of errors for the estimated relative position. In this way, the estimated position and orientation are updated until an updated relative position and orientation that provides an acceptable error is selected. However, it should be understood that the processing in block 2272 may apply other criteria to terminate the iterative subprocess, such as completing a certain number of iterations without finding an acceptable error.
[0277] Method 2200 is described in relation to DVS cameras 2120 and 2140, but similar calibrations may be performed for any pair of cameras used for stereoscopic imaging, or for any set of multiple cameras where relative position and orientation are desired.
[0278] Exemplary camera configuration
[0279] The headset 2100 incorporates components that provide fields of view and illumination fields to support multiple functions of the XR system. Figures 23A–23C are illustrative schematic diagrams of fields of view or illumination associated with the headset 2100 of Figure 21, according to several embodiments. Each illustrative schematic diagram depicts the field of view or illumination from a different orientation and distance from the headset. Figure 23A depicts the field of view or illumination at a distance of 1 meter from the headset, from an elevated off-axis line of sight. Figure 23A depicts the overlap between fields of view for DVS cameras 2120 and 2140, in particular how the DVS cameras 2120 and 2140 are angled so that fields of view 2121 and 2141 intersect the midline of the headset 2100. In the depicted configuration, field of view 2141 extends beyond field of view 2121, forming peripheral field of view 2160b. As depicted, the illumination fields for emitters 2130a and 2130b primarily overlap. Thus, emitters 2130a and 2130b may be configured to support imaging or depth measurement for objects within the central field of view 2150 under low ambient light conditions. Figure 23B depicts the field of view or illumination from an up-and-down line of sight, at a distance of 0.3 meters from the headset. Figure 23B depicts the overlap of fields of view 2121 and 2141 at 0.3 meters from the headset. However, in the depicted configuration, field of view 2141 does not extend very far beyond field of view 2121, limiting the range of the peripheral field of view 2160b and demonstrating an asymmetry between the peripheral field of view 2160a and 2160b. Figure 23C depicts the field of view or illumination from a front-view line of sight, at a distance of 0.25 meters from the headset. Figure 23C illustrates that the overlap between fields of view 2121 and 2141 exists at 0.25 meters from the headset. However, field of view 2141 is entirely contained within field of view 2121, and therefore, in the depicted configuration, peripheral field of view 2160b does not exist at this distance from the headset 2100.
[0280] As can be seen from Figures 23A-23C, the overlap of fields 2121 and 2141 creates a central field of view where stereoscopic imaging techniques can be employed using grayscale images output by cameras 2120 and 2140, with or without IR illumination from emitters 2130a and 2130b. In this central field of view, color information from camera 2140 may be combined with grayscale image information from camera 2120. In addition, there is a peripheral field of view where there is no overlap, but monocular grayscale image information or color image information is available from either camera 2120 or camera 2140, respectively. Different calculations may be performed on the image information obtained for the central and peripheral fields of view, as described herein.
[0281] World Model Generation
[0282] In some embodiments, image data output by the DVS camera 2120 and / or camera 2140 may be used to build or update a world model. Figure 24 is a simplified flowchart of a method 2400 for creating or updating a passable world model according to some embodiments. As disclosed above with respect to Figure 21, the XR system may be configured to use a processor to determine and update the passable world model. In some embodiments, the DVS camera 2120 may be configured to output image information representing the intensity level detected in each of a plurality of pixels, or to indicate pixels in which an intensity change exceeding a threshold has been detected. In some embodiments, the processor may determine and update the passable world model based on the output of the DVS camera 2120, representing the detected intensity, which can be used in conjunction with image information output from camera 2140 to stereoscopically determine the location of objects in the passable world. In some embodiments, the output of the DVS camera 2120, representing the detected change in intensity, may be used to identify areas of the world model to update based on the change in image information from those locations.
[0283] In some embodiments, the DVS camera 2120 may be configured to output image information that reflects intensity, along with a global shutter, while the camera 2140 may be configured with a roll shutter. The processor may therefore perform a compensation routine to compensate for roll shutter distortion in the image output by the camera 2140. In various embodiments, the processor may determine and update the passable world model without using emitters 2130a and 2130b. However, in some embodiments, the passable world model may be incomplete. For example, the processor may incompletely determine depth with respect to walls or other flat surfaces. As an additional embodiment, the passable world model may incompletely represent objects with many corners, curved surfaces, transparent surfaces, or large surfaces, such as windows, doors, balls, tables, and equivalents. The processor may be configured to identify such incomplete information, obtain additional information, and update the world model using the additional depth information. In some embodiments, emitters 2130a and 2130b may be selectively enabled to collect additional image information from which to build or update the passable world model. In some scenarios, the processor may be configured to perform object recognition in the acquired image, select a template for the recognized object, and add information to the passable world model based on the template. In this way, the wearable display system can improve the passable world model with little to no use of power-intensive components such as emitters 2130a and 2130b, thereby extending battery life.
[0284] Method 2400 may be initiated once or more during the operation of the wearable display system. The processor may be configured to create a passable world model when the user moves to a new environment, such as by turning on the system for the first time or walking into another room, or generally when the processor detects a change in the user's physical environment. Alternatively, or in addition, Method 2400 may be performed periodically during the operation of the wearable display system, or in response to user input such as an input indicating that a significant change in the physical world is detected or that the world model is not synchronized with the physical world.
[0285] In some embodiments, all or part of the passable world model may be stored, provided by other users of the XR system, or otherwise obtained. Thus, although the creation of the world model is described, it should be understood that method 2400 may be used for a portion of the world model, along with other parts of the world model derived from other sources.
[0286] In block 2405, the processor can perform a compensation routine to compensate for roll shutter distortion in the image output by camera 2140. As described above, images acquired by an image sensor with a global shutter, such as the image sensor in the DVS camera 2120, contain pixel values acquired simultaneously. In contrast, images acquired by an image sensor with a roll shutter contain pixel values acquired at different times. The relative movement of the headset and the environment between image acquisition by camera 2140 can introduce spatial distortion into the image. These spatial distortions can affect the accuracy of the method, which depends on comparing the image acquired by camera 2140 with the image acquired by DVS camera 2120.
[0287] The step of performing the compensation routine may include using a processor to compare the image output by the DVS camera 2120 with the image output by the camera 2140. The processor performs this comparison and identifies any distortion in the image output by the camera 2140. Such distortion may include distortion in at least a portion of the image. For example, if the image sensor in the camera 2140 is acquiring pixel values row by row from the top to the bottom of the image sensor while the headset 2100 is moving laterally, the appearance of an object or part of an object may be offset in consecutive rows of pixels by an amount depending on the speed of the translation and the time difference between acquiring the values for each row. Similar distortion may also occur when the headset is rotated. These distortions may result in overall distortion of the location and / or orientation of an object or part of an object in the image. The processor may be configured to perform a line-by-line comparison between the image output by the camera 2120 and the image output by the camera 2140 to determine the amount of distortion. The image output by camera 2140 can then be transformed to remove distortion (for example, to remove detected distortion).
[0288] In block 2410, a passable world model may be created. In the illustrated embodiment, the processor may use DVS cameras 2120 and 2140 to create the passable world model. As described above, when generating the passable world model, the processor may be configured to use the images output from DVS cameras 2120 and 2140 to stereoscopically determine depth information about objects in the physical world when constructing the passable world model. In some embodiments, the processor may receive color information from camera 2140. This color information may be used to distinguish objects or to identify surfaces associated with the same object. The color information may also be used to recognize objects. As disclosed above with respect to Figure 3A, the processor can create a passable world model by associating information about the physical world with information about the location and orientation of the headset 2100. In a non-limiting embodiment, the processor may be configured to determine the distance from the headset 2100 to features in a view (e.g., field of view 2121 and / or field of view 2141). The processor may be configured to estimate the current location and orientation of the view. The processor can be configured to store such distances along with location and orientation information. The location and orientation of a feature in the environment can be determined by triangulation of the distances to the feature obtained from multiple locations and orientations. In various embodiments, the passable world model can be a combination of raster images, point and descriptor sets, and polygon / geometric definitions that describe the location and orientation of such a feature in the environment. In some embodiments, the distance from the headset 2100 to a feature in the central field of view 2150 can be determined stereoscopically using image data output by the DVS camera 2120 and a compensatory image generated using image data output by camera 2140. In various embodiments, light field information can be used to complement or refine this determination.For example, the arrival angle information may be converted through calculation into distance information that indicates the distance to the object from which the light is reflected.
[0289] In block 2415, the processor may disable camera 2140 or reduce the frame rate of camera 2140. For example, the frame rate of camera 2140 may be reduced from 30Hz to 1Hz. As disclosed above, the color camera 2140 may consume more power than the DVS camera 2120, a grayscale camera. By disabling camera 2140 or reducing its frame rate, the processor can reduce power consumption and extend the battery life of the wearable display system. Therefore, the processor may conserve power by disabling camera 2140 or reducing its frame rate. A lower power state may be maintained until a condition is detected that indicates an update may be required within the world model. Such a condition may be detected over time or based on input such as from sensors that gather information about the user's surrounding environment or from the user.
[0290] Alternatively, or in addition, once the passable world model is sufficiently complete, the location and orientation of features in the physical environment can be sufficiently determined using images or image patches obtained using camera 2120. In a non-limiting embodiment, the passable world model may be identified as sufficiently complete based on the percentage of space surrounding the user's location represented in the model, or based on the amount of new image information that matches the passable world model. With respect to the latter approach, newly obtained images may be associated with locations in the passable world. The world model may be considered complete if the features in those images have features that match features identified as landmarks in the passable world model. Coverage or matching rate does not need to be 100% complete. Rather, a suitable threshold may be applied for each criterion, such as coverage of more than 95% or more than 90% of features that match previously identified landmarks. Regardless of how the passable world model is determined to be complete, once completed, the processor can use the existing passable world information to refine estimates of the location and orientation of features in the physical world. This process may reflect the assumption that features in the physical world, if any, change position and / or orientation slowly compared to the rate at which the processor processes the images output by the DVS camera 2120.
[0291] In block 2420, after creating a passable world model, the processor may identify surfaces and / or objects to update the passable world model. In some embodiments, the processor may use grayscale images or image patches output by the DVS camera 2120 to identify such surfaces or objects. For example, once a world model showing a surface at a particular location in the passable world is created in block 2410, grayscale images or image patches output by the DVS camera 2120 may be used to detect surfaces with substantially identical characteristics and determine that the passable world model should be updated by updating the location of that surface in the passable world model. A surface at substantially the same location with substantially identical shape to a surface in the passable world model may be considered, for example, comparable to that surface in the passable world model, and the passable world model may be updated as appropriate. In another embodiment, the locations of objects represented in the passable world model may be updated based on grayscale images or image patches output by the DVS camera 2120. As described herein, the DVS camera 2120 can detect events, which are associated with grayscale images or further with image patches. In response to the detection of such events, the DVS camera 2120 can be configured to update the passable world model using the grayscale images or image patches.
[0292] In some embodiments, the processor may use light field information acquired from the camera 2120 to determine depth information for objects in the physical world. For example, angle of arrival information may be used to determine depth information. This determination may be more accurate for objects in the physical world closer to the headset 2100. Therefore, in some embodiments, the processor may be configured to use light field information to update only the portion of the passable world model that satisfies the depth criterion. The depth criterion may be based on the maximum distinguishable distance. For example, the processor may be unable to distinguish objects at different distances from the headset 2100 when those distances exceed a threshold distance. The depth criterion may also be based on a maximum error threshold. For example, the error in the estimated distance may increase with increasing distance, with a specific distance corresponding to the maximum error threshold. In some embodiments, the depth criterion may be based on the minimum distance. For example, the processor may be unable to accurately determine distance information for objects within a minimum distance from the headset 2100, such as 15 cm. Therefore, the portion of the world model beyond 16 cm from the headset may satisfy the depth criterion. In some embodiments, the passable world model may consist of three-dimensional "bricks" of voxels. In such embodiments, the step of updating the passable world model may include the step of identifying voxel bricks for updating. In some embodiments, the processor may be configured to determine a viewing frustum, which may have a maximum depth such as 1.5 m. The processor may be configured to identify bricks within the viewing frustum. The processor may then update the passable world information relating to the voxels within the identified bricks. In some embodiments, the processor may be configured to update the passable world information relating to voxels using light field information obtained in step 2450, as described herein.
[0293] The process for updating the world model may differ based on whether the object is in the central or peripheral field of view. For example, the update may be performed for detected surfaces in the central field of view. In the peripheral field of view, the update may be performed only for objects for which the processor has a model such that the processor can verify that any update to the passable world model matches that object. Alternatively, or in addition, new objects or surfaces may be recognized based on processing of grayscale images. Even if such processing leads to a less accurate representation of the object or surface than the processing in block 2410, the trade-off of accuracy for faster and lower power processing may lead to a better overall system in some scenarios. Furthermore, the lower accuracy information may be periodically replaced with higher accuracy information by periodically repeating method 2400 so as to replace the portion of the world model generated using only monocular grayscale images with the portion generated stereoscopically through the use of the color camera 2140 in combination with the DVS camera 2120.
[0294] In some embodiments, the processor may be configured to determine whether the updated world model satisfies quality criteria. If the world model satisfies quality criteria, the processor may continue updating the world model with camera 2140 disabled or having a reduced frame rate. If the updated world model does not satisfy quality criteria, method 2400 may enable camera 2140 or increase the frame rate of camera 2140. Method 2400 may also return to step 2410 and recreate the passable world model.
[0295] In block 2425, after updating the passable world model, the processor may identify whether the passable world model contains incomplete depth information. Incomplete depth information can arise in any of several ways. For example, some objects do not bring detectable structures into the image. For example, very dark areas in the physical world may not be imaged with sufficient resolution to extract depth information from an image obtained using ambient illumination. In another embodiment, the surface of a window or glass table may not appear or be recognizable by computerized processing in a visible image. In yet another embodiment, a large, uniform surface such as the surface of a table or a wall may lack sufficient features to correlate in two stereoscopic images in order to enable stereoscopic image processing. As a result, it may be impossible for the processor to determine the location of such an object using stereoscopic processing. In these scenarios, if a “hole” exists in the world model, the process of attempting to determine the distance to the surface in a particular direction passing through the “hole” using the passable world model would be unable to obtain any arbitrary depth information.
[0296] If the passable world model does not contain incomplete depth information, method 2400 may return to the step of updating the passable world model using a grayscale image or image patch acquired from the DVS camera 2120.
[0297] Following the identification of incomplete depth information, the processor controlling method 2400 may take one or more actions to obtain additional depth information. Method 2400 may proceed to block 2431, block 2433, and / or block 2435. In block 2431, the processor may enable emitters 2130a and / or emitter 2130b. As disclosed above, one or more of cameras 2120 and 2140 may be configured to detect light emitted by emitters 2130a and / or emitter 2130b. The processor may then obtain depth information by causing emitters 2130a and / or 2130b to emit light that can improve the acquired image of an object in the physical world. When DVS cameras 2120 and 2140 are sensitive to emitted light, for example, the images output by DVS cameras 2120 and 2140 may be processed to extract stereoscopic information. Other analytical techniques may also be used, as an alternative or in addition, to acquire depth information when emitters 2130a and / or emitters 2130b are enabled. Time-of-flight measurement and / or structured optical techniques may be used, as an alternative or in addition, in some embodiments.
[0298] In block 2433, the processor may determine additional depth information from previously obtained depth information. In some embodiments, for example, the processor may be configured to identify objects in images formed using the DVS camera 2120 and / or camera 2140 and fill in any holes in the passable world model based on the model of the identified objects. For example, the process may detect planar surfaces in the physical world. Planar surfaces may be detected using existing depth information obtained using the DVS camera 2120 and / or camera 2140, or depth information stored in the passable world model. Planar surfaces may be detected in response to a determination that a portion of the world model contains incomplete depth information. The processor may be configured to estimate additional depth information based on the detected planar surfaces. For example, the processor may be configured to extend the identified planar surfaces through the region of incomplete depth information. In some embodiments, when extending the planar surfaces, the processor may be configured to interpolate missing depth information based on the surrounding portion of the passable world model.
[0299] In some embodiments, as an additional embodiment, the processor may be configured to detect objects within a portion of the world model, including incomplete depth information. In some embodiments, this detection may involve a step of recognizing the object using a neural network or other machine learning tool. In some embodiments, the processor may be configured to access a database of stored templates and select an object template corresponding to the identified object. For example, if the identified object is a window, the processor may be configured to access a database of stored templates and select a corresponding window template. In a non-limiting embodiment, the template may be a three-dimensional model representing a class of objects, such as a window, door, ball, or equivalent. The processor may construct an instance of the object template based on an image of the object in the updated world model. For example, the processor may scale, rotate, and translate the template to match the detected location of the object in the updated world model. The additional depth information may then be estimated based on the boundaries of the constructed template, which represent the surface of the recognized object.
[0300] In block 2435, the processor can obtain light field information. In some embodiments, this light field information can be obtained along with the image and may include angle of arrival information. In some embodiments, camera 2120 is configured as a plenooptic camera and can obtain this light field information.
[0301] After blocks 2431, 2433, and / or 2435, method 2400 may proceed to block 2440. In block 2490, the processor may update the passable world model using additional depth information obtained in blocks 2431 and / or 2473. For example, the processor may be configured to blend additional depth information obtained from measurements made using active IR illumination into the existing passable world model. Similarly, additional depth information determined from light field information using triangulation, for example, based on arrival angle information, can be blended into the existing passable world model. As an additional embodiment, the processor may be configured to blend interpolated depth information obtained by extending detected planar surfaces into the existing passable world model, or to blend additional depth information estimated from the boundaries of a configured template into the existing passable world model.
[0302] The information may be hybridized in one or more ways, depending on the nature of the additional depth information and / or the information in the passable world model. Hybridization may be performed, for example, by adding additional depth information collected with respect to locations where holes exist in the passable world model to the passable world model. Alternatively, the additional depth information may override the information at corresponding locations in the passable world model. Yet another alternative is that hybridization may involve a step of selection between information already present in the passable world model and the additional depth information. Such selection may be based, for example, on a step of selecting depth information that is either already present in the passable world model or is in the additional depth information that represents the surface closest to the camera used to collect the additional depth information.
[0303] In some embodiments, the passable world model may be represented by a mesh of connected points. Updating the world model may be done by calculating a mesh representation of an object or surface to be added to the world model, and then combining that mesh representation with the mesh representation of the world model. The inventors have recognized the value of performing the process in this order, as it may require less processing than adding the object or surface to the world model and then calculating a mesh for the updated model.
[0304] Figure 24 shows that the world model can be updated in both blocks 2420 and 2440. The processing in each block may be carried out in the same or different ways, for example, by generating mesh representations of objects or surfaces to be added to the world model and combining the generated meshes with the meshes of the world model. In some embodiments, this merging operation may be carried out once with respect to both the objects or surfaces identified in block 2420 and block 2440. Such combining processing may be carried out, for example, as described in relation to block 2440.
[0305] In some embodiments, method 2400 may return to block 2420 to repeat the process of updating the world model based on information obtained using the DVS camera 2120. The processing in block 2420 may be repeated at a higher rate because it can be performed on fewer and smaller images than the processing in block 2410. This processing may be performed at a rate of less than 10 times per second, such as 3 to 7 times per second.
[0306] Method 2400 may be repeated in this manner until a termination condition is detected. For example, Method 2400 may be repeated over a predetermined time period until user input is received or until a change of a particular type or size in a portion of the physical world model within the field of view of the headset 2100's camera is detected. Method 2400 may then terminate in block 2499. Method 2400 may be restarted so that new information about the world model, including that obtained using a higher-resolution color camera, is captured in block 2405. Method 2400 may be terminated and restarted so that the processing in block 2405 is repeated using the color camera at an average rate slower than the rate at which the world model is updated based solely on grayscale image information, thereby creating a portion of the world model. The processing using the color camera may be repeated, for example, once per second or at an average rate slower.
[0307] Head posture tracking
[0308] The XR system may track the position and orientation of the user's head while wearing the XR display system. Determining the user's head pose allows the information in the passable world model to be translated into a reference frame on the user's wearable display device so that the information in the passable world model can be used when rendering objects on the wearable display device. As the head pose is frequently updated, performing head pose tracking using only the DVS camera 2120 may offer power savings, reduced computation, or other advantages. Such tracking may be event-based, as described above with respect to Figure 4-16, and complete images and / or image patch data may be obtained. The XR system may therefore be configured to disable the color camera 2140 or reduce its frame rate as needed to balance head tracking accuracy with power consumption and computation requirements.
[0309] Figure 25 is a simplified flowchart of Method 2500 for head pose tracking in several embodiments. Method 2500 may include the steps of creating a world model, selecting a tracking methodology, tracking the head pose using the selected methodology, and evaluating the tracking quality. According to Method 2500, the processor may preferentially track the head pose using event-based acquisition of image patch data by the DVS camera 2120. If this preferred approach proves unsuitable, the processor may track the head pose using full-frame images periodically output by the DVS camera 2120. If this secondary approach proves unsuitable, the processor may stereoscopically track the head pose using images output by the DVS camera 2120 and camera 2140.
[0310] In block 2510, the processor may create a passable world model. In some embodiments, the processor may be configured to create a passable world model as described above with respect to blocks 2405-2415 of method 2400. For example, the processor may be configured to obtain images from DVS camera 2120 and camera 2140. In some implementations, the processor may compensate for roll shutter distortion in camera 2140. The processor may then use the images from DVS camera 2120 and the compensated images from camera 2140 to stereoscopically determine the depth of features in the physical world. Using these depths, the processor may create a passable world model. After creating the passable world model, in some embodiments, the processor may be configured to disable camera 2140 or reduce its frame rate. By disabling camera 2140 or reducing its frame rate after generating the passable world model, the XR system may reduce power consumption and computing requirements.
[0311] A processor may select features from a world model that correspond to stationary features, for example, as described above in relation to Figure 13. Image information indicating the location of stationary features relative to a camera mounted on a device worn on the user's head may be used to calculate changes in the position of the user's head relative to the world model. According to some embodiments, the processor may select a tracking methodology for tracking the relative positions of stationary features that both satisfy quality criteria and require low computational load compared to other tracking methods.
[0312] In block 2520, the processor may select a tracking methodology. The processor may preferentially select head pose tracking using events detected by the DVS camera 2120. The processor may continue to use this preferred approach as long as head pose tracking satisfies tracking quality criteria. For example, the difficulty of tracking head pose may depend on the location and orientation of the user's head and the content of the passable world model. Therefore, in some instances, it may be impossible or impossible for the processor to track head pose using only asynchronously acquired image data output by the DVS camera 2120.
[0313] If the processor determines that head pose tracking provided by the preferred approach does not satisfy the tracking quality criteria, the processor may select a secondary approach to track the head pose. For example, the processor may select to track the head pose using indications of events output by the DVS camera 2120 in combination with color information acquired using camera 2140. If the processor determines that this secondary approach does not satisfy the tracking quality criteria, the processor may select a tertiary approach to track the head pose. For example, the processor may select to track the head pose stereoscopically using images output by the DVS camera 2120 and camera 2140. The processor may continue using the selected approach as long as the head pose tracking satisfies the tracking quality criteria. Alternatively, the processor may revert to a more preferred approach after a predetermined duration, time, or number of head pose updates, or in response to the criteria being met.
[0314] In each approach, image information may be obtained for the entire field of view for each camera used. However, as described above in relation to Figure 13, image information may be collected only for patches of the image that correspond to the portion containing the features being tracked.
[0315] Figure 25 illustrates a first tracking methodology implemented in blocks 2530a and 2540a. In block 2530a, the processor may enable the DVS functionality of the DVS camera 2120 if it is not already enabled. This functionality may be enabled by setting intensity threshold changes associated with the movement of stationary features selected from a world model. In some embodiments, patches incorporating these features may also be set. In block 2540a, the processor may track head pose using patch data obtained in response to events detected by the DVS camera 2120. In some embodiments, the processor may be configured to calculate real-time or near-real-time user head pose from this patch data.
[0316] The secondary tracking methodology is illustrated in blocks 2530b and 2540b. In this embodiment, color image information may be used in combination with event information to track the relative position of features. In block 2530b, the processor may enable camera 2140 if camera 2140 is not already enabled. Camera 2140 may be enabled to provide images at a rate equal to or lower than the rate at which head pose updates are provided. For example, head pose updates may be provided through the use of asynchronous event data at an average rate of 30–60 Hz. Camera 2140 may be enabled to provide frames at a rate of less than 30 Hz, such as 5–10 Hz.
[0317] Color information may be used to increase the accuracy of tracking stationary features. For example, color information may be used to calculate the updated location of a tracked feature, which can be determined more accurately than if grayscale events were used alone. As a result, the location of the tracked patch may be updated or changed to encompass other features. Alternatively, or in addition, information from camera 2140 may be used to identify alternative features to track. Further alternatively, color information may allow the camera's relative position to be calculated based on the analysis of a surface, edge, or larger feature than those traced using DVS camera 2120.
[0318] The tertiary tracking methodology is illustrated in blocks 2530c and 2540c. In this embodiment, the tertiary tracking may be based on stereoscopic information. In block 2530b, the processor may disable the DVS functionality of the DVS camera 2120 so that the camera 2120 outputs intensity information rather than event information representing changes in intensity. In block 2540b, the processor may track head pose using images periodically output by the DVS camera 2120, which may be full frame images or images within the specific patch being tracked. In block 2540c, the processor may track head pose using stereoscopic image data obtained from images output by the DVS camera 2120 and camera 2140. For example, the processor may be configured to determine depth information from the stereoscopic image data. In some embodiments, the processor may be configured to calculate real-time or near-real-time user head pose from these images.
[0319] In block 2550, the processor may evaluate tracking quality according to tracking criteria. The tracking criteria may depend on factors such as the stability of the estimated head pose, the noise level of the estimated head pose, the consistency of the estimated head pose with a world model, or similar factors. In specific embodiments, the calculated head pose may be compared with other information that may indicate inaccuracies, such as the output of an inertial measurement unit or a model of the range of motion of the human head, so that errors in the head pose can be identified. Specific tracking criteria may vary depending on the tracking methodology used. For example, in methodologies using event-based information, correspondences between feature locations, such as those shown by event-based outputs compared to the locations of corresponding features in a complete frame image, may be used. Alternatively, or in addition, the visual discriminability of a feature relative to its surroundings may be used as a tracking criterion. For example, a tracking criterion for an event-based methodology may indicate poor tracking when the field of view is filled with one or more objects that make it difficult to identify the movement of a specific feature. The percentage of the field of view that is occluded is an example of a criterion that may be used. For example, a threshold such as exceeding 40% may be used as an indication of when to switch away from the use of image-based methodologies for head pose tracking. In a further embodiment, reprojection error may be used as a measure of head pose tracking quality. Such a criterion may be calculated by matching features in the available images with a previously determined passable world model. The location of the features in the images may be related to their location in the passable world model using a geometric transformation calculation based on head pose. For example, a deviation, expressed as the mean squared error between the calculated location and the feature in the passable world model, may indicate an error in head pose, so that the deviation can be used as a tracking criterion.
[0320] In some embodiments, the processor may be configured to calculate an error regarding the estimated head pose based on a world model. In calculating this error, the processor may be configured to estimate how the world model (or multiple features within the world model) should appear based on the estimated head pose. In some embodiments, this estimation may be compared to the world model (or features within the world model) to generate an error regarding the estimated head pose. Such an error may be calculated using linear algebra techniques. For example, the mean squared deviation between the calculated location and the actual location of each of the multiple features in the image may be used as a metric for the error. This metric may, in turn, be used as a measure of head pose tracking quality.
[0321] After evaluating the head pose tracking quality, method 2500 may return to block 2520, where the processor may select a tracking methodology using the measured head pose tracking quality. In scenarios where the tracking quality is low, such as below a threshold, an alternative tracking methodology may be selected.
[0322] Method 2500 may be repeated in this manner until a termination condition is detected. For example, Method 2500 may be repeated over a predetermined time period until user input is received or until a particular type or magnitude of change in a part of the physical world model within the field of view of the headset 2100's camera is detected. Method 2500 may then terminate in block 2599.
[0323] Method 2500 may be restarted so that new information about the world model, including that obtained using a higher-resolution color camera, is captured in block 2510. Method 2500 may be terminated and restarted so that the processing in block 2510 is repeated using the color camera, creating a portion of the world model at an average rate slower than the rate at which head pose tracking is performed in blocks 2520-2550. The processing using the color camera may be repeated, for example, at a rate of once per second or a slower average rate.
[0324] Other tracking methodologies may be used instead of, or in addition to, the tracking methodologies described above as embodiments. In some embodiments, the processor may be configured to calculate real-time or near-real-time user head pose from image information, which may include grayscale images and / or light field information such as angle of arrival information. Alternatively, or in addition, in scenarios where image-based head pose tracking methodologies have unacceptable quality metrics, a “dead reckoning” methodology may be selected, in which the motion of the user’s head, such as that measured by an inertial measurement unit, may be used to calculate the head pose.
[0325] Object Tracking
[0326] As described above, the processor of the XR system can track objects in the physical world and support realistic rendering of virtual objects relative to physical objects. Tracking has been described in relation to movable objects, for example, the user's hand in the XR system. For example, the XR system may track objects in the central field of view 2150, the peripheral field of view 2160a, and / or the peripheral field of view 2160b. Rapidly updating the position of movable objects enables realistic rendering of virtual objects, as rendering may reflect occlusion of physical objects by virtual objects or vice versa, or interactions between virtual objects within physical objects. In some embodiments, updates regarding the location of physical objects may be calculated at an average rate of at least 10 times / second, and in some embodiments at least 20 times / second, such as about 30 times / second. When the tracked object is the user's hand, tracking can enable gesture control by the user. For example, a gesture may correspond to a command to the XR system.
[0327] In some embodiments, the XR system may be configured to track objects that have features that provide high contrast when imaged using an image sensor sensitive to IR light. In some embodiments, objects with high-contrast features may be created by adding markers to the object. For example, a physical object may be equipped with one or more markers that appear as high-contrast regions when imaged using IR light. The markers may be passive markers that are highly reflective or highly absorbing to IR light. In some embodiments, at least 25% of the light over the frequency range of interest may be absorbed or reflected. Alternatively, or in addition, the markers may be active markers that emit IR light, such as IR LEDs. For example, by tracking such features using a DVS camera, information accurately representing the location of a physical object can be rapidly determined.
[0328] Similar to head pose tracking, the tracked object position is frequently updated, and therefore, performing object tracking using only the DVS camera 2120 may offer power savings, reduced computational requirements, or other advantages. The XR system may therefore be configured to disable the color camera 2140 or reduce its frame rate as needed to balance object tracking accuracy with power consumption and computational requirements. Furthermore, the XR system may be configured to track objects asynchronously in response to events generated by the DVS camera 2120.
[0329] Figure 26 is a simplified flowchart of Method 2600 for object tracking in several embodiments. According to Method 2600, the processor can perform object tracking differently depending on the field of view containing the object and the value of the tracking quality criterion. Furthermore, the processor may or may not use light field information depending on whether the object being tracked satisfies the depth criterion. Other criteria, such as available battery power or the operation of the XR system being performed and the need for those operations to track object locations or to track object locations with high accuracy, may be applied by the processor, either as an alternative or in addition, to dynamically select the object tracking methodology.
[0330] Method 2600 can begin with block 2601. In some embodiments, camera 2140 may have a disabled or reduced frame rate. The processor may disable camera 2140 or reduce the frame rate of camera 2140 to reduce power consumption and improve battery life. In various embodiments, the processor may track objects in the physical world (e.g., the user's hand). The processor may be configured to predict the next location or trajectory of an object based on one or more previous locations of the object.
[0331] Starting from block 2601, method 2600 can proceed to block 2610. In block 2610, the processor can determine the field of view that encompasses the object (e.g., field of view 2121, field of view 2141, peripheral field of view 2160a, peripheral field of view 2160b, or central field of view 2150). In some embodiments, the processor can base this determination on the object's current location (e.g., whether the object is currently in central field of view 2150). In various embodiments, the processor can base this determination on an estimate of the object's location. For example, the processor can determine that an object that is away from central field of view 2150 may enter peripheral field of view 2160a or peripheral field of view 2160b.
[0332] In block 2620, the processor can select an object tracking methodology. According to method 2600, when the object is within the field of view 2121 (e.g., within the peripheral field of view 2160a or the central field of view 2150), the processor may preferentially select object tracking using the DVS camera 2120. Furthermore, the processor may preferentially select event-based asynchronous object tracking.
[0333] In some embodiments, patch tracking as described above may be used with one or more patches established to encompass the features of the object being tracked, as described above. Patch tracking may be used for some or all of the object tracking methodologies and for some or all of the cameras. Patches may be selected to encompass the estimated location of the object being tracked within the field of view.
[0334] As described above in relation to Figure 25 and head pose tracking, the processor may dynamically select an appropriate tracking methodology for object tracking. The methodology may be selected to provide preferred tracking quality that requires low processing overhead compared to other methodologies. Therefore, if event-based asynchronous object tracking does not satisfy the object tracking criteria, the processor may use another methodology. In the embodiment of Figure 26, four methodologies are illustrated, shown in blocks 2640a, 2640b, 2640c, and 2640d. The methodologies are ordered, and method 2600 will select the first methodology in an order that satisfies the tracking quality criteria. The methodologies may be ordered to reflect trade-offs in, for example, accuracy, latency, and power consumption. The methodology with the shortest latency may be ordered first, for example, and methodologies with longer latency or higher power consumption may be ordered lower.
[0335] In block 2630a, the processor may enable the DVS functionality of the DVS camera 2120 if this functionality is not already enabled. In block 2640a, the processor may track objects using events detected by the DVS camera 2120. In some embodiments, the DVS camera 2120 may be configured to limit image acquisition to patches in the image sensor that encompass the location of objects within the field of view 2121. Changes in image data within a patch (e.g., caused by the movement of an object) may trigger an event. In response to an event, the DVS camera 2120 may acquire image data about the patch and update the object's position and / or orientation based on the acquired patch image data.
[0336] In some embodiments, events may indicate changes in intensity. For example, as described above in conjunction with Figures 12 and 13, changes in intensity may be tracked to track the motion of an object. In some embodiments, the DVS camera 2120 may be a plenooptic camera, or may be configurable to operate as such. When an object satisfies a depth criterion, the processor may be configured to obtain, in addition, or alternatively, angle of arrival information. The depth criterion may be identical or similar to the depth criterion described above with respect to block 2420 of method 2400. For example, the depth criterion may also be a maximum error rate or maximum distance beyond which the processor cannot distinguish between different distances between objects. Depth information may therefore be used to determine the location of an object and / or changes in the location of an object. The use of such plenooptic image information may be part of the methodology described herein, or may be a further methodology that can be used in conjunction with other methodologies.
[0337] A second methodology is shown in blocks 2630b and 2640b. In block 2630b, the processor may enable camera 2140 if camera 2140 is not already enabled. As described above in relation to block 2530b, the color camera may operate to acquire color images at a relatively low average rate. In block 2640b, color information may be used in conjunction with information from the color images to better identify features or their locations, primarily from events output by the DVS camera 2120.
[0338] The tertiary methodology is shown in blocks 2630c and 2640c. In block 2630c, the processor may enable camera 2140 if camera 2140 is not already enabled. Camera 2140 is described as being capable of outputting color image information. With respect to the tertiary methodology, color image information may be used, but in some embodiments or scenarios, only intensity information may be obtained from camera 2140, or only intensity information may be processed. The processor may also increase the frame rate of camera 2140 to a rate sufficient for object tracking (e.g., a frame rate of 40Hz to 120Hz). In some embodiments, this rate may match the sampling frequency of DVS camera 2120. The processor may also disable the DVS functionality of DVS camera 2120 if this functionality is not already disabled. In this configuration, the output of DVS camera 2120 may represent intensity information. In block 2640c, the processor may track objects using stereoscopic image data acquired from images output by the DVS cameras 2120 and 2140.
[0339] A fourth methodology is shown in blocks 2630d and 2640d. In block 2630d, the processor may enable a camera if it is not already enabled. Other cameras may be disabled. The camera to be enabled may be DVS camera 2120 or camera 2140. If DVS camera 2120 is used, it may be configured to output image intensity information. If camera 2140 is enabled, it may be enabled to output color information or simply grayscale intensity information. The processor may also increase the frame rate of the enabled camera to a rate sufficient for object tracking (e.g., a frame rate of 40Hz to 120Hz). In block 2640d, the processor may use the image output by the enabled camera to track an object.
[0340] In block 2650, the processor may evaluate tracking quality. One or more of the metrics described above in relation to block 2550 may be used in block 2650. However, in block 2650, those metrics would be applied to tracking features on objects rather than tracking stationary features in the environment.
[0341] After evaluating the object tracking quality, method 2600 may return to blocks 2610 and 2620, where the processor again determines the field of view encompassing the object (block 2610), and then, using the determined field of view and the measured object tracking quality, selects an object tracking methodology (block 2620). In the illustrated embodiment, a first methodology is selected in a sequence that provides quality above a threshold associated with desirable performance.
[0342] Method 2600 may be repeated in this manner until a termination condition is detected. For example, Method 2600 may be repeated over a predetermined time period until user input is received or until the tracked object leaves the field of view of the XR device. Method 2600 may then terminate in block 2699.
[0343] Method 2600 may be used to track any object, including a user's hand. However, in some embodiments, different or additional actions may be performed when tracking a user's hand. Figure 27 is a simplified flowchart of Method 2700 for hand tracking according to some embodiments. The object tracked in Method 2700 may be a user's hand. In various embodiments, the XR system can perform hand tracking using image data or images acquired from DVS camera 2120 and / or images acquired from camera 2140. These cameras may be configured to operate in one of several modes. For example, one or both may be configured to acquire image patch data. Alternatively, or in addition, DVS camera 2140 may be configured to output events as described above with respect to Figure 4-16. Alternatively, or in addition, camera 2140 may be configured to output color information or simply grayscale intensity information. Furthermore, camera 2140 may be configured to output plenooptic information as described above with respect to Figure 16-20. Similarly, the DVS camera 2120 may be configured to output plenooptic information in some embodiments. Any combination of these cameras and capabilities may be selected to generate information related to hand tracking.
[0344] Similar to object tracking described in relation to Figure 26, the processor may be configured to select a hand tracking approach based on the determined location of the hand and an assessment of the quality of hand tracking provided by the selected approach. For example, stereoscopic depth information may be acquired when the hand is within the central field of view 2150. In another embodiment, plenooptic information may have sufficient resolution only for tracking the hand within a certain angular range (e.g., + / - 20 degrees) relative to the center of the field of view of the plenooptic camera. Thus, techniques relying on stereoscopic or plenooptic image information may be used only when it is detected that the hand is within the appropriate field of view.
[0345] If necessary, the XR system can enable camera 2140 or increase its frame rate to enable tracking within the field of view 2160b. Thus, the wearable display system may be configured in this configuration to provide proper hand tracking using a reduced number of available cameras, enabling reduced power consumption and increased battery life.
[0346] Method 2700 may be performed under the control of the XR system's processor. The method may be initiated in response to the detection of an object to be tracked, such as a hand, as a result of analyzing images obtained using any one of the cameras on the headset 2100. This analysis may involve the step of recognizing the object as a hand based on areas of the image having photometric properties, which are characteristics of a hand. Alternatively, or in addition, depth information obtained from stereoscopic image analysis may be used to detect the hand. In a specific embodiment, the depth information may indicate the presence of an object having a shape that matches a 3D model of a hand. Detecting the presence of a hand in this way may also involve the step of setting the parameters of the hand model to match the orientation of the hand. In some embodiments, such a model may also be used for high-speed hand tracking by using photometric information from one or more grayscale cameras to determine how the hand has moved from its original position.
[0347] Other trigger conditions may initiate Method 2700, such as the XR system performing an action involving an object tracking step, such as rendering a virtual button that the user is likely to attempt to press with their hand, so that the user's hand is expected to enter the field of view of one or more cameras. Method 2700 may be repeated at a relatively high rate, such as 30 to 100 times per second, for example, 40 to 60 times per second. As a result, updated positional information about the tracked object can be made available with short latency for processing to render virtual objects interacting with the physical object.
[0348] Starting from block 2701, method 2700 may proceed to block 2710. In block 2710, the processor may determine the location of the potential hand. In some embodiments, the location of the potential hand may be the location of a detected object in the acquired image. In embodiments where the hand is detected based on matching depth information to a 3D model of the hand, the same information may be used as the initial position of the hand in block 2710.
[0349] In some embodiments, the image information used to construct the hand model may be dynamically selected. After block 2710, method 2700 may proceed to block 2720. In block 2720, the processor may select a processing approach for constructing the hand model. This selection may depend on the potential location of the hand and whether a hand model has already been constructed. In some embodiments, this selection may further depend on the desired accuracy and / or quality of the hand model and / or the processing computer power available, given other tasks being performed by the processor. As an example, when the hand model is initially constructed in response to hand detection, broader, but more precise, image information may be used at the expense of additional processing. For example, stereoscopic information may be used to initially construct the model. The model itself then provides information about the hand's position, since there are limits to the ways in which a human hand can move. Thus, as the XR system operates, processing to reconstruct the hand model and account for hand movement may be performed, at will, using image information that can be processed more quickly, even if less comprehensive, or when more comprehensive image information is unavailable. For example, if an object is within the central field of view 2150 and robust hand tracking or fine hand detail is required, or if other approaches prove unsuitable, the processor may perform stereoscopic hand tracking using images output by the DVS camera 2120 and camera 2140. Alternatively, if an object is not within the central field of view 2150, or robust hand tracking or fine hand detail is not required, the processor may be configured to preferentially construct a hand model using other image information, such as monocular color images or grayscale images. The processor may be configured to construct a hand model using images obtained from the DVS camera 2120 when the object is within the peripheral field of view 2160a. The processor may be configured to construct a hand model using images obtained from camera 2140 when the object is within the peripheral field of view 2160b.In some embodiments, the processor may be configured to use a minimum power consumption method that still meets quality standards, for example.
[0350] Figure 27 illustrates four approaches to gathering information for constructing a hand model. These four approaches are shown as four parallel paths through a flowchart, including paths through blocks 2740a, 2740b, 2740c, and 2740d.
[0351] In the first path, in block 2730a, when a potential hand is within the central field of view 2150 and hand tracking robustness or fine hand detail is required, the processor may enable camera 2140 if it is not already enabled. This path may also be selected in block 2720 for the initial configuration of the hand model. In block 2730a, the processor may also increase the frame rate of camera 2140 to a rate sufficient for hand tracking (e.g., a frame rate of 40Hz to 120Hz). In some embodiments, this rate may match the sampling frequency of DVS camera 2120. The processor may also disable the DVS functionality of DVS camera 2120 if it is not already disabled. As a result, both cameras may provide image information representing intensity. This image information may be grayscale, or, for cameras that support color imaging, may instead include color information.
[0352] In block 2740a, the processor may acquire depth information about the potential hand. The depth information may be acquired based on stereoscopic image analysis, from which the distance between the camera collecting image information and the segment or feature of the potential hand can be calculated. The processor may, for example, select a feature in the central field of view and determine depth information about the selected feature.
[0353] In some embodiments, the selected features may represent different segments of a human hand, such as those defined by bones and joints. Feature selection may be based on matching image information to a model of a human hand. Such matching may be performed heuristically, for example. The human hand may be represented by a finite number of segments, such as 16, and points in the image of the hand may be mapped to one of those segments so that features on each segment can be selected. Alternatively, or in addition, such matching may use a deep neural network or a classification / decision forest to apply a series of yes / no decisions in the analysis to identify different parts of the hand and select features that represent different parts of the hand. The matching may, for example, identify whether a particular point in the image belongs to the palm, back of the hand, non-thumb fingers, thumb, fingertips, and / or finger joints. Any suitable classifier can be used for this analysis stage. For example, a deep learning module or a neural network mechanism can be used instead of or in addition to a classification forest. In addition, a regression forest (e.g., using the Hough transform, etc.) can be used in addition to a classification forest.
[0354] Regardless of the specific number of features to be selected and the technique used to select those features, after block 2740a, method 2700 may proceed to block 2750a. In block 2750a, the processor may construct a hand model based on depth information. In some embodiments, the hand model may reflect structural information about a human hand, representing each of the bones in the hand as segments within the hand, and each joint, to define the range of likely angles between adjacent segments. Information about the position of the hand may be provided for subsequent processing by the XR system by assigning locations to each of the segments in the hand model based on the depth information of the selected features.
[0355] Regardless of how the 3D hand model is updated, the updated model may be refined based on photometric image information. The model may be used, for example, to generate a projection of the hand and represent how the image of the hand is expected to appear. The expected image may be compared with photometric image information obtained using an image sensor. The 3D model may be adjusted to reduce the error between the expected photometric information and the obtained photometric information. The adjusted 3D model then provides an indication of the hand's position. As this update process is repeated, the 3D model provides an indication of the hand's position as the hand moves.
[0356] In some embodiments, the processing in blocks 2740a and 2750a may be performed iteratively, with respect to which depth information is collected, and the feature selection is refined based on the configuration of the hand model. The hand model may include shape constraints and motion constraints that the processor may use to refine the feature selection that represents the components of the hand. For example, if a feature selected to represent a segment of the hand indicates the position or motion of that segment in violation of the constraints of the hand model, a different feature may be selected to represent that segment.
[0357] In block 2750a, the processor may select features in a patch or image output by the DVS camera 2120 or camera 2140 that represent the structure of a human hand. Such features may be identified heuristically or using AI techniques. For example, features may be heuristically selected by representing the human hand with a finite number of segments such that a feature can be selected on each segment, and mapping points in the image to the individual segments. Alternatively, or in addition, such matches may be selected by applying a series of yes / no decisions within the analysis using a deep neural network or classification / decision forest to identify different parts of the hand and select features that represent different parts of the hand. Any suitable classifier can be used for this analysis stage. For example, a deep learning module or neural network mechanism can be used instead of or in addition to a classification forest. In addition, a regression forest (e.g., using the Hough transform, etc.) can be used in addition to a classification forest. The processor may attempt to match the selected features and the motion of those selected features per image to a hand model without the benefit of depth information. This matching may yield less robust or less accurate information than that generated in block 2750a. Nevertheless, the information identified based on monocular data can provide useful information about the operation of the XR system.
[0358] In block 2750a, the processor can also evaluate the hand model. This evaluation may depend on the completeness of the match between the selected features and the hand model, the stability of the match between the selected features and the hand model, the noise level of the estimated hand location, the consistency between the location and orientation of the detected features and the constraints imposed by the hand model, or similar factors. In some embodiments, the processor can determine, based on the hand model, where the selected features should appear. In a specific example, the processor may parameterize a generic hand model, then check the photometric consistency of the model's edges, and compare those edges to those detected in the acquired image of the hand. In some embodiments, these estimates may be compared to the estimated location and orientation of the selected features to generate an error regarding the hand model. Such an error may be calculated using linear algebra techniques. For example, the mean squared deviation between the calculated location and the actual location of each of several selected features in the image may be used as a metric for the error. This metric may, in turn, be used as a measure of the hand model quality.
[0359] After matching the image portion to the hand model portion in block 2750a, the model may be used in one or more actions performed by the XR system. In one embodiment, the model may be used directly as an indication of the location of an object in the physical world to render a virtual object. In such an embodiment, the processing in block 2760 may be optionally omitted. Alternatively, or in addition, the model may be used to determine whether the user has made a gesture, such as a gesture indicating a command or interaction with a virtual object, using their hand. Thus, the processor may use the determined hand model information in block 2760 to recognize a hand gesture. This gesture recognition may be performed using a hand tracking method described in U.S. Patent Publication No. 2016 / 0026253 (which is incorporated herein by reference in all respects relating to hand tracking and the use of information about the hand obtained from image information in the XR system).
[0360] In some embodiments, gestures may be recognized without stereoscopic depth information. Gestures may be recognized based on, for example, a hand model constructed based on monocular image information. Thus, in some embodiments, gesture recognition may be performed even with respect to hands within peripheral fields 2160a and 2160b when stereoscopic depth information is unavailable. Alternatively, or in addition, monocular image information may be used when hand tracking is performed for gesture recognition rather than other functions that may require a more precise determination of the hand's location (e.g., rendering virtual objects to appear so as to realistically interact with the user's hand). In such embodiments, different cameras may be enabled and / or used to collect image information in blocks 2730a and 2740b.
[0361] In some embodiments, a series of iterations of the hand tracking process may be performed as the system operates. Gestures may be identified, for example, by a series of hand position determinations. Alternatively, or in addition, the series of iterations may be performed so that a constructed model of the user's hand matches the actual hand position as the hand position changes. The same or different imaging techniques may be used in each iteration.
[0362] In some embodiments, the source of image information may be selected for iteration based on one or more factors, including the application consisting of a hand model in block 2720 and / or the quality of the tracking performed using a specific tracking methodology. The quality of the tracking may be determined, for example, using an approach such as the one described above in relation to block 2650 (Figure 26), and the selection of the technique may also be done in such a way as to require the least amount of processing and power consumption on other computing resources to achieve the desired quality metric, as described above.
[0363] Therefore, Figure 27 shows that once an iteration is performed, method 2700 returns to block 2720, and a processing approach may be selected for further iterations. These alternative approaches may use image information instead of, or in addition to, stereoscopic depth information. In each iteration, the 3D model of the hand may be updated to reflect the hand's movement.
[0364] An approach requiring low processing power tracks the hand position based on event information. Such an approach may be selected in block 2720 in a scenario where, for example, a hand model of a suitable accuracy has already been calculated, for example, through the use of stereoscopic image information in a previous iteration. Such an approach may be implemented by branching to block 2730b. In block 2730b, the processor may enable the DVS functionality of the DVS camera 2120 if this functionality is not already enabled. In block 2740b, the processor may obtain image data in response to events detected by the DVS camera 2120. In some embodiments, the DVS camera 2120 may be configured to obtain information about a patch in the image sensor that contains the location of a potential hand within the field of view 2121. A change in image data within the patch (e.g., caused by potential hand movement) may trigger an event. In response to the event, the DVS camera 2120 may obtain image data. This information about the user's hand movement may then be used in block 2750b to update the hand model. The update may be performed using the techniques described above with respect to block 2750a. Such an update may take into account other information, including constraints on the previously calculated position of the hand and the movement of the human hand, as shown by the model.
[0365] Once updated, the hand model may be used by the system in the same manner as the initial hand model calculated in block 2750a, including determining the interaction between the virtual object and the user's hand and / or recognizing gestures in block 2760.
[0366] In some scenarios, event-based image information may be unsuitable even for updating the hand model. Such a scenario may occur, for example, when the user's hand fills the field of view of the DVS camera 2120. In such scenarios, the hand model may be updated based on color information, such as using color information from camera 2140, rather than event information. Such an approach may be implemented by branching to block 2730d when processing in block 2720 detects other characteristics indicating that the images within the field of view of the DVS camera 2120 have intensity fluctuations below a predetermined threshold, or that there is a lack of features that are clearly distinct enough for event-based tracking. In block 2730d, the processor may disable the DVS functionality of the DVS camera 2120 if this functionality has not already been disabled. Camera 2140 may be enabled to obtain color information.
[0367] In block 2740d, the processor may use color image information to identify the user's hand movements. As described above, this information regarding the user's hand movements may then be used in block 2750b to update the hand model.
[0368] In some scenarios, color information may be unavailable or unnecessary for tracking hand movements. For example, when the hand is within the peripheral field of view 2160a, color information may be unavailable, but intensity image information acquired using the DVS camera 2120, operating in a mode where DVS functionality is disabled, may be available and preferable. Alternatively, in some scenarios where event-based tracking produces a quality metric below a threshold, tracking using grayscale image information may produce a quality metric above the suitability threshold. Such an approach may be implemented by branching to block 2730c. In block 2730c, the processor may enable camera 2140 if it is not already enabled. The processor may also increase the frame rate of camera 2140 to a rate sufficient for hand tracking (e.g., a frame rate of 40Hz to 120Hz).
[0369] In block 2740d, the processor can use camera 2140 to obtain image data. In other scenarios, DVS camera 2120 may be configured to collect intensity information and may be used to collect monocular image information instead of camera 2140. Regardless of which camera is used, full frame information may be used, or patch tracking may be used to reduce the amount of image information processed, as described above.
[0370] Figure 27 illustrates how the process in block 2720 selects from three alternative approaches to track a hand. The approaches may be ordered according to the degree to which they meet one or more criteria, such as low processing or low power consumption. The process in block 2720 may select a processing approach by selecting a first approach in order that is operational in the detected scenario (e.g., hand location) and produces quality metrics that meet a threshold. Different or additional processing techniques may be included. For example, with respect to an XR system with a plenooptic camera that provides depth-indicating image information, the approach may be based on the use of that depth information alone or in combination with any other data source. In another variation, hand features may be tracked using color-assisted DVS tracking, as described above in relation to block 2640b (Figure 26).
[0371] Regardless of the approach selected or the set of approaches from which such selection is made, the hand model may be updated and used in block 2760 for XR functionality, such as rendering a virtual object that interacts with the user's hand or detecting gestures. After block 2760, method 2700 may terminate in block 2799. However, it should be understood that hand tracking may occur continuously during the operation of the XR system or during intervals in which the hand is within the field of view of one or more cameras. Thus, once one iteration of method 2700 is completed, another iteration may be performed, and this process may be performed over intervals in which hand tracking is being performed. In some embodiments, information used in one iteration may be used in subsequent iterations. In various embodiments, for example, the processor may be configured to estimate the updated location of the user's hand based on the previously detected hand location. For example, the processor may estimate where the user's hand will next be based on the previous location and the velocity of the user's hand. Such information may be used to narrow down the amount of image information that is processed to detect the location of an object, as described above in relation to patch tracking techniques.
[0372] Therefore, although some aspects of several embodiments have been described, it should be understood that various modifications, alterations, and improvements will be readily conceivable to those skilled in the art.
[0373] As one example, the embodiment is described in relation to an augmented reality (AR) environment. It should be understood that some or all of the techniques described herein may be applied to an mixed reality (MR) environment, or more generally, to other XR environments.
[0374] Furthermore, an embodiment of an image array is described in which one patch is applied to the image array and controls the selective output of image information about one movable object. It should be understood that there may be more than one movable object in a physical embodiment. Moreover, in some embodiments, it may be desirable to selectively acquire frequent updates of image information in areas other than where the movable object is located. For example, a patch may be configured to selectively acquire image information about the area of the physical world in which a virtual object should be rendered. Thus, some image sensors may be able to selectively provide information about two or more patches, with or without a network for tracking the trajectory of those patches.
[0375] In a further embodiment, the image array is described as outputting information related to the magnitude of incident light. The magnitude may be a representation of power across the spectrum of optical frequencies. The spectrum may have a relatively wide capture energy at frequencies corresponding to any color of visible light in a monochrome camera, etc. Alternatively, the spectrum may be narrow, corresponding to a single color of visible light. A filter may be used for this purpose to restrict the light incident on the image array to light of a specific color. If pixels are restricted to receive light of a specific color, different pixels may be restricted to different colors. In such embodiments, the outputs of pixels sensitive to the same color may be processed together.
[0376] A process for setting up a patch in an image array and then updating the patch for an object of interest has been described. This process may be performed, for example, for each movable object as it enters the field of view of the image sensor. The patch may be cleared when the object of interest leaves the field of view so that the patch is no longer tracked or image information is no longer output for the patch. It should be understood that the patch may be updated from time to time by determining the location of the object associated with the patch and setting the position of the patch to correspond to that location. Similar adjustments may be made to the calculated trajectory of the patch. Motion vectors for the object and / or motion vectors for the image sensor may be calculated from other sensor information and used to reset values programmed into the image sensor or other components for patch tracking.
[0377] For example, the location, motion, and other characteristics of an object may be determined by analyzing the output of a wide-angle video camera or a pair of video cameras with stereoscopic information. Data from these other sensors may be used to update the world model. In connection with the update, patch location and / or trajectory information may be updated. Such updates may occur at a lower rate than the patch location is updated by the patch tracking engine. The patch tracking engine may calculate the new patch location at a rate of, for example, about 1 to 30 times per second. Patch location updates based on other information may occur at even slower rates, such as 1 time per second to about 1 time per 30-second intervals.
[0378] As yet another embodiment of the modification, Figure 2 shows a system with a head-mounted display separate from the remote processing module. Image sensors such as those described herein can lead to a compact system design. Such sensors generate less data, which in turn leads to lower processing requirements and less power consumption. The lower processing and power requirements allow for size reduction, such as by reducing the size of the battery. Therefore, in some embodiments, the entire augmented reality system may be integrated into a head-mounted display without a remote processing module. The head-mounted display may be configured as a pair of goggles, or, as shown in Figure 2, may be similar in size and shape to a pair of eyeglasses.
[0379] Furthermore, embodiments in which the image sensor responds to visible light are described. It should be understood that the techniques described herein are not limited to operation with visible light. They may, alternatively or in addition, respond to “light” in other parts of the spectrum, such as IR light or UV light. Furthermore, image sensors such as those described herein respond to naturally occurring light. Alternatively or in addition, the sensor may be used in a system with an illumination source. In some embodiments, the sensitivity of the image sensor may be adjusted to the part of the spectrum from which the illumination source emits light.
[0380] In another embodiment, it is explained that a selected region of an image array, where changes should be output from an image sensor, is defined by defining a “patch” on which image analysis should be performed. However, it should be understood that the patch and the selected region may be of different sizes. The selected region may be larger than the patch, for example, to account for the motion of an object in the tracked image that deviates from a predicted trajectory, and / or to allow processing around the edges of the patch.
[0381] Furthermore, several processes such as passable world model generation, object tracking, head pose tracking, and hand tracking are described. These, and in some embodiments, other processes, may be executed by the same or different processors. The processors may be operated to enable the simultaneous operation of these processes. However, each process may be executed at a different rate. If different processes request data from an image sensor or other sensor at different rates, the acquisition of sensor data may be managed by another process, etc., to provide data to each process at a rate appropriate for its operation.
[0382] Such modifications, alterations, and improvements are intended to be part of the disclosure and to be within the spirit and scope of the disclosure. For example, in some embodiments, the color filter 102 of the image sensor pixels may not be a separate component, but instead be incorporated into one of the other components of the pixel subarray 100. For example, in an embodiment including a single pixel with both an angle-of-arrival / position intensity converter and a color filter, the angle-of-arrival / intensity converter may be a transmissive optical component formed from a material that filters specific wavelengths.
[0383] According to some embodiments, a wearable display system may be provided, comprising: a frame; a first camera mechanically coupled to the frame, which is configured to output image data satisfying a first field of view intensity change criterion with respect to the first camera; and a processor operably coupled to the first camera, which is configured to determine whether an object is in the first field of view and to track the motion of the object using image data received from the first camera with respect to one or more portions of the first field of view.
[0384] In some embodiments, the object may be a hand, and the step of tracking the object's movement may include updating corresponding parts of a hand model, including shape constraints and / or motion constraints, based on image data from a first camera that satisfies the intensity change criterion.
[0385] In some embodiments, the processor may further be configured to provide instructions to the first camera to restrict image data acquisition to one or more patches of the first field of view corresponding to objects.
[0386] In some embodiments, a second camera may be mechanically coupled to the frame to provide a second field of view that at least partially overlaps with the first field of view, and the processor may further be configured to determine whether an object satisfies occlusion criteria with respect to the first field of view, enable the second camera or increase the frame rate of the second camera, use the second camera to determine depth information about the object, and use the determined depth information to track the object.
[0387] In some embodiments, depth information may be determined stereoscopically using images output by a first camera and a second camera.
[0388] In some embodiments, depth information may be determined using light field information output by a second camera.
[0389] In some embodiments, the object may be a hand, and the step of tracking the hand's movement may include using the determined depth information by selecting a point in a first field of view, associating the selected point with depth information, generating a depth map using the selected point, and matching a portion of the depth map to a corresponding portion of a hand model, including both shape constraints and motion constraints.
[0390] In some embodiments, the step of tracking the motion of an object may include the step of updating the object's location in a world model, and the interval between updates may have a duration of 1 ms to 15 ms.
[0391] In some embodiments, a second camera may be mechanically coupled to the frame to provide a second field of view that at least partially overlaps with the first field of view, and a processor may be operably coupled to the second camera and further configured to create a world model using images output by the first and second cameras, and to update the world model using light field information output by the second camera.
[0392] In some embodiments, the processor may be mechanically coupled to the frame.
[0393] In some embodiments, the display device, which is mechanically coupled to the frame, may include a processor.
[0394] In some embodiments, the local data processing module may include a processor, and the local data processing module is operably coupled to a display device via a communication link, and the display device is mechanically coupled to a frame.
[0395] According to some embodiments, a wearable display system may be provided, comprising a frame, two cameras mechanically coupled to the frame, one first camera and one second camera, each configured to output image data satisfying an intensity change criterion, the first and second cameras being positioned to provide an overlapping view of the central field of view, and a processor operably coupled to the first and second cameras.
[0396] In some embodiments, the processor may further determine that an object is within the central field of view, track the object using image data output by the first camera, determine whether the object tracking meets quality criteria, enable the second camera or increase the frame rate of the second camera, and, if the object tracking does not meet quality criteria, track the object using the first and second cameras.
[0397] In some embodiments, the first camera may be configured to selectively output image frames or image data that satisfy an intensity change criterion, and the processor may further be configured to provide instructions to the first camera to limit image data acquisition to one or more portions of the central field of view corresponding to an object, and the image data output by the first camera may relate to one or more portions of the central field of view.
[0398] In some embodiments, the first camera may be configured to selectively output image frames or image data that satisfy an intensity change criterion, the second camera may include a plenooptic camera, and the processor may further be configured to determine whether an object satisfies a depth criterion, and the step of tracking an object using the first and second cameras may include, when the depth criterion is met, tracking the object using light field information obtained from the plenooptic camera, and when the depth criterion is not met, tracking the object using depth information stereoscopically determined from images output by the first and second cameras.
[0399] In some embodiments, the prenoptic camera may include a transmission diffraction mask.
[0400] In some embodiments, the plenoptic camera may have a horizontal field of view of 90 to 140 degrees, and the central field of view may extend to 40 to 80 degrees.
[0401] In some embodiments, the first camera may provide a grayscale image, and the second camera may provide a color image.
[0402] In some embodiments, the first camera may be configured to selectively output image frames or image data that satisfy an intensity change criterion, the first camera may include a global shutter, and the second camera may include a roll shutter, and the processor may further be configured to compare a first image obtained using the first camera with a second image obtained using the second camera, to detect distortion in at least a portion of the second image, to adjust at least a portion of the second image, and to compensate for the detected distortion.
[0403] In some embodiments, the step of comparing a first image obtained using a first camera with a second image obtained using a second camera may include the step of performing a line-by-line comparison between the first image and the second image.
[0404] In some embodiments, the processor may be mechanically coupled to the frame.
[0405] In some embodiments, the frame includes a display device that is mechanically coupled to the processor.
[0406] In some embodiments, the local data processing module may include a processor, and the local data processing module may be operably coupled to a display device via a communication link, and the frame may include a display device.
[0407] In some embodiments, the processor may further be configured to determine whether an occlusion criterion is met for the first camera, enable the second camera or increase its frame rate, and track an object using the image data output by the second camera.
[0408] In some embodiments, the occlusion criterion may be satisfied when the object occupies more than a threshold amount of the field of view with respect to the first camera.
[0409] In some embodiments, the object may be a stationary object in the environment.
[0410] In some embodiments, the object may be the hand of a user of the wearable display system.
[0411] In some embodiments, the wearable display system further includes an IR emitter that is mechanically coupled to the frame.
[0412] In some embodiments, the IR emitter is configured to be selectively activated to provide IR illumination.
[0413] Furthermore, while the advantages of this disclosure are shown, it should be understood that not all embodiments of this disclosure include all described advantages. Some embodiments may not implement any of the features described as advantageous herein. Therefore, the foregoing description and drawings are merely examples.
[0414] The embodiments described above in this disclosure can be implemented in any of a number of ways. For example, embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can run on any suitable processor or set of processors, whether provided within a single computer or distributed across multiple computers. Such processors may be implemented as an integrated circuit, together with one or more processors in an integrated circuit component, including, to name a few, commercially available integrated circuit components known in the art, such as CPU chips, GPU chips, microprocessors, microcontrollers, or coprocessors. In some embodiments, the processor may be implemented in a custom circuit such as an ASIC, or in a semi-custom circuit resulting from constituting a programmable logic device. As a further alternative, the processor may be part of a larger circuit or semiconductor device, whether commercial, semi-custom, or custom. In specific examples, some commercially available microprocessors have multiple cores such that one or a subset of their cores may constitute a processor. However, the processor may be implemented using circuitry in any suitable format.
[0415] Furthermore, it should be understood that a computer can be embodied in any of several forms, such as a rack-mount computer, desktop computer, laptop computer, or tablet computer. In addition, a computer may be embodied in a device that is not generally considered a computer but has suitable processing capabilities, including a personal digital assistant (PDA), a smartphone, or any suitable portable or fixed electronic device.
[0416] Furthermore, a computer may have one or more input and output devices. These devices can, among other things, be used to present a user interface. Embodiments of output devices that may be used to provide a user interface include a printer or display screen for visual presentation of output, or a speaker or other sound-generating device for audible presentation of output. Embodiments of input devices that may be used for a user interface include a keyboard and pointing devices such as a mouse, touchpad, and digitized tablet. In another embodiment, a computer may receive input information through speech recognition or in other audible formats. In the illustrated embodiments, the input / output devices are illustrated as physically separate from the computing device. However, in some embodiments, the input and / or output devices may be physically integrated into the same unit as the processor or into other elements of the computing device. For example, a keyboard may be implemented as a soft keyboard on a touchscreen. In some embodiments, the input / output devices may be completely disconnected from the computing device and functionally integrated through a wireless connection.
[0417] Such computers may be interconnected by one or more networks of any preferred form, including local area networks or wide area networks, such as corporate networks or the Internet. Such networks may be based on any preferred technology, operate according to any preferred protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0418] Furthermore, the various methods and processes outlined herein may be coded as software that can run on one or more processors employing any one of various operating systems or platforms. In addition, such software may be written using any of several suitable programming languages and / or programming or scripting tools, and may be compiled as executable machine language code or intermediate code that runs on a framework or virtual machine.
[0419] In this regard, the Disclosure may be embodied as a computer-readable storage medium (or more computer-readable media) (e.g., computer memory, one or more floppy disks, compact disks (CDs), optical disks, digital video disks (DVDs), magnetic tape, flash memory, circuit configurations in field-programmable gate arrays or other semiconductor devices, or other tangible computer storage media) encoded with one or more programs that, when executed on one or more computers or other processors, implement the various embodiments of the Disclosure discussed above. As will be apparent from the embodiments described above, the computer-readable storage medium may retain information for a period of time sufficient to provide computer-executable instructions in a non-transient form. Such a computer-readable storage medium or more media may be transportable, as described above, so that one or more programs stored thereon can be loaded onto one or more different computers or other processors to implement the various aspects of the Disclosure. As used herein, the term “computer-readable storage medium” includes only computer-readable media that can be considered a manufacture (i.e., a product) or machine. In some embodiments, the present disclosure may be embodied as a computer-readable medium other than a computer-readable storage medium such as a propagating signal.
[0420] The terms “program” or “software” are used herein in a general sense to refer to any type of computer code or set of computer executable instructions that may be employed to program a computer or other processor to implement various aspects of the Disclosure, as described above. In addition, it should be understood that, according to one aspect of this embodiment, when executed to perform the methods of the Disclosure, one or more computer programs do not need to reside on a single computer or processor, but can be distributed in a modular manner among several different computers or processors to implement various aspects of the Disclosure.
[0421] Computer executable instructions can take many forms, such as program modules, which are executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. Typically, the functionality of program modules may be combined or distributed as desired in various embodiments.
[0422] Furthermore, the data structure may be stored in a computer-readable medium in any preferred form. For simplicity of illustration, it may be shown that the data structure has fields that are related through locations within the data structure. Such relationships may also be achieved by allocating storage for the fields, along with locations in the computer-readable medium that convey the relationships between the fields. However, any preferred mechanism, including the use of pointers, tags, or other mechanisms for establishing relationships between data elements, may be used to establish relationships between the information in the fields of the data structure.
[0423] Various aspects of this disclosure may be used individually, in combination, or in various arrangements not specifically discussed in the embodiments described above, and therefore their applications are not limited to the details and arrangements of components described in the above description or illustrated in the drawings. For example, an aspect described in one embodiment may be combined in any way with an aspect described in another embodiment.
[0424] Furthermore, the disclosure may be embodied as a method, in which its embodiments are provided. Actions performed as part of the method may be ordered in any preferred manner. Thus, while the illustrative embodiments are shown as a series of actions, embodiments may be constructed in which the actions are performed in a different order than those illustrated, which may include performing several actions simultaneously.
[0425] The use of sequential terms such as “first,” “second,” “third,” etc. in a claim to modify a claim element does not, by itself, imply any priority, precedence, or sequence, or chronological order in which any action of method is performed of a single claim element compared to another element; however, sequential terms are used only as identifiers to distinguish claim elements, (for the purpose of using sequential terms) from one claim element having a certain name to another element having the same name.
[0426] Furthermore, the terms and technical descriptions used herein are for illustrative purposes only and should not be considered limiting. The use herein of “including,” “equipped with,” “having,” “containing,” “accompanying,” and variations thereof means that the items subsequently listed, and their equivalents and supplementary items, are included.
Claims
1. A portable device, wherein the portable device is A first camera configured to output an image frame or image data that meets an intensity change criterion, A second camera, A processor operably coupled to the first camera and the second camera Equipped with, The first camera and the second camera are positioned to provide an overlapping view of the central field of view. The aforementioned processor, Using depth information stereoscopically determined from images output by the first camera and the second camera, a world model is created. Using the world model and the image data output by the first camera, a tracking routine is executed. The tracking routine determines whether the quality standards are met, When the tracking routine does not meet the quality criteria, enable the second camera or modulate the frame rate of the second camera. A portable device configured to perform the following actions.
2. A method for performing a tracking routine using a portable device, wherein the portable device is A first camera configured to output an image frame or image data that meets an intensity change criterion, A second camera, A processor operably coupled to the first camera and the second camera Equipped with, The first camera and the second camera are positioned to provide an overlapping view of the central field of view. The above method uses the processor, Using depth information stereoscopically determined from images output by the first camera and the second camera, a world model is created. Using the world model and the image data output by the first camera, a tracking routine is executed. The tracking routine determines whether the quality standards are met, When the tracking routine does not meet the quality criteria, enable the second camera or modulate the frame rate of the second camera. Methods that include...
3. A portable device, wherein the portable device is Frame and, A first camera mechanically coupled to the frame, wherein the first camera is configured to output image data that satisfies a first field of view intensity change criterion related to the first camera, A second camera, A processor operably coupled to the first camera and Equipped with, The aforementioned processor, To determine whether the point of interest is within the first field of view, Using image data received from the first camera with respect to one or more portions of the first field of view, a tracking routine is performed. The tracking routine determines whether the quality standards are met, When the tracking routine does not meet the quality criteria, enable the second camera or modulate the frame rate of the second camera. A portable device configured to perform the following actions.
4. A method for performing a tracking routine using a portable device, wherein the portable device is Frame and, A first camera mechanically coupled to the frame, wherein the first camera is configured to output image data that satisfies a first field of view intensity change criterion related to the first camera, A second camera, A processor operably coupled to the first camera and Equipped with, The above method uses the processor, To determine whether the point of interest is within the first field of view, Using image data received from the first camera with respect to one or more portions of the first field of view, a tracking routine is performed. The tracking routine determines whether the quality standards are met, When the tracking routine does not meet the quality criteria, enable the second camera or modulate the frame rate of the second camera. Methods that include...
5. A portable device, wherein the portable device is A first camera, wherein the first camera is configured to output image data that satisfies a first field of view intensity change criterion relating to the first camera, A second camera, A processor operably coupled to the first camera and the second camera Equipped with, The first camera and the second camera are positioned to provide an overlapping view of the central field of view. The aforementioned processor, Using image data received from the first camera with respect to one or more portions of the first field of view, a tracking routine is performed. The tracking routine determines whether the quality standards are met, When the tracking routine does not meet the quality criteria, enable the second camera or modulate the frame rate of the second camera. A portable device configured to perform the following actions.
6. The portable device according to claim 1, or the portable device according to claim 3, or the portable device according to claim 5, wherein the strength change criterion includes an absolute or relative strength change criterion.
7. The portable device according to claim 1, or the portable device according to claim 3, or the portable device according to claim 5, wherein the first camera is configured to output the image data asynchronously.
8. The portable device according to claim 7, wherein the processor is further configured to track objects asynchronously.
9. The portable device according to claim 1, wherein executing the tracking routine includes limiting the acquisition of image data to points of interest within the world model.
10. The first camera is configured to limit image acquisition to one or more portions of the first camera's field of view. The aforementioned tracking routine is Identifying points of interest within the aforementioned world model, Determining one or more first portions of the field of view of the first camera corresponding to the point of interest, To provide the first camera with a command to limit image acquisition to one or more first portions of the field of view. The portable device according to claim 9, further comprising:
11. The aforementioned tracking routine is Based on the movement of the point of interest relative to the world model or the movement of the portable device relative to the point of interest, one or more second portions of the field of view of the first camera corresponding to the point of interest are estimated. To provide the first camera with a command to limit image acquisition to one or more second portions of the field of view. The portable device according to claim 10, further comprising:
12. The portable device further includes an inertial measuring unit, The portable device according to claim 9, wherein executing the tracking routine includes estimating the updated relative position of one of the plurality of points of interest based at least in part on the output of the inertial measurement unit.
13. The tracking routine includes repeatedly calculating the location of the point of interest within the world model, The portable device according to claim 9, wherein the aforementioned iterative calculation is performed with a transient resolution greater than 60 Hz.
14. The portable device according to claim 13, wherein the interval between the repeated calculations is a duration of 1 ms to 15 ms.
15. The portable device according to claim 1, further comprising a display device mechanically coupled to the processor.
16. The portable device according to claim 1, further comprising an IR emitter.
17. The portable device according to claim 16, wherein the processor is configured to selectively enable the IR emitter to enable head attitude tracking under low light conditions.
18. The method according to claim 2, or the method according to claim 4, wherein the intensity change criterion includes an absolute or relative intensity change criterion.
19. The method according to claim 2, or the method according to claim 4, wherein the first camera is configured to output the image data asynchronously.
20. The method according to claim 19, wherein the processor is further configured to asynchronously track head posture.
21. The method according to claim 2, wherein the processor is further configured to restrict image data acquisition to a plurality of points of interest within the world model by executing a tracking routine.
22. The first camera is configured to limit image acquisition to one or more portions of the first camera's field of view. The aforementioned tracking routine is Identifying points of interest within the aforementioned world model, Determining one or more first portions of the field of view of the first camera corresponding to the point of interest, To provide the first camera with a command to limit image acquisition to one or more first portions of the field of view. The method according to claim 21, including the method described in claim 21.
23. The aforementioned tracking routine is Based on the movement of the point of interest relative to the world model or the movement of the portable device relative to the point of interest, one or more second portions of the field of view of the first camera corresponding to the point of interest are estimated. To provide the first camera with a command to limit image acquisition to one or more second portions of the field of view. The method according to claim 22, further comprising:
24. The portable device further includes an inertial measuring unit, The method according to claim 21, wherein executing the tracking routine includes estimating the updated relative position of one of the plurality of points of interest based at least in part on the output of the inertial measurement unit.
25. The tracking routine includes repeatedly calculating the location of the point of interest within the world model, The method according to claim 21, wherein the iterative calculation is performed with a transient resolution greater than 60 Hz.
26. The method according to claim 25, wherein the interval between the repeated calculations is a duration of 1 ms to 15 ms.
27. The method according to claim 2, wherein the portable device comprises a display.
28. The method according to claim 2, wherein the portable device further comprises an IR emitter.
29. The method according to claim 28, wherein the processor is configured to selectively enable the IR emitter to enable the tracking routine to be executed under low light conditions.
Citation Information
Patent Citations
Solid-state imaging apparatus and camera system
JP2008172606A
Using dynamic vision sensors for motion detection in head mounted displays
US20180295337A1
Continuous time warp and binocular time warp for virtual and augmented reality display systems and methods
WO2018039586A1
Dynamic vision sensor architecture
WO2018122798A1
Laser scanner with real-time, online ego-motion estimation
WO2018140701A1