Eye tracking using images with different exposure times
The eye-tracking system addresses the challenges of VR, AR, and MR technologies by using different exposure times and frame rates to accurately predict gaze direction and enhance user experience through foveation rendering.
Patent Information
- Application Number
- JP2024014027
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-01-25
- Filing Date
- 2024-02-01
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2040-01-23
AI Technical Summary
Current VR, AR, and MR technologies face challenges in providing comfortable, natural, and rich presentations of virtual image elements among real-world image elements due to the complexity of human visual perception.
An eye-tracking system that uses an eye-tracking camera to capture images of the eye at different exposure times or frame rates, allowing for accurate gaze prediction and foveation rendering in wearable display systems.
The system enables accurate prediction of future gaze direction and supports foveation rendering, reducing rendering latency and improving user experience in VR, AR, and MR environments.
Smart Images

Figure 0007679503000001 
Figure 0007679503000002 
Figure 0007679503000003
Abstract
Description
Technical Field
[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority of U.S. Patent Application No. 62 / 797,072, filed on January 25, 2019, entitled "EYE - TRACKING USING IMAGES HAVING DIFFERENT EXPOSURE TIMES", which is incorporated herein by reference in its entirety.
[0002] This application also incorporates by reference in their entireties the following patent applications and publications: U.S. Patent Application No. 15 / 159,491, filed on May 19, 2016 and published as U.S. Patent Application Publication No. 2016 / 0344957 on November 24, 2016; U.S. Patent Application No. 15 / 717,747, filed on September 27, 2017 and published as U.S. Patent Application Publication No. 2018 / 0096503 on April 5, 2018; U.S. Patent Application No. 15 / 803,351, filed on November 3, 2017 and published as U.S. Patent Application Publication No. 2018 / 0131853 on May 10, 2018; U.S. Patent Application No. 15 / 841,043, filed on December 13, 2017 and published as U.S. Patent Application Publication No. 2018 / 0183986 on June 28, 2018; U.S. Patent Application No. 15 / 925,577, filed on March 19, 2018 and published as U.S. Patent Application Publication No. 2018 / 0278843 on September 27, 2018; U.S. Provisional Patent Application No. 62 / 660,180, filed on April 19, 2018; U.S. Patent Application No. 16 / 219,829, filed on December 13, 2018; U.S. Patent Application No. 16 / 219,847, filed on December 13, 2018; U.S. Patent Application No. 16 / 250,931, filed on January 17, 2019; and U.S. Patent Application No. 16 / 251,017, filed on January 17, 2019.
[0003] The present disclosure relates to display systems, virtual reality, and augmented reality imaging and visualization systems, and more particularly, to techniques for tracking a user's eyes in such systems.
Background Art
[0004] Modern computing and display technologies have facilitated the development of systems for so-called "virtual reality", "augmented reality", or "mixed reality" experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that appears or can be perceived as being real. Virtual reality, i.e., the "VR" scenario, typically involves the presentation of digital or virtual image information without transparency to other actual real-world visual inputs. Augmented reality, i.e., the "AR" scenario, typically involves the presentation of digital or virtual image information as an augmentation to the visualization of the actual world around the user. Mixed reality or "MR" relates to the fusion of the real and virtual worlds to create a new environment in which physical and virtual objects coexist and interact in real time. In conclusion, the human visual perception system is very complex, and it is difficult to produce VR, AR, or MR technologies that facilitate a comfortable, natural, and rich presentation of virtual image elements among other virtual or real-world image elements. The systems and methods disclosed herein address various challenges associated with VR, AR, and MR technologies.
Summary of the Invention
Means for Solving the Problems
[0005] The eye tracking system can include an eye tracking camera configured to acquire images of the eye at different exposure times or different frame rates. For example, an image of the eye taken at a longer exposure time may show iris or pupil features, and an image of the eye taken at a shorter exposure time (sometimes also referred to as a flash image) may show the peak of the flash reflected from the cornea. The flash image with shorter exposure is taken at a higher frame rate (HFR) than the image with longer exposure and can provide accurate gaze prediction. The flash image with shorter exposure can be analyzed to provide the flash location up to sub-pixel accuracy. The image with longer exposure can be analyzed for the pupil center or the center of rotation. The eye tracking system can predict the future gaze direction and can be used for foveation rendering by a wearable display system, such as an AR, VR, or MR wearable display system.
[0006] In various embodiments, the exposure time of the image with longer exposure can be in the range of 200 μs to 1,200 μs, for example, about 700 μs. The image with longer exposure can be taken at a frame rate in the range of 10 frames per second (fps) to 60 fps (e.g., 30 fps), 30 fps to 60 fps, or some other range. The exposure time of the flash image with shorter exposure can be in the range of 5 μs to 100 μs, for example, less than about 40 μs. The ratio of the exposure time of the image with longer exposure to the exposure time of the flash image can be in the range of 5 to 50, 10 to 20, or some other range. The flash image can be taken at a frame rate in the range of 50 fps to 1,000 fps (e.g., 120 fps), 200 fps to 400 fps, or some other range in various embodiments. The ratio of the frame rate of the flash image to the frame rate of the image with longer exposure can be in the range of 1 to 100, 1 to 50, 2 to 20, 3 to 10, or some other ratio.
[0007] In some embodiments, images with shorter exposure are analyzed by a first processor (which may be disposed within or on a head-mounted component of the wearable display system), and images with longer exposure are analyzed by a second processor (e.g., which may be disposed within or on a non-head-mounted component of the wearable display system such as a belt pack or on the head-mounted component). In some embodiments, the first processor includes a buffer in which a portion of the image with shorter exposure is temporarily stored to determine the flash location.
[0008] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, the drawings, and the claims. Neither this summary nor any of the following forms for carrying out the invention purports to define or limit the scope of the subject matter of the invention.
[0009] The present invention provides, for example, the following. (Item 1) A wearable display system, A head-mounted display configured to present virtual content by outputting light to the eyes of a wearer of the head-mounted display, At least one light source configured to direct light toward the eyes of the wearer, At least one eye-tracking camera, wherein the at least one eye-tracking camera A first image of the eyes of the wearer, the first image of the eyes of the wearer being captured at a first frame rate and a first exposure time, A second image of the wearer's eye, the second image of the wearer's eye being captured at a second frame rate that is greater than or equal to the first frame rate and at a second exposure time that is less than the first exposure time, the second image of the wearer's eye At least one eye-tracking camera configured to capture At least one hardware processor communicatively coupled to the head-mounted display and the at least one eye-tracking camera, the at least one hardware processor being Analyzing the first image to determine the pupil center of the eye; Analyzing the second image to determine the position of the reflection of the light source from the eye; Determining the line-of-sight direction of the eye from the pupil center and the position of the reflection; Estimating a future line-of-sight direction of the eye at a future line-of-sight time from the line-of-sight direction and previous line-of-sight direction data; Causing the head-mounted display to present the virtual content at the future line-of-sight time, at least in part based on the future line-of-sight direction; At least one hardware processor programmed to perform A wearable display system comprising (Item 2) The wearable display system of item 1, wherein the at least one light source comprises at least one infrared light source. (Item 3) The wearable display system of item 1 or item 2, wherein the at least one hardware processor is programmed to analyze the first image or the second image to determine a center of rotation or a line-of-sight center of the wearer's eye. (Item 4) To analyze the second image, the at least one hardware processor is programmed to apply a threshold to the second image, identify non-maximum values within the second image, or suppress or remove non-maximum values within the second image, the wearable display system according to any one of items 1-3. (Item 5) To analyze the second image, the at least one hardware processor is programmed to identify a search area for a reflection of the light source in the current image from the second image, at least in part, based on a position of a reflection of the light source in a previous image from the second image, the wearable display system according to any one of items 1-4. (Item 6) To analyze the second image, the at least one hardware processor determines a common velocity of a plurality of reflections of the at least one light source; determines whether a velocity of a reflection of the at least one light source differs from the common velocity by more than a threshold amount; and is programmed to perform the above, the wearable display system according to any one of items 1-5. (Item 7) To analyze the second image, the at least one hardware processor is programmed to determine whether a reflection of the light source is from a non-spherical portion of the cornea of the eye, the wearable display system according to any one of items 1-6. (Item 8) To analyze the second image, the at least one hardware processor is programmed to identify the presence of at least a partial occlusion of a reflection of the light source, the wearable display system according to any one of items 1-7. (Item 9) The at least one hardware processor is programmed to determine a position of the estimated pupil center, at least in part, based on a position of reflection of the light source and a flash-pupil relationship, for the wearable display system according to any one of items 1-8. (Item 10) The flash-pupil relationship includes a linear relationship between a flash position and a pupil center position, for the wearable display system according to item 9. (Item 11) The first exposure time is in a range of 200 μs to 1,200 μs, for the wearable display system according to any one of items 1-10. (Item 12) The first frame rate is in a range of 10 frames per second to 60 frames per second, for the wearable display system according to any one of items 1-11. (Item 13) The second exposure time is in a range of 5 μs to 100 μs, for the wearable display system according to any one of items 1-12. (Item 14) The second frame rate is in a range of 100 frames per second to 1,000 frames per second, for the wearable display system according to any one of items 1-13. (Item 15) A ratio of the first exposure time to the second exposure time is in a range of 5 to 50, for the wearable display system according to any one of items 1-14. (Item 16) A ratio of the second frame rate to the first frame rate is in a range of 1 to 100, for the wearable display system according to any one of items 1-15. (Item 17) The future gaze time is in a range of 5 ms to 100 ms, for the wearable display system according to any one of items 1-16. (Item 18) The at least one hardware processor is A first hardware processor, wherein the first hardware processor is disposed on a non-head-mounted component of the wearable display system, and the first hardware processor; A second hardware processor, wherein the second hardware processor is disposed within or on the head-mounted display, and the second hardware processor Comprising; The first hardware processor is utilized to analyze the first image, The second hardware processor is utilized to analyze the second image, The wearable display system according to any one of items 1-17. (Item 19) The second hardware processor includes or is associated with a memory buffer configured to store at least a portion of each of the second images, and the second hardware processor is programmed to delete at least a portion of each of the second images after determining the position of the reflection of the light source from the eye. The wearable display system according to item 18. (Item 20) The at least one hardware processor is programmed not to combine the first image and the second image. The wearable display system according to any one of items 1-19. (Item 21) A method for eye tracking, the method comprising: Capturing a first image of an eye at a first frame rate and a first exposure time via an eye tracking camera; Capturing a second image of the eye at a second frame rate greater than the first frame rate and a second exposure time less than the first exposure time via the eye tracking camera; Determining at least the pupil center of the eye from the first image; At least, determining a position of reflection of a light source from the eye from the second image; determining a line-of-sight direction of the eye from the pupil center and the position of the reflection; A method comprising: (Item 22) The method according to Item 21, further comprising rendering virtual content on a display at least partially based on the line-of-sight direction. (Item 23) The method according to Item 21 or Item 22, further comprising estimating a future line-of-sight direction at a future line-of-sight time based at least partially on the line-of-sight direction and previous line-of-sight direction data. (Item 24) The method according to Item 23, further comprising rendering virtual content on a display at least partially based on the future line-of-sight direction. (Item 25) The method according to any one of Items 21-24, wherein the first exposure time is in the range of 200 μs to 1,200 μs. (Item 26) The method according to any one of Items 21-25, wherein the first frame rate is in the range of 10 frames / second to 60 frames / second. (Item 27) The method according to any one of Items 21-26, wherein the second exposure time is in the range of 5 μs to 100 μs. (Item 28) The method according to any one of Items 21-27, wherein the second frame rate is in the range of 100 frames / second to 1,000 frames / second. (Item 29) The method according to any one of Items 21-28, wherein a ratio of the first exposure time to the second exposure time is in the range of 5 to 50. (Item 30) The method according to any one of Items 21-29, wherein a ratio of the second frame rate to the first frame rate is in the range of 1 to 100. (Item 31) A wearable display system, A head-mounted display, wherein the head-mounted display is configured to present virtual content by outputting light to the eyes of the wearer of the head-mounted display, a head-mounted display, A light source, wherein the light source is configured to direct light toward the eyes of the wearer, a light source, An eye-tracking camera, wherein the eye-tracking camera is configured to capture an image of the eyes of the wearer, and the eye-tracking camera alternately, At a first exposure time, capturing a first image; At a second exposure time shorter than the first exposure time, capturing a second image And an eye-tracking camera configured to perform the above, A plurality of electronic hardware components, at least one of which comprises a hardware processor communicatively coupled to the head-mounted display, the eye-tracking camera, and at least one other electronic hardware component within the plurality of electronic hardware components, the hardware processor being configured to: Receive each first image of the eyes of the wearer captured at the first exposure time by the eye-tracking camera and relay it to the at least one other electronic hardware component; Receive the pixels of each second image of the eyes of the wearer captured at the second exposure time by the eye-tracking camera and store them in a buffer; Analyze the pixels stored in the buffer and identify the locations where the reflection of the light source is present in each second image of the eyes of the wearer captured at the second exposure time by the eye-tracking camera; Transmit location data indicating the locations to the at least one other electronic hardware component And a plurality of electronic hardware components programmed to perform the above, Comprising, The at least one other electronic hardware component is Analyze each first image of the wearer's eyes captured by the eye-tracking camera at the first exposure time, and identify the location of the pupil center of the eyes; Determine the line-of-sight direction of the wearer's eyes from the location of the pupil center and the location data received from the hardware processor; A wearable display system configured to perform the above. (Item 32) The wearable display system according to item 31, wherein the first exposure time is within a range of 200 μs to 1,200 μs. (Item 33) The wearable display system according to item 31 or item 32, wherein the second exposure time is within a range of 5 μs to 100 μs. (Item 34) The wearable display system according to any one of items 31 - 33, wherein the eye-tracking camera is configured to capture the first image at a first frame rate within a range of 10 frames per second to 60 frames per second. (Item 35) The wearable display system according to any one of items 31 - 34, wherein the eye-tracking camera is configured to capture the second image at a second frame rate within a range of 100 frames per second to 1,000 frames per second. (Item 36) The wearable display system according to any one of items 31 - 35, wherein the ratio of the first exposure time to the second exposure time is within a range of 5 to 50. (Item 37) The wearable display system according to any one of items 31 - 36, wherein the at least one other electronic hardware component is disposed within or on the non-head-mounted component of the wearable display system. (Item 38) The wearable display system according to item 37, wherein the non-head-mounted component includes a belt pack. (Item 39) The wearable display system according to any one of Items 31-38, wherein the pixels of each second image of the eye include fewer pixels than all of the pixels of each second image. (Item 40) The wearable display system according to any one of Items 31-39, wherein the pixels include an array of n×m pixels, and n and m are each integers within the range of 1 to 20. (Item 41) The plurality of electronic hardware components further estimating a future line-of-sight direction of the eye at a future line-of-sight time from the line-of-sight direction and the previous line-of-sight direction data; and causing the head-mounted display to present virtual content at the future line-of-sight time, at least partially based on the future line-of-sight direction. The wearable display system according to any one of Items 31-40, configured to perform the above. (Item 42) The wearable display system according to any one of Items 31-41, wherein the hardware processor is programmed to apply a threshold to the pixels stored in the buffer, identify non-maximum values within the pixels stored in the buffer, or suppress or remove non-maximum values within the pixels stored in the buffer. (Item 43) The hardware processor determines a common speed of reflection of the light source; and determines whether the speed of reflection of the light source differs from the common speed by more than a threshold amount. The wearable display system according to any one of Items 31-42, programmed to perform the above. (Item 44) The wearable display system according to any one of items 31 - 43, wherein the hardware processor is programmed to determine whether the location of reflection of the light source is from an aspherical portion of the cornea of the eye. (Item 45) The wearable display system according to any one of items 31 - 44, wherein the hardware processor is programmed to identify the presence of at least a partial occlusion of the reflection of the light source. (Item 46) The wearable display system according to any one of items 31 - 45, wherein the at least one other electronic hardware component is programmed to identify the location of the pupil center based at least in part on the location of reflection of the light source and the flash - pupil relationship. (Item 47) The wearable display system according to any one of items 18 - 20, wherein the second hardware processor is configured to identify a flash and provide flash position information regarding the flash to the first hardware processor. (Item 48) The wearable display system according to item 47, wherein the first hardware processor is configured to determine the line of sight from the flash position. (Item 49) The wearable display system according to any one of items 18 - 20, wherein the second hardware processor is configured to identify a flash candidate and provide position information regarding the flash candidate to the first hardware processor. (Item 50) The wearable display system according to item 49, wherein the first hardware processor is configured to identify a subset of the flash candidates and use the subset of the flash candidates to perform one or more operations. (Item 51) The wearable display system according to item 50, wherein the first hardware processor is configured to determine the line of sight based on the flash candidates. (Item 52) The hardware processor is programmed to identify a flash candidate and an associated location of the flash candidate, and transmit location data indicating the location of the flash candidate to the at least one other electronic hardware component, according to any one of Items 31 - 46 of the wearable display system described above. (Item 53) The at least one other electronic hardware component is configured to identify a subset of the flash candidates and perform one or more operations using the subset of the flash candidates, according to Item 52 of the wearable display system described above. (Item 54) The at least one other electronic hardware component is configured to determine a line of sight direction of the wearer's eye from the subset of the flash candidates, according to Item 53 of the wearable display system described above.
Brief Description of the Drawings
[0010]
Figure 1
[0011]
Figure 2
[0012]
Figure 3
[0013]
Figure 4
[0014]
Figure 5
[0015]
Figure 6
[0016]
Figure 7
[0017]
Figure 8
[0018]
Figure 9A
Figure 9B
Figure 9C
[0019]
Figure 10A
[0020]
Figure 10B
[0021]
Figure 11
[0022]
Figure 12
[0023]
Figure 13A
[0024]
Figure 13B
[0025]
Figure 14A
Figure 14B
[0026]
Figure 15
[0027]
Figure 16
[0028]
Figure 17
[0029]
Figure 18A
Figure 18B
Figure 18C
Figure 18D
[0030]
Figure 19
DETAILED DESCRIPTION OF THE INVENTION
[0031] Throughout the drawings, reference numerals may be reused to indicate correspondence between referenced elements. Unless otherwise indicated, the drawings are schematic and are not necessarily drawn to scale. The drawings are provided to illustrate the exemplary embodiments described herein and are not intended to limit the scope of the disclosure. (Overview)
[0032] For example, a wearable display system, such as an AR, MR, or VR display system, can track a user's eyes in order to project virtual content towards the location where the user is looking. The eye tracking system can include an inward-facing eye tracking camera and a light source (e.g., an infrared light emitting diode) that provides a reflection (referred to as a flash) from the user's cornea. A processor can analyze an image of the user's eyes captured by the eye tracking camera, obtain the positions of the flash and other eye features (e.g., the pupil or iris), and determine the eye line of sight from the flash and the eye features.
[0033] Not only the flash, but also an eye image that is sufficient to indicate eye features can be captured with a relatively long exposure time (e.g., several hundred to a thousand μs). However, the flash can be saturated in such an image with a longer exposure, which can make it difficult to accurately identify the position of the flash center. For example, the uncertainty at the flash position can be 10 to 20 pixels, which can introduce a corresponding error in the line of sight direction of about 20 to 50 minutes.
[0034] Therefore, various embodiments of the eye-tracking systems described herein acquire images of the eye at different exposure times or different frame rates. For example, an image of a longer exposure of the eye taken at a longer exposure time may show iris or pupil features, and an image of a shorter exposure may show the peak of the flash reflected from the cornea. Images of shorter exposure are sometimes also referred to herein as flash images because they can be used to identify the coordinate position of the flash within the image. In some implementations, flash images of shorter exposure may be taken at a high frame rate (HFR) (e.g., a frame rate higher than the frame rate for images of longer exposure) for accurate gaze prediction. Flash images of shorter exposure can be analyzed to provide the location of the flash to sub-pixel accuracy, which leads to an accurate prediction of the gaze direction (e.g., within a few minutes or better). Images of longer exposure can be analyzed for the pupil center or the center of rotation.
[0035] In some implementations, at least a portion of the flash image is temporarily stored in a buffer, and that portion of the flash image is analyzed to identify the location of one or more flashes that may be located within that portion. For example, the portion may include a relatively small number of pixels, rows, or columns of the flash image. In some cases, the portion may include an n×m portion of the flash image, where n and m are integers that may be in the range of about 1 to 20. After the location of the flash is identified, the buffer may be cleared. Additional portions of the flash image may then be stored in the buffer for analysis until either the entire flash image is processed or all (generally four) flashes are identified. The flash location (e.g., Cartesian coordinates) may be used for subsequent actions in the eye-tracking process, and after the flash location is stored or communicated to a suitable processor, the flash image may be deleted from memory (buffer memory or other volatile or non-volatile storage device). Such a buffer advantageously enables high-speed processing of the flash image to identify the flash location or may reduce the memory requirements of the eye-tracking process because the flash image can be deleted after use.
[0036] Thus, in some embodiments, shorter exposure images cannot be combined with longer exposure images to obtain high-dynamic range (HDR) images that are used for eye tracking. Instead, in some such embodiments, shorter exposure images and longer exposure images are processed separately and used to determine different information. For example, a shorter exposure image may be used to identify a flash position (e.g., coordinates of the flash center) or an eye gaze direction. After the flash position is determined, the shorter exposure image may be deleted from memory (e.g., a buffer). Longer exposure images may be used to determine a pupil center or a center of rotation, extract iris features for biometric security purposes, determine an eyelid shape or occlusion of the iris or pupil by the eyelid, measure a pupil size, determine rendering camera parameters, and so on. In some implementations, different processors perform the processing of shorter and longer exposure images. For example, a processor within a head-mounted display may process shorter exposure images, and a processor within a non-head-mounted unit (e.g., a belt pack) may process longer exposure images.
[0037] Accordingly, various embodiments of the multiple exposure time techniques described herein can obtain the advantages of HDR photometry provided by both shorter and longer exposure images collectively, without combining, synthesizing, merging, or otherwise processing such short and long exposure images together (e.g., as an HDR image). Thus, various embodiments of a multiple exposure eye tracking system do not use such short and long exposure images to generate or otherwise obtain an HDR image.
[0038] In various embodiments, the exposure time of the longer exposure image may be in the range of 200 μs to 1,200 μs, for example, about 700 μs. The longer exposure image can be captured at a frame rate in the range of 10 frames per second (fps) to 60 fps (e.g., 30 fps), 30 fps to 60 fps, or some other range. The exposure time of the flash image may be in the range of 5 μs to 100 μs, for example, less than about 40 μs. The ratio of the exposure time for the longer exposure image to the exposure time for the shorter exposure flash image can be in the range of 5 to 50, 10 to 20, or some other range. The flash image can be captured at a frame rate in the range of 50 fps to 1,000 fps (e.g., 120 fps), 200 fps to 400 fps, or some other range in various embodiments. The ratio of the frame rate for the flash image to the frame rate for the longer exposure image can be in the range of 1 to 100, 1 to 50, 2 to 20, 3 to 10, or some other ratio.
[0039] Some wearable systems may utilize a foveation rendering technique where virtual content can be primarily rendered in the direction the user is looking. Embodiments of the eye-tracking system can accurately estimate the future line-of-sight direction (e.g., about 50 ms into the future), which can be used by the rendering system to prepare virtual content for future rendering, which can advantageously reduce rendering latency and improve the user experience. (Example of a 3D display of a wearable system)
[0040] A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present a 2D or 3D virtual image to a user. The image may be a still image, a frame of a video, or a video in a combination or equivalent. At least a portion of the wearable system can be implemented on a wearable device that can present a VR, AR, or MR environment, alone or in combination, for user interaction. The wearable device can be used synonymously with an AR device (ARD). Further, for the purposes of this disclosure, the term “AR” is used synonymously with the term “MR”.
[0041] FIG. 1 depicts an illustration of a composite reality scenario with a virtual reality object and a physical object as viewed by a person. In FIG. 1, an MR scene 100 is depicted, and to a user of MR technology, a real-world park-like setting 110 is visible, featuring people, trees, buildings in the background, and a concrete platform 120. In addition to these items, a user of MR technology also “sees” a robotic image 130 standing on the real-world platform 120 and a flying, cartoon-like avatar character 140 that appears as an anthropomorphic bumblebee, although these elements do not exist in the real world.
[0042] It may be desirable for a 3D display to generate a perspective adjustment response corresponding to the virtual depth for each point within the field of view of the display in order to generate a true sense of depth, more specifically, a simulated sense of surface depth. If the perspective adjustment response for a display point does not correspond to the virtual depth of that point such that it is determined by both binocular depth cues of convergence and stereopsis, the human eye experiences a vergence conflict, resulting in unstable imaging, harmful eye strain, headaches, and in the absence of vergence information, a nearly complete lack of surface depth.
[0043] VR, AR, and MR experiences can be provided by a display system having a display that provides an image corresponding to a plurality of depth planes to a viewer. The images may differ for each depth plane (e.g., providing a somewhat different presentation of a scene or object), and are separately focused by the viewer's eyes, thereby based on the eye accommodation required to focus on different image features regarding scenes located on different depth planes, or based on observing different image features on different depth planes that are out of focus, which can help provide depth cues to the user. As discussed anywhere in this specification, such depth cues provide a perception of reliable depth.
[0044] Figure 2 illustrates an embodiment of a wearable system 200, which can be configured to provide an AR / VR / MR scene. The wearable system 200 may also be referred to as the AR system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems to support the functions of the display 220. The display 220 may be coupled to a frame 230, which can be worn by a user, wearer, or viewer 210. The display 220 can be positioned in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can include a head-mounted display (HMD) that is worn on the user's head.
[0045] In some embodiments, speaker 240 is coupled to frame 230 and positioned adjacent to the user's external auditory canal (in some embodiments, another speaker, not shown, is positioned adjacent to the user's other external auditory canal to provide stereo / formable sound control). Display 220 can include an audio sensor (e.g., a microphone) 232 to detect an audio stream from the environment and capture ambient sound. In some embodiments, one or more other audio sensors, not shown, are positioned to provide stereo sound reception. Stereo sound reception can be used to determine the location of the sound source. Wearable system 200 can perform speech or utterance recognition on the audio stream.
[0046] Wearable system 200 can include an outward-facing imaging system 464 (shown in FIG. 4) that observes the world within the environment around the user. Wearable system 200 can also include an inward-facing imaging system 462 (shown in FIG. 4) that can track the user's eye movements. The inward-facing imaging system can track either the movement of one eye or the movement of both eyes. Inward-facing imaging system 462 may be attached to frame 230 and may communicate electrically with a processing module 260 or 270 that can process the image information obtained by the inward-facing imaging system and determine, for example, the pupil diameter or orientation, eye movement, or eye pose of the user's 210 eyes. Inward-facing imaging system 462 may include one or more cameras. For example, at least one camera may be used to image each eye. The images obtained by the cameras may be used to determine the pupil size or eye pose separately for each eye, thereby enabling the presentation of image information to each eye to be dynamically adjusted with respect to that eye.
[0047] As an example, the wearable system 200 can obtain an image of the user's posture using an outward-facing imaging system 464 or an inward-facing imaging system 462. The image may be a still image, a video frame, or a video.
[0048] The display 220 can be operably coupled to a local data processing module 260 that can be mounted in various configurations, such as fixed to the frame 230 by a wired conductor or wireless connection, fixed to a helmet or hat worn by the user, built into headphones, or removably attached to the user 210 in another manner (e.g., in a backpack configuration, in a belt attachment configuration). (250)
[0049] The local processing and data module 260 may comprise a digital memory such as a hardware processor and a non-volatile memory (e.g., flash memory), both of which may be utilized to assist in the processing, caching, and storing of data. The data may include a) data captured from sensors such as an image capture device (e.g., a camera within an inward-facing or outward-facing imaging system), an audio sensor (e.g., a microphone), an inertial measurement unit (IMU), an accelerometer, a compass, a global positioning system (GPS) unit, a wireless device, or a gyroscope (e.g., operably coupled to the frame 230 or otherwise attachable to the user 210), or b) possibly data obtained and / or processed using the remote processing module 270 or the remote data repository 280 for passage to the display 220 after processing or reading. The local processing and data module 260 may be operably coupled to the remote processing module 270 or the remote data repository 280 via a communication link 262 or 264, such as a wired or wireless communication link, such that these remote modules are available as resources to the local processing and data module 260. Additionally, the remote processing module 280 and the remote data repository 280 may be operably coupled to each other.
[0050] In some embodiments, the remote processing module 270 may comprise one or more processors configured to analyze and process data or image information. In some embodiments, the remote data repository 280 may comprise a digital data storage facility, which may be available through other networking configurations in an Internet or "cloud" resource configuration. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, enabling full autonomy from the remote modules. (Exemplary components of a wearable system)
[0051] FIG. 3 schematically illustrates exemplary components of a wearable system. FIG. 3 shows a wearable system 200, which can include a display 220 and a frame 230. Stretch view 202 schematically illustrates various components of the wearable system 200. In one implementation, one or more of the components illustrated in FIG. 3 can be part of the display 220. The various components can collect various data (e.g., auditory or visual data, etc.) associated with the user or the user's environment of the wearable system 200, either alone or in combination. It should be understood that other embodiments may have additional or fewer components depending on the use for which the wearable system is employed. Note that FIG. 3 provides some of the various components and a basic concept of the types of data that can be collected, analyzed, and stored through the wearable system.
[0052] FIG. 3 shows an exemplary wearable system 200, which can include a display 220. The display 220 can include a display lens 226 that can be mounted on a housing or frame 230 corresponding to the user's head or frame 230. The display lens 226 may include one or more transparent mirrors positioned in front of the user's eyes 302, 304 by the housing 230, bounce the projected light 338 into the eyes 302, 304, and be configured to allow transmission of at least some light from the local environment while facilitating beam shaping. The wavefront of the projected light beam 338 may be bent or focused to match the desired focal length of the projected light. As shown, two wide field of view machine vision cameras 316 (also referred to as world cameras) are coupled to the housing 230 and can image the environment around the user. These cameras 316 can be dual capture visible light / non-visible (e.g., infrared) light cameras. The cameras 316 may be part of an outward facing imaging system 464 shown in FIG. 4. Images obtained by the world cameras 316 can be processed by a pose processor 336. For example, the pose processor 336 can implement one or more object recognition devices 708 (e.g., shown in FIG. 7) and identify the pose of the user or another person within the user's environment or identify physical objects within the user's environment.
[0053] Continuing to refer to FIG. 3, shown is a pair of scanning laser shaped wavefront (e.g., for depth) light projection modules with a display mirror and optical system configured to project light 338 into eyes 302, 304. The depicted figure also shows two small infrared cameras 324 paired with a light source 326 (such as a light emitting diode “LED”) configured to track the user's eyes 302, 304 and support rendering and user input. The light source 326 can emit light within the infrared (IR) portion of the optical spectrum so that the eyes 302, 304 will not perceive a light source that would shine into the user's eyes in a way that the user would be sensitive to IR light and find uncomfortable. The cameras 324 may be part of an inward facing imaging system 462 as shown in FIG. 4. The wearable system 200 can further feature a sensor assembly 339, which has X, Y, and Z axis accelerometer capabilities and a magnetic compass and X, Y, and Z axis gyroscope capabilities and preferably can provide data at a relatively high frequency such as 200 Hz. The sensor assembly 339 may be part of an IMU as described with reference to FIG. 2A. The depicted system 200 can also include a head pose processor 336 such as an ASIC (application specific integrated circuit), FPGA (field programmable gate array), or ARM processor (advanced reduced instruction set machine), which may be configured to calculate from wide field of view image information output from a capture device 316 a real-time or near real-time user head pose. The head pose processor 336 can be a hardware processor and can be implemented as part of the local processing and data module 260 shown in FIG. 2.
[0054] The wearable system can also include one or more depth sensors 234. The depth sensors 234 can be configured to measure the distance between objects in the environment and the wearable device. The depth sensors 234 may include a laser scanner (e.g., LIDAR), an ultrasonic depth sensor, or a depth-sensing camera. In one implementation where the camera 316 has depth-sensing capabilities, the camera 316 can also be regarded as a depth sensor 234.
[0055] Also shown is a processor 332 that is configured to perform digital or analog processing and derive the posture from gyroscope, compass, or accelerometer data from the sensor assembly 339. The processor 332 may be part of the local processing and data module 260 shown in FIG. 2. The wearable system 200 can also include a positioning system such as, for example, a GPS 337 (Global Positioning System) as shown in FIG. 3, which can assist with posture and positioning analysis. Additionally, the GPS may further provide remote-based (e.g., cloud-based) information about the user's environment. This information may be used to recognize objects or information in the user's environment.
[0056] The wearable system may combine data obtained by the GPS 337 and a remote computing system (e.g., remote processing module 270, another user's ARD, etc.), which can provide more information about the user's environment. As an example, the wearable system can determine the user's location based on GPS data and read out a world map that includes virtual objects associated with the user's location (e.g., by communicating with the remote processing module 270). As another example, the wearable system 200 can use the world camera 316 (which may be part of the outward-facing imaging system 464 shown in FIG. 4) to monitor the environment. Based on the images obtained by the world camera 316, the wearable system 200 can detect objects in the environment. The wearable system can further interpret the character using the data obtained by the GPS 337.
[0057] The wearable system 200 may also include a rendering engine 334, which can be configured to provide rendering information local to the user for the view of users worldwide and to facilitate the operation of the scanner and the imaging into the user's eyes. The rendering engine 334 may be implemented by a hardware processor (e.g., a central processing unit or a graphics processing unit, etc.). In some embodiments, the rendering engine is part of the local processing and data module 260. The rendering engine 334 may include a light field rendering controller 618, which will be described with reference to FIGS. 6 and 7. The rendering engine 334 can be communicatively coupled to other components of the wearable system 200 (e.g., via a wired or wireless link). For example, the rendering engine 334 can be coupled to the eye camera 324 via a communication link 274 and to the projection subsystem 318 (which can project light into the user's eyes 302, 304 via a scanning laser array in a manner similar to a retinal scanning display) via a communication link 272. The rendering engine 334 can also communicate with other processing units, such as the sensor attitude processor 332 and the image attitude processor 336, via links 276 and 294, respectively.
[0058] The camera 324 (e.g., a small infrared camera) may be utilized to track eye poses and support rendering and user input. Some exemplary eye poses may include where the user is looking or the depth at which the user is focused (which may be estimated using the vergence of the eyes). The camera 324 and the infrared light source 326 can be used to provide data for the multiple exposure time eye tracking technique described herein. The GPS 337, gyroscope, compass, and accelerometer 339 may be utilized to provide gross or high-speed pose estimation. One or more of the cameras 316 can obtain images and poses, which, in conjunction with data from associated cloud computing resources, may be used to map the local environment and share the user view with others.
[0059] The exemplary components depicted in FIG. 3 are for illustrative purposes only. Multiple sensors and other functional modules are shown together for ease of illustration and explanation. Some embodiments may include only one or a subset of these sensors or modules. Further, the locations of these components are not limited to the positions depicted in FIG. 3. Some components may be mounted or housed within other components such as belt-mounted components, handheld components, or helmet components. As an example, the image pose processor 336, the sensor pose processor 332, and the rendering engine 334 may be positioned within a belt pack and configured to communicate with other components of the wearable system via wireless communication such as ultra-wideband, Wi-Fi, Bluetooth®, or via wired communication. The depicted housing 230 is preferably head-mountable and wearable by the user. However, some components of the wearable system 200 may be worn on other parts of the user's body. For example, the speaker 240 may be inserted into the user's ear to provide sound to the user.
[0060] Regarding the projection of light 338 into the eyes 302, 304 of a user, in some embodiments, the camera 324 may generally be utilized to measure the location where the center of the user's eyes geometrically converges, which generally coincides with the position of the focus of the eyes or the "depth of focus". The three-dimensional surface of all the points where the eyes converge may be referred to as the "horopter". The focal distance may take a finite number of depths or may vary infinitely. The light projected from the convergence / divergence movement distance appears to be focused on the target eyes 302, 304, while the light in front of or behind the convergence / divergence movement distance is blurred. Examples of the wearable system and other display systems of the present disclosure are also described in U.S. Patent Publication No. 2016 / 0270656, which is incorporated herein by reference in its entirety.
[0061] The human visual system is complex, and it is difficult to provide a realistic perception of depth. The viewer of an object can perceive the object in three dimensions due to the combination of convergence / divergence movement and accommodation. The convergence / divergence movement of the two eyes relative to each other (e.g., the rotational movement of the pupils towards or away from each other to converge the lines of sight of the eyes and fixate on an object) is closely associated with the focusing (or "accommodation") of the eye lenses. Under normal conditions, the change in the focus of the eye lenses or the accommodation of the eyes to change the focus from one object to another object at a different distance will automatically cause a corresponding change in convergence / divergence movement at the same distance under the relationship known as the "accommodation-convergence / divergence reflex". Similarly, a change in convergence / divergence movement will induce a corresponding change in accommodation under normal conditions. A display system that provides a better match between accommodation and convergence / divergence movement can form a more realistic and comfortable simulation of a three-dimensional image.
[0062] Furthermore, spatially coherent light with a beam diameter of less than about 0.7 millimeters can be correctly resolved by the human eye regardless of where the eye is focused. Thus, in order to create an illusion of appropriate depth of focus, the eye's convergence / divergence movements may be tracked using camera 324, and rendering engine 334 and projection subsystem 318 may be utilized to focus and render all objects on or near the single viewing trajectory and to render all other objects with a variable degree of defocus (e.g., using intentionally created blur). Preferably, system 220 renders to the user at a frame rate of about 60 frames per second or greater. As described above, preferably, camera 324 may be utilized for eye tracking, and the software may be configured to take into account not only the convergence / divergence movement geometry but also a focus location queue for serving as user input. Preferably, such a display system is configured with brightness and contrast suitable for daytime or nighttime use.
[0063] In some embodiments, the display system preferably has a latency of less than about 20 milliseconds, an angular alignment of less than about 0.1 degrees, and a resolution of about 1 arcminute for visual object alignment, which, while not limited by theory, is thought to be near the limits of the human eye. Display system 220 may be integrated with a location system, which may involve a GPS element, optical tracking, a compass, an accelerometer, or other data sources and may assist in position and orientation determination. The location information may be utilized to facilitate accurate rendering within the user's view of the relevant world (e.g., such information would facilitate glasses knowing their location relative to the real world).
[0064] In some embodiments, the wearable system 200 is configured to display one or more virtual images based on the user's eye accommodation. Different from the conventional 3D display approach that forces the user to focus on the location where the image is projected, in some embodiments, the wearable system automatically varies the focus of the projected virtual content and is configured to enable more comfortable viewing of the one or more images presented to the user. For example, if the user's eye has a current focus of 1 m, the image may be projected to match the user's focus. If the user shifts the focus to 3 m, the image is projected to match the new focus. Thus, rather than forcing a predetermined focus on the user, the wearable system 200 of some embodiments allows the user's eyes to function in a more natural manner.
[0065] Such a wearable system 200 can eliminate or reduce the incidence of eye strain, headaches, and other physiological symptoms typically observed with virtual reality devices. To achieve this, various embodiments of the wearable system 200 are configured to project virtual images at variable focal distances through one or more variable focus elements (VFEs). In one or more embodiments, 3D perception may be achieved through a multi-plane focus system that projects the image onto a fixed focal plane from the user. Other embodiments employ variable plane focus, and the focal plane is reciprocally moved in the z-direction to match the current state of the user's focus.
[0066] In both multi-plane focus systems and variable plane focus systems, the wearable system 200 may employ eye tracking, determine the convergence / divergence movement of the user's eyes, determine the user's current focus, and project a virtual image onto the determined focus. In other embodiments, the wearable system 200 includes a light modulator that variably projects a light beam of variable focus in a raster pattern across the retina through a fiber scanner or other light generating source. Thus, the display capabilities of the wearable system 200 that project an image at a variable focal distance not only facilitate depth adjustment for the user to visually perceive an object in 3D, but may also be used to compensate for the user's eye abnormalities, as further described in U.S. Patent Publication No. 2016 / 0270656, which is incorporated herein by reference in its entirety. In some other embodiments, the spatial light modulator may project an image to the user through various optical components. For example, as further described below, the spatial light modulator may project an image onto one or more waveguides, which then transmit the image to the user. (Waveguide stack assembly)
[0067] FIG. 4 illustrates an example of a waveguide stack for outputting image information to a user. The wearable system 400 includes a stack or stacked waveguide assembly 480 of waveguides that can be utilized to provide a three-dimensional perception to the eye / brain using a plurality of waveguides 432b, 434b, 436b, 438b, 440b. In some embodiments, the wearable system 400 may correspond to the wearable system 200 of FIG. 2, and FIG. 4A schematically shows some portions of that wearable system 200 in more detail. For example, in some embodiments, the waveguide assembly 480 may be integrated within the display 220 of FIG. 2.
[0068] Continuing to refer to FIG. 4, waveguide assembly 480 may also include a plurality of features 458, 456, 454, 452 between the waveguides. In some embodiments, features 458, 456, 454, 452 may be lenses. In other embodiments, features 458, 456, 454, 452 may not be lenses. Rather, they may simply be spacers (e.g., a cladding layer or structure for forming an air gap).
[0069] Waveguides 432b, 434b, 436b, 438b, 440b or a plurality of lenses 458, 456, 454, 452 may be configured to deliver image information to the eye using various levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a particular depth plane and may be configured to output image information corresponding to that depth plane. Image input devices 420, 422, 424, 426, 428 may each be utilized to input image information into waveguides 440b, 438b, 436b, 434b, 432b such that the incident light is dispersed across each individual waveguide for output toward the eye 410. Light exits from the output surfaces of image input devices 420, 422, 424, 426, 428 and is input into the corresponding input edges of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) is input into each waveguide and outputs an entire field of cloned collimated beams that are directed toward the eye 410 at a particular angle (and divergence amount) corresponding to a particular depth plane associated with the particular waveguide.
[0070] In some embodiments, the image input devices 420, 422, 424, 426, 428 are discrete displays that each generate image information for input into their respective waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the image input devices 420, 422, 424, 426, 428 are the output ends of a single multiplexed display that can send image information to each of the image input devices 420, 422, 424, 426, 428, for example, via one or more optical waveguides (such as optical fiber cables).
[0071] The controller 460 controls the operation of the stacked waveguide assembly 480 and the image input devices 420, 422, 424, 426, 428. The controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that adjusts the timing and provision of image information to the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the controller 460 may be a single integrated device or a distributed system connected by a wired or wireless communication channel. The controller 460 may, in some embodiments, be part of the processing module 260 or 270 (shown in FIG. 2).
[0072] Waveguides 440b, 438b, 436b, 434b, 432b may be configured to propagate light within each individual waveguide by total internal reflection (TIR). The waveguides 440b, 438b, 436b, 434b, 432b may each be planar, or have another shape (e.g., curved), with a major top surface and a bottom surface and an edge extending between their major top and bottom surfaces. In the illustrated configuration, the waveguides 440b, 438b, 436b, 434b, 432b each include light extraction optical elements 440a, 438a, 436a, 434a, 432a configured to extract light from the waveguide by redirecting the light propagating within each individual waveguide out of the waveguide and outputting the image information to the eye 410. The extracted light may also be referred to as external coupled light, and the light extraction optical elements may also be referred to as external coupling optical elements. The beam of the extracted light is output by the waveguide at the location where the light propagating within the waveguide impinges on the light redirecting element. The light extraction optical elements (440a, 438a, 436a, 434a, 432a) may be, for example, reflective or diffractive optical features. For ease of explanation and clarity of the drawings, they are shown disposed on the bottom major surface of the waveguides 440b, 438b, 436b, 434b, 432b, but in some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be disposed on the top or bottom major surface, or may be disposed directly within the volume of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be attached to a transparent substrate and formed within a layer of the material forming the waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, the waveguides 440b, 438b, 436b, 434b, 432b may be a monolithic piece of material, and the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be formed on and / or within the surface of that piece of material.
[0073] Continuing to refer to FIG. 4, as discussed herein, each of the waveguides 440b, 438b, 436b, 434b, 432b is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye may be configured to deliver collimated light to the eye 410 as it is input into such waveguide 432b. The collimated light may represent an optical infinity focal plane. The next upper waveguide 434b may be configured to output collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 may be configured to generate some convex wavefront curvature such that the eye / brain interprets the light arising from its next upper waveguide 434b as arising from a first focal plane that is closer inwardly toward the eye 410 from optical infinity. Similarly, the third upper waveguide 436b passes its output light through both the first lens 452 and the second lens 454 before reaching the eye 410. The combined refractive power of the first and second lenses 452 and 454 may be configured to generate another incremental amount of wavefront curvature such that the eye / brain interprets the light arising from the third waveguide 436b as arising from a second focal plane that is closer inwardly toward the person from optical infinity than the light from the next upper waveguide 434b was.
[0074] Other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, and the top waveguide 440b in the stack sends its output through all of the lenses between it and the eye for the collective focusing power that represents the focal plane closest to the person. When viewing / interpreting light originating from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 may be disposed on top of the stack to compensate for the stack of lenses 458, 456, 454, 452. (The compensating lens layer 430 and the stacked waveguide assembly 480 may be configured such that light originating from the world 470 is transmitted to the eye 410 with substantially the same level of divergence (or collimation) as the light had when it was first received by the stacked waveguide assembly 480.) Such a configuration provides the same number of perceived focal planes as there are available waveguide / lens pairs. Both the light extraction optical elements of the waveguides and the focusing sides of the lenses may be static (e.g., not dynamic or electroactive). In some alternative embodiments, one or both may be dynamic using electroactive features.
[0075] Continuing to refer to FIG. 4, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be configured to redirect light out of their respective waveguides for a particular depth plane associated with the waveguide and output the light with an appropriate amount of divergence or collimation. As a result, waveguides having different associated depth planes may have light extraction optical elements of different configurations that output light with different amounts of divergence depending on the associated depth plane. In some embodiments, as discussed herein, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be three-dimensional or surface features configured to output light at a specific angle. For example, the light extraction optical elements 440a, 438a, 436a, 434a, 432a may be a volume hologram, a surface hologram, and / or a diffraction grating. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published Jun. 25, 2015, which is incorporated herein by reference in its entirety.
[0076] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffraction features or “diffractive optical elements” (also referred to herein as “DOEs”) that form a diffraction pattern. Preferably, the DOE has a relatively low diffraction efficiency such that only a portion of the light of the beam is deflected toward the eye 410 at each intersection of the DOE while the remainder continues to travel through the waveguide via total internal reflection. The light carrying the image information is thus split into several associated output beams that exit the waveguide at multiple locations, resulting in a very uniform pattern of output emission toward the eye 304 with respect to this particular collimated beam that bounces within the waveguide.
[0077] In some embodiments, one or more DOEs may be switchable between an “on” state where they actively diffract and an “off” state where they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer dispersed liquid crystal where the microdroplets have a diffraction pattern in the host medium, and the refractive index of the microdroplets may be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract the incident light), or the microdroplets may be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts the incident light).
[0078] In some embodiments, the number and distribution of depth planes or depth of field may be dynamically varied based on the pupil size or orientation of the viewer's eye. The depth of field may vary inversely with the pupil size of the viewer. As a result, as the pupil size of the viewer's eye decreases, the depth of field increases such that a single plane that was indistinguishable because its location was beyond the depth of focus of the eye becomes distinguishable, and may appear more in focus with the reduction in pupil size and the corresponding increase in depth of field. Similarly, the number of separated depth planes used to present different images to the viewer may be decreased with a decreased pupil size. For example, a viewer may not be able to clearly perceive the details of both a first depth plane and a second depth plane at one pupil size without adjusting the eye's focusing from one depth plane to the other. However, these two depth planes may simultaneously be sufficiently in focus for the user at another pupil size without changing the focusing adjustment.
[0079] In some embodiments, the display system may vary the number of waveguides that receive image information based on a determination of pupil size or orientation, or in response to receiving an electrical signal indicative of a particular pupil size or orientation. For example, if the user's eye is indistinguishable between two depth planes associated with two waveguides, the controller 460 (which may be an embodiment of the local processing and data module 260) can be configured or programmed to stop providing image information to one of these waveguides. Advantageously, this can reduce the processing burden on the system, thereby increasing the responsiveness of the system. In embodiments where the DOE for the waveguide is switchable between on and off states, the DOE may be switched to the off state when the waveguide receives image information.
[0080] In some embodiments, it may be desirable to satisfy the condition that the exit beam has a diameter less than the diameter of the viewer's eye. However, satisfying this condition can be difficult in light of the variability of the viewer's pupil size. In some embodiments, this condition is satisfied over a wide range of pupil sizes by varying the size of the exit beam in response to a determination of the viewer's pupil size. For example, as the pupil size decreases, the size of the exit beam may also decrease. In some embodiments, the exit beam size may be varied using a variable aperture.
[0081] The wearable system 400 can include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 can be referred to as the field of view (FOV) of the world camera, and the imaging system 464 is sometimes also referred to as the FOV camera. The FOV of the world camera may or may not be the same as the FOV of the viewer 210 and includes a portion of the world 470 that the viewer 210 perceives at a given time. For example, in some situations, the FOV of the world camera can be larger than the field of view of the viewer 210 of the wearable system 400. The entire area available for viewing or imaging by the viewer can be referred to as the field of regard (FOR). The FOR may include a solid angle of 4π steradians surrounding the wearable system 400 so that the wearer can move their body, head, or eyes and perceive substantially any direction in space. In other contexts, the movement of the wearer may be more restricted, and accordingly, the wearer's FOR can contact a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used to track gestures (e.g., hand or finger gestures) made by the user and detect objects within the world 470 in front of the user, etc.
[0082] The wearable system 400 includes an audio sensor 232, for example, a microphone, and can capture ambient sound. As described above, in some embodiments, one or more other audio sensors can be positioned to provide stereo sound reception useful for determining the location of the source of speech. The audio sensor 232, as another example, can comprise a directional microphone, which can also provide such useful directional information regarding the location where the audio source is located. The wearable system 400 can use information from both the outward-facing imaging system 464 and the audio sensor 230 when locating the source of speech or determining the active speaker at a particular instant. For example, the wearable system 400 can use speech recognition, alone or in combination with a reflected image of the speaker (e.g., as seen in a mirror), to determine the identification of the speaker. As another example, the wearable system 400 can determine the location of a speaker in the environment based on the sound obtained from a directional microphone. The wearable system 400 can use a speech recognition algorithm to analyze the sound resulting from the location of the speaker, determine the content of the speech, and use speech recognition techniques to determine the identification of the speaker (e.g., name or other demographic information).
[0083] The wearable system 400 can also include an inward-facing imaging system 466 (e.g., comprising a digital camera) that observes the movement of the user, such as eye movement (e.g., for eye tracking) and face movement. The inward-facing imaging system 466 can capture an image of the eye 410 and may be used to determine the size and / or orientation of the pupil of the eye 304. The inward-facing imaging system 466 can be used to obtain an image for use in determining the direction the user is looking (e.g., eye pose), or for biometric identification of the user (e.g., via iris identification). The inward-facing imaging system 426 can be used to provide input images and information for the multiple exposure time eye tracking technique described herein. In some embodiments, at least one camera is utilized to separately determine the pupil size or eye pose of each eye independently for each eye, thereby enabling the presentation of image information to each eye to be dynamically adjusted for that eye. In some other embodiments, only the pupil diameter or orientation of a single eye 410 (e.g., using only a single camera per pair of eyes) is determined and assumed to be similar for both eyes of the user. The image obtained by the inward-facing imaging system 466 may be analyzed to determine the eye pose or mood of the user, which can be used by the wearable system 400 to determine the audio or visual content to be presented to the user. The wearable system 400 can also use sensors such as an IMU, accelerometer, gyroscope, etc. to determine the head pose (e.g., head position or head orientation). The inward-facing imaging system 426 can comprise a camera 324 and a light source 326 (e.g., IR LED) as described with reference to FIG. 3.
[0084] The wearable system 400 can include a user input device 466 through which a user can input commands to the controller 460 and interact with the wearable system 400. For example, the user input device 466 can include a trackpad, a touch screen, a joystick, a multi-degree of freedom (DOF) controller, a capacitance sensing device, a game controller, a keyboard, a mouse, a directional pad (D-pad), a wand, a tactile device, a totem (e.g., functioning as a virtual user input device), and the like. A multi-DOF controller can sense user input in translational (e.g., left / right, forward / backward, or up / down) or rotational (e.g., yaw, pitch, or roll) movements that are possible for some or all of the controller. A multi-DOF controller that supports translational movement can be referred to as 3DOF, while a multi-DOF controller that supports both translational and rotational movement can be referred to as 6DOF. In some cases, the user may use a finger (e.g., the thumb) to press or swipe on a touch sensor-based input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input device 466 may be held by the user's hand during use of the wearable system 400. The user input device 466 can communicate with the wearable system 400 either wired or wirelessly. (Example of eye image)
[0085] FIG. 5 illustrates an image of an eye 500 with an eyelid 504, a sclera 508 (“white of the eye”), an iris 512, and a pupil 516. Curve 516a indicates the pupil boundary between the pupil 516 and the iris 512, and curve 512a indicates the edge boundary between the iris 512 and the sclera 508. The eyelid 504 includes an upper eyelid 504a and a lower eyelid 504b. The eye 500 is illustrated in a natural rest position (e.g., oriented such that both the user's face and line of sight will be directed toward a distant object straight ahead of the user). The natural rest position of the eye 500 is in a natural rest position (e.g., out of plane immediately, with respect to the eye 500 shown in FIG. 5), and in the present embodiment, can be indicated by a natural rest direction 520 which is the direction orthogonal to the surface of the eye 500 when centered within the pupil 516.
[0086] As the eye 500 moves to look toward different objects, the eye pose will change with respect to the natural rest direction 520. The current eye pose is a direction orthogonal to the surface of the eye (and centered within the pupil 516), but can be determined with reference to an eye pose direction 524 which is oriented toward the object at which the eye is currently directed. Referring to the exemplary coordinate system shown in FIG. 5A, the pose of the eye 500 can be represented as two angular parameters that both indicate the azimuth deviation and zenith deviation of the eye pose direction 524 of the eye with respect to the natural rest direction 520 of the eye. For illustrative purposes, these angular parameters can be represented as θ (azimuth deviation, determined from a reference azimuth) and φ (zenith deviation, sometimes also referred to as polar deviation). In some implementations, the angular roll of the eye around the eye pose direction 524 can be included in the determination of the eye pose, and the angular roll can be included in eye tracking. In other implementations, other techniques for determining the eye pose can be used, such as a pitch, yaw, and optionally, a roll system. Thus, the eye pose can be provided as a 2DOF or 3DOF orientation.
[0087] The light source 326 can illuminate the eye 500 (e.g., within the IR), and the reflection of the light source from the eye (typically from the cornea) is referred to as a flash. FIG. 5 schematically shows an embodiment where there are four flashes 550. The position, number, brightness, etc. of the flashes 550 can depend on the position and number of the light source 326, the posture of the eye, etc. As will be further described below, the eye tracking camera 324 can acquire an eye image, and the processor can analyze the eye image to determine the position and movement of the flash for eye tracking. In some embodiments, multiple eye images with different exposure times or different frame rates can be used to provide high-accuracy eye tracking.
[0088] Eye images can be obtained from a video using any suitable process, e.g., a video processing algorithm that can extract the image from one or more sequential frames (or non-sequential frames). The inward-facing imaging system 426 of FIG. 4 or the camera 324 and light source 326 of FIG. 3 can be utilized to provide a video or image of one or both eyes. The posture of the eye can be determined from the eye image using various eye tracking techniques, e.g., the multiple exposure time technique for accurate corneal flash detection described herein. For example, the eye posture can be determined by considering the lens effect of the cornea on the provided light source. Any suitable eye tracking technique can be used to determine the eye posture in the eyelid shape estimation technique described herein. (Example of an eye tracking system)
[0089] FIG. 6 illustrates a schematic diagram of a wearable system 600 that includes an eye tracking system 601. The wearable system 600 may be an embodiment of the wearable systems 200 and 400 described with reference to FIGS. 2-4. The wearable system 600, in at least some embodiments, includes components located within a head-mounted unit 602 and components located within a non-head-mounted unit 604. The non-head-mounted unit 604 may be, by way of example, a belt-mounted component, a handheld component, a component within a backpack, a remote component, etc. Incorporating some of the components of the wearable system 600 within the non-head-mounted unit 604 can help reduce the size, weight, complexity, and cost of the head-mounted unit 602. In some implementations, some or all of the functionality described as being implemented by one or more components of the head-mounted unit 602 and / or the non-head-mounted unit 604 may be provided using one or more components included anywhere within the wearable system 600. For example, some or all of the functionality described below in connection with the CPU 612 of the head-mounted unit 602 may be provided using the CPU 616 of the non-head-mounted unit 604, and vice versa. In some examples, some or all of such functionality may be provided using a peripheral device of the wearable system 600. Further, in some implementations, some or all of such functionality may be provided using one or more cloud computing devices or other remotely located computing devices in a manner similar to that described above with reference to FIG. 2.
[0090] As shown in FIG. 6, the wearable system 600 can include an eye tracking system 601 that captures an image of the user's eye 610, including a camera 324. Optionally, the eye tracking system may also include light sources 326a and 326b (such as light emitting diodes "LEDs"). The light sources 326a and 326b can generate a flash (e.g., a reflection from the user's eye that appears in the image of the eye captured by the camera 324). A schematic example of the flash 550 is shown in FIG. 5. The positions of the light sources 326a and 326b relative to the camera 324 can be known, such that the position of the flash in the image captured by the camera 324 can be used in tracking the user's eye (as will be described in more detail below). In at least one embodiment, there may be one light source 326 and one camera 324 associated with one of the user's eyes 610. In another embodiment, there may be one light source 326 and one camera 324 associated with each of the user's eyes 610. In yet other embodiments, there may be one or more cameras 324 and one or more light sources 326 associated with one or each of the user's eyes 610. As a specific example, there may be two light sources 326a and 326b and one or more cameras 324 associated with each of the user's eyes 610. As another example, there may be three or more light sources such as 326a and 326b and one or more cameras 324 associated with each of the user's eyes 610.
[0091] The eye tracking module 614 may receive an image from the eye tracking camera 324, analyze the image, and extract various information. As described herein, the image from the eye tracking camera may include a shorter exposure (flash) image and a longer exposure image. As an example, the eye tracking module 614 may detect the user's eye posture, the three-dimensional position of the user's eyes with respect to the eye tracking camera 324 (and the head-mounted unit 602), the direction of one or both of the user's focused eyes 610, the user's convergence / divergence movement depth (e.g., the depth from the user at which the user is focused), the position of the user's pupils, the position of the user's corneas and corneal spheres, the respective centers of rotation of the user's eyes, or the respective line-of-sight centers of the user's eyes. As shown in FIG. 6, the eye tracking module 614 may be a software module implemented using the CPU 612 within the head-mounted unit 602.
[0092] Data from the eye tracking module 614 may be provided to other components within the wearable system. As an example, such data may be transmitted to components within a non-head-mounted unit 604, such as the CPU 616, including software modules for the light field rendering controller 618 and the alignment observer 620.
[0093] As further described herein, in some implementations of the multiple exposure time eye tracking technology, the functionality may be implemented differently than that shown in FIG. 6 (or FIG. 7), which is illustrative and not intended to be limiting. For example, in some implementations, the flash images of shorter exposures can be processed by the CPU 612 (which may be disposed within the camera 324) within the head-mounted unit 602, and the images of longer exposures can be processed by the CPU 616 (or GPU 621) within the non-head-mounted unit 604 (e.g., within a belt pack). In some such implementations, some of the eye tracking functionality implemented by the eye tracking module 614 may be implemented by a processor (e.g., CPU 616 or GPU 621) within the non-head-mounted unit 604 (e.g., a belt pack). This can be advantageous because some of the eye tracking functionality can be CPU-intensive and, in some cases, can be implemented more efficiently or quickly by a more powerful processor disposed within the non-head-mounted unit 604.
[0094] The rendering controller 618 may adjust the image displayed to the user using the information from the eye tracking module 614 by the rendering engine 622 (which may be a software module within the GPU 620 and can provide the image to the display 220, a rendering engine). As an example, the rendering controller 618 may adjust the image displayed to the user based on the center of rotation or the center of gaze of the user. In particular, the rendering controller 618 may use the information regarding the center of gaze of the user to simulate a rendering camera (e.g., simulate the collection of an image from the user's line of sight), and based on the simulated rendering camera, adjust the image displayed to the user. Further details regarding the creation, adjustment, and use of the rendering camera in the rendering process are discussed in "METHODS AND Provided in U.S. Patent Application No. 15 / 274,823, entitled "SYSTEMS FOR DETECTING AND COMBINING STRUCTURAL FEATURES IN 3D RECONSTRUCTION" (which is hereby expressly incorporated herein by reference in its entirety).
[0095] In some embodiments, one or more modules (or components) of system 600 (e.g., light field rendering controller 618, rendering engine 620, etc.) may determine the position and orientation of a rendering camera within a rendering space based on the position and orientation of a user's head and eyes (e.g., as determined based on head pose and eye tracking data, respectively). For example, system 600 may effectively map the position and orientation of the user's head and eyes to specific locations and angular positions within a 3D virtual environment, place and orient a rendering camera at the specific locations and angular positions within the 3D virtual environment, and render virtual content for the user as would be captured by the rendering camera. Further details regarding the real-world / virtual-world mapping process are provided in U.S. Patent Application No. 15 / 296,869, entitled "SELECTING VIRTUAL OBJECTS IN A THREE-DIMENSIONAL SPACE" (which is hereby expressly incorporated herein by reference in its entirety). As an example, rendering controller 618 may adjust the depth at which an image is displayed by selecting the depth plane (or depth planes) to be utilized at any given time for displaying the image. In some implementations, such depth plane switching may be performed through adjustment of one or more intrinsic rendering camera parameters.
[0096] The alignment observer 620 may identify whether the head-mounted unit 602 is properly positioned on the user's head using information from the eye tracking module 614. As an example, the eye tracking module 614 may provide eye location information such as the position of the center of rotation of the user's eyes indicating the three-dimensional position of the user's eyes relative to the camera 324, and the head-mounted unit 602 and the eye tracking module 614 may use the location information to determine whether the display 220 is properly aligned within the user's field of view, or whether the head-mounted unit 602 (or headset) has slipped or is otherwise misaligned with the user's eyes. As an example, the alignment observer 620 may determine whether the head-mounted unit 602 has slipped from the user's nasal bridge and thus moved the display 220 away from and downward from the user's eyes (which may not be desirable), whether the head-mounted unit 602 has moved above the user's nasal bridge and thus moved the display 220 closer to and upward from the user's eyes, whether the head-mounted unit 602 has shifted left or right relative to the user's nasal bridge, whether the head-mounted unit 602 has been lifted above the user's nasal bridge, or whether the head-mounted unit 602 has moved away from the desired position or range of positions in these or other ways. Generally, the alignment observer 620 may generally be able to determine whether the head-mounted unit 602, and in particular the display 220, is properly positioned in front of the user's eyes. In other words, the alignment observer 620 may determine whether the left display in the display system 220 is properly aligned with the user's left eye and whether the right display in the display system 220 is properly aligned with the user's right eye. The alignment observer 620 may determine whether the head-mounted unit 602 is properly positioned by determining whether the head-mounted unit 602 is positioned and oriented within the desired range of positions and / or orientations relative to the user's eyes.Exemplary alignment observation and feedback techniques that may be utilized by alignment observer 620 are described in U.S. Patent Application No. 15 / 717,747, filed September 27, 2017, titled "PERIOCULAR TEST FOR MIXED REALITY CALIBRATION", and U.S. Patent Application No. 16 / 251,017, filed January 17, 2019, titled "DISPLAY SYSTEMS AND METHODS FOR DETERMINING REGISTRATION BETWEEN A DISPLAY AND A USER’S EYES", both of which are hereby incorporated by reference in their entirety.
[0097] Rendering controller 618 can receive eye tracking information from eye tracking module 614 and may provide an output to rendering engine 622, which can generate an image to be displayed for viewing by a user of wearable system 600. As an example, rendering controller 618 may receive vergence / accommodation motion depth, left and right eye centers of rotation (and / or line of sight centers), and other eye data such as blink data, saccade data, etc. Vergence / accommodation motion depth information and other eye data based on such data can cause rendering engine 622 to convey content with a particular depth plane (e.g., at a particular depth of field or focal distance) to the user. As discussed in connection with FIG. 4, the wearable system may include a plurality of discrete depth planes formed by a plurality of waveguides, each transmitting image information with a variable level of wavefront curvature. In some embodiments, the wearable system may include one or more variable depth planes, such as optical elements, that transmit image information with a level of wavefront curvature that varies over time. Rendering engine 622 can, in part, convey content to the user at a selected depth based on the user's vergence / accommodation motion depth (e.g., cause rendering engine 622 to instruct display 220 to switch depth planes).
[0098] The rendering engine 622 can generate content by simulating cameras at the positions of the user's left and right eyes and generating the content based on the lines of sight of the simulated cameras. As discussed above, the rendering camera may be a simulated camera for use in rendering virtual image content from a database of objects within a virtual world. The objects may have locations and orientations relative to the user or wearer, and possibly relative to real objects within an environment surrounding the user or wearer. The rendering camera may be included within the rendering engine to render a virtual image based on a database of virtual objects for presentation to the eyes. The virtual image may be rendered as if it were captured from the line of sight of the user or wearer. For example, the virtual image may be rendered as if it were captured by a camera (corresponding to the "rendering camera") having an aperture, lens, and detector that views the objects within the virtual world. The virtual image is captured from the line of sight of such a camera having the position of the "rendering camera". For example, the virtual image may be rendered as if it were captured from the line of sight of a camera having a specific location relative to the user's or wearer's eyes such that the virtual image provides an image that appears to be from the line of sight of the user or wearer. In some implementations, the image is rendered as if it were captured from the line of sight of a camera having an aperture at a specific location relative to the user's or wearer's eyes (such as a line-of-sight center or a center of rotation or other location as discussed herein). (Examples of Eye Tracking Modules)
[0099] A detailed block diagram of the exemplary eye tracking module 614 is shown in FIG. 7. As shown in FIG. 7, the eye tracking module 614 may include various different sub-modules, may provide various different outputs, and may utilize various available data when tracking a user's eyes. By way of example, the eye tracking module 614 may utilize available data including the geometric arrangement of the light source 326 and the eye tracking camera 324 relative to the head-mounted unit 602, an assumed eye dimension 704 such as a typical distance of about 4.7 mm between the center of the user's corneal curvature and the average center of rotation of the user's eye or a typical distance between the user's center of rotation and the line-of-sight center, and per-user calibration data 706 such as the inter-pupillary distance of a particular user, among other incidental and intrinsic properties of eye tracking. Additional examples of incidental properties, intrinsic properties, and other information that may be employed by the eye tracking module 614 are described in U.S. Patent Application No. 15 / 497,726, filed Apr. 26, 2017, published as U.S. Patent Publication No. 2018 / 0018515, entitled "IRIS BOUNDARY ESTIMATION USING CORNEA CURVATURE" (incorporated herein by reference in its entirety). Exemplary eye tracking modules and techniques that may be implemented as or otherwise utilized by the eye tracking module 614 are described in U.S. Patent Application No. 16 / 250,931, filed Jan. 17, 2019, published as U.S. Patent Publication No. 2019 / 0243448, entitled "EYE CENTER OF ROTATION DETERMINATION, DEPTH PLANE SELECTION, AND RENDER CAMERA POSITIONING IN DISPLAY SYSTEMS" (incorporated herein by reference in its entirety).
[0100] The image preprocessing module 710 may receive an image from an eye camera such as the eye camera 324, and may perform one or more preprocessing (e.g., adjustment) operations on the received image. As an example, the image preprocessing module 710 may apply Gaussian blur to the image, may downsample the image to a lower resolution, may apply an unsharp mask, may apply an edge sharpening algorithm, or may apply other suitable filters that assist in subsequent detection, localization, and labeling of flashes, pupils, or other features within the image from the eye camera 324. The image preprocessing module 710 may apply a low-pass filter or a morphological filter such as an open filter that can remove noise, such as high-frequency noise from the pupil boundary 516a (see FIG. 5), thereby interfering with pupil and flash determination. The image preprocessing module 710 may output the preprocessed image to the pupil identification module 712 and the flash detection and labeling module 714.
[0101] The pupil recognition module 712 may receive the pre - processed image from the image pre - processing module 710 and may identify regions of those images that contain the user's pupils. In some embodiments, the pupil recognition module 712 may determine the coordinates of the position of the user's pupil in the eye - tracking image from the camera 324, i.e., the coordinates of the center or centroid. In at least some embodiments, the pupil recognition module 712 may identify the contour (e.g., the contour of the pupil - iris boundary) in the eye - tracking image, identify the contour moments (e.g., the center of mass), apply the starburst pupil detection and / or Canny edge detection algorithm, exclude outliers based on intensity values, identify sub - pixel boundary points, correct for eye camera distortion (e.g., the distortion in the images captured by the eye camera 324), apply the random sample consensus (RANSAC) iterative algorithm, fit an ellipse to the boundary in the eye - tracking image, apply a tracking filter to the image, and identify the sub - pixel image coordinates of the user's pupil centroid. The pupil recognition module 712 may output pupil recognition data, which may indicate the region of the pre - processed image module 712 identified as showing the user's pupil, to the flash detection and labeling module 714. The pupil recognition module 712 may provide the 2D coordinates of the user's pupil (e.g., the 2D coordinates of the centroid of the user's pupil) in each eye - tracking image to the flash detection module 714. In at least some embodiments, the pupil recognition module 712 may also provide the same type of pupil recognition data to the coordinate system normalization module 718.
[0102] Pupil detection techniques that may be utilized by the pupil recognition module 712 are described in U.S. Patent Publication No. 2017 / 0053165, published on February 23, 2017, and U.S. Patent Publication No. 2017 / 0053166, published on February 23, 2017, each of which is incorporated herein by reference in its entirety.
[0103] The flash detection and labeling module 714 may receive the pre - processed image from module 710 and the pupil identification data from module 712. The flash detection module 714 may use this data to detect and / or identify a flash (e.g., the reflection of light from the light source 326 from the user's eye) within the region of the pre - processed image that indicates the user's pupil. As an example, the flash detection module 714 may search for bright regions in the eye - tracking image that are near the user's pupil or iris, sometimes also referred to herein as "blobs" or local intensity maxima. In at least some embodiments, the flash detection module 714 may re - scale (e.g., enlarge) the pupil ellipse to include additional flashes. The flash detection module 714 may filter the flashes by size or intensity. The flash detection module 714 may also determine the 2D position of each flash within the eye - tracking image. In at least some examples, the flash detection module 714 may determine the 2D position of the flash relative to the user's pupil, which may also be referred to as the pupil - flash vector. The flash detection and labeling module 714 may label the flashes and output the pre - processed image with the labeled flashes to the 3D corneal center estimation module 716. The flash detection and labeling module 714 may also convey data such as the pre - processed image from module 710 and the pupil identification data from module 712. In some implementations, the flash detection and labeling module 714 may determine the light source (e.g., among the plurality of light sources of the system, including the infrared light sources 326a and 326b) that produced each identified flash. In these examples, the flash detection and labeling module 714 may label the flashes with information identifying the associated light source and output the pre - processed image with the labeled flashes to the 3D corneal center estimation module 716.
[0104] Pupil and flash detection, such as that implemented by modules such as modules 712 and 714, can use any suitable technique. As an example, edge detection can be applied to an eye image to identify a flash, pupil, or iris. Edge detection can be applied by various edge detectors, edge detection algorithms, or filters. For example, a Canny edge detector can be applied to an image to detect edges such as lines in the image. The edges may include points located along a line corresponding to a local maximum derivative. For example, the pupil boundary 516a or the iris (limbus) boundary 512a (see FIG. 5) can be located using a Canny edge detector. Once the location of the pupil or iris is determined, various image processing techniques can be used to detect the "pose" of the pupil 116. The pose may also be referred to as the line of sight, the direction being looked at, or the orientation of the eye. For example, the pupil may be looking left towards an object, and the pose of the pupil can be classified as a left-facing pose. Other methods can also be used to detect the location of the pupil or flash. For example, concentric rings can be located within an eye image using a Canny edge detector. As another example, an integral differential operator may be used to find the limbus boundary of the pupil or iris. For example, a Daugman integral differential operator, a Hough transform, or other iris segmentation techniques can be used to return a curve that estimates the boundary of the pupil or iris. Modules 712, 714 can be applied to the flash detection techniques described herein that can use multiple eye images captured using different exposure times or different frame rates.
[0105] The 3D corneal center estimation module 716 may receive a pre - processed image including the detected flash data and pupil (or iris) identification data from modules 710, 712, 714. The 3D corneal center estimation module 716 may use this data to estimate the 3D position of the user's cornea. In some embodiments, the 3D corneal center estimation module 716 may estimate the 3D position of the center of the corneal curvature of the eye or the center of the spherical surface of the user's cornea, e.g., generally, the center of an imaginary spherical surface having a surface portion co - extensive with the user's cornea. The 3D corneal center estimation module 716 may provide data indicating the estimated 3D coordinates of the corneal spherical surface and / or the user's cornea to the coordinate system normalization module 718, the optical axis determination module 722, and / or the light field rendering controller 618. Techniques for estimating the position of eye features such as the cornea or corneal spherical surface, which may be utilized by the 3D corneal center estimation module 716 and other modules within the wearable system of the present disclosure, are discussed in U.S. Patent Application No. 15 / 497,726, filed on April 26, 2017 (incorporated herein by reference in its entirety).
[0106] The coordinate system normalization module 718 may optionally be included within the eye tracking module 614 (as indicated by its dashed outline). The coordinate system normalization module 718 may receive data indicating the estimated 3D coordinates of the center of the user's cornea (and / or the center of the user's corneal sphere) from the 3D corneal center estimation module 716, and may also receive data from other modules. The coordinate system normalization module 718 may normalize the eye camera coordinate system, which may help compensate for slippage of the wearable device (e.g., slippage of a head-mounted component from its normal resting position on the user's head, which may be identified by the alignment observer 620). The coordinate system normalization module 718 may rotate the coordinate system and align the z-axis of the coordinate system (e.g., the convergence / divergence motion depth axis) with the corneal center (e.g., as indicated by the 3D corneal center estimation module 716), and may translate the camera center (e.g., the origin of the coordinate system) to a predetermined distance away from the corneal center, such as 30 mm (e.g., the module 718 may expand or contract the eye tracking image depending on whether the eye camera 324 is determined to be closer or farther than the predetermined distance). By using this normalization process, the eye tracking module 614 may be able to establish consistent orientation and distance within the eye tracking data, relatively independently of variations in the headset positioned on the user's head. The coordinate system normalization module 718 may provide the 3D coordinates of the center of the cornea (and / or corneal sphere), pupil identification data, and the preprocessed eye tracking image to the 3D pupil center locator module 720.
[0107] The 3D pupil center locator module 720 may receive data including the 3D coordinates of the center of the user's cornea (and / or corneal sphere), pupil location data, and pre - processed eye - tracking images, in a normalized or non - normalized coordinate system. The 3D pupil center locator module 720 may analyze such data to determine the 3D coordinates of the user's pupil center in a normalized or non - normalized eye camera coordinate system. The 3D pupil center locator module 720 may determine the location of the user's pupil in three dimensions based on the 2D position of the pupil centroid (as determined by module 712), the 3D position of the corneal center (as determined by module 716), assumed eye dimensions 704 such as the size of the typical user's corneal sphere and the typical distance from the corneal center to the pupil center, and the optical properties of the eye such as the refractive index of the cornea (relative to the refractive index of air), or any combination of these. Techniques for estimating the location of eye features such as the pupil, which may be utilized by the 3D pupil center locator module 720 and other modules within the wearable system of the present disclosure, are discussed in U.S. Patent Application No. 15 / 497,726, filed on April 26, 2017, which is incorporated herein by reference in its entirety.
[0108] The optical axis determination module 722 may receive data indicating the 3D coordinates of the user's cornea and the user's pupil center from modules 716 and 720. Based on such data, the optical axis determination module 722 may identify a vector from the location of the corneal center (e.g., from the center of the corneal sphere) to the user's pupil center that may define the optical axis of the user's eye. As an example, the optical axis determination module 722 may provide an output defining the user's optical axis to modules 724, 728, 730, and 732.
[0109] The center of rotation (CoR) estimation module 724 may receive data from module 722 that includes parameters of the optical axis of the user's eye (e.g., data indicating the direction of the optical axis in a coordinate system with a known relationship to the head-mounted unit 602). The CoR estimation module 724 may estimate the center of rotation of the user's eye (e.g., the point around which the user's eye rotates as the user's eye rotates left, right, up, and / or down). Assume that a single point may be sufficient even if the eye cannot rotate perfectly around a single point. In at least some embodiments, the CoR estimation module 724 may estimate the center of rotation of the eye by moving the pupil center (identified by module 720) or the center of curvature of the cornea (as identified by module 716) a specific distance along the optical axis (identified by module 722) towards the retina. This specific distance may be the assumed eye dimension 704. As an example, the specific distance between the center of curvature of the cornea and the CoR may be about 4.7 mm. This distance may be varied for a particular user based on any relevant data including the user's age, gender, vision prescription, other relevant characteristics, etc.
[0110] In at least some embodiments, the CoR estimation module 724 may refine over time the estimated value of the respective center of rotation of the user's eyes. As an example, over time, the user will ultimately rotate their eyes (to look at something else, closer, farther away, or sometimes left, right, up, or down), and will displace them along the respective optical axes of their eyes. The CoR estimation module 724 may then analyze the two (or more) optical axes identified by module 722 and localize the 3D point of intersection of those optical axes. The CoR estimation module 724 may then determine the center of rotation at the 3D point of intersection. Such techniques may provide an estimated value of the center of rotation with improved accuracy over time. Various techniques may be employed to increase the accuracy of the CoR estimation module 724 and the determined CoR positions of the left and right eyes. As an example, the CoR estimation module 724 may estimate the CoR by finding the average point of intersection of the optical axes determined over time for various different eye postures. As an additional example, module 724 may filter or average the estimated CoR positions over time, may calculate a moving average of the estimated CoR positions over time, and / or may apply a Kalman filter and the known dynamics of the eye and eye tracking system to estimate the CoR position over time. As a specific example, module 724 may calculate a weighted average of the determined point of intersection of the optical axes and the assumed CoR position (e.g., 4.7 mm behind the center of the corneal curvature of the eye) such that the determined CoR moves slowly over time to a slightly different location within the user's eye as eye tracking data for the user is acquired, thereby enabling per-user refinement of the CoR position.
[0111] The interpupillary distance (IPD) estimation module 726 may receive data indicating the estimated 3D positions of the centers of rotation of the user's left and right eyes from the CoR estimation module 724. The IPD estimation module 726 may then estimate the user's IPD by measuring the 3D distance between the centers of rotation of the user's left and right eyes. Generally, the distance between the estimated CoR of the user's left eye and the estimated CoR of the user's right eye may be approximately equal to the distance between the user's pupil centers when the user is looking at optical infinity (e.g., the optical axes of the user's eyes are substantially parallel to each other), which is the typical definition of the interpupillary distance (IPD). The user's IPD may be used by various components and modules within the wearable system. As an example, the user's IPD may be provided to the alignment observer 620 and used when assessing the degree to which the wearable device is aligned with the user's eyes (e.g., whether the left and right display lenses are appropriately spaced according to the user's IPD). As another example, the user's IPD may be provided to the convergence / divergence motion depth estimation module 728 and used when determining the user's convergence / divergence motion depth. Module 726 may employ various techniques such as those discussed in relation to the CoR estimation module 724 to increase the accuracy of the estimated IPD. As an example, the IPD estimation module 724 may apply filtering, averaging over time, weighted averaging, including an assumed IPD distance, a Kalman filter, etc. as part of the estimation of the user's IPD in an accurate manner.
[0112] In some embodiments, the IPD estimation module 726 may receive data indicating the estimated 3D positions of the user's pupils and / or corneas from the 3D pupil center locator module and / or the 3D corneal center estimation module 716. The IPD estimation module 726 may then estimate the user's IPD by referring to the distance between the pupil and the cornea. Generally, these distances will vary over time as the user rotates their eyes and changes the depth of their convergence / divergence movements. In some cases, the IPD estimation module 726 may look for the maximum measured distance between the pupils and / or corneas that should occur while the user is looking near optical infinity and that should generally correspond to the user's interpupillary distance. In other cases, the IPD estimation module 726 may fit the measured distance between the user's pupils (and / or corneas) to a mathematical relationship that describes how the interpupillary distance of a person changes as a function of the depth of their convergence / divergence movements. In some embodiments, using these or other similar techniques, the IPD estimation module 726 may be able to estimate the user's IPD without the observation that the user is looking at optical infinity (e.g., by extrapolating from one or more observations that the user was convergence / diverging at a distance closer than optical infinity).
[0113] The convergence / divergence motion depth estimation module 728 may receive data from various modules and sub-modules within the eye tracking module 614 (as shown in relation to FIG. 7). In particular, the convergence / divergence motion depth estimation module 728 may employ data indicative of the estimated 3D position of the pupil center (e.g., as provided by module 720 described above), one or more determined parameters of the optical axis (e.g., as provided by module 722 described above), the estimated 3D position of the center of rotation (e.g., as provided by module 724 described above), the estimated IPD (e.g., the Euclidean distance between the estimated 3D positions of the centers of rotation) (e.g., as provided by module 726 described above), and / or one or more determined parameters of the optical axis and / or the visual axis (e.g., as provided by module 722 and / or module 730 described below). The convergence / divergence motion depth estimation module 728 may detect or otherwise obtain a measurement of the user's convergence / divergence motion depth, which may be the distance from the user at which the user's eyes are focused. As an example, when the user is looking at an object 3 feet from their front, the user's left and right eyes have a convergence / divergence motion depth of 3 feet, while when the user is looking at a distant landscape (e.g., the optical axes of the user's eyes are substantially parallel to each other such that the distance between the user's pupil centers can be approximately equal to the distance between the centers of rotation of the user's left and right eyes), the user's left and right eyes have an infinite convergence / divergence motion depth. In some implementations, the convergence / divergence motion depth estimation module 728 may utilize data indicative of the estimated center of the user's pupil (e.g., as provided by module 720) and determine the 3D distance between the estimated centers of the user's pupils. The convergence / divergence motion depth estimation module 728 may obtain a measurement of the convergence / divergence motion depth by comparing such a determined 3D distance between the pupil centers with the estimated IPD (e.g., the Euclidean distance between the estimated 3D positions of the centers of rotation) (e.g., as indicated by module 726 described above).In addition to the 3D distance between pupil centers and the estimated IPD, the convergence / divergence motion depth estimation module 728 may calculate the convergence / divergence motion depth using known, assumed, estimated, and / or determined geometries. As an example, module 728 may combine the 3D distance between pupil centers, the estimated IPD, and the 3D CoR position in a triangulation calculation to estimate (e.g., determine) the user's convergence / divergence motion depth. In fact, the evaluation of such a determined 3D distance between pupil centers relative to the estimated IPD can serve to indicate a measured value of the user's current convergence / divergence motion depth relative to optical infinity. In some embodiments, the convergence / divergence motion depth estimation module 728 may simply receive or access data indicative of the estimated 3D distance between the estimated centers of the user's pupils for the purpose of obtaining such a measured value of the convergence / divergence motion depth. In some embodiments, the convergence / divergence motion depth estimation module 728 may estimate the convergence / divergence motion depth by comparing the user's left and right optical axes. In particular, the convergence / divergence motion depth estimation module 728 may estimate the convergence / divergence motion depth by locating the distance from the user at which the user's left and right optical axes intersect (or the projections of the user's left and right optical axes on a plane such as a horizontal plane intersect). Module 728 may utilize the user's IPD in this calculation by setting zero depth to be the depth at which the user's left and right optical axes are separated by the user's IPD. In at least some embodiments, the convergence / divergence motion depth estimation module 728 may determine the convergence / divergence motion depth by triangulating the eye tracking data with known or derived spatial relationships.
[0114] In some embodiments, the convergence / divergence motion depth estimation module 728 may estimate the user's convergence / divergence motion depth based on the intersection of the user's visual axes (instead of their optical axes), which may provide a more accurate indication of the distance at which the user is focused. In at least some embodiments, the eye tracking module 614 may include an optical axis / visual axis mapping module 730. As will be discussed in more detail in connection with FIG. 10, the user's optical axis and visual axis generally do not coincide. The visual axis is the axis along which a person is looking, while the optical axis is defined by the user's lens and pupil center and may pass through the center of the user's retina. In particular, the user's visual axis is generally offset from the center of the user's retina, thereby resulting in different optical and visual axes, as defined by the location of the user's fovea. In at least some of these embodiments, the eye tracking module 614 may include an optical axis / visual axis mapping module 730. The optical axis / visual axis mapping module 730 may correct for the difference between the user's optical axis and visual axis and provide information regarding the user's visual axis to other components within the wearable system, such as the convergence / divergence motion depth estimation module 728 and the light field rendering controller 618. In some examples, the module 730 may use an assumed eye dimension 704 that includes a typical offset of approximately 5.2° inward (nasally, towards the user's nose) between the optical axis and the visual axis. In other words, the module 730 may shift the user's left optical axis 5.2° to the right nasally (towards the nose), and the user's right optical axis 5.2° to the left nasally (towards the nose) to estimate the directions of the user's left and right optical axes. In other examples, the module 730 may utilize per-user calibration data 706 when mapping the optical axis (e.g., as indicated by the module 722 described above) to the visual axis. As an additional example, the module 730 may shift the user's optical axis nasally by any range formed by, for example, 4.0° - 6.5°, 4.5° - 6.0°, 5.0° - 5.4°, etc., or any of these values.In some arrays, module 730 may apply an offset, at least in part, based on the characteristics of a particular user, such as their age, gender, visual prescription, or other relevant characteristics, and / or at least in part, based on a calibration process for a particular user (e.g., to determine the optical axis - visual axis offset for a particular user). In at least some embodiments, module 730 may also offset the origin of the left and right optical axes and correspond to the user's CoP (as determined by module 732) instead of the user's CoR.
[0115] When an optional center of pupil (CoP) estimation module 732 is provided, it may estimate the location of the user's left and right centers of pupil (CoP). The CoP is a useful location for the wearable system and, in at least some embodiments, may be the position directly in front of the pupil. In at least some embodiments, the CoP estimation module 732 may estimate the location of the user's left and right centers of pupil based on the 3D location of the user's pupil center, the 3D location of the center of the user's corneal curvature, or such suitable data, or any combination thereof. As an example, the user's CoP may be approximately 5.01 mm in front of the center of the corneal curvature (e.g., 5.01 mm in the direction along the optical axis from the center of the corneal sphere towards the cornea of the eye) and may be approximately 2.97 mm behind the outer surface of the user's cornea along the optical or visual axis. The user's center of pupil may be directly in front of the center of their pupil. As an example, the user's CoP may be less than approximately 2.0 mm from the user's pupil, less than approximately 1.0 mm from the pupil, less than approximately 0.5 mm from the user's pupil, or any range between these values. As another example, the center of pupil may correspond to a location within the anterior chamber of the eye. As another example, the CoP may be between 1.0 mm and 2.0 mm, approximately 1.0 mm, between 0.25 mm and 1.0 mm, between 0.5 mm and 1.0 mm, or between 0.25 mm and 0.5 mm.
[0116] (As the potentially desirable position of the pinhole of the rendering camera and the anatomical position within the user's eye), the line-of-sight center described herein can be a position that serves to reduce and / or eliminate undesirable parallax shifts. In particular, the optical system of the user's eye generally approximates the theoretical system formed by the front pinhole of the lens projecting onto the screen, and the pinhole, lens, and screen generally correspond to the user's pupil / iris, lens, and retina, respectively. Further, when two point light sources (or objects) at different distances from the user's eye rotate precisely about the opening of the pinhole (e.g., rotated along a radius of curvature equal to its individual distance from the opening of the pinhole), it may be desirable for there to be little or no parallax shift. Thus, the CoP would be expected to be located at the center of the pupil of the eye (and such a CoP may be used in some embodiments). However, the human eye includes a cornea that, in addition to the pinhole of the lens and pupil, imparts additional refractive power to light propagating towards the retina. Thus, the anatomical equivalent of the pinhole within the theoretical system described in this paragraph can be the region of the user's eye located between the outer surface of the user's eye's cornea and the center of the user's eye's pupil or iris. For example, the anatomical equivalent of the pinhole can correspond to the region within the anterior chamber of the user's eye. For various reasons discussed herein, it may be desirable to set the CoP to such a position within the anterior chamber of the user's eye.
[0117] As discussed above, the eye tracking module 614 may provide data such as the estimated 3D positions of the left and right eye centers of rotation (CoR), the vergence / accommodation depth, the left and right eye optical axes, the 3D position of the user's eyes, the 3D positions of the left and right centers of corneal curvature of the user, the 3D positions of the left and right pupil centers of the user, the 3D positions of the left and right line-of-sight centers of the user, the user's IPD, etc. to other components such as the light field rendering controller 618 and the alignment observer 620 within the wearable system. The eye tracking module 614 may also include other sub-modules that detect and generate data associated with other aspects of the user's eyes. As an example, the eye tracking module 614 may include a blink detection module that provides a flag or other alert each time the user blinks, and a saccade detection module that provides a flag or other alert each time the user's eyes saccade (e.g., rapidly shift focus to another point). (Example of Flash Positioning Using an Eye Tracking System)
[0118] FIG. 8A is a schematic cross-sectional view of an eye showing the cornea 820, iris 822, lens 824, and pupil 826 of the eye. The sclera (the white of the eye) surrounds the iris 822. The cornea can have a generally spherical shape, shown by the corneal sphere 802, which has a center 804. The eye optical axis is a line (shown by the solid line 830) that passes through the pupil center 806 and the corneal center 804. The user's line-of-sight direction (shown by the dashed line 832 and sometimes also referred to as the line-of-sight vector) is the user's visual axis and is generally at a small offset angle from the eye optical axis. The offset, which is specific to each particular eye, can be determined by user calibration from the eye image. The user-measured offset can be stored by the wearable display system and used to determine the line-of-sight direction from measurements of the eye optical axis.
[0119] Light sources 326a and 326b (such as light emitting diodes, LEDs, etc.) can illuminate the eye and generate a flash (e.g., specular reflection from the user's eye) that is imaged by camera 324a. A schematic example of flash 550 is shown in FIG. 5. The positions of light sources 326a and 326b relative to camera 324 are known, and as a result, the position of the flash within the image captured by camera 324 can be used when tracking the user's eye and modeling corneal sphere 802 to determine its center 804. FIG. 8B is a photograph of an eye showing an example of four flashes 550 produced by four light sources (for this exemplary photograph). Generally, light sources 326a, 326b produce infrared (IR) light so that the light sources are not visible to the user and can avoid distracting the user.
[0120] In some eye tracking systems, a single image of the eye is used to determine information about the pupil (e.g., to determine its center) and the flash (e.g., its position within the eye image). As will be further explained below, measurements of both pupil information and flash information from a single exposure can lead to errors in determining the eye optical axis or line of sight because, for example, the image of the flash can be saturated within an image that also shows sufficient detail to extract pupil information.
[0121] Examples of such errors are illustrated schematically in FIGS. 9A-9C. In FIGS. 9A-9C, the actual line-of-sight vector is shown as solid line 830 passing through pupil center 806 and center 804 of corneal sphere 802, and the line-of-sight vector extracted from the eye image is shown as dashed line 930. FIG. 9A shows an example of the error in the extracted line-of-sight vector when there is a small error 900a only in the measured position of the pupil center. FIG. 9B shows an example of the error in the extracted line-of-sight vector when there is a small error 900b only in the measured position of the corneal center. FIG. 9C shows an example of the error in the extracted line-of-sight vector when there is a small error 900a in the measured position of the pupil center and a small error 900b in the measured position of the corneal center. In some embodiments of the wearable system 200, the error in line-of-sight determination can be about 20 minutes per pixel in the determination of the flash in the eye image, while using some embodiments of the multiple exposure time eye tracking technique described herein, the error can be reduced to less than about 3 minutes per pixel. (Eye imaging using multiple images with different exposure times)
[0122] FIG. 10A shows an example of the uncertainty at the flash position when the flash is acquired from a single long-exposure eye image 1002a. In this example, the exposure time for the eye image 1002a was approximately 700 μs. One of the flashes within the image 1002a is shown in the enlarged image 1006a, and the contour plot 1008a shows the intensity contour of the flash. The enlarged image 1006a shows that the flash is saturated and lacks a peak in flash intensity. The dynamic range of the intensity within the image 1006a can be relatively high in order to capture both the lower-intensity pupil and iris features and the higher-intensity flash. The flash covers many pixels within the image 1006a, which can make it difficult to accurately determine the center of the flash, especially when the peak intensity of the flash is lacking. The full width at half maximum (FWHM) of the flash 1006a is approximately 15 pixels. For some embodiments of the wearable system 200, the error of each pixel in determining the center (or centroid) of the flash corresponds to an error of about 20 minutes in determining the line-of-sight direction or the direction of the eye optical axis. Embodiments of eye imaging using a plurality of eye images captured using different exposure times can advantageously identify the flash position more accurately and advantageously reduce the error in the line-of-sight direction or the optical axis direction to within a few minutes or less.
[0123] FIG. 10B shows an example of the uncertainty of the flash position when a technique using a plurality of images taken at different exposure times is used. In this example, two sequential images, namely, the first longer-exposure image 1002b and the second shorter-exposure image 1004, are captured by the eye-tracking camera. The total capture time for the two images 1002b, 1004 is approximately 1.5 ms. The labels "first" and "second" as applied to the images 1002b and 1004 are not intended to indicate the chronological order in which the images were taken, but rather are for convenience in simply referring to each of the two images. Thus, the first image 1002b can be taken before, after, or the two exposures may at least partially overlap.
[0124] The longer exposure image 1002b can have an exposure time similar to the image 1002a shown in FIG. 10A, which, in this embodiment, is about 700 μs. The first longer exposure image 1002b can be used to determine pupil (or iris) characteristics. For example, the longer exposure image 1002b can be analyzed to determine the pupil center or center of rotation (CoR), extract iris characteristics for biometric security applications, determine eyelid shape or occlusion of the iris or pupil by the eyelid, measure pupil size, determine rendering camera parameters, etc. However, as described with reference to FIG. 10A, the flash can be saturated in such longer exposure images, which can lead to relatively large errors at the flash location.
[0125] In some implementations, the pupil contrast in the longer exposure image can be increased by using an eye-tracking camera having a better modulation transfer function (MTF), e.g., an MTF closer to the diffraction-limited MTF. For example, a better MTF for the imaging sensor or a better MTF for the imaging lens of the camera can be selected to improve pixel contrast. Exemplary imaging sensing devices and techniques that can be implemented as an eye-tracking camera having such an MTF or otherwise employed in one or more of the eye-tracking systems described herein were filed on Dec. 13, 2018, as "GLOBAL SHUTTER PIXEL CIRCUIT AND" Described in U.S. Patent Application No. 16 / 219,829, titled "METHOD FOR COMPUTER VISION APPLICATIONS", and U.S. Patent Application No. 16 / 219,847, titled "DIFFERENTIAL PIXEL CIRCUIT AND METHOD OF COMPUTER VISION APPLICATIONS", filed on December 13, 2018 (both of which are incorporated herein by reference in their entirety). In various embodiments, the eye-tracking camera 324 can produce an image with a pupil contrast, such as 5 - 6 pixels, 1 - 4 pixels, etc. (e.g., measured at the transition between the pupil and the iris). Additional exemplary imaging sensing devices and techniques that can be implemented as the eye-tracking camera 324 or employed in one or more of the eye-tracking systems otherwise described herein are described in U.S. Patent Application No. 15 / 159,491, titled "SEMI-GLOBAL SHUTTER IMAGER", filed on May 19, 2016 (incorporated herein by reference in its entirety).
[0126] The exposure time for the second, shorter exposure image 1004 is substantially less than the exposure time of the first image 1002b and can reduce the likelihood of saturating the flash 550 (e.g., causing a lack of a flash peak in the image). As described above, the second, shorter exposure image 1004 is sometimes also referred to as a flash image herein because it can be used to accurately identify the flash position. In this embodiment, the exposure time of the second image 1004 was less than 40 μs. The flash is perceptible within the second image 1004, but the pupil (or iris) features are not readily perceptible, which is noted as the reason why the longer exposure first image 1002b can be used for pupil center extraction or CoR determination. Image 1006b is an enlarged view of one of the flashes within the second image 1004. The smaller size of the flash within image 1006b (compared to the size within image 1006a) is readily apparent, demonstrating that the flash is not saturated within image 1006b. Contour plot 1008b shows a much smaller contour pattern for the flash (compared to the contour shown in contour plot 1008a). In this embodiment, the location of the center of the flash can be determined to sub-pixel accuracy, e.g., up to about 1 / 10 of a pixel (corresponding to an error of only about 2 minutes in the line-of-sight or optical axis direction in this embodiment). In some implementations, the location of the pixel center can be determined very accurately from the second, shorter exposure image 1004, e.g., by fitting a two-dimensional (2D) Gaussian (or other bell-shaped curve) to the flash pixel values. In some embodiments, the location of the mass center of the flash is determined and can be relied upon in a role similar to that of the location of the center of the flash. The exposure time of the flash image used to determine the flash location can be selected to be only long enough to image the peak of the flash.
[0127] The eye imaging technique can thus utilize the capture of a longer exposure image (which can be used for extraction of pupil properties) and a shorter exposure flash image (which can be used for extraction of flash position). As described above, the longer and shorter images can be taken in any order, or the exposure times can at least partially overlap. The exposure time for the first image may be in the range of 200 μs to 1,200 μs, while the exposure time for the second image may be in the range of 5 μs to 100 μs. The ratio of the exposure time for the first image to the exposure time for the second image can be in the range of 5 to 50, 10 to 20, or some other range. The frame rate at which the longer and shorter exposure images can be captured is described with reference to FIG. 11.
[0128] A potential advantage of determining the flash position from the second shorter exposure image is that the flash covers a relatively small number of pixels (e.g., comparing image 1006b with image 1006a), and finding the flash center can be carried out computationally quickly and efficiently. For example, the search area for the flash can have a diameter of only about 2 to 10 pixels in some embodiments. Further, since the flash covers a relatively small number of pixels, only a relatively small portion of the image needs to be analyzed or stored (e.g., in a buffer), which can provide substantial memory savings. Additionally, the dynamic range of the flash image can be low enough such that images of ambient environmental light sources (e.g., room lights, sky, etc.) are not perceivable due to the short exposure time, which advantageously means that such ambient light sources will not interfere with the flash or be misinterpreted as such.
[0129] Furthermore, using a relatively short exposure time can also serve to reduce the presence of motion blur in an image of an eye engaged in saccades or otherwise moving rapidly. Exemplary eye-tracking and saccade detection systems and techniques and associated exposure time switching and adjustment schemes are described in U.S. Provisional Patent Application No. 62 / 660,180, filed Apr. 19, 2018, entitled "SYSTEMS AND METHODS FOR ADJUSTING OPERATIONAL PARAMETERS OF A HEAD-MOUNTED DISPLAY SYSTEM BASED ON USER SACCADES" (incorporated herein by reference in its entirety). In some implementations, one or more of such exemplary systems, techniques, and schemes may be implemented as, or otherwise utilized by, one or more of the systems and techniques described herein (e.g., by eye-tracking module 614). Additional details of the switching between shorter-duration and longer-duration exposures are described below with reference to FIG. 11.
[0130] As described above, in some implementations, at least a portion of the flash image is temporarily stored in a buffer, and that portion of the flash image is analyzed to identify the location of one or more flashes that may be located within that portion. For example, the portion may include a relatively small number of pixels, rows, or columns of the flash image. In some cases, the portion may include an n×m portion of the flash image, where n and m are in the range of about 1 to 20. For example, a 5×5 portion of the flash image may be stored in the buffer.
[0131] After the position of the flash has been identified, the buffer may be cleared. Additional portions of the flash image may then be stored in the buffer for analysis until either the entire flash image is processed or all (generally four) flashes are identified. In some cases, more than one flash of the image is in the buffer simultaneously for processing. In some cases, most of the flashes of the image are in the buffer simultaneously for processing. In some cases, all of the flashes of the image are in the buffer simultaneously for processing. The flash position (e.g., Cartesian coordinates) may be used for subsequent actions in the eye-tracking process, and after the flash position has been stored or communicated to a suitable processor, the flash image may be deleted from memory (buffer memory or other volatile or non-volatile storage device). Such a buffer advantageously enables fast processing of the flash image to identify the flash position or may reduce the memory requirements of the eye-tracking process since the flash image can be deleted after use.
[0132] In some implementations, the flash images of the shorter exposures are processed by a hardware processor within the wearable systems 200, 400, or 600. For example, the flash image may be processed by the CPU 612 of the head-mounted unit 602, described with reference to FIG. 6. The CPU 612 may include or communicate with a buffer 615, which can be used to temporarily store at least a portion of the flash image for processing (e.g., identifying the flash location). In some implementations, the CPU 612 may output the identified flash location to another processor (e.g., the CPU 616 or GPU 620, described with reference to FIG. 6) that analyzes the image of the longer exposure. In some implementations, this other processor may be remote from the head-mounted display unit. The other processor may be, for example, within a belt pack. Thus, in some implementations, a hardware processor such as the CPU 612 may output the identified flash location (e.g., the Cartesian coordinates of the flash peak), possibly along with data indicating the intensity of the identified flash (e.g., the peak intensity of the flash), to another processor (e.g., the CPU 616 or GPU 620, described with reference to FIG. 6). In some embodiments, the CPU 612 and the buffer 615 are within a hardware processor that is within or associated with the eye-tracking camera 324. Such an arrangement can provide increased efficiency because the flash images of the shorter exposures can be processed by the camera circuitry and do not need to be communicated to another processing unit (either on or off the head-mounted unit 602). The camera 324 may simply output the flash location (e.g., the Cartesian coordinates of the flash peak) to another processing unit (e.g., a hardware processor that analyzes the images of the longer exposure, such as the CPU 612, 616, or GPU 620).Thus, in some embodiments, a hardware processor within or associated with camera 324 may output the identified flash location (e.g., the Cartesian coordinates of the flash peak), possibly along with data indicating the respective intensity of the identified flash (e.g., the peak intensity of the flash), to another processor (e.g., CPU 612, 616 or GPU 620 described with reference to FIG. 6). In some implementations, flash images with shorter exposures are processed by a hardware processor (e.g., CPU 612), which may output the location of the identified flash candidates (e.g., the Cartesian coordinates of the flash peak) that could potentially be flash images, possibly along with data indicating the intensity of the identified flash candidates (e.g., the peak intensity of the flash), to another processor. This other processor may identify a subset of the identified flash candidates as flashes. This other processor may also perform one or more operations (e.g., one or more operations for determining the line-of-sight direction) based on the subset of the identified flash candidates (e.g., the subset of the identified flash candidates considered as flashes by the other processor). Thus, in some implementations, flash images with shorter exposures are processed by a hardware processor (e.g., CPU 612, a hardware processor within or associated with camera 324, etc.), which may output the location of the identified flash candidates (e.g., the Cartesian coordinates of the flash peak), possibly along with data indicating the intensity of the identified flash candidates (e.g., the peak intensity of the flash), to another processor (e.g., CPU 616 or GPU 620 described with reference to FIG. 6), and this other processor (e.g., CPU 616 or GPU 620 described with reference to FIG. 6) may perform one or more operations (e.g., one or more operations for determining the line-of-sight direction) based on the subset of the identified flash candidates (e.g., the subset of the identified flash candidates considered as flashes by the other processor).In some implementations, a hardware processor that processes the flash images of shorter exposures may further output data indicating the intensities (e.g., peak intensity of the flash) of different identified flash candidates to another processor (e.g., CPU 616 or GPU 620 described with reference to FIG. 6). In some embodiments, this other processor may select a subset of the identified flash candidates based on one or both of the position of the identified flash candidates and the intensity of the identified flash candidates. Optionally, one or more operations (e.g., one or more operations for determining the line-of-sight direction) may be further performed based on the selected subset of the identified flash candidates. For example, in at least some of these implementations, another processor (e.g., CPU 616 or GPU 620 described with reference to FIG. 6) may determine the line-of-sight direction of the eye based at least in part on the selected subset of the identified flash candidates and the pupil center of the eye (e.g., as determined by another processor based on one or more images of longer exposures). In some implementations, the subset of the identified flash candidates selected by another processor may include only one flash candidate for each infrared light source employed by the system, while the amount of flash candidates identified and communicated to another processor may exceed the amount of infrared light sources employed by the system. Other configurations and approaches are also possible.
[0133] In various implementations, the images of longer exposures may be processed by CPU 612 or communicated to non-head-mounted unit 604 for processing by, for example, CPU 616 or GPU 621. As described, in some implementations, CPU 616 or GPU 621 may be programmed to analyze the flash images stored by buffer 615 to obtain the flash positions identified from the images of shorter exposures from another processor (e.g., CPU 612). (Frame Rate for Multiple Exposure Time Eye Imaging)
[0134] As described above, capturing multiple eye images using different exposure times can advantageously provide an accurate determination of the flash center (e.g., from an image with a shorter exposure) and an accurate determination of the pupil center (e.g., from an image with a longer exposure). The position of the flash center can be used to determine the pose of the cornea.
[0135] As the user's eye moves, the flash will also move correspondingly. To track eye movement, multiple short flash exposures can be taken to capture the movement of the flash. Thus, embodiments of the eye tracking system can capture flash exposures at a relatively high frame rate (e.g., compared to the frame rate of images with longer exposures) in the range of, for example, about 100 frames per second (fps) to 500 fps. Therefore, the time period between consecutive short flash exposures can be in the range of about 1 - 2 ms to a maximum of about 5 - 10 ms in various embodiments.
[0136] Images with longer exposures for pupil center determination may be captured at a frame rate lower than the frame rate for flash images. For example, the frame rate for images with longer exposures can be in the range of about 10 fps to 60 fps in various embodiments. This is not a requirement, and in some embodiments, the frame rates of both shorter and longer exposure images are the same.
[0137] Figure 11 shows an example of a combined operation mode in which longer exposure images are captured at a frame rate of 50 fps (e.g., 20 ms between consecutive images), and flash images are captured at a frame rate of 200 fps (e.g., 5 ms between consecutive images). The combined operation mode may be implemented by an eye tracking module 614, which is described with reference to FIGS. 6 and 7. The horizontal axis in FIG. 11 is time in milliseconds. The top row in FIG. 11 schematically illustrates the illumination of the eye (e.g., by an IR light source) and the capture of the corresponding image, e.g., by an eye tracking camera 324. The flash image 1004 is illustrated as a sharp triangle, while the longer exposure image 1002 is illustrated as a trapezoid with a wider flat top. In this embodiment, the flash image is captured before the subsequent longer exposure image, but the order can be reversed in other implementations. In this embodiment, since the frame rate for the flash image 1004 is higher (e.g., 4 times higher in this embodiment) than the frame rate for the longer exposure image 1002, FIG. 11 shows four flash images (separated by approximately 5 ms each) that are captured before the next longer exposure image is taken (e.g., 20 ms after the first longer exposure image is captured).
[0138] In other implementations, the flash image frame rate or the longer exposure image frame rate may be different from that shown in FIG. 11. For example, the flash image frame rate can be in the range of about 100 fps to about 1000 fps, and the longer exposure image frame rate can be in the range of about 10 fps to about 100 fps. In various embodiments, the ratio of the flash image frame rate to the longer exposure frame rate can be in the range of about 1 to 20, in the range of 4 to 10, or in some other range.
[0139] The middle row of FIG. 11 schematically illustrates the reading of image data of a flash and a longer exposure (the reading occurs after the image is captured), and the row below FIG. 11 indicates the operation mode of the eye tracking system (e.g., flash priority using a flash image, or pupil priority using an image of a longer exposure). Thus, in this embodiment, the flash position is determined every 5 ms from the flash image of the short exposure, and the pupil center (or CoR) is determined every 20 ms from the image of the longer exposure. The exposure time of the image of the longer exposure is long enough to capture pupil and iris features and has a sufficient dynamic range, but can be short enough so that the eye does not move substantially during the exposure (e.g., to reduce the likelihood of blur in the image). As described above, the exposure time of the image of the longer exposure can be about 700 μs in some embodiments, and the exposure time of the flash image can be less than about 40 μs.
[0140] In some embodiments, the eye tracking module 614 may dynamically adjust, for example, the exposure time between a shorter exposure time and a longer exposure time. For example, the exposure time may be selected (e.g., from a range of exposure times) based on the type of information being determined by the display system. As an example, the eye tracking module 614 may determine the occurrence of saccades. However, in other embodiments, the wearable display system 200 may acquire one or more images and determine the line of sight associated with the user's eye. For example, the display system may utilize the geometry of the user's eye, as described with reference to FIG. 8A, to determine a vector extending from the user's fovea or visual axis of the eye. The display system may thus select a shorter exposure time to reduce the presence of motion blur. Additionally, the display system may perform a biometric authentication process based on an image of the user's eye. For example, the display system may compare known eye features of the user's eye with the user's eye features identified in the image. Thus, the display system may likewise select a shorter exposure time to reduce the presence of motion blur.
[0141] When dynamically adjusting the exposure time, the eye tracking module 614 may alternate between acquiring an image at a first exposure time (e.g., long exposure) and acquiring an image at a second exposure time (e.g., short exposure). For example, the eye tracking module may acquire an image at the first exposure time, determine whether the user is performing saccades, and then subsequently acquire an image at the second exposure time. Additionally, certain operating conditions of the wearable display system 200 may provide knowledge of whether an image should be acquired at the first or second exposure time.
[0142] In some embodiments, the eye tracking module 614 may dynamically adjust one or more of the exposure times. For example, the eye tracking module may increase or decrease the exposure time used for saccade detection or flash position. In this example, the eye tracking module may determine that a measurement associated with motion blur is too high or too low. That is, the measurement may not accurately detect saccades or may over-detect saccades due to the exposure time. For example, the eye tracking module may be configured to perform saccade detection using both motion blur detection and comparison between continuously captured image frames. Assuming that the comparison between image frames provides a more accurate determination of the occurrence of saccades, the results provided by comparing multiple images may be used as a reference, and the motion blur detection may be adjusted until a desired (e.g., high) level of agreement is reached between the results of the two schemes for saccade detection. If the image frame comparison indicates that saccades are under-detected, the eye tracking module may be configured to increase the exposure time. Conversely, if saccades are erroneously detected, the exposure time may be decreased.
[0143] The eye tracking module may also, or alternatively, dynamically adjust the exposure time for longer exposure images used to determine pupil or iris properties (e.g., pupil center, CoR, etc.). For example, if the longer exposure image has a high dynamic range such that iris details are saturated, the eye tracking module may reduce the exposure time.
[0144] In some embodiments, the eye tracking module may also adjust the exposure time when performing a biometric authentication process or when determining the user's line of sight. For example, the display system may dynamically reduce the exposure time to reduce motion blur, or the eye tracking module may increase the exposure time if the acquired image is not properly exposed (e.g., if the image is too dark).
[0145] In some embodiments, the eye tracking module may utilize the same camera (e.g., camera 324) for each image of the user's eye obtained. That is, the eye tracking module may utilize the camera facing the eye of a particular user. When the user performs saccades, both eyes can move in a corresponding manner (e.g., similar speed and amplitude). Thus, the eye tracking module can reliably determine whether saccades are being performed using images of the same eye. Optionally, the eye tracking module may utilize the camera facing each eye of the user. In such embodiments, the eye tracking module may optionally utilize the same camera to acquire images of the same eye or may select a camera for use. For example, the eye tracking module may select a camera that is not currently in use. The eye tracking module may acquire images of the user's eyes for purposes other than determining the occurrence of saccades. As an example, the eye tracking module may perform gaze detection, etc. (e.g., the eye tracking module may determine a 3D point at which the user is fixating, a prediction of the future gaze direction (e.g., for foveation rendering), biometric authentication (e.g., the eye tracking module may determine whether the user's eye matches a known eye)). In some embodiments, when the eye tracking module provides a command that an image should be captured, one of the cameras may be in use. Thus, the eye tracking module may select a camera that is not in use to acquire an image for use in saccade detection, gaze direction determination, blink identification, biometric authentication, etc.
[0146] Optionally, the eye tracking module may trigger both cameras and acquire images simultaneously. For example, each camera may acquire an image at an individual exposure time. In this way, the eye tracking module acquires a first image of the first eye and determines a measurement of motion blur, a flash position, etc., while acquiring a second image of the second eye and determining other information (e.g., information used for gaze detection, pupil center determination, authentication, etc.). Optionally, both images may be utilized to determine whether the user is performing a saccade. For example, the eye tracking module may determine a deformation of a feature (e.g., the user's eye, flash, etc.) shown in the first image by comparing it with the same feature as shown in the second image. Optionally, the eye tracking module may alternate between two exposure values for each camera, for example, such that they are out of phase with each other. For example, the first camera may acquire an image at a first exposure value (e.g., a shorter exposure time), while simultaneously the second camera may acquire an image at a second exposure value (e.g., a longer exposure time). Subsequently, the first camera may acquire an image at the second exposure value, and the second camera may acquire an image at the first exposure value. (Flash motion detector)
[0147] FIG. 12 schematically illustrates an example of how the use of short-exposure flash images captured at a relatively high frame rate can provide robust flash detection and tracking as the eye moves. In frame #1, four exemplary flashes are shown. Frame #2 is a flash image taken shortly after frame #1 (e.g., about 5 ms later at a 200 fps rate), illustrating the extent to which the initial flash position (labeled as #1) has moved to the position labeled by #2. Similarly, frame #3 is another flash image taken shortly after frame #2 (e.g., about 5 ms later at a 200 fps rate), illustrating the extent to which the initial flash position has continued to move to the position labeled by #3. Note that in frame #2, only the flash at the position labeled by #2 appears in the image, and the flash at the position labeled by #1 is shown for reference, and the same is true for frame #3. Thus, FIG. 12 schematically shows the extent to which an exemplary constellation or pattern of flashes has moved from frame to frame.
[0148] Since flash images can be captured at a relatively high frame rate, determining the flash position in an earlier frame can assist in determining the expected position of the flash in subsequent frames when the flash does not move significantly between frames when captured at a relatively high frame rate (e.g., the flash movement reference depicted in FIG. 12). Such flash imaging enables the eye tracking system to limit the search area per frame to a small number of pixels (e.g., 2-10 pixels) around the previous flash position, which can advantageously improve processing speed and efficiency. For example, as described above, only a portion of the flash image (e.g., a 5×5 group of pixels) may be stored in the temporary memory buffer 615 for processing by the associated CPU 612. The flash present within the image portion can be quickly and efficiently identified, and its position can be determined. The flash position can be used for subsequent processing (e.g., using an image with a longer exposure), and the flash image may be deleted from memory.
[0149] A further advantage of such flash imaging is that since the eye tracking system can follow slight movement of the flash from frame to frame, each labeling of the flash is less likely to introduce an error (e.g., mislabeling the top left flash as the top right or bottom left flash). For example, the “constellation” of four flashes depicted in FIG. 12 may tend to move at a substantially common speed from frame to frame as the eye moves. The common speed may represent the representative or average speed of the constellation. Thus, the system can check that all identified flashes are moving at approximately the same speed. If one (or more) of the flashes within the “constellation” is moving at a substantially different speed (e.g., different by more than a threshold amount from the common speed), then that flash may be misidentified or the flash may be reflected from a non-spherical portion of the cornea.
[0150] Thus, a slight change in the flash position from frame to frame can be relied upon with relatively high confidence, for example, if all four flashes exhibit comparable changes in position. Thus, in some implementations, the eye tracking module 614 may be able to detect saccade and microsaccade movements with relatively high confidence based on just two image frames. For example, in these implementations, the eye tracking module may determine whether a global change in the flash position from one frame to the next exceeds a threshold, and in response to determining that such a global change does not actually exceed the threshold, may be configured to determine that the user's eye is engaged in a saccade (or microsaccade) movement. This may advantageously be utilized, for example, to perform depth plane switching as described in U.S. Patent Application No. 62 / 660,180, incorporated above and filed on April 19, 2018.
[0151] Generally, similar to FIG. 8A, FIG. 13A illustrates a situation where one of the flashes is reflected from the aspherical portion of the cornea and the eye tracking system can determine the corneal modeling center of the corneal sphere 1402 from the flashes from light sources 326a and 326b. For example, the reflections from light sources 326a, 326b can be projected back to the corneal modeling center 804 where both dashed lines 1302 meet at a common point (shown as dashed line 1302). However, when the system projects towards the center of the flash from light source 326c, the dashed line 1304 does not meet at the corneal modeling center 804. If the eye tracking system attempts to find the center 804 of the corneal sphere 802 using the flash from light source 326c, an error will be introduced into the center position (because line 1304 does not intersect the center 804). Thus, by tracking the flash positions within the flash image, the system can identify situations where the flash is likely to be reflected from the aspherical region of the cornea (e.g., because the flash speed is different from the common speed of the constellation) and remove the flash from the corneal sphere modeling calculation (or reduce the weight assigned to that flash). The flash speed of the flash within the aspherical corneal region is often much higher than the flash speed for the spherical region, and this increase in speed can be used, at least in part, to identify when the flash originates from the aspherical corneal region.
[0152] When a user blinks, the user's eyelids can cover at least partially a portion of the cornea from which light would be reflected. The flash resulting from this area can have an image shape that has a lower intensity or is substantially non-spherical (which can introduce an error in determining its position). FIG. 13B is an image showing an example of a flash 550a where there is a partial occlusion. The flash imaging can also be used to identify when the flash is in a state of being at least partially occluded by monitoring the intensity of the flash within the image. In the absence of occlusion, each flash can have approximately the same intensity for each frame, while in the presence of at least partial occlusion, the intensity of the flash will rapidly decrease. Thus, the eye tracking system can monitor the flash intensity as a function of time (e.g., for each frame), determine when partial or complete occlusion occurs, and remove the flash from the eye tracking analysis (or reduce the weight assigned to that flash). Further, the constellations of flashes can have approximately similar intensities, and thus, the system can monitor whether there is a difference in the intensity of a particular flash from the common intensity of the constellation. In response to such a determination, the eye tracking system can remove the flash from the eye tracking analysis (or reduce the weight assigned to that flash).
[0153] The flash detection and labeling module 714, described with reference to FIG. 7, thus advantageously includes a flash motion detector that can monitor the flash velocity or flash intensity for each frame in the flash imaging and provide a more robust determination of the flash position, which can provide a more robust determination of the center of the corneal sphere (determined, for example, by the 3D corneal center estimation module 716 of FIG. 7). (Exemplary flash-pupil motion relationship)
[0154] The applicant has determined that there is a relationship between the movement of the flash and the movement of the pupil within the eye image. FIGS. 14A and 14B are graphs of an example of flash movement versus pupil movement in a Cartesian (x,y) coordinate system, where the x-axis is horizontal and the y-axis is vertical. FIG. 14A shows the flash x-location (in pixel units) versus the pupil x-location (in pixel units), and FIG. 14B shows the flash y-location (in pixel units) versus the pupil y-location (in pixel units). In these examples, four flashes (A, B, C, and D) were tracked, and there was no movement of the user's head relative to the eye tracking (ET) system during the measurements. As can be seen from FIGS. 14A and 14B, there is a strong linear relationship between the flash movement and the pupil movement. For example, the flash movement has approximately half the speed of the pupil movement. The specific values of the linear relationship and slope in FIGS. 14A and 14B may be due to the geometry of the eye tracking system used for the measurements, and a similar relationship between flash movement and pupil movement is expected to exist for other eye tracking systems (and can be determined by analysis of the eye tracking images).
[0155] The flash-pupil relationship can be used to provide robustness to the determination of the flash center, along with the robustness provided by flash imaging, and can also be used to provide robustness to the determination of the pupil center (or center of rotation, CoR). Further, leveraging the flash-pupil relationship to determine the pupil position can advantageously provide a computational savings for pupil identification and positioning techniques, as described above with reference to one or more of the modules of FIG. 7, such as module 712 or module 720. For example, the flash position can be robustly determined from the flash image (e.g., to sub-pixel accuracy), and an estimate of the position of the pupil center (or CoR) can be predicted, at least in part, based on the flash-pupil relationship (see, e.g., FIGS. 14A and 14B). Over a relatively short time period during which the flash image is captured (e.g., every 2-10 ms), the pupil center (or CoR) does not change significantly. The flash position can be tracked and averaged over a plurality of flash frame captures (e.g., 1-10 frames). Using the flash-pupil relationship, the average flash position provides an estimate of the pupil position (or CoR). In some embodiments, this estimate can then be used in the analysis of an image with a longer exposure to more accurately and reliably determine the pupil center (or CoR). In other embodiments, this estimate can, by itself, serve to provide the system with knowledge of the pupil center. Thus, in at least some of these embodiments, the eye tracking module does not need to rely on an image with a longer exposure to determine the pupil center (or CoR). Rather, in such embodiments, the eye tracking module can rely on an image with a longer exposure for other purposes, or in some examples, can entirely omit the capture of such an image with a longer exposure. (Gaze Prediction for Foveation Rendering)
[0156] As described with reference to FIG. 6, the rendering controller 618 can use information from the eye tracking module 614 to adjust the image to be displayed to the user by the rendering engine 622 (e.g., a software module within the GPU 620 that can provide an image to the display 220 of the wearable system 200). As an example, the rendering controller 618 may adjust the image to be displayed to the user based on the center of rotation or the center of gaze of the user.
[0157] In some systems, the pixels within the display 220 are rendered at a higher resolution or frame rate closer to the line of sight than in regions of the display 220 farther from the line of sight (in some cases, they may not be rendered at all). This is sometimes also referred to as foveated rendering and can provide a substantial computational performance gain, mainly because only pixels substantially within the line of sight can be rendered. Foveated rendering can provide an increase in rendering bandwidth and a reduction in power consumption by the system 200. An example of a wearable system 200 that utilizes foveated rendering is described in U.S. Patent Publication No. 2018 / 0275410, entitled "Depth Based Foveated Rendering for Display Systems", which is incorporated herein by reference in its entirety.
[0158] FIG. 15 schematically illustrates an example of foveation rendering. The field of view (FoV) of a display (e.g., display 220) with the original location in the line-of-sight direction (also referred to as the fovea) is shown as circular 1502. Arrow 1504 represents a high-speed saccade of the eye (e.g., about 300 deg / sec). Dashed circular 1506 represents an area of uncertainty that may be the case when the line-of-sight direction is a high-speed saccade. Pixels of the display within area 1506 may be rendered, while pixels of the display outside display area 1506 may not need to be rendered (or may be rendered at a lower frame rate or resolution). The area of area 1506 increases as approximately the square of the time between when the line of sight moves to a new direction and when the image can actually be rendered by the rendering pipeline (sometimes also referred to as the time scale from motion to image drawing). In various embodiments of the wearable system 200, the time scale from motion to image drawing is in the range of about 10 ms to 100 ms. Thus, it can be advantageous for the eye-tracking system to be able to predict the future line-of-sight direction (ahead by approximately the time from motion to image drawing for display) so that the rendering pipeline can start generating image data for display when the user's eye moves to a future line-of-sight direction. Such prediction of the future line-of-sight direction for foveation rendering can provide advantages to the user, such as increased apparent responsiveness of the display, reduced latency in image generation, etc.
[0159] FIG. 16 schematically illustrates an exemplary timing diagram 1600 for a rendering pipeline that utilizes an embodiment of multiple exposure time eye imaging for eye tracking. The rendering pipeline first is idle 1602 (waiting for the next flash exposure), and then a flash exposure 1604 is captured (the exposure time may be less than 50 μs). The flash image is read by the camera stack 1606, and the read time can be reduced by reading only the peak of the flash image. The flash and the flash position are extracted at 1608, and the eye tracking system determines the line of sight at 1610. The renderer (e.g., rendering controller 618 and rendering engine 622) renders the image at 1612 for display.
[0160] FIG. 17 is a block diagram of an exemplary line of sight prediction system 1700 for foveation rendering. System 1700 can be implemented as part of system 600, described, for example, with reference to FIGS. 6 and 7. System 1700 has two pipelines: an image pipeline for acquiring a flash image and a flash pipeline for processing the flash and predicting the future line of sight direction for a line of sight prediction time. As described above, the line of sight prediction time for foveation rendering may be comparable to the time from motion to image rendering for system 600 (e.g., 10 ms to 100 ms in various implementations). For some rendering applications, the line of sight prediction time is in the range of about 20 ms to 50 ms, for example, about 25 ms to 30 ms. The blocks shown in FIG. 17 are illustrative, and in other embodiments, one or more of the blocks can be combined, rearranged, omitted, or additional blocks can be added to system 1700.
[0161] In block 1702, system 1700 receives a flash image. An example of a flash image is image 1004 shown in FIG. 10B where the peak of the flash is imaged. In block 1706, the system can threshold process the image, for example, set the intensity value of the lowest image pixel value to a lower threshold (e.g., zero). In block 1710, non-maximum values within the flash image are suppressed or removed, which can assist in finding only the flash peak. For example, system 1700 can scan across a row (or column) of the image to find the maximum value. When proceeding to block 1710, near maximum values can be processed to remove the lower maximum values of a group of closely spaced maximum values (e.g., maximum values within a threshold pixel distance from each other, e.g., 2 - 20 pixels). For example, a user may wear contact lenses. Reflection 326 of the light source can occur at both the front of the contact lens and the cornea, resulting in two closely spaced flashes. Non-maximum suppression in block 1710 can exclude the lower maximum values so that only the primary or first specular reflection is retained for further processing.
[0162] The flash pipeline receives the processed flash image from the image pipeline. Flash tracking and classification (e.g., labeling) can be performed in block 1714. As described with reference to FIG. 12, information about previously identified flashes (e.g., position, intensity, etc.) can be received from block 1716, and block 1714 can use this information to identify search regions within the flash image where each of the flashes is likely to be found (which can reduce the search time and increase the processing efficiency). The previous flash information from block 1716 can also be useful in identifying whether a flash is at least partially occluded or whether the flash is originating from a non-spherical portion of the cornea. Such flashes may be removed from flash tracking in block 1714.
[0163] In block 1718, system 1700 calculates the current line of sight (e.g., the eye optical axis or the line of sight vector) from the flash information received from block 1714. In block 1720, the previous line of sight information (e.g., determined from the previous flash image) is input into block 1722, where the future line of sight direction is calculated. For example, system 1700 can extrapolate from the current line of sight (block 1718) and one or more previous lines of sight (block 1720) to predict the line of sight at a future line of sight time (e.g., 10 ms to 100 ms in the future). The line of sight predicted from block 1722 is provided to the rendering engine and can enable foveation rendering.
[0164] In some embodiments, the pipeline within system 1700 activates flash imaging at a high frame rate (e.g., 100 fps to 400 fps).
[0165] Figures 18A - 18D illustrate the results of an experiment to predict future gaze using an embodiment of the gaze prediction system 1700. In this exemplary experiment, the flash image frame rate was 160 fps. Figure 18A shows the flash paths for every four flashes when the user's eye was moving across a field of view that was approximately 40 minutes by 20 minutes. The paths of the four flashes are indicated by differently shaded dots, and the path of the average value of these flashes is indicated by an unshaded square. Figure 18B shows a comparison of the gaze prediction (short dashed line) versus the path of the average value of the flashes (which is considered the ground truth for this experiment) within the unshaded square. The future prediction time was 37.5 ms, and flashes from 18.75 ms back in the past were used for the prediction. As can be seen from Figure 18B, much of the prediction, if not most of it, accurately tracks the average path of the eye. Figure 18C is a plot of the angular eye velocity (in degrees per second) versus the frame, with the velocity histogram 1820 accompanying the left side of the figure. Figure 18C shows that most of the eye movements occur at relatively low angular velocities (e.g., below about 25 degrees per second), are continuous, but are accompanied by sporadic or random movements up to about 200 degrees per second for a few moments. Figure 18D is a receiver operating characteristic (ROC) plot showing the prediction percentile versus the error between the prediction and the ground truth (GT). The different lines are for different future prediction times from 6.25 ms to 60 ms. The closer the line is to vertical (towards the left), the more accurate the prediction. For a prediction time of about 6.25 ms (which is the reciprocal of the frame rate for this experiment), the error is very small. The error increases as the prediction time gets longer, but even at a prediction time of 50 ms, over 90 percent of the gaze predictions have an error of less than 5 pixels. (Exemplary Method for Eye Tracking)
[0166] FIG. 19 is a flowchart illustrating an exemplary method 1900 for eye tracking. The method 1900 can be implemented, for example, by an embodiment of the wearable display system 200, 400, or 600 using the eye tracking system 601 described with reference to FIGS. 6 and 7. In various embodiments of the method 1900, the blocks described below can be implemented in any suitable order or sequence, the blocks can be combined or rearranged, or other blocks can be added.
[0167] In block 1904, the method 1900 captures an image of the eye with a longer exposure, for example, using the eye tracking camera 324. The exposure time of the image with the longer exposure can be in the range of 200 μs to 1,200 μs, for example, about 700 μs. The exposure time for the image with the longer exposure can depend on the nature of the eye tracking camera 324 and can be different in other embodiments. The exposure time for the image with the longer exposure should have a sufficient dynamic range to capture pupil or iris features. The image with the longer exposure can be captured at a frame rate in the range of 10 fps to 60 fps in various embodiments.
[0168] In block 1908, method 1900 captures an image of a shorter exposure (flash) of the eye, for example, using eye tracking camera 324. The exposure time of the flash image may be in the range of 5 μs to 100 μs, for example, less than about 40 μs. The exposure time for the flash image with a shorter exposure may be less than the exposure time for the image with a longer exposure captured in block 1904. The exposure time for the flash image may depend on the nature of the eye tracking camera 324 and may be different in other embodiments. The exposure time for the flash image should be sufficient to image the peak of the flash from light source 326. The exposure time may be sufficient to enable the sub-pixel location of the flash center, such that the flash is not saturated and has a width of several (e.g., 1 to 5) pixels. The ratio of the exposure time for the image with a longer exposure to the exposure time for the flash image may be in the range of 5 to 50, 10 to 20, or some other range. The flash image can be captured at a frame rate in the range of 100 fps to 1,000 fps in various embodiments. The ratio of the frame rate for the flash image to the frame rate for the image with a longer exposure may be in the range of 1 to 100, 1 to 50, 2 to 20, 3 to 10, or some other ratio.
[0169] In block 1912, method 1900 (e.g., using CPU 612) can determine the pupil center (or center of rotation, CoR) from the image with a longer exposure obtained in block 1904. Method 1900 can analyze the image with a longer exposure for other eye features (e.g., iris code) for other biometric applications.
[0170] In block 1916, method 1900 (e.g., using CPU 612) can determine the flash position from the flash image obtained at block 1908. For example, method 1900 can fit a 2-D Gaussian to the flash image to determine the flash position. Other functionality can be implemented at block 1916. For example, as described with reference to FIG. 17, method 1900 can threshold the flash image and remove non-maximum or near-maximum values from the flash image or the like. Method 1900 can track the "constellation" of the flash and use the average speed of the constellation to assist in identifying the estimated flash position, whether the flash is at least partially occluded by being reflected from a non-spherical region of the cornea, etc. In block 1916, method 1900 can utilize the previously determined position of the flash (e.g., within a previous flash image) to assist in localizing or labeling the flash within the current flash image.
[0171] In block 1920, method 1900 (e.g., using CPU 612) can determine the current line of sight using information determined from the image with a longer exposure and the flash image with a shorter exposure. For example, the flash position obtained from the flash image and the pupil center obtained from the image with a longer exposure can be used to determine the line-of-sight direction. The line-of-sight direction can be represented as two angles, e.g., θ (azimuth deviation determined from the base azimuth angle) and φ (zenith deviation, sometimes also referred to as polar deviation), as described with reference to FIG. 5A.
[0172] In block 1924, method 1900 (e.g., using CPU 612) can predict a future line of sight from the current line of sight and one or more previous lines of sight. As described with reference to FIGS. 15 - 17, the future line of sight can advantageously be used for foveation rendering. The future line of sight can be predicted via extrapolation techniques. Block 1924 can predict the future line of sight at a future line-of-sight time (e.g., 10 ms to 100 ms after the time when the current line of sight was calculated).
[0173] In block 1928, method 1900 (e.g., using rendering controller 618 or rendering engine 622) can render virtual content for presentation to a user of the wearable display system. As described above, for foveated rendering, knowledge of the user's future line-of-sight direction can be used to more efficiently initiate preparation of virtual content for rendering when the user is looking in the future line-of-sight direction, which can advantageously reduce latency and improve rendering performance.
[0174] In some embodiments, the wearable display system may not utilize foveated rendering techniques and method 1900 may not predict future line-of-sight directions. In such embodiments, block 1924 is optional. (Additional Example) (Part I)
[0175] (Example 1) A wearable display system comprising: a head-mounted display configured to present virtual content by outputting light to the eyes of a wearer of the head-mounted display; a light source configured to direct light towards the eyes of the wearer; an eye-tracking camera configured to capture a first image of the wearer's eyes captured at a first frame rate and a first exposure time, and a second image of the wearer's eyes captured at a second frame rate greater than the first frame rate and a second exposure time less than the first exposure time; and a hardware processor communicatively coupled to the head-mounted display and the eye-tracking camera, the hardware processor being programmed to analyze the first image to determine the pupil center of the eye, analyze the second image to determine the position of the reflection of the light source from the eye, determine the line-of-sight direction of the eye from the pupil center and the position of the reflection, estimate the future line-of-sight direction of the eye at a future line-of-sight time from the line-of-sight direction and previous line-of-sight direction data, and program the head-mounted display to present virtual content at the future line-of-sight time based at least in part on the future line-of-sight direction.
[0176] (Example 2) The system according to wearable display Example 1, wherein the light source comprises an infrared light source.
[0177] (Example 3) The wearable display system according to Example 1 or Example 2, wherein the hardware processor is programmed to analyze the first image or the second image to determine the center of rotation or the line-of-sight center of the wearer's eyes.
[0178] (Example 4) The wearable display system according to any one of Examples 1-3, wherein the hardware processor is programmed to apply a threshold to the second image, identify non-maximum values within the second image, or suppress or remove non-maximum values within the second image in order to analyze the second image.
[0179] (Example 5) To analyze the second image, the hardware processor is programmed, at least in part, to identify a search area for the reflection of the light source in the current image from the second image based on the position of the reflection of the light source in the previous image from the second image, for the wearable display system according to any one of embodiments 1-4.
[0180] (Embodiment 6) To analyze the second image, the hardware processor is programmed to determine a common speed of the reflections of a plurality of light sources and to determine whether the speed of the reflection of the light source exceeds a threshold amount and differs from the common speed, for the wearable display system according to any one of embodiments 1-5.
[0181] (Embodiment 7) To analyze the second image, the hardware processor is programmed to determine whether the reflection of the light source is from an aspherical portion of the cornea of the eye, for the wearable display system according to any one of embodiments 1-6.
[0182] (Embodiment 8) To analyze the second image, the hardware processor is programmed to identify the presence of at least partial occlusion of the reflection of the light source, for the wearable display system according to any one of embodiments 1-7.
[0183] (Embodiment 9) The hardware processor is programmed, at least in part, to determine the estimated position of the pupil center based on the position of the reflection of the light source and the flash-pupil relationship, for the wearable display system according to any one of embodiments 1-8.
[0184] (Embodiment 10) The flash-pupil relationship includes a linear relationship between the flash position and the pupil center position, for the wearable display system according to embodiment 9.
[0185] (Example 11) The wearable display system according to any one of Examples 1-10, wherein the first exposure time is within the range of 200 μs to 1,200 μs.
[0186] (Example 12) The wearable display system according to any one of Examples 1-11, wherein the first frame rate is within the range of 10 frames / second to 60 frames / second.
[0187] (Example 13) The wearable display system according to any one of Examples 1-12, wherein the second exposure time is within the range of 5 μs to 100 μs.
[0188] (Example 14) The wearable display system according to any one of Examples 1-13, wherein the second frame rate is within the range of 100 frames / second to 1,000 frames / second.
[0189] (Example 15) The wearable display system according to any one of Examples 1-14, wherein the ratio of the first exposure time to the second exposure time is within the range of 5 to 50.
[0190] (Example 16) The wearable display system according to any one of Examples 1-15, wherein the ratio of the second frame rate to the first frame rate is within the range of 1 to 100.
[0191] (Example 17) The wearable display system according to any one of Examples 1-16, wherein the future gaze time is within the range of 5 ms to 100 ms.
[0192] (Example 18) The hardware processor includes a first hardware processor disposed on a non-head-mounted component of the wearable display system and a second hardware processor disposed within or on the head-mounted display. The first hardware processor is used to analyze a first image, and the second hardware processor is used to analyze a second image. The wearable display system according to any one of Embodiments 1-17.
[0193] (Embodiment 19) The second hardware processor includes or is associated with a memory buffer configured to store at least a part of each of the second images. The second hardware processor is programmed to delete at least a part of each of the second images after determining the position of the reflection of the light source from the eye. The wearable display system according to Embodiment 18.
[0194] (Embodiment 20) The hardware processor is programmed not to combine the first image and the second image. The wearable display system according to any one of Embodiments 1-19.
[0195] (Embodiment 21) A method for eye tracking, including the steps of capturing a first eye image by an eye tracking camera at a first frame rate and a first exposure time, capturing a second eye image by the eye tracking camera at a second frame rate higher than the first frame rate and a second exposure time shorter than the first exposure time, determining the pupil center of the eye from at least the first image, determining the position of the reflection of the light source from the eye from at least the second image, and determining the line-of-sight direction of the eye from the pupil center and the position of the reflection.
[0196] (Embodiment 22) The method according to Example 21, further comprising the step of rendering virtual content on the display, at least in part, based on the line-of-sight direction.
[0197] (Example 23) The method according to Example 21 or Example 22, further comprising the step of estimating a future line-of-sight direction at a future line-of-sight time, at least in part, based on the line-of-sight direction and previous line-of-sight direction data.
[0198] (Example 24) The method according to Example 23, further comprising the step of rendering virtual content on the display, at least in part, based on the future line-of-sight direction.
[0199] (Example 25) The method according to any one of Examples 21-24, wherein the first exposure time is in the range of 200 μs to 1,200 μs.
[0200] (Example 26) The method according to any one of Examples 21-25, wherein the first frame rate is in the range of 10 frames per second to 60 frames per second.
[0201] (Example 27) The method according to any one of Examples 21-26, wherein the second exposure time is in the range of 5 μs to 100 μs.
[0202] (Example 28) The method according to any one of Examples 21-27, wherein the second frame rate is in the range of 100 frames per second to 1,000 frames per second.
[0203] (Example 29) The method according to any one of Examples 21-28, wherein the ratio of the first exposure time to the second exposure time is in the range of 5 to 50.
[0204] (Example 30) The method according to any one of Examples 21-29, wherein the ratio of the second frame rate to the first frame rate is within the range of 1 to 100.
[0205] (Example 31) A wearable display system, comprising: a head-mounted display configured to present virtual content by outputting light to the eyes of a wearer of the head-mounted display; a light source configured to direct light toward the eyes of the wearer; an eye-tracking camera configured to capture an image of the wearer's eyes, the eye-tracking camera being configured to alternately capture a first image at a first exposure time and a second image at a second exposure time less than the first exposure time; and a plurality of electronic hardware components, at least one of which includes a hardware processor communicatively coupled to the head-mounted display, the eye-tracking camera, and at least one other electronic hardware component within the plurality of electronic hardware components, the hardware processor receiving each first image of the wearer's eyes captured by the eye-tracking camera at the first exposure time and relaying it to at least one other electronic hardware component, receiving the pixels of each second image of the wearer's eyes captured by the eye-tracking camera at the second exposure time, storing them in a buffer, analyzing the pixels stored in the buffer, identifying the location where the reflection of the light source is present in each second image of the wearer's eyes captured by the eye-tracking camera at the second exposure time, and transmitting location data indicating the location to at least one other electronic hardware component, at least one other electronic hardware component being configured to analyze each first image of the wearer's eyes captured by the eye-tracking camera at the first exposure time, identify the location of the center of the pupil of the eye, and determine the line-of-sight direction of the wearer's eyes from the location of the center of the pupil and the location data received from the hardware processor.
[0206] (Example 32) The first exposure time is within the range of 200 μs to 1,200 μs, the wearable display system described in Example 31.
[0207] (Example 33) The second exposure time is within the range of 5 μs to 100 μs, the wearable display system described in Example 31 or Example 32.
[0208] (Example 34) The eye-tracking camera is configured to capture a first image at a first frame rate within the range of 10 frames per second to 60 frames per second, the wearable display system according to any one of Examples 31 - 33.
[0209] (Example 35) The eye-tracking camera is configured to capture a second image at a second frame rate within the range of 100 frames per second to 1,000 frames per second, the wearable display system according to any one of Examples 31 - 34.
[0210] (Example 36) The ratio of the first exposure time to the second exposure time is within the range of 5 to 50, the wearable display system according to any one of Examples 31 - 35.
[0211] (Example 37) At least one other electronic hardware component is disposed within or on a non-head-mounted component of the wearable display system, the wearable display system according to any one of Examples 31 - 36.
[0212] (Example 38) The non-head-mounted component includes a belt pack, the wearable display system described in Example 37.
[0213] (Example 39) The pixels of each second image of the eye are included in fewer pixels than all of the second images, the wearable display system according to any one of Examples 31-38.
[0214] (Example 40) The pixels include an array of n×m pixels, where n and m are each integers within the range of 1 to 20, the wearable display system according to any one of Examples 31-39.
[0215] (Example 41) The plurality of electronic hardware components are further configured to estimate the future line-of-sight direction of the eye at a future line-of-sight time from the line-of-sight direction and previous line-of-sight direction data, and cause the head-mounted display to present virtual content at the future line-of-sight time, at least partially based on the future line-of-sight direction, the wearable display system according to any one of Examples 31-40.
[0216] (Example 42) The hardware processor is programmed to apply a threshold to the pixels stored in the buffer, identify non-maximum values within the pixels stored in the buffer, or suppress or remove non-maximum values within the pixels stored in the buffer, the wearable display system according to any one of Examples 31-41.
[0217] (Example 43) The hardware processor is programmed to determine a common speed of reflection of the light source and determine whether the speed of reflection of the light source exceeds a threshold amount and differs from the common speed, the wearable display system according to any one of Examples 31-42.
[0218] (Example 44) The hardware processor is programmed to determine whether the location of reflection of the light source is from a non-spherical portion of the cornea of the eye, the wearable display system according to any one of Examples 31-43.
[0219] (Example 45) The wearable display system according to any one of Examples 31 - 44, wherein the hardware processor is programmed to identify the presence of at least a partial occlusion of the reflection of the light source.
[0220] (Example 46) The wearable display system according to any one of Examples 31 - 45, wherein at least one other electronic hardware component is programmed to identify the location of the pupil center based at least in part on the location of the reflection of the light source and the flash - pupil relationship.
[0221] (Example 47) The wearable display system according to any one of Examples 18 - 20, wherein the second hardware processor is configured to identify a flash and provide flash position information regarding the flash to the first hardware processor.
[0222] (Example 48) The wearable display system according to Example 47, wherein the first hardware processor is configured to determine the line of sight from the flash position.
[0223] (Example 49) The wearable display system according to any one of Examples 18 - 20, wherein the second hardware processor is configured to identify a flash candidate and provide position information regarding the flash candidate to the first hardware processor.
[0224] (Example 50) The wearable display system according to Example 49, wherein the first hardware processor is configured to identify a subset of the flash candidates and use the subset of the flash candidates to perform one or more operations.
[0225] (Example 51) The wearable display system according to Example 50, wherein the first hardware processor is configured to determine a line of sight based on the flash candidate.
[0226] (Example 52) The wearable display system according to any one of Examples 31-46, wherein the hardware processor is programmed to identify a flash candidate and an associated location of the flash candidate and transmit location data indicating the location of the flash candidate to at least one other electronic hardware component.
[0227] (Example 53) The wearable display system according to Example 52, wherein the at least one other electronic hardware component is configured to identify a subset of the flash candidates and perform one or more operations using the subset of the flash candidates.
[0228] (Example 54) The wearable display system according to Example 53, wherein the at least one other electronic hardware component is configured to determine a line-of-sight direction of the wearer's eye from the subset of the flash candidates. (Part II)
[0229] (Example 1) A wearable display system, a head-mounted display configured to present virtual content by outputting light to the eyes of a wearer of the head-mounted display; at least one light source configured to direct light towards the eyes of the wearer; a first image of the wearer's eye captured at a first frame rate and a first exposure time; a second image of the wearer's eye captured at a second frame rate greater than the first frame rate and a second exposure time less than the first exposure time; at least one eye-tracking camera configured to capture; At least one hardware processor communicably coupled to a head-mounted display and at least one eye-tracking camera, analyzes a first image and determines the pupil center of the eye, analyzes a second image and determines the position of the reflection of the light source from the eye, determines the line-of-sight direction of the eye from the pupil center and the position of the reflection, estimates the future line-of-sight direction of the future eye at the future line-of-sight time from the line-of-sight direction and the previous line-of-sight direction data, causes the head-mounted display to present virtual content at the future line-of-sight time, at least in part, based on the future line-of-sight direction, and at least one hardware processor programmed to: A wearable display system comprising:
[0230] (Example 2) The wearable display system according to Example 1, wherein the at least one light source comprises at least one infrared light source.
[0231] (Example 3) The wearable display system according to Example 1 or Example 2, wherein the at least one hardware processor is programmed to analyze the first image or the second image and determine the center of rotation or the line-of-sight center of the wearer's eye.
[0232] (Example 4) The wearable display system according to any one of Examples 1-3, wherein, to analyze the second image, the at least one hardware processor is programmed to apply a threshold to the second image, identify non-maximum values within the second image, or suppress or remove non-maximum values within the second image.
[0233] (Example 5) To analyze the second image, at least one hardware processor is programmed to identify a search area for a reflection of a light source in the current image from the second image, based at least in part on the position of the reflection of the light source in the previous image from the second image, of the wearable display system according to any one of embodiments 1-4.
[0234] (Embodiment 6) To analyze the second image, at least one hardware processor
[0235] determines a common velocity of a plurality of reflections of at least one light source,
[0236] and determines whether the velocity of the reflection of at least one light source differs from the common velocity by more than a threshold amount, of the wearable display system according to any one of embodiments 1-5.
[0237] (Embodiment 7) To analyze the second image, at least one hardware processor is programmed to determine whether the reflection of the light source is from an aspherical portion of the eye's cornea, of the wearable display system according to any one of embodiments 1-6.
[0238] (Embodiment 8) To analyze the second image, at least one hardware processor is programmed to identify the presence of at least partial occlusion of the reflection of the light source, of the wearable display system according to any one of embodiments 1-7.
[0239] (Embodiment 9) At least one hardware processor is programmed to determine a position of an estimated pupil center, at least in part, based on a position of a reflection of a light source and a flash-pupil relationship, of the wearable display system according to any one of embodiments 1-8.
[0240] (Example 10) The flash-pupil relationship is the wearable display system according to Example 9, including a linear relationship between the flash position and the pupil center position.
[0241] (Example 11) The first exposure time is within the range of 200 μs to 1,200 μs, and the wearable display system according to any one of Examples 1-10.
[0242] (Example 12) The first frame rate is within the range of 10 frames / second to 60 frames / second, and the wearable display system according to any one of Examples 1-11.
[0243] (Example 13) The second exposure time is within the range of 5 μs to 100 μs, and the wearable display system according to any one of Examples 1-12.
[0244] (Example 14) The second frame rate is within the range of 100 frames / second to 1,000 frames / second, and the wearable display system according to any one of Examples 1-13.
[0245] (Example 15) The ratio of the first exposure time to the second exposure time is within the range of 5 to 50, and the wearable display system according to any one of Examples 1-14.
[0246] (Example 16) The ratio of the second frame rate to the first frame rate is within the range of 1 to 100, and the wearable display system according to any one of Examples 1-15.
[0247] (Example 17) The wearable display system according to any one of Embodiments 1-16, wherein the future gaze time is within the range of 5 ms to 100 ms.
[0248] (Embodiment 18) At least one hardware processor includes a first hardware processor disposed on a non-head-mounted component of the wearable display system, and a second hardware processor disposed within or on the head-mounted display, and is provided with The first hardware processor is used to analyze the first image, The second hardware processor is used to analyze the second image, The wearable display system according to any one of Embodiments 1-17.
[0249] (Embodiment 19) The second hardware processor includes or is associated with a memory buffer configured to store at least a portion of each of the second images, and the second hardware processor is programmed to delete at least a portion of each of the second images after determining the position of the reflection of the light source from the eye. The wearable display system according to Embodiment 18.
[0250] (Embodiment 20) The wearable display system according to any one of Embodiments 1-19, wherein at least one hardware processor is programmed not to combine the first image and the second image.
[0251] (Embodiment 21) A method for eye tracking, comprising: capturing a first image of the eye at a first frame rate and a first exposure time via an eye tracking camera; Capturing a second image of the eye at a second frame rate that is greater than or equal to a first frame rate and at a second exposure time that is less than a first exposure time via an eye tracking camera; Determining the pupil center of the eye from at least the first image; Determining the position of the reflection of the light source from the eye from at least the second image; Determining the line-of-sight direction of the eye from the pupil center and the position of the reflection; A method comprising the above steps.
[0252] (Example 22) The method according to Example 21, further comprising rendering virtual content on the display at least partially based on the line-of-sight direction.
[0253] (Example 23) The method according to Example 21 or 22, further comprising estimating a future line-of-sight direction at a future line-of-sight time based at least partially on the line-of-sight direction and previous line-of-sight direction data.
[0254] (Example 24) The method according to Example 23, further comprising rendering virtual content on the display at least partially based on the future line-of-sight direction.
[0255] (Example 25) The method according to any one of Examples 21 - 24, wherein the first exposure time is in the range of 200 μs to 1,200 μs.
[0256] (Example 26) The method according to any one of Examples 21 - 25, wherein the first frame rate is in the range of 10 frames per second to 60 frames per second.
[0257] (Example 27) The method according to any one of Examples 21 - 26, wherein the second exposure time is in the range of 5 μs to 100 μs.
[0258] (Example 28) The method according to any one of Examples 21-27, wherein the second frame rate is in the range of 100 frames / second to 1,000 frames / second.
[0259] (Example 29) The method according to any one of Examples 21-28, wherein the ratio of the first exposure time to the second exposure time is in the range of 5 to 50.
[0260] (Example 30) The method according to any one of Examples 21-29, wherein the ratio of the second frame rate to the first frame rate is in the range of 1 to 100.
[0261] (Example 31) A wearable display system, A head-mounted display configured to present virtual content by outputting light to the eyes of a wearer of the head-mounted display, A light source configured to direct light toward the eyes of the wearer, An eye-tracking camera configured to capture an image of the wearer's eyes, and alternately, Capturing a first image at a first exposure time, Capturing a second image at a second exposure time less than the first exposure time, An eye-tracking camera configured as such, A plurality of electronic hardware components, at least one of which comprises a hardware processor communicatively coupled to the head-mounted display, the eye-tracking camera, and at least one other electronic hardware component within the plurality of electronic hardware components, the hardware processor being Receiving each first image of the wearer's eyes captured at the first exposure time by the eye-tracking camera and relaying it to at least one other electronic hardware component, Receiving the pixels of each second image of the wearer's eyes captured at the second exposure time by the eye-tracking camera, storing them in a buffer, Analyzing the pixels stored in the buffer to identify the location where the reflection of the light source exists within each second image of the wearer's eyes captured at the second exposure time by the eye-tracking camera, Transmitting location data indicating the location to at least one other electronic hardware component, A plurality of electronic hardware components programmed to: Comprising, At least one other electronic hardware component: Analyzing each first image of the wearer's eyes captured at the first exposure time by the eye-tracking camera to identify the location of the center of the pupil of the eye, Determining the line-of-sight direction of the wearer's eyes from the location of the center of the pupil and the location data received from the hardware processor, A wearable display system configured to:
[0262] (Example 32) The wearable display system according to Example 31, wherein the first exposure time is within a range of 200 μs to 1,200 μs.
[0263] (Example 33) The wearable display system according to Example 31 or Example 32, wherein the second exposure time is within a range of 5 μs to 100 μs.
[0264] (Example 34) The wearable display system according to any one of Examples 31 - 33, wherein the eye-tracking camera is configured to capture the first image at a first frame rate within a range of 10 frames per second to 60 frames per second.
[0265] (Example 35) The eye-tracking camera is configured to capture a second image at a second frame rate within a range of 100 frames per second to 1,000 frames per second, of the wearable display system according to any one of Examples 31-34.
[0266] (Example 36) The ratio of the first exposure time to the second exposure time is within a range of 5 to 50, of the wearable display system according to any one of Examples 31-35.
[0267] (Example 37) At least one other electronic hardware component is disposed within or on a non-head-mounted component of the wearable display system, of the wearable display system according to any one of Examples 31-36.
[0268] (Example 38) The non-head-mounted component includes a belt pack, of the wearable display system according to Example 37.
[0269] (Example 39) Each pixel of each second image of the eye includes fewer pixels than all of each second image, of the wearable display system according to any one of Examples 31-38.
[0270] (Example 40) The pixels include an array of n×m pixels, where n and m are each integers within a range of 1 to 20, of the wearable display system according to any one of Examples 31-39.
[0271] (Example 41) The plurality of electronic hardware components further estimate a future eye gaze direction at a future gaze time from the gaze direction and previous gaze direction data, The head-mounted display is configured to present virtual content at least partially based on a future line-of-sight direction at a future line-of-sight time. A wearable display system according to any one of Examples 31-40, configured as such.
[0272] (Example 42) The hardware processor is programmed to apply a threshold to pixels stored in a buffer, identify non-maximum values within pixels stored in the buffer, or suppress or remove non-maximum values within pixels stored in the buffer. A wearable display system according to any one of Examples 31-41.
[0273] (Example 43) The hardware processor determines a common speed of reflection of a light source, and determines whether the speed of reflection of the light source exceeds a threshold amount and differs from the common speed. A wearable display system according to any one of Examples 31-42, programmed as such.
[0274] (Example 44) The hardware processor is programmed to determine whether the location of reflection of a light source is from a non-spherical portion of the eye's cornea. A wearable display system according to any one of Examples 31-43.
[0275] (Example 45) The hardware processor is programmed to identify the presence of at least partial occlusion of reflection of a light source. A wearable display system according to any one of Examples 31-44.
[0276] (Example 46) At least one other electronic hardware component is programmed to identify the location of the pupil center, at least in part, based on the location of reflection of the light source and the flash-pupil relationship, of the wearable display system according to any one of embodiments 31-45.
[0277] (Example 47) A second hardware processor is configured to identify a flash and provide flash position information regarding the flash to a first hardware processor, of the wearable display system according to any one of embodiments 18-20.
[0278] (Example 48) The first hardware processor is configured to determine a line of sight from the flash position, of the wearable display system according to Example 47.
[0279] (Example 49) A second hardware processor is configured to identify a flash candidate and provide position information regarding the flash candidate to a first hardware processor, of the wearable display system according to any one of embodiments 18-20.
[0280] (Example 50) The first hardware processor is configured to identify a subset of the flash candidates and use the subset of the flash candidates to perform one or more operations, of the wearable display system according to Example 49.
[0281] (Example 51) The first hardware processor is configured to determine a line of sight based on the flash candidate, of the wearable display system according to Example 50.
[0282] (Example 52) The wearable display system according to any one of embodiments 31 - 46, wherein the hardware processor is programmed to identify a flash candidate and an associated location of the flash candidate, and transmit location data indicating the location of the flash candidate to at least one other electronic hardware component.
[0283] (Embodiment 53) The wearable display system according to embodiment 52, wherein the at least one other electronic hardware component is configured to identify a subset of the flash candidates and perform one or more operations using the subset of the flash candidates.
[0284] (Embodiment 54) The wearable display system according to embodiment 53, wherein the at least one other electronic hardware component is configured to determine a line of sight direction of the wearer's eye from the subset of the flash candidates.
[0285] (Embodiment 55) The wearable display system according to any one of embodiments 1 - 20 and 47 - 51, wherein the at least one eye - tracking camera includes one eye - tracking camera.
[0286] (Embodiment 56) The wearable display system according to any one of embodiments 1 - 20, 47 - 51, and 55, wherein the at least one light source includes a plurality of light sources. (Additional Considerations)
[0287] The processes, methods, and algorithms described in this specification and / or depicted in the accompanying figures can each be embodied in one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, thereby being fully or partially automated. For example, a computing system can include a general-purpose computer (e.g., a server) or a dedicated computer, a dedicated circuit, etc., programmed with specific computer instructions. The code modules can be installed in a dynamic link library that can be compiled and linked into an executable program, or can be written in an interpreted type programming language. In some implementations, certain operations and methods can be implemented by circuits specific to a given function.
[0288] Furthermore, because the functional implementations of the present disclosure are sufficiently mathematically, computationally, or technically complex, special-purpose hardware or one or more physical computing devices (utilizing appropriate specialized executable instructions) may be required to implement the functionality, for example, due to the amount or complexity of the calculations involved or to provide the results in substantially real time. For example, a video or a video clip can contain many frames, each frame can have millions of pixels, and specifically programmed computer hardware is required to process the video data to provide the desired image processing tasks or applications within a commercially reasonable amount of time. Additionally, real-time eye tracking for AR, MR, VR wearable devices is computationally difficult, and multiple exposure time eye tracking techniques may utilize an efficient CPU, GPU, ASIC, or FPGA.
[0289] A code module or any type of data can be stored on any type of non-transitory computer-readable medium, such as a physical computer storage device including a hard drive, solid state memory, random access memory (RAM), read only memory (ROM), optical disk, volatile or non-volatile storage device, combinations of the same, and / or equivalents. The methods and modules (or data) can also be transmitted as data signals generated on various computer-readable transmission media, including wireless-based and wired / cable-based media (e.g., as part of a carrier wave or other analog or digital propagated signal), and can take various forms (e.g., as part of a single or multiplexed analog signal or as multiple discrete digital packets or frames). The results of the disclosed process or process steps can be persistently or otherwise stored within any type of non-transitory tangible computer storage device or communicated via a computer-readable transmission medium.
[0290] Any process, block, state, step, or functionality in a flowchart described in and / or depicted in the accompanying figures herein is to be understood as potentially representing a code module, segment, or portion of code that includes one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in a process. The various processes, blocks, states, steps, or functionality can be combined, rearranged, added to the exemplary embodiments provided herein, deleted therefrom, modified, or otherwise changed therefrom. In some embodiments, additional or different computing systems or code modules can implement some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the associated blocks, steps, or states can be performed in other sequences that are appropriate, e.g., sequentially, in parallel, or in some other manner. Tasks or events can be added to or removed from the disclosed exemplary embodiments. Further, the separation of the various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems can generally be integrated together in a single computer product or packaged in multiple computer products. Many implementation variations are possible.
[0291] The present process, method, and system can be implemented in a network (or distributed) computing environment. The network environment can include an enterprise-wide computer network, an intranet, a local area network (LAN), a wide area network (WAN), a personal area network (PAN), a cloud computing network, a cloud source computing network, the Internet, and the World Wide Web. The network can be a wired or wireless network or any other type of communication network.
[0292] The systems and methods of the present disclosure each have several innovative aspects, none of which alone contribute to or are required for the desirable attributes disclosed herein. The various features and processes described above can be used independently of one another or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. Various modifications of the implementations described in this disclosure may be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other implementations without departing from the spirit or scope of the present disclosure. Accordingly, the claims are not intended to be limited to the implementations shown herein but are to be accorded the widest scope consistent with the present disclosure, the principles, and the novel features disclosed herein.
[0293] In the context of separate implementations, certain features described herein can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented separately in multiple implementations or in any suitable sub-combination. Further, features may be described above as acting in a certain combination and may further be claimed as such, but one or more features from the claimed combination can, in some cases, be deleted from the combination and the claimed combination can be directed to a sub-combination or a variation of a sub-combination. No single feature or group of features is necessary or essential to every embodiment.
[0294] In particular, conditional statements used herein such as "can", "could", "might", "may", "e.g.", and equivalents, generally convey that while one embodiment includes a certain feature, element, or step, another embodiment does not, unless specifically stated otherwise or understood otherwise within the context in which they are used. Thus, such conditional statements are not generally intended to imply that a feature, element, and / or step is required in any way for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps are to be included or implemented in any particular embodiment, regardless of the author's input or prompting. The terms "comprising", "including", "having", and equivalents are synonyms and are used inclusively in a non-limiting manner, excluding additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not in its exclusive sense), so that, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Additionally, the articles "a", "an", and "the" as used in this application and the appended claims should be construed to mean "one or more" or "at least one" unless otherwise defined.
[0295] As used herein, the phrase that refers to a list of items "at least one of" refers to any combination of those items, including a single element. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Connective phrases such as "at least one of X, Y, and Z" are generally understood in a context such that they are used to convey that an item, term, etc. can be at least one of X, Y, or Z, unless specifically stated otherwise. Thus, such connective phrases are generally not intended to suggest that an embodiment requires that at least one of each of X, at least one of Y, and at least one of Z be present respectively.
[0296] Similarly, operations may be depicted in the drawings in a particular order, but it should be recognized that this is not necessary for achieving the desired result, and that such operations may be performed in the particular order in which they are shown, or in a sequential order, or that all of the illustrated operations need not be performed. Additionally, the drawings may schematically depict one or more exemplary processes in the form of a flowchart. However, other operations not depicted may also be incorporated into the exemplary methods and processes schematically illustrated. For example, one or more additional operations may be performed before, after, concurrently with, or during any of the illustrated operations. In addition, the operations may be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing may be advantageous. Further, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the program components and systems described may generally be integrated together in a single software product or packaged into multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result.
Claims
1. 1. A method for eye tracking, the method comprising: capturing a first image of the eye via an eye tracking camera at a first frame rate and a first exposure time; capturing a second image of the eye via an eye tracking camera at a second frame rate greater than or equal to the first frame rate and at a second exposure time less than the first exposure time, wherein the second image is captured before the first image is captured or the second image is captured after the first image is captured; determining a pupil center of the eye from at least the first image; determining a location of a reflection of a light source from the eye from at least the second image; and determining a gaze direction of the eye from the pupil center and the position of the reflection; Including, The method wherein the second image is captured every 1 ms to 10 ms.
2. The method of claim 1 , further comprising causing a display to render virtual content based at least in part on the line of sight direction.
3. The method of claim 1 or claim 2, further comprising estimating a future gaze direction at a future gaze time based at least in part on the gaze direction and previous gaze direction data.
4. The method of claim 3 , further comprising causing a display to render virtual content based at least in part on the future gaze direction.
5. The method according to any one of claims 1 to 4, wherein the first exposure time is in the range of 200 μs to 1,200 μs.
6. The method of any one of claims 1 to 5, wherein the first frame rate is in the range of 10 frames / second to 60 frames / second.
7. The method according to any one of claims 1 to 6, wherein the second exposure time is in the range of 5 μs to 100 μs.
8. The method of any one of claims 1 to 7, wherein the ratio of the first exposure time to the second exposure time is in the range of 5 to 50.
9. The method of any one of claims 1 to 8, wherein the ratio of the second frame rate to the first frame rate is in the range 1-100.
Citation Information
Patent Citations
Visual axis detector having integrating time control means
JP1992242627A
Sight line detection device and sight line detection method
JP2017111746A
Line-of-sight detection computer program, line-of-sight detection device and line-of-sight detection method
JP2018197974A
Gaze Point Tracking Using Polarized Light
US20110170061A1
Dual LED usage for glint detection
US8971570B1