Systems and methods for operating a head-mounted display system based on user identity
By identifying users through iris scanning and authentication, and adjusting the wavefront divergence of the image output from the display, the problem of uncalibrated depth plane selection for users in virtual reality and augmented reality technologies is solved, thereby improving the comfort and adaptability of the user experience.
Patent Information
- Application Number
- CN202080095898.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-09
- Filing Date
- 2020-12-09
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2040-12-09
AI Technical Summary
Existing virtual reality, augmented reality, and mixed reality technologies suffer from insufficient comfort and adaptability in terms of user identification and depth plane selection, especially for uncalibrated users.
User identity is identified through iris scanning or authentication, and the wavefront divergence of the image output to the display is adjusted based on the identification results to achieve depth plane selection and user calibration, including different processing for registered and unregistered users.
It improves the comfort and adaptability of the user experience, ensures the depth perception realism and stability of virtual content, and is especially friendly to uncalibrated users.
Smart Images

Figure CN115053270B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 62 / 945,517, filed December 9, 2019, entitled “SYSTEMS AND METHODS FOR OPERATING A HEAD-MOUNTED DISPLAY SYSTEM BASED ON USER IDENTITY,” the entire contents of which are incorporated herein by reference.
[0003] Incorporation by Reference
[0004] The present application incorporates by reference the entirety of the following patent applications and patent publications: U.S. Application No. 4 / 555,585, filed November 27, 2014, published as U.S. Publication No. 2015 / 0205126 on July 23, 2015; U.S. Application No. 14 / 690,401, filed April 18, 2015, published as U.S. Publication No. 2015 / 0302652 on October 22, 2015; U.S. Application No. 14 / 212,961, filed March 14, 2014, now U.S. Patent No. 9,417,452, issued August 16, 2016; U.S. Application No. 14 / 331,218, filed July 14, 2014, published as U.S. Publication No. 2015 / 0309263 on October 29, 2015; U.S. Application No. 15 / 927,808, filed March 21, 2018; U.S. Application No. 15 / 291,929, filed October 12, 2016, published as U.S. Publication No. 2017 / 0109580 on April 20, 2017; U.S. Application No. 15 / 408,197, filed January 17, 2017, published as U.S. Publication No. 2017 / 0206412 on July 20, 2017; U.S. Application No. 15 / 469369, filed March 24, 2017, published as U.S. Publication No. 2017 / 0276948 on September 28, 2017; U.S. Provisional Application No. 62 / 618,559, filed January 17, 2018; U.S. Application No. 16 / 250,931, filed January 17, 2019; U.S. Application No. 14 / 705 / 741, filed May 6, 2015, published as U.S. Publication No. 2016 / 0110920 on April 21, 2016; U.S. Patent Publication No. 2017 / 0293145, published October 12, 2017; U.S. Patent Application No. 15 / 993,371, filed May 30, 2018, published as U.S. Publication No. 2018 / 0348861 on December 6, 2018; U.S. Patent Application No. 15 / 155,013, filed May 14, 2016, published as U.S. Publication No. 2016 / 0358181 on December 8, 2016; U.S. Patent Application No. 15 / 934,941, filed March 23, 2018, published as U.S. Publication No. 2018 / 0276467 on September 27, 2018; U.S. Provisional Application No. 62 / 644,321, filed March 16, 2018; U.S. Application No. 16 / 251,017, filed January 17, 2019; U.S. Provisional Application No. 62 / 702,866, filed July 24, 2018; International Patent Application PCT / US2019 / 043096, filed July 23, 2019; U.S. Provisional Application No. 62 / 714,649, filed August 3, 2018;U.S. Provisional Application No. 62 / 875,474, filed July 17, 2019; and U.S. Application No. 16 / 530,904, filed August 2, 2019. TECHNICAL FIELD
[0005] The present disclosure relates to display systems, virtual reality and augmented reality imaging and visualization systems, and more specifically to depth plane selection based in part on a user’s identity. BACKGROUND
[0006] Modern computing and display technologies have facilitated the development of systems for so called “virtual reality” or “augmented reality” experiences, wherein digitally reproduced images or portions thereof are presented to a user in a manner wherein they seem to be, or can be perceived as, real. A virtual reality, or “VR,” scenario typically involves presentation of digital or virtual image information without transparency to other actual real-world visual input; an augmented reality, or “AR,” scenario typically involves presentation of digital or virtual image information in a manner wherein other actual real-world visual input is at least somewhat perceptible.
[0007] The human visual perception system is quite complex, and producing VR, AR or MR technology that facilitates comfortable, natural-feeling presence in virtual environments that are also rich in sensory detail is challenging. SUMMARY
[0008] Various examples of systems and methods for depth plane selection in a display system, such as an augmented reality display system, are disclosed, wherein the display system comprises a mixed reality display system.
[0009] A wearable system can include one or more displays, one or more cameras, and one or more processors. The one or more displays can be configured to present virtual image content to one or both eyes of a user via image light, where the one or more displays are configured to output the image light to the one or both eyes of the user having different amounts of wavefront divergence corresponding to different depth planes at different distances from the user. The one or more cameras can be configured to capture images of the one or both eyes of the user. The one or more processors can be configured to: obtain one or more images of the one or both eyes of the user captured by the one or more cameras; generate an indication based on the obtained one or more images of the one or both eyes of the user, the generated indication indicating whether the user is identified; and control the one or more displays to output the image light to the one or both eyes of the user having the different amounts of wavefront divergence based at least in part on the generated indication of whether the user is identified.
[0010] In some system embodiments, generating the indication indicating whether the user is identified can include performing one or more iris scanning operations or one or more iris authentication operations.
[0011] In some system embodiments, the generated indication indicates that the user is identified as a registered user, and in response to the generated indication indicating that the user is identified as a registered user, the one or more processors can be configured to perform controlling the one or more displays to output image light to the one or both eyes of the user includes controlling the one or more displays to output the image light to the one or both eyes of the user having the different amounts of wavefront divergence based at least in part on settings associated with the registered user.
[0012] In some system embodiments, the generated indication indicates that the user is not identified as a registered user, and in response to the generated indication indicating that the user is not identified as a registered user, the one or more processors can be configured to perform controlling the one or more displays to output the image light to the one or both eyes of the user includes controlling the one or more displays to output the image light to the one or both eyes of the user having the different amounts of wavefront divergence based at least in part on a set of default settings.
[0013] In some system embodiments, the generated indication indicates that the user is not identified as a registered user, and in response to the generated indication indicating that the user is not identified as a registered user, the one or more processors can be configured to: perform a set of one or more operations to enroll the user to iris authentication or generate a calibration profile for the user; and determine settings for the user based at least in part on information obtained by performing the set of one or more operations. Controlling the one or more displays to output image light to one or both eyes of the user can include controlling the one or more displays to output the image light to one or both eyes of the user, the image light having the different amounts of wavefront divergence based at least in part on the determined settings for the user.
[0014] In some system embodiments, the one or more processors can be configured to: control the one or more displays to present a virtual target to the user. Obtaining one or more images of one or both eyes of the user captured by the one or more cameras can include: obtaining the one or more images of one or both eyes of the user captured by the one or more cameras while presenting the virtual target to the user.
[0015] In some system embodiments, the one or more processors can be configured to: perform one or more operations to attempt to improve accuracy or reliability of user identification in generating an indication of whether the user is identified.
[0016] A method can include presenting, by one or more displays, virtual image content to one or both eyes of a user via image light, where the one or more displays are configured to output the image light to one or both eyes of the user, the image light having different amounts of wavefront divergence corresponding to different depth planes at different distances from the user. The method can include capturing, by one or more cameras, images of one or both eyes of the user; and obtaining one or more images of one or both eyes of the user captured by the one or more cameras. The method can include generating an indication based on the obtained one or more images of one or both eyes of the user, the generated indication indicating whether the user is identified; and controlling the one or more displays to output the image light to one or both eyes of the user, the image light having different amounts of wavefront divergence based at least in part on the generated indication of whether the user is identified.
[0017] In some method embodiments, generating an indication of whether the user is identified can include: performing one or more iris scanning operations or one or more iris authentication operations.
[0018] In some method embodiments, the generated indication indicates that the user is identified as a registered user. The method can include, in response to the generated indication indicating that the user is identified as a registered user, performing controlling the one or more displays to output image light to one or both eyes of the user that includes controlling the one or more displays to output the image light to one or both eyes of the user having the different amounts of wavefront divergence based at least in part on settings associated with the registered user.
[0019] In some method embodiments, the generated indication indicates that the user is not identified as a registered user. The method can include, in response to the generated indication indicating that the user is not identified as a registered user, performing controlling the one or more displays to output the image light to one or both eyes of the user that includes controlling the one or more displays to output the image light to one or both eyes of the user having the different amounts of wavefront divergence based at least in part on a set of default settings.
[0020] In some method embodiments, the generated indication indicates that the user is not identified as a registered user. The method can include, in response to the generated indication indicating that the user is not identified as a registered user, performing a set of one or more operations to enroll the user to iris authentication or generate a calibration profile for the user; and determining settings for the user based at least in part on information obtained by performing the set of one or more operations. Controlling the one or more displays to output image light to one or both eyes of the user can include controlling the one or more displays to output the image light to one or both eyes of the user having the different amounts of wavefront divergence based at least in part on the determined settings for the user.
[0021] In some method embodiments, the method can include controlling the one or more displays to present a virtual target to the user. Obtaining one or more images of one or both eyes of the user captured by the one or more cameras can include obtaining one or more images of one or both eyes of the user captured by the one or more cameras while presenting the virtual target to the user.
[0022] In some method embodiments, the method can include, in generating the indication of whether the user is identified, performing one or more operations to attempt to improve accuracy or reliability of user identification.
[0023] A non-transitory computer-readable medium that can store instructions that, when executed by one or more processors of a wearable system, cause the wearable system to perform a method. The method can include presenting, by one or more displays of the wearable system, virtual image content to one or both eyes of a user via image light, wherein the one or more displays are configured to output the image light to the one or both eyes of the user having different amounts of wavefront divergence corresponding to different depth planes at different distances from the user. The method can include capturing, by one or more cameras of the wearable system, images of the one or both eyes of the user; and obtaining one or more images of the one or both eyes of the user captured by the one or more cameras. The method can include generating an indication based on the obtained one or more images of the one or both eyes of the user, the generated indication indicating whether the user is identified; and controlling the one or more displays to output the image light to the one or both eyes of the user having different amounts of wavefront divergence based at least in part on the generated indication of whether the user is identified.
[0024] In some non-transitory computer-readable medium embodiments, generating the indication of whether the user is identified can include performing one or more iris scanning operations or one or more iris authentication operations.
[0025] In some non-transitory computer-readable medium embodiments, the generated indication indicates that the user is identified as a registered user. The method can include, in response to the generated indication indicating that the user is identified as a registered user, performing the controlling the one or more displays to output image light to the one or both eyes of the user includes controlling the one or more displays to output the image light to the one or both eyes of the user having the different amounts of wavefront divergence based at least in part on settings associated with the registered user.
[0026] In some non-transitory computer-readable medium embodiments, the generated indication indicates that the user is not identified as a registered user. The method can include, in response to the generated indication indicating that the user is not identified as a registered user, performing the controlling the one or more displays to output the image light to the one or both eyes of the user includes controlling the one or more displays to output the image light to the one or both eyes of the user having the different amounts of wavefront divergence based at least in part on a set of default settings.
[0027] In some non-transitory computer-readable medium embodiments, the generated indication indicates that the user is not identified as a registered user. The method can include: responsive to the generated indication indicating that the user is not identified as a registered user, performing a set of one or more operations to enroll the user to iris authentication or generate a calibration profile for the user; and determining settings for the user based at least in part on information obtained by performing the set of one or more operations. Controlling the one or more displays to output image light to one or both eyes of the user can include controlling the one or more displays to output the image light to one or both eyes of the user, the image light having the different amounts of wavefront divergence based at least in part on the determined settings for the user.
[0028] In some non-transitory computer-readable medium embodiments, the method can include controlling the one or more displays to present a virtual target to the user. Obtaining one or more images of one or both eyes of the user captured by the one or more cameras can include obtaining one or more images of one or both eyes of the user captured by the one or more cameras while presenting the virtual target to the user. BRIEF DESCRIPTION OF DRAWINGS
[0029] FIG. 1 A diagram depicts a mixed reality scene with certain virtual reality objects and certain physical objects viewed by a person.
[0030] FIG. 2 An example of a wearable system is shown schematically.
[0031] FIG. 3 Example components of a wearable system are shown schematically.
[0032] FIG. 4 An example of a waveguide stack of a wearable device for outputting image information to a user is shown schematically.
[0033] FIG. 5 An example of an eye and an example coordinate system for determining an eye pose of the eye are shown schematically.
[0034] FIG. 6 A diagram of a wearable system including an eye tracking system.
[0035] FIG. 7A A block diagram of a wearable system that can include an eye tracking system.
[0036] FIG. 7B A block diagram of a rendering controller in a wearable system.
[0037] FIG. 7Cis a block diagram of a registration observer in a head-mounted display system.
[0038] FIG. 8A is a schematic diagram of an eye showing a corneal sphere of the eye.
[0039] FIG. 8B shows an example corneal glint detected by an eye tracking camera.
[0040] FIG. 8C-8E shows example stages of locating a corneal center of a user by an eye tracking module in a wearable system.
[0041] FIG. 9A-9C shows example normalization of a coordinate system of an eye tracking image.
[0042] FIG. 9D-9G shows example stages of locating a pupil center of a user by an eye tracking module in a wearable system.
[0043] FIG. 10 shows an example of an eye including an optical axis and a visual axis of the eye and a center of rotation of the eye.
[0044] FIG. 11 is a process flow diagram of an example of a method for using eye tracking in rendering content and providing feedback about registration in a wearable device.
[0045] FIG. 12A and FIG. 12B shows a nominal position of a display element relative to an eye of a user and shows a coordinate system used to describe the position of the display element and the eye of the user relative to each other.
[0046] FIG. 13 is a set of example plots showing how a wearable system switches depth planes in response to eye movements of a user.
[0047] FIG. 14 is a process flow diagram of an example of a method for selecting a depth plane using an existing calibration, dynamic calibration, and / or content-based switching scheme.
[0048] FIG. 15 is a process flow diagram of an example of a method for selecting a depth plane based at least in part on an interpupillary distance of a user.
[0049] FIG. 16A shows a top-down view of a representation of a user viewing content presented by a display system configured to switch depth planes by detecting a user gaze within one of a plurality of regions partitioning a field of view of the user along a horizontal axis.
[0050] FIG. 16B showsFIG. 16A perspective view of a representation.
[0051] FIG. 17A A top-down view of a representation of a user viewing content presented by a display system configured to switch depth planes by detecting a user gaze within discrete marker volumes within a display frustum is shown.
[0052] FIG. 17B A perspective view of a representation. FIG. 16A
[0053] A flowchart showing an example process of selecting depth planes according to content-based switching is shown. FIG. 18
[0054] A flowchart showing another example process of adjusting regions according to content-based switching is shown. FIG. 19
[0055] A flowchart showing an example process for operating a head-mounted display system based on a user identity is shown. FIG. 20 In all of the drawings, reference numbers can be repeated to indicate corresponding or analogous elements throughout the several figures. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the disclosure.
[0056] DETAILED DESCRIPTION
[0057] As described herein, a display system (e.g., an augmented reality or virtual reality display system) can render virtual content to be presented to a user at different perceived depths from the user. In an augmented reality display system, the virtual content can be projected with different depth planes, where each depth plane is associated with a particular perceived depth from the user. For example, a stack of waveguides can be utilized that are configured to output light having different wavefront divergence, where each depth plane has a corresponding wavefront divergence and is associated with at least one waveguide. As the virtual content moves around the field of view of the user, the virtual content can be adjusted along three discrete axes. For example, the virtual content can be adjusted along the X, Y, and Z axes so that it can be presented at different perceived depths from the user. The display system can switch between depth planes as if the virtual content is moving further away or closer to the user in perception. It will be understood that switching depth planes can involve changing the wavefront divergence of the light forming the virtual content in discrete steps. In a waveguide-based system, such depth plane switching can involve switching waveguides that output light to form the virtual content, in some embodiments.
[0058] In some embodiments, the display system can be configured to monitor the gaze of the user's eyes and determine a three-dimensional gaze point at which the user is looking. The gaze point can be determined based on the distance between the user's eyes and the gaze direction of each eye, among other things. It will be understood that these variables can be understood to form a triangle, with the gaze point at one corner of the triangle and the eyes at the other corners. It will also be understood that a calibration can be performed to accurately track the orientation of the user's eyes and determine or estimate the gaze of those eyes to determine the gaze point. Thus, after a full calibration is performed, the display device can have a calibration file or calibration information for the primary user of the device. The primary user is also referred to herein as a calibrated user. More details regarding calibration and eye tracking can be found, for example, in U.S. Patent Application No. 15 / 993,371 entitled "EYE TRACKING CALIBRATION TECHNIQUES," the entirety of which is incorporated by reference herein.
[0059] The display system can occasionally be used by a visitor user who has not completed a full calibration. Furthermore, these visitor users can not have the time or desire to perform a full calibration. However, if the display system does not actually track the gaze point of the visitor user, then the depth plane switching can not be suitable for providing a realistic and comfortable viewing experience for the user.
[0060] Accordingly, it will be understood that the current user of the display system can be classified as a calibrated user or a visitor user. In some embodiments, the display system can be configured to classify the current user by performing an identification or authentication process, for example, to determine whether information provided by and / or obtained from the current user matches information associated with a calibrated user, to determine whether the current user is a calibrated user. For example, the authentication or identification process can be one or more of the following: asking for and verifying a username and / or password, performing an iris scan (e.g., by comparing a current image of the user's iris to a reference image), performing voice recognition (e.g., by comparing a current sample of the user's voice to a reference voice file), and IPD matching. In some embodiments, the IPD matching can include determining whether the IPD of the current user matches the IPD of a calibrated user. In some embodiments, if there is a match, then the current user can be assumed to be a calibrated user. In some embodiments, if there is no match, then the current user can be assumed to not be a calibrated user (e.g., to be a visitor user). In some other embodiments, multiple authentication processes can be performed for the current user to improve the accuracy of determining whether the current user is a calibrated user.
[0061] In some embodiments, if it is determined that the current user is not a calibrated user (e.g., a guest user), the display system can use content-based depth plane switching; for example, rather than determining a gaze point of the user’s eye, based on the depth plane in which the gaze point lies, the display system can be configured to display content having an appropriate amount of wavefront divergence for a location in 3D space at which content (e.g., a virtual object) is designated to be placed. It will be understood that a virtual object can have an associated location or coordinate in a three-dimensional volume surrounding the user, and the display system can be configured to present the object in that three-dimensional volume relative to the user using light having an amount of wavefront divergence appropriate for the depth of the object.
[0062] In the case where multiple virtual objects at different depths are to be displayed, content-based depth plane switching can involve a coarse determination of which virtual object the gaze is on, and then using the location of that virtual object to establish the plane to which the depth plane switch should switch; for example, in the case where it is determined that the user’s gaze is approximately on a particular virtual object, the display system can be configured to output light having a wavefront divergence corresponding to the depth plane associated with that virtual object. In some embodiments, such a coarse determination of whether the user’s gaze is on an object can involve determining whether the gaze point is within a display system-defined volume that is unique to that object, and if so, switching to the depth plane associated with that object regardless of the determined depth of the gaze point (so if the volume spans multiple depth planes, the display system will switch to the depth plane associated with the object).
[0063] In some embodiments, if it is determined that the current user is not a calibrated user (e.g., a guest user), the display system can perform a coarse calibration by measuring the inter-pupillary distance (IPD) of the guest user. This coarse calibration can also be referred to as dynamic calibration. Preferably, the IPD is measured when the user’s eyes are pointed at or focused on an object at optical infinity. In some embodiments, this IPD value can be understood to be a maximum IPD value. The maximum IPD value can be used, or a smaller value selected within a distribution of sampled values (e.g., a value at the 95th percentile of sampled IPD values) can be used as a reference value for determining gaze points. For example, this IPD value can constitute one leg (e.g., the base leg) of an imaginary triangle in which the gaze point forms a corner (e.g., the apex).
[0064] Accordingly, in some embodiments, the display system can be configured to monitor whether a user is a primary user or a guest user. If the user is a primary user, a calibration file can be retrieved. If the user is a guest user, content-based depth plane switching can be used, and / or a coarse calibration involving determining an IPD can be performed to establish a reference IPD value. The display system can be configured to use the reference IPD value to determine or estimate a gaze point of the guest user, and thus to decide when to switch depth planes (e.g., when to switch wavefront divergence of light used to form virtual content).
[0065] In some embodiments, the display system can transition from performing content-based depth plane switching to performing depth plane switching based on dynamic calibration of a current user. For example, such a transition can occur if data obtained in connection with the content-based depth plane switching scheme is deemed unreliable (e.g., values of virtual objects at which the user gazes generally have a high degree of uncertainty or variability), or if content is provided across different depth ranges of multiple depth planes (e.g., in which virtual content spans more than a threshold number of depth planes).
[0066] Reference will now be made to the drawings wherein like numerals refer to like components throughout. The drawings are schematic and are not necessarily drawn to scale unless otherwise specified.
[0067] Examples of 3D displays of wearable systems
[0068] A wearable system (also referred to herein as an augmented reality (AR) system) can be configured to present 2D or 3D virtual images to a user. The images can be still images, frames of a video, or a video, or a combination, etc. At least a portion of the wearable system can be implemented on a wearable device that can individually or in combination present a VR, AR, or MR environment for user interaction. The wearable device can interchangeably be used as an AR device (ARD). Further, for the purposes of the present disclosure, the term "AR" is used interchangeably with the term "MR."
[0069] FIG. 1 An illustration depicting a mixed reality scene with certain virtual reality objects and certain physical objects viewed by a person. In FIG. 1 In the middle, an MR scene 100 is depicted, in which a user of the MR technology sees a real-world park-like setting 110 featuring people, trees, buildings, and a concrete platform 120 in the background. In addition to these items, the user of the MR technology also perceives that he "sees" a robot statue 130 standing upon the real-world platform 120, as well as a flying cartoon-like avatar character 140, which seems to be an avatar of a bumble bee, even though these elements do not exist in the real world.
[0070] To produce a realistic sense of depth, and more specifically, a simulated sense of surface depth, for a 3D display, it can be desirable for each point in the field of view of the display to generate an accommodation response corresponding to its virtual depth. If the accommodation response to a display point does not conform to the virtual depth of that point (as dictated by binocular depth convergence cues and stereopsis), the human eye can experience an accommodation conflict, resulting in unstable imagery, harmful eye strain, headaches, and a nearly complete lack of surface depth in the absence of accommodation cues.
[0071] VR, AR, and MR experiences can be provided by a display system having a display in which images corresponding to multiple depth planes are provided to a viewer. The images can be different for each depth plane (e.g., providing slightly different presentations of a scene or object) and can be respectively focused by the viewer’s eyes, thereby helping to provide depth cues to the user based on the accommodation required by the eyes to focus different image features based on viewing different depth planes. Such depth cues provide a credible perception of depth, as discussed elsewhere herein.
[0072] FIG. 2 An example of a wearable system 200 that can be configured to provide AR / VR / MR scenes is shown. The wearable system 200 can also be referred to as an AR system 200. The wearable system 200 includes a display 220 and various mechanical and electronic modules and systems that support the functioning of the display 220. The display 220 can be coupled to a frame 230 that can be worn by a user, wearer, or viewer 210. The display 220 can be positioned in front of the eyes of the user 210. The display 220 can present AR / VR / MR content to the user. The display 220 can include a head-mounted display (HMD) that is worn on the head of the user.
[0073] In some embodiments, a speaker 240 is coupled to the frame 230 and positioned adjacent to the ear canal of the user (another speaker, not shown, can be positioned adjacent to the other ear canal of the user to provide stereo / shapeable sound control, in some embodiments). The display 220 can include an audio sensor (e.g., microphone) 232 to detect audio streams from the environment and capture ambient sound. In some embodiments, one or more other audio sensors, not shown, are positioned to provide stereo reception. Stereo reception can be used to determine the location of a sound source. The wearable system 200 can perform sound or speech recognition on the audio streams.
[0074] The wearable system 200 can include an outward-facing imaging system 464 FIG. 4As shown), the imaging system 464 observes the world in the environment surrounding the user. The wearable system 200 may also include an inward-facing imaging system 462 that can be used to track the user's eye movements. FIG. 4 (As shown). An inward-facing imaging system can track the movement of one eye or both eyes. The inward-facing imaging system 462 can be attached to frame 230 and electrically communicated with processing module 260 or 270, which can process image information acquired by the inward-facing imaging system to determine, for example, the pupil diameter or orientation of user 210's eyes, eye movement, or eye pose. The inward-facing imaging system 462 may include one or more cameras. For example, at least one camera can be used to image each eye. Images acquired by the cameras can be used to determine the pupil size or eye pose of each eye separately, thereby allowing image information to be presented to each eye dynamically adapted to that eye.
[0075] As an example, the wearable system 200 can use either an externally oriented imaging system 464 or an internally oriented imaging system 462 to acquire images of the user's posture. The images can be still images, video frames, or videos.
[0076] The display 220 can be operatively coupled to the local data processing module 260, for example via a wired lead or a wireless connection 250. The local data processing module 260 can be installed in various configurations, such as being fixedly attached to the frame 230, fixedly attached to a helmet or hat worn by the user, embedded in headphones, or otherwise detachably attached to the user 210 (e.g., in a backpack configuration or a belt-coupled configuration).
[0077] The local processing and data module 260 can include a hardware processor and a digital memory, e.g., nonvolatile memory (e.g., flash memory), both of which can be used to assist in the processing, caching, and storage of data. The data can include: a) data captured from sensors (which can be, e.g., operatively coupled to the frame 230 or otherwise attached to the user 210), such as image capture devices (e.g., cameras in an inward-facing imaging system or an outward-facing imaging system), audio sensors (e.g., microphones), inertial measurement units (IMUs), accelerometers, compasses, global positioning system (GPS) units, radio devices, or gyros; or b) data retrieved or processed by the remote processing module 270 or the remote data repository 280, possibly after such processing or retrieval, for delivery to the display 220. The local processing and data module 260 can be operatively coupled to the remote processing module 270 or remote data repository 280 by communication links 262 or 264 (e.g., via wired or wireless communication links) such that these remote modules are available as resources to the local processing and data module 260. Additionally, the remote processing module 270 and remote data repository 280 can be operatively coupled to each other to allow data to be shared between them.
[0078] In some embodiments, the remote processing module 270 can include one or more processors configured to analyze and process data or image information. In some embodiments, the remote data repository 280 can be a digital data storage facility that can be used over the Internet or other network configuration in a “cloud” resource configuration. In some embodiments, all data is stored and all computations are performed in the local processing and data module, allowing for fully autonomous use from the remote modules.
[0079] Example components of wearable systems
[0080] FIG. 3 Example components of a wearable system are schematically illustrated. FIG. 3 A wearable system 200 is shown, which can include a display 220 and a frame 230. An exploded view 202 schematically illustrates various components of the wearable system 200. In certain implementations, FIG. 3 One or more of the illustrated components can be part of the display 220. The various components, alone or in combination, can collect various data (e.g., audio or visual data) associated with a user of the wearable system 200 or the user’s environment. It will be appreciated that other embodiments can have more or less components depending on the application for which the wearable system is used. Nonetheless, FIG. 3 A basic idea of some of the various components and the types of data that can be collected, analyzed, and stored by the wearable system is provided.
[0081] FIG. 3An example wearable system 200, which may include a display 220, is shown. The display 220 may include a display lens 226 that can be mounted to a user's head or a housing or frame 230, corresponding to the frame 230. The display lens 226 may include one or more transparent lenses positioned by the housing 230 in front of the user's eyes 302, 304, and can be configured to deflect projected light 338 into the eyes 302, 304 and facilitate beam shaping, while also allowing at least some light from the local environment to pass through. The wavefront of the projected beam 338 can be bent or focused to match the desired focal length of the projected light. As shown, two wide-field-of-view machine vision cameras 316 (also referred to as world cameras) may be coupled to the housing 230 to image the environment surrounding the user. These cameras 316 may be dual-capture visible / invisible (e.g., infrared) light cameras. FIG. 4 This is part of the outward-facing imaging system 464 shown. Images acquired by the world camera 316 can be processed by the pose processor 336. For example, the pose processor 336 can implement one or more object recognizers 708 (e.g., shown in FIG. 7) to identify the pose of a user or another person in the user's environment or to identify physical objects in the user's environment.
[0082] Continue to refer to FIG. 3 The diagram illustrates a pair of scanning laser-shaped wavefront (e.g., for depth) light projector modules with display mirrors and optics, configured to project light 338 into eyes 302, 304. The depicted view also shows two miniature infrared cameras 324 paired with an infrared light source 326 (e.g., a light-emitting diode "LED"), configured to track the user's eyes 302, 304 to support rendering and user input. The cameras 324 may be... FIG. 4 This is part of the inward-facing imaging system 462 shown. The wearable system 200 is further characterized by a sensor assembly 339, which may include X, Y, and Z-axis accelerometer capabilities and a magnetic compass, as well as X, Y, and Z-axis gyroscope capabilities, preferably providing data at a relatively high frequency (e.g., 200 Hz). The sensor assembly 339 may serve as a reference. FIG. 2 A is part of the IMU described. The depicted system 200 may also include a head pose processor 336, such as an ASIC (Application-Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), or ARM processor (Advanced Simplified Instruction Set Machine), which can be configured to calculate real-time or near-real-time user head pose based on wide-field-of-view image information output from the capture device 316. The head pose processor 336 may be a hardware processor and may be implemented as... FIG. 2 Part of the local processing and data module 260 shown in A.
[0083] The wearable system may also include one or more depth sensors 234. The depth sensor 234 may be configured to measure the distance between an object in the environment and the wearable device. The depth sensor 234 may include a laser scanner (e.g., LiDAR), an ultrasonic depth sensor, or a depth-sensing camera. In some embodiments, where the camera 316 has depth-sensing capabilities, the camera 316 may also be considered as the depth sensor 234.
[0084] A processor 332 is also shown, configured to perform digital or analog processing to derive attitude from gyroscope, compass, or accelerometer data from sensor assembly 339. Processor 332 may be... FIG. 2 This is part of the local processing and data module 260 shown. FIG. 3 The wearable system 200 shown may also include a positioning system such as GPS 337 (Global Positioning System) to assist in posture and positioning analysis. Additionally, GPS can further provide remote (e.g., cloud-based) information about the user's environment. This information can be used to identify objects or information in the user's environment.
[0085] The wearable system can combine data acquired by GPS 337 and a remote computing system (e.g., remote processing module 270, another user's ARD, etc.) to provide more information about the user's environment. As an example, the wearable system can determine the user's location based on GPS data and retrieve a world map including virtual objects associated with the user's location (e.g., by communicating with remote processing module 270). As another example, the wearable system 200 can use a world camera 316 (which may be...) FIG. 4 As part of the externally facing imaging system 464 shown, the wearable system 200 monitors the environment. Based on images acquired by the world camera 316, the wearable system 200 can detect objects in the environment (e.g., by using one or more object recognizers 708 shown in Figure 7). The wearable system can further use data acquired by GPS 337 to interpret roles.
[0086] The wearable system 200 may also include a rendering engine 334, which can be configured to provide user-local rendering information to facilitate the operation of the scanner and imaging into the user's eyes for the user to view the world. The rendering engine 334 may be implemented by a hardware processor (e.g., a central processing unit or a graphics processing unit). In some embodiments, the rendering engine is part of a local processing and data module 260. The rendering engine 334 may be communicatively coupled (e.g., via wired or wireless links) to other components of the wearable system 200. For example, the rendering engine 334 may be coupled to an eye camera 324 via communication link 274 and to a projection subsystem 318 (which can project light into the user's eyes 302, 304 via a scanned laser arrangement in a manner similar to a retinal scanning display) via communication link 272. The rendering engine 334 may also communicate with other processing units, such as a sensor pose processor 332 and an image pose processor 336, via links 276 and 294, respectively.
[0087] Camera 324 (e.g., a miniature infrared camera) can be used to track eye pose to support rendering and user input. Some example eye poses may include where the user is looking or the depth at which he or she is focusing (this can be estimated via eye vergence). GPS 337, a gyroscope, a compass, and an accelerometer 339 can be used to provide coarse or rapid pose estimation. One or more of cameras 316 can acquire images and poses, which, along with data from associated cloud computing resources, can be used to map the local environment and share the user view with other users.
[0088] FIG. 3 The example components shown are for illustrative purposes only. For ease of explanation and description, multiple sensors and other functional modules are shown together. Some embodiments may include only one or a subset of these sensors or modules. Furthermore, the placement of these components is not limited to... FIG. 3 The locations shown are as indicated. Some components may be mounted to or housed within other components, such as a belt-mounted component, a handheld component, or a helmet component. As an example, the image pose processor 336, the sensor pose processor 332, and the rendering engine 334 may be housed in a belt pouch and configured to communicate with other components of the wearable system via wireless communication (e.g., ultra-wideband, Wi-Fi, Bluetooth, etc.) or via wired communication. The depicted housing 230 is preferably head-mounted and wearable. However, some components of the wearable system 200 may be worn on other parts of the user's body. For example, a speaker 240 may be inserted into the user's ear to provide sound to the user.
[0089] With respect to the projection of light 338 into the user's eyes 302, 304, in some embodiments, the camera 324 can be used to measure where the center of the user's eyes are geometrically verged, which generally coincides with the focal point position or "depth of focus" of the eyes. The three-dimensional surface of all points to which the eyes are verged can be referred to as the "horopter." The focal distance can have a finite amount of depth, or can vary infinitely. Light projected from vergence distances appears to be focused to the subject's eyes 302, 304, while light before or after the vergence distance becomes blurred. Examples of wearable devices and other display systems of the present disclosure are also described in U.S. Patent Publication No. 2016 / 0270656, which is incorporated by reference herein in its entirety.
[0090] The human visual system is complex, and providing a realistic sense of depth is challenging. A viewer of an object can perceive the object in three dimensions due to the combination of vergence and accommodation. Vergence movements of the two eyes relative to each other (i.e., rolling movements of the pupils toward or away from each other to converge the lines of sight of the eyes to fixate on an object) are closely associated with focusing (or "accommodation") of the lenses of the eyes. Under normal circumstances, changing the focus of the lenses of the eyes, or accommodating the eyes, to change focus from one object to another object at a different distance will automatically cause a matching change in vergence to the same distance, under a relationship known as the "accommodation-vergence reflex." Likewise, a change in vergence will trigger a matching change in accommodation under normal circumstances. Display systems that provide a better match between accommodation and vergence can create more realistic and comfortable simulations of three-dimensional imagery.
[0091] The human eye can resolve spatially coherent light correctly with a beam diameter of less than about 0.7 mm, regardless of where the eye is focused. Thus, to create the illusion of proper depth of focus, the camera 324 can be used to track eye vergence, and the rendering engine 334 and projection subsystem 318 can be utilized to render all objects in focus on or near the horopter, and all other objects with varying degrees of defocus (e.g., using intentionally created blur). Preferably, the system 220 renders to the user at a frame rate of approximately 60 frames per second or higher. As noted above, the camera 324 can be used for eye tracking, and the software can be configured to pick up not only vergence geometry, but also focal point position cues for use as user input. Preferably, such display systems are configured with luminance and contrast suitable for use in daylight or at night.
[0092] In some embodiments, the display system preferably has a visual object alignment latency of less than about 20 milliseconds, an angular alignment of less than about 0.1 degrees, and a resolution of about 1 arc minute, which is not limited by theory to be approximately the limit of the human eye. The display system 220 can be integrated with a positioning system, which can involve GPS elements, optical tracking, compass, accelerometer, or other data sources, to help determine position and pose; the positioning information can be used to facilitate accurate rendering of the user's view of the relevant world time (e.g., such information would help the glasses know their position relative to the real world).
[0093] In some embodiments, the wearable system 200 is configured to display one or more virtual images based on the accommodation of the user's eyes. In some embodiments, unlike existing 3D display methods that force the user to focus on where the image is projected, the wearable system is configured to automatically vary the focus of the projected virtual content to allow for more comfortable viewing of the one or more images presented to the user. For example, if the user's eye is currently focused at 1 m, the image can be projected so that it is consistent with the user's focus. If the user shifts focus to 3 m, the projected image is consistent with the new focus. Thus, the wearable system 200 of some embodiments does not force the user to reach a predetermined focus, but rather allows the user's eyes to act in a more natural manner.
[0094] Such a wearable system 200 can eliminate or reduce the occurrence of eye strain, headaches, and other physiological symptoms commonly observed with respect to virtual reality devices. To achieve this, various embodiments of the wearable system 200 are configured to project virtual images at varying focal distances through one or more variable focus elements (VFEs). In one or more embodiments, 3D perception can be achieved through a multi-plane focus system that projects images on a fixed focal plane that is distanced from the user. Other embodiments employ a variable plane focus, in which the focal plane moves back and forth in the z-direction to be consistent with the user's current focus state.
[0095] In multi-plane focusing systems and variable plane focusing systems, the wearable system 200 can use eye tracking to determine the vergence of the user’s eyes, determine the user’s current focus, and project virtual images with the determined focus. In other embodiments, the wearable system 200 includes a light modulator that projects a beam of light across the retina in a raster pattern with different foci in a variable manner by a fiber optic scanner or other light producing source. Thus, as further described in U.S. Patent Publication No. 2016 / 0270656 (which is hereby incorporated by reference in its entirety), the ability of the display of the wearable system 200 to project images with varying focal distances not only makes it easy for the user to accommodate to view 3D objects, but can also be used to compensate for user eye abnormalities. In some other embodiments, a spatial light modulator can project images to a user through various optical components. For example, as further described below, a spatial light modulator can project images onto one or more waveguides, which then send the images to the user.
[0096] Waveguide stack assembly
[0097] FIG. 4 An example of a waveguide stack for outputting image information to a user is shown. The wearable system 400 includes a stack of waveguides or waveguide assembly 480, which can be used to provide three-dimensional perception to the eye / brain using multiple waveguides 432b, 434b, 436b, 438b, 440b. In some embodiments, the wearable system 400 can correspond to the wearable system 200 of FIG. 2 FIG. 4 Some portions of the wearable system 200 are shown in more detail schematically. For example, in some embodiments, the waveguide assembly 480 can be integrated into the display 220 of FIG. 2
[0098] With continued reference to FIG. 4 The waveguide assembly 480 can also include a plurality of features 458, 456, 454, 452 between the waveguides in some embodiments. In some embodiments, the features 458, 456, 454, 452 can be lenses. In some embodiments, the lenses can be variable focus elements (VFEs). For example, in some embodiments, the waveguide assembly 480 can simply include two variable focus elements and one or more waveguides between the two variable focus elements. An example of a waveguide assembly with VFEs is disclosed in U.S. Patent Publication No. 2017 / 0293145, published October 12, 2017, the entire disclosure of which is hereby incorporated by reference. In other embodiments, the features 458, 456, 454, 452 can not be lenses. Rather, they can simply be spacers (e.g., cladding layers or structures to form air gaps).
[0099] The waveguides 432b, 434b, 436b, 438b, 440b or the plurality of lenses 458, 456, 454, 452 can be configured to send image information to the eye with various levels of wavefront curvature or light ray divergence. Each waveguide level can be associated with a particular depth plane and can be configured to output image information corresponding to that depth plane. Image injection devices 420, 422, 424, 426, 428 can be used to inject image information into the waveguides 440b, 438b, 436b, 434b, 432b, each of which can be configured to distribute incoming light across each respective waveguide to output to the eye 410. Light exits the output surface of the image injection devices 420, 422, 424, 426, 428 and is injected into the corresponding input edge of the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, a single beam of light (e.g., a collimated beam) can be injected into each waveguide to output an entire field of cloned collimated beams pointed toward the eye 410 at particular angles (and amounts of divergence) corresponding to the depth plane associated with a particular waveguide.
[0100] In some embodiments, the image injection devices 420, 422, 424, 426, 428 are discrete displays that each produce image information to inject into a corresponding waveguide 440b, 438b, 436b, 434b, 432b, respectively. In some other embodiments, the image injection devices 420, 422, 424, 426, 428 are output ends of a single multiplexed display that can deliver image information to each of the image injection devices 420, 422, 424, 426, 428, e.g., via one or more optical conduits (e.g., fiber optic cables).
[0101] A controller 460 controls operation of the stacked waveguide assembly 480 and the image injection devices 420, 422, 424, 426, 428. The controller 460 includes programming (e.g., instructions in a non-transitory computer-readable medium) that modulates timing and provides image information to the waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, the controller 460 can be a single integrated device or a distributed system connected through wired or wireless communication channels. In some embodiments, the controller 460 can be part of the processing modules 260 or 270 (shown in FIG. 2B). FIG. 2
[0102] Waveguides 440b, 438b, 436b, 434b, 432b can be configured to propagate light within each respective waveguide by total internal reflection (TIR). Waveguides 440b, 438b, 436b, 434b, 432b can each be planar or have another shape (e.g., curved), and can have major top and bottom surfaces and edges extending between those major top and bottom surfaces. In the illustrated configuration, waveguides 440b, 438b, 436b, 434b, 432b can each include light extraction optical elements 440a, 438a, 436a, 434a, 432a configured to extract light from the waveguides to output image information to eye 410 by redirecting the light propagating within the respective waveguides. The extracted light can also be referred to as outcoupled light, and the light extraction optical elements can also be referred to as outcoupling optical elements. An extracted light beam is output by the waveguide at a location where light propagating in the waveguide hits the light redirecting element. The light extraction optical elements (440a, 438a, 436a, 434a, 432a) can for example be reflective or diffractive optical features. While shown on the bottom major surfaces of waveguides 440b, 438b, 436b, 434b, 432b for ease of description and drawing clarity, in some embodiments light extraction optical elements 440a, 438a, 436a, 434a, 432a can be disposed on the top major surface or the bottom major surface, or can be disposed directly in the volume of waveguides 440b, 438b, 436b, 434b, 432b. In some embodiments, light extraction optical elements 440a, 438a, 436a, 434a, 432a can be formed in a layer of material that is attached to a transparent substrate to form waveguides 440b, 438b, 436b, 434b, 432b. In some other embodiments, waveguides 440b, 438b, 436b, 434b, 432b can be a monolithic piece of material, and light extraction optical elements 440a, 438a, 436a, 434a, 432a can be formed on a surface of or inside the piece of material.
[0103] With continued reference to FIG. 4As described herein, each waveguide 440b, 438b, 436b, 434b, 432b is configured to output light to form an image corresponding to a particular depth plane. For example, the waveguide 432b closest to the eye can be configured to deliver collimated light injected into such waveguide 432b to the eye 410. The collimated light can represent an optically-infinite focal plane. The next up waveguide 434b can be configured to emit collimated light that passes through a first lens 452 (e.g., a negative lens) before reaching the eye 410. The first lens 452 can be configured to produce a slight convex wavefront curvature, so that the eye / brain interprets light coming from this next up waveguide 434b as coming from a first focal plane that is inwardly closer to the eye 410 from optical infinity. Similarly, the third up waveguide 436b has its output light pass through the first and second lenses 452, 454 before reaching the eye 410. The combined optical power of the first and second lenses 452, 454 can be configured to produce another incremental amount of wavefront curvature, so that the eye / brain interprets light coming from the third waveguide 436b as coming from a second focal plane that is inwardly closer to the person than the light from the next up waveguide 434b from optical infinity.
[0104] The other waveguide layers (e.g., waveguides 438b, 440b) and lenses (e.g., lenses 456, 458) are similarly configured, with the highest waveguide 440b in the stack sending its output through all the lenses between it and the eye for representing the aggregate focal power of the closest-to-the-person focal plane. To compensate for the lens stack 458, 456, 454, 452 when viewing / interpreting light from the world 470 on the other side of the stacked waveguide assembly 480, a compensating lens layer 430 can be provided at the top of the stack to compensate for the aggregate focal power of the underlying lens stack 458, 456, 454, 452. (The compensating lens layer 430 and the stacked waveguide assembly 480 as a whole can be configured so that light from the world 470 is delivered to the eye 410 with substantially the same level of divergence (or collimation) as the light had when it was initially received by the stacked waveguide assembly 480). Such a configuration provides as many perceived focal planes as there are waveguide / lens pairs available. The light extraction optical elements of the waveguides and the focusing aspects of the lenses can both be static (e.g., not dynamic or electrically-activated). In some alternative embodiments, one or both can be dynamic, where electrically-activated features are used.
[0105] With continued reference to FIG. 4The light extraction optical elements 440a, 438a, 436a, 434a, 432a can be configured to both redirect light out of their respective waveguides and output that light with an appropriate amount of divergence or collimation for the particular depth plane associated with the waveguide. As a result, waveguides having different associated depth planes can have different configurations of light extraction optical elements that output light at different amounts of divergence depending on the associated depth plane. In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a can be volume or surface features that can be configured to output light at particular angles, as discussed herein. For example, the light extraction optical elements 440a, 438a, 436a, 434a, 432a can be volume holograms, surface holograms, and / or diffraction gratings. Light extraction optical elements such as diffraction gratings are described in U.S. Patent Publication No. 2015 / 0178939, published June 25, 2015, which is incorporated by reference herein in its entirety.
[0106] In some embodiments, the light extraction optical elements 440a, 438a, 436a, 434a, 432a are diffractive features that form a diffraction pattern or a diffractive optical element (also referred to herein as a "DOE"). Preferably, the DOE has a relatively low diffraction efficiency, such that with each interaction with the DOE, only a portion of the light of the light beam is deflected toward the eye 410, while the remainder continues to move through the waveguide by total internal reflection. The light carrying the image information can thus be split into a plurality of related exit beams that exit the waveguide at a plurality of locations, and the result is a fairly uniform pattern of exit emission toward the eye 304 for that particular collimated light beam that bounces around within the waveguide.
[0107] In some embodiments, one or more DOEs can switch between an "on" state in which they actively diffract, and an "off state in which they do not significantly diffract. For example, a switchable DOE can include a polymer dispersed liquid crystal layer in which microdrops contain a diffractive pattern in a host medium, and the refractive index of the microdrops can be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdrops can be switched to a refractive index that does not match the refractive index of the host medium (in which case the pattern actively diffracts incident light).
[0108] In some embodiments, the number and distribution of depth planes or depth of field can be dynamically varied based on the pupil size or orientation of the viewer's eye. The depth of field can be inversely proportional to the pupil size of the viewer. As a result, as the pupil size of the viewer's eye decreases, the depth of field increases, such that planes that would not be resolvable due to their location beyond the eye's depth of focus can become resolvable and appear more in focus as the pupil size decreases and the depth of field increases. Likewise, as the pupil size decreases, the number of spaced apart depth planes used to present different images to the viewer can be reduced. For example, at one pupil size, the viewer can not be able to clearly perceive the details of both a first depth plane and a second depth plane without adjusting accommodation of the eye from one depth plane to the other. However, the two depth planes can be simultaneously sufficiently in focus for a user at another pupil size without changing accommodation.
[0109] In some embodiments, the display system can change the number of waveguides that receive image information based on a determination of the pupil size or orientation, or based on receiving an electrical signal indicating a particular pupil size or orientation. For example, if the user's eye is not able to distinguish between two depth planes associated with two waveguides, the controller 460 (which can be an embodiment of the local processing and data module 260) can be configured or programmed to stop providing image information to one of the waveguides. Advantageously, this can reduce the processing burden on the system, increasing responsiveness of the system. In embodiments where the DOEs of the waveguides are switchable between an on and off state, when the waveguide does receive image information, the DOEs can be switched to the off state.
[0110] In some embodiments, it can be desirable for the exit beam to satisfy the condition that the diameter is less than the diameter of the viewer's eye. However, given the variability of the viewer's pupil size, it can be challenging to satisfy this condition. In some embodiments, the condition is satisfied over a wide range of pupil sizes by varying the size of the exit beam in response to a determination of the viewer's pupil size. For example, as the pupil size decreases, the size of the exit beam can also decrease. In some embodiments, a variable aperture can be used to vary the size of the exit beam.
[0111] The wearable system 400 can include an outward-facing imaging system 464 (e.g., a digital camera) that images a portion of the world 470. This portion of the world 470 can be referred to as the field of view (FOV) of the world camera, and the imaging system 464 is sometimes referred to as a FOV camera. The FOV of the world camera can be the same as or different from the FOV of the viewer 210, which encompasses the portion of the world 470 that the viewer 210 perceives at a given time. For example, in some cases, the FOV of the world camera can be larger than the FOV of the viewer 210 of the wearable system 400. The entire area that can be viewed by the viewer or imaged can be referred to as the field of regard (FOR). The FOR can include 4p steradians of solid angle around the wearable system 400, as the wearer can move his body, head, or eyes to perceive substantially any direction in space. In other contexts, the wearer's movements can be more limited, and thus the wearer's FOR can subtend a smaller solid angle. Images obtained from the outward-facing imaging system 464 can be used to track gestures made by the user (e.g., hand or finger gestures), detect objects in the world 470 in front of the user, and so on.
[0112] The wearable system 400 can include an audio sensor 232 (e.g., a microphone) to capture ambient sound. As described above, in some embodiments, one or more other audio sensors can be positioned to provide stereo reception useful in determining the location of a source of speech. As another example, the audio sensor 232 can include a directional microphone that can also provide directional information useful about where an audio source is located. The wearable system 400 can use information from the outward-facing imaging system 464 and the audio sensor 230 in localizing a source of speech, or determining an active speaker at a particular time, and so on. For example, the wearable system 400 can use speech recognition alone or in combination with a reflected image of the speaker (e.g., as seen in a mirror) to determine the identity of the speaker. As another example, the wearable system 400 can determine the location of a speaker in the environment based on sound acquired from a directional microphone. The wearable system 400 can parse sound from the location of the speaker using a speech recognition algorithm to determine the content of the speech, and use speech recognition techniques to determine the identity of the speaker (e.g., a name or other demographic information).
[0113] The wearable system 400 can also include an inward-facing imaging system 466 (e.g., a digital camera) that observes the movements of the user, e.g., eye movements and facial movements. The inward-facing imaging system 466 can be used to capture images of the eyes 410 to determine the size and / or orientation of the pupils of the eyes 304. The inward-facing imaging system 466 can be used to acquire images for determining the direction in which the user is looking (e.g., eye pose) or for biometrically identifying the user (e.g., via iris identification). In some embodiments, at least one camera can be utilized per eye to independently determine the pupil size or eye pose of each eye, allowing image information to be presented for each eye to be dynamically tailored to that eye. In some other embodiments, only the pupil diameter or orientation of a single eye 410 is determined (e.g., only a single camera is used per pair of eyes), and the pupil diameter or orientation is considered to be similar for both eyes of the user. The images obtained by the inward-facing imaging system 466 can be analyzed to determine the eye pose or emotions of the user, which the wearable system 400 can use to determine which audio or visual content should be presented to the user. The wearable system 400 can also use sensors such as IMUs, accelerometers, gyroscopes, etc. to determine head pose (e.g., head position or head orientation).
[0114] The wearable system 400 can include user input devices 466 through which a user can input commands to the controller 460 to interact with the wearable system 400. For example, the user input devices 466 can include touchpads, touchscreens, gamepads, multi-degree-of-freedom (DOF) controllers, capacitive sensing devices, game controllers, keyboards, mice, directional pads (D-pads), sticks, haptic devices, totems (e.g., that act as virtual user input devices), and the like. A multi-DOF controller can sense user input in some or all of the possible translations (e.g., left / right, forward / back, or up / down) or rotations (e.g., yaw, pitch, or roll) of the controller. A multi-DOF controller that supports translational movement can be referred to as 3DOF, while a multi-DOF controller that supports translations and rotations can be referred to as 6DOF. In some cases, a user can use a finger (e.g., a thumb) to press or swipe on a touch-sensitive input device to provide input to the wearable system 400 (e.g., to provide user input to a user interface provided by the wearable system 400). The user input devices 466 can be held in a user’s hand during use of the wearable system 400. The user input devices 466 can be in wired or wireless communication with the wearable system 400.
[0115] Other components of wearable systems
[0116] In many implementations, in addition to or instead of the components of the wearable system described above, the wearable system can include other components. The wearable system can for example include one or more haptic devices or components. The haptic devices or components are operable to provide a sense of touch to the user. For example, when touching virtual content (e.g., virtual objects, virtual tools, other virtual constructs), the haptic devices or components can provide a sense of pressure or texture. The sense of touch can replicate the feel of a physical object that the virtual object represents, or can replicate the feel of an imagined object or character (e.g., a dragon) that the virtual content represents. In some implementations, the user can wear the haptic devices or components (e.g., a user-wearable glove). In some implementations, the haptic devices or components can be held by the user.
[0117] For example, the wearable system can include one or more physical objects that the user can manipulate to allow input or interaction with the wearable system. These physical objects can be referred to herein as totems. Some totems can take the form of inanimate objects, e.g., a piece of metal or plastic, a wall, a table surface. In certain implementations, the totem can actually have no physical input structure (e.g., buttons, toggles, joysticks, trackballs, rocker switches) at all. Instead, the totem can simply provide a physical surface, and the wearable system can render a user interface so as to appear to the user to be on one or more surfaces of the totem. For example, the wearable system can cause images of a computer keyboard and trackpad to be rendered to appear to reside on one or more surfaces of a totem. For example, the wearable system can cause a virtual computer keyboard and virtual trackpad to be rendered to appear on a surface of a thin aluminum rectangular plate that serves as a totem. The rectangular plate itself can have no physical buttons, trackpad, or sensors. But the wearable system can detect user operations or interactions or touches of the rectangular plate as selections or inputs through the virtual keyboard or virtual trackpad. The user input device 466 (shown in FIG. 4) can be one embodiment of a totem, which can include a trackpad, touchpad, toggle, joystick, trackball, rocker, or virtual switch, mouse, keyboard, multi-degree-of-freedom controller, or other physical input device. The user can use the totem alone or in combination with gestures to interact with the wearable system or other users. FIG. 4
[0118] Examples of haptic devices and totems of the present disclosure for use with wearable devices, HMDs, and display systems are described in U.S. Patent Publication No. 2015 / 0016777, the entire contents of which are incorporated herein by reference.
[0119] Examples of eye images
[0120] FIG. 5 An image of an eye 500 is shown having an eyelid 504, a sclera 508 (the "white" of the eye), an iris 512, and a pupil 516. A curve 516a shows the pupil boundary between the pupil 516 and the iris 512, and a curve 512a shows the limbus boundary between the iris 512 and the sclera 508. The eyelid 504 includes an upper eyelid 504a and a lower eyelid 504b. The eye 500 is shown in a natural resting pose (e.g., where the user's face and gaze are both oriented toward a distant object directly in front of the user). The natural resting pose of the eye 500 can be indicated by a natural resting direction 520, which is the direction normal to the surface of the eye 500 when in the natural resting pose (e.g., directly in front of the eye 500, outside the plane of the eye 500 shown in the middle of the figure), and in this example, at the center of the pupil 516. FIG. 5
[0121] FIG. 5 The pose of the eye 500 can be expressed as two angular parameters, which indicate an azimuthal deflection and a zenithal deflection of the eye pose direction 524 of the eye, both relative to the natural resting direction 520 of the eye. For purposes of illustration, these angular parameters can be denoted as Θ (azimuthal deflection from a fiducial azimuth) and Φ (zenithal deflection, sometimes also referred to as polar deflection). In some implementations, an angular roll of the eye about the eye pose direction 524 can be included in the determination of the eye pose, and can be included in the following analysis. In other implementations, other techniques for determining the eye pose can use, for example, a pitch, yaw, and optionally roll system.
[0122] An eye image can be obtained from a video using any appropriate process, for example, using a video processing algorithm that can extract an image from one or more consecutive frames. The pose of the eye can be determined from the eye image using a variety of eye tracking techniques. For example, the eye pose can be determined by considering the lensing action of the cornea on a provided light source. Any suitable eye tracking technique can be used to determine the eye pose in the eyelid shape estimation techniques described herein.
[0123] Examples of eye tracking systems
[0124] FIG. 6 A schematic diagram of a wearable system 600 including an eye-tracking system is shown. In at least some embodiments, the wearable system 600 may include components located in a head-mounted unit 602 and components located in a non-head-mounted unit 604. The non-head-mounted unit 604 may be, for example, a belt-mounted component, a handheld component, a component in a backpack, a remote component, etc. Incorporating some components of the wearable system 600 into the non-head-mounted unit 604 can help reduce the size, weight, complexity, and cost of the head-mounted unit 602. In some embodiments, some or all of the functions described as being performed by one or more components of the head-mounted unit 602 and / or the non-head-mounted unit 604 may be provided by one or more components included elsewhere in the wearable system 600. For example, some or all of the functions described below associated with the CPU 612 of the head-mounted unit 602 may be provided by the CPU 616 of the non-head-mounted unit 604, and vice versa. In some examples, some or all of such functions may be provided by peripheral devices of the wearable system 600. Furthermore, in some implementations, computing can be performed via one or more cloud computing devices or other remote computing devices in a manner similar to the above reference. FIG. 2 The described method provides some or all of these functions.
[0125] like FIG. 6 As shown, the wearable system 600 may include an eye-tracking system comprising a camera 324 that captures images of a user's eyes 610. If desired, the eye-tracking system may also include light sources 326a and 326b (such as light-emitting diodes, "LEDs"). Light sources 326a and 326b may produce flashes (i.e., reflections that appear in the image of the eye captured by the camera 324 and are reflected away by the user's eyes). The positions of the light sources 326a and 326b relative to the camera 324 may be known, and therefore, the position of the flashes within the image captured by the camera 324 can be used to track the user's eyes (as will be discussed in more detail below in conjunction with Figures 7-11). In at least one embodiment, one light source 326 and one camera 324 may be associated with one of the user's eyes 610. In another embodiment, one light source 326 and one camera 324 may be associated with each of the user's eyes 610. In still other embodiments, one or more cameras 324 and one or more light sources 326 may be associated with one or each of the user's eyes 610. As a specific example, there may be two light sources 326a and 326b associated with each of the user's eyes 610, and one or more cameras 324. As another example, there may be three or more light sources (e.g., light sources 326a and 326b) associated with each of the user's eyes 610, and one or more cameras 324.
[0126] The eye tracking module 614 can receive images from the eye tracking camera 324 and can analyze the images to extract various pieces of information. As an example, the eye tracking module 614 can detect the user's eye pose, the three-dimensional position of the user's eyes relative to the eye tracking camera 324 (and relative to the head-mounted unit 602), the direction in which the user's eye or eyes 610 are focused, the user's vergence depth (i.e., the depth at which the user is focusing), the position of the user's pupils, the position of the user's corneas and cornea spheres, the center of rotation of each of the user's eyes, and the center of perspective of each of the user's eyes. The eye tracking module 614 can extract such information using the techniques described below in connection with FIGS. 7-11. As shown, the eye tracking module 614 can be a software module implemented using the CPU 612 in the head-mounted unit 602. Further details discussing the creation, adjustment, and use of eye tracking module components are provided in U.S. Patent Application No. 15 / 993,371, entitled "EYE TRACKING CALIBRATION TECHNIQUES," which is incorporated by reference herein in its entirety. FIG. 6
[0127] Data from the eye tracking module 614 can be provided to other components in the wearable system. For example, such data can be sent to components in the non-head-mounted unit 604, such as the CPU 616 including software modules for the light field rendering controller 618 and the registration observer 620.
[0128] The rendering controller 618 can use information from the eye tracking module 614 to adjust images displayed to the user by the rendering engine 622 (e.g., a rendering engine that can be a software module in the GPU 620 and that can provide images to the display 220). As an example, the rendering controller 618 can adjust images displayed to the user based on the user's center of rotation or center of perspective. In particular, the rendering controller 618 can use information about the user's center of perspective to simulate a rendering camera (i.e., to simulate collecting images from the user's perspective) and can adjust images displayed to the user based on the simulated rendering camera.
[0129] A "rendering camera," sometimes also referred to as a "pinhole perspective camera" (or simply "perspective camera") or "virtual pinhole camera" (or simply "virtual camera"), is a simulated camera used in rendering virtual image content that can come from a database of objects in a virtual world. The objects can have a position and orientation relative to a user or wearer and possibly relative to real objects in an environment surrounding the user or wearer. In other words, the rendering camera can represent a perspective within a rendering space from which the user or wearer will view 3D virtual content (e.g., virtual objects) of the rendering space. The rendering camera can be managed by a rendering engine to render a virtual image based on the database of virtual objects to be presented to the eyes. The virtual image can be rendered as if taken from the user's or wearer's point of view. For example, the virtual image can be rendered as if captured by a pinhole camera (corresponding to the "rendering camera") having a particular intrinsic parameter set (e.g., focal length, camera pixel size, principal point coordinates, skew / distortion parameters, etc.) and a particular extrinsic parameter set (e.g., translation and rotation components relative to the virtual world). The virtual image is taken from the perspective of this camera having a position and orientation (e.g., extrinsic parameters of the rendering camera) of the rendering camera. Thus, the system can define and / or adjust intrinsic and extrinsic rendering camera parameters. For example, the system can define a particular extrinsic rendering camera parameter set so that the virtual image can be rendered as if captured from the perspective of a camera having a particular position relative to the user's or wearer's eyes to provide an image that looks as if from the user's or wearer's point of view. The system can then adjust the extrinsic rendering camera parameters on-the-fly dynamically to maintain registration with the particular position. Similarly, intrinsic rendering camera parameters can be defined and adjusted dynamically over time. In some implementations, the image is rendered as if captured from the perspective of a camera having an aperture (e.g., pinhole) at a particular position (e.g., center of perspective or center of rotation or elsewhere) relative to the user's or wearer's eyes.
[0130] In some embodiments, the system can create or dynamically reposition and / or reorient one rendering camera for the user's left eye, and another rendering camera for the user's right eye, as the user's eyes are physically separated from each other and, thus, are always positioned at different locations. Thus, in at least some implementations, virtual content rendered from the perspective of a rendering camera associated with the user's left eye can be presented to the user through the eyepiece on the left side of a head-mounted display (e.g., head-mounted unit 602), and virtual content rendered from the perspective of a rendering camera associated with the user's right eye can be presented to the user through the eyepiece on the right side of such a head-mounted display. Further details regarding the creation, adjustment, and use of rendering cameras in the rendering process are discussed in U.S. Patent Application No. 15 / 274,823, entitled "METHODS AND SYSTEMS FOR DETECTING AND COMBINING STRUCTURAL FEATURES IN 3D RECONSTRUCTION," which is expressly incorporated by reference herein in its entirety for all purposes.
[0131] In some examples, one or more modules (or components) of system 600 (e.g., light field rendering controller 618, rendering engine 620, etc.) can determine the position and orientation of a rendering camera within a rendering space based on the position and orientation of the user's head and eyes (e.g., as determined based on head pose and eye tracking data, respectively). For example, system 600 can effectively map the position and orientation of the user's head and eyes to a particular position and angular position within a 3D virtual environment, place and orient a rendering camera at the particular position and angular position within the 3D virtual environment, and render virtual content for the user as if captured by the rendering camera. Further details regarding the real-world to virtual-world mapping process are discussed in U.S. Patent Application No. 15 / 296,869, entitled "SELECTING VIRTUAL OBJECTS IN A THREE-DIMENSIONAL SPACE," which is expressly incorporated by reference herein in its entirety for all purposes. As an example, rendering controller 618 can adjust the depth of a displayed image by selecting which depth plane (or multiple depth planes) to utilize in displaying the image at any given time. In some implementations, such depth plane switching can be performed by adjusting one or more intrinsic rendering camera parameters.
[0132] The registration observer 620 can use information from the eye-tracking module 614 to identify whether the head-mounted unit 602 is correctly positioned on the user's head. As an example, the eye-tracking module 614 can provide eye position information, such as the position of the user's eye rotation center, indicating the three-dimensional position of the user's eyes relative to the camera 324 and the head-mounted unit 602. The eye-tracking module 614 can use this position information to determine whether the display 220 is correctly aligned in the user's field of view, or whether the head-mounted unit 602 (or headset) has slipped or is misaligned with the user's eyes. As an example, the registration observer 620 can determine whether: the head-mounted unit 602 has slid down along the user's nose, thereby moving the display 220 away from the user's eyes and downwards (which may be undesirable); whether the head-mounted unit 602 has moved up along the user's nose, thereby moving the display 220 closer to the user's eyes and upwards; whether the head-mounted unit 602 has moved to the left or right relative to the user's nose, whether the head-mounted unit 602 has been raised above the user's nose, or whether the head-mounted unit 602 has been moved away from the desired position or position range in these or other ways. Generally, the registration observer 620 can determine whether the head-mounted unit 602, and especially the display 220, is correctly positioned in front of the user's eyes. In other words, the registration observer 620 can determine whether the left display in the display system 220 is correctly aligned with the user's left eye, and whether the right display in the display system 220 is correctly aligned with the user's right eye. The registration observer 620 can determine whether the head-mounted unit 602 is correctly positioned by determining whether the head-mounted unit 602 is positioned and oriented within the desired range of position and / or orientation relative to the user's eyes.
[0133] In at least some embodiments, the registration observer 620 may generate user feedback in the form of alarms, messages, or other content. Such feedback may be provided to the user to inform them of any misalignment of the head-mounted unit 602, as well as optional feedback on how to correct the misalignment (such as suggestions for adjusting the head-mounted unit 602 in a particular manner).
[0134] An example registration observation and feedback technique that can be used by the registration observer 620 is described in U.S. Patent Application No. 15 / 717,747 (Attorney’s No. MLEAP.052A2), filed on September 27, 2017, the entire contents of which are incorporated herein by reference.
[0135] Examples of eye tracking modules
[0136] exist FIG. 7A A detailed block diagram of an example eye-tracking module 614 is shown. FIG. 7AAs shown, eye tracking module 614 can include a variety of different sub-modules, can provide a variety of different outputs, and can utilize a variety of available data to track a user's eyes. As an example, eye tracking module 614 can utilize available data including eye tracking extrinsics and intrinsics, such as the geometric arrangement of eye tracking camera 324 relative to light source 326 and head-mounted unit 602; assumed eye dimensions 704, such as the typical distance between a user's corneal curvature center and the average center of rotation of a user's eye of approximately 4.7 mm, or the typical distance between a user's center of rotation and center of view; and per-user calibration data 706, such as a particular user's interpupillary distance. Additional examples of extrinsics, intrinsics, and other information that eye tracking module 614 can employ are described in U.S. Patent Application No. 15 / 497,726 (Attorney Docket No. MLEAP.023A7), filed April 26, 2017, which is incorporated by reference herein in its entirety.
[0137] Image pre-processing module 710 can receive images from an eye camera, such as eye camera 324, and can perform one or more pre-processing (i.e., conditioning) operations on the received images. As an example, image pre-processing module 710 can apply a Gaussian blur to the images, can down-sample the images to a lower resolution, can apply a non-sharpening mask, can apply an edge-sharpening algorithm, or can apply other suitable filters to aid in later detection, localization, and labeling of glints, pupils, or other features in images from eye camera 324. Image pre-processing module 710 can apply a low-pass filter or a morphological filter (e.g., an opening filter) that can remove high-frequency noise such as from pupil boundary 516a (see FIG. 6B), thus removing noise that can interfere with pupil and glint determination. Image pre-processing module 710 can output the pre-processed images to pupil identification module 712 and glint detection and labeling module 714. FIG. 5
[0138] The pupil identification module 712 can receive the pre-processed images from the image pre-processing module 710 and can identify regions of those images that include the user's pupils. In some embodiments, the pupil identification module 712 can determine the coordinates of the location or the coordinates of the center or centroid of the user's pupils in the eye tracking images from the camera 324. In at least some embodiments, the pupil identification module 712 can identify contours (e.g., contours of the pupil-iris boundary) in the eye tracking images, identify the contour moments (i.e., the centroid), apply a starburst pupil detection and / or Canny edge detection algorithm, reject outliers based on intensity values, identify sub-pixel boundary points, correct for eye camera distortion (i.e., distortions in the images captured by the eye camera 324), apply a random sample consensus (RANSAC) iterative algorithm to fit an ellipse to the boundary in the eye tracking images, apply a tracking filter to the images, and identify the sub-pixel image coordinates of the user's pupil centroid. The pupil identification module 712 can output pupil identification data to the glint detection and labeling module 714, which can indicate which regions of the pre-processing image module 712 are identified as showing the user's pupils. The pupil identification module 712 can provide the glint detection module 714 with 2D coordinates of the user's pupils within each eye tracking image (i.e., 2D coordinates of the centroid of the user's pupils). In at least some embodiments, the pupil identification module 712 can also provide the same kind of pupil identification data to the coordinate system normalization module 718.
[0139] Pupil detection techniques that can be utilized by the pupil identification module 712 are described in U.S. Patent Publication No. 2017 / 0053165, published February 23, 2017, and U.S. Patent Publication No. 2017 / 0053166, published February 23, 2017, each of which is incorporated by reference herein in its entirety.
[0140] The glint detection and labeling module 714 can receive the pre-processed images from module 710 and the pupil identification data from module 712. The glint detection module 714 can use this data to detect and / or identify glints (i.e., reflections of light from the light source 326 off of the user's eyes) within the region of the pre-processed images showing the user's pupils. As an example, the glint detection module 714 can search for bright regions (sometimes referred to herein as "blobs" or local intensity maxima) within the eye tracking images that are near the user's pupils. In at least some embodiments, the glint detection module 714 can rescale (e.g., magnify) the pupil ellipse to include additional glints. The glint detection module 714 can filter the glints by size and / or intensity. The glint detection module 714 can also determine the 2D position of each glint within the eye tracking images. In at least some examples, the glint detection module 714 can determine the 2D position of the glints relative to the user's pupils, which can also be referred to as the pupil glint vector. The glint detection and labeling module 714 can label the glints and output the pre-processed images with the labeled glints to the 3D corneal center estimation module 716. The glint detection and labeling module 714 can also pass data, such as the pre-processed images from module 710 and the pupil identification data from module 712.
[0141] Pupil and glint detection performed by modules such as modules 712 and 714 can use any suitable technique. As an example, edge detection can be applied to the eye images to identify glints and pupils. Edge detection can be applied by various edge detectors, edge detection algorithms, or filters. For example, a Canny edge detector can be applied to an image to detect edges in lines such as lines in the image. Edges can include points positioned along a line that correspond to local maxima of derivatives. For example, the pupil boundary 516a (see FIG. 5A) can be located using a Canny edge detector. In cases where the location of the pupil is determined, various image processing techniques can be used to detect the "pose" of the pupil 116. Determining the eye pose of an eye image can also be referred to as detecting the eye pose of an eye image. Pose can also be referred to as gaze, pointing direction, or orientation of the eye. For example, the pupil can be looking to the left at an object, and the pose of the pupil can be classified as a leftward pose. Other methods can be used to detect the location of the pupil or glints. For example, concentric rings can be placed in an eye image using a Canny edge detector. As another example, an integral differential operator can be used to find the pupil or limbus boundary of the iris. For example, a Daugman integral differential operator, a Hough transform, or other iris segmentation techniques can be used to return a curve that estimates the boundary of the pupil or iris. FIG. 5
[0142] 3D corneal center estimation module 716 can receive pre-processed images including detected glint data and pupil identification data from modules 710, 712, 714. 3D corneal center estimation module 716 can use this data to estimate the 3D position of the user's cornea. In some embodiments, 3D corneal center estimation module 716 can estimate the center of curvature of the eye's cornea or the 3D position of the user's corneal sphere, i.e., the center of an imaginary sphere having a surface portion that generally coextends with the user's cornea. 3D corneal center estimation module 716 can provide data indicating the estimated 3D coordinates of the corneal sphere and / or the user's cornea to coordinate system normalization module 718, optical axis determination module 722, and / or light field rendering controller 618. More details of the operation of 3D corneal center estimation module 716 are discussed herein in connection with FIG. 8A-8E Techniques for estimating the position of an eye feature, such as the cornea or corneal sphere, are discussed in U.S. Patent Application No. 15 / 497,726 (Attorney Docket MLEAP.023A7), filed April 26, 2017, the entirety of which is hereby incorporated by reference, which can be used by 3D corneal center estimation module 716 and other modules in the wearable system of the present disclosure.
[0143] Coordinate system normalization module 718 can optionally (as indicated by its dashed outline) be included in eye tracking module 614. Coordinate system normalization module 718 can receive from 3D corneal center estimation module 716 an estimate of the 3D coordinates of the center of the user's cornea (and / or the center of the user's corneal sphere), and can also receive data from other modules. Coordinate system normalization module 718 can normalize the eye camera coordinate system, which can help compensate for slippage of the wearable device (e.g., slippage of the head-mounted component from its normal, stationary position on the user's head, which can be identified by registration observer 620). Coordinate system normalization module 718 can rotate the coordinate system to align the z-axis of the coordinate system (i.e., the vergence depth axis) with the corneal center (e.g., as indicated by 3D corneal center estimation module 716), and can translate the camera center (i.e., the origin of the coordinate system) to a predetermined distance (such as 30 mm) from the corneal center (i.e., module 718 can magnify or demagnify the eye tracking images depending on whether the eye camera 324 is determined to be closer or farther than the predetermined distance). Through this normalization process, eye tracking module 614 is able to establish a consistent orientation and distance in the eye tracking data, relatively independent of changes in the head-mounted receiver positioning on the user's head. Coordinate system normalization module 718 can provide 3D pupil center localizer module 720 with the 3D coordinates of the center of the cornea (and / or corneal sphere), pupil identification data, and pre-processed eye tracking images. More details of the operation of coordinate system normalization module 718 are discussed herein in connection with FIG. 9A-9C Provided herein.
[0144] 3D pupil center locator module 720 can receive data in normalized or non-normalized coordinate systems, including 3D coordinates of the center of the user's cornea (and / or corneal sphere), pupil position data, and pre-processed eye tracking images. 3D pupil center locator module 720 can analyze such data to determine 3D coordinates of the center of the user's pupil in the normalized or non-normalized eye camera coordinate system. 3D pupil center locator module 720 can determine the position of the user's pupil in three dimensions based on the 2D position of the pupil centroid (as determined by module 712), the 3D position of the corneal center (as determined by module 716), the assumed eye dimensions 704 (e.g., the size of a typical user's corneal sphere and the typical distance from the corneal center to the pupil center), and the optical properties of the eye (e.g., the refractive index of the cornea (relative to the refractive index of air) or any combination of these. More details of the operation of 3D pupil center locator module 720 are discussed herein in connection with FIG. 9D-9G Techniques that can be utilized by 3D pupil center locator module 720 and other modules in the wearable system of the present disclosure for estimating the position of eye features such as the pupil are discussed in U.S. Patent Application No. 15 / 497,726 (Attorney Docket No. MLEAP.023A7), filed April 26, 2017, which is incorporated by reference herein in its entirety.
[0145] Optical axis determination module 722 can receive data from modules 716 and 720 indicating the 3D coordinates of the user's cornea and the center of the user's pupil. Based on such data, optical axis determination module 722 can identify a vector from the position of the corneal center (i.e., from the center of the corneal sphere) to the center of the user's pupil, which can define the optical axis of the user's eye. As an example, optical axis determination module 722 can provide an output specifying the user's optical axis to modules 724, 728, 730, and 732.
[0146] The center of rotation (CoR) estimation module 724 can receive data from the module 722 including parameters of the optical axis of the user's eye (i.e., data indicating the direction of the optical axis in a coordinate system having a known relationship to the head-mounted unit 602). The CoR estimation module 724 can estimate the center of rotation of the user's eye (i.e., the point around which the user's eye rotates when the user's eye rotates left, right, up, and / or down). While the eye can not be able to rotate perfectly around a singular point, it can be sufficient to assume one. In at least some embodiments, the CoR estimation module 724 can estimate the center of rotation of the eye by moving a particular distance along the optical axis (identified by the module 722) from the center of the pupil (identified by the module 720) or the center of curvature of the cornea (identified by the module 716) to the retina. This particular distance can be the assumed eye size 704. As one example, the particular distance between the center of curvature of the cornea and the CoR can be approximately 4.7 mm. This distance can be altered for a particular user based on any relevant data, including the user's age, gender, vision prescription, other relevant characteristics, etc.
[0147] In at least some embodiments, the CoR estimation module 724 can refine its estimate of the center of rotation of each of the user's eyes over time. For example, over the course of time, the user will eventually rotate the eyes (look at other places, look at closer things, look at farther things, or look left, right, up, or down at some point), causing the optical axis of each of the user's eyes to shift. The CoR estimation module 724 can then analyze the two (or more) optical axes identified by the module 722 and locate the 3D intersection of those optical axes. The CoR estimation module 724 can then determine that the center of rotation lies at that 3D intersection. Such a technique can provide an estimate of the center of rotation that improves in accuracy over time. Various techniques can be employed to improve the accuracy of the CoR estimation module 724 and the determined CoR locations of the left and right eyes. As an example, the CoR estimation module 724 can estimate the CoR by finding the average intersection of the optical axes determined over time for various different eye poses. As an additional example, the module 724 can filter or average the estimated CoR locations over time, can compute a moving average of the estimated CoR locations over time, and / or can apply a Kalman filter and known dynamics of the eye and eye tracking system to estimate the CoR location over time. As a particular example, the module 724 can compute a weighted average of the determined optical axis intersection points and the assumed CoR location (e.g., 4.7 mm from the center of curvature of the cornea of the eye) such that the determined CoR can slowly drift over time from the assumed CoR location (i.e., 4.7 mm behind the center of curvature of the cornea of the eye) to a slightly different location within the user's eye as eye tracking data for the user is obtained, enabling refinement of the CoR location per user.
[0148] The interpupillary distance (IPD) estimation module 726 can receive data from the CoR estimation module 724 indicating the estimated 3D positions of the centers of rotation of the user's left and right eyes. The IPD estimation module 726 can then estimate the user's IPD by measuring the 3D distance between the centers of rotation of the user's left and right eyes. Generally, when the user is gazing at optical infinity (i.e., the optical axes of the user's eyes are substantially parallel to each other), the distance between the estimated CoR of the user's left eye and the estimated CoR of the user's right eye can be approximately equal to the distance between the centers of the user's pupils, which is the typical definition of interpupillary distance (IPD). The user's IPD can be used by various components and modules in the wearable system. For example, the user's IPD can be provided to the registration observer 620 and used to assess the degree to which the wearable device is aligned with the user's eyes (e.g., whether the left and right display lenses are properly separated according to the user's IPD). As another example, the user's IPD can be provided to the vergence depth estimation module 728 and used to determine the user's vergence depth. The module 726 can employ various techniques, such as those discussed in connection with the CoR estimation module 724, to improve the accuracy of the estimated IPD. As an example, the IPD estimation module 724 can apply filtering, average over time, weighted average including a hypothesized IPD distance, Kalman filter, etc. as part of estimating the user's IPD in an accurate manner.
[0149] In some embodiments, the IPD estimation module 726 can receive data from the 3D pupil center localizer module and / or the 3D cornea center estimation module 716 indicating the estimated 3D positions of the user's pupils and / or corneas. The IPD estimation module 726 can then estimate the user's IPD with reference to the distances between the pupils and corneas. Generally, these distances will vary over time as the user rotates the eyes and changes their vergence depth. In some cases, the IPD estimation module 726 can look for the largest measured distance between the pupils and / or corneas, which should occur when the user is gazing at optical infinity and should generally correspond to the user's interpupillary distance. In other cases, the IPD estimation module 726 can fit the measured distances between the user's pupils (and / or corneas) to a mathematical relationship of how the interpupillary distance of a person varies as a function of their vergence depth. In some embodiments, using these or other similar techniques, the IPD estimation module 726 is able to estimate the user's IPD even in the absence of observations of the user looking at optical infinity (e.g., by extrapolating from one or more observations in which the user is converging at a distance closer than optical infinity).
[0150] The vergence depth estimation module 728 can receive data from various modules and sub-modules in the eye tracking module 614 (e.g., the CoR estimation module 724, the 3D pupil center localizer module, the 3D cornea center estimation module 716, etc.) and use this data to estimate the user's vergence depth.FIG. 7AThe illustrated). In particular, the vergence depth estimation module 728 can employ data indicative of the estimated 3D positions of the centers of the pupils (e.g., as provided by the module 720 described above), one or more determined optical axis parameters (e.g., as provided by the module 722 described above), the estimated 3D positions of the centers of rotation (e.g., as provided by the module 724 described above), the estimated IPD (e.g., the Euclidean distance between the estimated 3D positions of the centers of rotation) (e.g., as provided by the module 726 described above), and / or one or more determined optical and / or visual axis parameters (e.g., as provided by the module 722 and / or the module 730 described below). The vergence depth estimation module 728 can detect or otherwise obtain a measure of the vergence depth of the user, which can be the distance from the user at which the user’s eyes are in focus. For example, when the user is gazing at an object three feet in front of them, the vergence depth of the user’s left and right eyes is three feet; and when the user is gazing at a distant landscape (i.e., the optical axes of the user’s eyes are substantially parallel to each other such that the distance between the centers of the user’s pupils can be approximately equal to the distance between the centers of rotation of the user’s left and right eyes), the vergence depth of the user’s left and right eyes is infinite. In some implementations, the vergence depth estimation module 728 can utilize data indicative of the estimated centers of the user’s pupils (e.g., as provided by the module 720) to determine the 3D distance between the estimated centers of the user’s pupils. The vergence depth estimation module 728 can obtain a measure of the vergence depth by comparing such determined 3D distance between the centers of the pupils to the estimated IPD (e.g., the Euclidean distance between the estimated 3D positions of the centers of rotation) (e.g., as indicated by the module 726 described above). In addition to the 3D distance between the centers of the pupils and the estimated IPD, the vergence depth estimation module 728 can utilize known, assumed, estimated, and / or determined geometry to calculate the vergence depth. As an example, the module 728 can combine the 3D distance between the centers of the pupils, the estimated IPD, and the 3D CoR positions in a trigonometric calculation to estimate (i.e., determine) the vergence depth of the user. In effect, evaluating the thus determined 3D distance between the centers of the pupils according to the estimated IPD can be used to indicate a measure of the user’s current depth of focus relative to optical infinity. In some examples, the vergence depth estimation module 728 can simply receive or access data indicative of an estimated 3D distance between the estimated centers of the user’s pupils in order to obtain such a measure of the vergence depth. In some embodiments, the vergence depth estimation module 728 can estimate the vergence depth by comparing the user’s left and right optical axes. In particular, the vergence depth estimation module 728 can estimate the vergence depth by locating the distance from the user at which the user’s left and right optical axes intersect (or the projections of the user’s left and right optical axes on, e.g., a horizontal plane intersect).By setting the zero depth to be the depth at which the user's left and right optical axes are separated by the user's IPD, the module 728 can utilize the user's IPD in this calculation. In at least some embodiments, the vergence depth estimation module 728 can determine the vergence depth by triangulating the eye tracking data together with known or derived spatial relationships.
[0151] In some embodiments, the vergence depth estimation module 728 can estimate the user's vergence depth based on the intersection of the user's visual axes (rather than their optical axes), which can provide a more accurate indication of the distance at which the user is focusing. In at least some embodiments, the eye tracking module 614 can include an optical-to-visual axis mapping module 730. As discussed in further detail in connection with FIG. 7B, the user's optical and visual axes are typically not aligned. The visual axis is the axis along which a person is looking, while the optical axis is defined by the center of the person's lens and pupil, and can pass through the center of the person's retina. In particular, the user's visual axis is typically defined by the location of the user's fovea, which can be offset from the center of the user's retina, resulting in a difference between the optical and visual axes. In at least some of these embodiments, the eye tracking module 614 can include an optical-to-visual axis mapping module 730. The optical-to-visual axis mapping module 730 can correct for the difference between the user's optical and visual axes, and provide information about the user's visual axis to other components in the wearable system (e.g., the vergence depth estimation module 728 and the light field rendering controller 618). In some examples, the module 730 can use assumed eye dimensions 704, including a typical offset of about 5.2° inwards (nasally, towards the user's nose) between the optical and visual axes. In other words, the module 730 can shift the user's left optical axis (nasally) 5.2° to the right of the nose, and the user's right optical axis (nasally) 5.2° to the left of the nose, in order to estimate the directions of the user's left and right optical axes. In other examples, the module 730 can utilize per-user calibration data 706 in mapping the optical axes (e.g., as indicated by the module 722 described above) to the visual axes. As additional examples, the module 730 can shift the user's optical axes nasally between 4.0° and 6.5°, between 4.5° and 6.0°, between 5.0° and 5.4°, and so on, or any range formed by any of these values. In some arrangements, the module 730 can apply the shift based at least in part on characteristics of a particular user (e.g., their age, gender, vision prescription, or other relevant characteristics), and / or can apply the shift based at least in part on a calibration process for a particular user (i.e., to determine the optical-visual axis offset for the particular user). In at least some embodiments, the module 730 can also shift the origins of the left and right optical axes to correspond to the user's CoP (determined by the module 732) rather than to the user's CoR. FIG. 10
[0152] The optional center of perspective (CoP) estimation module 732 (when provided) can estimate the location of the left and right center of perspective (CoP) of the user. The CoP can be a useful location for the wearable system, and in at least some embodiments, is a location directly in front of the pupil. In at least some embodiments, the CoP estimation module 732 can estimate the location of the left and right center of perspective of the user based on the 3D location of the center of the pupil of the user, the 3D location of the center of curvature of the cornea of the user, or such suitable data or any combination. As an example, the CoP of the user can be approximately 5.01 mm in front of the center of curvature of the cornea (i.e., 5.01 mm from the center of the corneal sphere in a direction toward the cornea of the eye and along the optical axis) and can be approximately 2.97 mm along the optical or visual axis behind the outer surface of the cornea of the user. The center of perspective of the user can be directly in front of the center of their pupil. For example, the CoP of the user can be less than approximately 2.0 mm from the pupil of the user, less than approximately 1.0 mm from the pupil of the user, or less than approximately 0.5 mm from the pupil of the user, or any range between these values. As another example, the center of perspective can correspond to a location within the anterior chamber of the eye. As other examples, the CoP can be between 1.0 mm and 2.0 mm, approximately 1.0 mm, between 0.25 mm and 1.0 mm, between 0.5 mm and 1.0 mm, or between 0.25 mm and 0.5 mm.
[0153] The center of perspective described herein (as the potential desired location of the pinhole of the rendering camera and the anatomical location in the eye of the user) can be a location for reducing and / or eliminating undesirable parallax shifts. In particular, the optical system of the eye of the user closely approximates the theoretical system formed by a pinhole in front of a lens projecting onto a screen, with the pinhole, lens, and screen roughly corresponding to the pupil / iris, lens, and retina of the user, respectively. Furthermore, it can be desirable that there is little or no parallax shift when two point lights (or objects) at different distances from the eye of the user are rigidly rotated around the opening of the pinhole (e.g., rotated along a radius of curvature equal to their distance from the opening of the pinhole). Thus, it can seem that the CoP should be located at the center of the pupil of the eye (and in certain embodiments can use such a CoP). However, the human eye includes a cornea in addition to the pinhole of the pupil and the lens, which imparts additional optical power to light propagating to the retina. Thus, in the theoretical system described in this section, the anatomical equivalent of the pinhole can be a region of the eye of the user between the outer surface of the cornea of the eye of the user and the center of the pupil or iris of the eye of the user. For example, the anatomical equivalent of the pinhole can correspond to a region within the anterior chamber of the eye of the user. For various reasons discussed herein, it can be desirable to set the CoP to such a location within the anterior chamber of the eye of the user. The derivation and significance of the CoP are discussed in detail in the Appendices (Part I and Part II), which form a part of the present application.
[0154] As described above, the eye tracking module 614 can provide data, such as estimates of 3D positions of the centers of rotation (CoR) of the left and right eyes, the vergence depth, the optical axes of the left and right eyes, the 3D positions of the user's eyes, the 3D positions of the left and right corneal centers of curvature of the user's eyes, the 3D positions of the left and right pupil centers of the user, the 3D positions of the left and right centers of gaze of the user, the IPD of the user, etc., to other components in the wearable system, such as the light field rendering controller 618 and the registration observer 620. The eye tracking module 614 can also include other sub-modules that detect and generate data associated with other aspects of the user's eyes. As examples, the eye tracking module 614 can include a blink detection module that provides a flag or other alert each time the user blinks, and a saccade detection module that provides a flag or other alert each time the user's eyes saccade (i.e., quickly shift focus to another point).
[0155] Examples of rendering controllers
[0156] A detailed block diagram of an example light field rendering controller 618 is shown in FIG. 7B As shown in FIG. 6 and 7B The rendering controller 618 can receive eye tracking information from the eye tracking module 614, and can provide output to the rendering engine 622, which can generate images to be displayed for viewing by the user of the wearable system. As examples, the rendering controller 618 can receive the vergence depth, the centers of rotation (and / or centers of gaze) of the left and right eyes, and other eye data, such as blink data, saccade data, etc.
[0157] The depth plane selection module 750 can receive vergence depth information and other eye data, and based on such data can cause the rendering engine 622 to communicate content to the user at a particular depth plane (i.e., at a particular distance accommodation or focal distance). As described in connection with FIG. 4As discussed, the wearable system can include multiple discrete depth planes formed by multiple waveguides, each depth plane conveying image information at varying levels of wavefront curvature. In some embodiments, the wearable system can include one or more variable depth planes, e.g., optical elements that convey image information at varying levels of wavefront curvature over time. In these and other embodiments, the depth plane selection module 750 can cause the rendering engine 622 to convey content to the user at a selected depth based in part on the vergence depth of the user (i.e., cause the rendering engine 622 to direct the display 220 to switch depth planes). In at least some embodiments, the depth plane selection module 750 and the rendering engine 622 can render content at different depths and also generate and / or provide depth plane selection data to display hardware such as the display 220. The display hardware such as the display 220 can perform electrical depth plane switching in response to depth plane selection data (which can be control signals) generated and / or provided by modules such as the depth plane selection module 750 and the rendering engine 622.
[0158] In general, it can be desirable for the depth plane selection module 750 to select a depth plane that matches the current vergence depth of the user, thereby providing accurate accommodative cues to the user. However, it can also be desirable to switch depth planes in a discreet and unobtrusive manner. As an example, it can be desirable to avoid excessive switching between depth planes and / or to switch depth planes when the user is unlikely to notice the switch, e.g., during a blink or saccade.
[0159] The hysteresis band crossing detection module 752 can help avoid excessive switching between depth planes, particularly when the vergence depth of the user fluctuates at a midpoint or transition point between two depth planes. In particular, the module 752 can cause the depth plane selection module 750 to exhibit hysteresis in its selection of depth planes. As an example, the module 752 can cause the depth plane selection module 750 to switch from a first, more distant depth plane to a second, closer depth plane only after the vergence depth of the user exceeds a first threshold. Similarly, the module 752 can cause the depth plane selection module 750 (which in turn can direct a display such as the display 220) to switch to the first, more distant depth plane only after the vergence depth of the user exceeds a second threshold that is farther from the user than the first threshold. In an overlap region between the first threshold and the second threshold, the module 750 can cause the depth plane selection module 750 to retain whichever depth plane is currently selected as the selected depth plane, thereby avoiding excessive switching between depth planes.
[0160] The eye event detection module 750 can receive data from the eye tracking system 710, e.g., data indicating the vergence depth of the user and / or data indicating the accommodative state of the user. The eye event detection module 750 can also receive data from the display 220, e.g., data indicating the current depth plane of the display 220 and / or data indicating the accommodative state of the user. The eye event detection module 750 can also receive data from the user input module 720, e.g., data indicating the accommodative state of the user and / or data indicating the accommodative state of the user. FIG. 7AThe eye tracking module 614 receives other eye data, and can cause the depth plane selection module 750 to delay some depth plane switches until an ocular event occurs. As an example, the ocular event detection module 750 can cause the depth plane selection module 750 to delay a planned depth plane switch until a user blinks is detected; can receive data from blink detection components in the eye tracking module 614 indicating when the user currently blinks; and in response, can cause the depth plane selection module 750 to perform the planned depth plane switch during the blink event (e.g., by causing the module 750 to direct the display 220 to perform the depth plane switch during the blink event). In at least some embodiments, the wearable system can be able to shift content onto a new depth plane during a blink event such that the user is unlikely to perceive the shift. As another example, the ocular event detection module 750 can delay a planned depth plane switch until an eye saccade is detected. As discussed in connection with blinking, such an arrangement can facilitate discrete shifts of depth planes.
[0161] If desired, the depth plane selection module 750 can delay a planned depth plane switch for a limited period of time before performing the depth plane switch, even in the absence of an ocular event. Similarly, the depth plane selection module 750 can perform a depth plane switch when the user's vergence depth is substantially outside of the currently selected depth plane (i.e., when the user's vergence depth has exceeded a predetermined threshold that is beyond the regular threshold for a depth plane switch), even in the absence of an ocular event. These arrangements can help ensure that the ocular event detection module 754 does not indefinitely delay depth plane switches, and does not delay depth plane switches when there is a large accommodation error. Further details of the operation of the depth plane selection module 750 are provided herein in connection with the discussion of the depth plane selection module 750. FIG. 13 Further details of the operation of the depth plane selection module 750 are provided, as well as how the module times depth plane switches.
[0162] The render camera controller 758 can provide information to the rendering engine 622 indicating where the user's left and right eyes are. The rendering engine 622 can then generate content by simulating a camera at the location of the user's left and right eyes and generating content based on the perspective of the simulated camera. As described above, a render camera is a simulated camera used in rendering virtual image content that can come from a database of objects in a virtual world. The objects can have a position and orientation relative to the user or wearer and possibly relative to real objects in the environment around the user or wearer. The render camera can be included in the rendering engine to render a virtual image based on a database of virtual objects to be presented to the eyes. The virtual image can be rendered as if taken from the perspective of the user or wearer. For example, the virtual image can be rendered as if captured by a camera (corresponding to a "render camera") with an aperture, lens, and detector that looks at objects in a virtual world. The virtual image is taken from the perspective of such a camera with a position of the "render camera." For example, the virtual image can be rendered as if captured from the perspective of a camera with a particular position relative to the eyes of the user or wearer, providing an image that appears to be from the perspective of the user or wearer. In some implementations, the image is rendered as if captured from the perspective of a camera with an aperture at a particular position relative to the eyes of the user or wearer (e.g., a perspective center or center of rotation discussed herein or elsewhere).
[0163] The render camera controller 758 can determine the position of the left and right cameras based on the left and right eye centers of rotation (CoRs) determined by the CoR estimation module 724 and / or based on the left and right eye centers of perspective (CoPs) determined by the CoP estimation module 732. In some embodiments, the render camera controller 758 can toggle or discretely switch between CoR and CoP positions based on various factors. As an example, the render camera controller 758 can always register the render camera to the CoR position, always register the render camera to the CoP position, toggle or discretely switch between registering the render camera to the CoR position and registering the render camera to the CoP position over time based on various factors, or dynamically register the render camera to any position in a range of positions along an optical (or visual) axis between the CoR and CoP positions over time based on various factors in various modes. The CoR and CoP positions can optionally be passed through a smoothing filter 756 (in any of the aforementioned modes for render camera positioning) that can average the CoR and CoP positions over time to reduce noise in these positions and prevent jitter in the simulated render camera.
[0164] In at least some embodiments, the rendering camera may be simulated as a pinhole camera, wherein the pinhole is positioned at an estimated CoR or CoP location identified by the eye-tracking module 614. When the CoP offsets the CoR, the position of the rendering camera and its pinhole shift with the rotation of the user's eye whenever the rendering camera's position is based on the user's CoP. Conversely, the position of the rendering camera's pinhole does not move with eye rotation whenever the rendering camera's position is based on the user's CoR, although in some embodiments, the rendering camera (which is behind the pinhole) may move with eye rotation. In other embodiments where the rendering camera's position is based on the user's CoR, the rendering camera may not move (i.e., rotate) with the user's eye.
[0165] Examples of registration viewers
[0166] exist FIG. 7C A block diagram of an example registration observer 620 is shown. FIG. 6 , FIG. 7A and FIG. 7C As shown, the registration observer 620 can be connected from the eye-tracking module 614 ( FIG. 6 and 7A The registration observer 620 receives eye-tracking information. As an example, it may receive information about the user's left and right eye rotation centers (e.g., the three-dimensional positions of the user's left and right eye rotation centers, which may be in a common coordinate system or have a common reference frame with the head-mounted display system 600). As other examples, the registration observer 620 may receive display extrinsic parameters, adaptation tolerances, and eye-tracking validity indicators. Display extrinsic parameters may include information about the display (e.g., ...). FIG. 2 Information about the display (200), such as the display's field of view, the dimensions of one or more display surfaces, and the position of the display surfaces relative to the head-mounted display system 600. Adaptation tolerances may include information about the display registration volume, indicating how far a user's left and right eyes can be moved from their nominal positions before affecting display performance. Additionally, adaptation tolerances may indicate the expected amount of display performance impact based on the user's eye position.
[0167] like FIG. 7C As shown, the registration observer 620 may include a 3D position adaptation module 770. The position adaptation module 770 can acquire and analyze various data, including, for example, the 3D position of the left eye rotation center (e.g., left CoR), the 3D position of the right eye rotation center (e.g., right CoR), display extrinsic parameters, and adaptation tolerances. The 3D position adaptation module 770 can determine how far the user's left and right eyes are from their respective nominal left and right eye positions (e.g., it can calculate the left 3D error and the right 3D error), and can provide the error distances (e.g., the left 3D error and the right 3D error) to the device 3D adaptation module 772.
[0168] 3D position adaptation module 770 can also compare the error distance to display extrinsic parameters and adaptation tolerances to determine whether the user's eyes are within a nominal volume, a partially degraded volume (e.g., a volume in which the performance of display 220 is partially degraded), or in a fully or almost fully degraded volume (e.g., a volume in which display 220 is substantially unable to provide content to the user's eyes). In at least some embodiments, 3D position adaptation module 770 or 3D adaptation module 772 can provide an output that qualitatively describes the adaptation of the HMD to the user, such as the adaptation quality output shown in FIG. 6B. As an example, module 770 can provide an output that indicates whether the current adaptation of the HMD to the user is good, marginal, or failing. A good adaptation can correspond to an adaptation that enables the user to view at least a certain percentage of the image (e.g., 90%), a marginal adaptation can enable the user to view at least a lower percentage of the image (e.g., 80%), and a failing adaptation can be one in which the user can only see an even lower percentage of the image. FIG. 7C
[0169] As another example, 3D position adaptation module 770 and / or device 3D adaptation module 772 can compute a visible area metric, which can be a percentage of the total area (or pixels) of the image that is displayed by display 220 that is visible to the user. Modules 770 and 772 can compute the visible area metric by evaluating the positions of the user's left and right eyes relative to display 220 (e.g., which can be based on the centers of rotation of the user's eyes) and using one or more models (e.g., mathematical or geometric models), one or more lookup tables, or other techniques or combinations of these and other techniques to determine the percentage of the image that is visible to the user as a function of the positions of the user's eyes. In addition, modules 770 and 772 can determine which regions or portions of the image that display 220 is expected to display are visible to the user as a function of the positions of the user's eyes.
[0170] Registration observer 620 can also include device 3D adaptation module 772. Module 772 can receive data from 3D position adaptation module 770 and can also receive an eye tracking validity indicator, which can be provided by eye tracking module 614 and can indicate whether the eye tracking system is currently tracking the positions of the user's eyes or whether the eye tracking data is unavailable or in an error condition (e.g., determined to be unreliable). If desired, device 3D adaptation module 772 can modify the quality of the adaptation data received from 3D position adaptation module 770 as a function of the status of the eye tracking validity data. For example, if the data from the eye tracking system is indicated as being unavailable or in error, device 3D adaptation module 772 can provide a notification that there is an error and / or not provide an output to the user regarding the quality of the adaptation or the adaptation error.
[0171] In at least some embodiments, the registration observer 620 may provide the user with detailed feedback regarding the fit quality and the nature and magnitude of the error. As an example, the head-mounted display system may provide feedback to the user during the calibration or fitting process (e.g., as part of the setup process) and may provide feedback during operation (e.g., if the fit deteriorates due to slippage, the registration observer 620 may prompt the user to readjust the head-mounted display system). In some embodiments, registration analysis may be performed automatically (e.g., during use of the head-mounted display system), and feedback may be provided without user input. These are merely illustrative examples.
[0172] Examples of using eye tracking systems to locate a user’s cornea
[0173] FIG. 8A This is a schematic diagram of an eye, showing the cornea and the bulb. (Example) FIG. 8A As shown, the user's eye 810 may have a cornea 812, a pupil 822, and a lens 820. The cornea 812 may be approximately spherical, as shown by the corneal sphere 814. The corneal sphere 814 may have a central point 816 (also referred to as the corneal center) and a radius 818. The hemispherical cornea of the user's eye may be curved around the corneal center 816.
[0174] FIG. 8B-8E An example is shown of using a 3D corneal center estimation module 716 and an eye tracking module 614 to locate the user's corneal center 816.
[0175] like FIG. 8B As shown, the 3D corneal center estimation module 716 can receive an eye-tracking image 852 including a corneal flash 854. The 3D corneal center estimation module 716 can then simulate the known 3D positions of the eye camera 324 and the light source 326 in an eye camera coordinate system 850 (which may be based on data from an eye-tracking extrinsic and intrinsic parameter database 702, a hypothetical eye size database 704, and / or per-user calibration data 706) to project light ray 856 in the eye camera coordinate system. In at least some embodiments, the eye camera coordinate system 850 may have its origin at the 3D position of the eye-tracking camera 324.
[0176] exist FIG. 8C In this process, the 3D corneal center estimation module 716 simulates the corneal bulb 814a (which may be based on an assumed eye size from the database 704) and the corneal curvature center 816a at a first location. The 3D corneal center estimation module 716 can then check whether the corneal bulb 814a will correctly reflect the light from the light source 326 to the flash position 854. FIG. 8C As shown, the first position does not match because ray 860a does not intersect with light source 326.
[0177] Similarly, inFIG. 8D In particular, the 3D cornea center estimation module 716 simulates a corneal sphere 814b and a corneal curvature center 816b at the second position. The 3D cornea center estimation module 716 then checks to see if the corneal sphere 814b correctly reflects light from the light source 326 to the glint location 854. As shown in FIG. 8D The second position also does not match, as shown in
[0178] As shown in FIG. 8E The 3D cornea center estimation module 716 is finally able to determine that the correct position of the corneal sphere is the corneal sphere 814c and the corneal curvature center 816c. The 3D cornea center estimation module 716 confirms that the position shown is correct by checking that the glint 854 from the light source 326 will be correctly reflected off the corneal sphere and imaged by the camera 324 in the image 852. With this arrangement and the known 3D positions of the light source 326, the camera 324, and the optical characteristics (focal length, etc.) of the camera, the 3D cornea center estimation module 716 can determine the 3D position of the corneal curvature center 816 (relative to the wearable system).
[0179] At least in conjunction with FIG. 8C-8E The processes described herein can effectively be iterative, repetitive, or optimization processes to identify the 3D position of the corneal center of the user. As such, any of a variety of techniques (e.g., iterative techniques, optimization techniques, etc.) can be used to efficiently and quickly prune or reduce the search space of possible positions. Moreover, in some embodiments, the system can include two, three, four, or more light sources, such as the light source 326, and some of all of these light sources can be disposed at different positions, resulting in multiple glints, such as the glint 854 located at different positions on the image 852, and multiple light rays (e.g., the light ray 856) having different origins and directions. Such embodiments can enhance the accuracy of the 3D cornea center estimation module 716, as the module 716 can attempt to identify a corneal position that results in some or all of the glints and light rays being correctly reflected between their respective light sources and their respective locations on the image 852. In other words, and in these embodiments, the 3D corneal position determination (e.g., iterative, optimization techniques, etc.) process can rely on the positions of some or all of the light sources. FIG. 8B-8E
[0180] Examples of coordinate systems of normalized eye tracking images
[0181] FIG. 9A-9C It is shown that the 3D corneal position determination process can rely on the positions of some or all of the light sources, such as FIG. 7A This example demonstrates the normalization of the coordinate system of an eye-tracking image by components in a wearable system, such as the coordinate system normalization module 718. Normalizing the coordinate system of the eye-tracking image relative to the user's pupil position compensates for slippage of the wearable system relative to the user's face (i.e., head-mounted receiver slippage), and this normalization establishes a consistent orientation and distance between the eye-tracking image and the user's eyes.
[0182] like FIG. 9A As shown, the coordinate system normalization module 718 can receive the estimated 3D coordinates 900 of the user's corneal rotation center, and can also receive a non-normalized eye-tracking image such as image 852. As an example, eye-tracking image 852 and coordinates 900 can be in a non-normalized coordinate system 850 based on the position of eye-tracking camera 324.
[0183] As the first normalization step, such as FIG. 9B The coordinate system normalization module 718 can rotate coordinate system 850 into a rotated coordinate system 902 so that the z-axis (i.e., the convergence / divergence depth axis) of this coordinate system is aligned with the vector between the origin of the coordinate system and the coordinate 900 of the corneal curvature center. Specifically, the coordinate system normalization module 718 can rotate the eye-tracking image 850 into a rotated eye-tracking image 904 until the coordinate 900 of the user's corneal curvature center is perpendicular to the plane of the rotated image 904.
[0184] As a second normalization step, such as FIG. 9C As shown, the coordinate system normalization module 718 can transform the rotated coordinate system 902 into a normalized coordinate system 910, such that the corneal curvature center coordinate 900 is a standard normalized distance 906 from the origin of the normalized coordinate system 910. Specifically, the coordinate system normalization module 718 can transform the rotated eye-tracking image 904 into a normalized eye-tracking image 912. In at least some embodiments, the standard normalized distance 906 can be approximately 30 mm. If necessary, a second normalization step can be performed before the first normalization step.
[0185] Examples of using eye tracking systems to locate a user’s pupil centroid
[0186] FIG. 9D-9G This illustrates the use of a 3D pupil center locator module 720 and an eye tracking module 614 to locate the user's pupil center (i.e., as shown in the diagram). FIG. 8A The example shown is the center of the user's pupil (822).
[0187] like FIG. 9DAs shown, the 3D pupil center locator module 720 can receive a normalized eye-tracking image 912, which includes a pupil centroid 913 (i.e., the center of the user's pupil identified by the pupil identification module 712). The 3D pupil center locator module 720 can then simulate the normalized 3D position 910 of the eye camera 324 to project light rays 914 in the normalized coordinate system 910 through the pupil centroid 913.
[0188] exist FIG. 9E In the middle, the 3D pupil center locator module 720 can be based on data from the 3D corneal center estimation module 716 (and as combined with...) FIG. 8B-8E (Discussed in more detail) to simulate a corneal bulb 901 such as a corneal bulb with a center of curvature 900. As an example, it can be based on the combination of FIG. 8E The location of the curvature center 816c is identified and based on FIG. 9A-9C The normalization process positions the corneal bulb 901 in the normalized coordinate system 910. Additionally, as... FIG. 9E As shown, the 3D pupil center locator module 720 can identify the first intersection point 916 between the ray 914 (i.e., the ray between the origin of the normalized coordinate system 910 and the normalized position of the user's pupil) and the simulated cornea.
[0189] like FIG. 9F As shown, the 3D pupil center locator module 720 can determine the pupil ball 918 based on the corneal ball 901. The pupil ball 918 may share a common center of curvature with the corneal ball 901, but has a smaller radius. The 3D pupil center locator module 720 can determine the distance between the corneal center 900 and the pupil ball 918 (i.e., the radius of the pupil ball 918) based on the distance between the corneal center and the pupil center. In some embodiments, the distance between the pupil center and the corneal curvature center can be obtained from... FIG. 7A The assumed eye size 704 is determined from an eye-tracking extrinsic and intrinsic parameter database 702, and / or from per-user calibration data 706. In other embodiments, it can be determined from... FIG. 7A The per-user calibration data 706 determines the distance between the pupil center and the corneal curvature center.
[0190] like FIG. 9GAs shown, the 3D pupil center localizer module 720 can locate the 3D coordinates of the center of the user's pupil based on various inputs. As an example, the 3D pupil center localizer module 720 can utilize the 3D coordinates and radius of the pupil sphere 918, the 3D coordinates of the intersection 916 between the simulated corneal sphere 901 and the light ray 914 associated with the pupil centroid 913 in the normalized eye tracking image 912, information about the refractive index of the cornea, and other relevant information (such as the refractive index of air (which can be stored in the eye tracking extrinsic and intrinsic parameter database 702) to determine the 3D coordinates of the center of the user's pupil. In particular, the 3D pupil center localizer module 720 can bend the light ray 916 into a refracted light ray 922 in simulation based on the refractive difference between air (a first refractive index of approximately 1.00) and the corneal material (a second refractive index of approximately 1.38). After accounting for the refraction caused by the cornea, the 3D pupil center localizer module 720 can determine the 3D coordinates of the first intersection 920 between the refracted light ray 922 and the pupil sphere 918. The 3D pupil center localizer module 720 can determine the center of the user's pupil 920 to be located at the approximate first intersection 920 between the refracted light ray 922 and the pupil sphere 918. With this arrangement, the 3D pupil center localizer module 720 can determine the 3D position of the pupil center 920 (relative to the wearable system) in the normalized coordinate system 910. If desired, the wearable system can non-normalize the coordinates of the pupil center 920 to the original eye camera coordinate system 850. The pupil center 920 can be used with the corneal curvature center 900 to determine the user's optical axis, among other things, using the optical axis determination module 722, and to determine the user's vergence depth using the vergence depth estimation module 728.
[0191] Examples of differences between optical and visual axes
[0192] As discussed in connection with the optical-to-visual mapping module 730, FIG. 7A The user's optical axis and visual axis are typically misaligned, in part because the user's visual axis is defined by their fovea, and the fovea is typically not centered on the person's retina. Thus, when a person wishes to focus their attention on a particular object, the person aligns their visual axis with the object to ensure that light from the object falls on their fovea, while their optical axis (defined by the center of their pupil and the center of curvature of their cornea) is actually slightly offset from the object. FIG. 10 is an example of an eye 1000 showing the optical axis 1002 of the eye, the visual axis 1004 of the eye, and the offset between these axes. In addition, FIG. 10The pupil center of the eye 1006, the corneal curvature center of the eye 1008, and the average center of rotation (CoR) of the eye 1010 are shown. In at least some populations, the corneal curvature center of the eye 1008 can be located about 4.7 mm in front, as indicated by the dimension 1012 of the average center of rotation (CoR) of the eye 1010. Further, the center of perspective of the eye 1014 can be located about 5.01 mm in front of the corneal curvature center of the eye 1008, about 2.97 mm behind the outer surface 1016 of the cornea of the user, and / or directly in front of the pupil center 1006 of the user (e.g., corresponding to a location within the anterior chamber of the eye 1000). As further examples, the dimension 1012 can be between 3.0 mm and 7.0 mm, between 4.0 mm and 6.0 mm, between 4.5 mm and 5.0 mm, or between 4.6 mm and 4.8 mm, or any range between any of these values and any values within any of these ranges. The center of perspective of the eye (CoP) 1014 can be a useful location for a wearable system, as in at least some embodiments, rendering a camera registered at the CoP can help reduce or eliminate parallax artifacts.
[0193] FIG. 10 Also shown is a location within the human eye 1000 with which a pinhole of a rendering camera can be aligned. As shown, a pinhole of a rendering camera can be aligned with the location 1014 along the optical axis 1002 or the visual axis 1004 of the human eye 1000 that is closer to the outer surface of the cornea than (a) the center of the pupil or iris 1006 and (b) the corneal curvature center 1008 of the human eye 1000. For example, as shown, a pinhole of a rendering camera can be registered with the location 1014 along the optical axis 1002 of the human eye 1000 that is about 2.97 mm behind the outer surface of the cornea 1016 and about 5.01 mm in front of the corneal curvature center 1008. The location 1014 of the pinhole of the rendering camera and / or the anatomical region of the human eye 1000 to which the location 1014 corresponds can be considered to represent the center of perspective of the human eye 1000. As shown in FIG. 10 FIG. 10 Example process of rendering content and checking registration based on eye tracking
[0194] FIG. 11
[0195] FIG. 3 is a process flow diagram for an example method 1100 for using eye tracking in rendering content and providing feedback about registration in a wearable device. The method 1100 can be performed by the wearable systems described herein. Embodiments of the method 1100 can be used by a wearable system to render content and provide feedback about registration (i.e., fit of the wearable device to the user) based on data from an eye tracking system.
[0196] At block 1110, the wearable system can capture images of one or both eyes of the user. The wearable system can use one or more eye cameras 324 to capture the eye images, as shown in the examples at least in FIGS. 6A and 6B. If desired, the wearable system can also include one or more light sources 326 configured to shine IR light on the user’s eyes and produce corresponding glints in the eye images captured by the eye cameras 324. As discussed herein, the glints can be used by the eye tracking module 614 to derive various pieces of information about the user’s eyes, including where the eyes are looking. FIG. 7A
[0197] At block 1120, the wearable system can detect glints and pupils in the eye images captured in block 1110. As an example, block 1120 can include processing the eye images by a glint detection and labeling module 714 to identify two-dimensional locations of glints in the eye images and by a pupil identification module 712 to identify two-dimensional locations of pupils in the eye images.
[0198] At block 1130, the wearable system can estimate three-dimensional locations of the user’s left and right corneas relative to the wearable system. As an example, the wearable system can estimate locations of centers of curvature of the user’s left and right corneas and distances between these centers of curvature and the user’s left and right corneas. Block 1130 can involve the 3D cornea center estimation module 716 described herein at least in connection with FIG. 7A and 8A -8E.
[0199] At block 1140, the wearable system can estimate three-dimensional locations of the user’s left and right pupil centers relative to the wearable system. As an example, the wearable system and 3D pupil center localizer module 720 can estimate, among other things, locations of the user’s left and right pupil centers, as described at least in connection with FIG. 7A and 9D -9G.
[0200] At block 1150, the wearable system can estimate three-dimensional locations of the user’s left and right centers of rotation (CoRs) relative to the wearable system. As an example, the wearable system and CoR estimation module 724 can estimate, among other things, locations of the CoRs of the user’s left and right eyes, as described at least in connection with FIG. 7B and 10 As described. As a specific example, a wearable system can locate the CoR of the eye by tracing back along the optical axis from the center of curvature of the cornea to the retina.
[0201] At box 1160, the wearable system can estimate the user's IPD, depth of convergence, center of view (CoP), optical axis, visual axis, and other desired attributes from eye-tracking data. As an example, IPD estimation module 726 estimates the user's IPD by comparing the 3D positions of the left and right CoRs; depth of convergence estimation module 728 estimates the user's depth by finding the intersection (or near-intersection) of the left and right optical axes or the intersection of the left and right visual axes; optical axis determination module 722 identifies the left and right optical axes over time; optical axis-to-visual-axis mapping module 730 identifies the left and right visual axes over time; and CoP estimation module 732 identifies the left and right center of view, as part of box 1160.
[0202] At box 1170, the wearable system may render content in part based on eye-tracking data identified in boxes 1120-1160, and may optionally provide feedback on registration (i.e., the wearable system's adaptation to the user's head). As an example, the wearable system may identify the appropriate location of the rendering camera and then generate content for the user based on the rendering camera's location, such as in conjunction with... Examples of registration coordinate systems The light field rendering controller 618 and rendering engine 622 are discussed. As another example, a wearable system may determine whether it is correctly fitted to the user, or relative to the user having slipped from its correct position, and may provide the user with optional feedback indicating whether the device's fit needs adjustment, as discussed in conjunction with registration observer 620. In some embodiments, the wearable system may adjust the rendered content based on incorrect or less-than-ideal registration in an attempt to reduce, minimize, or compensate for the effects of incorrect or unregistered settings.
[0203] FIG. 12A-12B
[0204] FIG. 12A An example eye position coordinate system is shown, which can be used to define the three-dimensional positions of a user's left and right eyes relative to the display of the wearable system described herein. As an example, the coordinate system may include axes X, Y, and Z. The z-axis of the coordinate system may correspond to depth, such as the distance between the plane containing the user's eyes and the plane containing the display 220 (e.g., a direction perpendicular to the plane of the user's face). The x-axis of the coordinate system may correspond to the left-right direction, such as the distance between the user's left and right eyes. The y-axis of the coordinate system may correspond to the up-down direction, which can be the vertical direction when the user is upright.
[0205] FIG. 2 The user's eye 1200 and display surface 1202 are shown (which may be...) FIG. 12BA side view of a portion of the display 220, while FIG. 4 A top view of a user's eye 1200 and a display surface 1202 is shown. The display surface 1202 may be located in front of the user's eye and may output image light to the user's eye. As an example, the display surface 1202 may include one or more outgoing optical elements, active or pixel display elements, and may be part of a waveguide stack, such as... FIG. 12A The stacked waveguide assembly 480. In some embodiments, the display surface 1202 may be planar. In some other embodiments, the display surface 1202 may have other topologies (e.g., curved). It should be understood that the display surface 1202 may be the physical surface of the display, or simply a planar or other imaginary surface from which image light is understood to propagate from the display 220 to the user's eye.
[0206] like FIG. 12A As shown, the user's eye 1200 may have an actual position 1204 offset from the nominal position 1206, and the display surface 1202 may be in position 1214. FIG. 12A The corneal apex 1212 of the user's eye 1200 is also shown. The user's line of sight (e.g., their optical axis and / or visual axis) can be substantially along the line between the actual location 1204 and the corneal apex 1212. FIG. 12B and FIG. 14 As shown, the actual position 1204 may be offset from the nominal position 1206 by z-offset 1210, y-offset 1208, and x-offset 1209. The nominal position 1206 may represent a preferred position (sometimes referred to as the design position, which is typically centered within the required volume) for the user's eye 1200 relative to the display surface 1202. As the user's eye 1200 moves away from the nominal position 1206, the performance of the display surface 1202 may degrade, as illustrated herein by example. FIG. 12A The subject of discussion.
[0207] It will be appreciated that a point or volume associated with the user's eye 1200 can be used to represent the position of the user's eye in the registration analysis herein. The representing point or volume can be any point or volume associated with the eye 1200 and is preferably used consistently. For example, the point or volume may be on or in the eye 1200, or may be placed away from the eye 1200. In some embodiments, the point or volume is the center of rotation of the eye 1200. The center of rotation can be determined as described herein and can have the advantage of simplifying the registration analysis because it is arranged generally symmetrically on the respective axes within the eye 1200 and allows a single display registration volume aligned with the optical axis to be used for analysis.
[0208] Example plot of rendering content in response to user eye motionIt is also shown that display surface 1202 can be centered below the user's field of view (when looking straight ahead along the y-axis, their optical axis is parallel to the ground) and can be tilted (relative to the y-axis). In particular, display surface 1202 can be positioned some distance below the user's field of view so that when eye 1200 is in position 1206, the user will have to look down at approximately angle 1216 to see the center of display surface 1202. This can facilitate a more natural and comfortable interaction with display surface 1202, particularly when viewing content rendered at a short depth (or distance from the user), as the user can be more comfortable viewing content below the field of view than above the field of view. Furthermore, display surface 1202 can be tilted, for example, at angle 1218 (relative to the y-axis) so that when the user is gazing at the center of display surface 1202 (e.g., gazing slightly below the user's field of view), display surface 1202 is generally perpendicular to the user's line of sight. In at least some embodiments, display surface 1202 can also be shifted left or right (e.g., along the x-axis) relative to the user's eye's nominal position. As an example, the left eye display surface can be shifted right and the right eye display surface can be shifted left (e.g., display surfaces 1202 can be shifted toward each other) so that the user's line of sight hits the center of the display surface when focusing at some distance less than infinity, which can increase user comfort during typical use on a wearable device.
[0209] FIG. 13
[0210] FIG. 4 includes a set of example graphs 1200a-1200j that illustrate how a wearable system can switch depth planes in response to a user's eye movements. As discussed herein in connection with FIG. 13 and FIG. 7, a wearable system can include multiple depth planes, where various depth planes are configured to present content to a user at different simulated depths or with different accommodative cues (i.e., with various levels of wavefront curvature or light ray divergence). As an example, a wearable system can include a first depth plane configured to simulate a first depth range and a second depth plane configured to simulate a second depth range, and while these two ranges can overlap as needed to facilitate hysteresis in switching, the second depth range can generally extend to a greater distance from the user. In such embodiments, a wearable system can track a user's vergence depth, saccadic movements, and blinks to avoid excessive depth plane switching, excessive accommodative-vergence mismatch, and periods of accommodative-vergence mismatch that are too long and attempt to reduce the visibility of depth plane switching (i.e., by shifting depth planes during blinks and saccades) between the first depth plane and the second depth plane.
[0211] Graph 1200a shows an example of a user's vergence depth over time. Graph 1200b shows an example of a user's saccade signal or eye movement velocity over time.
[0212] Graph 1200c can show vergence depth data generated by eye tracking module 614, and in particular, data generated by vergence depth estimation module 728. As shown in graphs 1200c-1200h, eye tracking data can be sampled at a rate of approximately 60 Hz in eye tracking module 614. As shown between graphs 1200b and 1200c, eye tracking data within eye tracking module 614 can lag behind the user's actual eye movement by a delay 1202. For example, at time tl, the user's vergence depth can cross a lag threshold 1210a, but eye tracking module 614 can not recognize this event until time t2, which is after delay 1202.
[0213] Graph 1200c also shows various thresholds 1210a, 1210b, 1210c in the lag band, which can be associated with a transition between the first depth plane and the second depth plane (i.e., depth plane #1 and depth plane #0 in FIG. 6 In some embodiments, the wearable system can attempt to display content with depth plane #1 when the user's vergence depth is greater than threshold 1210b, and display content with depth plane #0 when the user's vergence depth is less than threshold 1210b. However, to avoid excessive switching, the wearable system can implement a hysteresis, whereby the wearable system will not switch from depth plane #1 to depth plane #0 until the user's vergence depth crosses outer threshold 1210c. Similarly, the wearable system can not switch from depth plane #0 to depth plane #1 until the user's vergence depth crosses outer threshold 1210a.
[0214] Graph 1200d shows an internal flag that can be generated by depth plane selection module 750 or lag band crossing detection module 752, indicating whether the user's vergence depth is in the volume typically associated with depth plane #1 or the volume typically associated with depth plane #2 (i.e., whether the user's vergence depth is greater than or less than threshold 1210b).
[0215] Graph 1200e shows an internal lag band flag that can be generated by the depth plane section module 750 or the lag band crossing detection module 752, indicating whether the vergence depth of the user has crossed an outer threshold, such as threshold 1210a or 1210c. In particular, graph 1200e shows a flag indicating whether the vergence depth of the user has crossed the lag band entirely and entered a region outside the volume of the active depth plane (i.e., entered a region associated with a depth plane other than the active depth plane), thereby potentially causing an unwanted accommodation-vergence mismatch (AVM).
[0216] Graph 1200f shows an internal AVM flag that can be generated by the depth plane selection module 750 or the lag band crossing detection module 752, indicating whether the vergence of the user has been outside the volume of the active depth plane for more than a predetermined time. Thus, the AVM flag can identify when the user can have experienced an unwanted accommodation-vergence mismatch for a near-excessive or excessive period of time. Additionally or alternatively, the internal AVM flag can also indicate whether the vergence of the user has reached a predetermined distance outside the volume of the active depth plane, thereby creating a potentially excessive accommodation-vergence mismatch. In other words, the AVM flag can indicate when the vergence of the user has exceeded other thresholds further from threshold 1210b than thresholds 1210a and 1210c.
[0217] Graph 1200g shows an internal blink flag generated by the eye event detection module 754, which can determine when the user has blinked or is blinking. As described herein, it can be desirable to switch depth planes when the user blinks to reduce the likelihood that the user perceives a depth plane switch.
[0218] Graph 1200h shows example output from the depth plane selection module 750. In particular, graph 1200h shows that the depth plane selection module 750 can output instructions to a rendering engine, such as rendering engine 622 (see FIG. 13 ), to use the selected depth plane, which can change over time.
[0219] Graphs 1200i and 1200j show the latency that can exist in the wearable system, including the latency of the rendering engine 622 to switch depth planes and the latency of the display 220, which can require providing light associated with a new image frame in the new depth plane to effectuate the change in depth plane.
[0220] Reference will now be made to the events shown in graphs 1200a-1200j at different times (t0-t 10 ).
[0221] At some time around time to, the user's vergence depth can cross threshold 1210a, which can be an outer lag threshold. After a delay associated with image capture and signal processing, the wearable system can generate a signal indicating that the user's vergence depth is within the lag band, as shown in graph 1200e. In the example of graph 1200e, eye tracking module 614 can present a lag band exceeded flag at approximately time ti in relation to the user's vergence depth crossing threshold 1210a.
[0222] From time to to approximately time t4, the user's vergence depth can continue to decrease and then can increase.
[0223] At time ti, the user's vergence depth can cross threshold 1210b, which can be a midpoint between two depth planes, such as depth plane #1 and depth plane #0. After a process delay 1202, eye tracking module 614 can change an internal flag indicating that the user's vergence depth has moved from a volume typically associated with depth plane #1 to a volume typically associated with depth plane #0, as shown in graph 1200d.
[0224] At time t3, eye tracking module 614 can determine that the user's vergence depth (as shown in graph 1200a) has moved completely past the lag band and crossed outer threshold 1210c. As a result, eye tracking module 614 can generate a signal indicating that the user's vergence depth is outside the lag band, as shown in graph 1200e. In at least some embodiments, eye tracking module 614 can only switch between a first depth plane and a second depth plane when the user's vergence depth is outside the lag band between the first depth plane and the second depth plane.
[0225] In at least some embodiments, eye tracking module 614 can be configured to switch depth planes at time t3. In particular, eye tracking module 614 can be configured to switch depth planes based on a determination that the vergence depth has moved from a volume of a currently selected depth plane (depth plane #1 as shown in graph 1200h) to a volume of another depth plane (depth plane #0) and completely across the lag band. In other words, eye tracking module 614 can implement a depth plane switch whenever the lag band is exceeded (graph 1200e is high) and an accommodation-vergence mismatch based on a mismatch time or a mismatch amount is detected (graph 1200f is high). In such embodiments, eye tracking module 614 can provide a signal to rendering engine 622 instructing rendering engine 622 to switch to another depth plane (depth plane #0). However, in FIG. 13In the example of FIG. 12a, the eye tracking module 614 can be configured to switch the depth plane. In particular, the eye tracking module 614 can determine that the user’s vergence has been in the volume associated with depth plane #0 for more than a predetermined time threshold (and optionally, also outside the hysteresis band for that time period). Examples of the predetermined time threshold include 5 seconds, 10 seconds, 20 seconds, 30 seconds, 1 minute, and 90 seconds, and any range between any of these values. In response to this determination, the eye tracking module 614 can generate the AVM flag, as shown in plot 1200f, and direct the rendering engine 622 to switch to depth plane #0, as shown in plot 1200h. In some embodiments, the eye tracking module 614 can generate the AVM flag and direct the rendering engine 622 to switch the depth plane if it detects that the user’s vergence depth is more than a threshold distance away from the currently selected depth volume.
[0226] At time t4 and after the delay 1204, the rendering engine 622 can begin rendering content at the newly selected depth plane #0. After a delay 1206 associated with rendering through the display 220 and delivering light to the user, the display 220 can fully switch to the newly selected depth plane #0 by time t6. Example process of calibration of depth plane selection In the example of FIG. 12a, the eye tracking module 614 can be configured to switch the depth plane. In particular, the eye tracking module 614 can determine that the user’s vergence has been in the volume associated with depth plane #0 for more than a predetermined time threshold (and optionally, also outside the hysteresis band for that time period). Examples of the predetermined time threshold include 5 seconds, 10 seconds, 20 seconds, 30 seconds, 1 minute, and 90 seconds, and any range between any of these values. In response to this determination, the eye tracking module 614 can generate the AVM flag, as shown in plot 1200f, and direct the rendering engine 622 to switch to depth plane #0, as shown in plot 1200h. In some embodiments, the eye tracking module 614 can generate the AVM flag and direct the rendering engine 622 to switch the depth plane if it detects that the user’s vergence depth is more than a threshold distance away from the currently selected depth volume.
[0227] At time t5 and after the delay 1204, the rendering engine 622 can begin rendering content at the newly selected depth plane #0. After a delay 1206 associated with rendering through the display 220 and delivering light to the user, the display 220 can fully switch to the newly selected depth plane #0 by time t6.
[0228] Thus, plots 1200a-j show how the system responds to the user’s changing vergence between times t0 and t6, and how it switches the depth plane after the user’s vergence has moved away from the previous depth volume for more than a predetermined period of time. Plots 1200a-j can show how the system responds to the user’s changing vergence between times t7 and t 10
[0229] At time t7, the eye tracking module 614 can detect that the user’s vergence depth has entered the hysteresis region between depth plane #0 and depth plane #1 (i.e., the user’s vergence depth has crossed the outer threshold 1210c). In response, the eye tracking module 614 can change the hysteresis flag, as shown in plot 1200e.
[0230] At time t8, the eye tracking module 614 can detect that the vergence depth of the user has crossed the threshold 1210b and has moved from the volume typically associated with depth plane #0 to the volume typically associated with depth plane #1. Accordingly, the eye tracking module 614 can change the depth volume flag, as shown in graph 1200d.
[0231] At time t9, the eye tracking module 614 can detect that the vergence depth of the user has crossed the threshold 1210a and moved out of the hysteresis volume into a volume typically associated with depth plane #1 only. In response, the eye tracking module 614 can change the hysteresis flag, as shown in graph 1200e.
[0232] At time t 10 At times t10 and t11, the user can blink, and the eye tracking module 614 can detect the blink. As one example, the eye event detection module 754 can detect the blink of the user. In response, the eye tracking module 614 can generate a blink flag, as shown in graph 1200h. In at least some embodiments, the eye tracking module 614 can implement a depth plane switch whenever the hysteresis band is exceeded (graph 1200e is high) and a blink is detected (graph 1200g is high). Accordingly, the eye tracking module 614 can instruct the rendering engine 622 to switch depth planes at time t 10 switches depth planes.
[0233] FIG. 2
[0234] As discussed herein, a head-mounted display such as display 220 can include multiple depth planes, each providing a different amount of wavefront divergence to provide different accommodation cues to the user's eyes. For example, the depth planes can be formed by optical elements such as waveguides 432b, 434b, 436b, and 440b that can be configured to send image information to the user's eyes with a desired level of wavefront divergence. FIG. 4 As discussed herein, a head-mounted display such as display 220 can include multiple depth planes, each providing a different amount of wavefront divergence to provide different accommodation cues to the user's eyes. For example, the depth planes can be formed by optical elements such as waveguides 432b, 434b, 436b, and 440b that can be configured to send image information to the user's eyes with a desired level of wavefront divergence. FIG. 7A
[0235] In at least some embodiments, a wearable system including display 220 can be configured to display image content with accommodation cues based on a current point of regard or vergence depth of a user's gaze (e.g., to reduce or minimize accommodation-vergence mismatch). In other words, the wearable system can be configured to identify a vergence depth of a user's gaze (e.g., using eye tracking module 614) and display image content with accommodation cues based on the vergence depth. FIG. 7A accommodation cues associated with the current vergence depth. Thus, when the user gazes at optical infinity, the wearable system can display image content on a first depth plane that provides accommodation cues for optical infinity. In contrast, when the user looks toward a near field (e.g., within one meter), the wearable system can display image content on a second depth plane that provides accommodation cues for within or at least close to the near field.
[0236] As previously noted, the vergence depth estimation that can be performed by the vergence depth estimation module 728 can be based in part on the inter-pupillary distance (IPD) of the current user. In particular and in some embodiments, determining the vergence depth can involve projecting the optical and / or visual axes of the user's left and right eyes (to determine their respective gazes) and determining where these axes intersect in space, thereby determining the user's point of gaze or the location of the vergence depth. Geometrically, the optical and / or visual axes are the legs of a triangle, the base of the triangle is the user's IPD, and the tip of the triangle is the user's point of gaze or vergence depth, and thus it should be appreciated that the user's IPD is useful in determining the vergence depth. FIG. 7A
[0237] In various embodiments, the wearable system can calibrate for a particular primary user. This calibration can include various processes and can include determining how the user's eyes move when the user focuses on objects at different locations and depths. The calibration can also include identifying the user's IPD (e.g., the distance between the user's pupils when the user focuses on optical infinity). The calibration can also include determining the user's pupil distance when the user focuses on objects that are closer than optical infinity, such as objects in a near field (e.g., less than 2.0 meters) and objects in a mid-field (e.g., between about 2.0 and 3.0 meters). Based on such calibration data, the wearable system can be able to determine the depth at which the user is looking by monitoring the user's pupil distance. In other words, when the user's pupil distance is at its maximum (e.g., equal to or close to the user's IPD), the wearable system can be able to infer that the user's vergence distance is at or close to optical infinity. In contrast, when the user's pupil distance is close to its minimum pupil distance, the wearable system can be able to infer that the user's vergence distance is close to the user as determined by this calibration.
[0238] In some embodiments, the wearable system can utilize one or more alternative processes to select a depth plane. As one example, the wearable system can implement a content-based switching scheme. It will be appreciated that virtual content can include information about a location in virtual space where the content should be located. Given this location, the virtual content can effectively specify an associated amount of wavefront divergence. Rather than determining a fixation point of a user's eye to switch depth planes (e.g., to switch the amount of wavefront divergence of light used to form a virtual object), the display system can be configured to switch depth planes based on the desired location in virtual space where the virtual content is placed.
[0239] In some embodiments, the wearable system can still determine whether the user is viewing virtual content in order to switch to a depth plane specified for that virtual content. For example, the display system can still track a user's gaze to determine whether they are viewing a virtual object, and once a determination is made, the depth information associated with the virtual content can be used to determine whether to switch depth planes. As another example, the wearable system can identify the most likely real or virtual object that the user is viewing based on an assumption that the user will be viewing a particular real or virtual object. For example, the wearable system can present a video to a user on a 2D virtual screen that is 1 meter away from the user. Although the user can be looking away from the screen to look at another object, it can be reasonable to assume that the user will be fixated on the video screen. In some embodiments, the wearable system can be configured to make an assumption that the user is viewing real or virtual content that has motion or visual changes, or more motion or changes than other real or virtual content; for example, the wearable system can assign a score to the amount of motion or changes in visual appearance of real or virtual content within the user's field of view, and assume that the user is viewing the real or virtual content with the highest score (e.g., the most motion or visual changes, such as a virtual screen displaying a video).
[0240] Another example of an alternative depth plane selection process is dynamic calibration. Dynamic calibration can be beneficial when the current user is not (or has not) performed a dedicated calibration process. As an example, dynamic calibration can be used when a guest user is wearing the device. In one example of a dynamic calibration system, the wearable system can collect eye tracking data in order to estimate the IPD of the current user, and then the estimated IPD (and its eye gaze direction, as discussed in relation to module 728 of FIG. 7) can be used to estimate the vergence depth of the user. The IPD estimate can be made by FIG. 7A FIG. 14 IPD estimation module 726 performs, and additional details and example embodiments are discussed herein in relation to module 726. Dynamic calibration can occur as a background process, without requiring specific action by the user. Further, dynamic calibration can continuously obtain samples or images of the user's eyes to further refine the IPD estimate. As discussed in further detail below, the wearable system can estimate the user's IPD as the 95th percentile (or other percentile) of all measured IPD values. In other words, the largest 5% of measured IPD values can be excluded, and then the largest remaining measured IPD value can be taken as the user's IPD. The IPD value calculated in this way is referred to herein as IPD 95.
[0241] It will be appreciated that the display system can be configured to continuously monitor IPD. In this way, the number of samples or individual IPD measurements used to determine a value associated with a particular percentile can increase over time, and potentially improve the accuracy of the IPD determination. In some embodiments, the IPD value (e.g., IPD 95) can be continuously or periodically updated, e.g., the IPD value can be updated after a predetermined amount of time has elapsed and / or after a predetermined number of individual IPD measurements have been taken.
[0242] In some embodiments, the dynamic calibration process can seek to identify the user's IPD (e.g., the user's maximum pupillary distance, such as when the user is looking at optically infinite). In such embodiments, the wearable system can be able to calibrate depth plane selection based on the IPD alone.
[0243] As a particular example, it has been determined that if a user's pupillary distance is reduced from their maximum IPD by 0.6 mm, then it is likely that the user will focus at a depth of approximately 78 mm. In some embodiments disclosed herein, 78 mm corresponds to a switching point between depth planes (e.g., the system can tend to use a first depth plane when the user is focusing at less than 78 mm, and a second depth plane when the user is focusing at greater than 78 mm). In embodiments in which switching points occur at different focal depths, the reduction in associated pupillary distance relative to their maximum IPD will change relative to the change in switching points (e.g., in such embodiments, from 78 mm to any switching point).
[0244] In some cases, the wearable system can refine its calculation of a user's vergence depth by considering not only the difference between the user's current pupillary distance and their maximum IPD, but also how that relationship varies relative to the user's maximum IPD. In particular, the 0.6 mm number discussed above can be an average number that applies to the general population and that accounts for various biases in the wearable system (e.g., on average, a user whose current pupillary distance is 0.6 mm less than their maximum IPD can be close to a distance of 78 mm). However, the actual IPD difference (between maximum and current) associated with a vergence distance of 78 mm (or other switch point distance as discussed in the previous paragraph) can be more or less than 0.6 mm, and can vary depending on the user's anatomical IPD. As particular examples, a person with an IPD of 54 mm can have a vergence distance of 78 mm when their current IPD is 0.73 mm less than their maximum IPD (e.g., when they are looking at a distance of at least 10 meters), a person with an IPD of 64 mm can have a vergence distance of 78 mm when their current IPD is 0.83 mm less than their maximum IPD, and a person with an IPD of 72 mm can have a vergence distance of 78 mm when their current IPD is 0.93 mm less than their maximum IPD. These numbers can differ from the 0.6 mm number for various reasons including, but not limited to, the 0.6 mm number failing to distinguish between users with different IPDs, and the 0.6 mm number referencing an IPD95 value (e.g., can reference an IPD value that is actually slightly lower than the user's anatomical IPD when looking at optical infinity). In some embodiments, the 0.6 mm number can vary depending on the user's maximum IPD. For example, a predetermined number of biases from the 0.6 mm number can be available and can be associated with different ranges of maximum IPD values.
[0245] Using such a relationship, the wearable system can be able to determine a user's current vergence depth by comparing the user's maximum IPD to their current pupillary distance (which can decrease as their vergence distance decreases) according to one or more mathematical functions. As an example, determining a user's maximum IPD can involve collecting data about a user's IPD over time and identifying the maximum IPD in the collected data. In other embodiments, the wearable system can use heuristics or other processes to determine a user's maximum IPD even when the user is not gazing at optical infinity. In particular, the wearable system can extrapolate a user's maximum IPD from a plurality of pupillary distances associated with closer vergence distances. The wearable system can also encourage the user to look at optical infinity by presenting virtual content at optical infinity and asking the user to focus their attention on the virtual content.
[0246] FIG. 14is a process flow diagram of an example method 1400 for selecting a depth plane using an existing calibration, a content-based switching scheme, and / or dynamic calibration. The method 1400 can be performed by the wearable system described herein. Embodiments of the method 1400 can be used by the wearable system to render content on a depth plane that generally reduces or minimizes any vergence-accommodation mismatch that can otherwise cause user discomfort and fatigue.
[0247] At block 1402, the wearable system can determine that it is not being worn by a user. The wearable system can use one or more sensors, such as an eye tracking system like the eye cameras 324, to determine that it is not being worn. As an example, the wearable system can determine that it is not being worn by a user based on a determination that there are no eyes in eye tracking images captured by the eye cameras 324. In particular, after failing to detect an eye in the eye tracking images for at least a given period of time (e.g., a predetermined period of time, a dynamically determined period of time, etc.), the wearable system can determine that the wearable system is not being worn by a user.
[0248] Upon detecting one or more eyes of a user in eye tracking images such as those captured by the cameras 324, the method 1400 can move to block 1404. In block 1404, the wearable system can determine that it is being worn. In some embodiments, some or all of the determinations associated with block 1402 and / or block 1406 can be made based at least in part on data from one or more other sensors of the wearable system (e.g., IMUs, accelerometers, gyroscopes, proximity sensors, touch sensors, etc.). For example, in these embodiments, the wearable system can monitor data from one or more IMUs, accelerometers, and / or gyroscopes for indications that the wearable system has been placed on or removed from the head of a user, can monitor data from one or more proximity sensors and / or touch sensors to detect the physical presence of a user, or both. For example, the wearable system can compare data received from one or more of these sensors, and the wearable system can have thresholds associated with respective sensors. In one example, the wearable system can then determine that the device has been removed from the head of a user based on a value from one or more proximity sensors reaching or exceeding a threshold, optionally in combination with data from an IMU, accelerometer, and / or gyroscope indicating sufficient motion (e.g., exceeding a threshold) to support a conclusion that the device has been removed.
[0249] At block 1406, the wearable system can attempt to identify the current user by performing an identification process. As one example, the wearable system can estimate the IPD of the current user to determine whether the IPD of the current user matches the IPD of the calibrated user (e.g., whether the two IPDs are within some threshold of each other). In some embodiments, the wearable system can determine that the current user is the calibrated user if the calibrated IPD and the IPD of the current user are within a threshold of, e.g., 0.5 mm, 1.0 mm, 1.5 mm, or 2.0 mm, of the IPD of the calibrated user. Generally, larger thresholds can facilitate faster determinations and help ensure that the calibrated user is identified and their calibration parameters are used; e.g., larger thresholds favor finding that the current user is the calibrated user. In at least some embodiments, the wearable system can be able to determine the IPD of the current user with relatively high accuracy (e.g., 95%) within a relatively short timeframe (e.g., 5 to 7 seconds of eye tracking data, which can correspond to between approximately 150 to 200 frames of eye tracking images).
[0250] If there is a match between the IPD of the current user and the IPD of the calibrated user, the wearable system can assume that the current user is the calibrated user and load the existing calibration at block 1408. The existing calibration can be the calibration parameters or data generated during the calibration process in which the calibrated user wore the wearable system. If the current user is not actually the calibrated user, but just has a similar IPD, the calibration data loaded at block 1408 can be an acceptable calibration to provide reasonable performance for the current user while allowing the current user to use the wearable system without performing a more detailed calibration.
[0251] In some embodiments, the wearable system can use measurements other than (or in addition to) the IPD of the user to identify the current user. As examples, the wearable system can ask the user for a username and / or password, the wearable system can perform an iris scan (e.g., compare a current image of the user’s iris to a reference image to determine whether there is a match, interpreting a match as indicating that the current user is the calibrated user), voice recognition (e.g., compare a current sample of the user’s voice to a reference voice file, interpreting a match as indicating that the current user is the calibrated user), or some combination of these and other authentication or identification techniques. In some embodiments, two or more identification processes can be performed to improve the accuracy of determining whether the current user is the calibrated user. For example, the current user can be assumed to be the calibrated user, but if all of the identification processes performed do not identify the current user as the calibrated user, then the current user can be determined not to be the calibrated user. As another example, the results from the various identification processes performed can be aggregated into a combined score, and the current user can be assumed to be the calibrated user unless the combined score exceeds a predetermined threshold.
[0252] In some embodiments, if the display system includes multiple calibration files for multiple users, additional criteria can be needed to select the appropriate calibration file, e.g., the user can be prompted to select the appropriate calibration file, and / or multiple of the identification schemes disclosed herein can be utilized.
[0253] In some embodiments, multiple identification schemes can be utilized to improve the accuracy of user identification. For example, it will be appreciated that IPD is a relatively coarse identification criterion. In some embodiments, IPD can be used as a first criterion to identify whether the current user is likely to be a calibrated user, and then a more precise or accurate identification scheme (e.g., iris scanning) can be used. Such a multi-step identification scheme advantageously can save processing resources, as more precise identification schemes can take more resources. Thus, processing resources can be saved by delaying use of the more accurate, resource-intensive identification scheme until IPD determines that there is a calibrated user present.
[0254] Continuing with reference to FIG. 16A through FIG. 18 If the IPD of the current user does not match the IPD of the calibrated user, the wearable device can perform dynamic calibration in block 1410. As discussed herein, dynamic calibration can involve monitoring eye tracking data to estimate how a user's interpupillary distance varies as a function of their vergence distance. As one example, dynamic calibration can involve estimating a user's maximum IPD, and then using the maximum IPD together with the current interpupillary distance to estimate the current vergence distance.
[0255] In block 1412, the wearable system can implement content-based switching. With content-based switching, the depth plane selection is based on a determination of the depth of any virtual content that the user is looking at (e.g., the most important or most interesting content being displayed, which can be identified by the content creator; based on the user's gaze from their eyes, etc.). In at least some embodiments, block 1412 can be performed whenever selected by the content creator or other designer of the wearable system and regardless of whether the current user is a calibrated user. In various embodiments, the content-based switching in block 1412 can be performed when there is no calibrated user. In such embodiments, block 1406 can be skipped if desired. Additionally, in some embodiments, the content-based switching of block 1412 can be performed regardless of whether blocks 1408 and / or 1410 (and blocks related to these blocks) are performed or available to the display system; e.g., in some embodiments, the display system can perform only block 1412 to determine depth plane switching.
[0256] As described herein, each virtual object can be associated with location information, such as 3D location information. The wearable system can then present each virtual object to the user based on this location information. For example, the location information of a particular virtual object can indicate the X, Y, and Z coordinates on which the object will be presented (e.g., the center or centroid of the object can be presented at these coordinates). Therefore, the wearable system can obtain information indicating the depth plane on which each virtual object will be presented.
[0257] As will be referred to below FIG. 14 In more detail, the wearable system can assign appropriate spatial volumes (also referred to herein as “regions” or markers) around each object. These spatial volumes can preferably not overlap. The wearable system can recognize the user’s gaze and identify the region that includes the user’s gaze. For example, the gaze can indicate the three-dimensional location (e.g., an approximate three-dimensional location) that the user is looking at. The wearable system can then present virtual content on a depth plane associated with the virtual object contained within the identified region. Thus, in some embodiments, once the display system determines which object the user is looking at, switching between depth planes can occur based on the location or depth plane associated with the virtual object rather than the user’s gaze point.
[0258] In box 1414, the wearable system can perform depth plane switching. Depth plane switching in box 1414 can be performed using configuration parameters loaded or generated in boxes 1408, 1410, or 1412. Specifically, if the current user is a calibrated user, box 1414 can perform depth plane switching based on configuration parameters generated during calibration with that user. If the current user is not identified as a calibrated user and dynamic calibration was performed in box 1410, box 1412 can involve depth plane switching based on calibration parameters generated in box 1410 as part of the dynamic calibration. In at least some embodiments, the calibration parameters generated at box 1410 can be the user's IPD, and depth plane switching at box 1414 can be based on the user's IPD. If the wearable system is implementing content-based switching in box 1412 (which can occur when there is no calibrated user and / or a content creator or user prefers content-based switching), depth plane switching can be performed based on the depth of the content the user is presumably looking at.
[0259] Continue to refer to FIG. 15It will be appreciated that the display system can be configured to continuously confirm the identity of the user. For example, after block 1414, the display system can be configured to return to block 1406 to identify the user. In some cases, the display system can lose track of a calibrated user that was previously detected in block 1406 as wearing the display device. For example, this can be because the calibrated user has removed the display device, or can be due to a latency issue or a sensor error that falsely indicates that the user is no longer wearing the display device, even if the calibrated user is still wearing the display device. In some embodiments, even if the display device detects that the calibrated user is no longer wearing the display device, the display system can continue to use the calibrated user's calibration profile for a predetermined amount of time or for a predetermined number of frames before switching to the content-based depth plane switching scheme of block 1412 or the dynamic calibration scheme of block 1410, respectively. As discussed herein, the display system can be configured to continuously perform block 1406, and can continue to fail to detect the calibrated user. In response to determining that the calibrated user is not detected for the predetermined amount of time or the predetermined number of frames, the system can switch to the content-based depth plane switching scheme of block 1412 or the dynamic calibration scheme of block 1410, respectively. Advantageously, because the calibrated user is typically the most likely user of the display device, by continuing to use the calibrated user's calibration file in the event that the system falsely fails to detect the calibrated user, the calibrated user can not experience a significant degradation in the user experience.
[0260] FIG. 15 Another example of a method for depth plane selection is shown in FIG. 15. FIG. 14 is a process flow diagram of an example method 1500 for depth plane selection based on a user's interpupillary distance. The method 1500 can be performed by the wearable system described herein. Examples of the method 1500 can be used by the wearable system to render content on a depth plane that generally reduces or minimizes any vergence-accommodation mismatch that can otherwise cause user discomfort and fatigue.
[0261] At block 1502, the wearable system can determine that it is not being worn by a user. As an example, the wearable system determines that it is not being worn after failing to detect the user's eyes by the eye tracking system for more than a threshold period of time (e.g., 5 seconds, 10 seconds, 20 seconds, etc.).
[0262] At block 1504, the wearable system can determine that it is being worn by a user. The wearable system can use one or more sensors to determine that it is being worn. As an example, the wearable system can include a proximity sensor or a touch sensor that is triggered when the wearable system is placed on the head of a user, and can determine that the system is being worn based on signals from such sensors. As another example, the wearable system can determine that it is being worn after recognizing the presence of a user's eyes in eye tracking images. In some cases, the wearable system can distinguish between 1) a failure to detect a user's eyes due to the user having closed their eyes, and 2) a failure to detect a user's eyes due to the user having removed the wearable system. Thus, when the wearable system detects a user's closed eyes, the wearable system can determine that the device is being worn. In some embodiments, some or all of the determinations associated with block 1502 and / or block 1504 can be made in connection with the above-referenced blocks 1402 and / or 1404, respectively. FIG. 17B The determinations associated with block 1402 and / or block 1404 are described above.
[0263] At block 1506, the wearable system can check to see if it has previously been calibrated for any user. In various embodiments, block 1506 can involve obtaining an IPD of a calibrated user as part of identifying whether the current user is a calibrated user (or, if there are multiple calibrated users, one of the calibrated users). In some embodiments, if the wearable system has not previously been calibrated for any user (or if any such calibration has been deleted or removed from the wearable system), the wearable system can implement content-based switching as discussed herein, and can activate content-based switching at block 1516. The wearable system can also implement content-based switching for particular content and / or at the request of a user or content creator. In other words, a content creator or user can be able to specify or request the use of content-based switching even if the wearable system has previously been calibrated, and the wearable system has been calibrated for the current user.
[0264] At block 1508, the wearable system can estimate the IPD of the current user, and at block 1510, the wearable system can accumulate eye tracking data regarding the IPD of the current user. The IPD data accumulated at block 1510 can be measurements of the inter-pupillary distance of the user over time, measurements of the rotational time of the user's left and right centers, indications of the maximum inter-pupillary distance measured over time, or any other relevant data. In at least some embodiments, the wearable system can be able to determine the IPD of the current user at block 1508 with relatively high accuracy (e.g., 95%) over a relatively short timeframe (e.g., 5 to 7 seconds of eye tracking data, which can correspond to about 150 to 200 frames of eye tracking images).
[0265] In some embodiments, the wearable system can estimate the user's IPD as a particular percentile (e.g., the 95th percentile) of all of the IPD values collected (e.g., during dynamic calibration at block 1518 and / or during IPD data accumulation at block 1510). For example, for the 95th percentile, the largest 5% of measured IPD values can be excluded, and then the largest remaining measured IPD value can be taken as the user's IPD. An IPD value calculated in this way can be referred to herein as IPD_95. One benefit of calculating the user's IPD value in this way is that outlier values that can be larger than the user's anatomical IPD can be excluded. Without being limited by theory, it is believed that there is a sufficient number of measured IPD values near or at the user's anatomical IPD such that the IPD_95 value accurately reflects the user's anatomical IPD. For example, if approximately 10% of IPD measurements are close to the user's anatomical IPD, then even if the largest 5% of values are excluded, the IPD_95 value should still reflect the user's anatomical IPD value. If desired, other IPD calculations can be used in which different percentages of the largest measured IPD values can be excluded. As an example, an IPD_100 value can be used in which no IPD values are excluded, or an IPD_98 value can be used in which only the largest 2% of values are excluded (which can be preferable in systems with eye tracking systems that produce relatively few outlier IPD measurements). As further examples, an IPD_90 value can be used, an IPD_85 value can be used, or other IPD values with a desired percentage of measured values excluded.
[0266] At block 1512, the wearable system can determine whether the current user's estimated IPD is within a threshold of the IPD of a calibrated user (or, if there are multiple such users, within a threshold of one of the IPD of a calibrated user). The threshold can be, for example, 0.5 mm, 1.0 mm, 1.5 mm, 2.0 mm, or 2.5 mm. As one particular example, the threshold can be 1.0 mm, which can be large enough to quickly and accurately identify a calibrated user as a calibrated user, while uncalibrated users can generally be identified as uncalibrated users. In general, it can be more preferable to falsely identify an uncalibrated user as a calibrated user than to risk not identifying a calibrated user as a calibrated user. In other words, because the expectation of optimal performance for a calibrated user can outweigh a slight sacrifice for an uncalibrated or guest user, and because a calibrated user can be a more likely or frequent user of a given display system, a false negative in user identification for calibration can be more serious than a false positive.
[0267] If the current user's IPD is within a threshold of the IPD of the calibrated user, the wearable system can assign the current user as a calibrated user and can load the associated calibration parameters at block 1514. As an example, the calibration parameters loaded at block 1514 can include the IPD of the calibrated user, the visual optical axis offset parameter, and other available calibration parameters.
[0268] If the current user's IPD is not within a threshold of the IPD of the calibrated user, the wearable system can assume that the current user is a guest user and can perform dynamic calibration at block 1518 (or perform content-based switching at block 1516). At block 1518, the wearable system can initiate dynamic calibration (e.g., dynamically generate calibration parameters), which, as discussed herein, can include on-the-fly calibration based on the estimated IPD of the current user.
[0269] At block 1516, the wearable system can implement content-based switching as described herein, such as when no prior calibration data is available. At this block, depth plane selection can be made according to the depth associated with the displayed virtual content, as described herein, as opposed to tracking the user's vergence depth. With respect to FIGS. 16- 18, the content-based switching can be performed according to the depth of the virtual content being displayed. Content-based switching Content-based depth plane switching is further discussed.
[0270] At block 1520, the wearable system can perform depth plane switching. The depth plane switching of block 1520 can be performed using the configuration parameters loaded or generated at block 1514, 1516, or 1518. In particular, if the current user is a calibrated user, block 1520 can perform depth plane switching according to the configuration parameters generated during calibration of that user. If the current user is not identified as a calibrated user, and dynamic calibration is performed at block 1518, then block 1520 can involve depth plane switching according to the calibration parameters generated at block 1520 as part of the dynamic calibration. In at least some instances, the calibration parameters generated at block 1518 can be the IPD of the user, and the depth plane switching at block 1520 can be based on the IPD of the user. If the wearable system is implementing content-based switching at block 1516 (which can occur when there is no calibrated user and / or the content creator or user elects to use content-based switching), then the depth plane switching of block 1520 can be performed according to the depth of the content determined to be being viewed by the user.
[0271] FIG. 16A through FIG. 17B
[0272] As described herein, a wearable system can present virtual content via a particular depth plane. When the virtual content is updated, e.g., when a virtual object moves and / or is replaced by a different virtual object, the wearable system can select a different depth plane on which to present the virtual object. As additional non-limiting examples, the wearable system can select a different depth plane on which to present the virtual object when the user of the wearable system focuses on a different virtual object or a different location of the user's field of view.
[0273] The following references are incorporated by reference in their entirety: FIG. 16A An example approach to selecting a depth plane is further discussed below, referred to herein as content-based switching. In content-based switching, the wearable system can select a depth plane on which to present virtual content based on a determination of which virtual object the user can be viewing. In some embodiments, this determination can be made according to whether the user is gazing within a particular region. For example, the wearable system can associate a region with each virtual object that is to be presented as virtual content to the user. The region can be, for example, a volume of space that surrounds the virtual object. The volume of space can be a sphere, a cube, a hyper-rectangle, a pyramid, or any arbitrary three-dimensional polyhedron (e.g., polyhedral). As will be described, these regions are preferably non-overlapping, and thus, the space encompassed by each region can be associated with only a single virtual object.
[0274] In some embodiments, the wearable system can monitor the gaze of the user, e.g., to identify a location that the user is viewing (e.g., within the user's field of view). For an example gaze, the wearable system can identify a region that includes the location that the user is viewing. In this example, having identified the region and there being only one virtual object in the region, the wearable system can select a depth plane that corresponds to the virtual object included in the identified region. The wearable system can then present virtual content at the selected depth plane, which is the depth plane associated with the virtual object.
[0275] Advantageously, the wearable system can utilize a higher latency eye tracking, or otherwise less accurate eye tracking scheme, as compared to an eye tracking scheme that can be required if depth plane switching depends on accurately determining the depth of the gaze point. In some embodiments, for a scene with multiple virtual objects displayed at different depths, a statistical probability or spatial correlation can be utilized to distinguish what the user can be observing or attending to. Thus, the wearable system can determine a gaze region or gaze volume (e.g., rather than an exact gaze point), and select a depth plane based on the gaze region or volume. In some embodiments, weighting factors such as the last application used or the degree of change in the user's gaze can adjust the statistical probability, and thus the size / shape of the region. In some embodiments, the size / shape of the region is related to the statistical probability that the user is gazing in a particular volume; for example, the size of the region can increase when there is a large amount of uncertainty (e.g., due to noise, low tolerance of the eye tracking system, tracking details of the last application used, speed of change in the user's gaze, etc. that cause the uncertainty to exceed a threshold), and the size of the region can decrease when there is less uncertainty.
[0276] As will be described in FIG. 14 some embodiments, the regions can contain respective angular distances (e.g., extending from the user's eyes to an infinite distance away from the user). In this example, the wearable system can thus require only sufficient accuracy to place the user's gaze within a particular portion of the X and Y planes.
[0277] In some embodiments, the display system can transition from performing content-based depth plane switching to performing depth plane switching based on dynamic calibration. This transition can occur, for example, when data obtained for performing a content-based depth plane switching scheme is determined to be unreliable, or in cases where content is provided across different depth ranges of multiple depth planes. As described herein, the amount of uncertainty calculated in the region of the user's gaze can vary based on various factors, including noise, tolerance of the eye tracking system, tracking details of the last application used, speed of change in the user's gaze, etc. In some embodiments, the display system can be configured to transition to dynamic calibration-based depth plane switching when the uncertainty (e.g., associated with determining the location of the user's gaze point) exceeds a threshold. In some other embodiments, particular virtual content can span multiple depth planes. For content that spans a threshold number of depth planes (e.g., three depth planes), the display system can be configured to transition to dynamic calibration-based depth plane switching.
[0278] In some embodiments, the display system can be configured to transition from performing content-based depth plane switching to performing dynamic calibration-based depth plane switching in response to determining that the uncertainty associated with the depth of the virtual content exceeds a threshold. For example, the display system can determine that the depth or position information for the virtual content being provided by a particular application running on the display system is relatively unreliable (e.g., the particular application indicates that the depth or position of the mixed reality virtual content relative to the user is static even when the position of the user changes or the user moves more than a predetermined level of position change or a predetermined level of movement, respectively), and in turn can transition to dynamic calibration-based depth plane switching. Such a transition can enable a more comfortable and / or realistic viewing experience.
[0279] Although primarily described in the context of depth plane switching, it should be appreciated that one or more techniques described herein can be utilized in any of a variety of different display systems capable of outputting light to a user’s eye with different amounts of wavefront divergence. For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). FIG. 15 and / or FIG. 14 For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). For example, in some embodiments, one or more techniques described herein can be employed in a head-mounted display system that includes one or more variable focus elements (VFEs). FIG. 15 and / or FIG. 4 For example, one or more of lenses 458, 456, 454, 452( FIG. 16A ) can be VFEs, as described herein. In these embodiments, the head-mounted display system can control operation of its one or more VFEs in real-time based at least in part on whether the wearer of the head-mounted display system is determined to be a calibrated user or a guest user. That is, the head-mounted display system can control or adjust the amount of wavefront divergence (the focal length of the light) that the light output to the wearer has based at least in part on whether the wearer of the head-mounted display system is determined to be a calibrated user or a guest user. Examples of architectures and control schemes for variable focus eyepieces are disclosed in U.S. Publication No. 2016 / 0110920, published April 21, 2016, and U.S. Publication No. 2017 / 0293145, published October 12, 2017, each of which is hereby incorporated by reference in its entirety. Other configurations are also possible.
[0280] FIG. 14A representation of a user's field of view 1602 is shown. The user's field of view 1602 (a three-dimensional frustum in the illustrated embodiment) can include a plurality of depth planes 1604A-D extending from the user's eye 1606 in the z-direction. The field of view 1602 can be divided into different regions 1608A-D. Each region can thus encompass a particular volume of space included in the field of view 1602. In this example, each region encompasses a volume of space that preferably includes all of the different depth planes 1604A-D. For example, region 1608A can extend from the user's eye 1606 in the z-direction to, for example, an infinite distance away from the eye 1606. Along an orthogonal direction such as the x-direction, region 1608A can extend from the end of the field of view 1602 up to the boundary between regions 1608A and 1608B.
[0281] In the illustrated embodiment, region 1608B includes a tree virtual object 1610. The tree virtual object 1610 is shown as being rendered at depth plane C 1604C. The tree virtual object 1610 can be associated with (e.g., stored or otherwise made available to the display system) location information, for example. As described above, this location information can be used to identify a three-dimensional location at which the tree virtual object 1610 is to be rendered. As shown, the display system can also display a book as a virtual object 1612. Region 1608C includes the book virtual object 1612. The wearable system thus renders virtual content that includes the tree virtual object 1610 and the book virtual object 1612.
[0282] The user's eye 1606 can move around the field of view 1602, for example, to gaze or look at different locations in the field of view 1602. If the user's eye 1606 gazes on the tree virtual object 1610 or on depth plane C, the wearable system can determine to select depth plane C 1604C to render virtual content. Similarly, if the user's eye 1606 gazes on the book virtual object 1612, more on depth plane B, the wearable system can determine to select depth plane B 1604B to render virtual content. As described in at least FIG. 16A the wearable system can use different schemes to cause selection of a depth plane.
[0283] With respect to content-based switching, the wearable system can identify a region that includes a location at which the user's eye 1606 is gazing. The wearable system can then select a depth plane that corresponds to a virtual object included in the identified region. For example, FIG. 16AAn example gaze point 1614 is shown. The wearable system can determine that the user is looking at gaze point 1614 based on the user's eyes 1606. The wearable system can then identify that this gaze point 1614 is included in region 1608C. Since this region 1608C includes the book virtual object 1612, the wearable system can cause virtual content to be presented at depth plane C 1604C. Without being limited by theory, this is believed to provide a comfortable viewing experience because the virtual object at which one can gaze in region 1608C is the book virtual object 1612, and thus, the depth plane for that book virtual object 1612 is appropriately switched to.
[0284] Thus, if the user adjusts gaze to gaze point 1616 or 1618, the wearable system can maintain presentation at depth plane C 1604B. However, if the user adjusts gaze to within region 1608B, the wearable system can select depth plane C 1604C on which to present virtual content. Although FIG. 16A The example includes four regions, but it should be understood that fewer regions or a greater number of regions can be utilized. Optionally, the number of regions can be the same as the number of virtual objects, and each virtual object has a unique region, and in some embodiments, the number of regions can dynamically change as the number of virtual objects changes (preferably, so long as the regions can extend from the front to the back of the display frustum without more than one object occupying the same region, e.g., without two or more objects being located generally on the same line of sight relative to the user). For example, in an example in which FIG. 16B Two regions can be used in an example of the example in which there are two virtual objects that do not have overlapping lines of sight for the user. Optionally, the wearable system can dynamically adjust the spatial volume that each region encompasses. For example, as gaze detection accuracy improves or degrades, the spatial volume can increase, decrease, or the number of regions can adjust.
[0285] FIG. 16A A perspective view is shown in FIG. 16B. The user is looking at the book virtual object 1612. The wearable system can determine that the user is looking at the book virtual object 1612 based on the user's eyes 1606. The wearable system can then identify that this gaze point 1614 is included in region 1608C. Since this region 1608C includes the book virtual object 1612, the wearable system can cause virtual content to be presented at depth plane C 1604C. Without being limited by theory, this is believed to provide a comfortable viewing experience because the virtual object at which one can gaze in region 1608C is the book virtual object 1612, and thus, the depth plane for that book virtual object 1612 is appropriately switched to. FIG. 17AAn example of the spatial volume contained by each region of the user field of view 1602. As shown, each region 1608A, 1608B, 1608C, and 1608D extends in the z-direction from the front to the back of the display frustum 1602. As shown, in the x-y axes, the regions extend vertically to partition or divide the field of view along the x-axis. In some other embodiments, the regions can extend horizontally in the x-y axes to partition the field of view along the y-axis. In such embodiments, the illustrated frustum can be effectively rotated 90° in the x-y plane. In yet other embodiments, in the x-y axes, the regions can partition or divide the field of view in both the x and y axes, forming a grid in the x-y axes. In all of these embodiments, each region still preferably extends from the front to the back of the display frustum 1602 (e.g., from the nearest plane to the farthest plane at which the display system can display content).
[0286] FIG. 17A Another representation of a user field of view 1602 is shown for content-based switching. In this example, as shown in FIG. 16A FIG. 16A through FIG. 16B The user field of view 1602 includes depth plane A 1604A to depth plane D 1604D. However, the spatial volume associated with each virtual object is different from the spatial volume of FIG. 17A For example, the region 1702 in which the book virtual object 1612 is shown to enclose a shape, such as a spherical spatial volume, a cylindrical spatial volume, a cube, a polyhedron, etc. Preferably, the region 1702 extends less than the entire depth of the display frustum 1602 in the z-axis. Advantageously, this arrangement of regions allows for differentiation between objects that can be in similar lines of sight for the user.
[0287] Optionally, the wearable system can adjust the regions of FIG. 17A based on eye tracking accuracy exceeding one or more thresholds. For example, the size of the regions can decrease as the tracking accuracy confidence increases. As another example, it will be understood that each virtual object has an associated region. The wearable system can adjust the regions of FIG. 17A to avoid overlapping regions. For example, if a new virtual object is to be displayed within a region corresponding to a different virtual object, the size of the region for one or both virtual objects can be decreased to avoid overlap. For example, if the book virtual object 1612 in FIG. 17A moves in front of the tree virtual object 1610 (e.g., in region 1608A), the wearable system can adjust the regions in FIG. 17A to be closer to the virtual object (e.g., as shown in FIG. 17A to prevent overlap.
[0288] Continuing with the example of FIG. 17B The tree virtual object 1610 is shown as being included in a first region 1706 and a second region 1708. Optionally, the first region 1706 can represent a relatively smaller volume of space surrounding the tree virtual object 1610. The second region 1708 can represent a current volume of space surrounding the tree virtual object 1610. For example, as virtual objects move around within the field of view 1602, the volume of space included by the regions can be adjusted (e.g., in real-time). In this example, the adjustments can ensure that each region does not overlap with any other region. Thus, if the book virtual object 1612 moves close to the tree virtual object 1610, the wearable system can reduce the volume of space of the region surrounding the tree virtual object 1610, the book virtual object 1612, or both. For example, the second region 1708 can be reduced to be more volumetrically close to the first region 1706.
[0289] As noted herein, in content-based depth plane switching, the depth plane associated with a virtual object dictates the depth plane rather than the gaze point. An example of this is shown with respect to the gaze point 1704. As shown, an example region 1702 of the book virtual object 1612 can include a volume of space that includes portions defined by the depth plane B 1604B and the depth plane 1604C. In some cases, the wearable system determines that the user is gazing at the point 1704 in the depth plane C, and also identifies that gaze point 1704 as being included in the region 1702. Because the depth plane associated with the book virtual object 1612 dictates the depth plane switching, the wearable system switches to the depth plane B of the book virtual object 1612 rather than the depth plane C of the gaze point 1704. FIG. 17A An example of a perspective view of a representation of a tree is shown. FIG. 18 An example of a perspective view of a representation of a tree is shown.
[0290] FIG. 16A through FIG. 17B A flowchart of an example process of selecting a depth plane based on content-based switching is shown. For convenience, the process 1800 will be described as being performed by a wearable system of one or more processors (e.g., a wearable device such as the wearable system 200 described above).
[0291] At block 1802, the wearable system presents virtual content on a particular depth plane. As described above, the depth plane can be used to provide accommodation cues to a user of the wearable system. For example, each depth plane can be associated with an amount of wavefront divergence of light presented by the wearable system to the user. The virtual content can include one or more virtual objects.
[0292] At block 1804, the wearable system determines a point of gaze. The wearable system can use sensors such as cameras to estimate a three-dimensional location at which the user is gazing. These cameras can update at a particular rate such as 30 Hz, 60 Hz, etc. The wearable system can determine a vector extending from the user's eye (e.g., from the center of the eye or the pupil) and estimate a three-dimensional location at which the vector intersects. In some embodiments, the point of gaze can be estimated based on IPD and assuming that IPD from a predefined variation of the maximum IPD is related to a gaze on a particular depth plane. Additionally, optionally, in addition to IPD, an approximation of the location at which the user's gaze is further located at the point of gaze can be determined. This estimated location can have a particular error associated with it, so the wearable system can determine a volume of space in which the point of gaze can be located relative to the virtual content.
[0293] At block 1806, the wearable system identifies a region that includes the point of gaze. As discussed with respect to FIG. 19 The wearable system can associate regions with each virtual object. For example, a region can include a particular (e.g., single) virtual object. Preferably, regions for different virtual objects do not overlap.
[0294] At block 1808, the wearable system selects a depth plane associated with virtual content included in the identified region. The wearable system can identify virtual content associated with the region identified in block 1806. The wearable system can then obtain information indicating a depth plane on which the virtual content is to be presented. For example, the virtual content can be associated with location information indicating a three-dimensional location. The wearable system can then identify a depth plane associated with the three-dimensional location. This depth plane can be selected to present the virtual content.
[0295] FIG. 16A through FIG. 17B A flow diagram illustrating an example process of adjusting regions based on content-based switching is shown. For convenience, the process 1800 will be described as being performed by a wearable system of one or more processors (e.g., a wearable system such as the wearable system 200 described above).
[0296] At block 1902, the wearable system presents a virtual object. As described herein, the wearable system can have regions associated with the presented virtual object. For example, a region can encompass a volume of space and include a particular virtual object (e.g., a single virtual object). These regions can be any shape or polyhedron and in some embodiments can extend infinitely in one or more directions. FIG. 16A Examples of regions are shown in FIGS. 16A-16C.
[0297] At block 1904, the wearable system adjusts the number of virtual objects or adjusts the position of the virtual objects. The wearable system can present additional virtual objects, e.g., 5 virtual objects can be presented at block 1902. At block 1904, 8 virtual objects can be presented. Alternatively, the wearable system can update the position of the virtual objects. For example, the virtual objects presented in block 1902 can be bees. The virtual objects can thus travel around the field of view of the user.
[0298] At block 1906, the wearable system updates the regions associated with the one or more virtual objects. For the example of presenting additional virtual objects, the virtual objects can become closer together. In some embodiments, because each region can only be allowed to include a single virtual object, the increase in the number of virtual objects can require adjustment of the regions. For example, FIG. 16A Two virtual objects are shown. If additional virtual objects are included, they can be included in a region that either of the two virtual objects is also included in. Thus, the wearable system can adjust the regions— e.g., adjust the spatial volume assigned to each region (e.g., decrease the spatial volume). In this way, each region can include a single virtual object. With respect to the example of moving virtual objects, a moving virtual object can come close to another virtual object. Thus, the moving virtual object can extend into a region associated with the other virtual object. Similar to the above, the wearable system can adjust the regions to ensure that each virtual object is included in its own region.
[0299] As another example of updating, the wearable system can associate regions with the virtual objects identified in block 1902. For example, the regions can be similar to regions 1608A-D shown in FIG. 16B. The wearable system can adjust the regions to contain a smaller spatial volume. For example, the wearable system can update the regions to be similar to regions 1702, 1706 shown in FIG. 17B. FIG. 17A FIG. 17A As another example of updating, the regions can be similar to the regions of FIG. 17A. With respect to the example of moving virtual objects, a moving virtual object can come close to another virtual object. Thus, the moving virtual object can extend into a region associated with the other virtual object. Similar to the above, the wearable system can adjust the regions to ensure that each virtual object is included in its own region.
[0300] As another example of updating, the regions can be similar to the regions of FIG. 17A. With respect to the example of moving virtual objects, a moving virtual object can come close to another virtual object. Thus, the moving virtual object can extend into a region associated with the other virtual object. Similar to the above, the wearable system can adjust the regions to ensure that each virtual object is included in its own region. FIG. 17A FIG. 17A If the book virtual object 1612 moves close to the tree virtual object 1610, the region 1708 around the tree virtual object 1610 can include the book virtual object 1612. Thus, the wearable system can update the region 1708 to decrease the spatial volume it contains. For example, the region 1708 can be adjusted to be the region 1706. As another example, the region 1708 can be adjusted to be closer to the region 1706. As shown in FIG. 17B, the region 1706 can optionally reflect the smallest region around the tree virtual object 1610. FIG. 20
[0301] FIG. 14 A flow diagram illustrating an example process for operating a head-mounted display system based on a user identity is shown. For convenience, the method 2000 will be described as performed by a wearable system of one or more processors (e.g., a wearable system described above, such as wearable system 200). In various implementations of method 2000, the below-described blocks can be performed in any suitable order or sequence, and individual blocks can be combined or rearranged, or other blocks added, as desired. In some implementations, method 2000 can be performed by a head-mounted system that includes one or more cameras configured to capture images of one or both eyes of a user and one or more processors coupled to the one or more cameras. In at least some such implementations, some or all of the operations of method 2000 can be performed, at least in part, by the one or more processors of the system.
[0302] At blocks 2002-2006 (i.e., blocks 2002, 2004, and 2006), method 2000 can begin with the wearable system determining whether it is not being worn by a user. The wearable system can use one or more sensors (e.g., an eye tracking system, such as eye cameras 324) to determine that it is not being worn. In some embodiments, some or all of the operations associated with block 2004 can be similar or substantially the same as those described above with reference to block 1402. In some embodiments, some or all of the operations associated with block 2006 can be similar or substantially the same as those described above with reference to block 1502. FIG. 15 and FIG. 14 In some examples, method 2000 can remain in the hold mode until the wearable system determines at block 2004 that it is being worn by a user. In response to determining at block 2004 that the user is wearing the wearable system, method 2000 can proceed to block 2008.
[0303] At block 2008, method 2000 can include the wearable system performing one or more iris authentication operations. In some embodiments, some or all of the operations associated with block 2008 can be similar or substantially the same as those described above with reference to block 1404, respectively. FIG. 15 and FIG. 14The operations associated with one or more of block 1406 and / or block 1506 described are similar or substantially the same for implementations using iris scanning or authentication. In some examples, the one or more operations associated with block 2008 can include the wearable system presenting a virtual target to the user (e.g., through one or more displays), capturing one or more images of one or both eyes of the user while the virtual target is presented to the user (e.g., through one or more eye tracking cameras), generating one or more iris codes or tokens based on the one or more images, evaluating the one or more generated iris codes or tokens relative to a set of one or more predetermined iris codes or tokens associated with the registered user, and generating an authentication result based on the evaluation. In at least some of these examples, the iris of each of one or both eyes of the user can be displayed in some or all of the one or more images. In some implementations, the virtual target can be presented to attract the user’s attention or to otherwise direct the user’s gaze in a certain direction and / or toward a certain location in three-dimensional space. Thus, the virtual target may, for example, include multi-colored and / or animated virtual content, can be presented to the user along with accompanying audio, haptic feedback, and / or other stimuli or combinations thereof. In at least some of these implementations, the virtual target can be presented to the user along with one or more prompts instructing or suggesting that the user look at or otherwise gaze at the virtual target. In some embodiments, instead of or in addition to generating and evaluating one or more iris codes or tokens, the one or more operations associated with block 2008 can include the wearable system can evaluate at least one of the one or more images with one or more reference images to determine whether a match exists, and generate an authentication result based on the evaluation. Examples of iris imaging, analysis, identification, and authentication operations are disclosed in U.S. Patent Application No. 15 / 934,941, filed March 23, 2018, published as U.S. Publication No. 2018 / 0276467 on September 27, 2018; U.S. Patent Application No. 15 / 291,929, filed October 12, 2016, published as U.S. Publication No. 2017 / 0109580 on April 20, 2017; U.S. Patent Application No. 15 / 408,197, filed January 17, 2017, published as U.S. Publication No. 2017 / 0206412 on July 20, 2017; and U.S. Patent Application No. 15 / 155,013, filed May 14, 2016, published as U.S. Publication No. 2016 / 0358181 on December 8, 2016, each of which is incorporated by reference herein in its entirety. In some embodiments, at block 2008, the method 2000 can include the wearable system performing one or more of the iris imaging, analysis, identification, and authentication operations described in the above patent applications.Furthermore, in some embodiments, one or more of the iris imaging, analysis and identification, and authentication operations described in the aforementioned patent application may be performed in association with one or more other logic blocks of method 2000, such as one or more of blocks 2010, 2012, and 2018 described below. Other configurations are possible.
[0304] At box 2010, method 2000 may include the wearable system determining whether an authentication result (e.g., obtained by performing one or more iris authentication operations at box 2008) indicates whether the user currently wearing the wearable system is a registered user. In some examples, a registered user may be similar to or substantially the same as a calibration user (e.g., a user associated with a pre-existing calibration profile). In some implementations, at box 2010, method 2000 may further include the wearable system determining whether a calibration profile exists for the current user of the device. In some embodiments, some or all of the operations associated with box 2004 may be related to those referenced above respectively. FIG. 15 and FIG. 14 The operations of one or more of the boxes 1406 and / or 1506 described are similar or substantially the same.
[0305] In response to the wearable system determining in block 2010 that the user currently wearing the wearable system is indeed a registered user, method 2000 may proceed to block 2020. In block 2020, method 2000 may include the wearable system selecting settings associated with the registered user. In some embodiments, some or all of the operations associated with block 2020 may be referenced as described above. FIG. 15 and FIG. 14 One or more of the associated operations in boxes 1408 and / or 1514 described are similar or substantially the same.
[0306] On the other hand, method 2000 may proceed from box 2010 to box 2012 in response to the wearable system (i) determining at box 2010 that the user currently wearing the wearable system is not a registered user or (ii) failing to determine at box 2010 that the user currently wearing the wearable system is a registered user. At box 2012, method 2000 may include the wearable system determining whether a confidence value associated with the authentication result exceeds a predetermined threshold. In some embodiments, at box 2012, method 2000 may include the wearable system generating or otherwise obtaining a confidence value associated with the authentication result and comparing the confidence value to a predetermined threshold. In some embodiments, the confidence value obtained at box 2012 may indicate the level of confidence that the user currently wearing the wearable system is not a registered user. In some examples, the confidence value may be determined at least in part based on whether any iris code or token was successfully generated or otherwise extracted at box 2008. Various different factors and environmental conditions may affect the wearable system's ability to successfully generate or otherwise extract one or more iris codes or tokens at box 2008. Examples of such factors and conditions may include the wearable system's adaptation or registration relative to the current user, the position and / or orientation of the iris of one or both eyes of the user relative to the position and / or orientation of one or more cameras when capturing one or more images, lighting conditions, or combinations thereof. Therefore, such factors and conditions may also affect the confidence that the wearable system can determine that the current user is not a registered user.
[0307] In response to the wearable system determining at box 2012 that the confidence value associated with the authentication result does indeed exceed a predetermined threshold, method 2000 may proceed to box 2014. In some examples, this determination at box 2012 may indicate that the wearable system has determined with a high confidence that the user currently wearing the wearable system is not a registered user. At box 2014, method 2000 may include the wearable system presenting a prompt to the user (e.g., via one or more displays) suggesting that the user register for iris authentication and / or participate in one or more eye calibration procedures to generate a calibration profile for the user. In some embodiments, such a prompt may request the user to indicate whether they agree to continue the suggested registration procedure. At box 2016, method 2000 may include the wearable system receiving user input in response to the prompt and determining whether the received user input indicates that the user has agreed to continue the suggested registration procedure. In some examples, at box 2016, method 2000 may monitor the received user input and / or pause until the wearable system receives user input in response to the prompt indicating whether the user agrees to continue the suggested registration procedure.
[0308] In response to the wearable system determining in box 2016 that the received user input in response to the prompt indicates that the user has indeed agreed to continue some or all of the suggested registration procedures, method 2000 may proceed to box 2018. In box 2018, method 2000 may include the wearable system performing some or all of the suggested registration procedures based on the received user input in response to the prompt. In some embodiments, performing such a registration procedure may include the wearable system performing one or more operations associated with iris authentication registration and / or eye calibration. As described further in detail below, in some embodiments, such a registration procedure associated with box 2018 may include one or more procedures for generating an eye-tracking calibration profile for the current user, one or more operations for registering the current user to iris authentication, or a combination of the above. Upon successful completion of the registration procedure associated with box 2018, method 2000 may proceed to box 2020.
[0309] On the other hand, method 2000 may proceed from box 2016 to box 2022 in response to the wearable system (i) determining in box 2016 that the received user input in response to the prompt indicates that the user has not agreed to continue some or all of the suggested registration procedures, or (ii) failing in box 2016 to determine that the received user input in response to the prompt indicates that the user has agreed to continue some or all of the suggested registration procedures. In box 2022, method 2000 may include the wearable system selecting default settings. In some embodiments, some or all of the operations associated with box 2022 may be related to those associated with the above references. FIG. 15 and FIG. 11 The operations of one or more of the boxes described in 1410, 1412, 1508-1512, 1516 and / or 1518 are similar or substantially the same.
[0310] In response to the wearable system determining in box 2012 that the confidence value associated with the authentication result does not exceed a predetermined threshold, method 2000 may proceed from box 2012 to box 2028. In some examples, such as the determination in box 2012, the wearable system may indicate that its confidence in determining that the user currently wearing the wearable system is not a registered user is relatively low. In box 2028, method 2000 may include the wearable system determining whether it has attempted to authenticate the current user more than a predetermined threshold number of times (e.g., by performing one or more iris authentication operations in box 2008). In response to the wearable system determining in box 2028 that it has indeed attempted to authenticate the current user more than the predetermined threshold number of times, method 2000 may proceed to box 2022.
[0311] In response to the wearable system determining in box 2028 that it has not attempted to authenticate the current user more than a predetermined threshold number of times, method 2000 may proceed to boxes 2030 to 2032. In box 2032, method 2000 may include the wearable system performing one or more operations to attempt to improve the accuracy and / or reliability of one or more iris authentication operations performed in conjunction with box 2008, so as to be able to make a determination about the user's identity with increased confidence. As mentioned above, various factors and environmental conditions may affect the wearable system's ability to successfully generate or otherwise extract one or more iris codes or tokens. By performing one or more operations associated with box 2032, the wearable system may attempt to counteract or otherwise compensate for one or more such factors and conditions. Therefore, in box 2032, method 2000 may include the wearable system performing one or more operations to increase the likelihood of successfully generating or otherwise extracting iris codes or tokens from images of one or both of the user's eyes. In some implementations, at block 2032, method 2000 may include the wearable system performing one or more adaptation or registration-related operations, followed by one or more operations presented at block 2008 for changing the position of the virtual target, one or more operations for determining whether the user's eyes are visible, or a combination thereof.
[0312] In some embodiments, one or more such adaptation or registration-related operations may correspond to the above references. FIG. 14 This refers to one or more of the operations described in Registration Observer 620 and / or box 1170. That is, in these embodiments, the wearable system may perform one or more operations to evaluate and / or attempt to improve the fit or registration of the wearable system relative to a user. To this end, the wearable system may provide feedback and / or suggestions to the user to adjust the fit, positioning, and / or configuration of the wearable system (e.g., nose pad, headrest, etc.). Other examples of operations related to fit and registration are disclosed in U.S. Provisional Application No. 62 / 644,321, filed March 16, 2018; U.S. Application No. 16 / 251,017, filed January 17, 2019; U.S. Provisional Application No. 62 / 702,866, filed July 24, 2018; and International Patent Application PCT / US2019 / 043096, filed July 23, 2019, each of which is incorporated herein by reference in its entirety. In some embodiments, at block 2032, method 2000 may include the wearable system performing one or more of the iris adaptation and registration-related operations described in the aforementioned patent application. Other configurations are possible.
[0313] As described above, in some embodiments, method 2000 may include a wearable system performing one or more operations to change the position of a virtual target subsequently presented in block 2008. In these examples, method 2000 may continue from block 2032 to block 2008, and, while performing one or more iris authentication operations in conjunction with block 2008, present the virtual target to the user at one or more locations that differ from the locations where the virtual target was presented to the user in previous iterations of method 2000. In some examples, the virtual target may be presented to the user in a manner that increases the chance of capturing a frontal image of the iris of one or both of the user's eyes. That is, in such examples, the virtual target may be presented to the user at one or more locations in three-dimensional space, which may require the user to redirect one or both of their eyes toward one or more cameras in order for the user to view and / or continue to gaze at it. In some embodiments, one or more other characteristics of the virtual target (e.g., size, shape, color, etc.) may be changed. Also as described above, in some embodiments, at block 2032, method 2000 may include a wearable system performing one or more operations to determine whether the user's eyes are visible. In some such implementations, method 2000 may proceed to block 2022 in response to the wearable system (i) determining in block 2032 that the user’s eyes are not visible or (ii) failing to determine in block 2032 that the user’s eyes are visible.
[0314] Upon reaching and performing the operation associated with box 2020 or box 2022, method 2000 may proceed to box 2024. In box 2024, method 2000 may include the wearable system operating according to settings selected in box 2020 or box 2022. In some embodiments, some or all of the operations associated with box 2024 may be related to those associated with the operations described above. FIG. 15 and FIG. 14 The operations of one or more of boxes 1414 and / or 1520 described are similar or substantially the same. In some embodiments, when transitioning from box 2020 or box 2022 to box 2024, method 2000 may include the wearable system entering an "unlocked" state or otherwise granting the current user access to certain features. In box 2026, method 2000 may include the wearable system determining whether it is not being worn by a user. In some embodiments, some or all of the operations associated with box 2004 may be similar to those associated with the aforementioned references. FIG. 15 , FIG. 20 and Computer vision to detect objects in ambient environment Those operations described in one or more of boxes 1402, 1404, 1502, 1504 and / or 2004 are similar or substantially the same.
[0315] In response to determining in box 2026 that a user is wearing the wearable system, method 2000 may proceed to box 2024. That is, while reaching boxes 2024-2026, the wearable system may continue to operate according to the settings selected in box 2020 or box 2022 until the wearable system determines in box 2026 that it is no longer being worn by the user. In response to determining in box 2026 that the user is no longer wearing the wearable system, method 2000 may transition back to box 2004. In some embodiments, during the transition from box 2026 to box 2004, method 2000 may include the wearable system entering a “locked” state or otherwise restricting access to certain features. In some embodiments, during the transition from box 2026 to box 2004, method 2000 may include the wearable system pausing or interrupting the performance of one or more operations described above with reference to 2024.
[0316] As described above, in some embodiments, at block 2018, method 2000 may include a wearable system performing a registration procedure, which may include performing one or more operations to generate an eye-tracking calibration profile for the current user, registering the current user to iris authentication, or a combination of the above operations. In some embodiments, at block 2018, method 2000 may include the wearable system performing calibration for the current user or otherwise performing one or more operations to generate an eye-tracking calibration profile for the current user. Calibration may include various processes and may include determining how the user's eyes move as the user focuses on objects at different locations and depths. In some embodiments, the wearable system may be configured to monitor the user's eye gaze and determine the three-dimensional gaze point the user is looking at. The gaze point may be determined based on the distance between the user's eyes and the gaze direction of each eye, etc. It is understood that these variables may be understood as forming a triangle, where the fixed point is at one corner of the triangle and the eyes are at the other corners. It will also be understood that calibration may be performed to accurately track the direction of the user's eyes and determine or estimate the gaze of those eyes to determine the gaze point. Calibration may also include identifying the user's IPD (e.g., the distance between the user's pupils when the user is focused on optical infinity). Calibration can also include determining the user's interpupillary distance (IPD) when the user is focusing on an object closer than optical infinity, such as an object in the near field (e.g., less than 2.0 meters) and an object in the intermediate field (e.g., between approximately 2.0 and 3.0 meters). Based on such calibration data, the wearable system may be able to determine the depth the user is viewing by monitoring the user's IPD. In other words, when the user's IPD is at its maximum (e.g., equal to or close to the user's IPD), the wearable system may be able to infer that the user's vergence distance is at or near optical infinity. Conversely, when the user's IPD is close to its minimum, the wearable system may be able to infer that the user's vergence distance is close to the user, as determined by calibration. Therefore, after a complete calibration, the wearable system can store calibration files and / or calibration information for the current user. Further details regarding calibration and eye tracking can be found in: U.S. Patent Application No. 15 / 993,371, filed May 30, 2018, published December 6, 2018, with U.S. Publication No. 2018 / 0348861; U.S. Provisional Application No. 62 / 714,649, filed August 3, 2018, and U.S. Provisional Application No. 62 / 875474, filed July 17, 2019; and U.S. Application No. 16 / 530,904, filed August 2, 2019, each of which is incorporated herein by reference in its entirety. In some embodiments, at box 2018, method 2000 may include a wearable system performing one or more of the calibration and eye-tracking operations described in the aforementioned patent applications. Other configurations are possible.
[0317] In some implementations, at block 2018, method 2000 may include the wearable system performing one or more operations for registering the current user to iris authentication. In some examples, such operations may include the wearable system presenting a virtual target to the user (e.g., via one or more displays), capturing one or more images of one or both eyes of the user while presenting the virtual target (e.g., via one or more eye-tracking cameras), generating a representation of the iris of each of the user's one or both eyes based on the one or more images (e.g., one or more iris codes or tokens), and storing the generated representation. In some implementations, instead of or in addition to generating and storing representations of the iris of each of the user's one or both eyes (e.g., one or more iris codes or tokens), such operations may include storing at least one of the one or more images as a reference or template image. In some implementations, the wearable system may perform one or more operations associated with block 2032, such as one or more of those associated with block 2018, while performing one or more operations for registering the current user to iris authentication, to increase the likelihood of successfully generating or otherwise extracting an iris code or token from images of one or both of the user's eyes. In some embodiments, the wearable system may perform one or more operations for registering the current user in iris authentication while simultaneously performing the calibration process described above for the current user. In other embodiments, the wearable system may perform one or more such operations, registering the current user in iris authentication, and the calibration process described above for the current user, at different times.
[0318] In some examples, at box 2018, method 2000 may include the wearable system storing calibration files or calibration information and the generated representation (e.g., one or more iris codes or tokens) in association. In some embodiments, instead of or in addition to storing calibration files or calibration information and the generated representation in association, at box 2018, method 2000 may include the wearable system storing calibration files or calibration information in association with at least one of one or more images. In some embodiments, some or all of the operations associated with box 2018 may be performed independently of method 2000. For example, the wearable system may perform one or more operations for registering the user in iris authentication and / or one or more associated operations for performing calibration for the user, based on a user's request or when the wearable system is initially turned on or set up.
[0319] In some embodiments, in response to the wearable system determining in block 2012 that the confidence value associated with the authentication result does not exceed a predetermined threshold, method 2000 may proceed from block 2012 to one or more logical blocks other than block 2028. For example, in these embodiments, in response to the wearable system determining in block 2012 that the confidence value associated with the authentication result does not exceed a predetermined threshold, the wearable system may present one or more alternative authentication options to the current user. For example, the wearable system may offer the current user the option to verify their identity by providing a personal identification number (PIN), password, and / or other credentials as input to the wearable system. In some embodiments of these embodiments, method 2000 may proceed to block 2020 if the current user is successfully identified as a registered user using one or more of the aforementioned alternative authentication options. Similarly, in some embodiments of these embodiments, method 2000 may proceed to block 2016 if the current user is not successfully identified as a registered user using one or more of the aforementioned alternative authentication options. At such a moment, method 2000 can proceed to box 2022 or to one or more logic boxes similar to box 2018, where the wearable system performs one or more operations for registering the current user.
[0320] In some examples, such operations for registering the current user may include one or more processes for generating an eye-tracking calibration profile for the current user, one or more operations for registering the current user in iris authentication, one or more operations for registering the current user to one or more of the aforementioned alternative authentication options, or a combination of the above operations. In some embodiments, at block 2018, method 2000 may further include the wearable system performing one or more operations to register the current user to one or more of the aforementioned alternative authentication options. In these embodiments, the information generated or otherwise obtained by the wearable system in connection with performing one or more operations for registering the current user to one or more of the aforementioned alternative authentication options may be
[0321] In these embodiments, at block 2018, method 2000 may include the wearable system storing in association information generated or otherwise obtained by the wearable system in connection with performing operations for registering the current user to one or more of the aforementioned alternative authentication options, calibration files or information of the current user, one or more iris representations (e.g., iris codes or tokens) generated for the current user, one or more images of one or both of the user's eyes, or a combination thereof. Additional information, such as information manually provided by the user at block 2018 (e.g., name, email address, interests, user preferences, etc.), may also be stored in association with one or more of the aforementioned information. Other configurations are possible.
[0322] Machine learning
[0323] As discussed above, a display system can be configured to detect objects or features in the environment surrounding the user. As discussed herein, various technologies, including a variety of environmental sensors (e.g., cameras, audio sensors, temperature sensors, etc.), can be used to accomplish this detection.
[0324] In some embodiments, computer vision techniques can be used to detect objects present in the environment. For example, as disclosed herein, a forward-facing camera of the display system can be configured to image the surrounding environment, and the display system can be configured to perform image analysis on the image to determine the presence of objects in the surrounding environment. The display system can analyze images acquired by an outward-facing imaging system to perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, or image restoration, etc. As another example, the display system can be configured to perform face and / or eye recognition to determine the presence and location of faces and / or eyes in the user's field of view. One or more computer vision algorithms can be used to perform these tasks. Non-limiting examples of computer vision algorithms include: Scale Invariant Feature Transform (SIFT), Speed-Up Robust Features (SURF), Oriented Fast and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypad (BRISK), Fast Retina Keypad (FREAK), Viola-Jones algorithm, Eigenfaces method, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, Visual Simultaneous Localization and Mapping (vSLAM) technique, Sequential Bayesian estimators (e.g., Kalman filter, Extended Kalman filter, etc.), Bundlelead adjustment, Adaptive thresholding (and other thresholding techniques), Iterative Nearest Point (ICP), Semi-Global Matching (SGM), Semi-Global Block Matching (SGBM), Feature Point Histogram, various machine learning algorithms (e.g., Support Vector Machine, k-Nearest Neighbors, Naive Bayes, Neural Networks (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.), and so on.
[0325] One or more of these computer vision techniques can also be used in conjunction with data acquired from other environmental sensors (e.g., microphones) to detect and determine various characteristics of objects detected by the sensors.
[0326] As discussed in this paper, objects in the surrounding environment can be detected based on one or more criteria. When the display system detects the presence or absence of criteria in the surrounding environment using computer vision algorithms or data received from one or more sensor components (which may or may not be part of the display system), the display system can subsequently signal the presence of the object.
[0327] Other considerations
[0328] Various machine learning algorithms can be used to learn to identify the presence of objects in the surrounding environment. Once trained, these algorithms can be stored by the display system. Some examples of machine learning algorithms may include: supervised or unsupervised machine learning algorithms, including regression algorithms (e.g., ordinary least squares regression), instance-based algorithms (e.g., learning vector quantization), decision tree algorithms (e.g., classification and regression trees), Bayesian algorithms (e.g., Naive Bayes), clustering algorithms (e.g., k-means clustering), association rule learning algorithms (e.g., prior algorithms), artificial neural network algorithms (e.g., perceptrons), deep learning algorithms (e.g., deep Boltzmann machines, or deep neural networks), dimensionality reduction algorithms (e.g., principal component analysis), holistic algorithms (e.g., stacked generalization), and / or other machine learning algorithms. In some embodiments, individual models can be customized for different datasets. For example, wearable devices can generate or store a base model. The base model can be used as a starting point to generate additional models specific to data types (e.g., specific users), datasets (e.g., a collection of additional images), conditional conditions, or other variations. In some embodiments, the display system can be configured to utilize a variety of techniques to generate models for analyzing aggregated data. Other techniques may include using predefined thresholds or data values.
[0329] The criteria used to detect objects may include one or more threshold conditions. If analysis of data acquired by environmental sensors indicates that a threshold condition has been met, the display system can provide a signal indicating the presence of an object in the surrounding environment. Threshold conditions may involve quantitative and / or qualitative measurements. For example, threshold conditions may include a score or percentage associated with the probability of a reflection and / or the presence of an object in the environment. The display system can compare a score calculated based on data from the environmental sensors with a threshold score. If the score is above the threshold level, the display system can detect the presence of a reflection and / or an object. In some other embodiments, if the score is below the threshold, the display system can signal the presence of an object in the environment. In some embodiments, threshold conditions may be determined based on the user's emotional state and / or the user's interaction with the surrounding environment.
[0330] In some embodiments, threshold conditions, machine learning algorithms, or computer vision algorithms may be tailored to a specific environment. For example, in a diagnostic environment, a computer vision algorithm may be dedicated to detecting certain responses to stimuli. As another example, as discussed herein, a display system may execute facial recognition algorithms and / or event tracking algorithms to sense a user's response to stimuli.
[0331] It will be appreciated that each process, method, and algorithm depicted in the description and / or figures herein can be embodied in and wholly or partially automated by one or more physical computing systems, hardware computer processors, special-purpose circuits, and / or electronic hardware configured to execute specialized and specific computer instructions. For example, a computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer, special-purpose circuitry, etc., programmed with specific computer instructions. Code modules may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language. In some embodiments, specific operations and methods may be performed by circuitry specific to a given function.
[0332] Furthermore, some embodiments of the functionality disclosed herein are mathematically, computationally, or technically complex enough that dedicated hardware or one or more physical computing devices (using appropriate dedicated executable instructions) may be required to perform the functionality, for example, due to the amount or complexity of the calculations involved, or to provide results substantially in real time. For example, video may comprise many frames, each with millions of pixels, and requires specially programmed computer hardware to process the video data to provide the required image processing tasks or applications within a commercially reasonable timeframe.
[0333] Code modules or any type of data can be stored on any type of non-transitory computer-readable medium, such as physical computer memory, including hard disk drives, solid-state memory, random access memory (RAM), read-only memory (ROM), optical disks, volatile or non-volatile storage devices, combinations thereof, and / or similar memories. In some embodiments, the non-transitory computer-readable medium can be part of one or more of a local processing and data module (140), a remote processing module (150), and a remote data repository (160). The methods and modules (or data) can also be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) over various computer-readable transmission media, including wireless and wired / cable-based media, and can take many forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed processes or process steps can be permanently or otherwise stored in any type of non-transitory tangible computer memory, or can be transmitted via a computer-readable transmission medium.
[0334] Any process, block, state, step, or function described herein and / or depicted in the accompanying drawings should be understood as potentially representing a code module, code segment, or code portion, including one or more executable instructions for implementing a specific function (e.g., logic or arithmetic) or step in the process. Various processes, blocks, states, steps, or functions may be combined, rearranged, added to, removed from, modified, or otherwise altered in the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functions described herein. The methods and processes described herein are not limited to any particular order, and the associated blocks, steps, or states may be performed in other suitable orders (e.g., serial, parallel, or some other manner). Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the embodiments described herein is for illustrative purposes and should not be construed as requiring such separation in all embodiments. It should be understood that the described program components, methods, and systems may generally be integrated together in a single computer product or packaged into multiple computer products.
[0335]
[0336] Each process, method, and algorithm depicted in the description and / or figures herein may be embodied in, and wholly or partially automated by, one or more physical computing systems, hardware computer processors, special-purpose circuits, and / or electronic hardware configured to execute specialized and specific computer instructions. For example, a computing system may include a general-purpose computer (e.g., a server) or a special-purpose computer, special-purpose circuitry, etc., programmed with specific computer instructions. Code modules may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language. In some implementations, specific operations and methods may be performed by circuitry specific to a given function.
[0337] Furthermore, some implementations of the functions disclosed herein are mathematically, computationally, or technically complex enough that dedicated hardware or one or more physical computing devices (using appropriate dedicated executable instructions) may be required to perform the functions, for example, due to the amount or complexity of the computations involved, or to provide results substantially in real time. For example, animation or video may include many frames, each with millions of pixels, and requires specially programmed computer hardware to process the video data to provide the required image processing tasks or applications within a commercially reasonable timeframe.
[0338] Code modules or any type of data can be stored on any type of non-transitory computer-readable medium, such as physical computer memory, including hard disk drives, solid-state memory, random access memory (RAM), read-only memory (ROM), optical disks, volatile or non-volatile storage devices, and combinations thereof and / or similar memories. Methods and modules (or data) can also be transmitted as generated data signals (e.g., as part of a carrier wave or other analog or digital propagation signal) on various computer-readable transmission media, including wireless and wired / cable-based media, and can take many forms (e.g., as part of a single or multiplexed analog signal, or as multiple discrete digital packets or frames). The results of the disclosed process or process steps or actions can be permanently or otherwise stored in any type of non-transitory tangible computer memory, or can be transmitted via a computer-readable transmission medium.
[0339] Any process, block, state, step, or function described herein and / or depicted in the accompanying drawings should be understood as potentially representing a code module, code segment, or code section, including one or more executable instructions for implementing a specific function (e.g., logic or arithmetic) or step in the process. Various processes, blocks, states, steps, or functions may be combined, rearranged, added to, removed from, modified, or otherwise altered in the illustrative examples provided herein. In some embodiments, additional or different computing systems or code modules may perform some or all of the functions described herein. The methods and processes described herein are not limited to any particular order, and the associated blocks, steps, or states may be performed in other suitable orders (e.g., serial, parallel, or some other manner). Tasks or events may be added to or removed from the disclosed exemplary embodiments. Furthermore, the separation of various system components in the implementations described herein is for illustrative purposes and should not be construed as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems can generally be integrated together in a single computer product or packaged into multiple computer products. Many implementation variations are possible.
[0340] Processes, methods, and systems can be implemented in networked (or distributed) computing environments. Networked environments include enterprise-wide computer networks, intranets, local area networks (LANs), wide area networks (WANs), personal area networks (PANs), cloud computing networks, crowdsourced computing networks, the Internet, and the World Wide Web. This network can be wired or wireless, or any other type of communication network.
[0341] The systems and methods disclosed herein each have several innovative aspects, none of which bear sole responsibility or claim for the desired properties disclosed herein. The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. Various modifications to the implementations described herein will be apparent to those skilled in the art, and the general principles defined herein can be applied to other implementations without departing from the spirit or scope of this disclosure. Therefore, the claims are not intended to be limited to the implementations shown herein, but should be given the broadest scope consistent with the invention, principles, and novel features disclosed herein.
[0342] Some features described in this specification in the context of a single implementation may also be implemented in combination within a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any suitable sub-combination. Moreover, although features may be described above as functioning in certain combinations and even initially claimed to be so, in some cases one or more features from the claimed combination may be removed, and the claimed combination may be for sub-combinations or variations thereof. For each embodiment, no single feature or set of features is necessary or indispensable.
[0343] The conditional language used herein, particularly words such as “can,” “will,” “may,” “may,” “for example,” etc., unless explicitly stated otherwise, is understood in the context in which it is used to generally convey that certain embodiments include certain features, elements, and / or steps that are not included in other embodiments. Therefore, such conditional language is not generally intended to imply that features, elements, and / or steps are necessary in any way for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether such features, elements, and / or steps are included or will be performed in any particular embodiment, with or without author input or prompting. The terms “comprising,” “including,” “having,” etc., are synonyms, used inclusively in an open-ended manner, and do not exclude additional elements, features, actions, operations, etc. Furthermore, the term “or” is used in its inclusive sense (rather than in its exclusive sense), and thus, for example, when used to connect lists of elements, the term “or” means one, some, or all of the elements in the list. Additionally, the words “a,” “an,” and “the” used in this application and the appended claims should be interpreted as meaning “one or more” or “at least one,” unless otherwise stated.
[0344] As used herein, the phrase “at least one” in a list of items refers to any combination of those items, including individual members. For example, “at least one of A, B, or C” is intended to cover: A, B, C, A and B, A and C, B and C, and A, B, and C. Unless otherwise specifically stated, words such as the phrase “at least one of X, Y, and Z” should be understood in the context in which items, terms, etc., are generally used to convey that an item, term, etc., can be at least one of X, Y, or Z. Therefore, such combined language is generally not intended to imply that certain embodiments require the presence of at least one of X, at least one of Y, and at least one of Z.
[0345] Similarly, although operations may be depicted in the accompanying drawings in a specific order, it should be understood that it is not necessary to perform such operations in the specific or sequential order shown, or to perform all the shown operations to achieve the desired result. Furthermore, the drawings may schematically depict one or more example processes in the form of flowcharts. However, other operations not shown may be combined with the schematically shown example methods and processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any of the shown operations. Additionally, in other implementations, operations may be rearranged or reordered. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above implementations should not be construed as requiring such separation in all implementations, and it should be understood that the described program components and systems can generally be integrated into a single software product or packaged into multiple software products. Additionally, other implementations are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and the desired result may still be achieved.
Claims
1. A system comprising: A display is configured to present an image light to a user, the image light including virtual image content, wherein presenting the image light includes: presenting the image light having a wavefront divergence corresponding to a depth plane at a distance from the user; Camera; and One or more processors are configured to: The camera captures an image of the user's eyes. Based on the image of the user's eyes, determine whether the user corresponds to a registered user; Based on the determination of the user corresponding to the registered user: Determine the settings associated with the registered user, the settings corresponding to the first pre-wave divergence; and The image light having the first wavefront divergence is presented to the user via the display; and Based on the determination that the user does not correspond to the registered user: The image light, having a second wavefront divergence different from the first wavefront divergence, is presented to the user via the display.
2. The system according to claim 1, wherein, Determining whether the user corresponds to the registered user includes performing one or more of an iris scanning operation and an iris authentication operation.
3. The system according to claim 1, wherein, The second wavefront divergence corresponds to the default setting.
4. The system according to claim 1, wherein, The one or more processors are further configured to: determine, based on the fact that the user does not correspond to the registered user. Obtain the user's iris information; and Based on the iris information, the second wavefront divergence is determined.
5. The system according to claim 1, wherein, The one or more processors are also configured to: The virtual target is presented to the user via the display. The step of acquiring the image of the user's eyes includes: acquiring the image of the user's eyes while presenting the virtual target to the user.
6. The system according to claim 1, wherein, The one or more processors are further configured to: improve the accuracy or reliability of user identification based on the determination that the user corresponds to the registered user.
7. A method comprising: Presenting an image light to a user via a display, the image light including virtual image content, wherein presenting the image light includes: presenting the image light having a wavefront divergence corresponding to a depth plane at a distance from the user; The user's eyes are captured via a camera; Based on the determination that the user corresponds to the registered user: Determine the settings associated with the registered user, the settings corresponding to the first pre-wave divergence; and The image light having the first wavefront divergence is presented to the user via the display; and Based on the determination that the user does not correspond to the registered user: The image light, having a second wavefront divergence different from the first wavefront divergence, is presented to the user via the display.
8. The method according to claim 7, wherein, Determining whether the user corresponds to the registered user includes performing one or more of an iris scanning operation and an iris authentication operation.
9. The method according to claim 7, wherein, The second wavefront divergence corresponds to the default setting.
10. The method of claim 7, further comprising: Based on the determination that the user does not correspond to the registered user. Obtain the user's iris information; as well as Based on the iris information, the second wavefront divergence is determined.
11. The method of claim 7, further comprising: The virtual target is presented to the user via the display. The step of acquiring the image of the user's eyes includes: acquiring the image of the user's eyes while presenting the virtual target to the user.
12. The method of claim 7, further comprising: Based on the determination that the user corresponds to the registered user, the accuracy or reliability of user identification is improved.
13. A non-transitory computer-readable medium storing instructions, said instructions, when executed by one or more processors, causing said one or more processors to perform a method, said method comprising: An image light is presented to a user by a display, the image light including virtual image content, wherein presenting the image light includes: presenting the image light having a wavefront divergence corresponding to a depth plane at a distance from the user; The user's eyes are captured via a camera; Based on the determination that the user corresponds to the registered user: Determine the settings associated with the registered user, the settings corresponding to the first pre-wave divergence; and The image light having the first wavefront divergence is presented to the user via the display; and Based on the determination that the user does not correspond to the registered user: The image light, having a second wavefront divergence different from the first wavefront divergence, is presented to the user via the display.
14. The non-transient computer-readable medium according to claim 13, wherein, Determining whether the user corresponds to the registered user includes performing one or more of an iris scanning operation and an iris authentication operation.
15. The non-transient computer-readable medium according to claim 13, wherein, The second wavefront divergence corresponds to the default setting.
16. The non-transient computer-readable medium according to claim 13, wherein, The method further includes: determining that the user does not correspond to the registered user. Obtain the user's iris information; and Based on the iris information, the second wavefront divergence is determined.
17. The non-transient computer-readable medium according to claim 13, wherein, The method further includes: The virtual target is presented to the user via the display. The step of acquiring the image of the user's eyes includes: acquiring the image of the user's eyes while presenting the virtual target to the user.
18. The non-transient computer-readable medium according to claim 13, wherein, The method includes: improving the accuracy or reliability of user identification based on the determination that the user corresponds to the registered user.
Citation Information
Patent Citations
Systems and methods for augmented and virtual reality
US10262462B2
Selecting virtual objects in a three-dimensional space
US10521025B2
Methods and systems for detecting and combining structural features in 3D reconstruction
US10559127B2
Periocular test for mixed reality calibration
US10573042B2
Virtual and augmented reality systems and methods
US10698215B2