Virtual Object Movement Speed Curve for Virtual and Augmented Reality Display Systems

The head-mounted display system addresses user discomfort and inaccurate eye tracking in VR/AR by moving virtual objects with an S-shaped velocity curve, enhancing comfort and accuracy in eye tracking calibration.

JP7712279B2Active Publication Date: 2025-07-23MAGIC LEAP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022547886
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-14
Filing Date
2021-02-11
Publication Date
2025-07-23
Estimated Expiration
2041-02-11

AI Technical Summary

Technical Problem

Existing virtual reality and augmented reality systems face challenges in providing a comfortable and accurate presentation of virtual objects due to inconsistencies between convergence and accommodation movements, leading to user discomfort and inaccurate eye tracking calibration.

Method used

A head-mounted display system that moves virtual objects using an S-shaped velocity curve, aligning with the user's eye tracking and focusing times to enhance comfort and accuracy during calibration processes.

Benefits of technology

The S-shaped velocity curve improves user experience by reducing overshoot and increasing accuracy in eye tracking calibration, providing a more natural and responsive viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007712279000004
    Figure 0007712279000004
  • Figure 0007712279000005
    Figure 0007712279000005
  • Figure 0007712279000006
    Figure 0007712279000006
Patent Text Reader

Abstract

Systems and methods are described for adjusting the speed of movement of a virtual object presented by a wearable system. The wearable system may present three-dimensional (3D) virtual content that moves, for example, laterally across a user's field of view and / or within the user's perceived depth. The speed of movement may follow an S-curve profile, gradually increasing to a maximum speed and then gradually decreasing until the end of the movement is reached. The speed decrease may be more gradual than the speed increase. This speed curve may be used in moving the virtual object for eye-tracking calibration. The wearable system may track the position of a virtual object (eye-tracking target) that moves at a certain speed, following an S-curve.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority of U.S. Provisional Application No. 62 / 976,977, filed on February 14, 2020, entitled "VIRTUAL OBJECT MOVEMENT SPEED CURVE FOR VIRTUAL AND AUGMENTED REALITY DISPLAY SYSTEMS", which is incorporated herein by reference in its entirety under 35 U.S.C. § 119(e).

[0002] The present disclosure relates to virtual reality and augmented reality imaging and visualization systems, and more particularly, to speed curves for the movement of virtual objects.

Background Art

[0003] Modern computing and display technologies have facilitated the development of systems for so - called "virtual reality" or "augmented reality" experiences, in which digitally reproduced images or portions thereof are presented to a user in a manner that appears or is perceived to be real. Virtual reality, i.e., "VR" scenarios, typically involve the presentation of digital or virtual image information without transparency to other actual real - world visual inputs, and augmented reality, i.e., "AR" scenarios, typically involve the presentation of digital or virtual image information as an augmentation to the visualization of the actual world around the user. Mixed reality or "MR" scenarios are a type of AR scenario that typically involves virtual objects integrated into and responsive to the natural world. For example, an MR scenario may include AR image content that is perceived to be blocked by or otherwise interact with objects within the real world.

[0004] Referring to FIG. 1, an AR scene 10 is depicted. To the user of the AR technology, a real-world park-like setting 20 featuring people, trees, buildings in the background, and a concrete platform 30 is visible. The user also "sees" "virtual content" such as a robot image 40 standing on the real-world platform 30 and an avatar character 50 like a flying comic that appears anthropomorphic like a bumblebee. These elements 50, 40 are "virtual" in that they do not exist in the real world. The human visual perception system is complex, and it is difficult to produce AR technology that promotes a comfortable, natural, and rich presentation of virtual image elements among other virtual or real-world image elements.

Summary of the Invention

Means for Solving the Problems

[0005] Various embodiments of techniques for moving virtual content are disclosed.

[0006] In some embodiments, a head-mounted display system is disclosed. The head-mounted display system comprises a display configured to present virtual content to the user, and a hardware processor in communication with the display, the hardware processor programmed to cause the display system to display a virtual object at a first location and to move the virtual object at a variable speed to a second location based on an S-shaped velocity curve.

[0007] In some embodiments, a method for moving virtual content is disclosed. The method includes the steps of displaying a virtual object at a first location on a display system and moving the virtual object at a variable speed to a second location on the display system following an S-shaped velocity curve.

[0008] In some embodiments, a wearable system for eye tracking calibration is disclosed. The system includes a display configured to display an eye calibration target to a user, and a hardware processor that communicates with the non-transitory memory and the display system, the hardware processor causing the display to display the eye calibration target at a first target location, identifying a second location different from the first target location, determining a distance between the first target location and the second target location, determining a total allocation time for moving the eye calibration target from the first target location to the second target location, calculating a target movement speed curve based on the total allocation time, the target movement speed curve being an S-curve, and programming the hardware processor to move the eye calibration target over a total time according to the target movement speed curve.

[0009] In some embodiments, a method for eye tracking calibration is disclosed. The method includes displaying an eye calibration target at a first target location within a user's environment, identifying a second location different from the first target location, determining a distance between the first target location and the second target location, determining a total allocation time for moving the eye calibration target from the first target location to the second target location, calculating a target movement speed curve based on the total allocation time, the target movement speed curve being an S-curve, and moving the eye calibration target over a total time according to the target movement speed curve.

[0010] In some embodiments, a wearable system for eye tracking calibration is described. The system includes an image capture device configured to capture an eye image of one or both eyes of a user of the wearable system, a non-transitory memory configured to store the eye image, a display system through which the user can perceive an eye calibration target within the user's environment, and a hardware processor that communicates with the non-transitory memory and the display system. The hardware processor is programmed to make the eye calibration target perceptible at a first target location within the user's environment via the display system, identify a second target location within the user's environment that is different from the first target location, determine a dioptric distance between the first target location and the second target location, determine a total time for moving the eye calibration target from the first target location to the second target location based on the distance, at least partially interpolate a target position based on the reciprocal of the dioptric distance, and move the eye calibration target as a function of time over the total time according to the interpolated target position.

[0011] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will be apparent from the description, the drawings, and the claims. Neither this summary nor the following detailed description purports to define or limit the scope of the subject matter of the invention. The present invention provides, for example, the following. (Item 1) A head-mounted display system, A display configured to present virtual content to a user, A hardware processor that communicates with the display, the hardware processor Causes the display system to display a virtual object at a first location, Causes the display system to move the virtual object to a second location at a variable speed based on an S-shaped speed curve A hardware processor programmed to perform A head-mounted display system comprising (Item 2) One or more hardware processors Determine the diopter distance between the first location and the second location, Determine the distance between the focal plane of the first location and the focal plane of the second location, Determine the total time for moving the virtual object from the first location to the second location, at least in part, based on the diopter distance and the distance between the focal plane of the first location and the focal plane of the second location The head-mounted display system according to Item 1, configured to perform (Item 3) To determine the total time, the one or more hardware processors Determine at least one tracking time associated with at least one eye of the user based on the diopter distance, Determine at least one focusing time associated with at least one eye of the user based on the focal plane distance, Select the total time from the at least one tracking time and the at least one focusing time The head-mounted display system according to Item 2, configured to perform (Item 4) The at least one tracking time includes the time for at least one eye of the user to move angularly over the diopter distance, the head-mounted display system according to Item 3. (Item 5) The at least one tracking time includes a tracking time associated with the left eye and a tracking time associated with the right eye, the head-mounted display system according to Item 3. (Item 6) The head-mounted display system according to item 3, wherein the at least one focusing time includes a time for at least one eye of the user to focus over a distance between a focal plane at the first location and a focal plane at the second location. (Item 7) The head-mounted display system according to item 3, wherein the at least one focusing time includes a focusing time associated with the left eye and a focusing time associated with the right eye. (Item 8) The head-mounted display system according to item 2, wherein the one or more hardware processors are configured to determine the variable speed based on the total time. (Item 9) The head-mounted display system according to item 2, wherein the one or more hardware processors are configured to interpolate the position of the virtual object as a function of time based on the total time or one or more parameters associated with the diopter distance. (Item 10) The head-mounted display system according to item 1, wherein the variable speed is based on the reciprocal of the diopter distances at the first and second locations. (Item 11) The head-mounted display system according to item 1, wherein the one or more parameters include uniform and non-uniform parameters, and the non-uniform parameters are configured to vary in value at a variable rate associated with one or more spline curves as a function of the uniform parameters. (Item 12) The head-mounted display system according to item 11, wherein the uniform parameter is at least partially based on the total time. (Item 13) The head-mounted display system according to item 1, comprising an image capture device configured to capture an eye image of one or both eyes of a user of the wearable system. (Item 14) The head-mounted display system according to item 1, wherein the virtual content includes an eye calibration target. (Item 15) A method of moving virtual content, the method comprising: displaying a virtual object at a first location on a display system; and moving the virtual object to a second location at a variable speed following an S-shaped speed curve on the display system. A method as described above. (Item 16) The method further comprises: determining a diopter distance between the first location and the second location. Determining a distance between a focal plane at the first location and a focal plane at the second location; Determining a total time for moving the virtual object from the first location to the second location, at least in part, based on the dioptric distance and the distance between the focal plane at the first location and the focal plane at the second location; The method according to item 15, comprising the above. (Item 17) Determining the total time comprises: Determining at least one tracking time associated with at least one eye of the user based on the dioptric distance; Determining at least one focusing time associated with at least one eye of the user based on the focal plane distance; Selecting the total time from the at least one tracking time and the at least one focusing time; The method according to item 16, comprising the above. (Item 18) The method according to item 17, wherein the at least one tracking time comprises a time for at least one eye of the user to move angularly over the dioptric distance. (Item 19) The method according to item 17, wherein the at least one tracking time comprises a tracking time associated with the left eye and a tracking time associated with the right eye. (Item 20) The method according to item 17, wherein the at least one focusing time comprises a time for at least one eye of the user to focus over the distance between the focal plane at the first location and the focal plane at the second location. (Item 21) The method according to item 17, wherein the at least one focusing time comprises a focusing time associated with the left eye and a focusing time associated with the right eye. (Item 22) The method according to item 17, further comprising determining the variable speed based on the total time. (Item 23) The method according to item 17, wherein the variable speed is based on the reciprocal of the dioptric distance. (Item 24) Interpolating the position of the virtual object as a function of time, which comprises interpolating an S-curve function using one or more parameters based on the total time or the dioptric distance, the method according to item 15. (Item 25) The one or more than one parameter includes uniform and non-uniform parameters, and the non-uniform parameters are configured to change values at a variable rate associated with one or more spline curves as a function of the uniform parameters, according to the method of item 24. (Item 26) The uniform parameter is at least partially based on the total time, according to the method of item 24. (Item 27) The virtual content includes an eye calibration target, according to the method of item 15. (Item 28) A wearable system for eye tracking calibration, the system A display configured to display an eye calibration target to a user, A hardware processor communicating with a non-transitory memory and a display system, the hardware processor Cause the eye calibration target to be displayed at a first target location on the display, Identify a second location different from the first target location, Determine the distance between the first target location and the second target location, Determine a total allocation time for moving the eye calibration target from the first target location to the second target location, Calculating a target movement speed curve based on the total allocation time, the target movement speed curve being an S-curve, Moving the eye calibration target over the total time according to the target movement speed curve A hardware processor programmed to perform A system comprising (Item 29) The hardware processor Identify the user's eye line of sight based on data obtained from the image capture device, Determine whether the user's eye line of sight matches the eye calibration target at the second target location, In response to determining that the user's eye line of sight matches the eye calibration target at the second location, instruct the image capture device to capture the eye image and start storing the eye image in the non-transitory memory A wearable system according to item 28, configured to perform (Item 30) A method for eye tracking calibration, the method Display an eye calibration target at a first target location within the user's environment, Identify a second location different from the first target location, Determine the distance between the first target location and the second target location, Determining a total allocation time for moving the eye correction target from the first target location to the second target location; Calculating a target movement speed curve based on the total allocation time, wherein the target movement speed curve is an S-curve; Moving the eye correction target over the total time according to the target movement speed curve; A method comprising. (Item 31) The method includes: Identifying the user's eye gaze based on data obtained from the image capture device; Determining whether the user's eye gaze aligns with the eye correction target at the second target location; In response to determining that the user's eye gaze aligns with the eye correction target at the second location, instructing the image capture device to capture the eye image and start storing the eye image in the non-transitory memory; The method according to item 30, comprising. (Item 32) A wearable system for eye tracking calibration, the system comprising: An image capture device configured to capture an eye image of one or both eyes of a user of the wearable system; A non-transitory memory configured to store the eye image; A display system through which the user can perceive an eye correction target in the user's environment; A hardware processor in communication with the non-transitory memory and the display system, the hardware processor: Making the eye correction target perceivable at a first target location in the user's environment via the display system; Identifying a second target location in the user's environment different from the first target location; Determining a dioptric distance between the first target location and the second target location; Determining a total time for moving the eye correction target from the first target location to the second target location based on the distance; Interpolating target positions based at least in part on the reciprocal of the dioptric distance; Moving the eye correction target as a function of time over the total time according to the interpolated target positions; A hardware processor programmed to perform; A system comprising. (Item 33) The hardware processor is Identifying the user's eye gaze based on data obtained from the image capture device; Determining whether the user's eye gaze aligns with the eye calibration target at the second target location; In response to determining that the user's eye gaze aligns with the eye calibration target at the second location, instructing the image capture device to capture the eye image and initiate storage of the eye image in the non-transitory memory; The wearable system according to item 32, configured to perform the above.

Brief Description of the Drawings

[0012]

Figure 1

[0013]

Figure 2

[0014]

Figure 3

[0015]

Figure 4A

[0016]

Figure 4B

[0017]

Figure 4C

[0018]

Figure 4D

[0019]

Figure 5

[0020]

Figure 6

[0021]

Figure 7

[0022]

Figure 8

[0023]

Figure 9A

[0024]

Figure 9B

[0025]

Figure 9C

[0026]

Figure 9D

[0027]

Figure 9E

[0028]

Figure 10

[0029]

Figure 11A

[0030]

Figure 11B

[0031]

Figure 11C

[0032]

Figure 11D

[0033]

Figure 12A

[0034]

Figure 12B

[0035]

Figure 12C

[0036]

Figure 13A

[0037]

Figure 13B

[0038]

Figure 14A

[0039]

Figure 14B

[0040]

Figure 15A

Figure 15B

[0041]

Figure 16

[0042]

Figure 17

[0043]

Figure 18

[0044]

Figure 19

[0045]

Figure 20

[0046] The drawings are provided to illustrate exemplary embodiments described in this specification and are not intended to limit the scope of the present disclosure.

Mode for Carrying Out the Invention

[0047] Detailed Description Wearable display devices such as head-mounted displays (HMDs) can present virtual content (or virtual objects) within a two-way VR / AR / MR environment. The virtual content may include data elements that can be interacted with by the user through various poses such as head pose, eye gaze, or body pose. In the context of user interaction using eye gaze, the wearable display system may collect eye data such as eye images (e.g., via an eye camera within an imaging system facing inward of the wearable display system). The wearable display system (also simply referred to herein as the wearable system) may calculate the user's eye gaze direction based on a mapping matrix that provides an association between the user's eye gaze and a gaze vector (which may indicate the user's line of sight direction). To improve the user experience, the wearable display system may perform an eye tracking calibration process that calibrates the mapping matrix and takes into account the uniqueness of each person's eyes, the specific orientation of the wearable display system with respect to the user when worn, current environmental conditions (e.g., lighting conditions, temperature, etc.), combinations thereof, or equivalents.

[0048] During an eye-tracking calibration process, a wearable display system may present various virtual targets and instruct the user to look at these virtual targets while collecting information regarding the user's eye gaze. However, the eye calibration process can be uncomfortable and fatiguing for the user. If the target moves at a constant speed such that it suddenly starts at one location within the user's environment and stops at a second location, the user's eyes may overshoot, resulting in a more redundant and uncomfortable eye calibration. The more redundant the eye calibration becomes, the higher the likelihood that the user will become distracted, thereby providing poor results and / or the user may completely stop participating in the calibration process. If the user does not look at the target as instructed, the wearable display system may collect data that does not accurately reflect the user's line of sight, which can introduce inaccuracies into the calibration and cause an incorrect mapping matrix to be generated. As a result of inaccuracies in the calibration process, if the wearable display system is to use eye gaze as an input, for example, as an interaction input, the user may be unable to accurately target and interact with an object, which can lead to a less satisfying user experience. What is disclosed herein is a system and method that provides for a more comfortable movement of a virtual object between two locations, whether or not it is used for eye-tracking calibration. This virtual object movement may advantageously be used to improve the accuracy and comfort associated with the eye calibration process.

[0049] In some embodiments, the virtual object is moved at a speed that follows an S-curve. For example, if the X-axis represents time and the Y-axis represents distance, the plot of the movement of the virtual object from a first position to a second position may define a generally S-shape. From the first position, the virtual object may gradually increase its speed to a maximum value and then gradually decrease its speed until it stops at the second position. In some embodiments, the decrease in speed may be more gradual than the increase in speed. The duration and rate at which the speed decreases may be selected, for example, to prevent the eye from overshooting the second position, which may have advantages for accurate assessment of eye tracking. Additionally, in some embodiments, using an S-shaped speed curve, the time required to move the virtual object from the first position to the second position can be shortened compared to the time required to perform a similar movement where the virtual object moves at a constant speed. This can also shorten the time required to perform eye tracking calibration. In some embodiments, assuming a set amount of time (travel time) to move from the first position to the second position, the S-shaped speed curve may be adapted within that duration. In some embodiments, the travel time may be determined based on user characteristics (e.g., the natural speed at which the user's eye can move or change focus, and / or the user's age, gender, eye health, etc.) or the magnitude of the distance between the first location and the second location. As used herein, it should be understood that the speed curve is a plot of time versus the distance traversed by the virtual object. As discussed herein, the curve is preferably generally shaped like the letter S.

[0050] In some embodiments, a particular value of the velocity along the S-shaped velocity curve may be selected based on the location within the virtual space where the first and second locations are positioned. For example, the value of the velocity may exceed the value of the velocity for the movement of the virtual object across the depth plane (e.g., the lateral movement of the virtual object) with respect to the movement of the virtual object within a single depth plane (e.g., the movement of the virtual object closer to or farther away from the user or vice versa). Additionally, at locations closer to the user, the velocity of the virtual object may be lower than the velocity of the virtual object at locations farther from the user.

[0051] Advantageously, the movement of the virtual object with an S-shaped velocity curve can provide a more comfortable viewing experience for the user. For example, such movement can reduce or prevent the eye from overshooting the position where the virtual object stops. In some applications such as eye tracking or eye tracking calibration, the reduction in overshoot can increase the accuracy of eye tracking or calibration, which can improve the user experience. For example, the increased accuracy can provide a more natural and responsive viewing experience.

[0052] In addition to eye calibration, the virtual object movement method described herein can be applicable to any movement of virtual objects within a 2D or 3D environment. For example, a display system may move a virtual game object according to the movement system and method described herein to prompt the user to more easily track the virtual object. In another example, a display system may move a target for the purpose of physical therapy using the movement system and method. In another example, a display system may, in one example, use the movement system and method to move a target for the purpose of collecting data about the user's environment or for the purpose of guiding the user to experience its 3D environment. However, while one application is described herein, any number of applications of the target movement system and method are possible.

[0053] Reference is made to the drawings, where like reference numerals refer to like parts throughout. Unless otherwise indicated, the drawings are schematic and are not necessarily drawn to scale. A. Exemplary Wearable 3D Display System

[0054] FIG. 2 illustrates a conventional display system for simulating a three-dimensional image for a user. It should be understood that when the user's eyes are spaced apart and looking at a real object in space, each eye has a slightly different view of the object and can form an image of the object at different locations on the retina of each eye. This may be referred to as binocular disparity and can be utilized by the human visual system to provide a perception of depth. The conventional display system simulates binocular disparity by presenting two distinct different images 190, 200 with slightly different views of one identical virtual object for each of the eyes 210, 220, corresponding to the views of the virtual object that would be seen by each eye as if the virtual object were a real object at the desired depth. These images provide binocular cues that can be interpreted by the user's visual system to derive a perception of depth.

[0055] Continuing to refer to FIG. 2, the images 190, 200 are spaced apart from the eyes 210, 220 by a distance 230 on the z-axis. The z-axis is parallel to the optical axis of the viewer in a state where the eye is fixating on an object at optical infinity directly in front of the viewer. The images 190, 200 are flat and at a fixed distance from the eyes 210, 220. Based on the slightly different views of the virtual object in the images presented to the eyes 210, 220 respectively, the eyes can necessarily rotate so that the images of the object come to corresponding points on the respective retinas of the eyes and a single binocular vision is maintained. This rotation can converge the respective lines of sight of the eyes 210, 220 onto a point in the space where the virtual object is perceived to exist. As a result, the provision of a three-dimensional image has conventionally involved manipulating the convergence and divergence movements of the user's eyes 210, 220 and providing binocular cues that are interpreted by the human visual system to provide a perception of depth.

[0056] However, the generation of realistic and comfortable perception of depth is difficult. It should be understood that light from objects at different distances from the eye has wavefronts with different divergence amounts. FIGS. 3A - 3C illustrate the relationship between distance and the divergence of light rays. The distances between the object and the eye 210 are represented in the order of decreasing distances R1, R2, and R3. As shown in FIGS. 3A - 3C, the light rays diverge more as the distance to the object decreases. Conversely, as the distance increases, the light rays become more collimated. In other words, it can be said that the light field generated by a point (object or part of an object) has a spherical wavefront curvature that is a function of the distance the point is away from the user's eye. The curvature increases as the distance between the object and the eye 210 decreases. Only the monocular eye 210 is illustrated in FIGS. 3A - 3C and various other figures herein for clarity of illustration, but the discussion regarding the eye 210 can be applied to both eyes 210 and 220 of the viewer.

[0057] Continuing to refer to FIGS. 3A - 3C, light from an object on which a viewer's eye is fixated can have different wavefront divergences. Due to the different amounts of wavefront divergence, the light can be focused differently by the eye's lens, which in turn can require the lens to assume different shapes to form a focused image on the eye's retina. If the focused image is not formed on the retina, the resulting retinal blur acts as a cue for accommodation, causing a change in the shape of the eye's lens until the focused image is formed on the retina. For example, the cue for accommodation triggers relaxation or contraction of the ciliary muscle surrounding the eye's lens, thereby modulating the force applied to the zonular fibers that hold the lens, and thus changing the shape of the eye's lens until the retinal blur of the fixated object is eliminated or minimized, thereby forming a focused image of the fixated object on the eye's retina (e.g., the fovea). The process by which the eye's lens changes shape can be referred to as accommodation, and the shape of the eye's lens required to form a focused image of a fixated object on the eye's retina (e.g., the fovea) can be referred to as the accommodative state.

[0058] Referring now to FIG. 4A, the representation of the accommodation-convergence / divergence response of the human visual system is illustrated. The movement of the eyes to fixate on an object causes the eyes to receive light from the object, and the light forms an image on each of the retinas of the eyes. The presence of retinal blur in the images formed on the retinas can provide a cue for accommodation, and the relative location of the images on the retinas can provide a cue for convergence / divergence. The cue for accommodation causes accommodation to occur, resulting in the eye's lens assuming a particular accommodation state in which a focused image of the object is formed on the retina of the eye (e.g., the fovea). On the other hand, the cue for convergence / divergence causes convergence / divergence movement (rotation of the eyes) such that the images formed on the retinas of each eye are at corresponding retinal points that maintain single binocular vision. At these positions, it can be said that the eyes are in a particular convergence / divergence state. Continuing to refer to FIG. 4A, accommodation can be understood as the process by which the eyes achieve a particular accommodation state, and convergence / divergence can be understood as the process by which the eyes achieve a particular convergence / divergence state. As shown in FIG. 4A, the accommodation and convergence / divergence states of the eyes can change when the user fixates on another object. For example, the accommodated state can change when the user fixates on a new object at a different depth along the z-axis.

[0059] Without being limited by theory, it is believed that the viewer of an object can perceive the object as "three-dimensional" due to the combination of convergence / divergence and accommodation. As described above, the convergence / divergence movement of the two eyes relative to each other (e.g., the rotation of the eyes such that the pupils move towards or away from each other and converge the lines of sight of the eyes to fixate on an object) is closely associated with the accommodation of the eye's lens. Under normal conditions, a change in the shape of the eye's lens to change focus from one object to another object at a different distance will automatically cause a corresponding change in convergence / divergence to the same distance under a relationship known as the "accommodation-convergence / divergence reflex". Similarly, a change in convergence / divergence will, under normal conditions, induce a corresponding change in the shape of the lens.

[0060] Referring now to FIG. 4B, examples of different accommodation and convergence / divergence movement states of the eyes are illustrated. The pair of eyes 222a fixates on an object at optical infinity, while the pair of eyes 222b fixates on an object 221 at less than optical infinity. It should be noted that the convergence / divergence movement states of each pair of eyes are different, with the pair of eyes 222a being directed straight, while the pair of eyes 222 converges onto the object 221. The accommodation states of the eyes forming each pair of eyes 222a and 222b are also different, as represented by the different shapes of the lenses 210a, 220a.

[0061] Unfortunately, many users of conventional "3-D" display systems find such conventional systems uncomfortable or do not perceive any sense of depth at all due to the inconsistency between accommodation and convergence / divergence movement states in these displays. As described above, many stereoscopic or "3-D" display systems display a scene by providing slightly different images to each eye. Such systems are uncomfortable for many viewers because they simply provide different presentations of the scene, causing a change in the convergence / divergence movement state of the eyes, but without a corresponding change in the accommodation state of those eyes. Rather, the images are presented at a fixed distance from the eyes by the display such that the eyes view all the image information in a single accommodation state. Such an arrangement goes against the "accommodation-convergence / divergence reflex" by causing a change in the convergence / divergence movement state without a corresponding change in the accommodation state. This inconsistency is thought to cause viewer discomfort. A display system that provides better alignment between accommodation and convergence / divergence movement can create a more realistic and comfortable simulation of three-dimensional images.

[0062] While not limited by theory, the human eye is typically thought to be able to interpret a finite number of depth planes and provide depth perception. As a result, a highly realistic simulation of the perceived depth can be achieved by providing different presentations of images corresponding to each of these limited number of depth planes to the eye. In some embodiments, the different presentations may provide both cues for convergence / divergence motion and matching cues for accommodation, thereby providing physiologically correct accommodation-convergence / divergence motion matching.

[0063] Continuing to refer to FIG. 4B, two depth planes 240 corresponding to different distances in space from eyes 210, 220 are illustrated. For a given depth plane 240, convergence / divergence motion cues may be provided by appropriately displaying images of different viewpoints for each of eyes 210, 220. Additionally, for a given depth plane 240, the light forming the images provided to each eye 210, 220 may have a wavefront divergence corresponding to a light field generated by a point at the distance of that depth plane 240.

[0064] In the illustrated embodiment, the distance along the z-axis of depth plane 240 containing point 221 is 1 m. As used herein, the distance or depth along the z-axis may be measured using a zero point located at the exit pupil of the user's eye. Thus, a depth plane 240 located at a depth of 1 m corresponds to a distance 1 m away from the exit pupil of the user's eye on the optical axis of those eyes in a state where the eyes are directed towards optical infinity. As an approximation, the depth or distance along the z-axis may be measured from the front display of the user's eye (e.g., the surface of a waveguide), and a value related to the distance between the device and the exit pupil of the user's eye may be added. That value is referred to as the interpupillary distance and may correspond to the distance between the exit pupil of the user's eye and the front display of the user's eye worn by the user. In practice, the value for the interpupillary distance may generally be a normalized value used for all viewers. For example, the interpupillary distance may be assumed to be 20 mm, and the depth plane at a depth of 1 m may be at a distance of 980 mm in front of the display.

[0065] Referring now to FIGS. 4C and 4D, examples of consistent vergence-accommodation movement distances and inconsistent vergence-accommodation movement distances are illustrated, respectively. As shown in FIG. 4C, the display system may provide an image of the virtual object to each of the eyes 210, 220. The image may cause the eyes 210, 220 to assume a vergence-accommodation state in which the eyes converge on a point 15 on the depth plane 240. In addition, the image may be formed by light having a wavefront curvature corresponding to the real object in the depth plane 240. As a result, the eyes 210, 220 assume an accommodation state in which the image is focused on their retinas. Thus, the user may perceive the virtual object as being at the point 15 on the depth plane 240.

[0066] It should be understood that the accommodation and vergence-accommodation states of the eyes 210, 220 are each associated with a specific distance on the z-axis. For example, an object at a specific distance from the eyes 210, 220 causes those eyes to assume a specific accommodation state based on the distance of the object. The distance associated with a specific accommodation state may be referred to as the accommodation distance A d Similarly, there exists a specific vergence distance V d associated with a specific vergence-accommodation state or the eyes in a particular position relative to each other. When the accommodation distance and the vergence distance are consistent, the relationship between accommodation and vergence can be said to be physiologically correct. This is regarded as the most comfortable scenario for the viewer.

[0067] However, in a stereoscopic display, the focusing distance and the vergence / accommodation movement distance may not always match. For example, as shown in FIG. 4D, the images displayed on eyes 210, 220 may be displayed with wavefront divergence corresponding to the depth plane 240, and eyes 210, 220 may assume a specific focusing state in which points 15a, 15b on that depth plane are in focus. However, the images displayed on eyes 210, 220 may provide a cue for vergence / accommodation movement that converges eyes 210, 220 on points 15 that are not located on the depth plane 240. As a result, in some embodiments, the focusing distance corresponds to the distance from the exit pupils of eyes 210, 220 to the depth plane 240, while the vergence / accommodation movement distance corresponds to a greater distance from the exit pupils of eyes 210, 220 to points 15. The focusing distance is different from the vergence / accommodation movement distance. As a result, there is a focusing-vergence / accommodation movement mismatch. Such a mismatch is considered undesirable and can cause discomfort to the user. The mismatch corresponds to a distance (e.g., V d -A d ) and can be characterized using diopters.

[0068] It should be understood that in some embodiments, as long as the same reference point is used for the focusing distance and the vergence / accommodation movement distance, a reference point other than the exit pupils of eyes 210, 220 may be used to determine the distance for determining the focusing-vergence / accommodation movement mismatch. For example, the distance can be measured from the cornea to the depth plane, from the retina to the depth plane, from the eyepiece (e.g., the waveguide of the display device) to the depth plane, etc.

[0069] Although not limited by theory, it is believed that a user may perceive vergence-accommodation mismatches of up to about 0.25 diopters, up to about 0.33 diopters, and up to about 0.5 diopters as physiologically correct without the mismatch itself causing significant discomfort. In some embodiments, a display system disclosed herein (e.g., display system 250, FIG. 6) presents an image having a vergence-accommodation mismatch of about 0.5 diopters or less to a viewer. In some other embodiments, the vergence-accommodation mismatch of an image provided by the display system is about 0.33 diopters or less. In yet other embodiments, the vergence-accommodation mismatch of an image provided by the display system is about 0.25 diopters or less, including about 0.1 diopters or less.

[0070] FIG. 5 illustrates a side view of an approach for simulating a three-dimensional image by modifying wavefront divergence. The display system includes a waveguide 270 configured to receive light 770 encoded with image information and output the light to the user's eye 210. The waveguide 270 may output light 650 with a defined amount of wavefront divergence corresponding to the wavefront divergence of the light field generated by a point on a desired depth plane 240. In some embodiments, the same amount of wavefront divergence is provided for all objects presented on that depth plane. Additionally, the user's other eye would be illustrated as being provided with image information from a similar waveguide.

[0071] In some embodiments, a single waveguide may be configured to output light with a set amount of wavefront divergence corresponding to a single or limited number of depth planes, and / or the waveguide may be configured to output light within a limited range of wavelengths. As a result, in some embodiments, multiple or stacked waveguides may be utilized to provide different amounts of wavefront divergence for different depth planes and / or output light within different ranges of wavelengths. It should be understood that as used herein, a depth plane may be a plane or may follow the contour of a curved surface.

[0072] FIG. 6 illustrates an example of a stack of waveguides for outputting image information to a user. The display system 250 includes a stack or stacked waveguide assembly 260 of waveguides 270, 280, 290, 300, 310 that can be utilized to provide three-dimensional perception to the eye / brain. It should be understood that the display system 250 may be considered a light field display in some embodiments. Additionally, the waveguide assembly 260 may also be referred to as an eyepiece.

[0073] In some embodiments, the display system 250 may be configured to provide a substantially continuous cue for convergence / divergence movement and a plurality of discrete cues for depth adjustment. The cue for convergence / divergence movement may be provided by displaying different images to each of the user's eyes, and the cue for depth adjustment may be provided by outputting light that forms an image with a selectable discrete amount of wavefront divergence. In other words, the display system 250 may be configured to output light with a variable level of wavefront divergence. In some embodiments, each discrete level of wavefront divergence may correspond to a particular depth plane and may be provided by a particular one of the waveguides 270, 280, 290, 300, 310.

[0074] Continuing to refer to FIG. 6, waveguide assembly 260 may also include a plurality of features 320, 330, 340, 350 between the waveguides. In some embodiments, features 320, 330, 340, 350 may be one or more lenses. Waveguides 270, 280, 290, 300, 310 and / or the plurality of lenses 320, 330, 340, 350 may be configured to transmit image information to the eye using various levels of wavefront curvature or ray divergence. Each waveguide level may be associated with a particular depth plane and may be configured to output image information corresponding to that depth plane. Image input devices 360, 370, 380, 390, 400 may function as light sources for the waveguides and may be utilized to input image information into waveguides 270, 280, 290, 300, 310, and may each be configured to disperse incident light across the respective waveguide to output it towards the eye 210, as described herein. Light exits from the output surfaces 410, 420, 430, 440, 450 of image input devices 360, 370, 380, 390, 400 and is input into the corresponding input surfaces 460, 470, 480, 490, 500 of waveguides 270, 280, 290, 300, 310. In some embodiments, input surfaces 460, 470, 480, 490, 500 may each be the edge of the corresponding waveguide or a portion of the major surface of the corresponding waveguide (i.e., one of the waveguide surfaces facing directly towards the world 510 or the viewer's eye 210). In some embodiments, a single beam of light (e.g., a collimated beam) may be input into each waveguide and output an entire field of cloned collimated beams directed towards the eye 210 at a particular angle (and amount of divergence) corresponding to the depth plane associated with the particular waveguide. In some embodiments, a single one of image input devices 360, 370, 380, 390, 400 may be associated with and input light into a plurality (e.g., three) of waveguides 270, 280, 290, 300, 310.

[0075] In some embodiments, the image input devices 360, 370, 380, 390, 400 are discrete displays that each generate image information for input into their respective waveguides 270, 280, 290, 300, 310. In some other embodiments, the image input devices 360, 370, 380, 390, 400 are the output ends of a single multiplexed display that can send image information to each of the image input devices 360, 370, 380, 390, 400, for example, via one or more optical conduits (such as optical fiber cables). It should be understood that the image information provided by the image input devices 360, 370, 380, 390, 400 may include light of different wavelengths or colors (e.g., different primary colors as discussed herein).

[0076] In some embodiments, the light input into waveguides 270, 280, 290, 300, 310 is provided by a light projection system 520, which includes an optical module 530, which may include a light emitter such as a light emitting diode (LED). The light from the optical module 530 may be directed and modified by a light modulator 540, such as a spatial light modulator, via a beam splitter 550. The light modulator 540 may be configured to vary the perceived intensity of the light input into waveguides 270, 280, 290, 300, 310 and to encode the light with image information. Examples of spatial light modulators include liquid crystal displays (LCDs), including liquid crystal on silicon (LCOS) displays. In some other embodiments, the spatial light modulator may be a MEMS device such as a digital light processing (DLP) device. Image input devices 360, 370, 380, 390, 400 are shown schematically, and in some embodiments, these image input devices may represent different optical paths and locations within a common projection system configured to output light into the associated ones of waveguides 270, 280, 290, 300, 310. In some embodiments, the waveguides of the waveguide assembly 260 may function as an ideal lens while relaying the light input into the waveguides to the user's eye. In this concept, the object may be the spatial light modulator 540, and the image may be an image on a depth plane.

[0077] In some embodiments, the display system 250 may be a scanning fiber display comprising one or more scanning fibers configured to project light in various patterns (e.g., raster scan, helical scan, Lissajous pattern, etc.) into one or more waveguides 270, 280, 290, 300, 310 and ultimately into the viewer's eye 210. In some embodiments, the illustrated image input devices 360, 370, 380, 390, 400 may schematically represent a single scanning fiber or a bundle of scanning fibers configured to input light into one or more waveguides 270, 280, 290, 300, 310. In some other embodiments, the illustrated image input devices 360, 370, 380, 390, 400 may schematically represent a plurality of scanning fibers or a plurality of bundles of scanning fibers, each configured to input light into an associated one of the waveguides 270, 280, 290, 300, 310. It should be understood that one or more optical fibers may be configured to transmit light from the optical module 530 to one or more waveguides 270, 280, 290, 300, 310. It should be understood that one or more intervening optical structures may be provided between the scanning fiber or fibers and the one or more waveguides 270, 280, 290, 300, 310, for example, to redirect light exiting the scanning fiber into one or more of the waveguides 270, 280, 290, 300, 310.

[0078] Controller 560 controls the operation of one or more of the stacked waveguide assemblies 260, including the operation of the image input devices 360, 370, 380, 390, 400, the light source 530, and the light modulator 540. In some embodiments, controller 560 is part of the local data processing module 140. Controller 560 includes programming (e.g., instructions in a non-transitory medium) that adjusts the timing and provision of image information to waveguides 270, 280, 290, 300, 310, for example, according to any of the various schemes disclosed herein. In some embodiments, the controller may be a single integrated device or a distributed system connected by wired or wireless communication channels. Controller 560 may be part of processing module 140 or 150 (FIG. 9E) in some embodiments.

[0079] Continuing to refer to FIG. 6, the waveguides 270, 280, 290, 300, 310 may be configured to propagate light within each individual waveguide by total internal reflection (TIR). The waveguides 270, 280, 290, 300, 310 may each be planar, or have another shape (e.g., curved), with a major top surface and a bottom surface and an edge extending between their major top and bottom surfaces. In the illustrated configuration, the waveguides 270, 280, 290, 300, 310 each include external coupling optical elements 570, 580, 590, 600, 610 configured to extract light from the waveguide by redirecting the light propagating within each individual waveguide out of the waveguide and outputting the image information to the eye 210. The extracted light may also be referred to as external coupled light, and the external coupling optical elements may also be referred to as light extraction optical elements. The beam of extracted light may be output by the waveguide at the location where the light propagating within the waveguide impinges on the light extraction optical element. The external coupling optical elements 570, 580, 590, 600, 610 may be gratings, for example, including diffractive optical features as further discussed herein. For ease of explanation and clarity of the drawings, the external coupling optical elements 570, 580, 590, 600, 610 are shown disposed on the bottom major surfaces of the waveguides 270, 280, 290, 300, 310, but in some embodiments, the external coupling optical elements 570, 580, 590, 600, 610 may be disposed on the top and / or bottom major surfaces and / or directly within the volume of the waveguides 270, 280, 290, 300, 310, as further discussed herein. In some embodiments, the external coupling optical elements 570, 580, 590, 600, 610 may be attached to a transparent substrate and formed within a layer of the material forming the waveguides 270, 280, 290, 300, 310. In some other embodiments, the waveguides 270, 280, 290, 300, 310 may be a monolithic piece of material, and the external coupling optical elements 570, 580, 590, 600, 610 may be formed on and / or within the surface of the piece of material.

[0080] Continuing to refer to FIG. 6, as discussed herein, each of the waveguides 270, 280, 290, 300, 310 is configured to output light and form an image corresponding to a particular depth plane. For example, the waveguide 270 closest to the eye may be configured to deliver collimated light (input into such waveguide 270) to the eye 210. The collimated light may represent an optically infinite focal plane. The next upper waveguide 280 may be configured to emit collimated light that passes through a first lens 350 (e.g., a negative lens) before reaching the eye 210. Such a first lens 350 may be configured to generate a slight convex wavefront curvature so that the eye / brain interprets the light arising from the next upper waveguide 280 as arising from a first focal plane that is closer inwardly from the optically infinite towards the eye 210. Similarly, the third upper waveguide 290 passes its output light through both the first lens 350 and the second lens 340 before reaching the eye 210. The combined refractive power of the first lens 350 and the second lens 340 may be configured to generate another incremental amount of wavefront curvature so that the eye / brain interprets the light arising from the third waveguide 290 as arising from a second focal plane that is even closer inwardly from the optically infinite towards the person than the light from the next upper waveguide 280 was.

[0081] The other waveguide layers 300, 310 and lenses 330, 320 are similarly configured, and the top waveguide 310 in the stack sends its output through all of the lenses between it and the eye for the converging focusing power that represents the focal plane closest to the person. When viewing / interpreting light originating from the world 510 on the other side of the stacked waveguide assembly 260, a compensation lens layer 620 may be disposed on top of the stack to compensate for the stack of lenses 320, 330, 340, 350. Such a configuration provides the same number of perceived focal planes as there are available waveguide / lens pairs. Both the external coupling optical elements of the waveguides and the focusing sides of the lenses may be static (i.e., not dynamic or electroactive). In some alternative embodiments, one or both may be dynamic using electroactive features.

[0082] In some embodiments, two or more of the waveguides 270, 280, 290, 300, 310 may have the same associated depth plane. For example, a plurality of waveguides 270, 280, 290, 300, 310 may be configured to output images set to the same depth plane, or a plurality of subsets of waveguides 270, 280, 290, 300, 310 may be configured to output images set to the same plurality of depth planes, with one set per depth plane. This can provide the advantage of forming tiled images to provide an extended field of view at those depth planes.

[0083] Continuing to refer to FIG. 6, the external coupling optical elements 570, 580, 590, 600, 610 may be configured to redirect light from their respective waveguides for a particular depth plane associated with the waveguide and output the light with an appropriate amount of divergence or collimation. As a result, waveguides having different associated depth planes may have different configurations of the external coupling optical elements 570, 580, 590, 600, 610, which output light with different amounts of divergence depending on the associated depth plane. In some embodiments, the light extraction optical elements 570, 580, 590, 600, 610 may be three-dimensional or surface features configured to output light at a specific angle. For example, the light extraction optical elements 570, 580, 590, 600, 610 may be volume holograms, surface holograms, and / or diffraction gratings. In some embodiments, features 320, 330, 340, 350 may not be lenses. Rather, they may simply be spacers (e.g., cladding layers and / or structures for forming voids).

[0084] In some embodiments, the external coupling optical elements 570, 580, 590, 600, 610 are diffraction features or “diffractive optical elements” (also referred to herein as “DOEs”) that form a diffraction pattern. Preferably, the DOE has a sufficiently low diffraction efficiency such that only a portion of the light of the beam is deflected towards the eye 210 at each intersection of the DOE, while the remainder continues to travel through the waveguide via TIR. The light carrying the image information is thus split into several associated output beams that exit the waveguide at various locations, resulting in a very uniform pattern of output emission towards the eye 210 with respect to this particular collimated beam that bounces within the waveguide.

[0085] In some embodiments, one or more DOEs may be switchable between an “on” state where they actively diffract and an “off” state where they do not significantly diffract. For example, a switchable DOE may comprise a layer of polymer-dispersed liquid crystal in which microdroplets have a diffraction pattern in a host medium, and the refractive index of the microdroplets may be switched to substantially match the refractive index of the host material (in which case the pattern does not significantly diffract incident light), or the microdroplets may be switched to a refractive index that does not match that of the host medium (in which case the pattern actively diffracts incident light).

[0086] In some embodiments, a camera assembly 630 (e.g., a digital camera including visible and infrared light cameras) may be provided to capture an image of the eye 210 and / or the tissue surrounding the eye 210, and, for example, detect user input and / or monitor the user's physiological state. As used herein, a camera may be any image capture device. In some embodiments, the camera assembly 630 may include an image capture device and a light source (e.g., infrared light) that projects light onto the eye and then the light can be reflected by the eye and detected by the image capture device. In some embodiments, the camera assembly 630 may be attached to a frame or support structure 80 (FIG. 9E) and may communicate electrically with a processing module 140 and / or 150 that can process image information from the camera assembly 630. In some embodiments, one camera assembly 630 may be utilized per eye to monitor each eye separately.

[0087] In some embodiments, the camera assembly 630 may observe user movement such as eye movement of the user. As an example, the camera assembly 630 may capture an image of the eye 210 and determine the size, position, and / or orientation of the pupil of the eye 210 (or some other structure of the eye 210). The camera assembly 630 may, if desired, acquire an image (processed by a processing circuitry network of the type described herein) that is used to determine the direction in which the user is looking (e.g., eye pose or line of sight direction). In some embodiments, the camera assembly 630 may include multiple cameras, at least one of which is utilized per eye and may independently determine the eye pose or line of sight direction of each eye separately. In some embodiments, the camera assembly 630, in combination with a processing circuitry such as the controller 560 or the local data processing module 140, may determine the eye pose or line of sight direction based on a flash (e.g., reflection) of light (e.g., infrared light) reflected from a light source included within the camera assembly 630.

[0088] Referring now to FIG. 7, an example of an output beam output by a waveguide is shown. Although one waveguide is illustrated, it should be understood that other waveguides within the waveguide assembly 260 (FIG. 6) may function similarly, and the waveguide assembly 260 includes a plurality of waveguides. Light 640 is input into the waveguide 270 at the input surface 460 of the waveguide 270 and propagates within the waveguide 270 by TIR. At the point where the light 640 impinges on the DOE 570, a portion of the light exits the waveguide as the output beam 650. The output beam 650 is illustrated as being substantially parallel, but as discussed herein, and depending on the depth plane associated with the waveguide 270, it may be redirected to propagate at an angle (e.g., to form a diverging output beam) to the eye 210. It should be understood that a substantially parallel output beam may represent a waveguide with an external coupling optical element that externally couples the light to form an image that appears to be set in a depth plane at a long distance (e.g., optical infinity) from the eye 210. Other waveguides or other sets of external coupling optical elements may output a more divergent output beam pattern, which would require the eye 210 to focus at a closer distance and would be interpreted by the brain as light from a distance closer to the eye 210 than optical infinity.

[0089] In some embodiments, a full-color image may be formed on each depth plane by overlaying the image on each of the primary colors, e.g., three or more primary colors. FIG. 8 illustrates an example of a stacked waveguide assembly, and each depth plane includes an image formed using a plurality of different primary colors. The illustrated embodiment shows depth planes 240a-240f, although more or fewer depths may also be considered. Each depth plane may have three or more primary color images associated therewith, including a first image of a first color G, a second image of a second color R, and a third image of a third color B. Different depth planes are shown in the figure by different numbers related to the diopter (dpt) following the letters G, R, and B. As a mere example, the numbers following each of these letters indicate the diopter (1 / m), i.e., the inverse distance of the depth plane from the viewer, and each box in the figure represents an individual primary color image. In some embodiments, to account for differences in the focusing of light of different wavelengths by the eye, the exact location of the depth planes for different primary colors may vary. For example, the different primary color images for a given depth plane may be placed on depth planes corresponding to different distances from the user. Such an arrangement may increase visual acuity and user comfort and / or reduce chromatic aberration.

[0090] In some embodiments, the light of each primary color may be output by a single dedicated waveguide, and as a result, each depth plane may have a plurality of waveguides associated therewith. In such embodiments, each box in the figure, including those containing the letters G, R, or B, may be understood to represent an individual waveguide, and three waveguides may be provided for each depth plane, with three primary color images being provided for each depth plane. The waveguides associated with each depth plane are shown adjacent to each other in this drawing for ease of explanation, but it should be understood that in a physical device, all of the waveguides may be arranged in a stack with one waveguide per level. In some other embodiments, a plurality of primary colors may be output by the same waveguide, e.g., such that only a single waveguide is provided for each depth plane.

[0091] Continuing to refer to FIG. 8, in some embodiments, G is green, R is red, and B is blue. In some other embodiments, other colors associated with other wavelengths of light, including magenta and cyan, may also be used in addition to or in place of one or more of red, green, or blue.

[0092] It should be understood that references throughout this disclosure to the color of a given light include light of one or more wavelengths within the range of wavelengths of that given color as perceived by a viewer. For example, red light may include light of one or more wavelengths within the range of about 620 - 780 nm, green light may include light of one or more wavelengths within the range of about 492 - 577 nm, and blue light may include light of one or more wavelengths within the range of about 435 - 493 nm.

[0093] In some embodiments, the light source 530 (FIG. 6) may be configured to emit light of one or more wavelengths outside the viewer's visual perception range, such as infrared and / or ultraviolet wavelengths. Additionally, the internal coupling, external coupling, and other light redirection structures of the waveguide of the display 250 may be configured to direct and emit this light from the display towards the user's eye 210, for example, for imaging and / or user stimulation applications.

[0094] Referring now to FIG. 9A, in some embodiments, light that impinges on a waveguide may need to be redirected to internally couple that light into the waveguide. An internal coupling optical element may be used to redirect and internally couple the light into its corresponding waveguide. FIG. 9A illustrates a cross-sectional side view of an example of a stack 660 of a plurality or set of waveguides, each including an internal coupling optical element. Each waveguide may be configured to output light of one or more different wavelengths or one or more different wavelength ranges. Stack 660 may correspond to stack 260 (FIG. 6), and the illustrated waveguides of stack 660 may correspond to a portion of the plurality of waveguides 270, 280, 290, 300, 310, but it should be understood that light from one or more of the image input devices 360, 370, 380, 390, 400 is input into the waveguide from a position that requires the light to be redirected for internal coupling.

[0095] The illustrated set 660 of stacked waveguides includes waveguides 670, 680, and 690. Each waveguide includes an associated internal coupling optical element (which may also be referred to as an optical input area on the waveguide). For example, internal coupling optical element 700 is disposed on a major surface (e.g., the upper major surface) of waveguide 670, internal coupling optical element 710 is disposed on a major surface (e.g., the upper major surface) of waveguide 680, and internal coupling optical element 720 is disposed on a major surface (e.g., the upper major surface) of waveguide 690. In some embodiments, one or more of internal coupling optical elements 700, 710, 720 may be disposed on the bottom major surface of the respective waveguides 670, 680, 690 (in particular, one or more of the internal coupling optical elements are reflective deflecting optical elements). As illustrated, internal coupling optical elements 700, 710, 720 may be disposed on the upper major surface (or the upper portion of the next lower waveguide) of their respective waveguides 670, 680, 690, and in particular, those internal coupling optical elements are transmissive deflecting optical elements. In some embodiments, internal coupling optical elements 700, 710, 720 may be disposed within the body of the respective waveguides 670, 680, 690. In some embodiments, as discussed herein, internal coupling optical elements 700, 710, 720 are wavelength selective such that they selectively redirect one or more wavelengths of light while transmitting other wavelengths of light. Although illustrated on one side or corner of their respective waveguides 670, 680, 690, it should be understood that in some embodiments, internal coupling optical elements 700, 710, 720 may be disposed within other areas of their respective waveguides 670, 680, 690.

[0096] As shown, the internally coupled optical elements 700, 710, 720 may be laterally offset from each other in the direction of the light propagating through these internally coupled optical elements, as seen in the illustrated front view. In some embodiments, each internally coupled optical element may be offset so that it receives light without its light passing through another internally coupled optical element. For example, each of the internally coupled optical elements 700, 710, 720 may be configured to receive light from different image input devices 360, 370, 380, 390, and 400, as shown in FIG. 6, and may be separated (e.g., laterally spaced) from the other internally coupled optical elements 700, 710, 720 so as to substantially not receive light from the other internally coupled optical elements 700, 710, 720.

[0097] Each waveguide also includes an associated light dispersing element. For example, the light dispersing element 730 is disposed on a major surface (e.g., the upper major surface) of the waveguide 670, the light dispersing element 740 is disposed on a major surface (e.g., the upper major surface) of the waveguide 680, and the light dispersing element 750 is disposed on a major surface (e.g., the upper major surface) of the waveguide 690. In some other embodiments, the light dispersing elements 730, 740, 750 may each be disposed on the bottom major surface of the associated waveguides 670, 680, 690. In some other embodiments, the light dispersing elements 730, 740, 750 may each be disposed on both the upper and bottom major surfaces of the associated waveguides 670, 680, 690, or the light dispersing elements 730, 740, 750 may each be disposed on different ones of the upper and bottom major surfaces within different associated waveguides 670, 680, 690.

[0098] Waveguides 670, 680, 690 may be separated and isolated, for example, by a gas, liquid, and / or solid layer of material. For example, as shown, layer 760a may separate waveguides 670 and 680, and layer 760b may separate waveguides 680 and 690. In some embodiments, layers 760a and 760b are formed from a low refractive index material (i.e., a material having a refractive index lower than the material forming the immediate waveguides 670, 680, 690). Preferably, the refractive index of the material forming layers 760a, 760b is less than the refractive index of the material forming waveguides 670, 680, 690 by 0.05 or 0.10. Advantageously, the lower refractive index layers 760a, 760b may function as cladding layers that facilitate total internal reflection (TIR) of light (e.g., TIR between the upper and bottom major surfaces of each waveguide) through waveguides 670, 680, 690. In some embodiments, layers 760a, 760b are formed from air. It should be understood that although not shown, the top and bottom of the illustrated set 660 of waveguides may include an immediate cladding layer.

[0099] Preferably, for ease of manufacture and other considerations, the materials forming waveguides 670, 680, 690 are similar or identical, and the materials forming layers 760a, 760b are similar or identical. In some embodiments, the materials forming waveguides 670, 680, 690 may be different between one or more waveguides, and / or the materials forming layers 760a, 760b may still be different while maintaining the various refractive index relationships described above.

[0100] Continuing to refer to FIG. 9A, light rays 770, 780, 790 are incident on the set 660 of waveguides. It should be understood that light rays 770, 780, 790 may be introduced into waveguides 670, 680, 690 by one or more image input devices 360, 370, 380, 390, 400 (FIG. 6).

[0101] In some embodiments, the light rays 770, 780, 790 may have different properties, such as different wavelengths or different wavelength ranges, corresponding to different colors. The internal coupling optical elements 700, 710, 720 each deflect the incident light so that light propagates through an individual one of the waveguides 670, 680, 690 by TIR. In some embodiments, the internal coupling optical elements 700, 710, 720 each selectively deflect one or more specific wavelengths of light while transmitting other wavelengths to the underlying waveguide and the associated internal coupling optical element.

[0102] For example, the internal coupling optical element 700 may be configured to deflect the light ray 770 having the first wavelength or wavelength range while transmitting the light rays 780 and 790 having different second and third wavelengths or wavelength ranges, respectively. The transmitted light ray 780 impinges on the internal coupling optical element 710 configured to deflect the light of the second wavelength or wavelength range, and is thereby deflected. The light ray 790 is deflected by the internal coupling optical element 720 configured to selectively deflect the light of the third wavelength or wavelength range.

[0103] Continuing to refer to FIG. 9A, the deflected light rays 770, 780, 790 are deflected to propagate through the corresponding waveguides 670, 680, 690. That is, the internal coupling optical elements 700, 710, 720 of each waveguide deflect the light into its corresponding waveguide 670, 680, 690 and internally couple the light into the corresponding waveguide. The light rays 770, 780, 790 are deflected at an angle that causes the light to propagate through the individual waveguides 670, 680, 690 by TIR. The light rays 770, 780, 790 propagate through the individual waveguides 670, 680, 690 by TIR until they impinge on the corresponding light dispersion elements 730, 740, 750 of the waveguide.

[0104] Referring now to FIG. 9B, a perspective view of an embodiment of the plurality of stacked waveguides of FIG. 9A is illustrated. As described above, the internally coupled light rays 770, 780, 790 are each deflected by the internally coupled optical elements 700, 710, 720 and then each propagate by TIR within waveguides 670, 680, 690. The light rays 770, 780, 790 then each impinge on the light dispersing elements 730, 740, 750. The light dispersing elements 730, 740, 750 each deflect the light rays 770, 780, 790 so as to propagate towards the externally coupled optical elements 800, 810, 820.

[0105] In some embodiments, the light dispersing elements 730, 740, 750 are orthogonal pupil expanders (OPEs). In some embodiments, the OPEs deflect or disperse light to the external coupling optical elements 800, 810, 820, and in some embodiments, also increase the beam or spot size of the present light as it propagates to the external coupling optical elements. In some embodiments, the light dispersing elements 730, 740, 750 may be omitted, and the internal coupling optical elements 700, 710, 720 may be configured to deflect light directly to the external coupling optical elements 800, 810, 820. For example, referring to FIG. 9A, the light dispersing elements 730, 740, 750 may each be replaced with the external coupling optical elements 800, 810, 820. In some embodiments, the external coupling optical elements 800, 810, 820 are an exit pupil (EP) or an exit pupil expander (EPE) that directs light to the viewer's eye 210 (FIG. 7). It should be understood that the OPE may be configured to increase the size of the eyebox in at least one axis, and the EPE may increase the eyebox in an axis that intersects, for example, is orthogonal to, the axis of the OPE. For example, each OPE may be configured to redirect a portion of the light impinging on the OPE to the EPE of the same waveguide while allowing the remainder of the light to continue to propagate along the waveguide. In response to a collision with the OPE, again, another portion of the remaining light is redirected to the EPE, and the remainder of that portion continues to propagate further along the waveguide or the like. Similarly, in response to a collision with the EPE, a portion of the colliding light is directed from the waveguide towards the user, and the remainder of that light continues to propagate through the waveguide until it impinges on the EP again, at which point another portion of the colliding light is directed from the waveguide, etc. As a result, a single beam of internally coupled light is "replicated" each time a portion of that light is redirected by the OPE or EPE, thereby forming a beam field of cloned light as shown in FIG. 6. In some embodiments, the OPE and / or EPE may be configured to modify the size of the beam of light.

[0106] Thus, referring to FIGS. 9A and 9B, in some embodiments, the waveguide set 660 includes, for each primary color, waveguides 670, 680, 690, internal coupling optical elements 700, 710, 720, light dispersion elements (e.g., OPE) 730, 740, 750, and external coupling optical elements (e.g., EP) 800, 810, 820. The waveguides 670, 680, 690 may be stacked with a gap / cladding layer between each one. The internal coupling optical elements 700, 710, 720 redirect or deflect the incident light into their respective waveguides (using different internal coupling optical elements that receive light of different wavelengths). The light then propagates at an angle that will result in TIR within the individual waveguides 670, 680, 690. In the illustrated example, the light ray 770 (e.g., blue light) is polarized by the first internal coupling optical element 700 in the manner described above and then continues to bounce through the waveguide, interacting with the light dispersion element (e.g., OPE) 730 and then the external coupling optical element (e.g., EP) 800. The light rays 780 and 790 (e.g., green and red light, respectively) pass through the waveguide 670, and the light ray 780 collides with the internal coupling optical element 710 and is thereby deflected. The light ray 780 then bounces through the waveguide 680 via TIR and will proceed to its light dispersion element (e.g., OPE) 740 and then the external coupling optical element (e.g., EP) 810. Finally, the light ray 790 (e.g., red light) passes through the waveguide 690 and collides with the light internal coupling optical element 720 of the waveguide 690. The light internal coupling optical element 720 deflects the light ray 790 such that the light ray propagates by TIR to the light dispersion element (e.g., OPE) 750 and then by TIR to the external coupling optical element (e.g., EP) 820. The external coupling optical element 820 then finally externally couples the light ray 790 to the viewer, who also receives the externally coupled light from the other waveguides 670, 680.

[0107] FIG. 9C illustrates a top and bottom plan view of an embodiment of a plurality of stacked waveguides of FIGS. 9A and 9B. This top and bottom view can also be referred to as a front-on view, as seen in the direction of propagation of light towards the internal coupling optical elements 800, 810, 820, i.e., it should be understood that the top and bottom view is a view of the waveguides where the image light is incident normal to the page. As shown, waveguides 670, 680, 690 may be vertically aligned, along with their associated light dispersion elements 730, 740, 750 and associated external coupling optical elements 800, 810, 820. However, as discussed herein, the internal coupling optical elements 700, 710, 720 are not vertically aligned. Rather, the internal coupling optical elements are preferably non-overlapping (e.g., laterally spaced as seen in the top and bottom view). As further discussed herein, this non-overlapping spatial arrangement facilitates the injection of light from different sources into different waveguides on a one-to-one basis, thereby enabling a specific light source to be uniquely coupled to a specific waveguide. In some embodiments, an arrangement including non-overlapping spatially separated internal coupling optical elements may be referred to as a shifted pupil system, and the internal coupling optical elements within these arrangements may correspond to sub-pupils.

[0108] It should be understood that the spatially overlapping area may have a lateral overlap of 70% or more, 80% or more, or 90% or more of that area, as seen in the top and bottom view. On the other hand, the laterally shifted area has an overlap of less than 30%, less than 20%, or less than 10% of that area, as seen in the top and bottom view. In some embodiments, the laterally shifted area has no overlap.

[0109] FIG. 9D illustrates a top and bottom plan view of another embodiment of a plurality of stacked waveguides. As shown, waveguides 670, 680, 690 may be vertically aligned. However, compared to the configuration of FIG. 9C, separate optical dispersion elements 730, 740, 750 and associated external coupling optical elements 800, 810, 820 are omitted. Instead, the optical dispersion elements and the external coupling optical elements are, in effect, superimposed and occupy the same area as seen in the top and bottom views. In some embodiments, an optical dispersion element (e.g., OPE) may be disposed on one major surface of waveguides 670, 680, 690, and an external coupling optical element (e.g., EPE) may be disposed on the other major surface of those waveguides. Thus, each of waveguides 670, 680, 690 may collectively have superimposed optical dispersion and external coupling optical elements, each referred to as combined OPE / EPE 1281, 1282, 1283, respectively. Further details regarding such combined OPE / EPE may be found in U.S. Patent Application No. 16 / 221,359, filed on December 14, 2018, the entire disclosure of which is incorporated herein by reference. Internal coupling optical elements 700, 710, 720 internally couple light and direct it to combined OPE / EPE 1281, 1282, 1283, respectively. In some embodiments, as shown, internal coupling optical elements 700, 710, 720 may be laterally offset if they have an offset pupil spatial arrangement (e.g., they are laterally separated as seen in the top and bottom views shown). Similar to the configuration of FIG. 9C, this laterally offset spatial arrangement facilitates the injection of light of different wavelengths into different waveguides on a one-to-one basis (e.g., from different light sources).

[0110] FIG. 9E illustrates an embodiment of a wearable display system 60 into which various waveguides and related systems disclosed herein may be integrated. In some embodiments, display system 60 is the system 250 of FIG. 6, which schematically shows some parts of that system 60 in more detail. For example, the waveguide assembly 260 of FIG. 6 may be part of a display 70.

[0111] Continuing to refer to FIG. 9E, the display system 60 includes a display 70 and various mechanical and electronic modules and systems for supporting the functions of the display 70. The display 70 may be coupled to a frame 80, which is wearable by a display system user or viewer 90 and is configured to position the display 70 in front of the user 90's eyes. In some embodiments, the display 70 may be regarded as eyewear. The display 70 may include one or more waveguides, such as waveguide 270, configured to relay internally coupled image light and output the image light to the eyes of the user 90. In some embodiments, a speaker 100 is coupled to the frame 80 and configured to be positioned adjacent to the outer ear canal of the user 90 (in some embodiments, another speaker, not shown, may also be optionally positioned adjacent to the other outer ear canal of the user to provide stereo / formable sound control). The display system 60 may also include one or more microphones 110 or other devices and may detect sound. In some embodiments, the microphone is configured to enable the user to provide an input or command to the system 60 (e.g., selection of a voice menu command, natural language question, etc.) and / or enable audio communication with other persons (e.g., other users of a similar display system). The microphone may further be configured as a peripheral sensor and may collect audio data (e.g., sound from the user and / or the environment). In some embodiments, the display system 60 may further include one or more outwardly directed environmental sensors 112 configured to detect objects, stimuli, people, animals, locations, or other aspects of the world around the user. For example, the environmental sensor 112 may include one or more cameras, which may be positioned outwardly facing, for example, to capture an image similar to at least a portion of the normal field of view of the user 90.In some embodiments, the display system may also include a peripheral sensor 120a, which is separate from the frame 80 and may be attached on the body of the user 90 (e.g., the head, torso, limbs, etc. of the user 90). In some embodiments, the peripheral sensor 120a may be configured to obtain data characterizing the physiological state of the user 90. For example, the sensor 120a may be an electrode.

[0112] Continuing to refer to FIG. 9E, the display 70 is operably coupled to the local data processing module 140 by a communication link 130 such as a wired conductor or wireless connectivity, which is fixedly attached to the frame 80, fixedly attached to a helmet or hat worn by the user, built into headphones, or otherwise removably attached to the user 90 (e.g., in a backpack configuration, in a belt attachment configuration), etc., and may be mounted in various configurations. Similarly, the sensor 120a may be operably coupled to the local data processing module 140 by a communication link 120b, e.g., a wired conductor or wireless connectivity. The local processing and data module 140 may comprise digital memory such as a hardware processor and non-volatile memory (e.g., flash memory or hard disk drive), both of which may be utilized to assist in the processing, caching, and storage of data. Optionally, the local processing and data module 140 may include one or more central processing units (CPUs), graphics processing units (GPUs), dedicated processing hardware, etc. The data may include a) data captured from sensors (image capture devices (such as cameras), microphones, inertial measurement units, accelerometers, compasses, GPS units, wireless devices, gyroscopes, and / or other sensors disclosed herein (e.g., operably coupled to the frame 80 or otherwise attachable to the user 90)), and / or b) optionally, data obtained and / or processed using the remote processing module 150 and / or the remote data repository 160 (including data related to virtual content) for passage to the display 70 after processing or retrieval. The local processing and data module 140 may be operably coupled to the remote processing module 150 and the remote data repository 160 by communication links 170, 180 via a wired or wireless communication link, etc., such that these remote modules 150, 160 are operably coupled to each other and available as resources to the local processing and data module 140.In some embodiments, the local processing and data module 140 may include one or more than one of an image capture device, a microphone, an inertial measurement unit, an accelerometer, a compass, a GPS unit, a wireless device, and / or a gyroscope. In some other embodiments, one or more than one of these sensors may be attached to the frame 80 or may be an independent structure that communicates with the local processing and data module 140 via a wired or wireless communication path.

[0113] Continuing to refer to FIG. 9E, in some embodiments, the remote processing module 150 may include one or more processors configured to analyze and process data and / or image information, for example, including one or more central processing units (CPUs), graphics processing units (GPUs), dedicated processing hardware, etc. In some embodiments, the remote data repository 160 may include a digital data storage facility, which may be available through the Internet or other networking configurations in a “cloud” resource configuration. In some embodiments, the remote data repository 160 may include one or more remote servers that provide information, for example, information for generating virtual content, to the local processing and data module 140 and / or the remote processing module 150. In some embodiments, all data is stored and all calculations are performed in the local processing and data module, enabling fully autonomous use from the remote module. Optionally, an external system (e.g., one or more processors, one or more computer systems) including a CPU, GPU, etc. may perform at least a portion of the processing (e.g., generating image information, processing data), for example, providing information to and receiving information from modules 140, 150, 160 via a wireless or wired connection. B. Exemplary Operations and Interfaces of the Wearable Display System

[0114] The wearable system may employ various mapping-related techniques to provide content for the rendered light field. When mapping the virtual world, it is advantageous to identify reference features and points within the real world and accurately depict virtual objects in relation to the real world. To achieve this purpose, the FOV images captured from the user of the wearable system may be added to the world model by including new photographs that convey information about various points and features of the real world. For example, the wearable system may collect a set of map points (such as 2D points or 3D points), find new map points, and render a more accurate version of the world model. The world model of the first user may be communicated to the second user (e.g., via a network such as a cloud network) so that the second user can experience the world surrounding the first user.

[0115] FIG. 10 is a block diagram of an embodiment of an MR system 1000 that includes various features related to the operation of a wearable display system. The MR system 1000 may be configured to receive inputs (e.g., visual input 1002 from one or more wearable systems 1020, stationary input 1004 such as an indoor camera, sensory input 1006 from various sensors, gestures, totems, eye tracking, user input, etc. from a user input device 466) from one or more user wearable systems (such as wearable system 60 or display system 220) or stationary indoor systems (such as an indoor camera). The wearable system may use various sensors (e.g., accelerometers, gyroscopes, temperature sensors, motion sensors, depth sensors, GPS sensors, inward-facing imaging systems, outward-facing imaging systems, etc.) to determine the location and various other attributes of the user's environment. This information may be further supplemented with information from stationary cameras within the room that may provide images or various cues from different viewpoints. The image data acquired by the camera (such as an indoor camera or the camera of an outward-facing imaging system) may be reduced to a set of mapping points.

[0116] One or more object recognition devices 1008 may crawl through the received data (e.g., a set of points), recognize or map the points, tag the images, and associate semantic information with the objects using the map database 1010. The map database 1010 may comprise various points collected over time and their corresponding objects. The various devices and the map database may be interconnected through a network (e.g., LAN, WAN, etc.) as interconnected remote databases stored, for example, "in the cloud".

[0117] Based on the present information and the set of points in the map database, the object recognition devices 1008a - 1008n may recognize the objects in the environment. For example, the object recognition device may recognize a face, a person, a window, a wall, a user input device, a television, or other objects in the user's environment. One or more object recognition devices may be specialized for objects with a certain characteristic. For example, the object recognition device 1008a may be used to recognize a face, while another object recognition device may be used to recognize a totem.

[0118] Object recognition may be performed using various computer vision techniques. For example, a wearable system may analyze an image obtained by an outward-facing imaging system 112 (shown in FIG. 9E) and perform scene reconstruction, event detection, video tracking, object recognition, object pose estimation, learning, indexing, motion estimation, or image restoration, etc. One or more computer vision algorithms may be used to perform these tasks. Non-limiting examples of computer vision algorithms include Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), Oriented FAST and Rotated BRIEF (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retina Keypoints (FREAK), Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, Visual Simultaneous Localization and Mapping (vSLAM) techniques, sequential Bayesian estimators (e.g., Kalman filter, Extended Kalman filter, etc.), bundle adjustment, adaptive thresholding (and other thresholding techniques), Iterative Closest Point (ICP), Semi-Global Matching (SGM), Semi-Global Block Matching (SGBM), feature point histogram, various machine learning algorithms (e.g., support vector machine, k-nearest neighbor algorithm, naive Bayes, neural network, etc. (including convolutional or deep neural networks), or other supervised / unsupervised models, etc.).

[0119] Object recognition may additionally or alternatively be performed by various machine learning algorithms. Once trained, the machine learning algorithms may, in some embodiments, be stored on the wearable display system. Some examples of machine learning algorithms may include supervised or unsupervised machine learning algorithms, regression algorithms (e.g., ordinary least squares regression, etc.), instance-based algorithms (e.g., learning vector quantization, etc.), decision tree algorithms (e.g., classification and regression trees, etc.), Bayesian algorithms (e.g., naive Bayes, etc.), clustering algorithms (e.g., k-means clustering, etc.), association rule learning algorithms (e.g., Apriori algorithm, etc.), artificial neural network algorithms (e.g., Perceptron, etc.), deep learning algorithms (e.g., Deep Boltzmann Machine, i.e., deep neural network, etc.), dimensionality reduction algorithms (e.g., principal component analysis, etc.), ensemble algorithms (e.g., Stacked Generalization, etc.), and / or other machine learning algorithms. In some embodiments, individual models may be customized for individual datasets. For example, a wearable display device may generate or store a base model. The base model is used as a starting point and may generate additional models specific to a data type (e.g., a particular user within a telepresence session), a dataset (e.g., a set of acquired additional images of a user within a telepresence session), a conditional situation, or other variations. In some embodiments, the wearable HMD may be configured to generate a model for the analysis of aggregated data using multiple techniques. Other techniques may include using predefined thresholds or data values.

[0120] Based on the book information and the set of points in the map database, the object recognition devices 1008a - 1008n may recognize an object, complement the object with semantic information, and assign a life to the object. For example, when the object recognition device recognizes that a set of points is a door, the system may associate some semantic information (e.g., the door has hinges and has a 90 - degree movement around the hinges). When the object recognition device recognizes that a set of points is a mirror, the system may associate semantic information that the mirror has a reflective surface that can reflect images of objects in the room. Over time, the map database grows as the system (which may be resident locally or accessible through a wireless network) accumulates more data from the world. Once an object is recognized, the information may be transmitted to one or more wearable systems. For example, the MR system 1000 may include information about a scene generated in California. The system 1000 may be transmitted to one or more users in New York. Based on data received from the FOV camera and other inputs, the object recognition device and other software components may map points collected from various images, recognize objects, etc., so that the scene can be accurately "passed" to a second user who may be in a different part of the world. The environment 1000 may also use a topological map for location - specific purposes.

[0121] Figure 11A is a process flow diagram of an embodiment of a method 800 for rendering virtual content in relation to a recognized object. The method 800 describes a way in which a virtual scene can be presented to a user of a wearable system. The user may be geographically remote from the scene. For example, the user may be in New York but may desire to view a scene currently taking place in California, or may desire to go for a walk with a friend in California.

[0122] In block 1102, the wearable system may receive input regarding the user's environment from the user and other users. This may be accomplished through various input devices and knowledge already held within the map database. The user's FOV camera, sensors, GPS, eye tracking, etc. communicate information to the system in block 1102. The system may determine sparse points based on this information in block 1104. The sparse points may be used in determining pose data (e.g., head pose, eye pose, body pose, or hand gesture) that can be used to display and understand the orientation and position of various objects in the user's surroundings. The object recognition devices 708a - 708n may crawl through these collected points in block 1106 and use the map database to recognize one or more objects. This information may then be communicated to the user's individual wearable system in block 1108, and a desired virtual scene may be appropriately displayed to the user in block 1110. For example, a desired virtual scene (e.g., for a user in CA) may be displayed in appropriate orientation, position, etc. in relation to various objects and other surroundings of a user in New York.

[0123] Figure 11B is a block diagram of another embodiment of a wearable system (which may correspond to the display system 60 of FIG. 9E). In this embodiment, the wearable system 1101 comprises a map, which may include map data regarding the world. The map may be resident locally on the wearable system in part, and may be resident in part in a networked storage location (e.g., within a cloud system) accessible by a wired or wireless network. A pose process 1122 is executed on a wearable computing architecture (e.g., processing module 260 or controller 460) and may utilize data from the map to determine the position and orientation of the wearable computing hardware or the user. The pose data may be calculated from data collected on-the-fly as the user experiences the system and operates within its world. The data may comprise images regarding objects in the physical or virtual environment, data from sensors (e.g., an inertial measurement unit generally comprising accelerometer and gyroscope components), and surface information.

[0124] The sparse point representation may be the output of a simultaneous localization and mapping (SLAM or V-SLAM, referring to configurations where the input is image / vision only) process. The system may be configured to find not only the locations of various components within the world, but also what constitutes the world. Pose may be a building block that achieves many goals, including filling in the map and using data from the map.

[0125] In one embodiment, the sparse point positions may not be fully proper by themselves, and additional information may be required to generate a multi-focus AR, VR, or MR experience. Generally, a dense representation, which refers to depth map information, may be utilized to at least partially fill this gap. Such information may be calculated from a process referred to as stereo 1128, and the depth information is determined using techniques such as triangulation or time-of-flight sensing. Image information and an active pattern (such as an infrared pattern generated using an active projector) may serve as inputs to the stereo process 1128. A significant amount of depth map information may be fused together, and some of this may be summarized using surface representations. For example, a mathematically definable surface may be an efficient (e.g., compared to a large-scale point cloud) and summarizable input to other processing devices such as a game engine. Thus, the output of the stereo process (e.g., depth map) 1128 may be combined in the fusion process 1126. Pose may similarly be an input to this fusion process 1126, and the output of the fusion 1126 serves as an input to fill the map process 1124. Sub-surfaces may be interconnected in topographic mapping, etc., to form larger surfaces, and the map becomes a large-scale hybrid of points and surfaces.

[0126] To resolve various aspects in the composite reality process 1129, various inputs may be utilized. For example, in the embodiment depicted in FIG. 11B, game parameters may be inputs for determining that the user of the system is playing a monster battle game with one or more monsters in various locations, that the monsters are dead, that they are fleeing under various conditions (such as when the user attacks the monsters), walls or other objects in various locations, and the like. The world map may include information regarding where such objects exist relative to each other and serve as another useful input to the composite reality. The pose with respect to the world may similarly be an input and play an important role for almost any bidirectional system.

[0127] Control or input from the user is another input to the wearable system 1101. As described herein, user input may include visual input, gestures, totems, audio input, sensory input, and the like. For example, in order to move around or play a game, the user may need to instruct the wearable system 1101 as to what they want to do. There are various forms of user control that can be utilized, not only moving oneself within space. In one embodiment, a totem (e.g., a user input device), or an object such as a toy gun, may be held by the user and tracked by the system. The system will preferably be configured to sense that the user is holding the item and understand the type of interaction the user is having with the item (e.g., if the totem or object is a gun, the system may be equipped with sensors such as an IMU that can assist in determining what is happening, not only the location and orientation, but also whether the user is clicking a trigger or other sensing button or element, even when such activity is not within the field of view of any of the cameras).

[0128] Hand gesture tracking or recognition may also provide input information. The wearable system 1101 may be configured to track and interpret hand gestures for button presses, gestures such as left or right, stop, grip, hold, and the like. For example, in one configuration, the user may desire to flip through an email or calendar in a non-game environment, or perform a "fist bump" with another person or performer. The wearable system 1101 may be configured to utilize a minimal amount of hand gestures, which may or may not be dynamic. For example, the gestures may be simple static gestures such as spreading the hand to indicate stop, raising the thumb to indicate OK, lowering the thumb to indicate not OK, or flipping the hand left or right or up or down to indicate a directional command, etc.

[0129] Eye tracking is another input (e.g., tracking where the user is looking, controlling display technology, and rendering at a specific depth or range). In one embodiment, the convergence / divergence movement of the eyes may be determined using triangulation, and then the accommodation may be determined using a convergence / divergence / accommodation model developed for that particular person.

[0130] Regarding the camera system, the exemplary wearable system 1101 shown in FIG. 11B can include three pairs of cameras, namely, a relatively wide FOV or passive SLAM pair of cameras arranged on both sides of the user's face, and a different pair of cameras oriented in front of the user that handle the stereo imaging process 1128 and also capture hand gestures and the trajectory of totems / objects in front of the user's face. The FOV cameras and the pair of cameras for the stereo process 1128 may be part of an outward-facing imaging system 112 (shown in FIG. 9E). The wearable system 1101 can include an eye tracking camera (which may be part of the inward-facing imaging system 630 shown in FIG. 6) oriented towards the user's eyes to triangulate the eye vector and other information. The wearable system 1101 may also include one or more textured light projectors (such as an infrared (IR) projector), which may project texture into the scene.

[0131] FIG. 11C is a process flow diagram of an example of a method 1103 for determining user input to a wearable system. In this example, the user may interact with a totem. The user may have multiple totems. For example, the user may have one designated totem for a social media application, another totem for playing a game, etc. In block 1130, the wearable system may detect the movement of the totem. The movement of the totem may be recognized through an outward-facing system or detected through sensors (such as a touch glove, an image sensor, a hand tracking device, an eye tracking camera, a head pose sensor, etc.).

[0132] At least in part, based on the detected gesture, eye gesture, head gesture, or input through the totem, the wearable system detects, at block 1132, the position, orientation, and / or movement of the totem (or the user's eyes or head or gesture) relative to the reference frame. The reference frame may be a set of map points based on which the wearable system converts the movement of the totem (or the user) into an action or command. At block 1134, the user's interaction with the totem is mapped. Based on the mapping of the user interaction relative to the reference frame 1132, the system determines, at block 1136, the user input.

[0133] For example, the user may move a totem or physical object back and forth to scroll a virtual page, move to the next page, or move from one user interface (UI) display screen to another UI screen. As another example, the user may move their head or eyes to view different real or virtual objects within the user's FOR. If the user's fixation on a particular real or virtual object is longer than a threshold time, that real or virtual object may be selected as the user input. In some implementations, the user's eye convergence / divergence movement may be tracked, and a focus adjustment / convergence / divergence movement model may be used to determine the user's focus adjustment state with respect to the depth plane on which the user is in focus. In some implementations, the wearable system may use a ray casting technique to determine real or virtual objects along the direction of the user's head gesture or eye gesture. In various implementations, the ray casting technique may include projecting a thin beam of light with substantially little lateral width or projecting a beam of light with a substantial lateral width (e.g., a cone or frustum of a cone).

[0134] The user interface may be projected by a display system (such as display 220 in FIG. 2) as described herein. Also, it may be displayed using various other techniques such as one or more projectors. The projector may project the image onto a physical object such as a canvas or a sphere. Interaction with the user interface may be tracked using one or more cameras external to the system or part of the system (e.g., using an inward-facing imaging system 630 (FIG. 6) or an outward-facing imaging system 112 (FIG. 9E)).

[0135] FIG. 11D is a process flow diagram of an example of method 1105 for interacting with a virtual user interface. Method 1105 may be implemented by a wearable system as described herein, e.g., by display system 60 of FIG. 9E.

[0136] In block 1140, the wearable system may identify a particular UI. The type of UI may be provided by the user. The wearable system may identify that a particular UI needs to be captured based on user input (e.g., gesture, visual data, audio data, sensory data, direct command, etc.). In block 1142, the wearable system may generate data for the virtual UI. For example, data associated with the boundaries, general structure, shape, etc. of the UI may be generated. Additionally, the wearable system may determine map coordinates of the user's physical location so that the wearable system can display the UI in relation to the user's physical location. For example, if the UI is body-centered, the wearable system may determine the coordinates of the user's physical standing position, head pose, or eye pose so that a ring UI can be displayed around the user or a planar UI can be displayed on a wall or in front of the user. If the UI is hand-centered, the map coordinates of the user's hand may be determined. These map points may be derived through an FOV camera, data received through sensory input, or any other type of collected data.

[0137] In block 1144, the wearable system may send data from the cloud to the display, or the data may be sent from the local database to the display component. In block 1146, the UI is presented to the user based on the sent data. For example, a light field display may project a virtual UI into one or both of the user's eyes. Once the virtual UI is generated, the wearable system may, in block 1148, simply wait for commands from the user and generate more virtual content on the virtual UI. For example, the UI may be a body-centered ring around the user's body. The wearable system may then wait for commands (such as gestures, head or eye movements, input from a user input device, etc.), and if recognized (block 1150), the virtual content associated with the command may be presented to the user (block 1152). As an example, the wearable system may wait for a gesture of the user's hand before mixing multiple stem tracks.

[0138] Additional examples of wearable systems, UIs, and user experiences (UX) are described in U.S. Patent Publication No. 2015 / 0016777, which is incorporated herein by reference in its entirety. C. Examples of Eye Tracking Calibration

[0139] As described herein, a user may interact with a wearable display system using eye gaze, which may include the direction in which the user's eyes are directed. Eye gaze (sometimes also referred to herein as eye pose) may be measured from a base direction (typically, the forward direction in which the user's eyes are necessarily directed), and often is measured using two angles (e.g., elevation angle and azimuth angle relative to the base direction) or three angles (e.g., elevation angle, azimuth angle, and in addition, roll angle). To provide a realistic and intuitive interaction with objects in the user's environment using eye gaze, the wearable system may use eye tracking calibration to calibrate the wearable display system and incorporate the uniqueness of the user's eye characteristics and other conditions that may affect the eye measurements.

[0140] Eye tracking calibration involves a process that enables a computing device to learn how to associate the user's eye gaze (e.g., as identified within an eye image) with a line-of-sight point in 3D space. The eye gaze may be associated with a single point in 3D space. The eye gaze may also be associated with multiple points in 3D space (e.g., a series of points that describe the movement of a virtual avatar 140 as described above with reference to FIG. 1 or the movement of a virtual butterfly as described below with reference to FIG. 12B) that may describe the movement of a virtual object.

[0141] In some embodiments, a wearable system may determine a user's eye gaze based on an eye image. The wearable system may use sensors (e.g., an eye camera) within an inward-facing imaging system 630 (FIG. 6) to obtain an eye image. The wearable system may image one or both of the user's eyes while the user changes their eye gaze (e.g., when the user looks around, follows a calibration target that moves). To map the user's eye image and gaze points, the wearable system may present a virtual target for the user to look at. The virtual target may be associated with one or more known points on a line of sight within 3D space. While the user is looking at the target, the wearable system may obtain an eye image and associate the image with the gaze points. The wearable system may calculate a mapping matrix based on the association of the eye image and the gaze points associated with the target. The mapping matrix may provide an association between a measurement of the user's eye gaze and a gaze vector (which may indicate the user's line of sight direction).

[0142] The mapping matrix may be generated using various machine learning techniques, as described with reference to FIG. 11C. For example, components of the wearable system, such as the local processing and data module 140 and / or the remote processing module 150 (FIG. 9E), may receive the eye image and the target position as inputs and generate the mapping matrix as an output by analyzing the association of the eye image and the gaze points using machine learning techniques. Eye gaze calculation techniques that may be used include feature-based techniques that detect and locate image features (e.g., iris features or the shape of the pupil or edge boundary), or model-based approaches that do not explicitly identify features but rather calculate the best-fit eye model that matches the acquired eye image. Some techniques (e.g., starburst) are hybrid approaches that include aspects of both feature-based and model-based eye gaze techniques.

[0143] Once trained, the wearable system may apply a mapping matrix to determine the user's line of sight direction. For example, the wearable system may observe the eye gaze while the user interacts with a virtual object, input the eye gaze into the mapping matrix, and determine the user's line of sight point. The line of sight point can be used in ray casting to identify the object of interest that intersects with the user's line of sight direction. For example, the wearable system may project a light ray in the user's line of sight direction and identify and select the virtual object "struck" by the light ray. In some cases, the light ray may be a line with a negligible side width, while in other cases, the light ray may be a cone with a side width for a solid angle. The wearable system may thus enable the user to select or perform other user interface operations based on the determined object of interest.

[0144] The calibration result may reflect the uniqueness in each person's eyes. For example, the wearable system may generate a mapping matrix customized for one or both eyes of a specific individual. For example, the user may have different amounts of eye movement or eye gaze in response to a specific target. As a result, by generating a calibration result specific to each individual user, the wearable system may enable more accurate user interaction with the eye gaze.

[0145] FIG. 12A illustrates an exemplary target in an eye tracking calibration process. FIG. 12A illustrates nine virtual targets within a user's field of view (FOV) 1200. The user's FOV 1200 may include a portion of the user's field of regard (FOR) that the user can perceive at a given time. The nine targets 1202a-1202i may be rendered at different depths. For example, target 1202e is in a depth plane that appears closer to the user than target 1202a. As a result, target 1202e appears larger to the user than target 1202a. The nine targets may be sequentially rendered to the user during the eye tracking calibration process. For example, the wearable system may first render target 1202e, followed by target 1202c, then target 1202b, and so on. As further described below with reference to FIG. 12B, in some embodiments, a single target is displayed to the user and the target moves around the perimeter of the user's field of view (e.g., passes through or temporarily stops at positions 1202a-1202i during the movement of the target). The wearable system may obtain an image of the user's eyes while the user is looking at these targets. For example, the wearable system may obtain a first image while the user is looking at target 1202e, while also obtaining a second image while the user is looking at target 1202c and a third image while the user is looking at target 1202b, and so on. The wearable system may thus align the first image with the position of target 1202e, the second image with the position of target 1202c, the third image with the position of target 1202b, and so on. The nine targets are shown in FIG. 12A, but this is for illustration purposes and in other implementations, fewer or more targets (or target locations) may be used and their positions may be different from those shown.

[0146] The location of the target can be represented by a position within the rig space. The rig space may include a coordinate system that is fixed with respect to a wearable display system (e.g., the HMD described herein). The coordinate system may be represented as a Cartesian x-y-z coordinate system. In this embodiment, the horizontal axis (x) is represented by axis 1204 (also referred to as the azimuth angle), and the vertical axis (y) is represented by axis 1208 (also referred to as the elevation angle). The axis (z) associated with the depth from the user is not shown in FIG. 12A.

[0147] As shown, target 1202e is at the center of the nine virtual targets. Thus, the x-axis position of target 1202e may be calculated as 0.5 times the sum of the x-axis 1204 values of the leftmost virtual objects (e.g., objects 1202a, 1202d, 1202g) and the rightmost virtual objects (e.g., objects 1202c, 1202f, 1202i). Similarly, the y-axis position of target 1202e may be calculated as 0.5 times the sum of the y-axis 1208 values of the virtual objects above the FOV (e.g., objects 1202a, 1202b, 1202c) and the virtual objects below the FOV (e.g., objects 1202g, 1202h, 1202i).

[0148] The wearable system may present the target within various eye pose regions of the display 220. The target may be shown as a graphic (e.g., a realistic or animated butterfly or mariposa or avatar, etc.). The graphic may be a still image that appears at a location within the FOV, or may appear to move from one location to another within the FOV.

[0149] The target may be displayed in various eye pose regions of the display 220 until an eye image with sufficient eye image quality is obtained for one or more than one eye pose region of the display 220. For example, the quality of the eye image may be determined and compared with a quality threshold to determine that the eye image has a quality that can be used for biometric applications (e.g., generation of an iris code). If the eye image within a certain eye pose region does not exceed or meet the quality threshold, the display 220 may be configured to continue to display one or more graphics within that specific region until an eye image with sufficient eye image quality is obtained. The one or more graphics displayed in one specific region may be the same or different in different implementations. For example, the graphics may be displayed at the same or different locations or in the same or different orientations within that specific region.

[0150] The graphics may be displayed in various eye pose regions of the display 220 using a story mode or a mode that can direct or attract the wearer's one or both eyes towards different regions of the display 220. For example, in one embodiment described below with reference to FIG. 12B, a virtual avatar (e.g., a butterfly) may be shown moving across various regions of the display 220. Instances of the graphics displayed in various regions of the display 220 may have properties (e.g., different depths, colors, or sizes) that attract or direct the wearer's one or both eyes towards one or more than one eye pose region in which the instance of the graphics is displayed. In some embodiments, the graphics displayed in various regions of the display 220 may appear to have variable depth such that the wearer's one or both eyes are attracted towards the eye pose region in which the instance of the graphics is displayed.

[0151] Figure 12B schematically illustrates an exemplary scene 1250 on a display 220 of a head-mounted display system. As depicted in Figure 12B, the display 220 may display the scene 1250 with a moving graphic 1205. For example, as depicted, the graphic 1205 may be a butterfly that is displayed to the user as flying throughout the scene 1250. The graphic 1205 may be displayed across or as part of a background image or scene (not shown in Figure 12B). In various embodiments, the graphic may be an avatar (e.g., a personification of a figure, animal, or object such as the butterfly or the honeybee 140 shown in FIG. 1 as an example), or any other image or video configured to be displayed in a particular eye gesture region of the display 220. The graphic 1205 may be adjusted for the user (e.g., based on age, anxiety level, maturity, interests, etc.). For example, to avoid causing anxiety to children, the graphic 1205 may be a children's character (such as a butterfly or a friendly honeybee 50 (FIG. 1)). As another example, for a user who likes cars, the graphic 1205 may be a car such as a race car. Thus, when moving within various regions of the display 220, the graphic 1205 may be displayed as a video and may appear to the wearer 210 using the wearable display system 200 as such. The graphic 1205 may start at an initial position 1210a and proceed along a path 1215 to a final position 1210b. For example, as depicted, the graphic 1205 may move across the display (e.g., along the dotted line) in a clockwise manner into different regions of the display 220. As another example, the graphic 1205 may appear to move in a zigzag or random manner across different regions of the display 220. One possible zigzag pattern may be regions 1220r1, 1220r2, 1220r4, 1220r0, 1220r3, 1220r5, 1220r7, and 1220r8.

[0152] In FIG. 12B, for illustrative purposes only, display 220 is shown as having nine regions 1220r0 - 1220r8 of the same size. The number of regions 1220r0 - 1220r8 of display 220 can be different in different implementations. Any number of regions of the display may be used to capture an eye image as the graphic progresses from one region to another to direct the eye towards that individual region. For example, the number of eye pose regions may be 2, 3, 4, 5, 6, 9, 12, 18, 24, 36, 49, 64, 128, 256, 1,000, or more. The eye image may be captured for some or all of the eye pose regions. The shapes of regions 1220r0 - 1220r8 of display 220 can be different in different implementations, such as rectangular, square, circular, triangular, oval, rhombus, etc. In some embodiments, the sizes of different regions of display 220 can be different. For example, regions closer to the center of display 220 can be smaller or larger than regions farther from the center of display 220. As another example, the eye pose regions may comprise halves, quadrants, or any segmentation of display 220.

[0153] Path 1215 may cross or move around within the eye pose region where it is desirable to obtain a good quality eye image, and may avoid the eye pose region where the eye image is undesirable (e.g., generally of poor quality) or not required (e.g., for certain biometric applications). For example, biometric applications (e.g., iris code generation) may tend to use an eye image where the user's eye is directed straight ahead (e.g., through eye pose region 1220r0). In such cases, graphic 1205 may tend to mainly move within eye pose region 1220r0 and not (or not very frequently) move within eye pose regions 1220r1 - 1220r8. Path 1215 may be more concentrated towards the center of scene 1250 compared to the peripheral region of scene 1250. In other biometric applications (e.g., diagnosis of the retina of the eye), it may be desirable to obtain an eye image of the user's eye looking in a direction away from region 1220r0 (e.g., away from the natural rest eye pose) such that an image of the inner or outer region of the retina (away from the fovea) is obtained. In such applications, graphic 1205 may tend to move around the periphery of scene 1250 (e.g., regions 1220r1 - 1220r8) compared to the center of the scene (e.g., region 1220r0). Path 1215 may be more concentrated around the periphery of the scene and tend to avoid the center of the scene (e.g., similar to path 1215 shown in FIG. 12).

[0154] The eye pose regions 1220r0 - 1220r8 of display 220 are depicted as being separated within display 220 by horizontal and vertical dotted lines for illustrative purposes only. Such eye pose regions 1220r0 - 1220r8 may represent, for the sake of convenience of explanation, the regions of display 220 towards which the wearer's eyes should be directed so that an eye image can be obtained. In some implementations, the horizontal and vertical dotted lines shown in FIG. 12B are invisible to the user. In some implementations, the horizontal or dotted lines shown in FIG. 12B may be visible to the user in order to direct one or both of the wearer's eyes towards a particular region of display 220.

[0155] The path 1215 shown in FIG. 12B is illustrative and not intended to be limiting. The path 1215 may have a shape different from that shown in FIG. 12B. For example, the path 1215 may cross, recross, or avoid one or more of the eye pose regions 1220r0 - 1220r1 and may be straight, polygonal, curved, or the like. The speed of the moving graphic 1215 may be substantially constant or variable. For example, the graphic 1205 may decelerate or stop within a certain eye pose region (e.g., where one or more eye images are captured), or the graphic 1205 may accelerate or skip through other eye pose regions (e.g., where an eye image is not needed or desired). The path 1215 may be continuous or intermittent (e.g., the graphic 1205 may skip over or around a certain eye pose region). For example, referring to FIG. 12B, when the graphic 1205 is at the position 1210b within the eye pose region 1220r4 and the biometric application requires an eye image with the user's eye directed towards the eye pose region 1220r8, the display system may display the graphic 1205 such that it continuously moves to the region 1220r8 (e.g., like a butterfly flying across the scene from the region 1220r4 through the region 1220r0 into the region 1220r8), or the display system may simply stop displaying the graphic 1205 within the region 1220r4 and then start displaying the graphic 1205 within the region 1220r8 (e.g., the butterfly would appear to jump from the region 1220r4 to 1220r8).

[0156] The eye pose region is a connected subset of the real two - dimensional coordinate space R 2 or the two - dimensional coordinate space of positive integers (N >0 ) 2 and can be regarded as such, which defines the eye pose region from the perspective of the angular space of the wearer's eye pose. For example, in one embodiment, the eye pose region is defined by a specific θ in the azimuthal deviation (e.g., the horizontal axis 1204 in FIG. 12A) min and a specific θmax and a particular φ in elevation deflection (e.g., the vertical axis 1208 in FIG. 12A). min and the particular φ max may be therebetween. Additionally, the eye pose regions may be associated with a particular region assignment. Such a region assignment may not appear on the display 220 for the wearer 210, but is shown in FIG. 12B for illustrative purposes. The regions may be assigned in any suitable manner. For example, as depicted in FIG. 12B, the central region may be assigned region 1220r0. In the depicted embodiment, the regions may be numbered in a sequential manner substantially horizontally, starting with region 1220r0 assigned to the central region and ending with the lower right region to which region 1220r8 is assigned. Such regions 1220r0-1220r8 may be referred to as eye pose regions. In other implementations, the regions may be numbered or referenced differently than those shown in FIG. 12B. For example, the upper left region may be assigned region 1220r0 and the lower right region may be assigned region 1220r8.

[0157] Scene 1250 may be presented by the wearable display system in the VR mode of the display. To the wearer 210, graphic 1205 is visible, but the outside world is not. Alternatively, scene 1250 may be presented in the AR / VR / MR mode of the display. To the wearer 210, visual graphic 1205 superimposed on the outside world is visible. While the graphic 1205 is being displayed within a certain eye gesture area, the eye image may be captured by an image capture device (e.g., the inward-facing imaging system 630 in FIG. 6) coupled to the wearable display system 200. Although only one example, one or more eye images may be captured within one or more of the eye gesture areas 1220r0 - 1220r8 of the display 220. For example, as depicted, the graphic 1205 may start at the initial position 1210a and move within the upper left eye gesture area (e.g., area 1220r1) of the display 220. As the graphic 1205 moves within its upper left eye gesture area, the wearer 210 may direct their eye towards that area of the display 220. While the graphic 1205 is within the upper left eye gesture area of the display 220, one or more eye images captured by the camera may include the eye in a certain eye gesture when looking in that direction.

[0158] Continuing with this embodiment, graphic 1205 may move along path 1215 to the central upper eye posture region (e.g., region 1220r2), and an eye image with an eye posture directed to the central upper region may be captured. While graphic 1205 may move along various eye posture regions 1220r0 - 1220r8 of display 220, the eye image is captured intermittently or continuously during this process until graphic 1205 reaches the final position 1210b within region 1220r4. One or more than one eye image may be captured for each region, or the eye image may be captured within all or fewer regions through which graphic 1205 moves. Thus, the captured eye image may include at least one image of the eye in one or more different eye postures. The eye posture may be represented as an expression of two angles, as will be further explained below.

[0159] Graphic 1205 may also remain within the eye posture region of display 220 until an image of a certain image quality is acquired or captured. As described herein, various image quality metrics are available to determine whether a certain eye image exceeds an image quality threshold (Q). For example, the image quality threshold may be a threshold corresponding to an image metric level for generating an iris code. Thus, if the eye image captured while graphic 1205 is within a certain eye posture region of display 220 exceeds the image quality threshold, graphic 1205 may remain within that eye posture region (or return to that eye posture region) until an image that meets or exceeds the image quality threshold is acquired. The image quality threshold may also be defined for a particular eye posture region of the display. For example, certain biometric applications may require dimming of a certain region of display 220. Thus, the image quality threshold for those regions may be higher than the image quality threshold for non - dimmed regions. During this image collection process, graphic 1205 may continue in a story mode or a video that continuously directs the wearer's eye towards that region.

[0160] The eye image collection routine may also be used to correct the vulnerable bits in the iris code. Vulnerable bits refer to bits of the iris code that are inconsistent between eye images (for example, there is a substantial probability that a bit is zero for some eye images and one for other images of the same iris). More specifically, the vulnerable bits can be bits that are ambiguously defined within the iris code of an eye image and that may represent empirical unreliability in the measurement. The vulnerable bits may be quantified, for example, using a Bayesian model for uncertainty in the parameters of a Bernoulli distribution. The vulnerable bits may also be identified, for example, as those bits that represent areas that are typically covered by an eyelid or occluded by an eyelash. The eye image collection routine may use Graphic 1205 to actively guide the eye into different eye poses, thereby reducing the impact of the vulnerable bits on the resulting iris code. Although only one example, Graphic 1205 may guide the eye into an eye pose area that is not occluded by an eyelid or an eyelash. Additionally, or alternatively, a mask may be applied to the eye image to reduce the impact of the vulnerable bits. For example, the mask may be applied such that eye regions identified as producing vulnerable bits (for example, the upper or lower portions of the iris where occlusion is more likely to occur) may be ignored for iris generation. As yet another example, Graphic 1205 may return to eye pose areas that are more likely to produce vulnerable bits and acquire more eye images from those areas, thereby reducing the impact of the vulnerable bits on the resulting iris code.

[0161] Graphic 1205 may also stay (or return) within the eye gesture region of display 220 until several images are captured or acquired with respect to a particular eye gesture region. That is, instead of comparing the image quality metric of each eye image with the image quality threshold "on the fly" or in real time, a number of eye images may be acquired from each eye gesture region. Each of the eye images acquired with respect to that eye gesture region may then be processed to obtain an image quality metric, which in turn is compared with an individual image quality threshold. As can be seen from the figures, the eye gesture regions of the eye image collection process may be implemented in parallel or sequentially, depending on the needs or requirements of the application.

[0162] During this eye image collection routine, the graphic may be displayed in one or more than one eye gesture region of display 220 in various modes. For example, the graphic may be displayed within a particular eye gesture region of the display (or across two or more than two eye gesture regions) in a random mode, a flight mode, a blinking mode, a varying mode, or a story mode. The story mode may contain various videos with which the graphic may interact. As an example only of a story mode, a butterfly may emerge from a cocoon and fly around the perimeter of a particular region of display 220. As the butterfly flies around, flowers may appear from which the butterfly may collect nectar. As can be seen from the figures, the story of the butterfly may be displayed within a particular region of display 220 or across two or more than two regions of display 220.

[0163] In the variable mode, as the butterfly wings flit around within a specific area of the display 220, they may appear to vary in size. In the random mode, the exact location of the graphic 1205 within a specific area may be randomized. For example, the graphic 1205 may simply appear in different locations within the upper left area. As another example, the graphic 1205 may start from the initial position 1210a and move within the upper left eye gesture area in a partially random pattern. In the blink mode, a butterfly or a group of butterflies may appear to blink across a specific area of the display 220 or across two or more areas. Various modes are conceivable within the various eye gesture areas of the display 220. For example, the graphic 1205 may appear within the upper left area at the initial position 1210a in the story mode, while the graphic 1205 may appear within the central left area at the final position 1210b using the blink mode.

[0164] The graphic may also be displayed throughout the entire eye gesture area 1220r0 - 1220r8 of the display 220 in various modes. For example, the graphic may appear in a random or sequential pattern (referred to as the random mode or the sequential mode, respectively). As described herein, the graphic 1205 may move across the various areas of the display 220 in a sequential pattern. Continuing with that example, the graphic 220 may move along the path 1215 using the intermediate video between the eye gesture areas of the display 220. As another example, the graphic 1205 may appear in different areas of the display 220 without an intermediate video. As yet another example, a first graphic (e.g., a butterfly) may appear within a first eye gesture area, while another graphic (e.g., a maruhana bachi) may appear within a second eye gesture area.

[0165] Different graphics can appear sequentially from one area to the next. Or, in another embodiment, various graphics may be used in a story mode such that different graphics appear within different eye pose regions to tell a story. For example, a cocoon can appear within one eye pose region and then a butterfly can appear within another region. In various implementations, different graphics may also be randomly dispersed through the eye pose regions as the eye image collection process can direct the eye from one eye pose region to another using different graphics that appear within each eye pose region.

[0166] Eye images may also be acquired in a random fashion. Thus, graphic 1205 may also be displayed in a random fashion within the various eye pose regions of display 220. For example, graphic 1205 can appear within the upper central region and once the eye image is acquired with respect to that region, graphic 1205 can then appear within the lower right eye pose region (e.g., region 1220r8 is assigned) of display 220 in FIG. 12B. As another example, graphic 1205 is displayed in an ostensibly random manner and may be displayed at least once on each eye pose region without replication over individual regions until graphic 1205 is displayed in other regions. Such a pseudo-random fashion of the display can occur until a sufficient number of eye images for image quality threshold or some other purpose are acquired. Thus, various eye poses regarding one or both eyes of the wearer may be acquired in a random fashion rather than a sequential fashion.

[0167] In some cases, if the eye image cannot be obtained for a certain eye pose region after a threshold number of trials (e.g., the three eye images captured for the eye pose region do not exceed the image quality threshold), the eye image collection routine may first obtain the eye image from one or more other eye pose regions while skipping or pausing the collection over that eye pose region for a certain time period. In one embodiment, the eye image collection routine may not obtain the eye image for a certain eye pose region if the eye image cannot be obtained after a threshold number of trials.

[0168] The eye pose can be described with respect to a natural rest pose (e.g., both the user's face and line of sight are oriented as if they would be directed towards a distant object in front of the user). The natural rest pose of the eye can be indicated by a natural rest position that is in the direction orthogonal to the surface of the eye when in the natural rest pose (e.g., straight out from the plane of the eye). As the eye moves to look towards different objects, the eye pose changes with respect to the natural rest position. Thus, the current eye pose may be measured with reference to the eye pose direction, which is the direction orthogonal to the surface of the eye (and centered within the pupil) but oriented towards the object that the eye is currently directed towards.

[0169] With reference to an exemplary coordinate system, eye pose can be expressed as two angular parameters, both indicating the azimuth and zenith deflections of the eye's eye attitude direction relative to the natural resting position of the eye. These angular parameters can be expressed as θ (azimuth deflection, measured from a base azimuth angle) and Φ (elevation deflection, sometimes also referred to as polar deflection). In some implementations, the angular roll of the eye around the eye attitude direction may be included in the measurement of eye attitude, and the angular roll may be included in the subsequent analysis. In other implementations, other techniques for measuring eye attitude may be used, such as pitch, yaw, and optionally roll systems. Using such expressions for eye attitude, eye attitude, expressed as azimuth and zenith deflections, may be associated with a particular eye attitude region. Thus, eye attitude may be determined from each eye image acquired during the eye image collection process. Such associations between eye poses and eye regions of the eye images may be stored in the data modules 260, 280 or made accessible to the processing modules 260, 270 (e.g., accessible via cloud storage).

[0170] Eye images may also be acquired selectively. For example, an eye image of a particular wearer may already be stored or accessible by the processing module 260, 270. As another example, an eye image for a particular wearer may already be associated with an eye posture region. In such a case, the graphic 1205 may appear in only one eye posture region or a particular eye posture region that does not have an eye image associated with the eye posture region or the particular eye posture region. By way of illustration, eye images may have been acquired for eye region numbers 1, 3, 6, and 8, but not for other eye posture regions 2, 4, 5, and 7. Thus, the graphic 1205 may appear in the latter posture regions 2, 4, 5, and 7 until an eye image for each individual eye posture region that exceeds the image quality metric threshold has been acquired.

[0171] A detailed example of eye image collection and analysis for eye gaze is described in U.S. Patent Application No. 15 / 408,277, filed on January 17, 2017, entitled "Eye Image Collection", the disclosure of which is hereby incorporated by reference in its entirety. D. Examples of Verifying Eye Gaze

[0172] The wearable system may obtain eye images during the eye tracking calibration process described with reference to FIGS. 12A and 12B. However, one issue in the eye tracking calibration process is that the user may not be looking at the target as expected. For example, when the wearable system renders a target (e.g., virtual butterfly 1205 or one of targets 1202a-i) within the rig space, the user may be looking at a different directional target instead of the graphic. For example, in one laboratory-based experiment, 10 percent of the users did not look at some of the targets even under laboratory test conditions during calibration. User compliance with the calibration protocol can be substantially lower when the user is alone in a home or office environment. As a result, the wearable system may not obtain accurate eye tracking results from calibration, and as a result, the user's visual experience using the wearable system may be affected.

[0173] To improve this problem and improve the quality of data obtained regarding the line of sight, the wearable system may verify the user's line of sight before adjusting the mapping matrix for calibration. During the line of sight verification, the wearable system may use the head pose (e.g., head position or rotation information) to verify that the user is actually looking at the target. FIG. 12C illustrates an example of using the user's head pose to verify whether the user is looking at the target. FIG. 12C illustrates three scenarios 1260a, 1260b, and 1260c. In these three scenarios, the user may perceive the reticle 1284 and the target 1282 via the display 220. The reticle 1284 represents a virtual object in the rig space, while the target 1282 represents a virtual or physical object at a given location within the user's environment. The location of the target 1290 may be represented by a position in world space associated with the world coordinate system. The world coordinate system may be with respect to the user's 3D space rather than the user's HMD. As a result, the objects within the world coordinate system may not necessarily align with the objects within the rig space.

[0174] During the eye gaze verification process, the user needs to align the reticle 1284 with the target 1290, and the wearable system may instruct the user to "aim" the reticle 1284 at the target 1290. As the reticle 1284 moves within the rig space, the user needs to move their head and eyes in order to be able to realign the reticle 1284 with the target again. The wearable system may check whether the reticle 1284 is aligned with the target 1290 (e.g., by comparing the measured user head pose or eye gaze with the known position of the target) and provide feedback to the user (e.g., indicating whether the reticle 1284 is aligned with the target 1290). Advantageously, in some embodiments, the wearable system may be configured to collect an eye image for eye tracking calibration only when there is sufficient alignment between the reticle 1284 and the target 1290. For example, the wearable system may determine that sufficient alignment exists when the offset between the positions of the target and the reticle is different by less than a threshold amount (e.g., less than an angular threshold such as less than 10°, less than 5°, less than 1°, etc.).

[0175] Referring to FIG. 12C, the head 1272 is initially at position 1276a and the eye 1274 is fixated on direction 1278a within scene 1260a. The user may perceive that the reticle 1284 is located at position 1286a via the display system 220. As shown in scene 1260a, the reticle 1284 is aligned with the target 1290.

[0176] During the calibration process, the wearable system may render the reticle 1284 at different locations within the user's FOV. In scene 1260b, the reticle 1284 is moved to position 1286b. As a result of this movement, the reticle 1284 is no longer aligned with the target 1290.

[0177] The user may need to rotate their eyeballs and / or move their head 1272 to realign the reticle 1284 with the target 1290. As depicted in scene 1260c, the user's head is tilted to position 1276c. In scene 1260c, the wearable system may analyze the user's head pose and line of sight and determine that the user's line of sight direction is currently in direction 1278c as compared to direction 1278a. Due to the movement of the user's head, the reticle 1284 is moved to position 1286c and aligned with the target 1290 as shown in scene 1260c.

[0178] In FIG. 12C, the location of the reticle 1284 may be associated with a position within the rig space. The location of the target 1290 may be associated with a position within the world space. As a result, the relative position between the reticle 1284 and the display 220 does not change even when the user's head pose changes within scenes 1260b and 1260c. The wearable system can align the reticle, and the target can align the position of the reticle within the rig space with the position of the reticle within the world space.

[0179] Advantageously, in some embodiments, the wearable system may utilize the user's vestibulo-ocular reflex to reduce the discomfort and eye strain caused by the calibration process. The wearable system may automatically track and infer the line of sight based on the head pose. For example, when the user's head moves to the right, the wearable system may track it and infer that the eyes will necessarily move to the left under the vestibulo-ocular reflex.

[0180] FIG. 13A illustrates an example of verifying a line of sight where the reticle is at the center of the user's FOV 1350. In FIG. 13A, three time-series scenes 1310, 1312, and 1314 are shown. In this example, the user can perceive an eye calibration target 1354 and a reticle 1352. The target 1354 (e.g., a diamond-shaped graphic) is displayed to be fixed within the three-dimensional space of the user's environment and is located away from the virtual reticle (e.g., offset from the center within the user's FOV). The reticle 1352 (e.g., a hoop or ring-shaped graphic) is displayed to be fixed at or near the center of the user's FOV 1350. For example, the center or the vicinity of the FOV may have an angular offset of less than 10°, less than 5°, less than 1° etc.

[0181] In scene 1310, the reticle 1352 is not aligned with the target 1354 and the reticle 1352 is slightly below the target 1354. As described with reference to FIG. 12C, the user may move their head to align the reticle 1352 and the target 1354. The wearable system may detect the user's head movement using the IMU described with reference to FIG. 2. In some embodiments, the head pose may be determined based on data obtained from other sources such as a reflection image of the user's head as observed by a sensor external to the HMD (e.g., a camera in the user's room) or an outward-facing imaging system 112 (FIG. 9E). As shown in scene 1312, the user may try to move their head upward to align the reticle 1352 and the target 1354. Once the reticle reaches the position as shown in scene 1314, the wearable system may determine that the reticle 1352 is properly aligned with the eye calibration target 1354 and thus the user's head is properly positioned to view the eye calibration target.

[0182] The wearable system may calculate the alignment between the reticle and the eye calibration target using various techniques. As one example, the wearable system may determine the relative position between the reticle and the eye calibration target. If the eye calibration target is within the reticle or a portion of the eye calibration target overlaps the reticle, the wearable system may determine that the reticle is aligned with the eye calibration target. The wearable system may also determine that the reticle and the target are aligned if the centers of the reticle and the target sufficiently coincide. In some embodiments, since the reticle is within the rig space while the target is within the world space, the wearable system may align the coordinate system associated with the rig space and the coordinate system associated with the world space and be configured to determine whether the reticle is aligned with the target. The wearable system may determine whether the reticle and the target overlap or coincide by determining that the relative offset between them is less than a threshold (e.g., an angular threshold as described above). In some examples, the present threshold may correspond to one or more thresholds associated with the user's head pose, as will be described in more detail below with reference to FIGS. 14A and 14B.

[0183] The wearable system may also identify a target head pose that represents the head pose at which alignment between the reticle and the eye calibration target occurs. The wearable system may compare the user's current head pose with the target head pose and verify that the user is actually looking at the target. The target head pose may be specific to the position of the reticle or the target within the 3D space. In some embodiments, the target head pose may be estimated based on data associated with the user or other people (e.g., a previous user in front of the wearable system, a user of one or more other similar wearable systems that communicate with one or more servers or other computing devices with which the wearable system communicates over a network, etc.).

[0184] In one embodiment, the wearable system may determine the alignment between the target and the reticle using ray casting or cone casting techniques. For example, the wearable system may project a ray or a cone (including a volume lateral to the ray) and determine the alignment by detecting a collision between the ray / cone and the target. The wearable system may detect a collision when a portion of the ray / cone intersects the target or when the target enters the volume of the cone. The direction of the ray / cone may be based on the user's head or line of sight. For example, the wearable system may project the ray from a location between the user's eyes. The reticle may reflect a portion of the ray / cone. For example, the shape of the reticle may match the shape on the distal end of the cone (e.g., the end of the cone away from the user). If the cone is a geometric cone, the reticle may have a circular or oval shape (which may represent a portion of the cone such as the cross-section of the cone). In one implementation, since the reticle is rendered in the rig space, as the user moves around, the wearable system may update the direction of the ray / cone even if the relative position between the ray and the user's HMD does not change.

[0185] Once the wearable system determines that the user is looking at the target (e.g., because the reticle is aligned with the target), the wearable system may start collecting eye line-of-sight data for calibration purposes, for example, using an imaging system 630 (FIG. 6) facing inward. In some examples, the wearable system may first store the output of one or more eye tracking sensors or processing modules (e.g., a local processing data module) in a temporary data storage device (e.g., a cache memory etc.) that is routinely flushed. In response to determining that the user is actually looking at the target, the wearable system may proceed to transfer the output data from the temporary data storage device to another data storage device, such as a disk or another memory location, for further analysis or long-term storage.

[0186] After the eye gaze data has been collected, the system can either complete the eye tracking calibration process or proceed to render another eye calibration target or reticle so that additional eye gaze data can be collected. For example, after the wearable system has collected eye data within scene 1314 shown in FIG. 13A, the reticle 1352 may be presented at different locations within the user's FOV 1350 as shown in scene 1320 of FIG. 13B. In some embodiments, the wearable system may evaluate each collected frame against a reference set to determine whether each frame represents data suitable for use in the eye tracking calibration process. For a given frame, such an evaluation may include, for example, steps of determining whether the user blinked during the collection of the frame, determining whether the target and reticle were properly aligned with each other during the collection of the frame, determining whether the user's eyes were detected normally during the collection of the frame, and the like. In these embodiments, the wearable system may determine whether a threshold amount of frames (e.g., 120 frames) that meet the reference set have been collected, and in response to determining that the threshold amount of frames has been met, the wearable system may complete the eye tracking calibration process. The wearable system may proceed to render another eye calibration target or reticle in response to determining that the threshold amount of frames has not yet been met.

[0187] Figure 13B illustrates an example of verifying the eye line of sight where the reticle is rendered at a location outside the center within the user's FOV 1350. The location of the virtual reticle in Figure 13B is different from the location of the virtual reticle in Figure 13A. For example, in Figure 13A, the location of the virtual reticle is at or near the center of the user's FOV, while in Figure 13B, the location of the virtual reticle is offset from the center of the user's FOV. Similarly, the location of the target is different in Figure 13A (e.g., towards the upper part of the FOV) from the location of the target in Figure 13B (e.g., at or near the center of the FOV). In Figure 13B, three time-series scenes 1320, 1322, and 1324 are shown. In this example, the reticle 1352 is rendered on the right side of the user's FOV 1350, and the target 1354 is rendered near the center of the user's FOV 1350. From scene 1314 to scene 1320, it can be seen that the location within the user's FOV 1350 where the reticle 1352 is rendered is updated, but the location within the environment where the target 1354 is rendered remains substantially the same. To align the reticle 1352 and the target 1354, the user can rotate their head to the left to align the reticle with the eye calibration target (see illustrative scenes 1322 and 1324). Once the wearable system determines that the target 1354 is within the reticle 1352, the wearable system may start collecting eye line of sight data in a manner similar to the example described above with reference to Figure 13A. If the user's eye line of sight moves (e.g., to the extent that the target and the reticle are no longer sufficiently aligned), the wearable system may stop collecting eye line of sight data because the user is no longer looking at the target and any acquired data would be of lower quality.

[0188] In one embodiment, the wearable system may calculate a target head pose at which the reticle 1352 is aligned with the target 1354. The wearable system may track the head pose of the user as the user moves. Once the wearable system determines that the user has taken the target head pose (e.g., the head pose shown in scenes 1314 or 1324), the wearable system may determine that the target 1354 and the reticle 1352 are aligned, and the wearable system may collect an eye image when the head is in the target head pose. E. Exemplary Process of Eye Tracking Calibration with Eye Gaze Verification

[0189] FIG. 14A illustrates an exemplary flowchart for an eye tracking calibration process with eye gaze verification. The exemplary process 1400 may be performed by one or more components of the wearable system 200, such as the remote processing module 270 or the local processing and data module 260, alone or in combination. The display 220 of the wearable system 200 may present a target or a reticle to the user, the imaging system 630 facing inward (FIG. 6) may acquire an eye image for eye gaze determination, and the IMU, accelerometer, or gyroscope may determine the head pose.

[0190] In block 1410, the wearable system may render an eye calibration target within the user's environment. The eye calibration target may be rendered within the world space (which can be represented by a coordinate system with respect to the environment). The eye calibration target may be represented in various graphical forms, including 1D, 2D, and 3D images. The eye calibration target may also include a still image or a moving image (e.g., a video, etc.). Referring to FIG. 13A, the eye calibration target is schematically represented by a rhombus.

[0191] In block 1420, the wearable system may identify a range of head postures associated with the rendered eye calibration target. The range of head postures may include a plurality of head postures (e.g., 2, 3, 4, 5, 10, or more). A head posture may describe the position and orientation of the user's head. The position may be represented by translational coordinate values. The orientation may be represented by angular values relative to the natural resting state of the head. For example, the angular values may represent forward and backward head tilts (e.g., pitch), left and right direction changes (e.g., yaw), and lateral tilts (e.g., roll). The wearable system may identify a range of head positions and a range of head orientations, which together may define a range of head postures in which the reticle and target are considered to be sufficiently aligned with each other. The boundaries of such a range may be considered to correspond to threshold values. Head postures that fall within this range may correspond to a target head posture for the user to align the target and reticle while the reticle appears in different regions of the user's FOV. Referring to FIGS. 13A and 13B, the range of head postures may include head postures 1314 and 1324, and the wearable system may determine that the head positions and orientations corresponding to head postures 1314 and 1324 fall within the identified ranges of head positions and head orientations, respectively, and thus meet one or more threshold values or other requirements for sufficient reticle-target alignment.

[0192] The wearable system may track the head posture using sensors inside or outside the HMD, such as an IMU or an outward-facing imaging system (e.g., for tracking the reflected image of the user's head), e.g., a camera mounted on a wall in the user's room. In block 1430, the wearable system may receive data indicating the user's current head posture. The data may include the current position and orientation of the user's head or the movement of the user's head in 3D space. For example, in FIG. 13A, as the user moves their head from the position shown in scene 1310 to the position shown in scene 1314, the wearable system may track and record the user's head movement.

[0193] In block 1440, the wearable system may determine whether the user is taking a head pose that falls within the range of the identified head pose based on the data obtained from block 1430. The wearable system may determine whether the user's head pose is in a position or orientation that can align the reticle with the target. As an example, the wearable system may determine whether both the head position and head orientation associated with the user's head pose fall within the range of the identified head position and the range of the identified head orientation. The wearable system can make such a determination by comparing the head position associated with the user's head pose with a threshold (e.g., translational coordinate values) that defines the boundary of the range of the identified head position, and by comparing the head orientation associated with the user's head pose with a threshold (e.g., angular values) that defines the boundary of the range of the identified head orientation. Referring to FIG. 13A, the wearable system may determine whether the user is taking the head pose shown at 1314. If the user is not taking a head pose that falls within the range of the identified head pose and thus not taking a head pose where the reticle and the target are considered to be sufficiently aligned with each other, the wearable system may continue to obtain and analyze the data associated with the user's head pose as shown in block 1430.

[0194] Optionally, at 1450, the wearable system may provide the user with feedback (e.g., visual, auditory, tactile, etc.) indicating that the user's head is properly positioned. For example, the visual feedback may include a color change or a blinking effect of the target or the reticle that can indicate that the user's head is properly positioned so that the reticle aligns with the target by blinking or changing the color of the reticle and / or the eye calibration target. In some embodiments, blocks 1410 - 1450 are part of an eye gaze verification process.

[0195] If it is determined that the user's head is within one of the identified head postures, at block 1460, the wearable system may receive and store data indicative of the user's eye gaze, in association with an eye calibration target. In the context of FIG. 13A, when the wearable system detects that the user head posture is at the position and orientation shown in scene 1314, the wearable system may receive and store data from one or more eye tracking sensors (e.g., an eye camera within the imaging system 630 (FIG. 6) facing inward).

[0196] At block 1470, the wearable system may determine whether additional data should be collected during eye tracking calibration. For example, the wearable system may determine whether an eye image in another eye gaze direction should be collected to update or complete the calibration process. If it is determined that additional eye calibration data should be collected, the wearable system may return to block 1410 and repeat process 1400. Referring to FIGS. 13A and 13B, for example, when the user 210 is at the position shown in scene 1314, after the wearable system has collected an eye image, it may render the target 1354 as shown in scene 1322.

[0197] In some embodiments, even when the user is actually looking at the target, the images obtained by the wearable system may be considered unsatisfactory (e.g., due to user blinking). As a result, the process may return to block 1460 to capture additional images.

[0198] If it is determined that no additional eye calibration data needs to be collected, at block 1480, the wearable system may complete process 1400 and use the stored eye gaze data for eye tracking calibration. For example, the stored data may be used to generate the mapping matrix described above.

[0199] FIG. 14B illustrates an exemplary eye gaze verification process. The exemplary process 1490 may be performed by one or more components of a wearable system, such as, for example, the remote processing module 270 and the local processing and data module 260, alone or in combination. The wearable system may include an HMD. The display 220 of the wearable system 200 may present a target or reticle to the user, and the inward-facing imaging system 630 (FIG. 6) may acquire an eye image for eye gaze determination, and the IMU, accelerometer, or gyroscope may determine the head pose.

[0200] In block 1492a, the wearable system may determine a target in a world space associated with the user's environment. The target may be fixed at a given location in the world space. The target may be a virtual object rendered by the display 220 or a physical object (e.g., a vase, shelf, flower pot, book, painting, etc.) in the user's environment. The virtual target may have various appearances as described with reference to FIGS. 12A, 12B, and 18. The world space may include the world map 920 shown in FIG. 11B. The location of the target in the world space may be represented by a position in a 3D world coordinate system.

[0201] In block 1492b, the wearable system determines a reticle in a rig space associated with the user's HMD. The reticle may be rendered by the HMD at a predetermined location within the user's FOV. The rig space may be associated with a coordinate system separate from the world coordinate system.

[0202] In block 1494, the wearable system may track the user's head pose. The wearable system may track the head pose based on an IMU within the user's HMD or an outward-facing imaging system. The wearable system may also track the head pose using other devices such as a webcam or a totem in the user's room that may be configured to image the user's environment. As the user's head pose changes, the relative position between the reticle and the target may also change.

[0203] In block 1496, the wearable system may update the relative position between the reticle and the target based on the head pose. For example, if the target is to the right of the reticle and the user turns their head to the right, the reticle may appear to move closer to the target. However, if the user turns their head to the left, the reticle may appear to move further away from the target.

[0204] In block 1498a, the wearable system may determine whether the target and the reticle are aligned. Alignment may be performed using ray / cone casting. For example, the wearable system may project a ray from the reticle and determine whether the target intersects the ray. If the target intersects the ray, the wearable system may determine that the target and the reticle are aligned. The wearable system may also determine an offset between a position in the rig space and a position in the world space based on the user's head pose. The wearable system may apply the offset to the reticle (or the target) and determine the position of the reticle that coincides with the position of the target, thereby aligning the location of the target in the world space and the location of the reticle in the rig space. In some situations, the offset may be used to translate the position of the reticle from the rig space to the corresponding position in the world space. The alignment between the reticle and the target may be determined based on the coordinate values of the reticle and the target with respect to the world space.

[0205] If the target and the reticle are misaligned, the wearable system may continue to track the head pose at block 1494. If the target and the reticle are aligned, the wearable system may determine that the user is actually looking at the target, and at block 1498b, may provide an indication that the user's line of sight direction has been verified. The indication may include an audio, visual, or tactile effect.

[0206] In some embodiments, the wearable system may present a series of reticles (e.g., each within a different line of sight region shown in FIG. 12B) for eye tracking calibration. As a result, after block 1498b, the wearable system may optionally resume at block 1492a and present the reticle at a new location within the rig space. The user may attempt to realign the reticle and the target at the new location by changing the user's head pose. F. Exemplary Object Movement

[0207] A long eye calibration process can cause fatigue to the user. Additionally, a long-term eye calibration process can lead to a state of diverting the user's attention during calibration, resulting in poor data quality regarding eye calibration. The length of the eye calibration process can be determined by the speed of movement of the eye calibration target and the user's ability to track the movement of the eye calibration target. Additionally, the human eye may be forced to make corrective eye movements to overshoot the eye calibration target and align the line of sight with the location of the eye calibration target. These corrective eye movements can add to the length of the eye calibration process and, in addition, increase user fatigue. Apart from the calibration target, a similar discomfort can also occur when the user is presented with a virtual object that moves, for example, at a certain speed and then suddenly stops.

[0208] FIG. 15A shows a graph of exemplary pupil velocity 1604 as a function of time for an eye, where a virtual object, i.e., an eye target, is presented. For example, the target may be displayed to the user by a head-mounted display of a display system or the like at one or more locations within its field of view. The target may move periodically from one location to another at a constant speed. As illustrated in FIG. 15A, the target may move with sudden starts and stops, which constitute a square wave as represented by line 1602. From the stop, the eye tracks the moving target with velocity 1604. This velocity 1604 may vary as the target is moved. As illustrated in FIG. 15A, the eye velocity 1604 may have some local peaks 1606 when the target is stationary. The peaks 1606 are understood to correspond to overshoot of the eye. Overshoot of the eye can occur when the eye continues to move past the location of the stationary target and then corrects and moves back to the location of the stationary target. In a calibration system where the target moves at a constant speed between target locations, the dwell time for the target (in other words, the time at each target location during target movement) may take into account the time it takes for the eye to perform this overshoot correction. As discussed above, the additional time and corrective eye movements resulting from overshoot can cause discomfort and fatigue to the user, result in a more inaccurate calibration, and / or lengthen the calibration process.

[0209] FIG. 15B illustrates the same graph as FIG. 15A, but provides labels for different portions of the graph to facilitate discussion of those portions. As shown in FIG. 15B, the target can move with a period represented by line 1602. The eye velocity 1604 that tracks the target can increase in velocity when the eye target first moves. For example, as shown in FIG. 15B, the eye velocity 1604 can have a peak 1608 when starting to proceed to a new target location. Without being bound by theory, as shown in FIG. 15B, the human eye is thought to move angularly fast and / or be able to increase velocity when first starting to move its line of sight to a new target location. However, also as shown in FIG. 15B, the human eye is thought not to respond as quickly by stopping at the steady end point of the target movement.

[0210] In some embodiments, systems and methods are provided for moving an eye target that can help reduce time and fatigue associated with viewing a virtual object that a user's eye is expected to track. For example, the eye target movement process disclosed herein may adjust the speed of movement of the virtual object based on the target position and / or depth. For example, rather than moving the virtual object at a constant speed, the display system may present the virtual object such that its movement gradually decreases in speed as it approaches an end position where the virtual object stops moving. Without being bound by theory, this is thought to enable better stability in eye movement and reduce the occurrence of eye corrections due to overshooting of the eye. Advantageously, this can result in a significant shortening of the total eye calibration duration, reduce eye fatigue, and / or help contribute to higher eye calibration quality. Additionally, shorter eye calibrations can result in less distraction and less eye fatigue, which in turn contributes to higher eye calibration quality.

[0211] Referring now to FIG. 16, an example of a virtual object movement process 1700 that may be implemented by a display system 60 (as illustrated in FIG. 9E) is shown. It should be understood that the virtual object may be at an initial target location, i.e., a first target location. In block 1702 of process 1700, the display system may identify a new target location for the virtual object to move to, i.e., a second target location. For example, the initial target location may be a first perceived 3D location in the user's 3D environment, and the new target location may be a new perceived 3D location in the user's 3D environment. Preferably, the new target location is within the user's field of view such that the user does not need to move their head to perceive the object at the new location.

[0212] Continuing to refer to FIG. 16, in block 1704, the display system may determine the distance and / or depth from the initial target location to the new target location for the virtual object. In some embodiments, the distance may include the angular distance between the initial target location and the new target location as perceived by the user. In some embodiments, a point associated with the user or the user's line of sight, such as an eye target whose apex, which is a reference point for defining the angle for determining the angular distance, is configured to be displayed between the user's eyes at the center of rotation of the user's eyes, or such a point, on a head-mounted display worn by the user, may be used. It should be understood that the angular distance may be used to define a distance that extends across (laterally and / or vertically) the user's field of view. Preferably, the angular distance defines a distance between locations on the same depth plane. In some embodiments, the angular distance defines a distance between locations that are displaced laterally and / or vertically within the user's field of view and across two or more depth planes.

[0213] In some embodiments, instead of the angular distance and the initial and new target locations crossing two or more depth planes, the distance between the initial location and the new target location may be referred to as the distance in diopter units (or simply, diopters). In some embodiments, depth may include the distance from a reference point (e.g., the user's eye) for distance determination. For example, the distance may be a linear distance from the reference point extending into the user's environment. As discussed herein, FIG. 17 illustrates examples of target distances and depths at first and second locations.

[0214] Continuing to refer to FIG. 16, at block 1706, the display system may determine a travel time for the eye target to move from the initial target location to the new target location. In some embodiments, the travel time may be an estimated amount of time it takes for the user's eye to comfortably move from the initial target location to the new target location. As described herein, the travel time may be estimated based on the diopter and / or angular difference between the two locations. In some embodiments, the travel time may be a desired travel time, and the speed curves or functions discussed herein may be adapted within this travel time to increase visual comfort, assuming the constraints of the desired travel time.

[0215] In block 1708, the display system may interpolate the velocity function or position of the eye target as a function of time. In some embodiments, the interpolation for the eye target may include a mathematical function that allows the target to gradually accelerate to a maximum speed and then gradually decelerate and come to rest at a new target location. As described herein, the plot of velocity interpolation preferably takes the form of an S-curve. In some embodiments, the velocity of the target may be interpolated from an S-curve using interpolation parameters that are associated with the travel time and / or dioptric distance or depth determined in block 1706. In some examples, the interpolation parameters associated with the dioptric distance may be a function of the reciprocal of the dioptric distance. Advantageously, although not limited by theory, this inverse correlation is thought to account for different responses of the eye to adapt to changes with distance from the user.

[0216] In block 1710 of process 1700, the display system may preferably move the eye target based on an interpolation (or velocity function) that defines an S-shaped velocity plot. For example, the display system may move the target from an initial target location to a new target location along a specified path. The specified path may be the shortest path between two points or a longer path (e.g., the specified path may be a linear path or a path with turns or curves in different directions). In an eye calibration example, the display system may move the target to a new point within the user's 3D environment that can be calculated by the eye calibration system to provide an accurate eye calibration. The movement may have a velocity and / or acceleration according to the interpolation calculated in block 1708. G. Exemplary Travel Time Calculation

[0217] As discussed herein, in some embodiments, the travel time for the movement of a virtual object from an initial (or first) target location to a new (or second) target location may be a predetermined set travel time in which an S-shaped velocity curve is fitted. Preferably, the predetermined set travel time is long enough such that the user's eye can be expected to move comfortably, or change the focus state, and track the virtual object. For example, the set travel time is long enough such that the eyes of most or all of the user population can move fast enough, or change the focus, so that they can comfortably traverse the desired distance between the first location and the second location.

[0218] It should be understood that in some other embodiments, the eyes of different users may necessarily move or change the focus at different speeds. As a result, the display system may customize the travel time for different users and determine the amount of time for the target to travel or move between the initial (or first) target location and the new (or second) target location. In some examples, the determined travel time may shorten the eye tracking time while maintaining a desired level of tracking accuracy. In some embodiments, as discussed herein, the calculation of the travel time may consider various parameters. Examples of parameters include whether the distance to be traversed by the virtual object is a distance that extends across the user's field of view (e.g., an angular distance), or a distance along the z-axis that extends along the user's line of sight (e.g., a depth distance away from or towards the user). Other parameters that may be used to adjust or provide a correction factor in the travel time calculation include the natural speed at which the user's eye can move or change the focus state, the magnitude of the distance between the initial target location and the new target location, and / or the user's age, gender, eye health, etc.

[0219] FIG. 17 illustrates an example of a general framework for explaining the movement of a virtual object between two locations. As discussed herein, the virtual object may be provided at a first target location 1502a and moved to a second target location 1502b. In some embodiments, θ1L and θ1R may respectively refer to the tangential angles of the user's left and right eyes with respect to the first target location 1502a. Similarly, θ2L and θ2R may respectively refer to the tangential angles of the left and right eyes with respect to the second target location 1502b. In some embodiments, d1L and d1R may respectively refer to the distances from the left and right eyes to the first target location 1502a. Similarly, d2L and d2R may respectively refer to the distances from the left and right eyes to the second target location 1502b.

[0220] It should be understood that the eye may require a certain amount of time to move angularly. f(θ) may represent a target tracking function associated with the time it takes for the eye to track a target from 0 to angle θ. As a result, the time for the user's left eye to track a target from the first target location 1502a to the second target location 1502b can be expressed as follows. t L =|f(θ 2L )-f(θ 1L )| (1) Similarly, the time for the user's right eye to track a target from the first target location 1502a to the second target location 1502b can be expressed as follows. t R =|f(θ 2R )-f(θ 1R )| (2)

[0221] In some embodiments, the eye may take a certain amount of time to focus on a target at a distance from the eye. In some examples, the eye may take longer to focus between the target and a location having a greater distance difference from the eye. g(ξ) may represent a function that measures the shortest time for the eye to focus from infinity to ξ = 1 / d, where d is the distance from the eye. Then, the time for the left eye to focus on a target that is moved from the first target location 1502a to the second target location 1502b can be expressed as follows. t d_L =|g(1 / d 2L ) - g(1 / d 1L )| (3) Similarly, the time for the right eye to focus on a target that is moved from the first target location 1502a to the second target location 1502b can be expressed as follows. t d_R =|g(1 / d 2R ) - g(1 / d 1R )| (4)

[0222] In some embodiments, the determined time for the user to track a target from the first target location 1502a to the second target location 1502b may take into account the amount of time for the user's left and right eyes to angularly move from the first target location 1502a to the second target location 1502b, and the amount of time for the user's left and right eyes to change the focus adjustment and focus on the target that is moved from the first target location 1502a to the second target location 1502b. For example, the target tracking time may be the maximum value of the relevant angular movement and focus time. t max = max(t L , t R , t d_L , t d_R ) (5)

[0223] In some examples, the target tracking time may be a different function of the angular movement and focus time. In some examples, the target tracking time may be a function of less or more input time related to eye tracking and / or focusing on the target. H. Exemplary Virtual Object Tracking Function

[0224] The virtual object or target tracking function f(θ) may be determined based on one or more parameters or assumptions associated with the eyes of the user or general population. Assumptions that may be utilized to determine the target tracking function may include that the time required for the eye to track the target from a first location to a second location monotonically increases. Another assumption that may be utilized to determine the target tracking function may include that time will be proportional to the angle θ for small angles of θ. Another assumption that may be utilized to determine the target tracking function may include that at larger angles, there is more distortion for the eye to track the target. Thus, the assumption may include that longer times may be required for tracking than would be considered in a simple linear relationship.

[0225] Figures 18A and 18B illustrate an exemplary target tracking function f(θ) in graph 1800 and its derivative f'(θ) in graph 1801, respectively. In some embodiments, the target tracking function f(θ) may have behavior similar to tan(θ) and its derivative sec2(θ). Thus, in some embodiments, the display system may have a first approximation of the target tracking function f(θ) as tan(θ).

[0226] In some embodiments, the time for the eye to change focus between two locations at different distances to the user may be approximated to be approximately proportional to the reciprocal of the distance between the two locations. In other words, the eye focus eye function g(ξ) may be inversely proportional to the distance or linearly proportional to ξ.

[0227] In some embodiments, the approximation of the target tracking function f(θ) or the focus function g(ξ) may be updated or adjusted based on parameters associated with the user. For example, the adjustment parameter may adjust the target tracking function to more closely model the comfortable pace of a particular user's eye movement and / or the duration for a change in focus. The adjustment parameter may be determined based on any number of observations, tests, or approximated parameters associated with the user's ability to track the target. For example, one or more adjustment parameters may be determined based on meta-information associated with the user, such as age, gender, eye health, or other parameters associated with the user's eyes. If the user's meta-information is associated with slower eye tracking, one or more adjustment parameters for the target tracking function may appropriately decelerate the target movement. If the user's meta-information is associated with faster eye tracking, one or more adjustment parameters for the target tracking function may appropriately accelerate the target movement. In some embodiments, the user may be categorized into one or more types of the user based on meta-information such as fast-tracking eyes, normal-tracking eyes, or slow-tracking eyes. However, other categorizations are also possible. For example, the user may be categorized into a mixed category such as fast-angle tracking and normal-focusing eyes or slow-angle tracking and normal-focusing eyes or other categorizations.

[0228] In some embodiments, the display system may determine the eye tracking time t based on the measurement. For example, by inverting t = f(θ), the display system may obtain θ = finv(t). The angular velocity ω may be obtained by finding the derivative of the inverse function finv(t) with respect to t. ω = P(θ) = df inv (t) / dt (6)

[0229] Figure 19 illustrates an exemplary modeled angular velocity 1900 as a function of θ, based on the inverse function. The display system may measure the angular velocity of one or more than one of the user's eyes from the user's direct front or a certain angle such as 0 or the origin to the point at the tangent angle θ. In some embodiments, the display system may measure the angular velocity of the user's eyes at incremental points from 0 to 90 degrees. However, other ranges of angles are also conceivable. The display system may measure any number of angular velocities associated with any number of tangent angles, such as three angular velocities at incremental angles from 0 to 90 degrees, five angular velocities at incremental angles from 0 to 90 degrees, seven angular velocities at incremental angles from 0 to 90 degrees, or any other combination of angular velocities, incremental angles, or ranges of angles.

[0230] The display system may adapt the modeled angular velocity based on the measured angular velocity. Since ω = dθ / dt, the display system may determine the time t = f(θ) relationship as follows.

Chemical formula

[0231] Over a certain period of virtual object travel time, such as the travel time or virtual object tracking time determined above, the display system may move the virtual object at a variable speed. In some embodiments, the display system may move the virtual object from an initial (or first) target location to a new (or second) target location at a pace approximating the natural and comfortable movement of the human eye. For example, the display system may move the virtual object (also referred to as the target) from the initial target location to the new target location with a relatively high initial speed and a deceleration as it reaches the second location.

[0232] Advantageously, as opposed to a constant speed, by utilizing a variable speed, movement with respect to nearby virtual objects in angular space and / or depth requires less time and thus can shorten the overall calibration time during a calibration session that utilizes a variable speed. For example, the maximum speed and average speed of variable speed movement can be higher than a constant speed, and / or overshooting can be avoided. Additionally, the time for movement may be adjusted according to the angular and / or distance movement of the virtual object. This can enable the display system to give surplus time for the user's comfort and remove unnecessary time as needed. Additionally, as the user approaches a predetermined location, the user can gain predictive ability at the final destination through deceleration of the virtual object movement. Advantageously, this can improve the user's comfort when tracking a virtual object and / or reduce overshooting of the virtual object as the predetermined location is reached, thus shortening the overall calibration time during the calibration session.

[0233] In some embodiments, the display system may estimate the position of a virtual object or target as a function of time through interpolation. The display system may interpolate the coordinates of the virtual object as a function of time based on one or more parameters associated with a variable speed. FIG. 20 illustrates an exemplary interpolation 2000. The interpolation function 2000 may be associated with an S-curve that may have interpolation parameters t and s corresponding to the horizontal and vertical axes, respectively. The interpolation maps the virtual object to an initial (or first) target location and a first time (s = 0 and t = 0, respectively) and a new (or second) target location and a second time (scaled to s = 1 and t = 1). As t changes uniformly, s changes non-uniformly. For example, s may change slowly, change more rapidly, then change slowly again and approach 1 as t changes uniformly from 0 to 1. In some embodiments, the behavior of s with respect to t may be understood as a slow and gradual initial increase up to the maximum speed of the movement of the virtual object at the initial target position, and then a stepwise and gradual decrease in speed as the virtual object comes to rest at the new target location. The relationship between t and s may be expressed as follows. s = B(t) Equation (8) where B(t) is a function determined by a curve as illustrated in FIG. 20.

[0234] Any number of curves, such as one or more spline curves, may be used for the interpolation function. For example, one or more spline curves may include one or more cubic Bezier spline curves, Hermite spline curves, or Cardinal spline curves. In the embodiment illustrated in FIG. 20, two cubic Bezier curves are used for the interpolation function. Point 2004 illustrates exemplary curve points at uniform intervals. Control points 2002 for one or more curves of the interpolation may be adjusted by the display system to smoothly connect one or more curves.

[0235] Since the eye provides more attention when the object is closer, it is more appropriate to approximate the interpolation simply according to the reciprocal of the distance rather than the distance. Therefore, the z - coordinate of the point interpolated with the parameter s can be approximated as follows. 1 / z=(1 - s)1 / z1 + s1 / z2 (Equation 9) Where z1 and z2 are the z - coordinates of the initial target location and the new target location, respectively. In some embodiments, in a display with a finite number of depth planes, z1 and z2 can correspond to the depth planes on which the initial target location and the new target location are placed. Even if z1 and z2 are intended to be perceived as being at different z - distances, in some embodiments where there are a limited number of depth planes such that they exist with wavefront divergence corresponding to the same depth plane, z1 and z2 can be understood to be equal. This can occur because it is expected that the same amount of wavefront divergence provides the same accommodation state such that the eye will not need to change its accommodation state between z1 and z2.

[0236] When the linear interpolation parameter u is defined as follows, z=(1 - u)z1 + uz2 (Equation 10) The linear parameter is obtained as follows.

Chemical formula

[0237] In some embodiments, the uniform parameter t may be based on the ratio of the time between the maximum determination time for tracking the virtual object and the measured actual time for the user to track the virtual object. For example, the maximum time required to move the virtual object from the initial target location to the new target location is t_max, as defined by equation (4). The actual time t_real for the user to track the virtual object moving from the initial target location to the new target location can be used in combination with t_max to determine the uniform parameter t.

Chemical formula

[0238] The interpolation parameters s and u may then be determined from the uniform parameter t. The point or position of the virtual object may then be approximated or determined based on interpolation using the obtained interpolation parameter u.

[0239] In some examples, the interpolation and / or the curve associated with the interpolation may be adjusted or regulated based on one or more parameters. In some examples, one or more parameters may vary the rate of increase, decrease, peak location, or other aspects of the speed of the virtual object as it moves from the first location to the second location. In some examples, the parameters may be adjusted for each user. In some examples, the parameters may be adjusted based on the population tendency or category of users within the population group. In some examples, one or more parameters are set to default and may be updated based on the user's usage or eye-tracking AR system measurements or other tracking parameters.

[0240] In some embodiments, one or more parameters may be personalized for a user or group of users. For example, an eye virtual object movement system may utilize collected eye movement data to analyze how people move their eyes within a test framework. The system may then update eye virtual object movement and help reduce sources of error in eye calibration such as overshoot based on the collected eye movement data. In some embodiments, the test framework may include angular velocity, focus velocity, eye movement range of motion, or other parameters related to eye tracking. In some embodiments, the test framework may include steps of measuring one or more parameters to help reduce measurement error. For example, a display system may measure a user's eye tracking parameters in conjunction with secondary user input related to calibration. The secondary user input may include, for example, pressing a virtual button indicating that the user's eyes are in a stationary state. Advantageously, the secondary user input may, for example, enable a display system to more accurately identify the total time for which a user tracks a virtual object moving than with eye tracking alone. J. Additional aspects

[0241] Included herein are additional aspects of a head-mounted display system and associated methods. Any of the aspects or embodiments disclosed herein may be combined, in whole or in part.

[0242] In a first aspect, a head-mounted display system is disclosed. The head-mounted display system includes a display configured to present virtual content to a user, and a hardware processor in communication with the display, the hardware processor programmed to cause the display system to display a virtual object at a first location and to move the virtual object at a variable speed to a second location based on an S-shaped velocity curve.

[0243] In the second aspect, which is the head-mounted display system described in the first aspect, one or more hardware processors determine the diopter distance between a first location and a second location, determine the distance between the focal plane of the first location and the focal plane of the second location, and are configured to determine the total time for moving a virtual object from the first location to the second location based at least in part on the diopter distance and the distance between the focal plane of the first location and the focal plane of the second location.

[0244] In the third aspect, which is the head-mounted display system according to any one of the first or second aspects, in order to determine the total time, one or more hardware processors determine at least one tracking time associated with at least one eye of the user based on the diopter distance, determine at least one focusing time associated with at least one eye of the user based on the focal plane distance, and are configured to select the total time from at least one tracking time and at least one focusing time.

[0245] In the fourth aspect, which is the head-mounted display system according to any one of the first to third aspects, at least one tracking time comprises the time for at least one eye of the user to move angularly over the diopter distance.

[0246] In the fifth aspect, which is the head-mounted display system according to any one of the first to fourth aspects, at least one tracking time comprises a tracking time associated with the left eye and a tracking time associated with the right eye.

[0247] In the sixth aspect, which is the head-mounted display system according to any one of the first to fifth aspects, at least one focusing time comprises the time for at least one eye of the user to focus over the distance between the focal plane of the first location and the focal plane of the second location.

[0248] In the seventh aspect, which is the head-mounted display system according to any one of aspects 1-6, at least one focusing time includes a focusing time associated with the left eye and a focusing time associated with the right eye.

[0249] In the eighth aspect, which is the head-mounted display system according to any one of aspects 1-7, one or more hardware processors are configured to determine a variable speed based on a total time.

[0250] In the ninth aspect, which is the head-mounted display system according to any one of aspects 1-8, the variable speed is based on the reciprocals of the diopter distances at the first and second locations.

[0251] In the tenth aspect, which is the head-mounted display system according to any one of aspects 1-9, one or more hardware processors are configured to interpolate the position of a virtual object as a function of time based on one or more parameters associated with the total time or the diopter distance.

[0252] In the eleventh aspect, which is the head-mounted display system according to any one of aspects 1-10, one or more parameters include uniform and non-uniform parameters, and the non-uniform parameters are configured to vary in value at a variable rate associated with one or more spline curves as a function of the uniform parameters.

[0253] In the twelfth aspect, which is the head-mounted display system according to any one of aspects 1-11, the uniform parameter is at least partially based on the total time.

[0254] In the thirteenth aspect, which is the head-mounted display system according to any one of aspects 1-12, an image capture device is provided that is configured to capture an eye image of one or both eyes of a user of the wearable system.

[0255] In the 14th aspect, which is the head-mounted display system according to any one of aspects 1-13, the virtual content includes an eye calibration target.

[0256] In the 15th aspect, a method for moving virtual content is disclosed. The method includes the steps of displaying a virtual object at a first location on the display system and moving the virtual object to a second location at a variable speed following an S-shaped speed curve on the display system.

[0257] In the 16th aspect, which is the method according to aspect 15, the method includes the steps of determining the diopter distance between the first location and the second location, determining the distance between the focal plane of the first location and the focal plane of the second location, and determining the total time for moving the virtual object from the first location to the second location based at least in part on the diopter distance and the distance between the focal plane of the first location and the focal plane of the second location.

[0258] In the 17th aspect, which is the method according to any one of aspects 15 or 16, the step of determining the total time includes determining at least one tracking time associated with at least one eye of the user based on the diopter distance, determining at least one focusing time associated with at least one eye of the user based on the focal plane distance, and selecting the total time from at least one tracking time and at least one focusing time.

[0259] In the 18th aspect, which is the method according to any one of aspects 15-17, at least one tracking time includes the time for at least one eye of the user to move angularly over the diopter distance.

[0260] In the 19th aspect, which is the method according to any one of aspects 15-18, at least one tracking time includes a tracking time associated with the left eye and a tracking time associated with the right eye.

[0261] In a 20th aspect, which is the method according to any one of aspects 15 - 19, at least one focusing time comprises a time for a user's at least one eye to focus over a distance between a focal plane at a first location and a focal plane at a second location.

[0262] In a 21st aspect, which is the method according to any one of aspects 15 - 20, at least one focusing time comprises a focusing time associated with the left eye and a focusing time associated with the right eye.

[0263] In a 22nd aspect, which is the method according to any one of aspects 15 - 21, the method includes a step of determining a variable speed based on a total time.

[0264] In a 23rd aspect, which is the method according to any one of aspects 15 - 22, the variable speed is based on the reciprocal of the dioptric distance.

[0265] In a 24th aspect, which is the method according to any one of aspects 15 - 23, the method includes a step of interpolating the position of a virtual object as a function of time, and the step of interpolating the position of the virtual object as a function of time includes a step of interpolating an S - curve function using one or more parameters based on the total time or the dioptric distance.

[0266] In a 25th aspect, which is the method according to any one of aspects 15 - 24, one or more parameters comprise uniform and non - uniform parameters, and the non - uniform parameters are configured to vary in value at a variable rate associated with one or more spline curves as a function of the uniform parameters.

[0267] In a 26th aspect, which is the method according to any one of aspects 15 - 25, the uniform parameters are at least partially based on the total time.

[0268] In a 27th aspect, which is the method according to any one of aspects 15 - 26, the virtual content comprises an eye calibration target.

[0269] In a 28th aspect, a wearable system for eye tracking calibration is disclosed. The system comprises a display configured to display an eye calibration target to a user, and a hardware processor that communicates with the non - transitory memory and the display system, the hardware processor being configured to cause the display to display the eye calibration target at a first target location, identify a second location different from the first target location, determine a distance between the first target location and the second target location, determine a total allocation time for moving the eye calibration target from the first target location to the second target location, calculate a target movement speed curve based on the total allocation time, the target movement speed curve being an S - curve, and being programmed to move the eye calibration target over a total time according to the target movement speed curve.

[0270] In a 29th aspect, which is the wearable system according to aspect 28, the hardware processor is configured to identify a user's eye line of sight based on data obtained from an image capture device, determine whether the user's eye line of sight is aligned with the eye calibration target at the second target location, and in response to the determination that the user's eye line of sight is aligned with the eye calibration target at the second location, instruct the image capture device to capture an eye image and start storing the eye image into the non - transitory memory.

[0271] In a 30th aspect, a method for eye tracking calibration is disclosed. The method includes the steps of displaying an eye calibration target at a first target location within a user's environment, identifying a second location different from the first target location, determining a distance between the first target location and the second target location, determining a total allocation time for moving the eye calibration target from the first target location to the second target location, calculating a target movement speed curve based on the total allocation time, the target movement speed curve being an S - curve, and moving the eye calibration target over a total time according to the target movement speed curve.

[0272] In a 31st aspect, which is the method described in the 30th aspect, the method includes: identifying a user's eye gaze based on data obtained from an image capture device; determining whether the user's eye gaze aligns with an eye calibration target at a second target location; and in response to determining that the user's eye gaze aligns with the eye calibration target at the second location, instructing the image capture device to capture an eye image and initiate storage of the eye image into a non-transitory memory.

[0273] In a 32nd aspect, a wearable system for eye tracking calibration is described. The system includes: an image capture device configured to capture an eye image of one or both eyes of a user of the wearable system; a non-transitory memory configured to store the eye image; a display system through which the user can perceive an eye calibration target within the user's environment; and a hardware processor that communicates with the non-transitory memory and the display system, the hardware processor being programmed to make the eye calibration target perceivable at a first target location within the user's environment via the display system, identify a second target location within the user's environment that is different from the first target location, determine a diopter distance between the first target location and the second target location, determine a total time for moving the eye calibration target from the first target location to the second target location based on the distance, at least partially interpolate a target position based on the reciprocal of the diopter distance, and move the eye calibration target as a function of time over the total time according to the interpolated target position.

[0274] In a 33rd aspect, which is the wearable system described in aspect 32, the hardware processor is configured to identify a user's eye gaze based on data obtained from an image capture device, determine whether the user's eye gaze aligns with an eye calibration target at a second target location, and in response to determining that the user's eye gaze aligns with the eye calibration target at the second location, instruct the image capture device to capture an eye image and begin storing the eye image in a non-transitory memory. K. Conclusion

[0275] The processes, methods, and algorithms described herein and / or depicted in the accompanying figures are each embodied in one or more physical computing systems, hardware computer processors, application-specific circuits, and / or electronic hardware configured to execute specific and particular computer instructions, thereby being fully or partially automated. For example, a computing system can include a general-purpose computer (e.g., a server) or a dedicated computer, a dedicated circuit, etc., programmed with specific computer instructions. Code modules can be installed in dynamic link libraries, can be compiled and linked into executable programs, or can be written in an interpreted type programming language. In some implementations, certain operations and methods can be performed by circuits specific to a given function.

[0276] Furthermore, the functional implementations of the present disclosure are sufficiently mathematically, computationally, or technically complex that a special-purpose hardware (utilizing appropriate specialized executable instructions) or one or more physical computing devices may be required to implement the functionality, e.g., due to the amount or complexity of the calculations involved or to provide the results substantially in real time. For example, a video may contain many frames, each frame may have millions of pixels, and specifically programmed computer hardware is required to process the video data to provide the desired image processing tasks or applications in a commercially reasonable amount of time. As another example, embodiments of the eye-tracking calibration techniques described herein may need to be implemented in real time while a user is wearing a head-mounted display system.

[0277] A code module or any type of data may be stored on any type of non-transitory computer-readable medium such as a physical computer storage device including a hard drive, solid-state memory, random access memory (RAM), read-only memory (ROM), optical disk, volatile or non-volatile storage device, combinations of the same, and / or equivalents. The methods and modules (or data) may also be transmitted as data signals generated on various computer-readable transmission media including wireless-based and wired / cable-based media (e.g., as part of a carrier wave or other analog or digital propagated signal) and may take various forms (e.g., as part of a single or multiplexed analog signal or as multiple discrete digital packets or frames). The results of the disclosed process or process steps may be persistently or otherwise stored within any type of non-transitory tangible computer storage device or communicated via a computer-readable transmission medium.

[0278] Any process, block, state, step, or functionality in a flowchart described herein and / or depicted in the accompanying figures is to be understood as potentially representing a code module, segment, or portion of code that includes one or more executable instructions for implementing a specific function (e.g., logical or arithmetic) or step in a process. The various processes, blocks, states, steps, or functionality may be combined, rearranged, added to, deleted from, modified from, or otherwise changed from the illustrative embodiments provided herein. In some embodiments, additional or different computing systems or code modules may implement some or all of the functionality described herein. The methods and processes described herein are also not limited to any particular sequence, and the associated blocks, steps, or states may be performed in other sequences that are appropriate, e.g., sequentially, in parallel, or in some other manner. Tasks or events may be added to or removed from the disclosed illustrative embodiments. Further, the separation of the various system components in the implementations described herein is for illustrative purposes and should not be understood as requiring such separation in all implementations. It should be understood that the described program components, methods, and systems may generally be integrated together in a single computer product or packaged in multiple computer products. Many implementation variations are also possible.

[0279] The present process, method, and system may be implemented in a network (or distributed) computing environment. The network environment may include an enterprise-wide computer network, an intranet, a local area network (LAN), a wide area network (WAN), a personal area network (PAN), a cloud computing network, a cloud source computing network, the Internet, and the World Wide Web. The network may be a wired or wireless network or any other type of communication network.

[0280] The systems and methods of the present disclosure each have several innovative aspects, none of which alone contribute to or are required for the desirable attributes disclosed herein. The various features and processes described above may be used independently of each other or combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. Various modifications to the implementations described in this disclosure may be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations without departing from the spirit or scope of the present disclosure. Accordingly, the claims are not intended to be limited to the implementations shown herein, but are to be accorded the widest scope consistent with the present disclosure, the principles, and the novel features disclosed herein.

[0281] In the context of separate implementations, certain features described herein may also be implemented in combination in a single implementation. Conversely, the various features described in the context of a single implementation may also be implemented separately in multiple implementations or in any suitable sub-combination. Further, features may be described above as acting in a certain combination and may further be initially claimed as such, but one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination. No single feature or group of features is necessary or essential to every embodiment.

[0282] In particular, conditional clauses used herein such as "can", "could", "might", "may", "e.g.", and the like, generally convey that while one embodiment includes certain features, elements, and / or steps, other embodiments do not, unless specifically stated otherwise or understood otherwise within the context in which they are used. Thus, such conditional clauses generally do not imply that features, elements, and / or steps are required in any way for one or more embodiments, or that one or more embodiments necessarily include logic for determining whether these features, elements, and / or steps should be included or implemented in any particular embodiment, regardless of the author's input or prompting. The terms "comprising", "including", "having", and the like are synonyms and are used inclusively in a non-limiting manner, and do not exclude additional elements, features, acts, operations, etc. Also, the term "or" is used in its inclusive sense (and not in its exclusive sense), and thus, for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. Additionally, as used in this application and the appended claims, the articles "a", "an", and "the" should be construed to mean "one or more than one" or "at least one" unless otherwise defined.

[0283] As used herein, the phrase referring to a list of items "at least one of" refers to any combination of those items, including a single element. As an example, "at least one of A, B, or C" is intended to cover A, B, C, A and B, A and C, B and C, and A, B, and C. Connective phrases such as "at least one of X, Y, and Z" are generally understood in a context such that they are used to convey that an item, term, etc. can be at least one of X, Y, or Z, unless specifically stated otherwise. Thus, such connective phrases generally are not intended to imply that an embodiment requires the presence of at least one of each of X, at least one of each of Y, and at least one of each of Z, respectively.

[0284] Similarly, operations may be depicted in the drawings in a particular order, but it should be recognized that such operations are not necessarily performed in the particular order shown, or in a sequential order, or that all illustrated operations are performed, to achieve a desired result. Additionally, the drawings may schematically depict one or more exemplary processes in the form of a flowchart. However, other operations not depicted may also be incorporated within the exemplary methods and processes schematically illustrated. For example, one or more additional operations can be performed before, after, simultaneously with, or between any of the illustrated operations. In addition, the operations may be rearranged or reordered in other implementations. In some situations, multitasking and parallel processing may be advantageous. Further, the separation of various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the program components and systems described generally may be integrated together in a single software product or packaged into multiple software products. Additionally, other implementations are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve a desired result.

Claims

1. A head-mounted display system, wherein the head-mounted display system comprises: a display configured to present virtual content to a user; one or more hardware processors in communication with the display and comprising: the one or more hardware processors are configured to: cause the display system to display a virtual object at a first location; determine a diopter distance between the first location and a second location; determine a distance between a focal plane of the first location and a focal plane of the second location; determine at least one tracking time associated with at least one eye of the user based on the diopter distance, the at least one tracking time including a time for the at least one eye of the user to angularly move over the diopter distance; determine at least one focusing time associated with at least one eye of the user based on the distance between the focal plane of the first location and the focal plane of the second location, the at least one focusing time including a time for the at least one eye of the user to focus over the distance between the focal plane of the first location and the focal plane of the second location; determine a total time for moving the virtual object from the first location to the second location based at least in part on the diopter distance and the distance between the focal plane of the first location and the focal plane of the second location, the total time being selected from the at least one tracking time and the at least one focusing time; cause the display system to move the virtual object to the second location at a variable speed based on an S-shaped speed curve representing the speed behavior of the movement of the virtual object over time A head-mounted display system programmed to perform the above.

2. The head-mounted display system according to claim 1, wherein the at least one tracking time includes a tracking time associated with the left eye of the user and a tracking time associated with the right eye of the user.

3. The at least one focusing time includes a focusing time associated with the user's left eye and a focusing time associated with the user's right eye, the head-mounted display system according to claim 1.

4. The one or more hardware processors are configured to determine the variable speed based on the total time, the head-mounted display system according to claim 1.

5. The one or more hardware processors are configured to interpolate the position of the virtual object as a function of time based on one or more parameters associated with the total time or the diopter distance, the head-mounted display system according to claim 1.

6. The variable speed is based on the reciprocal of the diopter distance, the head-mounted display system according to claim 1.

7. The one or more hardware processors are configured to interpolate the position of the virtual object based on the equation P(t) = (1 - u)P1 + uP2, where P1 and P2 are the coordinates of the virtual object at the first location and the second location, and the one or more parameters include a uniform parameter t and a linear parameter u, the uniform parameter t is based on the time ratio between the maximum time determined to track the virtual object and the time actually measured by the user to track the virtual object, the linear parameter u is an interpolation parameter related to the z coordinates of the initial target location and the new target location, and the linear parameter u is configured to change its value at a variable rate associated with one or more spline curves as a function of the uniform parameter t, the head-mounted display system according to claim 5.

8. The uniform parameter t is at least partially based on the total time, the head-mounted display system according to claim 7.

9. The head-mounted display system further comprises an image capture device configured to capture an eye image of at least one or both eyes of the user, the head-mounted display system according to claim 1.

10. The virtual content includes an eye calibration target, the head-mounted display system according to claim 1.

11. A method for moving virtual content, the method comprising: causing a virtual object to be displayed to a user at a first location on a display system; determining a diopter distance between the first location and a second location; determining a distance between a focal plane of the first location and a focal plane of the second location; determining at least one tracking time associated with at least one eye of the user based on the diopter distance, the at least one tracking time including a time for the at least one eye of the user to angularly move over the diopter distance; determining at least one focusing time associated with at least one eye of the user based on the distance between the focal plane of the first location and the focal plane of the second location, the at least one focusing time including a time for the at least one eye of the user to focus over the distance between the focal plane of the first location and the focal plane of the second location; determining a total time for moving the virtual object from the first location to the second location based at least in part on the diopter distance and the distance between the focal plane of the first location and the focal plane of the second location, the total time being selected from the at least one tracking time and the at least one focusing time; causing the display system to move the virtual object to the second location at a variable speed following an S-shaped speed curve representing the speed behavior of the movement of the virtual object over time A method comprising the above.

12. The method according to claim 11, wherein the at least one tracking time includes a tracking time associated with the left eye of the user and a tracking time associated with the right eye of the user.

13. The method according to claim 11, wherein the at least one focusing time includes a focusing time associated with the left eye of the user and a focusing time associated with the right eye of the user.

14. The method according to claim 11, further comprising determining the variable speed based on the total time.

15. The method according to claim 11, wherein the variable speed is based on the reciprocal of the diopter distance.

16. The method further includes interpolating the position of the virtual object as a function of time, and interpolating the position of the virtual object as a function of time is based on one or more parameters associated with the total time or the dioptric distance. The method according to claim 11.

17. The interpolating includes interpolating the position of the virtual object based on the formula P(t) = (1 - u)P1 + uP2, where P1 and P2 are the coordinates of the virtual object at the first location and the second location, and the one or more parameters include a uniform parameter t and a linear parameter u. The uniform parameter t is based on the time ratio between the maximum time determined for tracking the virtual object and the time actually measured by the user for tracking the virtual object. The linear parameter u is an interpolation parameter related to the z coordinates of the initial target location and the new target location, and the linear parameter u is configured to change its value at a variable rate associated with one or more spline curves as a function of the uniform parameter t. The method according to claim 16.

18. The uniform parameter t is at least partially based on the total time. The method according to claim 17.

19. The virtual content includes an eye calibration target. The method according to claim 11.

Citation Information

Patent Citations

  • Dynamic display calibration based on eye tracking

    JP2019501564A

  • Interacting with 3D Virtual Objects Using Attitude and Multiple DOF Controllers

    JP2019517049A

  • Eye tracking calibration techniques

    US20180348861A1

  • Information processing device, information processing method, and program

    WO2017022302A1