Visual axis registration

By displaying readable text in an extended reality system and using the eye tracking camera to estimate the visual axis, the problem of users needing explicit actions in traditional methods is solved, and interference-free and efficient visual axis registration is achieved.

CN120359487APending Publication Date: 2025-07-22APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380086198.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-08
Filing Date
2023-12-21
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the existing extended reality system, the visual axis registration process requires the user to make explicit actions, resulting in strong invasiveness and inability to complete the visual axis estimation and registration without interference.

Method used

By displaying a line of text that can be read by the user at a known vertical position and virtual depth, the user's image is captured using the eye tracking camera, estimating the stimulus plane, and calculating the view axis by error, avoiding explicit prompts.

Benefits of technology

A interference-free and more efficient visual axis registration process is achieved, reducing additional steps and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359487A_ABST
    Figure CN120359487A_ABST
Patent Text Reader

Abstract

During interference-free optical axis registration, a row of text or other content that is subsequently readable by a user is displayed at a known vertical position and virtual depth. The line text may be content that a user needs to read as part of a normal enrollment process. The eye tracking camera may capture an image of the eye as the user reads the line of text. This data may then be used to estimate a stimulation plane. The k-angle may then be estimated using an error between the estimated stimulus plane and a ground live stimulus plane (the actual position of the line text in the virtual space).
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Extended reality (XR) systems such as mixed reality (MR) or augmented reality (AR) systems combine computer-generated information, referred to as virtual content, with real-world images or real-world views to enhance or augment a user's perception of the world. Thus, XR systems can be utilized to provide an interactive user experience for a variety of applications, such as applications that add virtual content to a live view of the viewer's environment, applications that interact with a virtual training environment, gaming applications, applications for remotely controlling a drone or other mechanical system, applications for viewing digital media content, or applications for interacting with the Internet, among others. Summary of the Invention

[0002] Various embodiments of methods and apparatuses for performing visual axis registration on a device (e.g., a head-mounted device (HMD), including but not limited to an HMD used in extended reality (XR) applications and systems) are described. The HMD can include a wearable device such as a head-mounted unit, helmet, goggles, or glasses. The XR system can include an HMD that can include one or more cameras that can be used to capture static images or video frames of the user's environment. The HMD can include lenses positioned in front of the eyes through which the wearer can view the environment. In the XR system, virtual content can be displayed on or projected onto these lenses such that the virtual content is visible to the wearer while still being able to view the real environment through the lenses.

[0003] In at least some systems, the HMD can include gaze tracking technology. Note that gaze tracking can be performed on one eye or on both eyes. In such systems, during an initial calibration or registration process, a multi-dimensional personalized model of the eye can be generated based on one or more images of the user's eyes captured by an eye tracking camera. This can be done unobtrusively without prompting as the user moves their eyes. The personalized eye model generated based on images captured as the eye moves can include information such as a corneal surface model, iris and pupil models, eye center, entrance pupil, and the pupil axis or optical axis (a vector passing through the geometric center of the eye and the entrance pupil).

[0004] However, the actual gaze direction of the eye corresponds to the visual axis, which is offset from the calculated optical axis of the eye model. Thus, another part of the calibration or registration process is to estimate the visual axis, or the k angle between the optical axis and the visual axis.

[0005] To estimate the k-angle / line of sight, it may be necessary to have at least one ground truth point that the user is viewing, where additional images captured by an eye-tracking camera are used to estimate the k-angle and thus the line of sight. Conventionally, a point is displayed on an interface and the user is prompted to look at that point. However, this conventional method is intrusive; the user must be prompted to do something (look at the displayed point) to fully register the eye.

[0006] Embodiments of methods and apparatuses for unobtrusively and non-intrusively estimating and registering the line of sight of an eye are described. In an embodiment, instead of displaying a point and prompting the user to look at that point, a line of text (or other content) that the user can subsequently read is displayed at a known vertical position and virtual depth. The line of text can be something that the user is required to read as part of the registration process, such as "Please read and agree to the following terms of use". The line of text can be the only content being displayed at that time (or at a known virtual depth). In other words, in some embodiments, only a single line of horizontal content such as text needs to be displayed. When the user reads the line of text, the eye-tracking camera can capture an image of the eye. This data can then be processed to estimate the stimulation plane. Then, the error between the estimated stimulation plane and the ground truth stimulation plane (the actual position of the line of text in virtual space) can be determined and used to estimate the k-angle and thus the true line of sight of the eye.

[0007] This method of estimating the line of sight avoids the extra step in the eye registration process and is thus more efficient and less intrusive than previous methods. The user reads a line of text and the eye model and line of sight are registered. There is no need for an explicit eye registration process with a prompt to the user.

[0008] Then, during use of the device, the personalized eye model and the estimated k-angle / line of sight can be used in various algorithms, such as in the gaze estimation process for a gaze-based interface. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 An N-dimensional model of an eye is graphically illustrated according to some embodiments.

[0010] Figure 2 A non-intrusive line of sight registration process is generally illustrated according to some embodiments.

[0011] Figure 3 An eye registration process is graphically illustrated according to some embodiments.

[0012] Figure 4 A non-intrusive line of sight registration process is illustrated in more detail according to some embodiments.

[0013] Figure 5A is a high-level flowchart of an eye registration process according to some embodiments.

[0014] Figure 5B is a high - level flowchart of a non - intrusive line - of - sight registration process according to some embodiments.

[0015] Figures 6A to 6C Illustrates an example device in which a method according to some embodiments can be implemented Figures 1 to 5B is exemplified.

[0016] Figure 7 is a block diagram illustrating an example device according to some embodiments. The example device can include components as shown Figures 1 to 5B and implement the method illustrated in these figures.

[0017] This specification includes references to "one embodiment" or "embodiments". The phrases "in one embodiment" or "in embodiments" do not necessarily refer to the same embodiment. Particular features, structures, or characteristics can be combined in any suitable manner consistent with this disclosure.

[0018] "Comprising", this term is open - ended. As used in the claims, this term does not exclude additional structures or steps. Consider the following cited claim: "An apparatus comprising one or more processor units...". Such a claim does not exclude the apparatus from including additional components (e.g., a network interface unit, graphics circuitry, etc.).

[0019] "Configured to", various units, circuits, or other components can be described or recited as "configured to" perform one or more tasks. In such contexts, "configured to" is used to imply that the structure (e.g., circuitry) by indicating that the unit / circuit / component includes the structure that performs these one or more tasks during operation. Thus, the unit / circuit / component is alleged to be configured to perform the task even when the specified unit / circuit / component is currently inoperable (e.g., not powered on). Units / circuits / components used in conjunction with "configured to" language include hardware - such as circuits, memories storing program instructions capable of being executed to implement the operations, etc. Referring to a unit / circuit / component "configured to" perform one or more tasks is specifically intended not to invoke paragraph (f) of 35 U.S.C.§112 for that unit / circuit / component. Additionally, "configured to" can include a general structure (e.g., a general circuit) manipulated by software or firmware (e.g., an FPGA or a general - purpose processor executing software) to operate in a manner capable of performing the one or more tasks to be solved. "Configured to" can also include adjusting a manufacturing process (e.g., a semiconductor fabrication facility) to fabricate a device (e.g., an integrated circuit) suitable for implementing or performing the one or more tasks.

[0020] "First", "second", etc. As used herein, these terms serve as labels for the nouns that precede them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.). For example, a buffer circuit may be described herein as performing a write operation on a "first" value and a "second" value. The terms "first" and "second" do not necessarily imply that the first value must be written before the second value.

[0021] "Based on" or "depending on", as used herein, these terms are used to describe one or more factors that affect a determination. These terms do not exclude additional factors that may affect the determination. That is, the determination may be based solely on these factors or at least in part on these factors. Consider the phrase "determine A based on B". In this case, B is a factor that affects the determination of A, and such a phrase does not exclude that the determination of A may also be based on C. In other instances, A may be determined based solely on B.

[0022] "Or", when used in a claim, the term "or" is used as an inclusive or rather than an exclusive or. For example, the phrase "at least one of x, y, or z" means any one of x, y, and z and any combination thereof. Detailed Description

[0023] Describes various embodiments of methods and apparatuses for performing visual axis registration on a device (e.g., a head-mounted device (HMD), including but not limited to an HMD used in extended reality (XR) applications and systems). The HMD may include a wearable device such as a head-mounted unit, helmet, goggles, or glasses. The XR system may include an HMD that may include one or more cameras that may be used to capture static images or video frames of the user's environment. The HMD may include lenses positioned in front of the eyes through which the wearer may view the environment. In an XR system, virtual content may be displayed on or projected onto these lenses such that the virtual content is visible to the wearer while still being able to view the real environment through the lenses.

[0024] In at least some systems, the HMD may include gaze tracking technology. In an example gaze tracking system, one or more infrared (IR) light sources emit IR light towards the user's eyes. A portion of the IR light is reflected from the eye and captured by an eye tracking camera. The image captured by the eye tracking camera may be input into a flash and pupil detection process, which is implemented, for example, by one or more processors of the controller of the HMD. The result of this process is passed to a gaze estimation process (which is implemented, for example, by one or more processors of the controller) to estimate the user's current gaze point. This gaze tracking method may be referred to as PCCR (pupil center corneal reflection) tracking. Note that gaze tracking may be performed on one eye or on both eyes.

[0025] In such systems, during an initial calibration or registration process, a multi-dimensional personalized model of the eye can be generated based on one or more images of the user's eye captured by an eye tracking camera as described above. Figure 1 An N-dimensional model 100 of an eye is graphically illustrated according to some embodiments. Physical components of the eye can include a sclera 102, a cornea 104, an iris 106, and a pupil 108. In some embodiments, during an initial calibration or registration process, an N-dimensional model of the user's eye 100 can be generated based on one or more images of the eye 100. In an example method, one or more infrared (IR) light sources emit IR light toward the user's eye. A portion of the IR light is reflected from the eye and captured by the eye tracking camera. Two or more images captured by the eye tracking camera can be input into an eye model generation process, which can be implemented, for example, by one or more processors of a controller of the HMD. The process can determine the shape and relationships of the components of the eye at least in part based on the localization of the glints (reflections of point light sources) in the two or more captured images. This information can then be used to generate a personalized eye model. The personalized eye model can include information such as a corneal surface model, an iris and pupil model, an eye center 112, an entrance pupil 110, and a pupil axis or optical axis 120 (a vector passing through the eye center 112 and the entrance pupil 110). During use of the device, the personalized eye model can subsequently be used in various algorithms (e.g., during a gaze estimation process).

[0026] However, the actual gaze direction of the eye corresponds to the visual axis 122, which is offset from the calculated optical axis 120 of the eye model. Thus, another part of the calibration or registration process is to estimate the visual axis 122 or the k angle 124 between the optical axis 120 and the visual axis 122.

[0027] To estimate the k angle 124 / visual axis 122, it may be necessary to have at least one ground truth point that the user is viewing, where additional images captured by the eye tracking camera are used to estimate the k angle 124 and thus the visual axis 122. Conventionally, a point is displayed on an interface and the user is prompted to look at the point. However, this conventional method is intrusive; the user must be prompted to do something (look at the displayed point) to fully register the eye.

[0028] Embodiments of methods and apparatuses for unobtrusively and non-intrusively estimating and registering the visual axis 122 of an eye are described. In an embodiment, as Figure 2Rather than showing a single point and prompting the user to look at it, as illustrated, a relatively long line of text (or other content) 230 that the user can subsequently read is displayed at a known vertical position and virtual depth as a stimulus. This line of text 230 can include content that the user needs to read as part of an enrollment process, such as "Please read and agree to the following terms of use". This line of text 230 can be the only content displayed at that time (or the only content in the vicinity of that distance), and the alphanumeric characters in the line can have a relatively small vertical size. As the user reads this line of text 230, an eye tracking camera can capture an image of the eye 290. This image data can then be processed to estimate the stimulus plane. Then, the error between the estimated stimulus plane and the ground truth stimulus plane (the actual position of the line of text in virtual space) can be used to estimate the k angle and thus the visual axis of the eye 290.

[0029] By assuming good eye convergence, only the eye pose can be used to locate the fixation data with the strongest convergence. The displayed text 230 (the stimulus) is a thin line, which allows determination of the pitch component of the k angle. Using the known depth of this line of virtual text 230, the left / right yaw angle can be estimated under the constraint that the convergence point is always at the same depth for different fixation angles.

[0030] Then, during use of the device, the personalized eye model and the estimated k angle / visual axis can be used in various algorithms, such as in the fixation estimation process for a gaze-based interface.

[0031] This method of using an unobtrusive stimulus (e.g., a line of text to read) to estimate the visual axis avoids additional steps in the eye enrollment process and is thus more efficient and less invasive than previous methods. The user reads the line of text, and the eye model and visual axis are enrolled. No explicit eye enrollment process with prompts to the user is required.

[0032] Figure 3 An eye enrollment process according to some embodiments is graphically illustrated. At 300, an N-dimensional personalized eye model 304 can be generated based on two or more images of the eye (gaze input 302) captured by an eye tracking camera of the eye. This can be done in the background during the user enrollment process after the user puts on and / or turns on the device, and can be done without specifically prompting the user to do anything specific to eye enrollment. This personalized eye model 304 can include information such as a corneal surface model, an iris and pupil model, an eye center, an entrance pupil, and a pupil axis or optical axis 308 (a vector passing through the eye center and the entrance pupil).

[0033] However, the actual gaze direction of the eye corresponds to this visual axis, which is offset from the calculated optical axis 306 of the eye model 304. Thus, at 310, the estimated visual axis 308 or the k angle between the optical axis 306 and the visual axis 308 is estimated. During this process, a relatively long line of text (or other content) that the user can subsequently read is displayed at a known vertical position and virtual depth as a stimulus. When the user reads this line of text, the eye tracking camera captures an image of the eye (gaze input 312). This gaze input 312 can then be processed to estimate the stimulus plane. Then, the error between the estimated stimulus plane and the ground truth stimulus plane (the actual position of this line of text in virtual space) can be used to estimate the k angle and thus the visual axis of the eye.

[0034] By assuming good eye convergence, only the eye pose can be used to locate the gaze data with the strongest convergence. The displayed text (stimulus) is a thin line, which allows determination of the pitch component of the k angle. Using the known depth of this line of virtual text, the left / right yaw angle can be estimated under the constraint that the convergence point is always at the same depth for different gaze angles.

[0035] Then, the personalized eye model 304 and the visual axis 308 can be used in various algorithms, such as for the gaze-based interaction functions of the device. The eye tracking camera captures images of the user's eyes during these functions, and the gaze tracking algorithm of the device uses the eye model 304 and the visual axis 308 to process these images to determine the gaze direction of the eyes, where the gaze direction is corrected from the optical axis 306 to the visual axis 308 according to the k angle.

[0036] Figure 4 The non-invasive visual axis registration process according to some embodiments is illustrated in more detail. This process can start with an initial default k angle (e.g., 0). The stimulus plane (ground truth) 430 corresponds to this line of text (or other content) displayed at a known vertical position and virtual depth, and the circle represents the convergence point. When the user reads this line of text, the stimulus plane 432 is estimated based on the gaze tracking data (using the eye model and the default k angle). Then, the error 434 between the estimated stimulus plane 432 and the ground truth stimulus plane 430 can be estimated. This error 434 can be used to estimate the true k angle.

[0037] Figure 5A is a high-level flowchart of the eye registration process according to some embodiments. As indicated at 610, an N-dimensional personalized eye model is generated. As indicated at 620, the visual axis is estimated. Figure 5B More details of the visual axis estimation process are provided. As indicated at 630, optionally, the eye model and the visual axis (k angle) can be stored. As indicated at 640, the eye model and the k angle can be used to perform gaze tracking during the use of the device.

[0038] Figure 5B is a high - level flowchart of an unobtrusive line - of - sight registration process according to some embodiments. The process can start with an initial default k - angle (e.g., 0). As indicated at 622, a line of virtual text can be displayed as a stimulus at a known virtual depth and height. This line of text can be the only content being displayed at that time (or at the known virtual depth). In other words, in some embodiments, only a single horizontal line of content such as text needs to be displayed. As indicated at 624, when the user reads the line of text, an eye - tracking camera captures an image of the eye; eye - pose data is generated based on the captured image data. The stimulus plane is estimated based on the gaze - tracking data (using an eye model and the default k - angle). As indicated at 626, the k - angle is estimated at least in part based on data generated from the pose of the eye when the user reads the text. Then, the error 434 between the estimated stimulus plane and the ground - truth stimulus plane (the line of text) can be estimated; this error can then be used to estimate the true k - angle. The k - angle indicates the true line of sight of the eye.

[0039] Although embodiments are generally described and illustrated with reference to one eye, there can be eye - tracking cameras for both eyes, and eye registration and gaze tracking can be performed for both eyes, and thus the techniques described herein can be implemented for both the left and right eyes in an HMD.

[0040] Figures 6A to 6C Illustrates an example device in which the method according to some embodiments can be implemented Figures 1 to 5B is given. Note that the HMD 1000 illustrated as Figures 6A to 6C an example is given by way of illustration and is not intended to be limiting. In various embodiments, the shape, size, and other characteristics of the HMD 1000 can be different, and the position, number, type, and other characteristics of the components of the HMD 1000 and the eye - imaging system can be different. Figure 6A shows a side view of an example HMD 1000, and Figure 6B and Figure 6C shows an alternative front view of an example HMD 1000, where Figure 6A shows a device with a single lens 1030 covering both eyes, and Figure 6B shows a device with a right lens 1030A and a left lens 1030B.

[0041] The HMD 1000 may include a lens 1030 mounted in a wearable housing or frame 1010. The HMD 1000 may be worn on a user's head ("wearer") such that the lens is positioned in front of the wearer's eyes. In some embodiments, the HMD 1000 may implement any one of a variety of display technologies or display systems. For example, the HMD 1000 may include a display system that guides light forming an image (virtual content) through one or more waveguide layers in the lens 1020; an output coupler of the waveguide (e.g., a relief grating or a volume holographic grating) may output light toward the wearer to form an image at or near the wearer's eyes. As another example, the HMD 1000 may include a direct retinal projection system that directs light toward a reflective component of the lens; the reflective lens is configured to redirect the light to form an image at the wearer's eyes.

[0042] In some embodiments, the HMD 1000 may further include one or more sensors (e.g., eye or gaze tracking sensors) that collect information about the wearer's environment (video, depth information, lighting information, etc.) and information about the wearer. The sensors may include one or more of the following but are not limited to the following: one or more eye tracking cameras 1020 (e.g., infrared (IR) cameras) that capture a view of the user's eyes, one or more world-facing or PoV cameras 1050 (e.g., RGB cameras) that may capture an image or video of the real-world environment in the field of view in front of the user, and one or more ambient light sensors that capture lighting information of the environment. The cameras 1020 and 1050 may be integrated in the frame 1010 or attached to the frame. The HMD 1000 may also include one or more light sources 1080 that emit light (e.g., light in the IR portion of the spectrum) toward one or more of the user's eyes, such as LEDs or infrared point sources.

[0043] A controller 1060 for the XR system may be implemented in the HMD 1000 or, alternatively, may be at least partially implemented by an external device (e.g., a computing system or a portable device) communicatively coupled to the HMD 1000 via a wired interface or a wireless interface. The controller 1060 may include one or more of various types of processors, image signal processors (ISPs), graphics processing units (GPUs), encoders / decoders (codecs), systems on a chip (SOCs), CPUs, and / or other components for processing and rendering video and / or images. In some embodiments, the controller 1060 may render frames including virtual content (each frame including a left image and a right image) at least partially based on inputs obtained from the sensors and from an eye tracking system, and may provide the frames to the display system.

[0044] The memory 1070 for the XR system may be implemented in the HMD 1000 or, alternatively, may be implemented at least in part by an external device (e.g., a computing system) communicatively coupled to the HMD 1000 via a wired or wireless interface. The memory 1070 may be used, for example, to record video or images captured by one or more cameras 1050 integrated in or attached to the frame 1010. The memory 1070 may include any type of memory, such as dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM (including mobile versions of SDRAM, such as mDDR3, etc., or low-power versions of SDRAM, such as LPDDR2, etc.), RAMBUS DRAM (RDRAM), static RAM (SRAM), etc. In some embodiments, one or more memory devices may be coupled to a circuit board to form a memory module, such as a single inline memory module (SIMM), dual inline memory module (DIMM), etc. Alternatively, these devices may be mounted with the integrated circuits implementing the system in a chip stack configuration, a package stack configuration, or a multi-chip module configuration. In some embodiments, DRAM may be used as temporary storage for images or video for processing, but other storage options may also be used in the HMD to store processed data, such as flash memory or other "hard drive" technologies. This other storage device may be separate from the external coupled storage device mentioned below.

[0045] Although Figures 6A to 6C Only the light source 1080 and cameras 1020 and 1050 for one eye are shown, but embodiments may include light sources 1080 and cameras 1020 and 1050 for each eye, and gaze tracking may be performed for both eyes. Additionally, the light source 1080, the eye tracking camera 1020, and the PoV camera 1050 may be located elsewhere than shown.

[0046] As Figures 6A to 6CThe illustrated embodiment of the HMD 1000 can be used, for example, in augmented or mixed (AR) applications to provide an augmented or mixed reality view to a user. The HMD 1000 can include, for example, one or more sensors located on the outer surface of the HMD 1000 that collect information about the wearer's external environment (video, depth information, lighting information, etc.); the one or more sensors can provide the collected information to the controller 1060 of the XR system. The sensors can include one or more visible light cameras 1050 (e.g., RGB cameras) that capture video of the wearer's environment, which in some embodiments can be used to provide the wearer with a virtual view of their real environment. In some embodiments, the video stream of the real environment captured by the visible light camera 1050 can be processed by the controller 1060 of the HMD 1000 to render an augmented or mixed reality frame that includes virtual content overlaid on the view of the real environment, and the rendered frame can be provided to the display system. In some embodiments, the input from the eye tracking camera 1020 can be used during the PCCR gaze tracking process performed by the controller 1060 to track the gaze / pose of the user's eyes for rendering augmented or mixed reality content for display. Additionally, as Figures 1 to 5B One or more of the illustrated methods can be implemented in the HMD to provide non-intrusive visual axis registration for the HMD 1000.

[0047] Figure 7 is a block diagram illustrating an example device according to some embodiments, the example device can include components as Figures 1 to 5B illustrated and implement the methods as shown in these figures.

[0048] In some embodiments, the XR system can include a device 2000, such as a head-mounted device, helmet, goggles, or glasses. The device 2000 can implement any of a variety of display technologies. For example, the device 2000 can include: a transparent or semi-transparent display 2060 (e.g., glasses lenses) through which the user can view the real environment; and a medium integrated with the display 2060 through which light representing a virtual image is directed to the wearer's eyes to provide the wearer with an augmented reality view.

[0049] In some embodiments, the device 2000 may include a controller 2060 configured to implement the functions of the XR system and generate frames (each frame including a left image and a right image) provided to the display 2030. In some embodiments, the device 2000 may further include a memory 2070 configured to store software (code 2074) of the XR system executable by the controller 2060, and data 2078 that may be used by the XR system when executed on the controller 2060. In some embodiments, the memory 2070 may also be used to store video captured by the camera 2050. In some embodiments, the device 2000 may also include one or more interfaces (e.g., Bluetooth interface, USB interface, etc.), and the one or more interfaces (e.g., Bluetooth interface, USB interface, etc.) are configured to communicate with an external device (not shown) via a wired or wireless connection. In some embodiments, at least a portion of the functions described for the controller 2060 may be implemented by an external device. The external device may be or may include any type of computing system or computing device, such as a desktop computer, a notebook or laptop computer, a tablet or tablet device, a smart phone, a handheld computing device, a game controller, a game system, and so on.

[0050] In various embodiments, the controller 2060 may be a single-processor system including one processor, or a multi-processor system including several processors (e.g., two, four, eight, or another suitable number). The controller 2060 may include a central processing unit (CPU) configured to implement any suitable instruction set architecture and may be configured to execute instructions defined in that instruction set architecture. For example, in various embodiments, the controller 2060 may include a general-purpose processor or an embedded processor implementing any of a variety of instruction set architectures (ISAs) such as x86, PowerPC, SPARC, RISC, or MIPS ISA, or any other suitable ISA. In a multi-processor system, each processor may implement the same ISA jointly, but this is not necessary. The controller 2060 may adopt any microarchitecture, including scalar, superscalar, pipelined, superpipelined, out-of-order, in-order, speculative, non-speculative, etc., or a combination thereof. The controller 2060 may include circuitry implementing microcode technology. The controller 2060 may include one or more processing cores each configured to execute instructions. The controller 2060 may include one or more levels of cache, which may adopt any size and any configuration (set associative, direct mapped, etc.). In some embodiments, the controller 2060 may include at least one graphics processing unit (GPU), and the at least one graphics processing unit (GPU) may include any suitable graphics processing circuitry. Generally, the GPU may be configured to render objects to be displayed into a frame buffer (e.g., a frame buffer including pixel data for an entire frame). The GPU may include one or more graphics processors, and the one or more graphics processors may execute graphics software for performing some or all of the graphics operations or hardware acceleration for certain graphics operations. In some embodiments, the controller 2060 may include one or more other components for processing and rendering video and / or images, such as an image signal processor (ISP), an encoder / decoder (codec), etc.

[0051] Memory 2070 may include any type of memory, such as dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM (including mobile versions of SDRAM, such as mDDR3, etc., or low-power versions of SDRAM, such as LPDDR2, etc.), RAMBUS DRAM (RDRAM), static RAM (SRAM), etc. In some embodiments, one or more memory devices may be coupled to a circuit board to form a memory module, such as a single in-line memory module (SIMM), dual in-line memory module (DIMM), etc. Alternatively, these devices may be mounted with the integrated circuit implementing the system in a chip stack configuration, package stack configuration, or multi-chip module configuration. In some embodiments, DRAM may be used as temporary storage for images or video for processing, but other storage options may also be used to store the processed data, such as flash memory or other "hard disk" technologies.

[0052] In some embodiments, device 2000 may include one or more sensors that collect information about the user's environment (video, depth information, lighting information, etc.). The sensors may provide the information to the controller 2060 of the XR system. In some embodiments, the sensors may include, but are not limited to, at least one visible light camera (e.g., RGB camera) 2050, an ambient light sensor, and at least one eye tracking camera 2020. In some embodiments, device 2000 may also include one or more IR light sources; the light from the source reflected from the eye may be captured by the eye tracking camera 2020. The gaze tracking algorithm implemented by the controller 2060 may process the images or videos of the eyes captured by the camera 2020 to determine the eye pose and gaze direction. Additionally, one or more of the methods illustrated in Figure 1 FIGs. 4 to 5 may be implemented in device 2000 to provide unobtrusive visual axis registration for device 2000.

[0053] In some embodiments, device 2000 may be configured to render and display frames to provide the user with an enhanced or mixed reality (MR) view based at least in part on sensor inputs (including inputs from the eye tracking camera 2020). The MR view may include rendering the user's environment, including rendering real objects in the user's environment based on video captured by one or more cameras, and the one or more cameras capture high-quality, high-resolution video of the user's environment for display. The MR view may also include virtual content (e.g., virtual objects, virtual labels for real objects, the user's avatar, etc.) that is generated by the XR system and synthesized with the displayed view of the user's real environment.

[0054] Extended Reality

[0055] A real environment refers to an environment that a person can perceive (e.g., see, hear, feel) without using a device. For example, an office environment can include furniture such as desks, chairs, and filing cabinets; structural elements such as doors, windows, and walls; and objects such as electronic devices, books, and writing tools. A person in a real environment can perceive various aspects of the environment and can be able to interact with objects in the environment.

[0056] On the other hand, an extended reality (XR) environment is partially or fully simulated using an electronic device. For example, in an XR environment, a user can see or hear computer-generated content that partially or fully replaces the user's perception of the real environment. Additionally, a user can interact with the XR environment. For example, a user's movement can be tracked, and virtual objects in the XR environment can change in response to the user's movement. As another example, a device presenting the XR environment to the user can determine that the user is moving their hand towards the virtual location of a virtual object and can move the virtual object in response. Additionally, a user's head position and / or eye gaze can be tracked, and the virtual object can move to stay in the user's line of sight.

[0057] Examples of XR include augmented reality (AR), virtual reality (VR), and mixed reality (MR). XR can be considered a spectrum of realities, where on one hand VR fully immerses a user and replaces the real environment with virtual content, and on the other hand a user can experience the real environment without device assistance. In between are AR and MR, which blend virtual content with the real environment.

[0058] VR generally refers to an XR type that fully immerses a user and replaces the user's real environment. For example, a head-mounted device (HMD) can be used to present VR to a user, and the HMD can include a near-eye display for presenting a virtual visual environment to the user and a head-mounted earphone for presenting a virtual audible environment. In a VR environment, a user's movement can be tracked and cause a change in the user's observation of the environment. For example, a user wearing an HMD can walk in the real environment, and it will appear to the user as if they are walking in the virtual environment they are experiencing. Additionally, a user can be represented by an avatar in the virtual environment, and the HMD can use various sensors to track the user's movements to animate the user's avatar.

[0059] AR and MR refer to a class of XR that includes some mixture of the real environment and virtual content. For example, a user may hold a tablet computer that includes a camera that captures images of the user's real environment. The tablet computer may have a display that shows images of the real environment mixed with images of virtual objects. AR or MR may also be presented to the user via an HMD. The HMD may have an opaque display or may use a see-through display, which allows the user to see the real environment through the display while showing virtual content overlaid on the real environment.

[0060] In different embodiments, the methods described herein may be implemented in software, hardware, or a combination thereof. Additionally, the order of the blocks of the methods may be changed, and various elements may be added, reordered, combined, omitted, modified, etc. It will be apparent to those skilled in the art who benefit from this disclosure that various modifications and changes can be made. The various embodiments described herein are intended to be illustrative rather than limiting. Many variations, modifications, additions, and improvements are possible. Thus, multiple examples may be provided for components that are described as a single example in this document. The boundaries between various components, operations, and data repositories are to some extent arbitrary, and specific operations are shown in the context of a particular example configuration. Other allocations of functionality are anticipated, and they may fall within the scope of the appended claims. Finally, the structure and functionality presented as discrete components in an example configuration may be implemented as a combined structure or component. These and other variations, modifications, additions, and improvements may fall within the scope of the embodiments as defined in the following claims.

Claims

1. A device, the device comprising: a display subsystem configured to display virtual content to an eye; an eye-facing camera configured to capture an image of the eye; and a controller including one or more processors configured to: cause the display subsystem to display a line of horizontal content at a known vertical position and virtual depth; and process the image of the eye captured by the eye-facing camera while the eye scans the line of content to estimate a k-angle of the eye, where the k-angle indicates an optical axis of the eye.

2. The device according to claim 1, wherein prior to causing the display of the line of content, the controller is configured to generate a multi-dimensional personalized eye model based on the image of the eye captured by the eye-facing camera, where the eye model includes an optical axis of the eye, and where the k-angle is an offset from the optical axis to the optical axis of the eye.

3. The device according to claim 2, wherein in order to process the image of the eye captured by the eye-facing camera while the eye scans the line of content to estimate the k-angle of the eye, the controller is configured to: set an initial k-angle relative to the optical axis of the eye model; estimate a stimulation plane using the eye model and the initial k-angle based on eye pose data generated from the captured image; determine an error between the estimated stimulation plane and a ground truth stimulation plane corresponding to the displayed line of content; and estimate the k-angle of the eye based on the determined error between the estimated stimulation plane and the ground truth stimulation plane.

4. The device according to claim 1, wherein the line of content is a line of text to be read during a registration process of the device.

5. The device according to claim 1, wherein the line of content is the only content displayed at the known virtual depth.

6. The device according to claim 1, wherein causing the display of a line of horizontal content and processing the image of the eye captured by the eye-facing camera are performed during a registration process of the device.

7. The device according to claim 1, wherein the device further includes a second eye-facing camera configured to capture an image of a second eye, and wherein the controller is further configured to process the image of the second eye captured by the second eye-facing camera while the second eye also scans the line of content to estimate a k-angle of the second eye, where the k-angle indicates an optical axis of the second eye.

8. The device according to claim 1, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system.

9. A method, the method comprising: performing, by a controller including one or more processors, the following operations: causing a display subsystem to display a line of horizontal content at a known vertical position and virtual depth; and Estimate the k - angle of the eye at least in part based on an image of the eye captured by an eye - facing camera while the eye scans a line of content, where the k - angle indicates the visual axis of the eye.

10. The method according to claim 9, wherein the method further comprises: Before displaying the line of horizontal content, generate a multi - dimensional personalized model of the eye based on an image of the eye captured by the eye - facing camera, where the eye model includes the optical axis of the eye, and where the k - angle is the offset from the optical axis to the visual axis of the eye.

11. The method according to claim 10, wherein estimating the k - angle of the eye at least in part based on an image of the eye captured by an eye - facing camera while the eye scans a line of content comprises: Set an initial k - angle relative to the optical axis of the eye model; Use the eye model and the initial k - angle to estimate a stimulation plane based on eye pose data generated from the captured image; Determine the error between the estimated stimulation plane and a ground - truth stimulation plane corresponding to the displayed line of content; And Estimate the k - angle of the eye based on the determined error between the estimated stimulation plane and the ground - truth stimulation plane.

12. The method according to claim 9, wherein the line of content is a line of text to be read during the registration process of the device.

13. The method according to claim 9, wherein the line of content is the only content displayed at the known virtual depth.

14. The method according to claim 9, wherein the displaying of a line of horizontal content and the estimating of the k - angle of the eye are performed during the registration process of the device.

15. The method according to claim 9, wherein the method further comprises: Estimate the k - angle of the second eye.

16. The method according to claim 9, wherein the controller, the display system, and the eye - facing camera are components of a head - mounted device (HMD) of an extended reality (XR) system.

17. A system, the system comprising: A head - mounted device (HMD), the head - mounted device (HMD) comprising A display subsystem configured to display virtual content to the eye; An eye - facing camera configured to capture an image of the eye; And A controller, the controller including one or more processors configured to: Generate a multi - dimensional personalized eye model based on an image of the eye captured by the eye - facing camera, where the eye model includes the optical axis of the eye; Cause the display subsystem to display a line of horizontal content at a known vertical position and virtual depth; And Estimate the k - angle of the eye at least in part based on an image of the eye captured by an eye - facing camera while the eye scans the line of content, where the k - angle is the offset from the optical axis to the visual axis of the eye.

18. The system according to claim 17, wherein in order to estimate the k - angle of the eye, the controller is configured to: Set an initial k - angle relative to the optical axis of the eye model; Estimate a stimulation plane using the eye model and the initial k angle based on eye pose data generated from the captured images; Determine an error between the estimated stimulation plane and a ground truth stimulation plane corresponding to a displayed line of content; and Estimate the k angle of the eye based on the determined error between the estimated stimulation plane and the ground truth stimulation plane.

19. The system according to claim 17, wherein the displaying a line of horizontal content and the estimating the k angle of the eye are performed during a registration process of the device.

20. The device according to claim 17, wherein the line of content is a line of text to be read during a registration process of the device, and wherein the line of text is the only content displayed at the known virtual depth.