Method and system for eye tracking using structured light

By integrating deflectometry and triangulation methods for eye-tracking, the method significantly enhances gaze estimation accuracy and coverage, addressing the limitations of current technologies to enable precise eye-tracking in AR/VR applications.

WO2025212590A1PCT designated stage Publication Date: 2025-10-09THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/022464
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2025-04-01
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Current eye-tracking technologies suffer from low accuracy and limited coverage due to sparse surface samples, leading to significant estimation errors, especially in AR/VR applications, which are inadequate for sophisticated tasks.

Method used

Combining deflectometry and triangulation techniques to increase data density and coverage of the eye surface by using structured light patterns, projecting patterns on both specular and diffuse surfaces, and employing a hybrid approach to enhance gaze estimation accuracy.

Benefits of technology

Achieves precise gaze estimation errors below 0.1°, enabling high-quality graphics experiences and improved interaction in AR/VR headsets, and providing accurate health monitoring and data delivery in combat situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025022464_09102025_PF_FP_ABST
    Figure US2025022464_09102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, devices and system for accurate and efficient determination of the gaze direction of an eye are described. One example method for determining the gaze direction illuminates one or more sections of the eye with structured light patterns, which includes projecting a first structured light pattern onto a surface of the eye with diffuse reflection properties. In response to the illumination, a first reflected light from the surface of the eye having diffuse reflection properties and a second reflected light from the surface of the eye having specular reflection properties are received, and a first and a second set of surface normal vectors are obtained based on the first reflected light and the second reflected light, respectively. The first and the second set of surface normal vectors are then used to determine the gaze direction.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR EYE TRACKING USING STRUCTURED LIGHTCROSS-REFERENCE TO RELATED APPLICATION(S)

[0001] This application claims priority to the provisional application with serial number 63 / 572,870 titled “METHOD AND SYSTEM FOR EYE TRACKING USING STRUCTURED LIGHT,” filed April 1, 2025. The entire contents of the above noted provisional application are incorporated by reference as part of the disclosure of this document.TECHNICAL FIELD

[0002] The technology described in this patent document relates to method and devices that enable determination of the gaze direction of an eye.BACKGROUND

[0003] Eye-tracking plays a crucial role in modem human-computer interaction (HCI) and machine perception (MP) methods. Recent implementations focus on the application in virtual reality (VR) headsets, as well as use cases in neuroscience research, and psychology. If implemented successfully, eye-tracking has the potential to be become the “go-to” input method for computers, mobile phones, or general user interfaces in the future and can potentially replace mouse and touchpads. Therefore, it is important to enable accurate and reliable eye-tracking techniques.SUMMARY

[0004] The disclosed embodiments relate to methods and devices for determining the gaze direction of an eye more accurately and more efficiently. As a result, improved eye-tracking applications can be implemented and used in a variety of fields.

[0005] One example method for determining a gaze direction of an eye includes illuminating one or more sections of the eye with one or more structured light patterns, which includes projecting a first structured light pattern onto a surface of the eye having diffuse reflection properties, and illuminating the eye using a second structured light pattern. The method further includes, in response to the illumination by the one or more structured light patterns, receiving, at one or more detectors, a first reflected light from the surface of the eye having diffuse reflection properties, and a second reflected light from the surface of the eye having specular reflectionproperties. The method additionally includes obtaining a first set of surface normal vectors based on the first reflected light received at the one or more detectors, and a second set of surface normal vectors based on the second of reflected light received at the one or more detectors, and determining the gaze direction based on the first and the second set of normal vectors.BRIEF DESCRIPTION OF THE DRAWINGS

[0006] FIG. 1, in panels (a) to (d), illustrates the use of a deflectometry -based technique for measuring specular surfaces of the eye, including an experimental setup diagram and determination of surface normals.

[0007] FIG. 2, in panels (a) to (d), illustrates the use of an optimization-based inverse rendering technique deflectometry-based technique for determining rotation, translation and shape of an eye.

[0008] FIG. 3 illustrates an example configuration that can be used to determine a gaze direction and to estimate parameters of an eye in accordance with some embodiments.

[0009] FIG. 4 illustrates a configuration that can be used to determine a gaze direction and to estimate parameters of an eye in accordance with some example embodiments.

[0010] FIG. 5 illustrates a set of operations that can be carried out to determine a gaze direction of an eye in accordance with an example embodiment.DETAILED DESCRIPTION

[0011] In recent years, the need for a robust, fast, and precise solution to estimate the human gaze direction has evolved into a central unsolved problem in AR / VR headset research. In consumer devices, sophisticated eye tracking would enable a higher quality graphics experience through foveated rendering or would significantly improve the interaction with virtual avatars. In a battlefield context, a precise eye tracking solution in AR headsets can be used to estimate and monitor the health status of the soldier (e.g., level of fatigue) and can help to significantly increase the headset’s viewing comfort by continuously keeping track of the inter-pupillary distance (“accommodation convergence reflex”). Moreover, tracking the gaze of the soldier in a combat situation can deliver important data about which information displayed in the AR headset he is actually using. Ultimately, precise eye tracking can compensate for imperfections of other system sensors, e.g., for position / location estimation.

[0012] Despite the obvious need for a precise eye tracking solution, the current state of the art leaves much room for improvement. Current approaches either utilize two-dimensional (2D)features detected from 2D eye images (“image-based methods”) or exploit sparse reflections (“reflection-based methods”) of a few point light sources at the eye surface (“corneal / scleral reflections”). The latter retrieves three-dimensional (3D) surface information and leads to an improved gazing direction calculation, albeit the number of surface samples (-10 = number of point light sources) is small. For those methods, the density of measured surface points (2D features for image-based methods, point source reflections for glint tracking) is relatively low, and the acquired 3D information about the eye surface is limited. For this reason, current methods are typically not able to achieve gaze estimation accuracies better than 2°-4° for imagebased methods and -1° for reflection-based methods. This, in some example configurations, translates to an error that is between 60 and 200 pixels, which is far too inaccurate for sophisticated tasks (e.g., text editing), and far away from “mouse accuracy.” More measured surface samples are expected to lead to a more precise estimation of the gazing direction. To significantly increase the information content provided by corneal or scleral reflection to calculate the gazing direction, the number of light sources observed over the eye surface must be significantly increased.

[0013] One technique to address this problem utilizes deflectometry. Deflectometry is an established method in surface metrology to reconstruct the 3D surface of specular objects, such as freeform lenses, car windshields, or technical parts. The specular reflection of an extended screen displaying a known pattern (e.g., a sinusoid) is observed over the specular surface of an object under test (in this case, the human eye). From the deformation of the pattern in the camera image, the normal vectors of the surface (and eventually the surface shape via integration) can be calculated. In comparison to the current state of the art, an increase in data density by a factor >100,000 is easily achievable. For example, the reflection information from a 2Mpix screen compared to 10 sparse point reflections of the prior systems leads to an increase of measured points by a factor of 200,000. In some laboratory implementations of the above defl ectometric technique, factors >3,300 on realistic eye models and real human eyes in-vivo have been demonstrated, and gaze estimation errors <0.25° and <0.45° have been achieved in the following two approaches.

[0014] Classical single-shot deflectometry approach: This method uses a crossed fringe pattern on the screen to measure the surface normal map of the eye surface. The evaluation via a 2D continuous wavelet transform approach and stereo-deflectometry allows for a uniquereconstruction of the shape and normal map of the eye surface in single-shot. To estimate the gaze direction, the calculated surface normals are traced back towards the cornea and sclera centers (see FIG. 1 panels (b) to (d)). Due to the spherical shape of the eye’s cornea and sclera, but their vastly different radii, the back-traced normals intersect at two different points inside the eye (cornea and sclera center). Connecting these two points delivers the optical axis of the eye and the gaze direction. The approach also works if cornea and sclera are not perfectly spherical but rotationally symmetric, meaning that all back-traced normals would intersect along a line that coincides with the optical axis of the eye.

[0015] Notably, with respect to FIG. 1, panel (a) illustrates the general technique for measuring the specular surfaces of the eye using deflectometry, where a camera observes a screen with displayed pattern over the specular surface. The surface normal and shape can be calculated from the deformation of the pattern in the camera image. Panel (b) illustrates a simulation where the screen reflection covers cornea and sclera, and the captured normals are back-traced towards the eye center to retrieve cornea and sclera center. Their connecting vector represents the optical axis of the eye which can be used to retrieve the gaze direction. Panels (c) and (d) illustrate example experiments, where the gaze direction is retrieved from a realistic eye model, and from a human eye in-vivo, respectively.

[0016] Optimization-based inverse rendering approach: The operations associated with this method are illustrated in FIG. 2, panels (a) to (d), where panel (a) illustrates a prototype that includes an eye model, a screen to project a pattern (e.g., a sinusoidal pattern as illustrated) and a camera to capture the reflected light. Panel (b) shows screen-camera correspondences are found in the real eye measurements. Panel (c) illustrates simulated “digital twin” of the prototype setup and panel (d) shows the obtained screen camera correspondences of “digital twin” setup are compared with real measurements (loss) and eye's rotation, translation, and shape are obtained via differential rendering.

[0017] To further elaborate, this method uses the known geometry of a calibrated deflectometric setup, as illustrated in FIG. 2’s panel (a), to develop a PyTorch3D-based differentiable rendering pipeline that simulates a virtual computer-generated (CG) eye model under screen illumination (see FIG. 2, panel (c)). Panels (a) and (c) illustrate twin setups, one real setup (panel (a)) and one virtual setup (panel (c)); the virtual setup uses an eye model, andthe same camera and display positioned at the same geometric relation as the real setup of panel (a). Using an initial image obtained in the real environment, a camera-screen correspondence is obtained (panel (b)). In the virtual setup, an initial position, rotation and / or shape of the eye model’s shape is selected, a correspondence map is calculated and compared to the real correspondence. The position / rotation / shape of the eye are changed, and this process is repeated via an optimization technique to map the eye, including the determination of its shape and its rotation and translation parameters. For example, in panel (d), a loss function (L), which represents the distance between the camera-screen correspondence of the real and virtual setups, is obtained. This is done until the simulated setup with the CG eye produces images and correspondences that closely match the real measurements, and the gaze direction of the CG eye is eventually used as an estimate of the real eye’s gaze direction.

[0018] In some implementations, the screen illumination can be realized by a secondary infrared (IR) screen that is placed, e.g., on the inner side of a VR headset, to project patterns that are invisible to the human eye. The “screen” could also be a simple self-illuminated pattern that covers the inside of the headset. Moreover, it should be emphasized that any form of “pattern” can work, as long as the pattern is known. For example, the images displayed on the computer screen or VR headset themselves (video stream, video game, website in browser, and the like) can be used to perform both of the above-described methods by finding screen-camera correspondences over shift-invariant feature transform (SIFT) feature tracking.

[0019] The above-noted techniques can be further improved (in some cases significantly) by addressing some of the shortcomings of those approaches that include the following.

[0020] Limited coverage: Although deflectometric methods allow for a very high number of acquired datapoints, their distribution over the eye surface is confined to a small area on the sclera and cornea where the reflection of the display is observed, leading to a limited coverage of the eye surface. Notably, due to the large curvature of the eye, reflections from the entire eye surface cannot be practically obtained via deflectometric techniques in a single shot. A better coverage can lead to an even better gaze estimation accuracy, for example, below 0.1°. Increasing the coverage for standard deflectometric measurements is hard, because it strongly depends on geometrical constraints (such as the size of the display and the standoff distance of the eye) which often cannot be changed for specific applications.

[0021] High computational effort for inverse rendering approach: The optimization-based inverse rendering approach described earlier simultaneously estimates the eye’s rotation parameters, translation parameters, and shape parameters via inverse rendering. However, in some instances, this process can take several seconds, making the inverse rendering approach impractical for detecting fast eye movements in real-time.

[0022] The disclosed embodiments, among other features and benefits, increase the coverage by exploiting additional surface information from triangulation, and improve the inverse rendering performance (speed and precision) with additional shape and rotation clues.

[0023] As noted above, the geometrical constraints for specular reflections lead to the low coverage problem. While it is strictly necessary to exploit specular reflection information for the highly specular cornea surface of the eye, the situation is different for the surface of the sclera. As the sclera surface has a significant diffuse component, in some embodiments, triangulationbased approaches (such as fringe projection or line triangulation) are used to capture the 3D shape of sclera and iris. It should be noted that in defl ectometric techniques, the eye is illuminated using a pattern that is present on a screen and the reflection is captured by the camera. Using a triangulation-based approach, a pattern is projected onto the diffuse surface (instead of observing the reflection of (screen) points over a specular surface), and the 3D shape of the surface is calculated by observing the deformation of the pattern in the camera image.

[0024] Projecting the pattern is possible because the diffuse surface scatters light in all directions (and not only in one direction as specular surfaces do). This is the key reason why all diffuse sections of the eye can be covered with an on-projected pattern; when combined with the deflectometry-based techniques, it significantly increases the total combined coverage of the entire eye (i.e., coverage of those sections of the eye with specular properties - e.g., cornea - and sections with diffuse properties - e g., sclera). We emphasize that specular surface deflectometry and diffuse surface triangulation are two physically different approaches with different fundamental physical and information-theoretical limitations.

[0025] It should be further noted that a significant improvement in gaze detection accuracy is obtained by the disclosed embodiments due in-part to the use of the combined deflectometry and projection-based (triangulation) techniques. As such, very small gaze estimation errors can be achieved, which enable eye tracking close to the fundamental limit. For instance, for a retinalresolution display (e.g., 60 pix / deg), this error corresponds to only 6 pixels, which overlaps with “mouse accuracy.”

[0026] FIG. 3 illustrates an example configuration that can be used to determine a gaze direction and to estimate parameters of an eye in accordance with some embodiments. The configuration in FIG. 3 includes two cameras, a projector for projecting a pattern on one or more sections of the eye, and an extended illumination source (e.g., a display) that provides an illumination pattern such as a sinusoidal or a cross-sinusoidal (e.g., checkerboard) pattern.

[0027] In operation, both the projector and the display illuminate the eye with the known pattern. The illumination can be implemented in different ways; for example, by fast temporal alteration of the pattens, by color (wavelength) coding, and / or illumination in the ultraviolet (UV) region, which can be done simultaneously or sequentially. In some instances, appropriate filters in front of the cameras may be placed to enable the detection of appropriate spectral ranges. In some implementations, a single camera may be used.

[0028] In the example configuration of FIG. 3, the projector and camera 2 can be used for triangulation-based measurements, where deformation of the projected pattern on diffuse surfaces is observed with the camera, and shape (and later surface normals) is calculated from deformation of the known pattern. The surfaces can include sclera and iris, for example, and can be at least partially diffuse. Each camera can form a triangulation pair with projector based on the proper calibration where the locations of the cameras and the projector are known. In the triangulation method, the projector, the camera and the point on the object’s surface define a spatial triangle. Because the distance between the projector and the camera (the base of the triangle) is known, the angles between the base and each side, the apex of the triangle and thus the 3D coordinates of the object, can be calculated from the triangular relations. In some embodiments, the first camera and the projector are used as a triangulation pair. In another embodiment, the second camera and the projector are used a triangulation pair. In still another embodiment, both cameras and the projector are used to form two triangulation pairs. In yet other embodiments, three triangulation pairs can be used: a first triangulation pair based on the projector and the first camera, a second triangulation pair based on the projector and the second, and a third triangulation pair between the two cameras.

[0029] In the example configuration of FIG. 3, the display and camera 1 are used to conductdefl ectom etry -based measurements for specular surfaces of the eye. Deformation of the displayed pattern after reflection from the specular surfaces is observed with camera 1, and surface normals (and subsequently shape of the eye) are calculated from deformation of the known pattern.

[0030] FIG. 4 illustrates another example configuration that can be used to determine a gaze direction and to estimate parameters of an eye in accordance with some embodiments. FIG. 4 illustrates that additional cameras may be used, and all cameras can work together with all projectors and displays. The cameras, the projector and the screen can be coupled to a computer that controls their operations, and can receive and process the information that it receives from the cameras. As noted earlier, each camera can form a triangulation pair with the projector and the other camera (if present). In general, having additional triangulation pairs increases the coverage area, and can improve the accuracy of the measurements. Therefore, additional cameras may be added to the configuration. In some embodiments, only one camera (along with a screen and the projector) is used, which can provide sufficient coverage and accuracy for some applications.

[0031] Based on the above, the gazing direction, as well as various eye parameters (e.g., shape, rotation, translation) can be obtained using different techniques. For example, in one embodiment, the surface normals can be traced back to a common center (or a common line) for the different surfaces, and the vector connecting the centers (e g., as shown in FIG. 1, panel (b)) determines the gaze direction.

[0032] In another embodiment, we use the obtained shapes for a “refined” optimizationbased digital twin setup: the rough shape of sclera obtained via the triangulation technique provides a high accuracy initial shape and position of eyeball, which is input into the digital twin setup built in a virtual environment (described earlier). In this way, instead of guessing the starting shape, a more realistic estimate is used, which reduces the number of computations. As before, in the virtual environment, we adjust the eye shape and position and acquire the correspondence map; we compare the virtual correspondence map with the real correspondence map we obtained in the real experiment, and optimize the eye shape and position by minimizing the correspondence map difference. In yet other embodiments, the two prior embodiments can be combined.

[0033] The disclosed embodiments can further improve the inverse rendering performance (speed and precision) by using additional shape and rotation clues. Although the classical deflectometry approach and the optimization-based approach have been phrased as two different flavors of the basic idea, both approaches show great potential to complement each other to form the next generation approach. The basic idea is as follows: use shape, rotation, and translation estimation of the classical deflectometry approach to obtain very good starting values for the inverse rendering optimization. This significantly decreases the number of necessary iterations and simultaneously prevents the optimization from running into local minima. Even better, the additional triangulation information obtained from deflectometry and triangulation techniques can be seamlessly included in this process, as the additional triangulation data just leads to even better (or more complete coverage) starting values for shape, rotation and translation. This means that in the final workflow, the optimization can be seen as a second step after shape estimation. This second step significantly improves the results, and at least in some scenarios, can compensate for systematic errors of the errors of the system or when one of the components is moved (which can cause the system’s de-calibration). For example, based on having sufficient information about the scene and object (especially due to the available triangulation data), the optimization algorithm can additionally evaluate the position of all cameras, the projector, and the display. This process can be used to refine the calibration of the setup or to compensate for de-calibration.

[0034] FIG. 5 illustrates a set of operations that can be carried out to determine a gaze direction of an eye in accordance with an example embodiment. At 502, one or more sections of the eye are illuminated with one or more structured light patterns; this operation includes projecting a first structured light pattern onto a surface of the eye having diffuse reflection properties, and illuminating the eye using a second structured light pattern. At 504, in response to the illuminating, receiving at one or more detectors, a first reflected light is received from the surface of the eye having diffuse reflection properties, and a second reflected light is received from the surface of the eye having specular reflection properties. At 506, a first set of surface normal vectors is obtained based on the first reflected light received at the one or more detectors, and a second set of surface normal vectors is obtained based on the second of reflected light received at the one or more detectors. At 508, the gaze direction is determined based on the first and the second set of normal vectors.

[0035] In one example embodiment, the surface of the eye having diffuse reflection properties includes sclera and the surface of the eye having specular reflection properties includes cornea. In another example embodiment, determining the gaze direction includes tracing back the first set and the second set of surface normal vectors to two or more intersection points and determining a vector based on the one or more intersection points that designates the gaze direction. In yet another example embodiment, the two or more intersection points include a first intersection point that represents a center corresponding to the surface of the eye having diffuse reflection properties, and a second intersection point that represents a center corresponding to the surface of the eye having specular reflection properties. In still another example embodiment, determining the gaze direction includes connecting the first intersection point and the second intersection point.

[0036] According to another example embodiment, the one or more detectors include a camera. In one example embodiment, illuminating the one or more sections of the eye includes illumination with an infrared light. In another example embodiment, the one or more structured light patterns includes one, or a combination, of: a sinusoidal illumination pattern, a checkerboard illumination pattern, a striped illumination pattern, or a particular image. In still another example embodiment, obtaining the first set of surface normal vectors based on the first reflected light includes using a triangulation-based method to obtain the first set of surface normal vectors.

[0037] In another example embodiment, obtaining the second set of surface normal vectors based on the second reflected light includes using a deflectometric method to obtain an initial illumination-to-detector correspondence in response to illumination by the one or more structured light patterns, and using a virtual setup to iteratively change a position or orientation of a model eye to obtain a more accurate estimate of the correspondence compared to the initial correspondence. In another example embodiment, the initial correspondence is obtained based on a triangulation-based method that is used to obtain the first set of surface normal vectors.

[0038] In one example embodiment, the operations of FIG. 5 also include determining one or more of a shape, rotation, or translation estimate of the eye, and using the one or more shape, rotation, or translation estimate as one or more starting points of an iterative method to render a presentation of the eye. In another example embodiment, obtaining the second set of surface normal vectors based on the second reflected light includes using a deflectometric method to obtain the second set of normal vectors, wherein the deflectometric method comprises determining a set of correspondences between pixels or sections of a screen used to provide the second structuredlight pattern and pixels of the one or more detectors. In still another example embodiment, illuminating the eye using the second structured light pattern comprises using a screen to display the second structured light pattern.

[0039] Another aspect of the disclosed embodiments relates to a device that includes one or more illumination sources operable to illuminate one or more sections of an eye with one or more structured light patterns; the one or more illumination sources are coupled to or include a screen and a projector, where the projector is operable to project a first structured light pattern onto a surface of the eye having diffuse reflection properties and the screen is operable to illuminate the eye with a second structured light pattern. The device further includes one or more detectors positioned, to receive, in response to illumination by the projector and the screen, a first reflected light from the surface of the eye having diffuse reflection properties, and a second reflected light from the surface of the eye having specular reflection properties. The device additionally includes a processor and a memory having instructions stored thereon, wherein the instructions upon execution by the processor cause the processor to: obtain a first set of surface normal vectors based on the first reflected light received at the one or more detectors, and a second set of surface normal vectors based on the second of reflected light received at the one or more detectors, and determine a gaze direction associated with the eye based on the first and the second set of surface normal vectors.

[0040] In one example embodiment, the device comprises a virtual reality or an augmented reality device that includes the screen. In another example embodiment, the device is a headmounted device. In yet another example embodiment, the one or more illumination sources are operable in one more of an ultraviolet, a visible or an infrared wavelength range, and the one or more detectors are configured to detect light in the one or more ultraviolet, visible or infrared wavelength ranges. In still another example embodiment, the instructions upon execution by the processor cause the processor to determine the gaze direction that includes tracing back the first set and the second set of surface normal vectors to two or more intersection points and determining a vector based on the one or more intersection points that designates the gaze direction. For example, the two or more intersection points include a first intersection point that represents a center corresponding to the surface of the eye having diffuse reflection properties, and a second intersection point that represents a center corresponding to the surface of the eye having specular reflection properties. In one example embodiment, determining the gaze direction includesconnecting the first intersection point and the second intersection point.

[0041] According to another example embodiment, the instructions upon execution by the processor cause the processor to obtain the first set of surface normal vectors based on the first reflected light at least in part by using a triangulation-based method to obtain the first set of surface normal vectors. In another example embodiment, the instructions upon execution by the processor cause the processor to obtain the second set of surface normal vectors based on the second reflected light at least in part by using a deflectometric method to obtain an initial illumination-to-detector correspondence in response to illumination by the one or more structured light patterns, and using a virtual setup to iteratively change a position or orientation of a model eye to obtain a more accurate estimate of the correspondence compared to the initial correspondence. For example, the initial correspondence is obtained based on a triangulation-based method that is used to obtain the first set of surface normal vectors.

[0042] In yet another example embodiment, the instructions upon execution by the processor cause the processor to determine one or more of a shape, rotation, or translation estimate of the eye, and use the one or more shape, rotation, or translation estimate as one or more starting points of an iterative method to render a presentation of the eye. In another example embodiment, the instructions upon execution by the processor cause the processor to obtain the second set of surface normal vectors based on the second reflected light at least in part by: using a deflectometric method to obtain the second set of normal vectors, wherein the deflectometric method comprises determining a set of correspondences between pixels or sections of a screen used to provide the second structured light pattern and pixels of the one or more detectors.

[0043] It is understood that the various disclosed embodiments may be implemented individually, or collectively, in devices comprised of various components, including electronics hardware and / or software modules and optical components. These devices, for example, may comprise a processor, a memory unit, an interface that are communicatively connected to each other, and may range from desktop and / or laptop computers, to mobile devices and the like. The processor and / or controller can perform various disclosed operations based on execution of program code that is stored on a storage medium. The processor and / or controller can, for example, be in communication with at least one memory and with at least one communication unit that enables the exchange of data and information, directly or indirectly, through the communication link with other entities, devices and networks. The communication unit mayprovide wired and / or wireless communication capabilities in accordance with one or more communication protocols, and therefore it may comprise the proper transmitter / receiver antennas, circuitry and ports, as well as the encoding / decoding capabilities that may be necessary for proper transmission and / or reception of data and other information.

[0044] Various information and data processing operations described herein may be implemented in one embodiment by a computer program product, embodied in a computer- readable medium, including computer-executable instructions, such as program code, executed by computers in networked environments. A computer-readable medium may include removable and non-removable storage devices including, but not limited to, Read Only Memory (ROM), Random Access Memory (RAM), compact discs (CDs), digital versatile discs (DVD), etc. Therefore, the computer-readable media that is described in the present application comprises non-transitory storage media. Generally, program modules may include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of program code for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps or processes.

[0045] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.

Claims

CLAIMSWhat is claimed is:

1. A method for determining a gaze direction of an eye, comprising: illuminating one or more sections of the eye with one or more structured light patterns, wherein the illuminating comprises projecting a first structured light pattern onto a surface of the eye having diffuse reflection properties, and illuminating the eye using a second structured light pattern; in response to the illuminating, receiving, at one or more detectors, a first reflected light from the surface of the eye having diffuse reflection properties, and a second reflected light from the surface of the eye having specular reflection properties; obtaining a first set of surface normal vectors based on the first reflected light received at the one or more detectors, and a second set of surface normal vectors based on the second of reflected light received at the one or more detectors; and determining the gaze direction based on the first and the second set of normal vectors.

2. The method of claim 1, wherein the surface of the eye having diffuse reflection properties includes sclera and the surface of the eye having specular reflection properties includes cornea.

3. The method of claim 1, wherein determining the gaze direction includes tracing back the first set and the second set of surface normal vectors to two or more intersection points and determining a vector based on the one or more intersection points that designates the gaze direction.

4. The method of claim 3, wherein the two or more intersection points include a first intersection point that represents a center corresponding to the surface of the eye having diffuse reflection properties, and a second intersection point that represents a center corresponding to the surface of the eye having specular reflection properties.

5. The method of claim 4, wherein determining the gaze direction includes connecting the first intersection point and the second intersection point.

6. The method of claim 1, wherein the one or more detectors include a camera.

7. The method of claim 1, wherein illuminating the one or more sections of the eye includes illumination with an infrared light.

8. The method of claim 1, wherein the one or more structured light patterns includes one, or a combination, of: a sinusoidal illumination pattern, a checkerboard illumination pattern, a striped illumination pattern, or a particular image.

9. The method of claim 1, wherein obtaining the first set of surface normal vectors based on the first reflected light includes using a triangulation-based method to obtain the first set of surface normal vectors.

10. The method of claim 1, wherein obtaining the second set of surface normal vectors based on the second reflected light includes using a deflectometric method to obtain an initial illumination-to-detector correspondence in response to illumination by the one or more structured light patterns, and using a virtual setup to iteratively change a position or orientation of a model eye to obtain a more accurate estimate of the correspondence compared to the initial correspondence.

11. The method of claim 10, wherein the initial correspondence is obtained based on a triangulation-based method that is used to obtain the first set of surface normal vectors.

12. The method of claim 1, comprising: determining one or more of a shape, rotation, or translation estimate of the eye, andusing the one or more shape, rotation, or translation estimate as one or more starting points of an iterative method to render a presentation of the eye.

13. The method of claim 1, wherein obtaining the second set of surface normal vectors based on the second reflected light includes: using a deflectometric method to obtain the second set of normal vectors, wherein the defl ectometric method comprises determining a set of correspondences between pixels or sections of a screen used to provide the second structured light pattern and pixels of the one or more detectors.

14. The method of claim 1, wherein illuminating the eye using the second structured light pattern comprises using a screen to display the second structured light pattern.

15. A device, comprising: one or more illumination sources operable to illuminate one or more sections of an eye with one or more structured light patterns, the one or more illumination sources coupled to or include a screen and a projector, wherein the projector is operable to project a first structured light pattern onto a surface of the eye having diffuse reflection properties and the screen is operable to illuminate the eye with a second structured light pattern; one or more detectors positioned, to receive, in response to illumination by the projector and the screen,: a first reflected light from the surface of the eye having diffuse reflection properties, and a second reflected light from the surface of the eye having specular reflection properties; a processor and a memory having instructions stored thereon, wherein the instructions upon execution by the processor cause the processor to: obtain a first set of surface normal vectors based on the first reflected light received at the one or more detectors, and a second set of surface normal vectors based on the second of reflected light received at the one or more detectors, anddetermine a gaze direction associated with the eye based on the first and the second set of surface normal vectors.

16. The device of claim 15, wherein the device comprises a virtual reality or an augmented reality device that includes the screen.

17. The device of claim 15, wherein the device is a head-mounted device.

18. The device of claim 15, wherein the one or more illumination sources are operable in one more of an ultraviolet, a visible or an infrared wavelength range, and the one or more detectors are configured to detect light in the one or more ultraviolet, visible or infrared wavelength ranges.

19. The device of claim 15, wherein the instructions upon execution by the processor cause the processor to determine the gaze direction at least in part by tracing back the first set and the second set of surface normal vectors to two or more intersection points and determining a vector based on the one or more intersection points that designates the gaze direction.

20. The device of claim 19, wherein the two or more intersection points include a first intersection point that represents a center corresponding to the surface of the eye having diffuse reflection properties, and a second intersection point that represents a center corresponding to the surface of the eye having specular reflection properties.

21. The device of claim 20, wherein determining the gaze direction includes connecting the first intersection point and the second intersection point.

22. The device of claim 15, wherein the one or more structured light patterns includes one, or a combination, of a sinusoidal illumination pattern, a checkerboard illumination pattern, a striped illumination pattern, ora particular image.

23. The device of claim 15, wherein the instructions upon execution by the processor cause the processor to obtain the first set of surface normal vectors based on the first reflected light at least in part by using a triangulation-based method to obtain the first set of surface normal vectors.

24. The device of claim 15, wherein the instructions upon execution by the processor cause the processor to obtain the second set of surface normal vectors based on the second reflected light at least in part by using a deflectometric method to obtain an initial illumination-to-detector correspondence in response to illumination by the one or more structured light patterns, and using a virtual setup to iteratively change a position or orientation of a model eye to obtain a more accurate estimate of the correspondence compared to the initial correspondence.

25. The device of claim 24, wherein the initial correspondence is obtained based on a triangulation-based method that is used to obtain the first set of surface normal vectors.

26. The device of claim 15, wherein the instructions upon execution by the processor cause the processor to: determine one or more of a shape, rotation, or translation estimate of the eye, and use the one or more shape, rotation, or translation estimate as one or more starting points of an iterative method to render a presentation of the eye.

27. The device of claim 15, wherein the instructions upon execution by the processor cause the processor to obtain the second set of surface normal vectors based on the second reflected light at least in part by: using a deflectometric method to obtain the second set of normal vectors, wherein the deflectometric method comprises determining a set of correspondences between pixels or sections of a screen used to provide the second structured light pattern and pixels of the one or more detectors.

Citation Information

Patent Citations

  • Eye tracking using optical flow

    US20170131765A1

  • Eye center of rotation determination with one or more eye tracking cameras

    US20220253135A1

  • Gaze-Operated Point Designation on a 3D Object or Scene

    US20230370576A1