Eye tracking apparatus and method thereof
By receiving multiple images and determining the three-dimensional position of the eye using triangulation, the accuracy and noise problems of existing eye tracking devices when measuring six degrees of freedom movements are solved, achieving high-precision eye tracking.
Patent Information
- Application Number
- CN202510166305.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-29
- Filing Date
- 2020-01-27
- Publication Date
- 2025-06-03
AI Technical Summary
Existing eye tracking devices have difficulty accurately measuring the six degrees of freedom of the eye, especially in natural environments, and there are problems of noise and low accuracy.
By receiving at least two images, the region associated with the limbus of the eye is identified, the geometric representation of the limbic structure is determined, and the three-dimensional position and gaze direction of the eye are determined by triangulation.
Accurate eye tracking in the range of 1/6 degrees to 1/60 degrees is achieved, avoiding user calibration and being free from corneal distortion and environmental reflexes.
Smart Images

Figure CN120085754A_ABST
Abstract
Description
[0001] This application is a divisional application of an application with an application date of January 27, 2020, an application number of 202080011133.2, and an invention title of "Eye Tracking Device and Method". Technical Field
[0002] The present invention relates to an eye tracking device and method for determining an eye position and / or a viewing direction. Background Art
[0003] The following references are listed as being considered relevant to the presently disclosed subject matter as background:
[0004] 1. U.S. Patent Application No. 2013 / 120712;
[0005] 2. International Patent Application No. WO9418883;
[0006] 3. U.S. Patent Application No. 2014 / 180162;
[0007] 4. U.S. Patent Application No. 2001 / 035938.
[0008] 5. Constable PA, Bach M, Frishman LJ, Jeffrey BG, Robson AG, International Society for Clinical Electrophysiology of Vision. ISCEV Standard for clinical electro-oculography (2017 update). Doc Ophthalmol. 2017; 134(1): 1-9.
[0009] 6. McCamy, Michael & Collins, Niamh & Otero-Millan, Jorge & Al-Kalbani, Mohammed & Macknik, Stephen & Coakley, Davis & Troncoso, Xoana & Boyle, Gerard & Narayanan, Vinodh & R Wolf, Thomas & Martinez-Conde, Susana. (2013). Simultaneous recordings of ocular microtremor and microsaccades with a piezoelectric sensor and a video-oculography system. Peer J. 1.e14. 10.7717 / peerj.14.
[0010] The acknowledgement of the above references should not be construed in this text as meaning that these are in any way relevant to the patentability of the presently disclosed subject matter.
[0011] Background
[0012] It is known in the prior art that there are different eye-tracking devices such as remote eye-trackers or head-mounted eye-trackers and different eye-tracking methods. The first eye-tracker was built at the end of the 19th century. They were difficult to build and caused discomfort to the participants. Specially designed rings, contact lenses, and suction cups were attached to the eye to assist in eye movement measurement. The first photography-based eye-tracker, which examined the light reflected from different parts of the eye, was introduced only in the early 20th century. They were much less invasive and opened a new era in eye research. For most of the 20th century, researchers built their own eye-trackers, which were expensive and had limited availability. The first commercially available eye-trackers appeared only in the 1970s. Since around the 1950s, many different techniques developed by researchers are still in use today, such as contact lenses with mirrors (more precisely, suction cups), contact lenses with coils, electrooculography (EOG) [5], and piezoelectric sensors [6]. However, these methods are only capable of having angular eye movements and, due to their nature, cannot measure any lateral eye movements (mirrors, coils, and EOG can only measure angles). In addition, these systems are stationary and require the user's head to be stable, thus making them unsuitable for studying eye movements in more natural environments.
[0013] The second part of the 20th century was dominated by much less invasive illumination and light sensor-based methods. The main limitation of this approach was that cameras were relatively slow due to the exposure and processing capabilities required to extract eye movement data. Position sensing photodetectors (PSDs) or four-wire based methods were developed in the mid-20th century. Dual Purkinje image eye tracker systems were developed in the 1970s and are still in use today. However, dual Purkinje image eye tracker systems are very expensive, fixed, heavy camera-less (PSD-based) systems with relatively low recording angles. Since the 1980s, due to improvements in camera sensors and computer technology, the eye tracking market has been dominated by so-called video oculography (VOG) systems. VOG systems typically capture images of the user's eyes and determine certain features of the eyes based on the captured images. These systems are non-invasive and typically rely on infrared illumination of the eyes that does not cause interference or discomfort to the user. Some of them rely on small, lightweight wearable cameras, but more precise systems are fixed, like most of the systems mentioned above. Since good illumination (but below safety levels) is required and taking into account the exposure requirements of the camera sensors, the frame rate of these systems is limited.
[0014] The combined pupil / corneal reflection (first Purkinje) eye tracker method (both fixed high-end and wearable low-end) illuminates the eye with multiple infrared light-emitting diodes and images the surface of the eye with one or more (usually to increase the frame rate) cameras, segmenting the pupil (as the darkest part of the eye) and the first Purkinje image of the diodes. The change in the pupil position relative to the first Purkinje image of the IR diodes indicates eye movement. User calibration must be used to calculate the true angle. Typically, the user needs to focus on a certain target and move along a known path to calibrate the system.
[0015] The accuracy and precision of this method are relatively low because the pupil position is distorted by the cornea, pupil dilation further reduces the measurement precision, and environmental reflections from the cornea confound the image processing algorithm. Different eye colors, long eyelashes, and contact lenses are additional factors that further complicate the image processing system. Thus, typically these systems are noisy and provide less than 1 degree of precision for angular motion and no information about lateral eye movement. Additionally, recalibration is required after the system has moved relative to the head. There are many other less common methods such as imaging the retina, the bright pupil method, and even examining eye movement with an MRI machine. These methods have their limitations and are not very common. There is always a continuing need for improved accuracy, precision, low latency, and compact size with regard to eye trackers and eye tracking methods. Additionally, there is a need for very compact eye tracking devices (providing great flexibility) as eye trackers become more integrated in many devices such as computers, cars, virtual reality glasses, etc.
[0016] An eye tracking system can track in particular the limbus (the boundary between the sclera and the iris region) or the pupil-iris boundary to measure relative eye rotation. It should be noted that the limbus is not affected by the optical power of the cornea and thus tracking the limbus provides accurate results. In addition, the limbus creates a straight plane. The limbus defines the area of the connection between the muscles of the iris and thus there is a direct correlation between the three-dimensional parameters of the limbus and the three-dimensional parameters of the eye. Unfortunately, the limbus is more of a transition zone between the cornea and the sclera rather than a sharp boundary. Existing limbus trackers have not been successful because the limbus is a poor candidate for any technique that relies on so-called feature points or edge detection. Thus, the unclear edges and shapes that are distorted due to rotation result in difficulties in extracting the signal from the noise. As a result, techniques that rely on the detection of limbus edges or limbus feature points may lack the desired accuracy and precision. Although with known "pinpoint" limbus trackers using accurately aligned photodetectors precisely arranged along the iris / sclera interface, sufficient tracking response may also be possible, providing and / or maintaining such alignment adds additional system components and complexity, especially with respect to the variability of eye geometry between different patients. Some techniques for tracking the limbus rely on a set of two-quadrant detectors as position sensing photodetectors (PSDs). The simple fact that the iris surrounded by the limbus is darker than the sclera allows for the construction of a system consisting of one or more quadrant wires that examine the boundary of the iris / sclera region, i.e., the limbus. When the eye moves, the image of the iris also moves across the PSD, thus allowing the estimation of the eye movement that causes this. Although in fact it allows stabilizing the laser on the cornea region, it does not provide a direct indication of where the eye is looking and requires user calibration to obtain this data. In addition, many factors that can be controlled during eye surgery, such as eyelid movement, eyelashes, and even the slightest lighting changes, will affect the accuracy of this method and make it unsuitable for eye research. PSD-based limbus trackers measure the centroid of the spot position. This is because there is a contrast between the white sclera and the darker iris. If the eye moves relative to the system, the spot on the PSD also moves accordingly. Such a system only controls the image of the iris to be on the same spot on the PSD but does not measure and is not able to estimate the actual eye position.
[0017] Overview
[0018] The present disclosure provides the following:
[0019] 1). An eye tracking device including a processing unit, the processing unit being configured and operable to receive at least two images indicative of a user's eye, identify regions associated with the limbus in each image; determine a geometric representation of the limbus structure, and determine the three-dimensional position and gaze direction of the user's eye by triangulation of the geometric representation of the limbus structure in at least two images.
[0020] 2). The eye tracking device according to 1), wherein the geometric representation of the limbus structure includes an annular or elliptical structure.
[0021] 3). The eye tracking device according to 1) or 2), further comprising at least two imagers, each imager being configured to capture at least one image of the user's eye at different angles.
[0022] 4). The eye tracking device according to any one of the foregoing, wherein the processing unit includes a limbus detector, the limbus detector being configured and operable to receive data indicative of the limbus region and determine a geometric representation of the limbus structure by digital image preprocessing.
[0023] 5). The eye tracking system according to 4), wherein for each image, the limbus detector is configured and operable for digital image preprocessing, the digital image preprocessing including: calculating an image intensity gradient map of the limbus region; identifying at least one region of the limbus structure in which the local direction of the gradient is substantially consistent; processing the data indicative of the limbus structure by weighting the pixels of such regions; and generating a geometric representation of the limbus structure based on matching pixels associated with the eye limbus.
[0024] 6). The eye tracking device according to 4) or 5), further comprising a limbus triangulator, the limbus triangulator being configured and operable to receive data indicative of the geometric representation of the limbus structure for each image and perform triangulation of the geometric representations in at least two images to determine three-dimensional limbus parameters.
[0025] 7). The eye tracking system according to 6), wherein the processing unit is configured and operable to determine the three-dimensional position and gaze direction of the user's eye based on the three-dimensional limbus parameters.
[0026] 8). The eye tracking device according to 6) or 7), wherein the three-dimensional limbus parameters include at least one of the three coordinates of the center of the limbus, the direction of the normal to the limbus plane, the position of the limbus plane, and the size of the limbus.
[0027] 9). The eye tracking device according to any one of 6) to 8), wherein the limbus triangulator is configured and operable to determine the radius of the limbus and / or the torsional rotation of the limbus.
[0028] 10). The eye tracking device according to any one of the foregoing, wherein the processing unit includes a region detector configured and operable to receive each of the at least two images indicative of the user's eye and to identify image data indicative of the initial limbus region in each image by using an iterative pixel filtering process and generating data indicative of the initial limbus region.
[0029] 11). The eye tracking device according to 10), wherein the region detector is configured to identify image data indicative of the initial limbus region by at least one of the following operations: identifying image data indicative of eye features such as the pupil, eyelid, sclera, iris, and eyelashes and identifying the initial limbus region based on anatomical parameters.
[0030] 12). The eye tracking device according to 11), wherein the region detector is configured to identify image data indicative of eye features by using machine learning.
[0031] 13). The eye tracking device according to 11) or 12), wherein the region detector is configured and operable to identify the image data indicative of eye features by segmenting each image for identifying pixels associated with the pupil region.
[0032] 14). The eye tracking device according to 13), wherein the limbus triangulator is configured to perform triangulation between at least two geometric representations of the pupil to determine three-dimensional pupil parameters, the three-dimensional pupil parameters including at least one of the direction of the normal to the pupil plane, the coordinates of the center of the pupil, and the pupil diameter; estimate the position of the initial eye limbus region based on the three-dimensional pupil parameters; and generate data indicative of the position of the initial region of the limbus in each image.
[0033] 15). The eye tracking device according to any one of 6) to 14), wherein the limbus triangulator is configured to process the triangulated three-dimensional limbus, re-project the triangulated limbus region back to the image plane, and refine the limbus region at each image.
[0034] 16). An eye tracking method, the method comprising:
[0035] Receiving image data indicative of at least two images of the user's eye;
[0036] Identify the region associated with the limbus of the eye in each image;
[0037] Determine the geometric representation of the limbus structure;
[0038] Perform triangulation on the geometric representations of the limbus structures of at least two images to determine the three-dimensional position and gaze direction of the user's eye.
[0039] 17). The method as described in 16), including capturing at least two images of the user's eye at different angles.
[0040] 18). The method as described in 16) or 17), wherein the determination of the geometric representation of the limbus structure includes digital image preprocessing.
[0041] 19). The method as described in 18), wherein the digital image preprocessing includes: calculating an image intensity gradient map of the limbus based on the pupil direction; identifying at least one region of the limbus structure in each image, in which the local direction of the gradient is substantially consistent; processing the data indicating the limbus structure by weighting the pixels of such regions; and generating the geometric representation of the limbus structure based on the matching pixels associated with the limbus of the eye.
[0042] 20). The method as described in 19), including determining three-dimensional limbus parameters including at least one of the three coordinates of the center of the limbus and the direction of the limbus, and determining the three-dimensional position and gaze direction of the user's eye based on the three-dimensional limbus parameters.
[0043] 21). The method as described in 20), wherein the determination of the three-dimensional position and gaze direction of the user's eye includes processing the three-dimensional limbus parameters.
[0044] 22). The method as described in any one of 16) to 21), further including determining the radius of the limbus and / or the torsional rotation of the limbus.
[0045] 23). The method as described in any one of 16) to 22), wherein the geometric representation of the limbus structure includes an annular or elliptical structure.
[0046] 24). The method as described in any one of 16) to 23), further including identifying the image data indicating the initial limbus region in each image by using an iterative pixel filtering process and generating data indicating the initial limbus region.
[0047] 25). The method according to 24), wherein the image data for identifying the initial limbal region in each image includes at least one of the following operations: identifying image data indicative of eye features such as the pupil, eyelid, sclera, iris, and eyelashes, and identifying the initial limbal region based on anatomical parameters.
[0048] 26). The method according to 25), wherein the identifying of the image data indicative of eye features is performed by using machine learning.
[0049] 27). The method according to 25) or 26), wherein the identifying of the image data indicative of eye features includes segmenting each image to identify pixels associated with the pupil region.
[0050] 28). The method according to 27) further includes performing triangulation between at least two geometric representations of the pupil to determine three-dimensional pupil parameters, the three-dimensional pupil parameters including at least one of the direction of the normal to the pupil plane in each image, the coordinates of the center of the pupil, and the pupil diameter; estimating the position of the initial limbal region of the eye based on the three-dimensional pupil parameters; and generating data indicative of the position of the initial limbal region in each image.
[0051] 29). The method according to any one of 16) to 28) further includes processing the triangulated three-dimensional limbus, re-projecting the triangulated limbal region back to the image plane, and refining the limbal region at each image.
[0052] Eye tracking refers to the process of tracking the movement of an eye and determining where the user is looking. The different techniques described above roughly estimate the scale of eye movement and can indicate that movement has occurred. However, they do not measure the six degrees of freedom of eye movement. In addition, these systems do not track the visual features of the eye. One of the problems associated with tracking the visual features of the eye is that the most prominent features of the eye are located on the iris. Both the pupil and the iris are strongly distorted by the corneal lens. The shape of this lens varies from person to person, thus presenting a challenge in understanding exactly where these features are located. Astigmatism (corneal asymmetry) further complicates this problem. The limbus (i.e., the boundary between the white sclera and the darker iris) is not distorted by the cornea, but it appears very uniform, slightly blurred, and does not have any prominent features sought by classical image processing algorithms. Therefore, there is a need in the art for techniques that provide limbus-based eye tracking. The inventors have found that limbus eye tracking can be provided by triangulating the identified stable regions, which are real, i.e., directly corresponding to some directly visible anatomical or virtual features, i.e., the computed features.
[0053] In a broad aspect of the present invention, there is provided an eye tracking method, which includes receiving image data indicating at least two images of a user's eye, identifying regions associated with the limbus in each image; determining a geometric representation of the limbus structure, and determining the three-dimensional position and gaze direction of the user's eye (i.e., all six degrees of freedom of the eye) by triangulation of the geometric representations of the limbus structures of at least two images. Since the limbus maintains a fixed relationship and closed circularity relative to the lens of the cornea, the three-dimensional position and gaze direction of the user's eye can be obtained by triangulation of the geometric representations of the limbus structures of at least two images (i.e., two-dimensional eye limbus parameters). Shapes in at least two images are identified and matched. Triangulation can be used to determine the three-dimensional distance to the matching regions. Epipolar geometry techniques commonly used in this field are applied to match the corresponding regions, and optionally to determine three-dimensional eye limbus parameters. The three-dimensional position and gaze direction of the user's eye are determined based on the three-dimensional limbus parameters.
[0054] In this regard, it should be understood that typically stereo vision algorithms rely on matching corresponding points or edges between two images. There are two basic categories of methods: intensity-based techniques, in which the information in the intensity values of each line is matched at each pixel and the difference in intensity is minimized; and feature-based techniques, in which a set of pixel values and their spatial distribution are used to extract information. Commonly used features are simple features such as edge and corner detectors (Canny, Harris, LoG filters). Generally, eye tracking systems relying on the features mentioned above suffer from several drawbacks, such as different noises generated by each of the multiple photodetectors as a result of parasitic reflections on the cornea and distortion of iris and pupil features as a result of the curvature of the corneal lens. Therefore, eye tracking based on feature-based techniques does not provide sufficiently accurate results. In addition, methods based on matching corneal reflections with eye features require user calibration. Regarding limbus tracking, as mentioned above, the limbus does not have distinct features, and thus classical methods do not provide three-dimensional eye limbus parameters. It should be noted that the technique of the present invention does not require user calibration of the eye tracking device and provides accurate eye tracking with an accuracy and precision in the range of approximately 1 / 6 degree to 1 / 60 degree.
[0055] In some embodiments, the method includes capturing at least two images of the user's eye at different angles.
[0056] In some embodiments, determining a geometric representation of a limbal structure includes digital image preprocessing. The digital image preprocessing includes calculating an image intensity gradient map of the limbus based on the pupil orientation, identifying at least one region of the limbal structure in each image (where the local orientation of the gradient is substantially consistent), processing data indicative of the limbal structure by weighting pixels of such regions, and generating a geometric representation of the limbal structure based on matching pixels associated with the eye limbus.
[0057] In some embodiments, the method includes determining three-dimensional limbal parameters, which include three coordinates of the center of the limbus and at least one of a normal direction of the limbal plane, a position of the limbal plane, and a size of the limbus. Determining the three-dimensional position and gaze direction of the user's eye can be based on the three-dimensional limbal parameters.
[0058] In some embodiments, the method further includes determining a radius of the limbus and / or a torsional rotation of the limbus.
[0059] In some embodiments, the method further includes identifying image data indicative of an initial limbal region in each image by using an iterative pixel filtering process and generating data indicative of the initial limbal region.
[0060] In some embodiments, identifying image data indicative of an initial limbal region in each image includes at least one of the following operations: identifying image data indicative of eye features such as a pupil, eyelids, sclera, iris, and eyelashes and identifying the initial limbal region based on anatomical parameters. Identifying image data indicative of eye features can be performed by using machine learning.
[0061] In some embodiments, identifying image data indicative of eye features includes segmenting each image to identify pixels associated with the pupil region.
[0062] In some embodiments, the method further includes performing triangulation between at least two geometric representations of the pupil to determine three-dimensional pupil parameters, estimating a position of an initial eye limbal region based on the three-dimensional pupil parameters, and generating data indicative of the position of the initial region of the limbus in each image. The three-dimensional pupil parameters include at least one of a direction of a normal of the pupil plane in each image, coordinates of the center of the pupil, and a pupil diameter.
[0063] In some embodiments, the method further includes processing the triangulated three-dimensional limbus, re-projecting the triangulated limbal region back to the image plane, and refining the limbal region at each image.
[0064] According to another broad aspect of the present invention, there is provided an eye tracking device including a processing unit configured and operable to receive at least two images indicative of a user's eye, identify regions associated with the limbus of the eye in each image; determine a geometric representation of the limbus structure, and determine the three-dimensional position and gaze direction of the user's eye by triangulation of the geometric representations of the limbus structures of at least two images. The processing unit is configured and operable to determine the three-dimensional position and gaze direction of the user's eye based on three-dimensional limbus parameters. The eye tracking device of the present invention can be a head-mounted eye tracking device, such as an eye tracking glasses device or an eye tracker integrated in a helmet or also integrated in a head-mounted display device, in virtual reality glasses, in augmented reality glasses or other head-mounted devices. The eye tracking device can also be a remote eye tracker optionally integrated or coupled with other devices such as a computer, a display, a monitor, etc.
[0065] In some embodiments, the eye tracking device further includes at least two imagers, each imager being configured to capture at least one image of the user's eye at different angles.
[0066] The imager can be a capture unit and can include one or more cameras (i.e., digital cameras or video cameras), light sensors such as image sensors, e.g., a CCD sensor or a CMOS sensor. One imager can be used and moved quickly along a known trajectory at a known speed. Optionally, an optical image with a certain type of mirror can also be used. In some embodiments, the present invention uses a stereoscopic imager to use two or more imagers with precisely known relative offsets. Two or more imagers can capture images simultaneously or with a known time delay (e.g., in an interleaved mode).
[0067] Generally, the processing unit can be a processor, a controller, a microcontroller or any kind of integrated circuit. The eye tracking device can be associated with an optical system, which can include any optical elements, such as one or more lenses, prisms, beam splitters, mirrors, reflectors, light guides, collimators, etc. The eye tracking device is also associated with a light source configured to produce as much diffused illumination as possible, preferably a light source that does not produce any pattern on the eye surface or minimizes the number of patterns. The eye tracking device includes a processing unit that communicates with each of the at least two imagers to receive the at least two images.
[0068] The new technology of the present invention uses two or more captured images of an eye and detects three-dimensional limbal parameters in each captured image through an iterative pixel filtering process, thereby determining the orientation of the eyeball by triangulation of the two-dimensional limbal parameters of the eye obtained from the images. The new technology determines the eye orientation based on the recognition of the limbus in the user's eye in the image data. Two or more imagers can be used to capture two or more corresponding images of the user's eye. Each imager is configured to capture an image of the user's eye from a different angle. Each imager images the user's iris and adjacent and surrounding regions (including the pupil, sclera, eyelids, etc.). For triangulation of the limbal region, multiple images of the eye are preferably captured from two or more different angles in close temporal proximity (e.g., simultaneously).
[0069] In some embodiments, the processing unit includes a region detector that is configured to receive each of at least two images indicative of the user's eye and identify, in each image, image data indicative of an initial limbal region, i.e., some preferably initially undistorted (i.e., not distorted by the eye optics itself such as the cornea) region / zone (i.e., a real or virtual computational region) or a region having a known distortion corresponding to the same physical region of the eye, by using an iterative pixel filtering process and generating data indicative of the initial limbal region. The inventors have found that these identified stable regions can be used for triangulation calculations. The present invention allows for the generation of such stable regions corresponding to the same physical region of the eye and the identification of them on a stereo image. The region detector is configured to identify the image data indicative of the initial limbal region by at least one of the following operations: identifying the image data indicative of eye features such as the pupil, eyelids, sclera, iris, and eyelashes and identifying the initial limbal region based on anatomical parameters. These regions can be identified by projecting a geometric representation of an eye feature (e.g., the pupil) onto each image and then capturing the regions that match between the images at different angles. The larger the size of such a region on the image, the more precisely the position of the corresponding physical region is determined by triangulation, and thus the position of the eye itself is more precise. The region detector generates a function that defines a two-dimensional virtual object on the eye by using an iterative pixel filtering process capable of performing triangulation. Such triangulation defines the limbal orientation corresponding to the gaze direction. The limbal parameters are defined as a mathematical equation in which the variables (x, y, z, α, β, and δ) represent six degrees of freedom of movement of a rigid body in three-dimensional space. In particular, x, y, z define changes in position as forward / backward (fluctuation), up / down (heave), and left / right (roll) translations on three perpendicular axes, combined with changes in orientation by rotation about the three perpendicular axes, yaw (normal axis), pitch (transverse axis), and roll (longitudinal axis).
[0070] In this regard, it should be understood that determination of three-dimensional limbus parameters can locate such undistorted regions or regions with known distortion corresponding to the same physical region of the eye, since the size of the limbus is large and is not actually distorted by the cornea. However, as mentioned above, the limbus itself has a boundary that appears somewhat blurred and is not well-defined, which looks different on each image and which has no distinguishable features that can be used as an anchor for such a match. Thus, matching the limbus region itself across two or more different images poses a challenge. The inventors have found that rather than attempting to match the features of the limbus region itself, digital image preprocessing can be performed to allow calculation of some virtual regions based on the limbus three-dimensional parameters of each image that do indeed well correspond to the same physical region of the eye. Digital image preprocessing includes the use of complex algorithms such as classification, feature extraction, multi-scale signal analysis, pattern recognition or projective mathematical transformation. The virtual regions should correspond to the same physical region of the eye, independent of the angle at which the image was taken. The region detector is thus configured and operable to identify image data indicative of the limbus region in each image and generate data indicative of the limbus region. This can be achieved by estimating the limbus region by means of using the anatomical relationship between the identified virtual region corresponding to certain eye features and the limbus. For example, the position and center of the pupil can be calculated by triangulation of the pupil region (i.e., the region with a defined circular contrast) between two images, and then the limbus region can be defined as a circular region around the center of the pupil with a certain radius range (e.g., approximately 5.5 mm to 6.5 mm). The radius range can be based on the statistics of a Gaussian distribution defining the distance between the limbus and the center of the pupil.
[0071] In some embodiments, pattern recognition tools are used to identify the pixels of the pupil and the pixels of the eyelid in each image. Due to a certain contrast in the image, the image pixels of the pupil can be identified. The identified image region corresponding to the pupil is discarded, and the image region of the eyelid can be used to determine calibration data (e.g., the distance / angle, pupil / limbus center, and approximate orientation of each imager). More particularly, the data image related to the possible intersection between the eyelid and the estimated limbus region is discarded, thus further reducing the number of potential limbus pixels. Then, two-dimensional eye limbus parameters are identified based on the calibration data and the three-dimensional model of the eye. Then, an iterative filtering process is used to discard the pixels located at the boundary of the limbus in each of two or more captured images. The geometric representation of the limbus structure includes an annular or elliptical structure. When it is assumed that the limbus is circular, the limbus two-dimensional parameters obtained from each image are then used to define a complete limbus ring / ellipse structure function for each image (reconstructed after discarding the eyelid region that hides the limbus region), and the eye orientation is determined by triangulation of the defined limbus ring / ellipse structure obtained from each image. Filtering out the image data corresponding to certain eye features can discard certain pixels and improve the resolution. If the technique fails to determine the eye orientation, the process is repeated again while the limbus region is increased to potentially accommodate the limbus-related pixels.
[0072] In some embodiments, the method includes the following steps: First, isolate the limbus region, and then perform a mathematical transformation, such as finding the geometric representation of the limbus structure in each of the images (e.g., by projecting a set of elliptical functions through the middle of the limbus ring), and triangulating the corresponding circles in six degrees of freedom (6DOF) that are strongly aligned with the physical eye.
[0073] In some embodiments, the processing unit includes a limbus detector that is configured and operable to receive data indicating the limbus region and determine the geometric representation of the limbus structure through digital image preprocessing, such as performing a limbus recognition process on each image and / or performing a mathematical transformation, including using an intensity gradient map. More particularly, the digital image preprocessing may include calculating the image intensity gradient map of the limbus region, identifying at least one region of the limbus structure (where the local direction of the gradient is substantially consistent), and processing the data indicating the limbus structure by weighting the pixels of these regions and determining the geometric representation of the limbus structure based on the matching pixels related to the eye limbus.
[0074] More particularly, after the limbal region has been estimated, the annular region is projected onto the eye surface at the estimated limbal region. Then, based on the estimated anatomical properties of a normal eye and optionally the calculated three-dimensional pupil position, most of the pixels of the image are discarded, such as the pupil, most of the iris, most of the sclera, and most of the eyelids, leaving only the ring around the potential limbal position. However, due to the corneal distortion of the pupil image, additional estimation errors are possible. This distortion results in an increase in the initial estimate of the limbal region.
[0075] In some embodiments, the region detector is configured to identify image data indicative of eye features by using machine learning. An iterative pixel filtering process is used to train a neural network to identify the limbal region in each acquired image. The region detector can be configured and operable to identify image data indicative of eye features by segmenting each image to identify the pixels associated with the pupil region. Thus, the neural network can be used to directly segment the approximate limbal region. More particularly, to estimate the limbal region, the pupil can be tracked. Additionally or alternatively, the neural network can also be used to directly estimate the positions of other eye features (such as the eyelids and the sclera) without going through pupil tracking. The neural network can be applied to the image itself or on a mathematically transformed image, i.e., a gradient map. The neural network can be trained with traditional machine learning algorithms, for example, based on the results of other eye tracking techniques where the eye feature characteristics have been previously defined.
[0076] In some embodiments, the processing unit is configured and operable to determine the size of the limbus.
[0077] In some embodiments, the pixels of the identified eyelid region are triangulated to improve the estimation of the position of the eye.
[0078] In some embodiments, the eye tracking device further includes a limbus triangulator, which is configured and operable to receive data indicative of a geometric representation of the limbal structure for each image and perform triangulation of the geometric representations of at least two images to determine three-dimensional limbal parameters. The three-dimensional limbal parameters include at least one of the three coordinates of the center of the limbus, the direction of the normal to the plane of the limbus, the position of the plane of the limbus, and the size of the limbus. The limbus triangulator can be configured and operable to determine the radius of the limbus and / or the torsional rotation of the limbus.
[0079] In some embodiments, a limbus triangulator is configured to perform triangulation between at least two geometric representations of a pupil to determine three-dimensional pupil parameters, the three-dimensional pupil parameters including at least one of a direction of a normal of a pupil plane, coordinates of a center of the pupil, and a pupil diameter; estimate a position of an initial eye limbus region based on the three-dimensional pupil parameters; and generate data indicative of the position of the initial region of the limbus in each image.
[0080] In some embodiments, a limbus triangulator is configured to process a triangulated three-dimensional limbus, re-project the triangulated limbus region back onto an image plane, and refine the limbus region at each image. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] To better understand the subject matter disclosed herein and to illustrate how the subject matter may be carried out in practice, embodiments will now be described, by way of non-limiting example only, with reference to the accompanying drawings, in which:
[0083] Figure 1 is a schematic block diagram showing the main functional parts of an eye tracking device according to the present invention;
[0084] Figure 2 is a schematic flow chart showing the main steps of an eye tracking method according to the present invention;
[0085] Figure 3 shows two stereoscopic images of an eye captured at different angles;
[0086] Figures 4A - 4D is an example of a pupil detection stage according to some embodiments of the present invention;
[0087] Figure 5 is according to some embodiments of the present invention in Figure 3 is an example of an eye feature detection stage illustrated on an image of;
[0088] Figure 6 is according to some embodiments of the present invention in Figure 3 is an example of a projection of a limbus circle illustrated on an image of;
[0089] Figure 7 is a schematic block diagram showing possible functional parts of an eye tracking device according to some embodiments of the present invention;
[0090] Figure 8 represents a two-dimensional eye tracking trajectory obtained by using the teachings of the present invention;
[0091] Figure 9 represents two angular degrees of freedom (i.e., angular measurements) of an eye tracking orientation trajectory in terms of azimuth and pitch as a function of time; and
[0092] Figures 10A - 10B Represent five degrees of freedom of an eye tracking trajectory obtained by using the teachings of the present invention.
[0093] Detailed description of embodiments
[0094] Reference Figure 1 , which illustrates a block diagram of the main functional parts of the eye tracking device of the present invention. The eye tracking device 100 includes a processing unit 106, which is configured and operable to receive at least two images indicative of a user's eye, identify regions associated with the limbus in each image, determine a geometric representation of the limbus structure, and determine the three-dimensional position and gaze direction of the user's eye by triangulation of the geometric representations of the limbus structures of at least two images. In this regard, it should be noted that, as detailed above, the present invention triangulates on the shape (geometric representation) rather than on individual corresponding points or edges.
[0095] The processing unit 106 is generally configured as a computing / electronic utility, which particularly includes utilities such as data input and output modules / utilities 106A and 106B, a memory 106D (i.e., a non-volatile computer-readable medium), and an analyzer / data processing utility 106C. Thus, the utilities of the processing unit 106 can be implemented by suitable circuitry and / or by software and / or hardware components, the software and / or hardware components including computer-readable code configured to implement Figure 2 the operations of the method 200 shown and described below.
[0096] Features of the present invention may include general or special purpose computer systems, which include various computer hardware components discussed in more detail below. Features within the scope of the present invention also include computer-readable media for executing or having stored thereon computer-executable instructions, computer-readable instructions, or data structures. Such computer-readable media can be any available media accessible by a general or special purpose computer system. By way of example and not limitation, such computer-readable media may include physical storage media such as RAM, ROM, EPROM, flash drives, CD-ROMs, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of computer-executable instructions, computer-readable instructions, or data structures and that can be accessible by a general or special purpose computer system. Computer-readable media may include computer programs or computer applications downloadable to a computer system via a network such as a wide area network (WAN), such as the Internet.
[0097] In this specification and in the following claims, a "processing unit" is defined as one or more software modules, one or more hardware modules, or a combination thereof that work together to perform operations on electronic data. For example, the definition of a processing utility includes the hardware components of a personal computer as well as software modules such as the operating system of a personal computer. The physical layout of the modules is not relevant. A computer system can include one or more computers coupled via a computer network. Similarly, a computer system can include a single physical device in which internal modules (such as memory and a processor) work together to perform operations on electronic data. Although any computer system can be mobile, the term "mobile computer system" or the term "mobile computer device" as used herein particularly includes laptop computers, netbook computers, cellular phones, smartphones, wireless telephones, personal digital assistants, portable computers with touch-sensitive screens, and the like. The processing unit 106 can consist of a processor embedded therein or attached thereto that runs a computer program. A computer program product can be embodied in one or more computer-readable media having computer-readable program code embodied thereon. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer program code for performing the operations of aspects of the present invention can be written in any combination of one or more programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer via any type of network connection, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider). These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. The specified functions of the processor can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0098] The eye tracking device 100 can include at least two imagers 110, each imager being configured to capture at least one image of the user's eye at a different angle. Each imager can be focused on the user's iris. In this regard, it should be noted that the limbal tracking of the present invention is performed independently for each eye.
[0099] In some embodiments, the processing unit 106 includes a limbus detector 102 that is configured and operable to receive data indicative of a limbus region and determine a geometric representation of a limbus structure (e.g., an annular or elliptical structure) through digital image preprocessing. In a specific and non-limiting example, the limbus detector 102 is configured and operable to perform a limbus recognition process on each image and / or perform a mathematical transformation of the image, including using an intensity gradient map and then running a limbus region recognition process on the transformed image. In other words, the limbus detector 102 can transform the image into an intensity gradient map, and then the limbus region is searched for on the transformed image. Instead of or in addition to the gradient map, an entropy map can also be used, or the processing can be performed directly on the image. If an intensity gradient map is used, the limbus detector 102 can calculate an intensity gradient map of the limbus region of the image, identify at least one region of the limbus structure (where the local direction of the gradient is substantially consistent), process the data indicative of the limbus structure by weighting the pixels of such a region, and determine a geometric representation of the limbus structure based on the matching pixels associated with the limbus of the eye. In addition to collinearity, knowledge of the anatomical form of the limbus can also be used, i.e., the collinear vectors should extend radially from the center of the limbus and the entire limbus region is continuous. More particularly, the region of the limbus structure is determined by identifying local consistency, i.e., the gradient direction of each point is collinear only with its neighboring points.
[0100] Each module of the processing unit 106, such as the limbus detector 102, is not limited to a specific number of modules and can be configured to process multiple images independently, simultaneously, or not. The limbus triangulator 108 is configured and operable to perform triangulation of the geometric representation of the limbus structure in at least two images to determine three-dimensional limbus parameters. The limbus triangulator 108 can also be configured and operable to determine the radius of the limbus and / or the torsional rotation of the limbus. The three-dimensional limbus parameters define at least five degrees of freedom of the position and orientation of the limbus. The three-dimensional limbus parameters include the position of the plane of the limbus and / or the direction of the normal to the limbus plane and / or the size of the limbus and / or the center of the limbus or any combination thereof.
[0101] In some embodiments, the processing unit 106 includes a region detector 104 that is configured and operable to receive each of at least two images indicative of an eye of each user and to identify image data indicative of an initial limbal region in each image by using an iterative pixel filtering process and generating data indicative of the initial limbal region. This can be achieved by identifying image data indicative of eye features such as the pupil, eyelid, sclera, iris, and eyelashes and / or by identifying the initial limbal region based on anatomical parameters. More particularly, identifying the initial limbal region can include pupil segmentation and triangulation of the pupil and / or limbal region estimation based on anatomical and / or eyelid region estimation and / or limbal region estimation without an eyelid region. This can be achieved by a limbal triangulator 108 or any other processing module. However, the present invention is not limited to any particular module configuration of the processing unit 106. The processing unit 106 can include a pupil detector, a pupil triangulator, and other modules. For example, the identification of image data indicative of eye features can include pupil segmentation. Thus, the region detector 104 can be a pupil detector that is designed to provide an initial limbal ring. It is assumed that the two regions - the pupil and the limbus - are concentric and in the same plane. Pupil segmentation can be achieved by segmenting each image to identify pixels associated with the pupil region, performing triangulation between at least two images to determine three-dimensional pupil parameters, estimating the position of the initial eye limbal region based on the three-dimensional pupil parameters, and generating data indicative of the position of the initial limbal region in each image. The three-dimensional pupil parameters include the direction of the normal of the pupil plane and / or the coordinates of the center of the pupil and / or the pupil diameter. The three-dimensional pupil parameters are used to estimate three-dimensional limbal parameters and thus define an initial limbal region that can be further refined. The identification of image data indicative of eye features can also include the identification and filtering out of the eyelid from the image data because the image data associated with the eyelid blocks a portion of the limbal ring. The present invention is not limited to a particular manner of identifying image data indicative of the initial limbus. Neural network-based methods can be used to identify image data indicative of the initial limbus. For example, limbal region estimation can be performed with a neural network, but a pupil detector can be used to estimate the eyelid region with a classical algorithm. Optionally, limbal region estimation and eyelid region estimation can be performed with neural network-based methods. Optionally, the neural network can be used only for pupil segmentation.
[0102] In some embodiments, the region detector 104 is configured to identify image data indicative of eye features by using machine learning. Machine learning is capable of self-classifying / learning the characteristics of input data related to eye features. The region detector 104 may use a neural network-based data recognition model to predict three-dimensional limbus parameters of the eye. The network may be trained based on the segmentation results of classical methods or by using an existing system for training. Additionally or alternatively, the three-dimensional limbus parameters obtained by the eye tracking device of the present invention may be used to train the neural network. A deep learning network (DLN) implemented by the region detector 104, such as an artificial neural network (ANN), may generate a representation of the limbus based on a series of images. For example, the representation generated by the DLN may include the probability of the placement of the limbus. This representation is used to generate a model that adapts the recognition classifier to the limbus of a person.
[0103] After the limbus detector 102 has determined the three-dimensional limbus parameters, the region detector 104 may receive the three-dimensional limbus parameters to further generate more precise data indicative of the limbus region. For example, the limbus triangulator 108 is configured to process the triangulated three-dimensional limbus, re-project the triangulated limbus region back onto the image plane, and refine the limbus region at each image.
[0104] Reference Figure 2 , which illustrates a flowchart of the main steps of the eye tracking method of the present invention. The method 200 for eye tracking includes receiving, at 204, image data indicative of at least two images of a user's eyes; identifying, at 206, regions associated with the eye limbus in each image; determining, at 208, a geometric representation of the limbus structure; and performing, at 210, triangulation of the geometric representation of the limbus structure of at least two images, thereby determining, at 212, the three-dimensional position and gaze direction of the user's eyes.
[0105] In some embodiments, step 210 includes performing triangulation of the geometric representation of the limbus structure in at least two images to determine the three-dimensional limbus parameters as described above. Determining the three-dimensional limbus parameters may include determining the radius of the limbus and / or the torsional rotation of the limbus.
[0106] In some embodiments, after the triangulation of the geometric representation of the limbus structure of at least two images in 210, the method 200 may include determining the geometric representation of the limbus structure by processing the three-dimensional limbus parameters to generate more precise data indicative of the limbus region.
[0107] In some embodiments, the method 200 may further include capturing, at 202, at least two images of each user's eyes at different angles.
[0108] In some embodiments, determining a geometric representation of the limbal structure in 208 may include digital image preprocessing in 218, such as performing a limbus identification process on each image or performing a mathematical transformation of the image, including using an intensity gradient map, and then running a limbal region identification process on the transformed image. The digital image preprocessing step in 218 may include calculating an image intensity gradient map of the limbal region, identifying at least one region of the limbal structure (where the local direction of the gradient is substantially consistent), processing data indicative of the limbal structure by weighting the pixels of such regions, and determining a geometric representation of the limbal structure based on matching pixels associated with the limbus of the eye. More particularly, the digital image preprocessing may include determining regions of the limbal structure by identifying local consistency (i.e., the direction of the gradient at each point is collinear only with its neighboring points). As described above, an entropy map may also be used as the intensity gradient map (instead of or in addition to the gradient map), or the processing may be performed directly on the image.
[0109] In some embodiments, identifying regions associated with the limbus of the eye in each image in 206 may include identifying image data indicative of eye features such as the pupil, eyelids, sclera, iris, and eyelashes, and / or identifying an initial limbal region based on anatomical parameters by using an iterative pixel filtering process and generating data indicative of the initial limbal region. For example, eyelid region estimation may be implemented in parallel with pupil and limbal region estimation. The present invention is not limited to the order of estimation of different regions. Different eye feature regions may be determined based on any geometric shape model (such as an ellipse or a circle).
[0110] In some embodiments, identifying the initial limbal region may include pupil segmentation and triangulation of the pupil and / or limbal region estimation based on anatomical and / or eyelid region estimation and / or limbal region estimation without an eyelid region. The method may include segmenting each image to identify pixels associated with the pupil region, performing triangulation between at least two images to determine three-dimensional pupil parameters, estimating the position of the initial limbal region of the eye based on the three-dimensional pupil parameters, and generating data indicative of the position of the initial region of the limbus in each image. Identifying image data indicative of eye features may also include identifying the eyelids and filtering out such image data since the image data associated with the eyelids blocks a portion of the limbal ring. As described above, the present invention is not limited to any particular manner of identifying image data indicative of the initial limbus. Neural network-based methods may be used to identify image data indicative of the initial limbus. Neural network-based methods may also be used to perform limbal region estimation and eyelid region estimation. The neural network may also be used only for pupil segmentation.
[0111] In some embodiments, identifying image data indicative of eye features can be achieved by using machine learning. Determining (i.e., predicting) three-dimensional limbus parameters can include using a neural network-based data recognition model. The method can include segmenting based on classical methods or training the network by using an existing system for training. A representation of the limbus can be obtained from the series of images by using a deep learning network (DLN) such as an artificial neural network (ANN). For example, generating a representation of the limbus can include calculating probabilities regarding the placement of the limbus and / or generating a model of the limbus, adapting the recognition classifier to a person.
[0112] Reference Figure 3 , which shows two stereo images of a user's eye taken at different angles.
[0113] Reference Figures 4A - 4D , showing an example of data processing for pupil region detection according to some embodiments of the present invention Figure 3 of two stereo images. Figure 4A shows a segmentation stage, where each image is segmented to identify pixels related to the pupil region. Pixels not related to the pupil region are discarded, and the resolution is enhanced, as Figure 4B shown. Figure 4C shows the segmentation of the pupil boundary based on the intensity map of each image. Figure 4D shows the refinement of the pupil boundary based on the projection of a geometric representation of the pupil (e.g., a circular shape) on each image. Then, the region of Figure 4D is triangulated between the image and the pupil diameter, and the 3D position and orientation are determined.
[0114] Reference Figure 5 , which shows an example of data recognition of eye features of two stereo images according to some embodiments of the present invention Figure 3 . The marked regions represent regions where the local directions of the image intensity gradients are substantially consistent.
[0115] Reference Figure 6 , which shows an example of a limbus recognition process based on the projection of an elliptic function on each of two stereo images according to some embodiments of the present invention Figure 3 .
[0116] Reference Figure 7, which is a block diagram illustrating possible functional parts of the eye tracking device of the present invention. In this particular and non-limiting example, the eye tracking device 300 includes two imagers, referred to herein as cam 1 and cam 2, which are configured to capture at least one image of each eye of the user at different angles. Each imager focuses on the iris of the user. The eye tracking device 300 includes a processing unit 310, which includes a pupil detector 302 and a limbus detector 304. The pupil detector 302 is a specific example of the region detector defined above with respect to Figure 1 The pupil detector 302 includes a plurality of modules and is configured and operable to receive each of at least two images of each eye of the user and identify the image data indicative of the limbus region in each image and generate data indicative of the limbus region. This can be achieved by using eye feature detection such as pupil region detection or by using a neural network. Each image is segmented by a segmentation module to identify the pixels associated with the pupil region. The pixels not associated with the pupil region are discarded in each image and the resolution is enhanced. Then, each ellipse module Ell calculates the elliptical curve corresponding to the pupil region on each image to find the matching elliptical curve. A three-dimensional equation indicative of the three-dimensional pupil region is generated for each image. Then, when the distance between the two imagers capturing the two stereo images is known, a first triangulation of the three-dimensional equations of the two stereo images is performed to provide the direction of the pupil. After that, using the average size of the limbus and assuming that the pupil and the limbus are concentric and in the same plane, the algorithm produces an estimate of the three-dimensional limbus parameters. This three-dimensional limbus region is projected onto each image and is referred to herein as the mask equation as shown in Figure 3 . For example, the limbus region can be delimited by two elliptical curves. The mask equation is used by the collinearity mapping module that gives the first estimate of the limbus projection on each image. The mask equation is also used by the weight module.
[0117] To determine three-dimensional limbus parameters, a collinearity module, referred to herein as collinearity mapping, performs a mathematical transformation. In a particular and non-limiting example, such a mathematical transformation can include calculating an image intensity gradient map of the limbus based on the pupil orientation, identifying at least one region of the limbus structure (where the direction of the gradient is substantially consistent), processing data indicative of the limbus structure by weighting the pixels of such a region, and determining three-dimensional limbus parameters based on matching pixels associated with the eye limbus. Optionally, such a mathematical transformation can include using the entropy of the image. In a particular and non-limiting example, the collinearity mapping module generates a gradient collinearity map from the original image masked by the estimated limbus region (calculated from the pupil 3D parameters and the average limbus size). The limbus detector 304 receives a three-dimensional equation indicative of the three-dimensional limbus region and the orientation of the limbus. The annular region is then projected onto each eye image to produce the estimated limbus region. Then, based on the estimated anatomical properties of a normal eye, most of the pixels of the image are discarded by using a masking equation, such as the pupil, most of the iris, most of the sclera, and most of the eyelids, leaving only a ring around the potential limbus location.
[0118] Based on the module weights of the collinearity mapping module, a weight is assigned to each pixel indicating the probability that the pixel belongs to the limbus region. The module weights are then used to refine the potential limbus region and provide an input to the Ell module that searches for a matching ellipse estimate of the limbus.
[0119] To perform a second triangulation, several techniques can be implemented, such as triangulation of a cone or quadratic estimation (linear or non-linear). The processing unit 310 is configured and operable to determine the three-dimensional position and gaze direction of the user's eye by triangulating the geometric representation of the limbal structure of at least two images by means of using linear or non-linear quadratic estimation. In this particular and non-limiting example, the triangulation of the cone can be performed as follows: Each ellipse module Ell generates a mathematical function / equation that defines the geometric representation of the limbus that defines the anatomical relationship between the identified virtual regions corresponding to certain eye features (e.g., an ellipse having a center corresponding to the pupil center), and the limbus is calculated for each image. The elliptical structures of each image are compared, and the intersection between the cones defined by the elliptical structures defines an ellipse in space, the projection of which matches the elliptical structures of the two images. This defines the three-dimensional limbal parameters referred to herein as the limbus R, Cx, Cy, Cz, azimuth, and pitch. In this regard, it should be noted that in the present invention, the three-dimensional limbal parameters are obtained by triangulation of the ellipse rather than by pixel triangulation. As described above, pixel triangulation of the limbus is practically impossible because the correspondence between the pixels of different images is unknown (since most of the pixels along the epipolar line of the second image look too similar). In classical computer vision, special near-pixel region features, such as different colors, corners, line intersections, or texture patches, are used, which allow the characterization of the match between two pixels along the epipolar line from different images. Due to the rough boundaries of the limbus, which are usually defined as wide regions, such features are missing in the limbal region and thus do not allow the use of classical natural feature tracking algorithms. A mask is then used to discard the regions related to the eye features. An iterative pixel filtering process is then used, repeating the same process as described above to more accurately define the limbal region and reduce the limbal region size.
[0120] Optionally or in addition, the iterative pixel filtering process can be implemented by using linear quadratic estimation such as a Kalman filter or non-linear quadratic estimation such as an extended Kalman filter and an unscented Kalman filter. The Kalman filter produces an estimate of the limbal equation with five degrees of freedom (e.g., the elliptical structure) as the average of the predicted state of the eye and the average of the new measurements using a weighted average. The weights are calculated from the covariance - a measure of the estimated uncertainty of the prediction of the state of the eye. The result of the weighted average is a new state estimate that lies between the predicted state and the measured state and has a better estimated uncertainty than either state taken alone. This process is repeated at each time step, and the new estimate and its covariance provide information for the prediction used in the next iteration.
[0121] At the end of this process, the processing unit 310 determines the three-dimensional position and gaze direction of the user's eye based on the three-dimensional limbus parameters.
[0122] Reference Figure 8 , which illustrates a two-dimensional eye-tracking position trajectory obtained by using the teachings of the present invention, where different colors distinguish different parts of the trajectory, i.e., the change of color over time. The axes Xc and Yc are measured in millimeters.
[0123] Reference Figure 9 , which illustrates the Figure 8 eye-tracking orientation trajectory in terms of azimuth and pitch as a function of time (in milliseconds), i.e., the angular measurements (measured in degrees) of two angular degrees of freedom of the eye-tracking orientation trajectory. Figure 8 , 9 and 10 both show data from the same recording.
[0124] Reference Figures 10A - 10B , which illustrates the Figure 8 five degrees of freedom of the eye-tracking trajectory. More particularly, Figure 10A shows the two angular degrees of freedom of the eye-tracking orientation trajectory in terms of azimuth and pitch as a function of time. Figure 10A is an enlarged view of a partial segment of the graph above Figure 9 . Figure 10B shows the three-dimensional degrees of freedom of the eye-tracking position trajectory in a Cartesian coordinate space as a function of time (measured in milliseconds).
Claims
1. An eye tracking device, the eye tracking device comprising a processing unit, the processing unit including a limbus detector configured to and operative to receive and process image data indicative of at least one image of a user's eye, wherein, the processing of the image data of the at least one image includes: identifying in the at least one image at least one region associated with the limbus of the eye, the limbus of the eye being defined as the transition zone between the cornea and the sclera of the user's eye; and applying digital preprocessing to data indicative of the transition zone and determining a geometric representation of the limbus structure.
2. The eye tracking device according to claim 1, wherein, applying the digital preprocessing to the data indicative of the transition zone includes: performing a limbus identification process and generating data indicative of the limbus region; and analyzing the data indicative of the limbus region and determining the geometric representation of the limbus structure.
3. The eye tracking device according to claim 2, wherein, performing the limbus identification process includes directly searching for the limbus region on the at least one image.
4. The eye tracking device according to claim 2, wherein, performing the limbus identification process includes mathematically transforming the image using at least one of an intensity gradient map and an entropy map and searching for the limbus region on the transformed image.
5. The eye tracking device according to claim 1, wherein, the limbus detector is configured to and operative to perform the digital preprocessing, the digital preprocessing including: calculating an image intensity gradient map of the transition zone, identifying in the image intensity gradient map at least one region where the local direction of the image intensity gradient is substantially consistent as the limbus region, weighting the pixels of the at least one region, and determining the geometric representation of the limbus structure based on the matching pixels associated with the limbus region.
6. The eye tracking device according to any one of claims 1 to 5, further comprising a region detector configured to and operative to receive at least one image indicative of a user's eye, identify image data indicative of an initial limbus region in the at least one image by using an iterative pixel filtering process, and provide the data indicative of the initial limbus region to the limbus detector.
7. The eye tracking device according to claim 6, wherein, the region detector is configured to identify the image data indicative of the initial limbus region by at least one of the following: identifying image data indicative of eye features including a pupil, eyelids, sclera, iris, and eyelashes; and identifying the initial limbus region based on anatomical parameters.
8. The eye tracking device according to claim 7, wherein, the region detector is configured to identify the image data indicative of the eye features by using machine learning.
9. The eye tracking device according to claim 7, wherein, The area detector is configured and operable to identify the image data indicative of the eye feature by segmenting each image to identify pixels associated with the pupil area.
10. The eye tracking device according to any one of claims 1 to 5, wherein, the geometric representation of the limbal structure includes an annular or elliptical structure.
11. An eye tracking method, the method comprising: receiving image data indicative of at least one image of a user's eye; identifying at least one area in the image data associated with the limbus of the eye, the limbus of the eye being defined as the transition zone between the cornea and the sclera of the user's eye; and digitally preprocessing the data indicative of the transition zone to determine a geometric representation of the limbal structure.
12. The method according to claim 11, wherein, performing the digital preprocessing on the data indicative of the transition zone includes: performing a limbus identification process and generating data indicative of the limbal area; and analyzing the data indicative of the limbal area and determining the geometric representation of the limbal structure.
13. The method according to claim 12, wherein, performing the limbus identification process includes directly searching for the limbal area on the at least one image.
14. The method according to claim 12, wherein, performing the limbus identification process includes mathematically transforming the image using at least one of an intensity gradient map and an entropy map, and searching for the limbal area on the transformed image.
15. The method according to claim 11, wherein, performing the digital preprocessing includes: calculating an image intensity gradient map of the transition zone; identifying the limbal area in the image intensity gradient map as at least one area where the local direction of the image intensity gradient is substantially consistent, weighting the pixels of the at least one area, and determining the geometric representation of the limbal structure based on the matching pixels associated with the limbal area.
16. The method according to any one of claims 11 to 15, further comprising applying an iterative pixel filtering process to at least one image of the user's eye to identify image data indicative of an initial limbal area; and generating data indicative of the initial limbal area.
17. The method according to any one of claims 11 to 15, wherein, the geometric representation of the limbal structure includes an annular or elliptical structure.
18. The method according to any one of claims 11 to 15, wherein, identifying the at least one area in the image data associated with the limbus of the eye includes at least one of the following: identifying image data indicative of eye features including the pupil, eyelids, sclera, iris, and eyelashes; and identifying at least one area associated with the limbus of the eye based on anatomical parameters.
19. The method according to claim 18, wherein, identifying the image data indicative of the eye features includes performing machine learning.
20. The method according to claim 18, wherein, The image data for identifying the eye features includes a segmented image for identifying pixels associated with a pupil region and identifying at least one limbus region based on the pupil region.
Citation Information
Patent Citations
Hybrid tracking system
US20010035938A1
Method and device for determining the eye position
US20130120712A1
Method for determining the azimuthal orientation of a patient eye and eye surgical apparatus therefor
US20140180162A1
System compensating for lateral target movements
WO1994018883A1