Eye-tracking device and method
The eye tracking method triangulates limbus structures in multiple images to overcome lens and cornea distortions, achieving precise eye movement measurement without user calibration, enhancing accuracy and flexibility in diverse applications.
Patent Information
- Application Number
- JP2025028507
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-01-29
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-01-27
AI Technical Summary
Existing eye tracking technologies struggle to accurately measure the six degrees of freedom of eye movement due to the distortion caused by the crystalline lens and cornea, and they require user calibration, which limits their precision and accuracy, especially in non-stationary environments.
An eye tracking method that triangulates the geometric representations of the limbus structures in multiple images to determine the three-dimensional position and line of sight of the user's eyes, using digital image preprocessing and triangulation techniques to identify and match stable regions corresponding to anatomical features, eliminating the need for user calibration.
Achieves accurate eye tracking with precision in the range of approximately 1/6 degree to 1/60 degree without requiring user calibration, providing a compact and flexible solution suitable for various devices.
Smart Images

Figure 2025098000000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an eye tracking device and method for measuring the position of an eye and / or the direction of a line of sight.
Background Art
[0002] The following documents are considered relevant as background art to the subject matter of the present disclosure. 1. US Patent Application No. 2013 / 120712 2. International Publication No. 9418883 3. US Patent Application No. 2014 / 180162 4. US Patent Application No. 2001 / 035938 5. Constable PA, Bach M, Frishman LJ, Jeffrey BG, Robson AG, International Society for Clinical Electrophysiology of Vision. ISCEV Standard for clinical electro-oculography (2017 update). Doc Ophthalmol. 2017;134(1):1-9 6. McCamy, Michael & Collins, Niamh & Otero-Millan, Jorge & Al-Kalbani, Mohammed & Macknik, Stephen & Coakley, Davis & Troncoso, Xoana & Boyle, Gerard & Narayanan, Vinodh & R Wolf, Thomas & Martinez-Conde, Susana. (2013). Simultaneous recordings of ocular microtremor and microsaccades with a piezoelectric sensor and a video-oculography system. Peer J. 1.e14. 10.7717 / peerj.14
[0003] The description of the above documents in this specification should not be inferred to mean that they are in any way relevant to the patentability of the subject matter of the present disclosure.
[0004] Conventionally, various eye-tracking devices such as remote eye trackers and head-mounted eye trackers, and various eye-tracking methods have been known. The first eye tracker was created in the late 19th century. However, they were difficult to create and gave discomfort to the subjects. Rings, contact lenses, and suction cups specially designed to be worn on the eyes were used to assist in measuring eye movements. The first photo-based eye tracker, which examined light reflected from various parts of the eye, was introduced at the beginning of the 20th century. They were much less invasive, and a new era of eye research began. For most of the 20th century, researchers were creating their own eye trackers, but they were costly and difficult to obtain. Commercial eye trackers did not appear until the 1970s. Since around the 1950s, numerous different techniques that are still in use today, such as contact lenses with mirrors (precisely suction cups), contact lenses with coils, electrooculogram recording (EOG) [5], piezoelectric sensors [6], etc., have been developed by researchers. However, by their nature, those methods can only perform angular eye movements and cannot measure lateral eye movements (mirrors, coils, and EOG can only measure angles). Also, those systems are stationary and require the user's head to be stabilized, so they are not suitable for studying eye movements in a more natural environment.
[0005] In the second half of the 20th century, much less invasive lighting and light sensor-based approaches became mainstream. However, the main limitation of this approach is that the camera becomes relatively slow due to the exposure and processing capabilities required to extract eye movement data. Position-sensitive detectors (PSDs) and quad-based approaches were developed in the mid-20th century. Dual Purkinje image eye tracker systems were developed in the 1970s and are still in use today. However, dual Purkinje image eye tracker systems are very expensive, stationary, heavy, (PSD-based) systems without cameras, and have a relatively low recording angle. Since the 1980s, with the advancement of camera sensor and computer technology, the eye tracking market has been dominated by so-called video oculography (VOG) systems. VOG systems typically capture an image of the user's eye and identify specific features of the eye based on the captured image. These systems are non-invasive and typically rely on irradiating the eye with infrared light so as not to cause any hindrance or discomfort to the user. Some of these systems rely on wearable, small and lightweight cameras, while more sophisticated systems, like most of the systems mentioned above, are stationary. Considering the fact that good lighting (but below safe levels) is required and the exposure requirements of camera sensors, the frame rate of these systems is limited.
[0006] In the eye tracker approach that combines pupil / corneal reflection (the first Purkinje), both stationary high-end and wearable low-end types, the eye is illuminated with a number of infrared diodes, the surface of the eye is imaged with one or more cameras (usually to increase the frame rate), and the pupil and the first Purkinje image of the diodes are segmented (as the darkest parts of the eye). The change in the position of the pupil relative to the first Purkinje image of the IR diodes indicates eye movement. To calculate the actual angle, user calibration is required. Usually, the user needs to focus on several targets moving along known paths to calibrate the system.
[0007] The accuracy and precision of this method are relatively low because the position of the pupil is distorted by the cornea, the measurement accuracy is further reduced by pupil dilation, and ambient reflections from the cornea disrupt the image processing algorithm. Also, as additional factors, different eye colors, long eyelashes, or contact lenses can further complicate the image processing system. For this reason, these systems typically have a lot of noise, the accuracy of angular motion is less than one degree, and no information about horizontal eye movement can be obtained. Furthermore, if the system moves relative to the head, recalibration is required. There are also some less common approaches, such as approaches that image the retina, approaches that utilize bright pupils, and approaches that use MRI devices to examine eye movement. These approaches have limitations and are not very common. In general, eye trackers and eye tracking methods are always required to have accuracy, precision, low latency, miniaturization, etc. Furthermore, as eye trackers are integrated into various devices such as computers, automobiles, and virtual reality glasses, there is also a need to provide a very compact and flexible eye tracking device.
[0008] In an eye tracking system, to measure relative eye rotation, it may be necessary to track, in particular, the limbus (the boundary between the sclera and the iris region), or the boundary between the pupil and the iris. Since the limbus is not affected by the refractive power of the cornea, accurate results can be obtained by tracking the limbus. Furthermore, the limbus forms a linear plane. Since the limbus defines the connection region between the iris muscles, there is a direct correlation between the three-dimensional parameters of the limbus and the three-dimensional parameters of the eye. Unfortunately, the limbus is not a sharp boundary but a transition zone between the cornea and the sclera. Since the limbus is not sufficient as a candidate for methods relying on so-called feature points or edge detection, previous limbus trackers have not been successful. Therefore, it becomes difficult to extract the signal from the noise due to the clearly defined but non-existent edges and shapes that appear to be deformed by rotation. As a result, methods that rely on edge detection or feature point detection of the limbus may not achieve the required accuracy and precision. Using a known "pinpoint" limbus tracker with photodetectors accurately arranged and aligned along the edge of the iris / sclera boundary may enable an appropriate tracking response, but providing and / or maintaining such alignment adds additional system components and complexity, especially considering the variation in the geometric shape of the eyes among different patients. Among the techniques for tracking the limbus, there are those that rely on a set of two quadrant detectors as position-sensitive photodetectors (PSDs). From the simple fact that the iris surrounded by the limbus is darker than the sclera, a system can be constructed that consists of one or more quads that examine the boundary line of the iris / sclera region, i.e., the limbus. When the eye moves, the image of the iris also moves on the PSD, so the eye movement that causes this can be estimated. Despite the fact that a laser can be stabilized on the cornea, it cannot directly indicate where the eye is looking, and calibration by the user is required to obtain this data. Also, many factors that can be controlled during eye surgery, such as eyelid movement, eyelashes, and slight changes in lighting, affect the accuracy of this approach and are not suitable for eye research. The limbus tracker based on the PSD measures the centroid of the position of the light spot. This is because there is a contrast between the white sclera and the dark iris.When the eye moves relative to the system, the light spot on the PSD also moves accordingly. In such a system, by only controlling that the iris image is at the same location on the PSD, it is not possible to measure or estimate the actual eye position.
Summary of the Invention
[0009] Eye tracking refers to the process of tracking the movement of the eyeball and determining where the user is looking. The various techniques described above can roughly estimate the scale of eye movement and indicate that movement has occurred. However, they do not measure the six degrees of freedom of eye movement. Furthermore, those systems do not track the visual features of the eye. One of the problems in tracking the visual features of the eye is that the most characteristic part of the eye is the iris. Both the pupil and the iris are strongly distorted by the crystalline lens. Since the shape of this lens varies from person to person, it is difficult to accurately determine where those features are located. Astigmatism (asymmetry of the cornea) further complicates this problem. The limbus (the boundary between the white sclera and the dark iris) is not distorted by the cornea, but appears very uniform and slightly blurred, and does not have prominent features such as those required by classical image processing algorithms. Therefore, there is a need in the art to provide an eye tracking technique based on limbus tracking. The inventors have found that limbus eye tracking can be provided by triangulating identified stable regions that directly correspond to existing, i.e., directly visible, anatomical or virtual features (i.e., calculated features).
[0010] According to a broad aspect of the present invention, an eye tracking method is provided, the method comprising the steps of receiving image data showing at least two images of a user's eyes, identifying in each image a region associated with the corneal limbus, specifying a geometric representation of the limbus structure, and determining the three-dimensional position and line of sight (i.e., the full six degrees of freedom of the eye) of the user's eyes by triangulating the geometric representations of the limbus structures of at least two images. Since the limbus maintains a near-circularity with a certain relationship to the lens of the cornea, the three-dimensional position and line of sight of the user's eyes can be obtained by triangulating the geometric representations (i.e., two-dimensional corneal limbus parameters) of the limbus structures of at least two images. The shapes of at least two images are identified and matched. The three-dimensional distance to the matching region can be determined using triangulation. The corresponding regions can be matched and optionally three-dimensional corneal limbus parameters can be obtained by applying the epipolar geometry techniques commonly used in this field. Based on the three-dimensional limbus parameters, the three-dimensional position and line of sight of the user's eyes are deduced.
[0011] In this regard, it should be generally understood that stereo vision algorithms rely on the matching of corresponding points or edges between two images. There are two basic approaches to this. One is intensity-based, where the intensity value information of each line is collated at each pixel and the difference in intensity is minimized. The other is feature-based, where information is extracted using groups of pixel values and their spatial distribution. Commonly used features are simple features, such as edge and corner detectors (Canny, Harris, LoG filters). Generally, eye tracking systems that rely on the above-mentioned features have several drawbacks, such as different noises generated at each of a plurality of photodetectors as a result of parasitic reflections on the cornea, and distortion of iris and pupil features as a result of the curvature of the corneal lens. Therefore, in eye tracking based on feature-based methods, results with sufficient accuracy cannot be obtained. Also, methods based on the matching of corneal reflections and eye features require calibration by the user. Regarding the tracking of the limbus, as described above, since the limbus does not have clearly defined features, classical approaches cannot provide three-dimensional corneal limbus parameters. It should be noted that the technology of the present invention provides accurate eye tracking with an accuracy and precision in the range of approximately 1 / 6 degree to 1 / 60 degree without requiring user calibration of the eye tracking device.
[0012] In some embodiments, the method includes the step of capturing at least two images of the user's eye at different angles.
[0013] In some embodiments, identifying a geometric representation of the limbus structure includes digital image preprocessing. Digital image preprocessing includes calculating an image intensity gradient map of the limbus based on the pupil direction, identifying at least one region of the limbus structure in each image where the local direction of the gradient is substantially uniform, processing data representing the limbus structure by weighting the pixels of such regions, and generating a geometric representation of the limbus structure based on corresponding pixels related to the corneal limbus.
[0014] In some embodiments, the method includes determining three-dimensional pupil parameters including at least one of three coordinates of the center of the pupil, the direction of the normal to the pupil plane, the position of the pupil plane, and the size of the pupil. Determining the three-dimensional position and gaze direction of the user's eyes can be performed based on the three-dimensional pupil parameters.
[0015] In some embodiments, the method further includes determining the radius of the pupil and / or the torsional rotation of the pupil.
[0016] In some embodiments, the method further includes identifying image data indicating an initial pupil region in each image and generating data indicating the initial pupil region using an iterative pixel filtration process.
[0017] In some embodiments, identifying image data indicating an initial pupil region in each image includes at least one of identifying image data indicating eye features such as the pupil, eyelids, sclera, iris, and eyelashes, and identifying the initial pupil region based on anatomical parameters. Identifying image data indicating eye features can be performed using machine learning.
[0018] In some embodiments, identifying image data indicating eye features includes segmenting each image to identify pixels associated with the pupil region.
[0019] In some embodiments, the method further includes performing triangulation between at least two geometric representations of the pupil to determine three-dimensional pupil parameters, estimating the position of the initial pupil region based on the three-dimensional pupil parameters, and generating data indicating the position of the initial pupil region of the pupil in each image. The three-dimensional pupil parameters include at least one of the direction of the normal to the pupil plane, the coordinates of the center of the pupil, and the diameter of the pupil in each image.
[0020] In some embodiments, the method further includes processing the triangulated three-dimensional limbus, re-projecting the triangulated limbus region onto the image plane, and refining the limbus region in each image.
[0021] According to another broad aspect of the present invention, an eye tracking device including a processing unit is provided, the processing unit being configured to receive at least two images indicating a user's eyes, identify a region associated with the corneal limbus in each image, determine a geometric representation of the limbus structure, and determine the three-dimensional position and gaze direction of the user's eyes by triangulating the geometric representations of the limbus structures of at least two images, and being operable. The processing unit is configured to determine the three-dimensional position and gaze direction of the user's eyes based on the three-dimensional limbus parameters and is operable. The eye tracking device of the present invention may be a head-mounted eye tracking device such as an eye tracking glasses, or an eye tracker integrated into a helmet, and may also be integrated into a head-mounted display device, virtual reality glasses, augmented reality glasses or other head-mounted devices. Further, the eye tracking device may be a remote eye tracker optionally integrated or coupled with other devices such as a computer, a display, a monitor, etc.
[0022] In some embodiments, the eye tracking device further includes at least two imagers, each imager being configured to capture at least one image of the user's eyes at different angles.
[0023] The imager may be a capture unit and can include one or more cameras (i.e., digital cameras or video cameras), light sensors, and image sensors such as CCD sensors or CMOS sensors. One imager can be used to move at high speed along a known trajectory at a known speed. Alternatively, it is also possible to use an optical image with certain mirrors. In some embodiments, the present invention uses a stereo imager and uses two or more imagers with a precisely known relative offset. The two or more imagers can capture images simultaneously or with a known time delay (e.g., in an interlaced mode).
[0024] Generally, the processing unit may be a processor, a controller, a microcontroller, or any kind of integrated circuit. The eye tracking device can be associated with an optical system including any optical elements such as one or more lenses, prisms, beam splitters, mirrors, reflectors, optical guides, collimators, etc. Also, the eye tracking device is associated with an illumination source configured to generate as dispersed illumination as possible, which preferably does not create or minimizes a pattern on the surface of the eye. The eye tracking device includes a processing unit that communicates with each of at least two imagers for receiving at least two images.
[0025] The novel technology of the present invention uses two or more captured images of the eyes, and in each of the captured images, three-dimensional corneal limbus parameters are detected by an iterative pixel filtering process, whereby the orientation of the eyeball is identified by triangulation of the two-dimensional parameters of the corneal limbus obtained from the images. This novel technology identifies the eye orientation based on the identification of the limbus of the user's eyes in the image data. Two or more imagers can be used to capture two or more images of each of the user's eyes. Each imager is configured to capture the user's eyes from different angles. Each image images the user's iris and adjacent and peripheral regions such as the pupil, sclera, and eyelids. For triangulation of the limbus region, multiple images of the eye are captured from two or more different angles, preferably at temporally close timings (e.g., simultaneously).
[0026] In some embodiments, the processing unit includes an area detector that receives each of at least two images indicating the user's eyes, and in each image, using an iterative pixel filtering process, an initial contour area, preferably an area / section without initial distortion (i.e., not distorted by the eye optical system itself such as the cornea), or an area with known distortion corresponding to the same physical area of the eye, or image data indicating an area corresponding to the same physical area of the eye is identified, and data indicating the initial contour area is generated. The inventors have found that the stable areas thus identified can be used for triangulation calculations. In the present invention, such stable areas corresponding to the same physical area of the eye can be generated and identified on the stereo image. The area detector is configured to identify image data indicating the initial contour area by at least one of identifying image data indicating features of the eye such as the pupil, eyelid, sclera, iris, eyelashes, etc., and identifying the initial contour area based on anatomical parameters. These areas are identified by projecting a geometric representation of an eye feature (e.g., the pupil) onto each image, and then these areas are collated between images taken at different angles. The larger the size of such an area on the image, the more accurately the position of the corresponding physical area is triangulated, and as a result, the position of the eye itself is accurately determined. The area detector uses an iterative pixel filtering process that enables the execution of triangulation to generate a function that defines a two-dimensional virtual object on the eyeball. By this triangulation, the direction of the contour corresponding to the line of sight is defined. The contour parameters are defined as a mathematical formula where the variables represent the six degrees of freedom of a rigid body in three-dimensional space (x, y, z, α, β, δ). Specifically, x, y, z define the change in position due to front / back (surge), up / down (heave), left / right (sway) movement along three perpendicular axes, combined with the change in orientation due to rotation about three perpendicular axes, yaw (vertical axis), pitch (horizontal axis), roll (longitudinal axis).
[0027] In this regard, it should be understood that since the dimensions of the limbus are large and are not substantially distorted by the cornea, by specifying three-dimensional limbus parameters, it is possible to find a region without such distortion or a region having a known distortion corresponding to the same physical region of the eye. However, as described above, the limbus itself has a slightly blurred boundary and is not clearly defined, and it looks different for each image, and does not have distinguishable features that can function as such matching anchor points. Therefore, it is difficult to match the limbus region itself on two or more different images. The inventors have found that instead of matching the features of the limbus region itself, digital image preprocessing can be performed to calculate several virtual regions based on the three-dimensional parameters of the limbus of each image that well corresponds to the same physical region of the eye. Digital image preprocessing includes the use of complex algorithms such as classification, feature extraction, multiscale signal analysis, pattern recognition, or projective mathematical transformation. The virtual region needs to correspond to the same physical region of the eye regardless of the angle at which the image is taken. For this reason, the region detector is configured and operable to identify the image data indicating the limbus region in each image and generate data indicating the limbus region. This can be performed by estimating the limbus region using the anatomical relationship between the identified virtual region corresponding to a specific eye feature and the limbus. For example, by triangulating the pupil region (a region having a defined circular contrast) between two images, the position and center of the pupil can be calculated, and then the limbus region can be defined as a circular region having a certain radius range (for example, about 5.5 mm to 6.5 mm) centered on the center of the pupil. This radius range may be based on statistical data defining the Gaussian distribution of the distance between the limbus and the center of the pupil.
[0028] In some embodiments, a pattern recognition tool is used to identify the pupil pixels and eyelid pixels within each image. The image pixels of the pupil can be identified by a specific contrast within the image. The identified image region corresponding to the pupil is discarded, and the eyelid image region can be used to determine calibration data (such as the distance / angle of each imager, the center of the pupil / limbus, the overall direction, etc.). More specifically, since the data image related to the possible intersection between the eyelid and the presumed limbus region is discarded, the number of potential limbus pixels can be further reduced. Next, based on the calibration data and the three-dimensional model of the eye, two-dimensional corneal limbus parameters are identified. Thereafter, in each of two or more captured images, an iterative filtering process is used to discard the pixels located at the boundary of the limbus. The geometric representation of the limbus structure includes a ring-shaped or elliptical structure. If it is assumed that the limbus is circular, using the two-dimensional parameters of the limbus obtained from each image, for each image, a complete limbus ring / ellipse structure function (reconstructed by discarding the eyelid region hiding the limbus region) is defined, and the orientation of the eye is specified by the triangulation of the defined limbus ring / ellipse structure obtained from each image. By filtering the image data corresponding to specific eye features, specific pixels can be discarded and the resolution can be increased. If the orientation of the eye cannot be specified, this process is repeated while increasing the region of the limbus and potentially retaining the pixels related to the limbus.
[0029] In some embodiments, the method first performs a mathematical transformation, such as separating the limbus region and finding a geometric representation of the limbus structure (e.g., by projecting a set of ellipse functions passing through the center of the limbus ring) in each of the images, and then triangulates the corresponding circles in six degrees of freedom (6DOF) that are accurately aligned with the physical eye.
[0030] In some embodiments, the processing unit comprises a limbus detector that is configured and operative to receive data indicative of a limbus region and perform limbus recognition processing on each image and / or to identify a geometric representation of the limbus structure by digital image preprocessing, including performing a mathematical transformation that involves use of an intensity gradient map. More specifically, the digital image preprocessing can include calculating an image intensity gradient map of the limbus region, identifying at least one region of the limbus structure where the local direction of the gradient is substantially uniform, processing data indicative of the limbus structure by weighting pixels of such a region, and identifying a geometric representation of the limbus structure based on matching pixels associated with the corneal limbus.
[0031] More specifically, after estimating the limbus region, a ring-shaped region is projected onto the surface of the eye in the estimated limbus region. Thereafter, most of the pixels of the image, such as most of the pupil, iris, sclera, and eyelids, are discarded, leaving only a ring around the location of the potential limbus based on the estimated anatomical characteristics of a normal eye and optionally the calculated three-dimensional pupil position. However, due to distortion of the cornea in the pupil image, further estimation errors can occur. This distortion leads to an increase in the initial estimate of the limbus region.
[0032] In some embodiments, the region detector is configured to identify image data indicative of eye features using machine learning. The iterative pixel filtering process is used to train one or more neural networks for identifying the rim region in each acquired image. The region detector may be configured and operable to identify image data indicative of eye features by segmenting each image to identify pixels associated with the pupil region. To this end, one or more neural networks can be used to directly segment the approximate rim region. More specifically, the pupil can be tracked to estimate the rim region. Additionally or alternatively, a neural network can be used to directly estimate the positions of other eye features such as eyelids or sclera without tracking the pupil. The neural network can be applied to the image itself or to a mathematically transformed image, i.e., a gradient map. The neural network can be trained with conventional machine learning algorithms, for example, based on the results of other eye tracking techniques where characteristic properties of the eye are predefined.
[0033] In some embodiments, the processing unit is configured and operable to determine the size of the rim.
[0034] In some embodiments, to improve the estimation of the eye position, the pixels of the identified eyelid region are triangulated.
[0035] In some embodiments, the eye tracking device further comprises a rim triangulator that is configured and operable to receive data indicative of a geometric representation of the rim structure for each image and perform triangulation of the geometric representations of at least two images to determine three-dimensional rim parameters. The three-dimensional rim parameters include at least one of the three coordinates of the center of the rim, the direction of the normal to the rim plane, the position of the rim plane, and the size of the rim. The rim triangulator may be configured and operable to determine the radius of the rim and / or the torsional rotation of the rim.
[0036] In some embodiments, the limbus triangulator is configured to perform triangulation between at least two geometric representations of the pupil to identify three-dimensional pupil parameters, the three-dimensional pupil parameters including at least one of the direction of the normal to the pupil plane, the coordinates of the center of the pupil, and the diameter of the pupil, the limbus triangulator estimating the position of the initial corneal limbus region based on the three-dimensional pupil parameters and generating data indicating the position of the initial region of the limbus in each image.
[0037] In some embodiments, the limbus triangulator is configured to process the triangulated three-dimensional limbus and re-project the triangulated limbus region onto the image plane to refine the limbus region in each image. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] To better understand the subject matter disclosed herein and to illustrate how it may be actually implemented, embodiments will be described by way of non-limiting example only, with reference to the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
[0039] Referring to FIG. 1, a block diagram showing the main functional parts of the eye tracking device of the present invention is illustrated. The eye tracking device 100 is configured to receive at least two images showing a user's eyes, identify regions associated with the corneal limbus in each image, specify a geometric representation of the limbus structure, and measure the three-dimensional position and gaze direction of the user's eyes by triangulation of the geometric representations of the limbus structures of at least two images, and includes a processing unit 106 that is operable. In this regard, as described above, it should be noted that the present invention triangulates the shape (geometric representation), but does not triangulate individual corresponding points or edges.
[0040] The processing unit 106 is generally configured as a computing / electronic utility that particularly includes utilities such as data input / output modules / utilities 106A and 106B, a memory 106D (i.e., a non-volatile computer-readable medium), and an analyzer / data processing utility 106C. For this reason, the utilities of the processing unit 106 can be executed by appropriate circuitry and / or by software and / or hardware components including computer-readable code configured to execute the operations of method 200 shown in FIG. 2 described below.
[0041] The features of the present invention can include general-purpose or special-purpose computer systems, including various computer hardware components that will be described in more detail below. Also, features within the scope of the present invention include computer-readable media for executing, or having, computer-executable instructions, computer-readable instructions, or data structures stored therein. Such computer-readable media may be any available media accessible by a general-purpose or special-purpose computer system. By way of non-limiting example, such computer-readable media can include physical storage media such as RAM, ROM, EPROM, flash disk, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or other media that can be used to hold or store desired program code means in the form of computer-executable instructions, computer-readable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer system. The computer-readable media can include computer programs or computer applications that are downloadable to a computer system via a network such as a wide area network (WAN), for example, the Internet.
[0042] In this specification and the following claims, a "processing unit" is defined as one or more software modules, one or more hardware modules, or a combination thereof, which cooperate to perform operations on electronic data. For example, the definition of a processing utility includes software modules such as the operating system of a personal computer in addition to the hardware components of the personal computer. The physical layout of the modules is irrelevant. A computer system can include one or more computers coupled via a computer network. Similarly, a computer system can include a single physical device in which internal modules (such as memory and a processor) cooperate to perform operations on electronic data. Any computer system may be mobile, but the terms "mobile computer system" or "mobile computer device" as used herein specifically include, for example, laptop computers, netbook computers, mobile phones, smartphones, wireless phones, personal digital assistants, portable computers with touch sensor screens, etc. The processing unit 106 incorporates a processor that executes a computer program or has a processor attached thereto. A computer program product may be incorporated in one or more computer-readable media and have computer-readable program code incorporated therein. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer program code for performing operations related to aspects of the present invention may be described in any combination of one or more programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.In the latter case, the remote computer may be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). These computer program instructions can be provided to the processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, and the instructions executed via the processor of the computer or other programmable data processing apparatus can result in a machine that creates means for performing the functions / operations specified in one or more blocks of the flowchart and / or block diagram. The specified functions of the processor can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0043] The eye tracking device 100 can include at least two imagers 110, each imager being configured to capture at least one image of the user's eye at a different angle. Each imager can focus on the user's iris. In this regard, it should be noted that the ring tracking of the present invention is performed independently for each eye.
[0044] In some embodiments, the processing unit 106 includes a limbus detector 102 that is configured and operative to receive data indicative of a limbus region and to identify a geometric representation (e.g., a ring-shaped or elliptical structure) of the limbus structure by digital image preprocessing. In a specific non-limiting example, the limbus detector 102 is configured and operative to perform limbus recognition processing on each image and / or to perform a mathematical transformation of the image including the use of an intensity gradient map, and then to perform limbus region recognition processing on the transformed image. In other words, the limbus detector 102 can search for a limbus region on the transformed image after transforming the image into an intensity gradient map. Also, instead of or in addition to the gradient map, an entropy map can be used, or processing can be performed directly on the image. When an intensity gradient map is used, the limbus detector 102 calculates an image intensity gradient map of the limbus region, identifies at least one region of the limbus structure where the local direction of the gradient is substantially uniform, weights the pixels of such region, and can process data indicative of the limbus structure by identifying a geometric representation of the limbus structure based on corresponding pixels related to the corneal limbus. In addition to collinearity, knowledge of the anatomical form of the limbus (i.e., that the vectors of collinearity extend radially from the center of the limbus and that the limbus is continuous throughout) can be used. More specifically, the region of the limbus structure is identified by identifying local uniformity, i.e., that the direction of the gradient at each point is collinear only with adjacent points.
[0045] Each module of the processing unit 106, such as the wheel detector 102, is not limited to a specific number of modules and may be configured to process multiple images independently, simultaneously, or in other ways. The wheel triangulator 108 is configured and operable to perform triangulation of the geometric representation of the wheel structure in at least two images to obtain three-dimensional wheel parameters. Further, the wheel triangulator 108 may be configured and operable to obtain the radius of the wheel and / or the torsional rotation of the wheel. The three-dimensional wheel parameters define at least five degrees of freedom of the position and orientation of the wheel. The three-dimensional wheel parameters include any one of the position of the plane of the wheel and / or the direction of the normal to the plane of the wheel and / or the size of the wheel and / or the center of the wheel or a combination thereof.
[0046] In some embodiments, the processing unit 106 includes an area detector 104, which receives each of at least two images indicating the eyes of each user and is configured and operable to identify image data indicating an initial limbal region within each image by using an iterative pixel filtering process and generate data indicating the initial limbal region. This can be performed by identifying image data indicating features of the eye such as the pupil, eyelid, sclera, iris, and eyelashes, and / or by identifying the initial limbal region based on anatomical parameters. More specifically, the identification of the initial limbal region can include pupil segmentation and pupil triangulation, and / or estimation of the limbal region based on anatomical structures, and / or estimation of the eyelid region, and / or estimation of the limbal region excluding the eyelid region. This can be performed by the limbal triangulator 108 or by any other processing module. However, the present invention is not limited by a specific module configuration of the processing unit 106. The processing unit 106 can include a pupil detector, a pupil triangulator, and other modules. For example, the identification of image data indicating features of the eye can include pupil segmentation. Thus, the area detector 104 can be a pupil detector aimed at providing an initial limbal ring. It is assumed that both of these regions, the pupil and the limbus, are concentric and on the same plane. Pupil segmentation can be performed by segmenting each image to identify pixels related to the pupil region, performing triangulation between at least two images to determine three-dimensional pupil parameters, estimating the position of the initial corneal limbal region based on the three-dimensional pupil parameters, and generating data indicating the initial region position of the limbus in each image. The three-dimensional pupil parameters include the direction of the normal to the pupil plane, and / or the coordinates of the center of the pupil, and / or the diameter of the pupil. The three-dimensional pupil parameters are used to estimate the three-dimensional limbal parameters, thereby defining an initial limbal region that can be further refined.In addition, the identification of the image data indicating the eye features can include the identification of the eyelids and the filtering of this image data because the image data related to the eyelids blocks a part of the limbus ring. The present invention is not limited to a specific method for identifying the image data indicating the initial limbus. To identify the image data indicating the initial limbus, an approach based on a neural network can be used. For example, the limbus region estimation can be performed using a neural network, but the eyelid region can also be estimated using a classical algorithm with a pupil detector. Alternatively, the limbus region estimation and the eyelid region estimation can be performed with an approach based on a neural network. Alternatively, the neural network can be used only for pupil segmentation.
[0047] In some embodiments, the region detector 104 is configured to identify the image data indicating the eye features using machine learning. Machine learning can self-classify / learn the characteristics of the input data related to the eye features. The region detector 104 can use a data recognition model based on a neural network to predict the three-dimensional corneal limbus parameters. The network can be trained based on the segmentation results of classical approaches or using an existing system. Additionally or alternatively, the three-dimensional corneal limbus parameters obtained by the eye tracking device of the present invention can be used for training the neural network. A deep learning network (DLN), such as an artificial neural network (ANN) executed by the region detector 104, can generate a representation of the limbus based on a series of images. For example, the representation generated by the DLN can include probabilities regarding the placement of the limbus. This representation is used to generate a model of the limbus for adapting the recognition classifier to a person.
[0048] After the wheel detector 102 obtains the three-dimensional wheel parameters, the region detector 104 can receive the three-dimensional wheel parameters and generate more accurate data indicating the wheel region. For example, the wheel triangulator 108 is configured to process the triangulated three-dimensional wheel, re-project the triangulated wheel region onto the image plane, and refine the wheel region in each image.
[0049] Referring to FIG. 2, a flowchart diagram illustrating the main steps of the eye tracking method of the present invention is exemplified. The eye tracking method 200 includes, at 204, receiving image data showing at least two images of the user's eyes; at 206, identifying regions associated with the corneal limbus in each image; at 208, specifying a geometric representation of the limbus structure; and at 210, performing triangulation of the geometric representations of the limbus structures of at least two images, thereby, at 212, determining the three-dimensional position and gaze direction of the user's eyes.
[0050] In some embodiments, step 210 includes performing triangulation of the geometric representations of the limbus structures in at least two images to obtain the three-dimensional limbus parameters as described above. Obtaining the three-dimensional limbus parameters can include obtaining the radius of the limbus and / or the torsional rotation of the limbus.
[0051] In some embodiments, after performing triangulation of the geometric representations of the limbus structures in at least two images at 210, the method 200 can include specifying the geometric representation of the limbus structure by processing the three-dimensional limbus parameters to generate more accurate data indicating the limbus region.
[0052] In some embodiments, the method 200 can include, at 202, capturing at least two images of each user's eyes at different angles.
[0053] In some embodiments, identifying the geometric representation of the limbal structure at 208 may include performing limbal recognition processing on each image or performing a mathematical transformation of the image using an intensity gradient map at 218, and then performing limbal region recognition processing on the transformed image, including digital image preprocessing. The steps of the digital image preprocessing at 218 may include calculating an image intensity gradient map of the limbal region, identifying at least one region of the limbal structure where the local direction of the gradient is substantially uniform, weighting the pixels of such a region, and processing the data indicating the limbal structure by identifying the geometric representation of the limbal structure based on the matching pixels associated with the corneal limbus. More specifically, the digital image preprocessing may include identifying the region of the limbal structure by identifying local uniformity, i.e., that the direction of the gradient at each point is collinear only with its adjacent points. As described above, the entropy map can be used as (instead of or in addition to) the intensity gradient map, or the processing can be performed directly on the image.
[0054] In some embodiments, at 206, identifying the region associated with the corneal limbus in each image may include identifying image data indicating features of the eye, such as the pupil, eyelid, sclera, iris, and eyelashes, and / or identifying an initial limbal region based on anatomical parameters using iterative pixel filtering processing, and generating data indicating the initial limbal region. For example, an estimation of the eyelid region can be performed in parallel with the pupil region and the limbal region. The present invention is not limited to the order of estimation of various regions. The various eye feature regions may be identified based on any geometric shape model (e.g., ellipse or circle).
[0055] In some embodiments, identifying the initial limbus region can include pupil segmentation and pupil triangulation, and / or estimating the limbus region based on anatomical structures, and / or estimating the eyelid region, and / or estimating the limbus region excluding the eyelid region. The method can include segmenting each image to identify pixels related to the pupil region, performing triangulation between at least two images to obtain three-dimensional pupil parameters, estimating the position of the initial corneal limbus region based on the three-dimensional pupil parameters, and generating data indicating the position of the initial region of the limbus in each image. Also, identifying image data indicating eye features can include identifying the eyelids because the image data related to the eyelids obscures a part of the limbus ring, and filtering this image data. As described above, the present invention is not limited to a specific method for identifying image data indicating the initial limbus. A neural network-based approach can be used to identify image data indicating the initial limbus. Also, limbus region estimation and eyelid region estimation can be performed with a neural network-based approach. Also, the neural network can be used only for pupil segmentation.
[0056] In some embodiments, machine learning can be used to perform identification of image data indicating eye features. Obtaining (i.e., predicting) three-dimensional corneal limbus parameters can include using a neural network-based data recognition model. This method can include training the network based on the segmentation results of classical approaches or using an existing system for training. Generating a representation of the limbus based on a series of images can be achieved using a deep learning network (DLN) such as an artificial neural network (ANN). For example, generating a representation of the limbus can include calculating probabilities regarding the arrangement of the limbus and / or generating a model of the limbus, and adapting a recognition classifier to a person.
[0057] Referring to FIG. 3, two stereo images of the user's eyes taken at different angles are shown.
[0058] Referring to FIGS. 4A to 4D, examples of data processing of the two stereo images of FIG. 3 for pupil region detection according to some embodiments of the present invention are shown. FIG. 4A shows a segmentation stage of segmenting each image to identify pixels related to the pupil region. As shown in FIG. 4B, pixels not related to the pupil region are discarded and the resolution is increased. FIG. 4C shows the segmentation of the pupil boundary based on the intensity map of each image. FIG. 4D is a refinement of the pupil boundary based on the projection of the geometric representation (e.g., ring shape) of the pupil in each image. Thereafter, the regions in FIG. 4D are triangulated between the images to obtain the pupil diameter, 3D position, and orientation.
[0059] Referring to FIG. 5, an example of data identification of eye features of the two stereo images of FIG. 3 according to some embodiments of the present invention is shown. The marked regions represent regions where the local direction of the image intensity gradient is substantially uniform.
[0060] Referring to FIG. 6, an example of a rim recognition process based on the projection of an ellipse function onto each of the two stereo images of FIG. 3 according to some embodiments of the present invention is shown.
[0061] Referring to FIG. 7, a block diagram illustrating possible functional parts of the eye tracking device of the present invention is exemplified. In this specific non-limiting example, the eye tracking device 300 includes two imagers, referred to herein as imager 1 and imager 2, which are configured to capture at least one image of each user's eye at different angles. Each imager focuses on the iris of the user. The eye tracking device 300 includes a processing unit 310 including a pupil detector 302 and a limbus detector 304. The pupil detector 302 is a specific example of the region detector defined with respect to FIG. 1 described above. The pupil detector 302 includes a plurality of modules and is configured and operable to receive each of at least two images showing each user's eye, identify image data indicating the limbus region within each image, and generate data indicating the limbus region. This may be performed using eye feature detection such as pupil region detection or may be performed using a neural network. Each image is segmented by a segmentation module to identify pixels related to the pupil region. Pixels not related to the pupil region are discarded in each image, and the resolution is increased. Thereafter, each ellipse module Ell calculates an elliptical curve corresponding to the pupil region on each image and finds a matching elliptical curve. For each image, a three-dimensional equation indicating the three-dimensional pupil region is generated. Subsequently, when the distance between the two imagers that capture the two stereo images is known, a first triangulation that matches the three-dimensional equations of the two stereo images is performed, and the direction of the pupil is obtained. Subsequently, assuming that the average size of the limbus is used and that the pupil and the limbus are concentric and on the same plane, the algorithm results in an estimation of three-dimensional limbus parameters. This three-dimensional limbus region is projected onto each image, referred to herein as a mask equation, as shown in FIG. 3. For example, the limbus region may be surrounded by two elliptical curves. The mask equation is used by the collinearity map module to give a first estimate of the projection of the limbus on each image. The mask equation is also used in the weighting module.
[0062] To determine the three-dimensional limbus parameters, a collinearity module, referred to herein as a collinearity map, performs mathematical transformations. In a specific non-limiting example, such mathematical transformations include calculating an image intensity gradient map of the limbus based on the pupil direction, identifying at least one region of the limbus structure where the direction of the gradient is substantially uniform, processing data indicative of the limbus structure by weighting the pixels of such a region, and determining three-dimensional limbus parameters based on matching pixels associated with the corneal limbus. Alternatively, such mathematical transformations can include using the entropy of the image. In a specific non-limiting example, the collinearity map module generates a gradient collinearity map from the original image masked by an estimated limbus region (calculated from the pupil 3D parameters and the average limbus size). The limbus detector 304 receives a three-dimensional limbus region and a three-dimensional equation indicating the direction of the limbus. Thereafter, a ring-shaped region is projected onto each eye image, and an estimated limbus region is generated. Most pixels of the image, such as most of the pupil, iris, sclera, and eyelids, are discarded using a mask equation, leaving only a ring around the potential limbus location based on the estimated anatomical characteristics of a normal eye.
[0063] The weighting module assigns to each pixel a weight of the probability that the pixel belongs to the limbus, based on the collinearity map module. The weighting module is then used to narrow down the potential limbus region and provides an input to the Ell module to find an estimate of the matching ellipse of the limbus.
[0064] To perform the second triangulation, several techniques such as pyramid triangulation or quadratic estimation (linear or non-linear) can be executed. The processing unit 310 is configured and operable to identify the three-dimensional position and the gaze direction of the user's eyes by triangulating the geometric representation of the limbus structures of at least two images using linear or non-linear quadratic estimation. In this specific non-limiting example, pyramid triangulation can be performed as follows. That is, each ellipse module Ell generates a mathematical function / equation of the geometric representation of the limbus that defines the anatomical relationship between the identified virtual regions corresponding to specific eye features (e.g., an ellipse having a center corresponding to the center of the pupil), and the limbus is calculated for each image. The ellipse structures of each image are compared, and the intersection of the cones defined by the ellipse structures defines an ellipse in space, the projection of which matches the ellipse structures of the two images. Thereby, in this specification, three-dimensional limbus parameters called limbusR, Cx, Cy, Cz, azimuth, and pitch are defined. In this regard, it should be noted that in the present invention, the three-dimensional limbus parameters are obtained by triangulating the ellipse, not the pixel. As described above, pixel triangulation of the limbus is virtually impossible because the correspondence between pixels in different images is unknown (since most of the pixels along the epipolar line of the second image are very similar). In classical computer vision, the matching between two pixels along the epipolar line of different images can be characterized using special features of adjacent pixel regions such as different colors, corners, intersections of lines, or texture patches. Such features are lacking in the limbus region due to the approximate boundary of the limbus defined as a generally large region, and thus classical natural feature tracking algorithms cannot be used. Thereafter, a mask is used to discard the regions related to the eye features. Subsequently, the same procedure as above is repeated using an iterative pixel filtering process to more accurately define the limbus region and reduce the size of the limbus region.
[0065] Alternatively or additionally, the iterative pixel filtering process can be performed using linear quadratic estimation such as a Kalman filter, or non-linear quadratic estimation such as an extended Kalman filter and an unscented Kalman filter. The Kalman filter generates an estimate of the pupil equation with five degrees of freedom (such as an elliptical structure) as the average of the predicted state of the eye and the new measurement using a weighted average. The weights are calculated from the covariance, which is a measure of the estimated uncertainty of the prediction of the state of the eye. As a result of the weighted average, the estimated value of the new state is located between the predicted state and the measured state and has a better estimated uncertainty than either state when obtained alone. This process is repeated for each time step, and the new estimated value and its covariance inform the prediction used in the next iteration.
[0066] At the end of the process, the processing unit 310 determines the three-dimensional position and the gaze direction of the user's eyes based on the three-dimensional pupil parameters.
[0067] Referring to FIG. 8, the trajectory of the two-dimensional eye tracking position obtained using the teachings of the present invention is illustrated, with different colors distinguishing different parts of the trajectory, i.e., the color changes over time. The axes Xc, Yc are measured in millimeters.
[0068] Referring to FIG. 9, two angular degrees of freedom of the eye tracking direction trajectory of FIG. 8 (i.e., angular measurement values measured in degrees) are illustrated with respect to azimuth and pitch as a function of time (measured in milliseconds). FIGS. 8, 9, and 10 all show data from the same recording.
[0069] Referring to FIGS. 10A and 10B, five degrees of freedom of the eye tracking trajectory of FIG. 8 are illustrated. More specifically, FIG. 10A shows, as a function of time, two angular degrees of freedom of the eye tracking direction trajectory (i.e., angular measurement values measured in degrees) with respect to azimuth and pitch. FIG. 10A is an enlarged view of a part of the graph of FIG. 9. FIG. 10B shows the three-dimensional degrees of freedom of the eye tracking position trajectory in the Cartesian coordinate space as a function of time (measured in milliseconds).
Claims
1. An eye tracking device comprising a processing unit, An eye tracking device characterized in that the processing unit is configured and operable to receive at least two images showing a user's eyes, identify an area associated with a corneal limbus in each image, determine a geometric representation of the limbus structure, and determine a three-dimensional position and gaze direction of the user's eyes by triangulation of the geometric representations of the limbus structures in the at least two images.
2. 2. The eye tracking device of claim 1, 11. An eye tracking device, wherein the geometric representation of the limbus structure comprises a ring-like or elliptical structure.
3. 3. The eye tracking device according to claim 1, The eye tracking device further comprising at least two imagers, each imager configured to capture at least one image of a user's eye at a different angle.
4. 4. The eye tracking device according to claim 1, 11. An eye tracking device, comprising: a limbus detector configured and operable to receive data indicative of a limbus region and identify a geometric representation of a limbus structure by digital image pre-processing.
5. 5. The eye tracking device of claim 4, An eye tracking device characterized in that, for each image, the limbus detector is configured and operable for digital image pre-processing, the digital image pre-processing including: calculating an image intensity gradient map of the limbus region; identifying at least one region of the limbus structure in which the local direction of the gradient is substantially uniform; processing data indicative of the limbus structure by weighting pixels in such region; and identifying a geometric representation of the limbus structure based on matching pixels associated with the corneal limbus.
6. 6. The eye tracking device according to claim 4, An eye tracking device further comprising a limbal triangulation device configured and operable to receive data indicating a geometric representation of a limbal structure for each image and to perform triangulation of the geometric representations of at least two images to identify three-dimensional limbal parameters.
7. 7. The eye tracking device of claim 6, The eye tracking device, characterized in that the processing unit is configured and operable to determine a three-dimensional position and gaze direction of a user's eyes based on the three-dimensional limbus parameters.
8. 8. An eye tracking device according to claim 6, An eye tracking device, characterized in that the three-dimensional limbus parameters include at least one of the three coordinates of the center of the limbus, the direction of the normal to the limbus plane, the position of the limbus plane, and the size of the limbus.
9. 9. An eye tracking device according to claim 6, 11. An eye tracking device, wherein the limbal triangulation device is configured and operable to determine annular radius and / or annular torsional rotation.
10. 10. An eye tracking device according to claim 1 , The eye tracking device further comprises a region detector configured and operable to receive each of at least two images showing a user's eyes, identify image data indicative of an initial limbus region for each image by using an iterative pixel filtering process, and generate data indicative of the initial limbus region.
11. 11. The eye tracking device of claim 10, 13. An eye tracking device, comprising: a first eye tracking device configured to detect an initial limbus region based on an anatomical parameter; and a second eye tracking device configured to detect an initial limbus region based on an anatomical parameter; and a region detector configured to detect an initial limbus region based on an anatomical parameter; and
12. 12. The eye tracking device of claim 11, 13. An eye tracking device, wherein the region detector is configured to use machine learning to identify image data indicative of eye features.
13. 13. An eye tracking device according to claim 11 or 12, The eye tracking device, wherein the region detector is configured and operable to identify image data indicative of eye features by segmenting each image to identify pixels associated with a pupil region.
14. 14. The eye tracking device of claim 13, An eye tracking device comprising: a limbal triangulation unit configured to perform triangulation between at least two geometric representations of the pupil to determine three-dimensional pupil parameters, the three-dimensional pupil parameters including at least one of a direction of a normal to a pupil plane, a coordinate of a center of the pupil, and a diameter of the pupil; and a limbal triangulation unit configured to estimate a position of an initial limbal region based on the three-dimensional pupil parameters and generate data indicative of the position of the initial region of the limbus in each image.
15. 15. An eye tracking device according to any one of claims 6 to 14, 13. An eye tracking device, comprising: a limbal triangulator configured to process the triangulated three-dimensional limbus, re-project the triangulated limbal region onto an image plane, and refine the limbal region in each image.
16. 1. A method for eye tracking, comprising: receiving image data indicative of at least two images of a user's eyes; identifying in each image a region associated with the limbus; identifying a geometric representation of a limbal structure; and performing triangulation of geometric representations of limbal structures in at least two images to identify a three-dimensional position and gaze direction of the user's eyes.
17. 17. The method of claim 16, 11. A method comprising capturing at least two images of a user's eye from different angles.
18. 18. The method of claim 16 or 17, The method, wherein identifying the geometric representation of the limbal structure includes digital image pre-processing.
19. 20. The method of claim 18, The method, wherein the digital image pre-processing includes calculating an image intensity gradient map of the limbus based on pupil direction, identifying in each image at least one region of the limbus structure in which the local direction of the gradient is substantially uniform, processing data indicative of the limbus structure by weighting pixels in such region, and generating a geometric representation of the limbus structure based on corresponding pixels associated with the corneal limbus.
20. 20. The method of claim 19, A method comprising the steps of: determining three-dimensional limbal parameters including at least one of three coordinates of a center of the limbus and an orientation of the limbus; and determining a three-dimensional position and gaze direction of a user's eyes based on the three-dimensional limbal parameters.
21. 21. The method of claim 20, 11. A method according to claim 10, wherein determining a three-dimensional position and gaze direction of a user's eyes comprises processing the three-dimensional limbus parameters.
22. 22. The method according to any one of claims 16 to 21, The method further comprising determining an annulus radius and / or annulus torsional rotation.
23. 23. The method according to any one of claims 16 to 22, further comprising: The method, wherein the geometric representation of the limbal structure comprises a ring-like or elliptical structure.
24. 24. The method according to any one of claims 16 to 23, further comprising: The method further comprising the steps of: identifying image data indicative of an initial limbus region in each image using an iterative pixel filtering process; and generating data indicative of the initial limbus region.
25. 25. The method of claim 24, A method, characterized in that in each image, identifying image data indicative of an initial limbal region includes at least one of identifying image data indicative of eye features such as the pupil, eyelid, sclera, iris, and eyelashes, and identifying the initial limbal region based on anatomical parameters.
26. 26. The method of claim 25, A method, characterized in that the identification of image data indicative of eye features is performed using machine learning.
27. 27. The method of claim 25 or 26, A method, wherein identifying image data indicative of eye features comprises segmenting each image to identify pixels associated with a pupil region.
28. 28. The method of claim 27, A method comprising the steps of: performing triangulation between at least two geometric representations of the pupil to determine three-dimensional pupil parameters, the three-dimensional pupil parameters including at least one of a direction of a normal to a pupil plane, a coordinate of a center of the pupil, and a diameter of the pupil in each image; estimating a position of an initial corneal limbus region based on the three-dimensional pupil parameters; and generating data indicating the position of an initial region of the limbus in each image.
29. 29. The method of any one of claims 16 to 28, further comprising: The method further comprises the steps of processing the triangulated three-dimensional limbus, reprojecting the triangulated limbus region onto the image plane, and refining the limbus region in each image.
Citation Information
Patent Citations
Apparatus, method and system for processing information, program and data structure
JP2011113550A
Data processing system, data processing method and data processing program
JP2016062524A
Gesture recognition device, gesture recognition method, projector equipped with gesture recognition device and video signal supply device
WO2018150569A1
Image processing device, image processing method for image processing device, and program
WO2019012817A1