System and method for eye tracking
By calculating the orientation changes of human retinal images using a retinal camera and polarized light source, this method solves the problems of large errors and position sensitivity in existing gaze tracking methods, and realizes accurate real-time gaze tracking and biometric applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-01-05
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, video-based eye tracking methods are limited by changes in ambient light and pupil size, and are sensitive to camera position, resulting in large estimation errors of the gaze target and making it difficult to achieve accurate real-time gaze tracking.
By using a retinal camera and a polarized light source, the orientation changes between images on the human retina are calculated. Combined with a processor, image matching is performed to determine the gaze direction, reducing sensitivity to changes in camera position.
It achieves precise gaze tracking that is insensitive to changes in camera position, and can accurately determine the gaze direction and target under real-time conditions, making it suitable for biometric and medical applications.
Smart Images

Figure CN113260299B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to eye tracking based on images of the human eye and retina. Background Technology
[0002] Eye tracking (also known as gaze tracking) used to determine gaze direction can be useful in various fields, including human-computer interaction control of devices such as industrial machines, aviation and emergency room situations requiring hands to perform tasks other than computer operation, virtual, augmented, or extended reality applications, computer games, entertainment applications, and research to better understand object behavior and visual processing. In fact, gaze tracking methods can be used in all ways people use their eyes.
[0003] The human eye is a two-part unit consisting of an anterior segment and a posterior segment. The anterior segment comprises the cornea, iris, and lens. The posterior segment consists of the vitreous humor, retina, choroid, and a white outer shell called the sclera.
[0004] The pupil of the eye is an aperture located in the center of the iris that allows light to enter the eye. The diameter of the pupil is controlled by the iris. Light entering through the pupil falls onto the retina, the innermost photosensitive layer of the outermost tissue coating the eye. Small pits (fovea centralis) are located on the retina and are specifically designed for maximum visual acuity, which is essential for human activities that require visual detail, such as reading and object recognition.
[0005] When an image of the world is formed on the retina, an image of the object being gazed at is formed on the fovea. That is, the position of the fovea corresponds to the direction of gaze.
[0006] Video-based eye trackers exist. Typically, video-based tracking uses corneal reflection and the center of the pupil as features for reconstructing the optical axis of the eye and / or as features for tracking in order to measure eye movement.
[0007] These methods are limited by factors such as image quality and / or pupil size changes in response to ambient light. Furthermore, in these methods, measurements are sensitive to the camera's position relative to the eye. Therefore, even small camera movements (referred to as 'slides') can introduce errors in eye orientation estimation, and thus large errors in gaze target estimation (especially for targets far from eye positioning).
[0008] The U.S. Publication No. 2016 / 0320837 (now Patent No. 10,248,194), assigned to MIT, combines images of retinal echoes (RRs) to create a digital (or reconstructed) image of the human retina. RRs are images that are not part of the retina. Known gaze directions of the eye can be used to determine the precise location of each RR for image reconstruction. The MIT Publication describes capturing a sequence of RRs and comparing this sequence with a database of known RR sequences to calculate the gaze direction. The MIT Publication itself explains that calculating the gaze direction from only a single RR is difficult because any single RR may arise from more than one gaze direction. Therefore, this method is not suitable for real-time gaze tracking.
[0009] U.S. Publication 2017 / 0188822, assigned to the University of Rochester, describes a method for compensating for eye movements during a scanning laser ophthalmoscopy examination. This disclosure addresses situations where the subject frequently gazes at the same target. The method described in this disclosure does not calculate the actual orientation of the eye or the person's actual gaze direction, or any changes thereto (e.g., angularly), and is not suitable for determining the gaze direction of an unknown target.
[0010] Accurate real-time gaze tracking remains a challenging task to date. Summary of the Invention
[0011] Embodiments of the present invention provide a system and method for gaze tracking using a camera. The camera is typically positioned such that at least a portion of the retina of the eye is imaged, and changes in eye orientation can be calculated based on a comparison of two images. According to embodiments of the invention, the gaze direction (and / or other information, such as the gaze target and / or gaze line) is determined based on the calculated changes in orientation.
[0012] In embodiments of the invention, the gaze direction of a person in an image associated with an unknown gaze direction can be determined based on another image of the retina of a person associated with a known gaze direction. In some embodiments, the gaze direction is determined by finding a spatial transformation between an image of the retina in the known gaze direction and a matching image of the retina in the unknown gaze direction. In other embodiments, the gaze direction is determined by calculating the position of the person's fovea based on the image of the retina. Thus, according to embodiments of the invention, gaze tracking is less sensitive to changes in the position of the camera relative to the eye. Therefore, the system and method according to embodiments of the invention are less prone to errors due to small movements of the camera relative to the user's eye.
[0013] A system according to an embodiment of the present invention includes a retinal camera comprising an image sensor for capturing images of a human eye and a camera lens configured to focus light originating from the human retina onto the image sensor. The system also includes a light source and a processor, the light source generating light emitted from a location near the retinal camera, and the processor calculating changes in the orientation of the human eye between two images captured by the image sensor.
[0014] The processor may further calculate one or more of the following based on the change in the orientation of the human eye between the two images: the orientation of the human eye, the direction of the person's gaze, the gaze target, and the gaze line.
[0015] The system may further include a beam splitter (e.g., a polarizing beam splitter) configured to direct light from the light source toward the human eye. The light source may be a polarized light source, and the system may also include a polarizer configured to block light originating from the light source and reflected by specular reflection.
[0016] The method according to embodiments of the present invention is used to effectively match retinal images with reference images to provide efficient and accurate gaze tracking, and can be used in other applications such as biometrics and medical applications. Attached Figure Description
[0017] The invention will now be described with reference to the following illustrative drawings, in conjunction with certain examples and embodiments, so as to provide a fuller understanding of the invention. In the drawings:
[0018] Figure 1A , 1B Figure 1C schematically illustrates a system operable according to an embodiment of the present invention;
[0019] Figure 2A and 2B A method for calculating changes in eye orientation according to an embodiment of the present invention is illustrated schematically, from which the direction of a person's gaze can be determined;
[0020] Figure 3A and 3B An example of a method for biometric identification according to an embodiment of the present invention is illustrated schematically;
[0021] Figure 4A and 4B A diagram illustrating the association of an image with a central concave position according to an embodiment of the present invention is shown schematically;
[0022] Figure 5A and 5B This is a schematic diagram of a method for determining a person's gaze direction using the rotation of a sphere, according to an embodiment of the present invention;
[0023] Figure 6 A method for comparing a reference image and an input image according to an embodiment of the present invention is illustrated schematically; and
[0024] Figure 7 The illustration schematically depicts a method for determining a person's gaze direction based on input from their eyes, according to an embodiment of the invention. Detailed Implementation
[0025] As described above, when a person gazes at a target, light entering the eye through the pupil falls onto the inner retina, the inner layer of the eyeball. The portion of the image that falls on the fovea is the image of the gazed target.
[0026] The line of sight corresponding to a person's gaze (also called the "gaze line" or "gaze") includes the origin of the light rays and their direction. The origin of the light rays can be assumed to be at the optical center of the lens of the human eye (hereinafter referred to as the 'Tens center'), while the direction of the light rays is determined by a line connecting the origin of the light rays and the gaze target. Each of a person's two eyes has its own gaze line, and under normal conditions, the two eyes meet at the same gaze target.
[0027] The position of the lens center can be estimated using known methods, such as by identifying the center of the pupil in the image and measuring the size of the iris in the image.
[0028] The direction of the gaze line originates from the orientation of the eyes. When the gaze target is near the eyes (e.g., about 30 cm or less), both the origin and direction of the gaze line are important for obtaining angular accuracy when calculating the gaze target. However, when the gaze target is located further away from the eyes, the direction of the light ray becomes significantly more important than the origin of the light ray for obtaining angular accuracy when calculating the gaze target. In the extreme case, when gazing at infinity, the origin has no effect on angular accuracy.
[0029] As any rigid object in 3D space, a complete description of eye pose has six degrees of freedom: three positional degrees of freedom (e.g., x, y, z) involving translational motion and three directional degrees of freedom (e.g., yaw, pitch, roll) involving rotation. The orientation and position of the eye can be measured in any frame of reference. Typically, in embodiments of the invention, the orientation and position of the eye are measured in the camera's frame of reference. Therefore, throughout the specification, even if a frame of reference is not explicitly mentioned, the camera's frame of reference is used.
[0030] Embodiments of the present invention provide a novel solution for finding eye orientation from which gaze direction and / or gaze target and / or gaze line can be derived.
[0031] The following illustrates a system and method for determining the orientation of a human eye according to embodiments of the present invention.
[0032] In the following description, various aspects of the invention will be described. Specific configurations and details are set forth for purposes of explanation in order to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention can be practiced without the specific details presented herein. Furthermore, well-known features may be omitted or simplified so as not to obscure the invention.
[0033] Unless otherwise specified, as will be apparent from the following discussion, it should be understood that throughout this specification, discussions using terms such as “analysis,” “processing,” “operation,” “calculation,” “determining,” “detecting,” “identifying,” “creating,” “generating,” “predicting,” “finding,” “trying,” and “selecting” refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulate and / or convert data represented as physical quantities (e.g., electronic quantities) in the registers and / or memory of the computing system into other data similarly represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system. Unless otherwise stated, these terms refer to the automatic actions of the processor, independent of and without any human intervention.
[0034] In one embodiment, Figure 1A The system 100, schematically illustrated, includes one or more cameras 103 configured to acquire images of at least a portion of one or both of a human eye 104. Typically, camera 103 is a retinal camera, which acquires an image of a portion of the human retina through the pupil of the eye with minimal interference or limitation to the person's field of view (FoV). For example, camera 103 may be located a few centimeters from the eye, at the periphery of the eye (e.g., below, above, or to the side). In one embodiment, camera 103 is located less than 10 centimeters from the eye. For example, the camera may be located 2 or 3 centimeters from the eye. In other embodiments, the camera is located greater than 10 centimeters from the eye, such as tens of centimeters or even several meters away.
[0035] Camera 103 may include a CCD or CMOS or other suitable image sensor and an optical system, which may include, for example, a lens 107. Additional optical components, such as mirrors, filters, beam splitters, and polarizers, may be included in the system. In other embodiments, camera 103 may include, for example, a standard camera equipped with a mobile device such as a smartphone or tablet.
[0036] Camera 103 images the retina using a suitable lens 107, converting light from a specific point on the retina into pixels on the camera sensor. For example, in one embodiment, the camera is focused at infinity. If the eye is also focused at infinity, the light originating from the specific point on the retina leaves the eye as a collimated beam (because in another direction, the eye focuses the incident collimated beam onto a point on the retina). Each collimated beam is focused by the camera onto a specific pixel on the camera sensor, depending on the direction of the beam. If neither the camera nor the eye is focused at infinity, a sufficiently sharp image of the retina can still be formed, depending on the camera's optical parameters and the precise focus of each in the camera and eye.
[0037] Because the eye is not a perfect lens, the wavefront of light rays exiting the eye is primarily distorted by the eye's lens and cornea compared to the wavefront of a perfect lens. This distortion varies from person to person and also depends on the angle of the outgoing wavefront relative to the eye's optical axis. Distortion can reduce the sharpness of the image of the retina captured by camera 103. In some embodiments, camera 103 may include lens 107 optically designed to correct for aberrations of light originating from the human retina, i.e., aberrations caused by the eye. In other embodiments, the system includes a spatial light modulator designed to correct for aberrations of light originating from the human retina by the distortions of a typical eye or a particular person's eye, and to correct for anticipated distortions at the angle of the camera position relative to the eye. In other embodiments, aberrations can be corrected by using appropriate software.
[0038] In some embodiments, camera 103 may include a lens 107 having a wide depth of field or an adjustable focal length. In some embodiments, lens 107 may be a multi-element lens.
[0039] The processor 102 communicates with the camera 103 to receive image data from the camera 103 and calculates changes in the orientation of the person's eyes based on the received image data, and may determine the direction of the person's gaze. The image data may include data such as pixel values representing the intensity of light reflected from the person's retina, and partial or complete images or videos of the retina or a portion of the retina.
[0040] Typically, in addition to the retina visible through the pupil, the image acquired by camera 103 may include different parts of the eye and a person's face (e.g., the iris, reflections on the cornea, sclera, eyelids, and skin). In some embodiments, processor 102 may perform segmentation on the image to separate the pupil (through which the retina is visible) from other parts of the person's eye or face. For example, a convolutional neural network (CNN) (such as UNet) can be used to perform pupil segmentation from the image. This CNN can be trained on an image in which the pupil is artificially labeled.
[0041] The eye reflects light back in approximately the same direction as it entered. Therefore, embodiments of the invention provide well-positioned light sources to prevent light from failing to return to the camera, causing the retina to appear too dark for proper imaging. Some embodiments of the invention include a camera with an accompanying light source 105. Thus, system 100 may include one or more light sources 105 configured to illuminate a person's eye. Light source 105 may include one or more illumination sources and may be arranged, for example, as a circular array of LEDs surrounding camera 103 and / or lens 107. Light source 105 may illuminate at wavelengths imperceptible to the human eye (and therefore inconspicuous); for example, light source 105 may include IR LEDs or other suitable IR illumination sources. The wavelength of the light source (e.g., the wavelength of each individual LED in the light source) can be selected to maximize the contrast of features in the retina and obtain a rich image with detail.
[0042] In some embodiments, the miniature light source may be positioned close to the camera lens 107, such as in front of the lens, on the camera sensor (behind the lens), or inside the lens.
[0043] In one embodiment, its instance is Figure 1B The image schematically shows an LED 105' located near the lens 107 of the camera 103 illuminating the eye 104.
[0044] Flicker can sometimes be caused by specular reflection (which is mostly smooth and glossy) of light from the anterior segment 134 of the eye, which can obstruct the image from the retina 124 and reduce its usefulness. Embodiments of the present invention provide a method for obtaining a flicker-reduced human eye image by using polarized light. In one embodiment, a polarizing filter 13 is applied to an LED 105' to provide polarized light. In another embodiment, the light source is a naturally polarized laser.
[0045] Light polarized in a specific direction (linear or circular) and directed toward the eye (arrow B) will be reflected back from the front of the eye 134 via specular reflection (arrow C), which primarily preserves the polarization (or, in the case of circular polarization, reverses it). However, light reflected from the retina 124 will be reflected back via diffuse reflection (arrow D), which randomizes the polarization. Therefore, using a filter 11 that blocks light with the original polarization (or reverse polarization, if circular) will be able to receive light reflected from the retina 124, but will block most of the light reflected from the front of the eye 134, thus essentially eliminating flicker from the image of the eye.
[0046] In another embodiment, its instances are in Figure 1CAs schematically shown, a beam splitter 15 may be included in system 100 to achieve the effect of light appearing to be emitted from a camera lens. The beam splitter can align light from light source 105 with camera lens 107, even if light source 105 is not physically close to camera lens 107.
[0047] When beam splitter 15 is used to generate a beam that appears to originate from inside or near camera lens 107, some light may be reflected back from the beam splitter (arrow A) and may cause glare that could obstruct the view of the retina. Glare from the beam splitter can be reduced by using polarized light (e.g., by applying polarizing filter 13 to light source 105) and a polarizing filter 11 in front of camera lens 107. Using polarized light to reduce glare from beam splitter 15 will also reduce flicker from the eyes because light reflected from outside the eyes (arrow C) is polarized in the same direction as light reflected directly from the beam splitter to the camera (arrow A), and both are orthogonally polarized to polarizing filter 11 on or in front of camera lens 107.
[0048] Therefore, in one embodiment, the light source 105 includes a filter (e.g., polarizing filter 13) or optical component to generate light with a specific polarization for illuminating the human eye 104, or a natural polarization source such as a laser. In this embodiment, the system 100 includes a polarizing filter 11 that blocks polarized light reflected back to the camera 103 (arrow C) from the anterior segment 134 of the eye, but allows diffuse reflection (arrow D) from the retina 124 to pass through and reach the camera 103, thereby obtaining an image of the retina with less obstruction. For example, the polarizing filter 11 may only allow light polarized perpendicular to the light source 105.
[0049] In some embodiments, system 100 includes a polarization beamsplitter. Using a polarization beamsplitter (possibly in combination with an additional polarizer) can provide similar benefits.
[0050] In embodiments where the camera 103 is positioned away from the eye's focal plane (e.g., the camera may be located 3 cm from the eye, while the eye is focused on a screen 70 cm from the eye), the light source 105 can be placed very close to the camera lens 107, for example, as a ring around the lens, rather than appearing to originate from the camera lens 107. This is still effective because the image of the light source 105 on the retina will be blurred, allowing some light to return and reach the camera in slightly different directions.
[0051] The processor 102, which may be locally embedded or remote, may include, for example, one or more processing units, including a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), a microprocessor, a controller, a chip, a microchip, an integrated circuit (IC), or any other suitable multi-purpose or specific processing or control unit.
[0052] Processor 102 typically communicates with memory unit 112, which can store at least a portion of image data received from camera 103. Memory unit 112 can be locally embedded or remote. Memory unit 112 may include, for example, random access memory (RAM), dynamic RAM (DRAM), flash memory, volatile memory, non-volatile memory, cache memory, buffer, short-term memory unit, long-term memory unit, or other suitable memory units or storage cells.
[0053] In some embodiments, memory unit 112 stores executable instructions that, when executed by processor 102, facilitate the execution of operations of processor 102, as described herein.
[0054] Processor 102 can communicate with light source 105 to control light source 105. In one example, a portion of light source 105 (e.g., different LEDs in a circular array) can be individually controlled by processor 102. For example, the illumination intensity and / or timing of a portion of the light source (e.g., each LED in the array) can be individually controlled, for example, in synchronization with the operation of camera 103. Different LEDs with different wavelengths can be turned on or off to obtain illumination of different wavelengths. In one example, the amount of light emitted by light source 105 can be adjusted by processor 102 based on the brightness of the captured image. In another example, light source 105 is controlled to emit light of different wavelengths so that different frames can capture the retina at different wavelengths, thereby capturing more detail. In yet another example, light source 105 can be synchronized with the camera shutter. In some embodiments, light source 105 can emit short bursts of very bright light to prevent motion blur, rolling shutter effects, or reduce overall power consumption. Typically, such short bursts of light are emitted at frequencies higher than human perception (e.g., 120 Hz) so as not to disturb people.
[0055] In one embodiment, system 100 includes one or more mirrors. In one embodiment, camera 103 is aligned with a mirror. The mirror and camera 103 can be positioned such that light reflected from the eye reaches the mirror and is reflected back to camera 103. This allows the camera to be positioned at the periphery of the eye without obstructing the user's field of vision (FoV).
[0056] In another embodiment, the camera is aligned with several small mirrors. The mirrors and the camera are positioned such that light reflected from the eye hits the mirrors and is reflected back to the camera. The position of each mirror determines the camera's viewpoint entering the pupil. Using several mirrors allows for the simultaneous capture of several viewpoints entering the eye, providing a wider field of view on the retina. Therefore, in some embodiments, system 100 includes a plurality of mirrors designed and arranged to simultaneously capture several viewpoints on the human retina.
[0057] In another embodiment, system 100 includes a concave mirror. The concave mirror and camera are positioned such that light reflected from the eye shines on the mirror and is reflected back to the camera. Using a concave mirror has an effect similar to placing the camera closer to the eye, producing an image with a larger FoV (focal distance) on the retina.
[0058] In some embodiments, system 100 includes one or more mirrors designed to reflect light of a predetermined wavelength. For example, the system may include one or more IR mirrors that are transparent to visible light. Such mirrors can be placed in front of the eye without obstructing the viewer's line of sight. IR light directed at the mirror from a light source (e.g., light source 105) near the camera can be reflected into the eye. The light is reflected from the retina back to the mirror, and then from the mirror back to the camera. The light reflected to the camera leaves the eye at a smaller angle relative to the eye's optical axis, where the optical performance is generally better, resulting in a sharper image. Therefore, this setup allows the camera to capture a sharper image of the retina without obstructing the viewer's FoV.
[0059] In an alternative embodiment, camera 103 includes a light sensor, while light source 105 includes a laser scanner (e.g., a MEMS laser scanner). Images of the eye and retina can be acquired using a MEMS laser scanner located near the light sensor (close enough for the light sensor to sense light returning from the retina through the pupil). A laser beam from the laser scanner passes through a lens or other optics located in front of the laser scanner and is designed to focus the laser onto a small area of the retina to produce a clear image of small retinal details. The laser beam travels rapidly through the eye, illuminating a different point at each instant. The photoelectric sensor senses the light returning from the eye (and from the retina, through the pupil), and by correlating the sensed value with the direction of the laser beam at that moment, an image can be created. This image is then processed as described herein.
[0060] Therefore, in one embodiment, system 100 includes a laser scanner located near a light sensor and pointed towards the human eye. The laser scanner is designed to emit laser light in multiple directions. A photoelectric sensor senses the value of the light returning from the retina of the human eye. Processor 102 creates an image of the retina based on the values sensed by the photoelectric sensor in each of the multiple directions of the laser scanner, and calculates the change in eye orientation between the two images created by the processor.
[0061] In some embodiments, the camera 103 may be located at a considerable distance from the eye (e.g., tens of centimeters or several meters). At this distance, the portion of the retina visible to the camera may be too small to obtain sufficient detail for reliable image comparison. In such cases, a camera with a larger lens (and / or concave mirror, and / or waveguide) can be used to capture a wider angular range of light rays emanating from the retina through the pupil.
[0062] Other alternative methods for obtaining images when the camera is far from the eye are described in further detail below.
[0063] In another embodiment, system 100 includes a waveguide configured to direct light from a person's retina to camera 103. In some embodiments, near-eye AR or VR optical systems (e.g., Lumus) TM DK-Vision, Microsoft TM Hololens or Magic Leap One TM This invention utilizes waveguide technology to project virtual images onto a user's eyes while simultaneously transmitting external light to the eyes, allowing the user to see both virtual images and real-world scenes. Embodiments of the invention provide gaze tracking compatible with such AR or VR optical systems. For example, a light source 105 can be positioned (typically outside the user's FoV) such that light is transmitted to the user's eyes via a waveguide. As described herein, light reflected from the user's retina is transmitted via the waveguide to a camera (e.g., camera 103) to provide a processor 102 with an image of the retina, which is used to calculate changes in eye orientation and, optionally, the user's gaze direction. The light reflected from the retina will be incident on the waveguide at the same location in the eye where the light from light source 105 is directed. From there, the light propagates through the waveguide to camera 103. When using a waveguide, light rays entering at a mirror angle will merge. Therefore, light source 105 can be positioned at a mirror angle to camera 103 to achieve the effect of the light source illuminating from near the camera. Therefore, when providing gaze tracking to an AR or VR system, the light source 105 and the camera 103 do not need to be physically close to each other.
[0064] Retinal images acquired via waveguides in near-eye AR or VR optical systems typically encompass a large portion of the retina. Furthermore, retinal features may appear darker relative to other obstacles that might appear in the same area of the image. For example, due to the nature of waveguide imaging and how it processes non-collimated light, blurred images of the outer regions of the eye may overlay faint but sharp images of retinal features. In some embodiments, images can be generated using a camera with high bit depth, where these faint but sharp retinal features can be extracted, for example, by using a high-pass filter on the image.
[0065] In some embodiments, optical elements may be used before the light source 105 to collimate the beam before the waveguide. The collimated light appears to the user as if it were emitted from a distant point light source and will be focused by the eye onto a small area on the retina. This small area will appear brighter in the image acquired by the camera, thus improving its quality.
[0066] In some embodiments, processor 102 may communicate with a user interface device including a display (e.g., a monitor or screen) to display images, instructions, and / or notifications to a user (e.g., via text or other content displayed on the monitor). In other instances, processor 102 may communicate with a storage device such as a server, including, for example, volatile and / or non-volatile storage media such as hard disk drives (HDDs) or solid-state drives (SSDs). As further described herein, storage devices that can be connected locally or remotely (e.g., in the cloud) may store and allow processor 102 to access databases of reference images, graphs that associate reference images with known gaze directions and / or foveal positions, and so on.
[0067] The components of system 100 can be wired or wireless, and may include suitable ports and / or network hubs.
[0068] In one embodiment, system 100 includes a camera 103 to acquire a reference image and an input image. The reference image includes at least a portion of a person's retina associated with a known orientation of the eye and / or a known gaze direction and / or a known gaze target and / or a known gaze line. The input image includes at least a portion of a person's retina associated with an unknown orientation of the eye and / or an unknown gaze direction and / or an unknown gaze target and / or an unknown gaze line. For simplicity, and because each of these parameters (eye orientation, gaze direction, gaze target, and gaze line) is relevant, other parameters are included when referring to "gaze direction" herein.
[0069] The processor 102, which communicates with the camera 103, determines the gaze direction of a person in the input image based on a reference image, a known gaze direction, and the input image.
[0070] In one embodiment, processor 102 receives an input image and compares it with a reference image to find a spatial transformation (e.g., translation and / or rotation) of the input image relative to the reference image. Typically, a transformation is found that optimally matches or overlays the retinal features of the reference image onto the retinal features of the input image. The orientation change of the human eye corresponding to the spatial transformation is then calculated. The gaze direction associated with the input image can then be determined based on the calculated transformation (or based on the gaze direction of the human eye associated with the input image). The signal based on the orientation change and / or based on the gaze direction of the human eye associated with the input image can be output by processor 102.
[0071] In some embodiments, the calculated transformation is converted into a corresponding spherical rotation, and the gaze direction associated with the input image is determined by rotating a known gaze direction associated with a reference image using the spherical rotation.
[0072] A spherical rotation is defined as a linear mapping from 3D space to itself, preserving length and chirality. That is, this is equivalent to a 3×3 real matrix R such that RR . t =I and det(R) = 1.
[0073] The transformation of the reference image relative to the input image is the inverse of the transformation of the input image relative to the reference image. With necessary modifications, the gaze direction associated with the input image can also be determined based on the transformation of the reference image relative to the input image. Therefore, the term "transformation of the input image relative to the reference image" or similar terms used herein also implies the inclusion of the transformation of the reference image relative to the input image.
[0074] An image of the human eye can include the fovea and the retina, primarily featuring retinal characteristics such as patterns of blood vessels that provide information about the retina.
[0075] Throughout this specification, the calculation of eye orientation or gaze direction is explained with reference to the pixels of the fovea. The terms "pixel of the fovea" or "position of the fovea," or similar terms used herein, may refer to an actual pixel (if the fovea is visible) or a theoretical pixel (as described below) (if the fovea is not visible). A pixel is the camera projection of the gaze direction vector onto the camera sensor. The theoretical (or calculated) position of the fovea refers to a theoretical pixel outside the visible portion of the retina in an image, and may even (and usually) be outside the camera's FoV (i.e., the theoretical pixel's position is outside the image sensor). The calculated (theoretical) pixel position of the fovea corresponds to an angle in the camera's frame of reference. The gaze direction of a person can be calculated from this angle. For an image associated with a known gaze direction, the pixel position of the fovea is calculated from this angle.
[0076] The reference to the fovea pixels is intended to illustratively explain the relationship between gaze direction and retinal image. However, according to embodiments of the invention, spherical rotation and other spatial transformations can be applied directly to eye orientation or gaze direction (as a 3D vector) without first projecting it onto the fovea pixels (whether practical or theoretical).
[0077] There exists a one-to-one mapping between the location of retinal features in an image and the orientation of the eye relative to the camera that captured the image. This mapping is a function of eye orientation and camera features. It is independent of the eye's position relative to the camera. To explain, imagine that both the eye and the camera are focused at infinity. Light reflected from the retina is projected into space as an angular pattern. That is, each point on the retina corresponds to a collimated beam at a specific orientation. If the eye's orientation in space is fixed, then the orientation of each collimated beam is fixed regardless of the eye's spatial position. The camera, in turn, focuses each collimated beam onto a specific pixel on the camera sensor according to the beam's orientation. Therefore, for each eye orientation, each retinal feature corresponds to a collimated beam at the corresponding orientation, which is focused by the camera onto a specific location in the image.
[0078] In one embodiment, the direction of a light beam originating from a fovea can be calculated based on pixels in an image where a fovea appears. For example, in a pinhole camera with a focal length f and an optical center at pixel (0,0), the fovea position or theoretical position (x,y) on the camera sensor corresponds to the beam orientation described by a unit vector.
[0079] In one embodiment, the gaze direction can be determined from an image of the eye (e.g., the retina of the eye) based on a calculated foveal position. In this example, processor 102 receives a reference image from camera 103, which is obtained when the person is looking in a known direction. Since each gaze direction uniquely corresponds to a known foveal position, the foveal position on the camera sensor for each reference image can be determined. In other embodiments, as described above, the position of the fovea (actual or theoretical) corresponds to an angle in the camera reference frame. The person's gaze direction can be calculated based on this angle. Processor 102 compares an input image of the person's retina obtained from camera 103 while tracking the person's (unknown) gaze direction with a graph of the reference images. The graph associates portions of the person's retina (when they appear in the camera image) with the known position of the person's fovea. Processor 102 can identify one or more reference images that are associated with or match the input image, and can then determine a spatial transformation (e.g., translation and / or rotation) of the input image relative to the corresponding reference image. The determined translation and / or rotation can be used to calculate the position of the person's fovea in the input image to obtain the calculated foveal position. The calculated (theoretical) position of the fovea corresponds to the theoretical angle at which the camera is mapped onto a theoretically larger sensor. This theoretical angle is the angle of gaze of a person about the camera axis.
[0080] Therefore, the gaze direction of the person associated with the input image can be determined based on the calculated foveal position of the person. The processor 102 can then output a signal based on the determined gaze direction.
[0081] In some embodiments, the fovea can be directly identified in its visible image (e.g., when a person is gazing at or near the camera). For example, the fovea can be identified based on the different color / brightness of the fovea compared to the retina. Images in which the fovea is identified can be stitched together to a panoramic image (as further described below) to provide a panoramic image including the fovea.
[0082] exist Figure 2A In one example schematically illustrated, a method for determining a person's gaze direction, executable by processor 102, includes receiving an input image comprising at least a portion of the person's retina (step 22). In some embodiments, the input image is associated with an unknown gaze direction. The input image is compared with a reference image comprising at least a portion of the person's retina (step 24). Typically, the reference image is associated with a known gaze direction. Comparing the reference image with the input image may include, for example, finding a spatial transformation (e.g., a rigid transformation, such as translation and / or rotation) between the input and reference images. The orientation change of the human eye corresponding to the spatial transformation is then calculated (step 26), and a signal is output based on the orientation change (step 28).
[0083] The method according to embodiments of the invention can be performed using optimal computation (e.g., assuming the lens of the eye is a perfect lens and / or assuming the camera is a pinhole camera). However, since neither is perfect, in some embodiments, the obtained images (e.g., the input and reference images) are processed to correct distortions (e.g., geometric distortions) before finding the spatial transformation between the input and reference images. In one embodiment, the processing may include making the input and reference images geometrically undistorted. Thus, processor 102 can geometrically undistort the input image to obtain a distortion-free input image, and can obtain a comparison between the distortion-free input image and the distortion-free reference image (e.g., in step 24). The calculation of geometrically undistorted (correcting image distortion) is described below. Additionally, known software can be used for image processing to obtain distortion-free images.
[0084] The signal output in step 28 may be a signal indicating the gaze direction or, for example, a signal indicating a change in the gaze direction. In some embodiments, the output signal may be used to control a device, as further illustrated below.
[0085] exist Figure 2B In another embodiment schematically illustrated, a method for determining a person's gaze direction includes: receiving an input image of at least a portion of the person's eye (e.g., a portion of the retina) (step 202); and comparing the input image with a graph of a reference image, the reference image being an image of the portion of the person's eye associated with a known gaze direction (step 204). The person's gaze direction associated with the input image can be determined based on the comparison (step 206). In some embodiments, a signal can be output based on the person's gaze direction. The signal can be used to indicate a determined gaze direction, or, for example, to indicate a change in gaze direction. In some embodiments, the output signal can be used to control a device. Thus, in some embodiments, the device is controlled based on the person's gaze direction (step 208).
[0086] Reference images typically include a portion of the retina but exclude the fovea. Therefore, the known gaze direction associated with each reference image can be outside the camera's field of view. That is, projecting the gaze direction onto the image sensor will result in theoretical pixels outside the image frame.
[0087] Multiple transformed (e.g., translation and rotation) copies of the input image can be compared with a reference image using, for example, the Pearson correlation coefficient between the pixel values of each image or a similar function.
[0088] Each copy of the transformation is transformed differently (e.g., with different translations and / or rotations). The copy with the highest Pearson correlation to the reference image can be used to determine the transformations (e.g., translations and / or rotations) of the input image relative to the reference image and the corresponding orientation changes of the human eye.
[0089] In some cases, applying Pearson correlation to two dissimilar images can result in relatively high values, even if the images are not similar. This is because each image has high spatial autocorrelation (i.e., nearby pixels tend to have similar values). This can significantly increase the probability of accidental correlation between regions on two dissimilar images. To reduce this effect of false positives, Pearson correlation should only be applied to sparse descriptions of images with low autocorrelation. Sparse descriptions of retinal images can include, for example, the locations of extrema and sharp edges in the image.
[0090] Therefore, in some embodiments, retinal image comparisons are performed by using a sparse description of the image (e.g., when using Pearson correlation).
[0091] As described in step 208 above, the device can be controlled based on the person's gaze direction. For example, when the direction changes, the user interface device can be updated to indicate the person's gaze direction. In another embodiment, another device can be controlled to issue a warning or other signal (e.g., an audio signal or a visual signal) based on the person's gaze direction.
[0092] In some embodiments, devices controlled based on a person's gaze direction may include industrial machinery, devices used in navigation, aviation, or driving, devices used in medical procedures, devices used in advertising, devices using virtual or augmented reality, computer games, devices used in entertainment, etc. Alternatively or additionally, according to embodiments of the invention, devices for biometric user identification may be controlled based on a person's gaze direction.
[0093] In another example, embodiments of the invention include devices using virtual reality (VR) and / or augmented reality (AR) (which can be used in various applications such as gaming, navigation, aviation, driving, advertising, entertainment, etc.). The gaze direction of a person determined according to embodiments of the invention can be used to control the VR / AR device.
[0094] In one embodiment, the apparatus used in a medical procedure may include the components described in FIG1. For example, the system may include a retinal camera consisting of image sensors to capture images of a person's retina, each image associated with a different gaze direction of the person. The system may also include a processor for finding spatial transformations between the images and stitching the images together based on the spatial transformations to create a panoramic image of the person's retina.
[0095] In some cases, obtaining retinal images of a person's retina requires the person to gaze at extreme angles so that these images are correlated with the person's extreme gaze direction and a wider field of view of the retina can be obtained.
[0096] The processor can enable the panoramic image to be displayed, for example, for viewing by professionals.
[0097] In some embodiments, the processor may run a machine learning model for predicting the health status of a human eye based on panoramic images, the machine learning model being trained on panoramic images of the retina of the eye under different health states.
[0098] Viewing a panorama according to embodiments of the invention provides a non-invasive and patient-friendly medical procedure (e.g., a procedure to examine a person's optic nerve or other features of the retina), which to date involve administering pupil-dilating eye drops and shining bright light into the eye, and can produce a narrower field of view in the retina. Furthermore, since a panorama according to embodiments of the invention can be obtained using simple and inexpensive hardware, this non-invasive medical procedure can be used in home settings and / or remote areas.
[0099] In the following text Figure 3A and 3B In some embodiments illustrated schematically, the apparatus for biometric user identification or authentication may include the components described in FIG1.
[0100] In one example, a biometric device includes one or more cameras to acquire a reference image and a recognition image. The reference image includes at least a portion of a person's retina. Some or all of the reference images are acquired in a known gaze direction. In some embodiments, the reference images are associated with a known location of a person's fovea in the image, or with a theoretical pixel of the fovea. The recognition image includes at least a portion of a person's retina in a known gaze direction or in a different known gaze direction.
[0101] In some embodiments, the apparatus for biometric user identification includes two cameras, each arranged to acquire a reference image and a recognition image for each of the person's eyes. In some embodiments, two or more cameras are arranged to acquire a reference image and a recognition image for the same eye. Reference images and recognition images for different eyes of a person can also be obtained directly or via a mirror or waveguide using a single camera having both eyes in its field of view.
[0102] In one embodiment, the apparatus for biometric user identification includes a processor communicating with one or more cameras and a memory for maintaining reference images associated with a person's identity. The processor compares a person's identification image with the reference images and determines the person's identity based on the comparison.
[0103] In one embodiment, the identification image is compared by performing one or more spatial transformations (e.g., rigid transformations) on it and calculating the correlation strength between the transformed identification image and a reference image. The processor then determines the person's identity based on the correlation strength.
[0104] For example, the strength of the correlation can be determined by comparing the transformed recognition image with a reference image using the Pearson correlation coefficient. In some embodiments, a maximum Pearson coefficient is calculated, and the correlation strength can be measured by the probability that the retina of a random person will reach this maximum coefficient. This probability can be calculated, for example, using a database of retinal images of many people. Random people are examined against this database, and the probability can be calculated based on the number of correct (but incorrect) recognitions by the random person. Alternatively, the measurement of the correlation strength (e.g., the Pearson correlation coefficient) can be calculated for each random person, and the probability distribution function (pdf) measured among the random people can be estimated from these samples. Thus, the probability of a false match would be the integral of this pdf from the correlation strength of the recognition image to infinity.
[0105] Figure 3A An example of a device for controlling biometric user identification based on a determined gaze direction, according to an embodiment of the present invention, is illustrated schematically.
[0106] In the initial stage (stage I), a reference image of the human eye is obtained (step 302). For example, the reference image can be obtained by a camera 103 located less than 10 cm away from the human eye, with minimal interference or limitation on the human's field of vision.
[0107] In some embodiments, when a person gazes in a known direction X, the reference image includes at least a portion of the person's retina. The obtained reference image is stored (step 304) in, for example, a user identity database. Typically, multiple reference images of a single person (obtained in different known directions) and / or different people can be stored in the database, each of one or more reference images being associated with the identity of the specific person whose eye was imaged.
[0108] In the second recognition stage, which is typically the later stage (stage II), the person is directed to gaze in the direction Y and a recognition image is obtained (step 306). Typically, the recognition image is obtained by the same camera 103 used to obtain the reference image (or a camera that provides the same optical results).
[0109] In step (308), the processor (e.g., processor 102) determines whether the one or more identification images (obtained in step 306) correspond to the person based on one or more reference images (images obtained in step 302) associated with the person's identity. Based on this determination, the person's identification is issued (step 310).
[0110] In some embodiments, in the determination step (step 308), the processor calculates the gaze direction in the identification image using a reference image of individual I (e.g., by finding a rigid transformation of the identification image relative to the reference image of individual I and using the same gaze tracking technique described herein). The processor then compares the calculated gaze direction with a known gaze direction Y of the identification image. If the calculated gaze direction matches the known gaze direction Y, and the identification image with the rigid transformation has a high correlation with the reference image, then the person is affirmatively identified as individual I. In this document, high correlation (or “highly correlated”) refers to a result that is unlikely to occur when matching two unrelated retinal images. For example, high Pearson correlation between two images (or a sparse description of the images), or high Pearson correlation between pixels sampled from two images.
[0111] Figure 3B A method for biometric identification according to an embodiment of the present invention is illustrated schematically.
[0112] In step 312, while the person is looking in a first known direction, the processor (e.g., processor 102) receives a reference image 31 of at least a portion of the person's retina 33. The position 1 of the fovea 30 of the reference image is known.
[0113] In step 314, the unidentified person receives identification image 32. The identification image is obtained while the unidentified person is looking in the second gaze direction. The position 2 of the central fovea 30 of the identification image is known.
[0114] In step 316, a transformation (e.g., translation and / or rotation, indicated by the dashed curved arrow) of the identification image 32 relative to the reference image 31 is found. If the identification image includes a different retina (or a different portion of the retina) than the reference image, no suitable transformation will be found, the process will abort, and the person will not be identified.
[0115] In step 318, the processor calculates the predicted position 3 of the central fovea in the recognition image based on the known central fovea position 1 and by using the relative transformation (e.g., rotation and translation) calculated in step 316.
[0116] In step 320, the known foveal position 2 and the predicted foveal position 3 are compared. If the known foveal position 2 and the predicted foveal position 3 match, and the correlation strength of the transformed recognition image relative to the reference image is high (e.g., above a threshold), then the person is affirmatively identified (step 322). If the positions do not match and / or the correlation strength is below the threshold, then the person is not identified.
[0117] The threshold can be a predetermined threshold or, for example, a calculated threshold. For instance, a threshold could include the probability that a random person is identified with certainty (but incorrectly), as described above.
[0118] The above steps can be repeated using several reference images (e.g., as described herein). In some embodiments, the above steps can be repeated on a sequence of recognition images obtained while a person is gazing at a moving target.
[0119] The above steps can be performed using other embodiments of the invention, such as by transforming the reference image and the recognition image and calculating the gaze direction without calculating the position of the fovea, as described herein.
[0120] In some embodiments, a reference image and a recognition image may be obtained, and the calculations described above may be performed on the human eyes.
[0121] For example, an identification issued in step 310 and / or step 322 can be considered a strong positive biometric if one or more of the following conditions are met (for one or both eyes of a person):
[0122] 1. The retinal pattern in the human identification image is matched with the retinal pattern in the reference image associated with the identity of the same person.
[0123] 2. Based on a reference image obtained when the same person is looking in direction X, the predicted positions of the retinal pattern and / or fovea in the identified image obtained when the person is looking in direction Y are matched with the expected transformations (e.g., translation and rotation) of the predicted positions of the retinal pattern and / or fovea. Since the retina must respond to commands to gaze in different directions in a certain way, and this way is natural only for real people, for a particular person to be reliably identified, the portion of the retina of a specific person must be strongly correlated with a specific reference image. This step provides protection against malicious users exploiting images or recordings of, for example, another person's retina.
[0124] 3. The sequence of recognized images obtained when a person gazes at a moving target is matched with the sequence of desired images and / or reference images.
[0125] Reference images and recognition images can be analyzed according to embodiments of the present invention. In one embodiment, the analysis includes finding a spatial transformation between the recognition image and the reference image, calculating an eye orientation change based on the spatial transformation, and determining, based on the eye orientation change, that the reference image and the input image originate from the retina of the same person. The analysis may include determining whether the orientation change is consistent with a gaze target associated with each of the reference image and the input image.
[0126] In some embodiments, the reference image may include a stitched image of multiple retinal images (also referred to as a panorama) covering a larger area of the human retina than a single retinal image. The panorama may be associated with a known gaze direction (and / or a known location of the person's fovea). The identification image may include a portion of the retina obtained from the person when they are gazing in a known direction other than X, or when they are gazing in an unknown direction. The identification image may be compared to a panorama of the person's retina to search for matches, as described herein, and the person's identification may be published based on the matches.
[0127] If the orientation change (determined using known eye-tracking methods and / or by using the methods described herein) is consistent with each associated gaze target in the reference and input images (as described above), and the identified image is highly correlated with the reference image, then the person is identified with certainty.
[0128] In some embodiments, as described below, the reference image and the recognition image are processed (e.g., flattened) before the step of matching or associating the input image with the reference image.
[0129] In some embodiments, the biometrics described above can be used in conjunction with other identification methods such as RFID cards, fingerprints, and three-dimensional facial recognition. Combined, this can significantly reduce the chance of false identification. For example, firstly, a fingerprint scanner can quickly, but with some degree of uncertainty, identify a specific person (or one of a few candidates) from a large group of people. Then, as described above, the device compares the person's retinal image with the identified candidate to obtain a more definitive, positive identification.
[0130] Now for reference Figure 4A and 4B The diagrams and analyses illustrating an embodiment of the invention will be shown schematically, which associate a portion of a person's retina imaged by a camera with the direction of gaze.
[0131] For example, Figure 4A One option illustrated in the diagram includes creating a panoramic image of the human retina, which will be used as a reference image. Panoramic images typically cover all or most of the retina visible to the camera in a generally reasonable direction of gaze. In one embodiment, a panoramic image can be used to obtain a wider retinal field of view by requiring the person to gaze at an extreme angle.
[0132] The panoramic image may or may not include the fovea. In this embodiment, several images of the retina (possibly a portion thereof) are obtained, without necessarily knowing the gaze direction of the eye when each image is obtained. These images are then stitched together to construct a panoramic image 401 of the retina. Stitching can be done using standard techniques, such as feature matching and / or finding regions where two images share very similar areas (e.g., via Pearson correlation or difference of squares), merging two images into one image, and repeating until all images are merged into one. Prior to stitching, the images may be projected onto the surface of a 3D sphere, as further described below. The foveal position relative to the panoramic image 401 can be calculated based on one or more images of a portion of the retina (the images constituting the panoramic image), for which the gaze direction is known.
[0133] Input image 405 is acquired while tracking a person's gaze. The direction of the person's gaze is unknown when input image 405 is acquired. Typically, input image 405 includes only a portion 45 of the retina 42 and does not include the fovea.
[0134] The input image 405 is compared with the panoramic image 401 to find a match between the pattern of the retina 42 in the input image 405 and the pattern of the retina 43 in the panoramic image 401. Spatial transformations (e.g., translations and / or rotations) of the pattern of the retina 42 in image 405 relative to the pattern of the retina 43 in the panoramic image 401 can be found. For example, the processor 102 can identify features in the retina that appear in the input image 405 and the panoramic image 401. The spatial correlation between the features in the panoramic image 401 and the input image 405 can provide the spatial transformations (e.g., translations and / or rotations) required to align the input image 405 with the panoramic image 401. The position of the fovea 40 can be transformed accordingly (e.g., translated and / or rotated) to obtain the calculated position of the fovea 40'. The gaze direction associated with the input image 405 corresponds to the calculated position of the fovea 40'.
[0135] In some embodiments, the camera imaging the human eye is positioned at a distance from the eye that is insufficient to obtain sufficient detail for reliable image comparison (i.e., the image may accidentally match other unrelated areas of the retina incorrectly). In this case, several input images are acquired over a short period (e.g., 1 second) at a high frame rate (e.g., 30 frames per second or higher) as the eye or camera moves. The input images are stitched together. Typically, each consecutive input image will provide a slightly different retinal field of view (either due to slight eye movement or slight camera movement), thus stitching together several input images provides a single panoramic input image covering a wider area of the retina. This panoramic input image can then be compared with panoramic image 401, as described above, to determine the person's gaze direction associated with the time period during which the several input images were acquired.
[0136] In other embodiments, the system for imaging a human eye from a distance may include large-diameter lenses to capture a wider field of view. In another embodiment, the system images both eyes of a person (as described below), but from a distance, and determines the possible gaze direction for each eye, accepting only direction pairs that coincide at a single gaze target. In other embodiments, a first, less precise method for determining the gaze direction may be used, for example, using pupil position and / or pupil shape and / or corneal reflex. According to embodiments of the invention, a more precise method for determining the gaze direction using retinal images can be used on the image generated by the first method, thereby limiting the search space for spatial transformations between images, thus allowing operation from a distance.
[0137] Figure 4B Another exemplary option of the diagram schematically shown includes a calibration step of obtaining an image of a portion of the human retina in a known gaze direction. The diagram can be created using a camera to obtain several images of the retina (or a portion thereof) while the person views several potentially arbitrary but known directions. A different image is obtained for each different direction. Each image is associated with a known gaze direction. In summary, these images comprise a diagram of reference images. For example, reference image 411, including a portion of retina 421, corresponds to a first known gaze direction and a known or calculated foveal position (x1, y1). Reference image 412 includes a different portion of retina 422 and corresponds to a second (and different) known gaze direction and a known or calculated foveal position (x2, y2).
[0138] The input image 415 (obtained to track a person's gaze, with the gaze direction unknown) is then compared with reference images 411 and 412 to identify matching reference images, i.e., matching portions of the retina with known gaze directions. Figure 4BIn the example shown, input image 415 is matched with reference image 412 (e.g., based on similar features). Once a matching reference image (reference image 412) is found, a spatial transformation of input image 415 relative to the matching reference image 412 can be determined and applied to a known or calculated foveal position (x2, y2) to obtain a calculated position (x3, y3) of the foveal relative to input image 415. A third (and different) gaze direction can then be calculated for input image 415 based on the calculated foveal position (x3, y3).
[0139] Typically, the input and reference images are obtained using the same camera or a camera that provides similar optical results.
[0140] A diagram or other construct that maintains the association between a portion of the retina and the gaze direction and / or the known or calculated position of the person's fovea may be stored in a storage device and may be accessed by the processor 102 to determine the gaze direction according to an embodiment of the invention.
[0141] Therefore, some embodiments of the present invention use the foveal position of a person to determine the gaze direction and track the person's eye gaze.
[0142] As described above, the position of a person's fovea can be calculated based on spatial transformations between the input image and the reference image (e.g., translation and / or rotation of one or more features).
[0143] In one embodiment, a method includes calculating the position of a person's fovea in an input image based on a spatial transformation to obtain a calculated foveal position, and determining one or more of the following associated with the input image based on the calculated foveal position: the person's eye orientation, gaze direction, gaze target, and gaze line.
[0144] In one embodiment, multiple images of the retinal portion obtained in a known gaze direction (and optionally associated with a known foveal position) are used as training instances for a statistical model, such as a linear regression model, a multinomial regression model, a neural network, or a suitable machine learning model.
[0145] In one embodiment, the orientation change of a person's eyes can be calculated using a statistical model trained on multiple pairs of reference images and input images. The statistical model can be trained on multiple spatial transformations between the multiple pairs of reference images and input images. The training objective of the model is to predict the orientation change of a person's eyes for each pair.
[0146] In one embodiment, the processor receives an input image comprising at least a portion of a human retina and finds a spatial transformation between the input image and a reference image (which also comprises at least a portion of a human retina). The processor may provide the spatial transformation to a statistical model and optionally provide one or more of the following: a known orientation of the human eye associated with the reference image, a known gaze direction of the person associated with the reference image, a known gaze target associated with the reference image, and a known gaze line associated with the reference image. The statistical model is trained on a set of input parameters, each of which contains at least, for example, a spatial transformation between a training reference image and a training input image, and each of which is associated with output parameters comprising at least one of the following: a known orientation of the eye associated with the training input image, a known gaze direction associated with the training input image, a known gaze target associated with the training input image, and a known gaze line associated with the training input image.
[0147] The statistical model can then be used to estimate at least one of the following: changes in human eye orientation, human eye orientation associated with the input image, gaze direction associated with the input image, gaze target associated with the input image, and gaze line associated with the input image.
[0148] In some embodiments, the processed image (e.g., a flattened image as described below) is used as a training instance of the model. The model can then be used to predict the orientation of a person's eyes and / or changes in the gaze direction of the eyes in an input image (e.g., an image of a portion of the retina of an eye in an unknown gaze direction).
[0149] For example, a linear regression model can be used to predict gaze direction. The model is matrix A such that for a given set of input data x, the predicted gaze direction is Ax. Here, the input data is all the information fed into the linear regression model and includes data from the input image and the reference image, as well as, for example, the relative translation and rotation between them.
[0150] To train the model, a large set of images can be captured, and the user's associated gaze target can be recorded for each image. For each pair of captured retinal images, one image is used as the input image, and the other as the reference image. The two images are compared to find spatial transformations, where, if any, the two images are optimally aligned with each other. A large set of spatial transformations can be obtained. For each spatial transformation i, a column of input data x is generated. i and a column of target data y iInput data may include, for example, the location of the pupil center in the input image; the location of the pupil center in the reference image; the gaze target when obtaining the reference image; and parameters of spatial transformations (e.g., relative translation and rotation) between the input and reference images. Target data may include the gaze target when obtaining the input image. All columns x i Together they form the training input matrix X, while column y i The training target matrix Y is constructed. Linear regression is used to find the optimal mapping matrix A from X to Y. For an input image and a reference image that generate the input data x, the predicted gaze target is Ax. In another instance, the input data x... i It contains information about the eyes.
[0151] Similarly, models other than linear regression, such as multinomial regression, neural networks, machine learning models, or any other model trained on instances, can be used. All of these models are referred to herein as statistical models.
[0152] For a single input image, the process can be repeated with different reference images. For each reference image, the model outputs a prediction. The average of the output (or any other statistic) can be used as a better estimate of changes in the eye's gaze direction or orientation.
[0153] The input image typically includes a portion of the human retina but excludes the fovea (because people rarely look directly at a camera), while the reference image may or may not include the fovea, as described above.
[0154] Features that can be used to determine translation and rotation between an input image and a reference image may include lines, specks, edges, dots, and textures. These features may correspond to blood vessels and other biological structures visible in the retina.
[0155] As mentioned above, due to the way light is reflected and diffused within and from the eye's surface, the signals in an image corresponding to features on the retina can be very weak. The image of the retina may have uneven brightness levels. The image may include whiter and darker areas that are not part of the retinal pattern (which can be caused by uneven light from a light source, shadows from the iris, light reflections within the eye, refraction from the cornea, or other unknown causes). These areas do not move with visible features of the retina (such as blood vessels) when the gaze direction changes, and therefore should be ignored.
[0156] In some embodiments, to compare features between two retinal images, the images are processed, such as flattened, before the comparison. For example, a reference image and an input image may be flattened before finding the spatial transformation between them. In one instance, the image may be processed to remove regions of uneven brightness and flatten its intensity spatial distribution. One way to achieve this is by applying a high-pass filter to the original input image or by subtracting a blurred version of the original input image from the most original input image. Another way is by using a machine learning process (such as a CNN, e.g., UNet CNN) trained to extract flattened images of the retina from the original input image. The machine learning process can be trained with a training set consisting of typical original retinal images given as input to the learning process and flattened retinal images used as the target output of the learning process. The flattened retinal images in the training set can be drawn manually, manually identified on a high-pass image, or copied from, for example, a complete retinal image obtained with a specialized fundus camera.
[0157] Other processing methods can be used to obtain a flattened image.
[0158] In one embodiment, camera 103 acquires an input image of at least a portion of a person's retina, and processor 102 acquires a processed input image, for example, by flattening regions of uneven brightness in the input image. Processor 102 then compares the processed input image with a graph that associates a known gaze direction with a previously acquired and processed reference image of the person's retina to determine the person's gaze direction based on the comparison. The device can be controlled based on the person's gaze direction (e.g., as detailed above).
[0159] In some embodiments, comparing the processed image with a graph that associates a known gaze direction with a previously processed image to find the best-correlated image can be done by calculating the highest Pearson correlation coefficient, as described above.
[0160] In some embodiments, finding a spatial transformation between an input image and a reference image (which could be a processed (e.g., flattened) input image and a reference image to be processed) includes attempting multiple spatial transformations between the input and reference images, calculating a similarity metric between the input and reference images corresponding to each spatial transformation, and selecting a spatial transformation based on the similarity metric. For example, a spatial transformation with the highest similarity metric can be selected. The similarity metric can be a Pearson correlation coefficient, the average of squared differences, etc., and can be applied to the entire image or parts thereof, as described above.
[0161] In one embodiment, the input image and reference image (and images of other objects emitting the angular pattern) can be more precisely aligned by mapping one to another via 3D spherical rotation (which is equivalent to a 3D rotation matrix). Each pixel in the image corresponds to an orientation, a spatial direction. Each point on the retina also corresponds to an orientation, since the eye works like a camera, where the retina is the sensor. Each orientation can be considered a unit vector, and all possible unit vectors make up the surface of a unit sphere. The image from the camera can be considered to lie on a patch of this sphere. Thus, in one embodiment of the invention, a planar camera image is projected onto the sphere, and instead of (or in addition to) comparing the planar translations and rotations of the two retinal images, they are compared by the rotation of the sphere. The rotation of the sphere can be about any axis and is not limited to 'rolling'. Spherical rotation has three degrees of freedom, which are analogous to the two degrees of freedom of planar translation plus one degree of freedom of planar rotation.
[0162] The rotation (torsional motion) of the eye around the gaze direction does not change the gaze direction. Therefore, for gaze tracking purposes, finding the spherical rotation from the reference image to the input image is theoretically a two-degree-of-freedom problem. However, if the rotation is not restricted to having no torsional component, the non-torsional component of the spherical rotation can be found more precisely. Therefore, in some embodiments, a full three-degree-of-freedom spherical rotation is used to align the reference image with the input image. A full three-degree-of-freedom spherical rotation can also be used when stitching retinal images to form a panorama as described above.
[0163] exist Figure 5A In the exemplary embodiment schematically shown, a processor, such as processor 102, performs a method for determining a person's gaze direction, the method comprising receiving an input image including at least a portion of the retina of a person in an unknown gaze direction (step 50). The processor then finds a spatial transformation that best matches a reference image to the input image (step 52), the reference image including at least a portion of the retina of a person in a known gaze direction. The spatial transformation is then converted into a corresponding spherical rotation (step 54), and the spherical rotation is applied to the known gaze direction (step 56) to determine a gaze direction associated with the input image (step 58). Optionally, the spherical rotation may be applied to one or more of the following associated with the input image: the orientation of the person's eyes, the gaze direction, the gaze target, and the gaze line.
[0164] A signal can be output based on the person's gaze direction associated with the input image (step 59). The signal can be used to control the device, as described above.
[0165] This embodiment can be illustrated by the following equation:
[0166] R = f(T)
[0167] g i =Rg r
[0168] Where T is the spatial transformation between images, R is the 3D spherical rotation between eye orientations, f is a function, and g r It is the gaze direction associated with the reference image, and g i It is the gaze direction associated with the input image.
[0169] For example, T can be a rigid planar transformation or a 3D spherical rotation, in which case f(T) = T. An example of T being a 3D spherical rotation is as follows.
[0170] In this example, reference image I is obtained. r The reference image includes a human eye gazing at a known gaze target (known in the camera's frame of reference).
[0171] The position of a person's pupil center in the camera's reference frame can be estimated from an image using known techniques. The pupil center is used as the estimated origin of the gaze path. The vector connecting the pupil center and the gaze target is normalized. The normalized vector is the reference gaze direction g. r .
[0172] Then obtain the input image I from the human eye. i .
[0173] The coordinates of the camera sensor pixels are projected onto a sphere. For example, in the case of a pinhole camera, the projection is: Where f is the focal length of the camera, and (0,0) is the pixel corresponding to the optical center of the camera. In cameras with lens distortion, the pixel position can be corrected before being projected onto the sphere (e.g., without distortion), (x,y)→(x',y'), as in pinhole cameras. For example, (x',y')=((1+kr 2 )x,(1+kr 2 )y)where r 2 =x 2 +y 2 and k are the first radial distortion coefficients of the camera lens.
[0174] Inverse projection is
[0175] For a 3×3 matrix T representing candidate spherical rotations between images, I r with I i The following comparisons are made: I r Retinal features with pixel coordinates (x,y) and I i The middle pixel coordinate is P -1Compare the retinal features of TP(x,y).
[0176] Among all candidate T, for all relevant (x,y), seek a 3D spherical rotation such that I r The retinal features at (x,y) are similar to I. i In P -1 Maximize the consistency among retinal features at TP(x,y). Represent this maximization rotation as T. max All relevant pixels can be represented, for example, in I. r All (x,y) in the region containing retinal features.
[0177] g i =T max g r .
[0178] This is the calculated direction of a person's gaze, associated with the input image and represented in the camera's frame of reference.
[0179] Find I i The center of the pupil. (Similar to I) i The relevant gaze is emitted from the center of the pupil. i .
[0180] exist Figure 5B In the schematically illustrated embodiment, the input image and the reference image are planar images, and the processor can project the planar input image onto a sphere to obtain a spherical input image, and project the planar reference image onto a sphere to obtain a spherical reference image, and find the spherical rotation between the spherical reference image and the spherical input image.
[0181] In another embodiment, the processor can find a planar rigid transformation between a planar input image and a planar reference image, and convert the planar rigid transformation into a corresponding spherical rotation.
[0182] Reference image 505 is acquired by camera 103. Reference image 505 is planar and includes at least a portion of the retina of the eye in a known gaze direction. As described above, the reference image can be a panoramic view of a larger portion of the retina or a single image of a smaller portion of the retina. The reference image may or may not include the fovea. Reference image 505 is projected onto sphere 500 to create patch 55.
[0183] An input image 515 is acquired by camera 103 to track the gaze of the eye. The input image 515 includes a portion of the retina in an unknown gaze direction. The input image 515 is also projected onto sphere 500 to create a patch 56.
[0184] Projecting an image onto a sphere can be accomplished using, for example, the optics of camera 103 as described above, by representing each pixel in the image as a unit vector on the sphere.
[0185] Then, processor 102 determines the 3D rotation of sphere 500 required to overlap the spherical projection (pattern 56) of input image 515 with the spherical projection (pattern 55) of reference image 505. The determined 3D rotation of the sphere corresponds to the change in eye orientation from when the reference image 505 is acquired to when the input image 515 is acquired. Therefore, the rotation of the eye orientation between when image 505 is acquired and when image 515 is acquired can be determined based on the calculated rotation of the sphere.
[0186] The gaze direction is determined by the fovea and the center of the eye's lens, both of which are part of the eye and rotate with it; therefore, the gaze line rotates in the same way as the eye. Since the gaze direction of reference image 505 is known, the gaze direction of the input image can be determined based on the rotation of the eye.
[0187] Using 3D rotation of a sphere to search for a match between a reference image and an input image helps achieve a highly accurate overlap between the two images.
[0188] Furthermore, while the reference and input images may include the same area of the retina, this does not mean that the eye is gazing at the same target in both cases. If the area of the retina appearing in the reference and input images is the same but rotated relative to each other, the fovea will be in a different position in each case (unless the image rotation is around the fovea), and therefore the gaze direction will be different. Additionally, if the images are the same but appear at different positions on the image sensor (i.e., the pupil is in a different position relative to the camera), the theoretical position of the fovea will also be different, and therefore the gaze direction will be different. Therefore, according to embodiments of the invention, 3D rotation may be used to determine how spatial transformations of the image affect eye orientation, which is important for accurate and reliable gaze tracking using retinal images.
[0189] Comparing the input image with reference images can be time-consuming if all possible rigid transformations or all 3D spherical rotations are applied to the input image and compared with each reference image. Embodiments of the present invention provide a more time-efficient process for comparing reference and input images.
[0190] In the first step, a search is typically conducted by translating and rotating the input image to find a match between it and multiple planar reference images, and possible matching reference images are identified. In the second step, a match is searched between the spherical projection of the input image and the spherical projection of the reference images identified in the first step. Multiple iterations can be performed for each step.
[0191] In another embodiment, a sparse description of the retinal image is used for comparisons between images. Alternatively or additionally, rotation- and translation-invariant features (e.g., distances between unique retinal features) can be used to compare the input image with a reference image.
[0192] To find a match between an input image and a reference image, it might be necessary to compare the input image with many rigid transformations of the reference image. Each comparison might require going over every pixel in the input image. A typical retinal image contains thousands or millions of pixels. Therefore, a sparse description of an image containing only tens or hundreds of descriptors will take far less time to compare. Additionally, as mentioned above, using a sparse description of an image is useful for reducing false positives when searching for similar images.
[0193] Alternatively or additionally, comparing only rotation- and translation-invariant features eliminates the need for repeated comparisons involving numerous rotations and translations. These embodiments can be used to supplement or replace comparisons of the entire high-resolution retinal image. These embodiments can be used for an initial fast search in a large library of reference images. The best match from the initial search can be selected as a candidate. Later, a slower but more accurate method can be used to compare the candidate with the input image to verify the match between the input and reference images and to compute the spatial transformations (e.g., translations and rotations) between them.
[0194] These embodiments provide faster searching of planar images and enable efficient searching of large image libraries.
[0195] When calculating more precise transformations between retinal images (e.g., as...) Figure 5A and 5B (as described above), similar methods can be used to provide faster searches for 3D spherical rotations.
[0196] Embodiments of the present invention provide high-precision gaze direction determination, such as... Figure 6 The illustration is shown in the figure.
[0197] Processor 102 receives an input image from camera 103 (step 602) and compares the input image with a plurality of previously obtained reference images to find a reference image that at least partially overlaps with the input image (step 604). Processor 102 then determines a plurality of previously obtained reference images that are best correlated with the input image (step 606) and calculates the foveal position and / or gaze direction and / or orientation change of the eye corresponding to the input image based on the best correlated reference image (608). The mean, median, average, or other appropriate statistics of the gaze direction and / or orientation change calculated based on each reference image can be used as a more accurate approximation of the foveal position and / or gaze direction and / or orientation change. Therefore, statistics of multiple foveal position and / or gaze direction and / or orientation changes are calculated (step 610), and the person's gaze direction can be determined based on said statistics (step 612).
[0198] In another embodiment, processor 102 receives an input image from camera 103 and identifies multiple spatial transformations between the input image and multiple reference images. Processor 102 calculates statistics (e.g., mean, median, average, or other suitable statistics) of the multiple spatial transformations and determines, based on the statistics, one or more of the following associated with the input image: the orientation of the person's eyes, the direction of gaze, the object being gazed at, and the direction of gaze along the line of sight.
[0199] In some embodiments, gaze direction is determined based on input from a person's eyes. In one embodiment, a system for determining gaze direction may include two cameras, one for each eye. Alternatively, a single camera may acquire images of both eyes, for example, by being positioned far enough so that the eyes are within their FoV or by using a mirror.
[0200] In some embodiments, two or more cameras are deployed to capture images of the same eye. In one embodiment, two or more cameras are used to capture images of one eye, while no camera is used for the other eye. In another embodiment, two or more cameras are used for one eye, and one or more cameras are used to obtain images of the other eye.
[0201] In one embodiment, at least one camera is configured to capture an image of a person's left eye, and at least one camera is configured to capture an image of a person's right eye. A processor receiving images from the cameras calculates the orientation changes of the left and right eyes. The processor may combine the orientation changes of the left and right eyes to obtain a combined value, and may then calculate one or more of the following based on the combined value: the orientation of the person's eyes, the direction of the person's gaze, the gaze target, and the gaze line.
[0202] In one embodiment, the processor can associate the gaze line with each of the left and right eyes based on the orientation changes of the left and right eyes. The processor can then determine the spatial intersection of the gaze line associated with the left eye and the gaze line associated with the right eye. The distance of the intersection point from the human eye can be calculated, for example, to determine the distance / depth of the gaze target.
[0203] exist Figure 7 In one embodiment illustrated schematically, a method for determining gaze direction includes obtaining simultaneous input images of each of a person's two eyes (702), and calculating, for example, the orientation change of each eye separately using the method described above (704). A combined value (e.g., an average) of the orientation changes from both eyes is calculated (706), and the change in the orientation and / or gaze direction of the person's eyes is determined based on the combined value (708).
[0204] Therefore, for example, a method for determining a person's gaze direction based on their two eyes includes: obtaining a reference image for each eye, the reference image including at least a portion of the retina of the eye in a known gaze direction; and receiving an input image for each of the person's two eyes, the input image including at least a portion of the retina of the eye in an unknown gaze direction. The method further includes finding a spatial transformation of the input image relative to the reference image, and determining the gaze direction of each of the two eyes based on the transformation. A combined value of the gaze directions of the person's two eyes is then calculated, and the person's gaze direction is determined based on the combined value.
[0205] In some embodiments, a statistical model (e.g., as described above) can be used to predict the gaze direction of a first eye and the gaze direction of a second eye based on spatial transformations between a reference image including at least a portion of the retina of a first eye and an input image including at least a portion of the retina of a first eye, and between a reference image including at least a portion of the retina of a second eye and an input image including at least a portion of the retina of a second eye. A combined value of the first gaze direction and the second gaze direction can be calculated, and the gaze direction of the person can be determined based on the calculated combined value.
[0206] In another embodiment, a single statistical model is used to predict a person's gaze direction based on data from input and reference images from a first eye and data from input and reference images from a second eye. For example, a model accepts two spatial transformations as input, one transformation associated with each eye, and outputs a single prediction of the gaze direction. The model's input data may include, for example, the location of the pupil center in the input image for each eye; the location of the pupil center in the reference image for each eye; the gaze target in the reference image for each eye; a rigid transformation between the input and reference images for each eye, etc.
[0207] In some embodiments, a statistical model can be used to predict the gaze direction of each eye based on the spatial transformation of the input image relative to a reference image. A combined value of the gaze direction for each eye can then be calculated, and the person's gaze direction can be determined based on the combined value.
[0208] In another embodiment, the gaze direction is determined using a physical model of the head, including both eyes, with certain constraints. These constraints may include a fixed position of each eye relative to the head, both eyes gazing at the same point, and the eyes maintaining a fixed roll (i.e., they do not rotate around the gaze direction). Input parameters (such as pupil position) and a relative rigid transformation between the input image and a reference image are used as input to a probabilistic model that computes the most probable values of the physical model's parameters given the inputs and constraints.
Claims
1. A method for eye tracking, comprising: receiving an input image comprising at least a portion of a person's retina; finding a spatial transformation between the input image and a reference image, the reference image comprising at least a portion of the person's retina, the reference image being associated with a known gaze direction of the person; converting the transformation to a corresponding three degrees of freedom spherical rotation; using the spherical rotation to rotate the known gaze direction associated with the reference image to compute an orientation change of the person's eye corresponding to the spatial transformation; and outputting a signal based on the orientation change.
2. The method for eye tracking of claim 1, wherein the step of finding a spatial transformation between the input image and a reference image comprises: trying a plurality of spatial transformations between the input image and a reference image; computing a similarity measure between the input image and reference image corresponding to each spatial transformation; and selecting a spatial transformation based on the similarity measure.
3. The method for eye tracking of claim 1, wherein the input image and the reference image are planar images, the method comprising: projecting a planar input image to a sphere to obtain a spherical input image; projecting a planar reference image to a sphere to obtain a spherical reference image; and finding a spherical rotation between the spherical reference image and the spherical input image.
4. The method for eye tracking of claim 1, comprising: finding a planar rigid transformation between a planar input image and a planar reference image; and converting the planar rigid transformation to a corresponding spherical rotation.
5. The method for eye tracking of claim 1, comprising: finding a plurality of spatial transformations between the input image and a plurality of reference images; computing a statistic of the plurality of spatial transformations; and determining one or more of the following associated with the input image based on the statistic: an orientation of the person's eye, a gaze direction, a gaze target, and a gaze direction of a gaze line.
6. The method for eye tracking of claim 1, for matching an input image with a reference image to identify a user.
7. A system for eye tracking, comprising: at least one retinal camera, the retinal camera comprising: an image sensor for capturing an image of a person's eye; and a camera lens configured to focus light originating from the person's retina on the image sensor; a light source that generates light emanating from a location proximate to the retinal camera; and a processor for computing an orientation change of the person's eye between a first image and a second image captured by the image sensor, thereby computing a gaze direction of the person in the first image, by: finding a spatial transformation between the first image and the second image; converting the transformation to a corresponding three degrees of freedom spherical rotation; using the three degrees of freedom spherical rotation to rotate the known gaze direction associated with the second image. 8. The system for eye tracking of claim 7, wherein the processor is to calculate one or more of: an orientation of the human eye, a gaze direction of the human, a gaze target, and a gaze line based on the change in the orientation of the human eye between the first image and the second image.
9. The system for eye tracking of claim 7, wherein the light source is a polarized light source, the system further comprising a beam splitter configured to direct light from the light source to the human eye and a polarizer configured to block light originating from the light source and reflected by specular reflection.
10. The system for eye tracking of claim 7, wherein the light source comprises at least one infrared LED.
Citation Information
Patent Citations
Methods and apparatus for retinal retroreflection imaging
US10248194B2
Methods and Apparatus for Retinal Retroreflection Imaging
US20160320837A1
System And Method For Real-Time Eye Tracking For A Scanning Laser Ophthalmoscope
US20170188822A1
Method and device for measuring the position of an eye
US20170007446A1
Real time eye tracking for human computer interaction
US8885882B1