An eye tracking system based on hand-eye interaction and its calibration method
Through the hand-eye interaction eye tracking system, an infrared dual-camera module and multiple light sources are used to detect hand-eye interaction events and implicitly calibrate eye tracking. This solves the problems of poor usability and low accuracy caused by complex calibration in existing technologies and achieves high-precision eye tracking.
Patent Information
- Application Number
- CN202310973905.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-03
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-08-03
AI Technical Summary
Existing eye tracking devices require a complex calibration process, making them difficult to use and having low accuracy, especially for children or people with mental illnesses.
An eye tracking system based on hand-eye interaction is adopted, which uses an infrared dual-camera module and multiple infrared light sources in combination with a scene camera. By detecting hand-eye interaction events such as fingertip clicks, implicit calibration is achieved to solve the angle between the optical axis and the visual axis of the eyeball, avoiding the display of calibration points and special tools.
Implicit calibration is achieved without the need for subjective cooperation from the user, which improves recognition accuracy and meets usability requirements.
Smart Images

Figure CN116935491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of eye movement instruments and meters, and in particular to an eye movement tracking system based on hand-eye interaction and a calibration method thereof. Background Art
[0002] Eye tracking devices are instruments used to study human visual behavior, such as where the eye fixates, how long it remains fixed, how quickly it scans, how quickly it moves, and other visual parameters, as well as how people respond to different stimuli. They are currently widely used in fields such as psychology, neuroscience, human-computer interaction, and marketing, helping to study human cognition and behavior and improve the effectiveness of product design and marketing strategies.
[0003] Currently, there are three main methods for eye tracking technology: 2D polynomial regression-based method, 3D eye model-based method, and machine learning-based method.
[0004] The 2D polynomial regression-based method establishes a two-dimensional mapping relationship between pupil-corneal vector features and the gazed plane. A commonly used method is quadratic polynomial fitting, which generally requires a nine-point calibration. This means the user must sequentially gaze at nine calibration points on the screen before use, and recalibrate each time. Another calibration method uses smooth tracking, which obtains a set of calibration data by having the user gaze at a calibration point that moves along a specific trajectory. This calibration data is then used to establish a mapping relationship between eye features and gaze points, ultimately achieving eye tracking.
[0005] Methods based on 3D eye models use the optical properties of the cornea and other parameters of the human eye to determine the direction of the eye's gaze, thereby estimating the real-world location of the gaze point. Since the direction of a person's gaze is determined by the visual axis, methods based on 3D eye models can only determine the direction of the eye's optical axis. There is an angle between the optical axis and the visual axis, called the kappa angle. Different people have different kappa angles, so this method also requires a user calibration process.
[0006] Machine learning-based methods use neural networks, deep learning, and other methods to match features such as pupil parameters, pupil-corneal vectors, and the coordinates of the reflected light spot with the gaze direction. This method does not require user calibration, but it does require a large amount of training data and has relatively low accuracy, generally around 5°. Methods based on 2D polynomial regression and 3D eye models generally achieve around 1° after calibration. Furthermore, machine learning-based methods may suffer from overfitting, which can lead to low accuracy in gaze estimation.
[0007] In summary, current eye-tracking devices require a complex calibration process to achieve high recognition accuracy. Even calibration tasks involving fixating a single point can interfere with natural interaction. A key application area for eye tracking is medicine, such as research on infant cognitive behavior, depression, and other mental illnesses. The majority of users in these studies are children or individuals with mental illnesses, who struggle to accurately complete the calibration process and meet usability requirements. Therefore, developing an implicitly calibrated eye interaction technology that does not display calibration points, requires no dedicated calibration tools, and requires no user interaction is crucial for practical application. Summary of the Invention
[0008] The purpose of the present invention is to provide an eye tracking system based on hand-eye interaction and a calibration method thereof, so as to solve the problems of poor usability and low accuracy caused by the need to display calibration points, use special calibration tools or require subjective cooperation of the user in the above-mentioned prior art.
[0009] In one aspect, the present invention provides an eye tracking system based on hand-eye interaction, comprising a display, a first light source, a second light source, a third light source, a fourth light source, an infrared dual-camera module, a scene camera, a glasses bracket, and a controller, wherein the infrared dual-camera module consists of a first camera and a second camera, and the infrared dual-camera module is installed directly below the display or on the glasses bracket; the first light source is located in the middle of the right side of the display, the second light source is located directly above the display, the third light source is located in the middle of the left side of the display, and the fourth light source is located directly below the display; the scene camera is set on the glasses bracket;
[0010] The infrared dual-camera module is used to obtain the user's eye image; the first light source, the second light source, the third light source, and the fourth light source are used to form a reflection point; the scene camera and the glasses bracket constitute a head-mounted scene camera, which is used to collect the field of view image directly in front of the user; the controller includes an eye feature extraction module, a fingertip extraction module, an implicit calibration module, and a gaze space geometry module, wherein the eye feature extraction module is used to segment the eye movement image and convert it into a grayscale image, and process the converted image to extract the pupil center position and the reflection point position of the infrared light source in the eye at each moment; the fingertip extraction module is used to extract the fingertip coordinates in the user's field of view when a hand-eye interaction event occurs; the implicit calibration module is used to use the pupil center coordinates extracted by the eye feature extraction module and the fingertip coordinates provided by the fingertip extraction module to solve the optical axis direction of the initial position; the gaze space geometry module is used to solve the gaze point coordinates at the current moment according to the optical axis direction of the initial position output by the implicit calibration module.
[0011] Furthermore, the first light source, the second light source, the third light source, and the fourth light source are all infrared light sources.
[0012] Furthermore, the scene camera is connected to the display via a data cable or wirelessly.
[0013] On the other hand, the present invention provides an eye movement calibration method based on hand-eye interaction. The method is based on the eye movement calibration system based on hand-eye interaction of the present invention, and specifically includes the following steps:
[0014] Step S1, collecting eye images, segmenting the eye movement images and converting them into grayscale images, and processing the converted images to extract the pupil center position and the position of the infrared light source's reflection point on the eye at each moment;
[0015] Step S2, extracting the coordinates of the fingertip in the user's field of view;
[0016] Step S3, using the pupil center coordinates and the reflection point coordinates extracted in step S1 and the fingertip coordinates extracted in step S2, to solve the optical axis direction of the initial position;
[0017] Step S4: Calculate the current gaze point coordinates based on the optical axis direction of the initial position.
[0018] Furthermore, the step S1 specifically includes the following steps:
[0019] Step S10: The user puts on the glasses holder;
[0020] Step S11: The infrared dual-camera module is turned on, and the second camera and the first camera capture images to obtain eye movement images;
[0021] Step S12, segmenting the current frame image obtained in step S11, extracting the images captured by the second camera and the first camera respectively, defining the image captured by the second camera as the right image, and defining the image captured by the first camera as the left image; converting both the left image and the right image into grayscale images; performing eye region detection on the grayscale images corresponding to the right image and the left image respectively, to obtain an eye region image corresponding to the right image and an eye region image corresponding to the left image;
[0022] Step S13, performing threshold segmentation on the eye region image corresponding to the right image and the eye region image corresponding to the left image obtained in step S12, respectively, to obtain a binarized image corresponding to the right image and a binarized image corresponding to the left image;
[0023] Step S14: Obtain the left eye pupil coordinates and the coordinates of the two corresponding reflection points from the binarized image corresponding to the right image obtained in step S13; obtain the left eye pupil coordinates and the coordinates of the two corresponding reflection points from the binarized image corresponding to the left image; similarly, obtain the right eye pupil coordinates and the coordinates of the two corresponding reflection points corresponding to the left and right images;
[0024] Step S15: According to the result of step S14, the three-dimensional coordinates of the left eye pupil and the two reflection points corresponding to the current moment in the camera coordinate system are obtained.
[0025] Furthermore, the step S2 specifically includes the following steps:
[0026] Step S20, the scene camera is turned on and a scene camera image is captured;
[0027] Step S21, determining whether a hand-eye interaction event occurs, wherein the hand-eye interaction event refers to the user clicking the start button with their finger; if the hand-eye interaction event occurs, saving the image captured by the field of view camera when the event occurs; otherwise, continuing to wait for the hand-eye interaction event to occur;
[0028] Step S22: Perform fingertip detection on the stored image to obtain fingertip coordinates.
[0029] Furthermore, in step S22, the fingertip detection method is to perform fingertip detection through a deep learning method, or to obtain the fingertip position using a skin color segmentation method: first, the collected RGB image is converted into an HSV format, and the hand position is obtained using a skin color model in the H, S, and V spaces, and then binarized, and the center of mass of the binarized hand area is calculated, and the point farthest from the center of mass is used as the fingertip coordinate.
[0030] Furthermore, the step S3 specifically includes the following operations:
[0031] (1) Solve the current visual axis direction:
[0032] A three-dimensional coordinate system is established with the second camera 131 as the coordinate origin. The following four plane equations are established to solve the coordinates of the center of corneal curvature C. The direction determined by the line connecting the coordinates of the center of corneal curvature C and the current pupil center coordinate P is the current optical axis direction OD.
[0033] (L1-O1)×(u 11 -O1)·(C-O1)=0
[0034] (L2-O1)×(u 21 -O1)·(C-O1)=0
[0035] (L1-O2)×(u 12 -O2)·(C-O2)=0
[0036] (L2-O2)×(u 22 -O2)·(C-O2)=0
[0037] Where:
[0038] L1 and L2—any two light sources among the four light sources;
[0039] O1—the coordinates of the optical center of the first camera 132;
[0040] O2—the coordinates of the optical center of the second camera 131;
[0041] u 12 、u 11 、u 22 、u 21 —The coordinates of the images formed by the four reflection points on the camera imaging plane;
[0042] C—center of corneal curvature;
[0043] (2) Solve the optical axis OB at the initial position:
[0044] The rotation axis OL is obtained by cross-producting the visual axis OA at the initial position and the visual axis OC at the current moment. The dot product of OA and OC is divided by the product of the modulus of the visual axis OA at the initial position and the visual axis OC at the current moment to obtain β. The optical axis OD at the current moment is rotated around the rotation axis OL by an angle of -β to obtain the optical axis OB at the initial position.
[0045] Furthermore, the step S4 specifically includes the following steps:
[0046] Step 41, cross product the visual axis OA at the initial position with (OB-OD) to obtain the rotation axis OL at the current moment; then project the initial optical axis position OB and the optical axis position OD at the current moment onto the rotation axis OL at the current moment to obtain OG; the angle between GD and GB is the angle β corresponding to the rotation; Step 42, rotate the visual axis OA at the initial position around the rotation axis OL at the current moment by the corresponding angle β to obtain the visual axis OC at the current moment, and the intersection of the visual axis OC at the current moment and the display screen is the gaze point of the corresponding eye.
[0047] Compared with the prior art, the present invention has the following technical effects:
[0048] 1. Only a single calibration point is needed to determine the angle between the visual axis and the optical axis. This is because the system consisting of a dual-infrared camera module and an infrared light source can determine the direction of the optical axis of the human eye in a 3D eyeball model. This, combined with the geometric relationship between the visual axis and the optical axis, enables implicit calibration. This implicit calibration also improves recognition accuracy.
[0049] 2. Implicit calibration is achieved without displaying calibration points, using dedicated calibration tools, or requiring user interaction. By detecting hand-eye interaction events, such as when a fingertip clicks the start button, the user's gaze falls on the fingertip. By detecting the fingertip position, the required calibration points can be obtained, completing implicit calibration. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the eye tracking system of the present invention.
[0051] Figure 2 It is a software structure diagram of the present invention.
[0052] Figure 3 It is a flow chart of the eye movement calibration method based on hand-eye interaction of the present invention.
[0053] Figure 4 This is the dual-camera dual-light source geometric model of the present invention.
[0054] Figure 5 It is the geometric relationship diagram of the visual axis and the optical axis.
[0055] Figure 6 This is a schematic diagram of optical axis direction detection.
[0056] Figure 7 This is a schematic diagram of fingertip detection. DETAILED DESCRIPTION
[0057] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and specific embodiments. The following examples or drawings are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0058] like Figure 1 As shown, the eye tracking system based on hand-eye interaction provided by the present invention includes a display 110, a first light source 120, a second light source 121, a third light source 122, a fourth light source 123, an infrared dual-camera module 130, a scene camera 160, a glasses support 170, and a controller. The infrared dual-camera module 130 is composed of a first camera 131 and a second camera 132, and is mounted directly below the display 110 or on the glasses support 170. The first light source 120 is located in the middle of the right side of the display 110, the second light source 121 is located directly above the display 110, the third light source 122 is located in the middle of the left side of the display 110, and the fourth light source 123 is located directly below the display 110. The scene camera 160 is mounted on the glasses support 170.
[0059] The display 110 is used to display the test video, image and gaze point;
[0060] The infrared dual-camera module 130 is used to obtain user eye images;
[0061] The first light source 120, the second light source 121, the third light source 122, and the fourth light source 123 are all infrared light sources, used to form reflection points;
[0062] The scene camera 160 and the glasses bracket 170 form a head-mounted scene camera for capturing images of the user's field of view directly in front of the user.
[0063] Preferably, the display 110 is supported by a display bracket 140 .
[0064] Preferably, the scene camera 160 is connected to the display 110 via a data cable 150 or a wireless connection.
[0065] Preferably, the glasses bracket 170 is Figure 1 The style shown in .
[0066] like Figure 2 As shown, the controller includes an eye feature extraction module 220 , a fingertip extraction module 240 , an implicit calibration module 250 , and a gaze space geometry module 260 .
[0067] Eye feature extraction module 220, for segmenting the eye motion image and converting it into a grayscale image, and processing the converted image to extract the pupil center position and the position of the infrared light source's reflection point on the eye at each moment;
[0068] The fingertip extraction module 240 is used to extract the coordinates of the fingertips in the user's field of view when a hand-eye interaction event occurs;
[0069] The implicit calibration module 250 uses the pupil center coordinates extracted by the eye feature extraction module 220 and the fingertip coordinates provided by the fingertip extraction module to solve the optical axis direction of the initial position.
[0070] The gaze space geometry module 260 is used to solve the gaze point coordinates at the current moment according to the optical axis direction of the initial position output by the implicit calibration module 250.
[0071] refer to Figure 3 , which shows a flow chart of eye tracking calibration based on hand-eye interaction provided by an embodiment of the present invention. The method of this embodiment is based on Figure 1 The eye tracking system based on hand-eye interaction shown in the figure specifically includes the following steps:
[0072] Step S1, extracting eye features, specifically includes the following steps:
[0073] In step S10 , the user puts on the glasses frame 170 .
[0074] In step S11 , the infrared dual-camera module 130 is turned on, and the second camera 131 and the first camera 132 capture images to obtain eye movement images (video streams).
[0075] Step S12, segmenting the current frame image obtained in step S11, extracting the images captured by the second camera 131 and the first camera 132 respectively, defining the image captured by the second camera 131 as the right image, and defining the image captured by the first camera 132 as the left image; converting both the left image and the right image into grayscale images; performing eye area detection on the grayscale images corresponding to the right image and the left image respectively, to obtain an eye area image corresponding to the right image and an eye area image corresponding to the left image.
[0076] If adopted Figure 1 In the example shown, the infrared binocular camera 130 is installed below the display 110. The captured image range is relatively wide, and eye area detection is required. The detection method can adopt any of the following methods: (1) Eye area detection is achieved by posting markers on the glasses bracket 170. Since the position of the two eyes of the user is roughly the same relative to the position of the glasses bracket 170 after wearing the glasses bracket 170, the approximate eye area can be determined by detecting the markers. (2) Collect and mark multiple positive samples (pictures containing human eyes) and negative samples (pictures not containing human eyes), train the Haar-cascade classifier, and the trained classifier can be used to obtain the position of both eyes in the image. (3) Use deep learning methods to detect human eyes. For example, use the Yolo model, collect and mark positive samples and negative samples, perform model training, save the model parameters after training, and the position of both eyes in the image can be obtained by the trained classifier. (4) Use dlib face detection to detect 68 feature points of the face, extract the feature points of the eye area, and determine the coordinate range of the eye area.
[0077] Alternatively, the infrared binocular camera 130 can be mounted on the glasses bracket 170 so that the infrared dual-camera module 130 faces the eye area of the person to directly obtain the eye area image.
[0078] Step S13 , performing threshold segmentation on the eye region image corresponding to the right image and the eye region image corresponding to the left image obtained in step S12 , respectively, to obtain a binarized image corresponding to the right image and a binarized image corresponding to the left image.
[0079] Step S14: Obtain the left pupil coordinates from the binary image corresponding to the right image (x PR ,y PR ), the coordinates of the two reflective points are marked as (x GR1 ,y GR1 )、(x GR2 ,y GR2 ); The left pupil coordinates are obtained from the binary image corresponding to the left image and are marked as (x PL ,y PL ), the coordinates of the two reflective points are marked as (x GL1 ,y GL1 )、(xGL2 ,y GL2 ).
[0080] In step S15, based on the result of step S14, the three-dimensional coordinates of the left eye pupil and the two reflective points corresponding to the current moment in the camera coordinate system can be obtained using the binocular camera ranging principle. This operation is a routine operation in this field and will not be repeated here.
[0081] The above steps S13 to S15 are operations for the left eye. Alternatively, the same operations can be performed on the right eye because the left eye and the right eye are equivalent for solving the spatial geometric model. Specifically, taking the left eye in the right image as an example, the specific process of step 14 is as follows: First, a closed operation is performed on the binary image corresponding to the right image obtained in step S13 to obtain a connected domain set, and noise points with less than 50 pixels are removed. The circularity test is performed on the remaining connected domains, and the centroid coordinates and radius R of the connected domain with the highest circularity are calculated. P , take the centroid coordinates as the coordinates of the left eye pupil center (x PR ,y PR ), and (x PR ,y PR ) as the center, twice R P Reflective points are extracted in the reflective point extraction area with a radius of . Specifically, in the process of solving the gaze geometry model of the present invention, theoretically only two reflective point coordinates are needed to obtain the direction of the optical axis. However, during the rotation of the eyeball, the reflection point of the light source may exceed the cornea range of the human eye and cannot form a reflective point. Therefore, a total of four light sources are set around the display 110 to ensure that when gazing at different areas, there are always two or more reflective points in the cornea range of the human eye. Then, the two reflective points closest to the pupil center are used to solve the model. Since the brightness of the reflective points is relatively high, the 60 pixels with the highest brightness in the reflective point extraction area are taken (the specific value can be reset according to the different camera resolutions used). Similarly, these 60 pixels are closed to obtain a connected domain set, and the centroid coordinates of each connected domain are obtained. The coordinates (x PR ,y PR )The two nearest coordinates are labeled (x GR1 ,y GR1 )、(x GR2 ,y GR2 ), as the coordinates of the two reflection points. Similarly, use the same method to obtain the left eye pupil center coordinates (x PL ,y PL ) and the coordinates of the two reflection points (x GL1 ,y GL1 )、(x GL2 ,y GL2), and then use the binocular camera ranging principle to obtain the spatial three-dimensional coordinates of the left eye pupil and the two reflecting points corresponding to the current moment.
[0082] Step 2: Extract fingertip coordinates, specifically including the following steps:
[0083] In step S20 , the scene camera 160 is turned on to capture scene camera images.
[0084] Step S21 determines whether a hand-eye interaction event has occurred. A hand-eye interaction event, as defined herein, refers to a user's finger tapping a "Start" button, or other fingertip tapping action strongly associated with gaze. The "Start" button can be a physical button on display 110 or a virtual button displayed on the screen of display 110. If a hand-eye interaction event occurs, the image captured by field of view camera 160 at the time of the event is saved. Otherwise, the system continues waiting for the hand-eye interaction event to occur.
[0085] Step S22: perform fingertip detection on the saved image to obtain the fingertip coordinates. The fingertip detection method can be implemented in the following two ways: (1) Use skin color segmentation to obtain the fingertip position. First, convert the collected RGB image into HSV format, use the skin color model of H (hue), S (saturation), V (brightness) space to obtain the hand position, then binarize it, calculate the centroid of the binarized hand area, and take the point farthest from the centroid as the fingertip coordinate (x FT ,y FT (2) Fingertip detection is performed using deep learning methods. For example, using the YOLO model, positive samples (pictures containing fingertips) and negative samples (pictures not containing fingertips) are collected and marked, and the model is trained. After training, the model parameters are saved, and the fingertip position coordinates of the image can be obtained through the trained classifier.
[0086] Step S3, implicit calibration, specifically includes the following steps:
[0087] refer to Figure 4 Dual-camera dual-light source geometric model, O1 is the optical center of the first camera 132, O2 is the optical center of the second camera 131, L1 and L2 are two of the four infrared light sources. Because the two cameras are in different positions, the reflection points of the two infrared light sources captured by each camera on the human cornea will be different, so there will be four reflection points q 11 ,q 12 ,q 21 ,q 22 The corresponding image of these four reflection points on the camera imaging plane is u 12 、u 11 、u 22 、u 21. The images of the pupil center P formed on the two camera imaging planes are P1 and P2. O is the center of the eyeball, and C is the center of corneal curvature. When the position of the light source is known, the actual three-dimensional coordinates of the center of corneal curvature C and the center of pupil P can be solved through the image of the human eye collected by the infrared dual-camera module 130. The direction determined by CP is the direction of the optical axis. In the eyeball model, the actual gaze point of a person is the intersection of the visual axis direction and the screen, and the visual axis is generally defined as the line between the center of corneal curvature and the fovea on the retina. The fovea is not on the optical axis, but deviates from the optical axis a little, that is, there is a certain angle between the optical axis and the visual axis. This angle is called the kappa angle. The size of the kappa angle varies from person to person, and it cannot be directly solved from the image collected by the infrared dual-camera module 130, so it needs to be calibrated. The calibrated kappa angle is the solved kappa angle. The specific process is as follows:
[0088] (1) Solve the visual axis direction at the current moment
[0089] According to the reflection theorem of light, it can be determined that the light source, the optical center of the camera, the reflection point and the center of corneal curvature are on the same plane. Then, four plane equations are established based on the four reflection points on the human cornea:
[0090] (L1-O1)×(u 11 -O1)·(C-O1)=0
[0091] (L2-O1)×(u 21 -O1)·(C-O1)=0
[0092] (L1-O2)×(u 12 -O2)·(C-O2)=0
[0093] (L2-O2)×(u 22 -O2)·(C-O2)=0
[0094] With the second camera 131 as the coordinate origin, a three-dimensional space coordinate system is established. In this coordinate system, the light source coordinates L1 and L2, the optical center coordinates O1 and O2, and the image u formed by the four reflection points on the camera imaging plane are 12 、u 11 、u 22 、u 21 The coordinates of the corneal curvature center C are known quantities. Substituting them into the four equations above, the coordinates of the corneal curvature center C are solved. The pupil center coordinate P can be directly solved using the principle of binocular stereo imaging (i.e., the current pupil center coordinate is obtained in step S1). After obtaining the pupil center coordinate P and the corneal curvature center coordinate C, the direction determined by the line CP connecting the two points is the current visual axis direction.
[0095] (2) Solve the initial position optical axis direction
[0096] Listing's law describes the motion of the eyeball from its initial position to any arbitrary position. Specifically, the eyeball can move in two ways: the first is to simultaneously move horizontally (around the vertical axis) and vertically (around the horizontal axis) to reach a specific position; the second is to achieve this position through a single rotation about another axis. These possible single rotation axes form a plane called the Listing plane. Therefore, the direction of the visual axis in any state can be calculated simply by analyzing the relationship between the optical axis and visual axis at the initial position and at other positions.
[0097] Figure 5 Represents the relationship between the optical axis OB and the visual axis OA at the initial position and the optical axis OD and the visual axis OC at the current moment. OL is the rotation axis of the optical axis OB at the initial position and the visual axis OA at the initial position. OL is parallel to the Listing plane and perpendicular to OA and OC.
[0098] Among the four vectors OA, OB, OC, and OD, if any three of them are known, the fourth vector can be calculated. We assume that the user's head position is fixed relative to the display 110. Preferably, the present invention defines the visual axis OA at the initial position as the direction perpendicular to the screen of the display 110, so the visual axis OA at the initial position is a known quantity. When a hand-eye interaction event is detected in step S21, the dual-camera dual-light source geometric model (see Figure 4 ) The direction of the optical axis solved is OD. At the same time, when the hand-eye interaction event occurs, the focus of the human eye is the fingertip (the fingertip coordinates are extracted in step 2), so the visual axis direction OC at the current moment is also known. Since OA, OC, and OD are known, according to Figure 5 As shown in the spatial geometric relationship, the visual axis OA at the initial position is cross-producted with the visual axis OC at the current moment to obtain the rotation axis OL; the dot product of OA and OC is divided by the product of the modulus of the visual axis OA at the initial position and the visual axis OC at the current moment to obtain β; the optical axis OD at the current moment is rotated around the rotation axis OL by an angle of -β to obtain the optical axis OB at the initial position.
[0099] Step 4, solving the gaze point, includes the following operations:
[0100] Step 41: In step S3, we get the optical axis OB and visual axis OA at the initial position, then these two vectors become known quantities. When we look at a random position, the optical axis OD at the current moment is given by Figure 4The dual-lamp dual-light source geometric model shown is obtained in real time (i.e., a frame of image captured by the first camera and the second camera at the current moment is processed by step 1 to obtain the pupil center coordinates of a certain eye, thereby obtaining the CP line and determining the optical axis OD direction at the current moment), so OD is also a known quantity.
[0101] Step 42: Since OA, OB, and OD are known quantities, according to Figure 5 As shown in the spatial geometric relationship, the rotation axis OL at the current moment is perpendicular to the visual axis OA at the initial position, and is also perpendicular to the difference between the initial optical axis position OB and the optical axis position OD at the current moment. Therefore, the visual axis OA at the initial position and (OB-OD) are cross-producted here to obtain the rotation axis OL at the current moment; then the initial optical axis position OB and the optical axis position OD at the current moment are projected onto the rotation axis OL at the current moment to obtain OG; the angle between GD and GB is the angle β corresponding to the rotation; the visual axis OA at the initial position is rotated around the rotation axis OL at the current moment by the corresponding angle β to obtain the visual axis OC at the current moment, and the intersection of the visual axis OC at the current moment and the display screen is the gaze point of the corresponding eye.
[0102] In summary, the hand-eye interaction-based eye tracking device provided by the present invention first obtains the direction of the eye's optical axis through a dual-camera, dual-light source system. It then uses a scene camera to obtain the fingertip coordinates at the time of the hand-eye interaction event. Finally, the relationship between the optical axis and the visual axis at the random gaze position and the initial position, as described in Listing's law, is used to determine the visual axis direction, thereby obtaining the gaze point coordinates and achieving the purpose of line of sight tracking. Compared with existing technologies, this system can be calibrated without the user's awareness, making it highly user-friendly. Furthermore, the present invention improves recognition accuracy through implicit calibration.
[0103] In order to illustrate the feasibility and effectiveness of the present invention, the specific test process is introduced below:
[0104] The present invention sets four light sources in order to improve the stability of the algorithm. In fact, only two infrared light sources are needed to find the optical axis, so the actual configuration of the following example is two infrared light sources.
[0105] In the first step, the system is turned on, the infrared dual-camera module 130 detects the eye area, extracts the pupil center and the coordinates of the reflection point, and calculates the optical axis direction of the eye according to the dual-camera dual-light source model, as shown in the following example: Figure 6 As shown, the red dots are the eye corners extracted by the dlib face detection algorithm, which are used to locate the eye area. The green and blue dots are the reflection points formed by two infrared light sources, and the blue line is the direction of the optical axis.
[0106] The second step is to wait for the occurrence of a hand-eye interaction event, such as pressing the "Start" button on the screen with your finger. When the hand-eye interaction event is detected, the fingertip coordinates are extracted through the fingertip coordinate extraction module, such as Figure 7 shown.
[0107] The third step is that when hand-eye interaction is detected, the actual gaze point of the human eye can be approximately considered to be the fingertip, so the direction of the visual axis at this time can be obtained, and the direction of the optical axis at this time can also be obtained. Then, the kappa angle can be calculated using the listing law.
[0108] The fourth step is to calculate the kappa angle and combine it with the optical axis direction calculated in real time to calculate the visual axis direction in real time, thereby obtaining the actual gaze point.
Claims
1. An eye tracking system based on hand-eye interaction, characterized in that: The device comprises a display, a first light source, a second light source, a third light source, a fourth light source, an infrared dual-camera module, a scene camera, a glasses bracket, and a controller, wherein the infrared dual-camera module consists of a first camera and a second camera, and the infrared dual-camera module is installed directly below the display or on the glasses bracket; the first light source is located in the middle of the right side of the display, the second light source is located directly above the display, the third light source is located in the middle of the left side of the display, and the fourth light source is located directly below the display; the scene camera is set on the glasses bracket; The infrared dual-camera module is used to obtain the user's eye image; the first light source, the second light source, the third light source, and the fourth light source are used to form a reflection point; the scene camera and the glasses bracket constitute a head-mounted scene camera, which is used to collect the field of view image directly in front of the user; the controller includes an eye feature extraction module, a fingertip extraction module, an implicit calibration module, and a gaze space geometry module, wherein the eye feature extraction module is used to segment the eye movement image and convert it into a grayscale image, and process the converted image to extract the pupil center position and the reflection point position of the infrared light source in the eye at each moment; the fingertip extraction module is used to extract the fingertip coordinates in the user's field of view when a hand-eye interaction event occurs; the implicit calibration module is used to use the pupil center coordinates extracted by the eye feature extraction module and the fingertip coordinates provided by the fingertip extraction module to solve the optical axis direction of the initial position; the gaze space geometry module is used to solve the gaze point coordinates at the current moment according to the optical axis direction of the initial position output by the implicit calibration module; The implicit calibration module specifically implements the following operations: (1) Solve the current visual axis direction: Establish a three-dimensional coordinate system with the second camera as the coordinate origin, establish the following four plane equations, and solve the center of corneal curvature Coordinates of the center of corneal curvature The direction determined by the line between the coordinates of and the pupil center coordinates P at the current moment is the direction of the optical axis at the current moment ; Where: and —Any two of the four light sources; —The coordinates of the optical center of the first camera; —The coordinates of the optical center of the second camera; 、 、 、 —The coordinates of the images formed by the four reflection points on the camera imaging plane; —center of corneal curvature; (2) Determine the optical axis at the initial position : Set the initial position of the visual axis The visual axis with the current moment Do the cross product to get the rotation axis ; After doing the dot product of OA and OC, divide it by the visual axis of the initial position The visual axis with the current moment The product of the modulus lengths is ; Set the current optical axis Around the axis of rotation Rotation angle Get the optical axis at the initial position ; The gaze space geometry module specifically implements the following operations: Step 1: Set the initial position of the visual axis and( - ) to obtain the current rotation axis by cross product ; Then respectively set the initial optical axis position and the current optical axis position Projected onto the current rotation axis Get ; and The angle between them is the angle corresponding to the rotation ; Step 2, the visual axis of the initial position Rotation axis around the current moment The corresponding angle of rotation Get the current viewing axis , the visual axis at the current moment The intersection with the display screen is the gaze point of the corresponding eye.
2. The eye tracking system based on hand-eye interaction according to claim 1, wherein: The first light source, the second light source, the third light source, and the fourth light source are all infrared light sources.
3. The eye tracking system based on hand-eye interaction according to claim 1, wherein: The scene camera is connected to the display via a data cable or wirelessly.
4. An eye movement calibration method based on hand-eye interaction, characterized in that: The eye tracking system based on hand-eye interaction according to any one of claims 1 to 3 specifically comprises the following steps: Step S1, collecting eye images, segmenting the eye movement images and converting them into grayscale images, and processing the converted images to extract the pupil center position and the position of the infrared light source's reflection point on the eye at each moment; Step S2, extracting the coordinates of the fingertip in the user's field of view; Step S3, using the pupil center coordinates and the reflection point coordinates extracted in step S1 and the fingertip coordinates extracted in step S2, to solve the optical axis direction of the initial position; Step S4: Calculate the current gaze point coordinates based on the optical axis direction of the initial position.
5. The eye movement calibration method based on hand-eye interaction according to claim 4, characterized in that: The step S1 specifically includes the following steps: Step S10: The user puts on the glasses holder; Step S11: The infrared dual-camera module is turned on, and the second camera and the first camera capture images to obtain eye movement images; Step S12, segmenting the current frame image obtained in step S11, extracting the images captured by the second camera and the first camera respectively, defining the image captured by the second camera as the right image, and defining the image captured by the first camera as the left image; and converting both the left image and the right image into grayscale images; Performing eye region detection on the grayscale images corresponding to the right image and the left image respectively to obtain an eye region image corresponding to the right image and an eye region image corresponding to the left image; Step S13, performing threshold segmentation on the eye region image corresponding to the right image and the eye region image corresponding to the left image obtained in step S12, respectively, to obtain a binarized image corresponding to the right image and a binarized image corresponding to the left image; Step S14, obtaining the left eye pupil coordinates and the corresponding two reflection point coordinates from the binarized image corresponding to the right image obtained in step S13; obtaining the left eye pupil coordinates and the corresponding two reflection point coordinates from the binarized image corresponding to the left image; Similarly, obtain the right eye pupil coordinates and the corresponding two reflection point coordinates corresponding to the left and right images; Step S15: According to the result of step S14, the three-dimensional coordinates of the left eye pupil and the two reflection points corresponding to the current moment in the camera coordinate system are obtained.
6. The eye movement calibration method based on hand-eye interaction according to claim 4, characterized in that: The step S2 specifically includes the following steps: Step S20, the scene camera is turned on and a scene camera image is captured; Step S21, determining whether a hand-eye interaction event occurs, wherein the hand-eye interaction event refers to the user clicking the start button with their finger; If a hand-eye interaction event occurs, save the image captured by the field of view camera when the event occurs. Otherwise, continue to wait for the hand-eye interaction event to occur. Step S22: Perform fingertip detection on the stored image to obtain fingertip coordinates.
7. The eye movement calibration method based on hand-eye interaction according to claim 4, characterized in that: In step S22, the fingertip detection method is to perform fingertip detection through a deep learning method, or to obtain the fingertip position using a skin color segmentation method: first, the collected RGB image is converted into an HSV format, and the hand position is obtained using a skin color model in the H, S, and V spaces, and then binarized, and the center of mass of the binarized hand area is calculated, and the point farthest from the center of mass is used as the fingertip coordinate.
Citation Information
Patent Citations
Eye movement and facial expression normal form-based student learning state evaluation system and method
CN113486744A
Method of determining gaze direction
WO2021125993A1