An eye movement tracking method based on coupled cascade regression
Through the coupled cascade regression model combined with eye image features, simultaneous detection of eye status, line of sight estimation and pupil center is achieved, which solves the problem of low accuracy of multi-task eye tracking in the prior art, and is real-time and robust.
Patent Information
- Application Number
- CN202210504365.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-10
AI Technical Summary
Existing eye tracking technologies cannot detect eye status, line of sight estimation and detection of pupil center at the same time, and are affected by light conditions, differences in individual physiological characteristics and changes in head posture, so the accuracy is not high.
Using a method based on coupled cascade regression, the eye state, line of sight direction and pupil center position are estimated through three regression models, and iterative optimization is performed based on the local and global shape characteristics of the eye image.
Real-time and accurate multi-task eye tracking is achieved, the stability and adaptability of detection is improved, the impact of lighting and posture changes is overcome, and the real-time requirements are met.
Smart Images

Figure CN114973389B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent vision technology, and more specifically, to an eye movement tracking method based on coupled cascade regression. Background Art
[0002] Eye movement tracking refers to the process of automatically detecting the position of the pupil center or identifying the three-dimensional line-of-sight direction and fixation point. This technology can predict the intentions and attention of the measured object, and evaluate its state and needs, and has been widely used in fields such as human-computer interaction, intelligent driving, emotion computing, virtual reality, etc. In intelligent driving, the mental state and fatigue level of the driver can be evaluated through eye movement tracking; in emotion computing, the depressive state or disease degree of the patient can be evaluated through the eye movement condition of the patient; in the applications of virtual reality and augmented reality, the eye movement condition of the measured person provides the direction and intention of their attention, which is helpful for research and development in aspects such as image rendering and scene design. However, in actual application scenarios, there are still many limitations in eye movement tracking estimation. On the one hand, factors such as individual physiological characteristics, variable head postures, and lighting conditions affect the accuracy of eye movement tracking; on the other hand, it is very difficult and time-consuming to collect a large number of samples with different appearances, lighting conditions, and head postures.
[0003] Eye movement tracking technology has developed for many years and is mainly divided into two categories:
[0004] 1. Contact eye movement tracking method. This method mainly uses invasive hardware for assistance. For example, the scleral search coil method uses a contact lens with a coil moving in an electromagnetic field, and measures the horizontal and vertical movements of the eyeball through the signal generated by the principle of electromagnetic induction, so as to realize eye position detection; the infrared method requires installing an infrared photosensitive tube near the eye, and can measure eye movement according to the images reflected by different optical interfaces such as the pupil, sclera, and cornea. Because of the low accuracy and the invasive device will cause discomfort to the measured person, these methods have been gradually eliminated;
[0005] 2. Image-based eye movement tracking technology. Extract texture features from the input facial image, and then learn the mapping model between the features and the line-of-sight direction to perform line-of-sight estimation or eye position localization. However, such methods are affected by factors such as changes in lighting conditions, individual physiological characteristics, and head posture changes, with low accuracy and sensitivity. And the existing image-based eye movement tracking technology cannot meet the multi-task requirements of simultaneously detecting the state of the human eye, estimating the line-of-sight direction, and detecting the pupil center.
[0006] A patent for an eye movement tracking, recognition method and system based on the Galaxy Ruihua mobile operating system is disclosed in the prior art. First, the face detection module detects a face from the continuous images captured by the YROS platform camera for subsequent module use. Then, the eye / eyeball positioning module detects and locates the eyes and eyeballs of the face output by the face detection module. Then, based on the positioning of the eyes and eyeballs, the position of the eyeball is calculated. Then, the eye movement posture is calculated. This patent tracks the movement trajectory of the eyeball and recognizes various eye movement postures through the eye movement tracking and recognition method, providing strong support for the YROS platform to achieve the eye posture command coding recognition and control function. However, this patent does not involve any solution to the problem that eye movement tracking cannot simultaneously detect the eye state, gaze estimation, and detect the pupil center. Summary of the Invention
[0007] The present invention provides an eye movement tracking method based on coupled cascade regression, which solves the problem that eye movement tracking cannot simultaneously detect the eye state, gaze estimation, and detect the pupil center.
[0008] In order to achieve the above technical effects, the technical solution of the present invention is as follows:
[0009] An eye movement tracking method based on coupled cascade regression includes the following steps:
[0010] S1: Perform face detection on the input picture and align the face key points;
[0011] S2: Use the face key points extracted in step S1 to extract the eye picture, initialize the eye key points, and calculate the local image features and global shape features of the eyes;
[0012] S3: Use the local image features and global shape features obtained in step S2 to estimate the eye state through the first regression model;
[0013] S4: Use the features obtained in step S2 and the eye state obtained in step S3 to estimate the three-dimensional gaze direction vector through the second regression model;
[0014] S5: Use the combination of the features obtained in step S2, the eye state obtained in step S3, and the gaze direction vector obtained in step S4 to estimate the pupil center position through the third regression model and update the key point position;
[0015] S6: Update the local image features, global shape features, eye state, and gaze direction according to the updated key point position, and alternately iterate multiple times to output stable eye state, gaze direction, and human eye pupil center position.
[0016] Further, the process of performing face detection in step S1 is:
[0017] For the input image, Haar-like features are calculated by computing the difference in pixel values within the feature rectangle region, and the integral image is used to accelerate the evaluation of Haar-like features. Each type of feature is classified using an Adaboost classifier. Different Adaboost classifiers are repeatedly trained, and finally these different classifiers are cascaded to obtain a strong classifier, which identifies the location of the face.
[0018] Furthermore, the process of aligning the facial key points in step S1 is as follows:
[0019] Calculate the features based on the marked points of each mean face, then perform face alignment by calculating the offset between the estimated face and the real face, and finally output the positions of the facial key points.
[0020] Furthermore, the specific process of step S1 is as follows:
[0021] S11: Grayscale the image to be tested;
[0022] S12: Scan the grayscale image with a search window, and calculate the Haar-like feature values of the sub-windows through the integral image;
[0023] S13: The strong classifier trained by cascaded AdaBoost filters the feature values of the sub-windows, and the sub-windows that pass through all the strong classifiers are the regions where the faces are located;
[0024] S14: Crop the face image from the input image, and extract Scale-Invariant Feature Transform (SIFT) features from 51 feature points. 128 SIFT features are extracted from each feature point;
[0025] S15: Use the obtained SIFT features and the method of Supervised Descent Method (SDM) to optimize the objective function:
[0026]
[0027] d(x0 + Δx) represents the marked points of the input image, h represents the non-linear feature extraction function, and φ * represents the manually marked SIFT features. Finally, the initial features are regressed to the real shape features of the face.
[0028] Furthermore, the specific process of step S2 is as follows:
[0029] S21: Select the key points of the human eyes from the obtained facial key points, and crop the images of the left eye and the right eye respectively;
[0030] S22: Initialize the key point positions of the intercepted eye image using the average key point positions of the human eye. The key point positions include 2 eye corner key points, 2 eyelid key points, 2 pupil edge key points, and 1 pupil center position;
[0031] S23: Extract the SIFT features of the 7 eye key points to form the local eye image features;
[0032] S24: Calculate the positive and negative differences between the pairwise human eye key point positions to form the global eye shape features.
[0033] Further, in step S2, the local eye image features are the SIFT features of the eye image, where SIFT features refer to Scale - Invariant Feature Transform; the global shape features are the positive and negative differences between the pairwise human eye key point positions.
[0034] Further, in step S3, the eye state is the probability that the eye is open, and this probability is between 0 and 1. Initialized as 1, it also represents the degree of occlusion of the eyelid to the pupil;
[0035] In step S3, the first regression model f t establishes a mapping relationship between the local image features Φ(x, I) and the global shape information Ψ(x) and the updated value of the eye state Δp. We define the updated value of the eye state as Δp t = f t (I, x t-1 ; θ f ), where θ f is the parameter of the regression model, I is the eye image, x t-1 is the key point position obtained from the previous iteration, estimating the target updated value of the eye state, and then adding it to the previous eye state p t-1 to obtain the eye state based on the local image features and the global key point structure shape information. The regression model is:
[0036] f t : I, x t-1 →Δp t
[0037] p t = p t-1 +Δp t
[0038] Further, in step S4, the three - dimensional line - of - sight direction vector represents three angles of the line of sight, namely the yaw angle, the pitch angle, and the roll angle. The three - dimensional line - of - sight direction vector is initialized as a three - dimensional zero vector;
[0039] The second regression model g tA mapping relationship is established among the local image feature Φ(x, I), the global shape information Ψ(x), the eye state p, and the updated value Δv of the line-of-sight direction. The updated value of the line-of-sight direction vector is defined as Δv t = g t (I, x t-1 , p t ; θ g ), where θ g is the parameter of the regression model, I is the eye image, x t-1 is the key-point position obtained in the previous iteration, p t is the eye state, the target updated value of the estimated line-of-sight direction is obtained, and then added to the previous line-of-sight direction v t-1 to obtain the estimated line-of-sight direction. The regression model is:
[0040] g t : I, x t-1 , p t → Δv t
[0041] v t = v t-1 + Δv t
[0042] Furthermore, in the step S5, the third regression model h t establishes a mapping relationship among the local image feature Φ(x, I), the global shape information Ψ(x), the eye state p, the line-of-sight direction vector v, and the key-point displacement Δx. We define the updated value of the key-point displacement as Δx t = h t (I, x t-1 , p t , v t ; θ h ), where θ h is the parameter of the regression model, I is the eye image, x t-1 is the key-point position obtained in the previous iteration, p t is the eye state, v t is the line-of-sight direction vector, the target updated value of the key-point displacement is obtained, and then added to the key-point position x t-1 obtained in the previous iteration to obtain the estimated key-point position and the pupil center position. The regression model is:
[0043] h t : I, x t-1 , p t , v t → Δx t
[0044] x t = x t-1+Δx t
[0045] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0046] The present invention proposes a unified method to simultaneously achieve human eye state detection, gaze estimation, and pupil center detection, and accurately complete three eye movement tracking tasks in real time, better solving the problem that a single method cannot achieve multiple eye movement tracking; integrating the global shape structure information and local image features of the eye image, considering the internal corresponding relationship between the pupil position change and gaze change of the human eye, and analyzing the human gaze by extracting the relative position shape features of the pupil center and surrounding key eye points, better solving the problem of low detection accuracy caused by adverse factors such as changes in lighting conditions, individual physiological characteristics, and head pose changes; adopting a lightweight cascaded regression method, with small latency and high accuracy, meeting the real-time requirements of eye movement tracking; this method has the advantages of simplicity, high efficiency, and strong robustness, and can be competent for different target scenarios. The present invention overcomes the deficiencies of the existing eye movement tracking methods, makes full use of the local texture features and shape features of the eye, and combines the relevance between eye movement tracking tasks, effectively solving the problem that eye movement tracking cannot simultaneously detect the eye state, gaze estimation, and detect the pupil center, greatly promoting the research of existing eye movement tracking, and having great research significance and practical application value. Description of the Drawings
[0047] Figure 1 Flow chart of the method of the present invention;
[0048] Figure 2 Extraction process of face pictures and eye pictures;
[0049] Figure 3 SIFT feature extraction process;
[0050] Figure 4 Example of the iterative process of the present invention;
[0051] Figure 5 Iterative algorithm process of the method of the present invention. Detailed Embodiments
[0052] The drawings are only for illustrative purposes and should not be construed as limitations on this patent;
[0053] To better illustrate this embodiment, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the size of the actual product;
[0054] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0055] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0056] Example 1
[0057] As Figure 1 shown, an eye movement tracking method based on coupled cascade regression includes the following steps:
[0058] S1: Perform face detection on the input image and align the face key points;
[0059] S2: Use the face key points extracted in step S1 to extract the eye images, initialize the eye key points and calculate the local image features and global shape features of the eyes;
[0060] S3: Use the local image features and global shape features obtained in step S2 to estimate the eye state through the first regression model;
[0061] S4: Use the features obtained in step S2 and the eye state obtained in step S3 to estimate the three-dimensional line-of-sight direction vector through the second regression model;
[0062] S5: Use the combination of the features obtained in step S2, the eye state obtained in step S3, and the line-of-sight direction vector obtained in step S4 to estimate the pupil center position through the third regression model and update the key point positions;
[0063] S6: Update the local image features, global shape features, eye state, and line of sight according to the updated key point positions, and alternately iterate multiple times to output stable eye state, line of sight, and human eye pupil center position.
[0064] This method realizes human eye state detection, line-of-sight estimation, and pupil center detection, and accurately completes the three eye movement tracking tasks in real time, better solving the problem that a single method cannot achieve multiple eye movement tracking; it integrates the global shape structure information and local image features of the eye image, considers the internal corresponding relationship between the change of the human eye pupil position and the line of sight, and analyzes the human line of sight by extracting the relative position shape features of the pupil center and the surrounding eye key points, better solving the problem of low detection accuracy caused by adverse factors such as changes in lighting conditions, individual physiological characteristics, and head postures; it adopts a lightweight cascade regression method with small latency and high accuracy, meeting the real-time requirements of eye movement tracking. The method has the advantages of simplicity, high efficiency, strong robustness, etc., and can be competent for different target scenarios. The present invention overcomes the deficiencies of the existing eye movement tracking methods, makes full use of the local texture features and shape features of the eyes, and combines the relevance between eye movement tracking tasks, effectively solving the problem that eye movement tracking cannot simultaneously detect the eye state, line-of-sight estimation, and detect the pupil center, greatly promoting the research of existing eye movement tracking, and having great research significance and practical application value.
[0065] Example 2
[0066] An eye movement tracking method based on coupled cascade regression, the process is as Figure 1 shown, and the specific implementation steps are as follows:
[0067] Step 1, perform face detection on the input picture and perform face key point detection;
[0068] The face detection method is a method based on a cascade haar classifier, which is a commonly used face detection method. This method calculates the haar-like features by calculating the difference of pixel values in the feature rectangle area for the input image, and uses the integral image to evaluate the haar-like features for acceleration. Each type of feature is classified by an Adaboost classifier, and different Adaboost classifiers are repeatedly trained. Finally, these different classifiers are cascaded to obtain a strong classifier, which can effectively identify the face position.
[0069] The face key point detection method is a face alignment method based on SDM. This algorithm mainly uses the least squares method for face alignment. This algorithm calculates the features of the marked points based on each mean face, and then aligns the face by calculating the offset between the estimated face and the real face, and finally outputs the face key point position.
[0070] The face detection and face key point alignment specifically include the following steps:
[0071] Step 11, grayscale the picture to be measured;
[0072] Step 11, scan the image with a search window on the grayscale image, and calculate the haar-like feature values of the sub-windows through the integral image;
[0073] Step 12, the strong classifier trained by the cascaded AdaBoost filters the feature values of the sub-windows, and the sub-windows filtered by all the strong classifiers are the areas where the faces are located;
[0074] Step 13, crop the face image from the input image, and extract the SIFT (Scale-Invariant Feature Transform) features from 51 feature points. 128 SIFT features are extracted from each feature point;
[0075] Step 14, use the obtained SIFT features and use the method of SDM (Supervised Descent Method) to optimize the objective function, and regress the initial features to the real shape features of the face.
[0076] The objective function is as follows:
[0077]
[0078] The process of face detection and face key point alignment is asFigure 2 as shown
[0079] Step 2: Using the face key points extracted in Step 1, extract the eye images, initialize the eye key points, and calculate the local image features and global shape features of the eyes;
[0080] The local image features of the eyes are the SIFT features of the eye images. The SIFT feature refers to the Scale-Invariant Feature Transform, which is an algorithm for detecting local features. This algorithm obtains features by finding the descriptors of the key points in an image and performs image feature point matching. The generation steps of the descriptors of the key points are as Figure 3 .
[0081] The global shape feature is the positive and negative differences between the positions of the eye key points pairwise.
[0082] The specific steps for extracting eye features include the following:
[0083] Step 21: Select the eye key points from the obtained face key points, and respectively intercept the images of the left and right eyes;
[0084] Step 22: Initialize the key point positions of the intercepted eye images using the average key point positions of the eyes. The key point positions include 2 eye corner key points, 2 eyelid key points, 2 pupil edge key points, and 1 pupil center position;
[0085] Step 23: Extract the SIFT features of 7 eye key points to form the local image features of the eyes;
[0086] Step 24: Calculate the positive and negative differences between the positions of the eye key points pairwise to form the global shape features of the eyes.
[0087] Embodiment 3
[0088] An eye movement tracking method based on coupled cascade regression, the process is as Figure 1 shown, and the specific implementation steps are as follows:
[0089] Step 1: Perform face detection on the input image and perform face key point detection;
[0090] The face detection method is a method based on a cascade haar classifier. This method is a commonly used face detection method. This method calculates the haar-like features by calculating the difference in pixel values in the feature rectangle area for the input image, and uses the integral image to evaluate the Haar-like features for acceleration. Each type of feature is classified using an Adaboost classifier, and different Adaboost classifiers are repeatedly trained. Finally, these different classifiers are cascaded to obtain a strong classifier, and this strong classifier can effectively identify the face position.
[0091] The described face key point detection method is a face alignment method based on SDM. This algorithm mainly uses the least squares method for face alignment. The algorithm calculates the features of the marked points of each mean face, and then aligns the face by calculating the offset between the estimated face and the real face, and finally outputs the positions of the face key points.
[0092] The specific steps of face detection and face key point alignment are as follows:
[0093] Step 11, grayscale the image to be tested;
[0094] Step 11, scan the image with a search window on the grayscale image, and calculate the haar-like feature values of the sub-windows through the integral image;
[0095] Step 12, use the strong classifier trained by cascaded AdaBoost to screen the feature values of the sub-windows. The sub-windows screened by all strong classifiers are the areas where the faces are located;
[0096] Step 13, crop the face image from the input image, and extract the SIFT (Scale-Invariant Feature Transform) features from 51 feature points. 128 SIFT features are extracted from each feature point;
[0097] Step 14, use the obtained SIFT features and use the SDM (Supervised Descent Method) method to optimize the objective function, and regress the initial features to the real shape features of the face.
[0098] The objective function is as follows:
[0099]
[0100] The process of face detection and face key point alignment is as Figure 2 shown.
[0101] Step 2, use the face key points extracted in Step 1 to extract the eye images, initialize the eye key points and calculate the local image features and global shape features of the eyes;
[0102] The local image features of the eyes are the SIFT features of the eye images. The SIFT feature refers to the Scale-Invariant Feature Transform, which is an algorithm for detecting local features. This algorithm obtains the features by finding the descriptors of the key points in an image and performs image feature point matching. The generation steps of the descriptors of the key points are as Figure 3 .
[0103] The described global shape feature is the positive and negative differences between the positions of the eye key points pairwise.
[0104] The specific steps of extracting eye features are as follows:
[0105] Step 21: Select the eye key points from the obtained face key points, and respectively extract the images of the left eye and the right eye.
[0106] Step 22: Initialize the key point positions of the intercepted eye images with the average key point positions of the eyes. The key point positions include 2 eye corner key points, 2 eyelid key points, 2 pupil edge key points, and 1 pupil center position.
[0107] Step 23: Extract the SIFT features of the 7 eye key points to form the local image features of the eyes.
[0108] Step 24: Calculate the positive and negative differences between the pairwise eye key point positions to form the global shape features of the eyes.
[0109] Step 3: Use the local image features and global shape features obtained in Step 2 to estimate the eye state through the first regression model.
[0110] The eye state is the probability that the eye is open. This probability is between 0 and 1 and is initialized to 1, which also represents the degree of occlusion of the eyelid to the pupil.
[0111] The first regression model f t establishes the mapping relationship between the local image features Φ(x, I) and the global shape information Ψ(x) and the eye state update value Δp. We define the eye state update value as Δp t = f t (I, x t-1 ; θ f ), where θ f is the parameter of the regression model, I is the eye image, and x t-1 is the key point position obtained in the previous iteration, that is, the target update value for estimating the eye state. Then, adding it to the previous eye state p t-1 can obtain the eye state based on the local image features and the global key point structure shape information. The regression model is:
[0112] f t : I, x t-1 → Δp t
[0113] p t = p t-1 + Δp t
[0114] Step 4: Use the features obtained in Step 2 and the eye state obtained in Step 3 to estimate the three-dimensional line-of-sight direction vector through the second regression model.
[0115] The three-dimensional line-of-sight direction vector represents three angles of the line of sight, namely the yaw angle, the pitch angle, and the roll angle, and the three-dimensional line-of-sight direction vector is initialized as a three-dimensional zero vector.
[0116] The second regression model g t establishes a mapping relationship between the local image feature Φ(x, I), the global shape information Ψ(x), the eye state p, and the updated value Δv of the line-of-sight direction. We define the updated value of the line-of-sight direction vector as Δv t = g t (I, x t-1 , p t ; θ g ), where θ g is the parameter of the regression model, I is the eye image, x t-1 is the key point position obtained from the previous iteration, p t is the eye state, that is, the target updated value of the estimated line-of-sight direction. Then, adding it to the previous line-of-sight direction v t-1 can obtain the estimated line-of-sight direction. The regression model is:
[0117] g t : I, x t-1 , p t → Δv t
[0118] v t = v t-1 + Δv t
[0119] Step 5: Using the features obtained in Step 2, the eye state obtained in Step 3, and the combination of the line-of-sight direction vector obtained in Step 4, estimate the pupil center position and update the key point position through the third regression model.
[0120] The third regression model h t establishes a mapping relationship between the local image feature Φ(x, I), the global shape information Ψ(x), the eye state p, the line-of-sight direction vector v, and the key point displacement Δx. We define the updated value of the key point displacement as Δx t = h t (I, x t-1 , p t , v t ; θ h ), where θ h is the parameter of the regression model, I is the eye image, x t-1 is the key point position obtained from the previous iteration, p t is the eye state, v t is the line-of-sight direction vector, that is, the target updated value of the key point displacement can be obtained. Then, adding it to the key point position x obtained from the previous iterationt-1 The estimated key point position and pupil center position can be obtained by addition. The regression model is as follows:
[0121] h t :I,x t-1 ,p t ,v t →Δx t
[0122] x t =x t-1 +Δx t
[0123] Step 6: Update the local image features, global shape features, eye state, and gaze direction according to the updated key point positions, and alternately iterate multiple times to output stable eye state, gaze direction, and human eye pupil center position; An example of the iteration process is as Figure 4 shown; The iteration algorithm is as follows (as Figure 5 shown):
[0124] Input: Human eye region image I, initialize the positions x 0 of N human eye key points (including human eye pupil, corner points of the eye, upper and lower eyelids, etc.), gaze direction v 0 and the possibility p 0 of the human eye being open.
[0125] Given the current eye key point position x t-1 . Estimate the updated value Δp t of the eye state, and update the eye state:
[0126] f t :I,x t-1 →Δp t
[0127] p t =p t-1 +Δp t
[0128] Given the current eye key point position x t-1 and the eye state p t estimated in the previous step, estimate the updated value Δv t of the gaze vector, and update the gaze vector:
[0129] g t :I,x t-1 ,p t →Δv t
[0130] v t =v t-1 +Δv t
[0131] Given the current position x of the eye key points t-1 and the estimated eye state p at the previous step t and the estimated gaze direction v t , update the position of the eye key points:
[0132] h t :I,x t-1 ,p t ,v t →Δx t
[0133] x t =x t-1 +Δx t
[0134] Output: the eye opening probability p T , the position x of the human eye key points T and the gaze vector v T .
[0135] The same or similar reference numerals correspond to the same or similar components;
[0136] The position relationships described in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0137] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. An eye movement tracking method based on coupled cascaded regression, characterized in that It includes the following steps: S1: Perform face detection on the input image and align the facial key points; S2: Use the facial key points extracted in step S1 to extract the eye images, initialize the eye key points, and calculate the local image features and global shape features of the eyes; S3: Use the local image features and global shape features obtained in step S2 to estimate the eye state through the first regression model; S4: Use the features obtained in step S2 and the eye state obtained in step S3 to estimate the three-dimensional line-of-sight direction vector through the second regression model; The three-dimensional line-of-sight direction vector represents three angles of the line of sight, namely the yaw angle, the pitch angle, and the roll angle, and the three-dimensional line-of-sight direction vector is initialized as a three-dimensional zero vector; The second regression model g mentioned above t establishes the mapping relationship between the local image feature Φ(x, I), the global shape information Ψ(x), the eye state p, and the updated value Δv of the line-of-sight direction. The updated value of the line-of-sight direction vector is defined as Δv t = g t (I, x t-1 , p t ; θ g ), where θ g is a parameter of the regression model, I is the eye image, x t-1 is the key point position obtained in the previous iteration, p t is the eye state, the target update value of the estimated line of sight direction, and then add it to the previous line of sight direction v t-1 to obtain the estimated line of sight direction; S5: Use the combination of the features obtained in step S2, the eye state obtained in step S3, and the line-of-sight direction vector obtained in step S4 to estimate the pupil center position through the third regression model and update the key point positions; S6: Update the local image features, global shape features, eye state, and line-of-sight direction according to the updated key point positions, and alternately iterate multiple times to output stable eye state, line-of-sight direction, and human eye pupil center position.
2. The eye movement tracking method based on coupled cascade regression according to claim 1, wherein The process of performing face detection in step S1 is as follows: For the input image, calculate the haar-like features by calculating the difference in pixel values in the feature rectangle area, and use the integral image to evaluate the Haar-like features for acceleration. Each type of feature is classified by an Adaboost classifier. Different Adaboost classifiers are repeatedly trained, and finally these different classifiers are cascaded to obtain a strong classifier, which identifies the face position.
3. The eye movement tracking method based on coupled cascade regression according to claim 2, characterized in that, The process of aligning the facial key points in step S1 is as follows: Calculate the features based on the marked points of each mean face, and then perform face alignment by calculating the offset between the estimated face and the real face, and finally output the facial key point positions.
4. The eye movement tracking method based on coupled cascade regression according to claim 3, wherein The specific process of step S1 is as follows: S11: Grayscale the image to be tested; S12: Scan the image with a search window on the grayscale image, and calculate the haar-like feature values of the sub-windows through the integral image; S13: The strong classifier trained by cascaded AdaBoost filters the feature values of the sub-windows, and the sub-windows filtered by all strong classifiers are the areas where the faces are located; S14: Crop the face image from the input image, and extract the scale-invariant feature transform features SIFT from 51 feature points. 128 SIFT features are extracted from each feature point; S15, Use the obtained SIFT features and use the method of supervised descent method SDM to optimize the objective function: d(x0+Δx) represents the marked points of the input image, h represents the non-linear feature extraction function, and φ * represents the SIFT features of the manually marked ground; Finally, the initial features are regressed to the real shape features of the face.
5. The eye movement tracking method based on coupled cascade regression according to claim 4, wherein, The specific process of step S2 is as follows: S21: Select the eye key points from the obtained facial key points, and crop the images of the left and right eyes respectively; S22: Use the average key point positions of the eyes to initialize the key point positions of the cropped eye images. The key point positions include 2 corner key points, 2 eyelid key points, 2 pupil edge key points, and 1 pupil center position; S23: Extract the SIFT features of 7 eye key points to form the local image features of the eyes; S24: Calculate the positive and negative differences between the positions of the human eye key points in pairs to form the global shape features of the eyes.
6. The eye movement tracking method based on coupled cascade regression according to claim 5, characterized in that In the step S2, the local image features of the eyes are the SIFT features of the eye images, and the SIFT features refer to scale-invariant feature transform; the global shape features are the positive and negative differences between the positions of the human eye key points in pairs.
7. The eye movement tracking method based on coupled cascade regression according to claim 6, wherein In the step S3, the eye state is the probability that the eyes are open. This probability is between 0 and 1, initialized to 1, and also represents the degree of occlusion of the eyelids on the pupils.
8. The eye movement tracking method based on coupled cascade regression according to claim 7, wherein In the step S3, the first regression model f t establishes the mapping relationship between the local image feature Φ(x, I) and the global shape information Ψ(x) and the eye state update value Δp. We define the eye state update value as Δp t = f t (I, x t-1 ; θ f ), where θ f is a parameter of the regression model, I is the eye image, and x t-1 is the key point position obtained in the previous iteration, estimating the target update value of the eye state, and then adding it to the previous eye state p t-1 to obtain the eye state based on local image features and global key point structure shape information.
9. The eye movement tracking method based on coupled cascade regression according to claim 8, characterized in that In the step S5, the third regression model h t establishes a mapping relationship among the local image feature Φ(x, I), the global shape information Ψ(x), the eye state p, the line-of-sight direction vector v, and the key-point displacement Δx. We define the key-point displacement update value as Δx t = h t (I, x t-1 , p t , v t ; θ h ), where θ h is a parameter of the regression model, I is the eye image, x t-1 is the key point position obtained in the previous iteration, p t is the eye state, v t is the line-of-sight direction vector, that is, the target update value for obtaining the key point displacement, and then adding it to the key point position x t-1 obtained in the previous iteration gives the estimated key point position and the pupil center position.
Citation Information
Patent Citations
Intelligent eyelid detection method and system
CN114360039A