A Gaze Estimation Method Based on Uncertainty Modeling

Through the gaze estimation method of probability embedding and label modeling, the input and label uncertainty problems in the gaze estimation are solved, and more accurate gaze point estimation and error measurement are achieved, reducing computational complexity.

CN116311442BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310216242.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-07-18
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

Existing deep learning-based appearance gaze estimation methods ignore uncertainty in gaze estimation, resulting in inaccurate modeling, especially input uncertainty and label uncertainty are not effectively modeled.

Method used

Through probability embedding and probability label modeling, image features are extracted using convolutional neural networks and mapped to multivariate Gaussian distributions in probability space. Combined with Monte Carlo sampling and embedded feature smoothing modules, a gaze estimation model is built to solve input and label uncertainty.

Benefits of technology

More accurate gaze point estimation is achieved, which can effectively measure the confidence of estimation error, reduce the gaze estimation error, and reduce the calculation amount through difficult mining methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311442B_ABST
    Figure CN116311442B_ABST
Patent Text Reader

Abstract

The present invention relates to a gaze estimation method based on uncertainty modeling, belonging to the field of human eye gaze estimation. First, a convolutional neural network (CNN) and a fully connected layer are used to extract the input image information and key point features respectively, then the features are fused, and the fused features are embedded into a multivariate Gaussian distribution in the probability space. The present invention also proposes an embedded feature smoothing module, which uses triplet loss to constrain the distribution of labels and the distribution of embedded features to learn smoother and more ordered probability embedded features, making the label distribution and the embedded feature distribution more consistent. And a method for mining hard examples is proposed to solve the problem of the explosion of the dimension of triplet data. By mapping the fused features to the probability space to model the input uncertainty existing in gaze estimation, and by probabilizing the labels to model the label uncertainty, and by introducing an embedded feature smoothing module to smooth the probability space, a more accurate gaze point estimation can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human eye gaze estimation, and particularly to a gaze estimation method based on uncertainty modeling. Background Art

[0002] Gaze is a basic visual behavior of humans and an important clue for studying human social interaction. Gaze estimation aims to predict the three-dimensional gaze direction or two-dimensional gaze point based on facial or eye images. Gaze estimation is the basis for many computer vision tasks, such as saliency prediction, salient object detection, and saccade path prediction. In addition, gaze estimation is also widely applied in various fields, such as medical diagnosis, human-computer interaction, and medical education. Therefore, accurate gaze estimation methods are very important for downstream computer vision tasks and applications in various fields.

[0003] According to Hansen et al. in the article "D.W. Hansen and Q. Ji, 'In the eye of the beholder: A survey of models for eyes and gaze,' IEEE Trans. Pattern Anal. Mach. Intell., vol. 32, no. 3, pp. 478–500, Mar. 2009.", gaze estimation methods are roughly divided into two categories: model-based gaze estimation methods and appearance-based gaze estimation methods. Model-based gaze estimation methods usually use precise geometric information, such as pupil size, eyeball angle, and eye image reflection, to estimate the gaze direction. However, these precise geometric features must be collected by high-precision sensors, such as high-precision infrared cameras or RGB-D cameras. Appearance-based gaze estimation methods usually directly regress the gaze result using facial or eye images without precise geometric information and geometric modeling. Therefore, appearance-based gaze estimation can be implemented by low-cost RGB cameras (e.g., built-in RGB cameras in mobile devices and laptops), thus having broad application prospects. With the introduction of large-scale datasets, many deep learning methods are used for appearance-based gaze estimation tasks to reduce the estimation error.

[0004] There are inherent ambiguities and external noises in the appearance-based gaze estimation problem, which lead to uncertainties. However, current deep learning-based appearance gaze estimation methods ignore the existing uncertainties, resulting in inaccurate modeling of the gaze estimation problem itself. Therefore, the present invention analyzes two types of uncertainties existing in gaze estimation: input uncertainty and label uncertainty. Thus, probability embedding and probability label modeling are proposed to model input uncertainty and label uncertainty respectively, and an embedding feature smoothing module is proposed to smooth the probability embedding features for more accurate gaze estimation. Summary of the Invention

[0005] Technical problem to be solved

[0006] Although many appearance-based gaze estimation methods have been proposed in recent decades, the uncertainty existing in gaze estimation has been ignored, resulting in insufficient modeling of the gaze estimation problem itself. The present invention classifies the uncertainty of gaze estimation into two major categories: (1) Input uncertainty: Due to the lack of depth information in the collected two-dimensional images, different three-dimensional gaze behaviors have consistent or similar facial and eye images, resulting in inherent ambiguity in the input of gaze estimation. In addition, other external factors during the data collection process may also cause input uncertainty, such as the observer's blinking, motion blur caused by head rotation, and illumination changes. (2) Label uncertainty: Psychological research shows that people usually gaze at a blurred area rather than a definite point. Therefore, gaze has inherent ambiguity. In addition, the ground truth of the gaze estimation dataset is usually obtained through instructions or eye tracking devices. However, for instructions, due to the subjectivity of the observer, instructing the observer to gaze at a certain point cannot ensure that the observer gazes at the center of the point; for eye tracking devices, the obtained ground truth has systematic errors. These reasons introduce label uncertainty into the line-of-sight estimation task. However, existing deterministic representation learning methods are difficult to represent this input uncertainty, and existing deterministic label learning is difficult to model label uncertainty.

[0007] To avoid the deficiencies of the prior art, the present invention provides a gaze estimation method based on uncertainty modeling.

[0008] Technical solution

[0009] The solution of the present invention is as follows: First, recruit subjects to collect an eye movement estimation dataset, use existing detection methods to detect faces and human eyes in the images in the collected eye movement estimation dataset, and record the corresponding image frame coordinates. Using the face and human eye images, perform probability embedding on the fused features through the method of probability embedding, that is, represent the fused features using a high-dimensional Gaussian distribution, and then sample the high-dimensional Gaussian distribution through Monte Carlo sampling to obtain a fused feature sampling sequence. Use a regressor to perform regression on the fused features to obtain the estimated gaze point. During training, convert the original deterministic label into a probability label to model label uncertainty.

[0010] A gaze estimation method based on uncertainty modeling, characterized by comprising:

[0011] S1: Construct a gaze estimation dataset;

[0012] S2: Detect faces and human eyes in the images;

[0013] S3: Extract features from the detected faces and human eyes respectively;

[0014] S4: Model the input uncertainty using probabilistic embedding;

[0015] S5: Regress the fixation points.

[0016] A further technical solution of the present invention: Specifically, S1 is as follows: Recruit subjects to observe a mobile phone, and require the subjects to fixate on randomly displayed dots on the mobile phone screen through instructions; Record the video of the subjects fixating on the dots through the mobile phone camera, and at the same time record the position coordinates of the randomly displayed dots; Align the recorded video with the position coordinates of the dots in time, so as to obtain the image frames and their corresponding fixation true coordinates.

[0017] A further technical solution of the present invention: Specifically, S2 is as follows: For the image frames recorded by the mobile phone, first use existing detection methods to detect the left and right eyes and the face image respectively, and obtain the left and right eye images and the face image {I f ,I eyel ,I eyer}, and obtain the coordinate positions L of the detected left and right eye image frames and the face image frame in the original image; Mark the images without detected faces and do not enter the subsequent process; Uniformly transform the sizes of all human eye images to h1×w1, where h1 represents the height of the human eye image and w1 represents the width of the human eye image; Uniformly transform the sizes of all face images to h2×w2, where h2 represents the height of the human eye image and w2 represents the width of the human eye image.

[0018] A further technical solution of the present invention: Specifically, S3 is as follows: Use CNN as the feature extraction network to extract features from the detected human eyes and left and right eyes respectively. In order to extract left and right eye features equally, share the network parameters of the left and right eye feature extraction networks, and fuse the left and right eye features through the CNN model; For the key point feature L, use a stacked fully connected layer to extract its features; In summary, obtain the face feature F f , the human eye feature F e , the feature F after extracting the key point features L ; And fuse the above features {F f ,F e ,F L} to obtain the feature F a .

[0019] A further technical solution of the present invention: The CNN model includes VGG, ResNet, and MobileNet.

[0020] A further technical solution of the present invention: Specifically, S4 is as follows: Map the fused features to a latent variable z, and the latent variable z is a Gaussian distribution in the probability space Among them, the mean of the Gaussian distribution represents the most likely feature representation, and the covariance of the Gaussian distribution represents the measure of the uncertainty of the input.

[0021] A further technical solution of the present invention: Specifically, S5 is: performing Monte Carlo sampling on the latent variable z to obtain a feature sequence {z1, z2,..., z T}, and passing the feature sequence {z1, z2,..., z φ} through a regressor R composed of stacked fully-connected layers T to perform regression respectively to obtain predicted fixation points {g1, g2,..., g T}.

[0022] A further technical solution of the present invention: It also includes modeling label uncertainty using probability labels, specifically: for an original label that is a deterministic point, label uncertainty cannot be modeled. Probabilizing the label means converting the original deterministic point into a two-dimensional Gaussian distribution where the mean of the two-dimensional Gaussian distribution is the original deterministic coordinate, and the variance of the two-dimensional Gaussian distribution is the size of the label uncertainty, used to measure the label quality.

[0023] A further technical solution of the present invention: The specific training of the model is: introducing probability embedding and probability labels into the loss function to obtain a new loss function, that is, the probability-regularized mean squared error PNMSE loss function; using triplet loss to constrain the embedding probability distribution and the label distribution; combining the PNMSE loss function with the triplet loss function and the prior loss function to obtain the final loss function; taking minimizing the loss function as the optimization goal, and using the Adam algorithm to train the fixation estimation model.

[0024] A further technical solution of the present invention: To address the problem of exponential explosion in the number of triplets, a hard example mining method is used to reduce the computational amount.

[0025] Beneficial effects

[0026] A fixation estimation method based on uncertainty modeling provided by the present invention, that is, using the images recorded by the camera during the fixation process, detecting the human face and human eyes in the images, and then respectively extracting the human face and human eye features and fusing them. By mapping the fused features to the probability space to model the input uncertainty existing in fixation estimation, and by probabilizing the labels to model the label uncertainty, and by introducing an embedded feature smoothing module to smooth the probability space, thereby achieving more accurate fixation point estimation. It has the following advantages:

[0027] 1) The present invention analyzes two types of uncertainties existing in gaze estimation, namely input uncertainty and label uncertainty. And the present invention proposes a method of probability embedding to model the input uncertainty, embedding the input into a multivariate Gaussian distribution in the probability space, where the mean of the multivariate high-dimensional Gaussian distribution represents the most likely feature representation, and the covariance of the multivariate Gaussian distribution measures the uncertainty of the input. And a method of probability label is proposed to model the label uncertainty.

[0028] 2) The present invention proposes an embedded feature smoothing module, which uses the triplet loss to constrain the distributions of the label and the embedded feature to learn a smoother and more ordered probability-embedded feature, making the label distribution and the embedded feature distribution more consistent. And a method of hard example mining is proposed to solve the problem of the explosion of the dimensionality of triplet data.

[0029] 3) The probability embedding proposed by the present invention can obtain the uncertainty estimation of the sample, which can effectively be used as a measure of the confidence of the gaze estimation error, solving the problem that the confidence of the result is unknown in the estimation by the previous deterministic methods, which is not conducive to practical applications.

[0030] 4) The uncertainty estimation of the sample obtained by the present invention is strongly positively correlated with the estimation error. Therefore, it can be used as a measure of the sample estimation error. The measure of the error estimation can be used as a criterion for sample screening. Samples with uncertainties greater than a certain value are rejected for recognition, thereby further reducing the gaze estimation error. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation to the present invention. Throughout the drawings, the same reference signs denote the same components.

[0032] Figure 1 is the overall flowchart for the implementation of the present invention;

[0033] Figure 2 is the overall framework diagram of the model of the present invention;

[0034] Figure 3 Schematic diagram of the extraction of the semantic vector sequence in the present invention;

[0035] Figure 4 Schematic diagram of the decoder LSTM network in the present invention for mining semantic relationships. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0037] The present invention provides a gaze estimation method based on uncertainty modeling. First, a convolutional neural network (CNN) and a fully connected layer are used to extract the input image information and key point features respectively, then the features are fused, and the fused features are embedded into a multivariate Gaussian distribution in the probability space. The mean of the multivariate high-dimensional Gaussian distribution represents the most likely feature representation, and the covariance of the multivariate Gaussian distribution measures the uncertainty of the input, thereby modeling the input uncertainty. The present invention proposes an embedded feature smoothing module, which uses triplet loss to constrain the distribution of labels and the distribution of embedded features to learn a smoother and more ordered probability embedded feature, making the label distribution and the embedded feature distribution more consistent. And a method for mining hard examples is proposed to solve the problem of the explosion of triplet data dimensions.

[0038] The specific steps are as follows:

[0039] Step 1, construct a gaze estimation dataset

[0040] Recruit subjects to observe a mobile phone, and require the subjects to fixate on a randomly displayed dot on the mobile phone screen through instructions. The video of the subjects fixating on the dot is recorded by the mobile phone camera, and the position coordinates of the randomly displayed dot are recorded at the same time. The recorded video and the position coordinates of the dot are aligned in time, so as to obtain the image frame and its corresponding gaze ground truth coordinates.

[0041] Step 2, perform face and eye detection on the image

[0042] Use existing detection methods to perform left and right eye and face image detection on the image respectively, and obtain the left and right eye images and the face image {I f , I eyel , I eyer}. Unify the sizes of all the eye images {I eyel , I eyer} to the size of h1×w1, where h1 represents the height of the eye image and w1 represents the width of the eye image. Unify the sizes of all the face images I f to the size of h2×w2, where h2 represents the height of the eye image and w2 represents the width of the eye image. Obtain the coordinate positions of the left and right eye image detection frames and the face image detection frame in the original image from the detection results, and add the mobile phone placement mode information to form the key point feature L:

[0043] L = {B e , B f , O}

[0044] where B e represents the position of the left and right eye image detection frame, B fIndicates the position of the face image detection frame, and O indicates the orientation information of the mobile phone placement.

[0045] Mark the images for which no face or human eyes are detected, and these images do not enter the subsequent processing flow.

[0046] Step 3: Extract features from the detected face and human eyes respectively

[0047] Use CNN to extract features from the detected human eye image and left and right eye images respectively. To equally extract human eye features, the left eye image and the right eye image feature extraction networks share network parameters, and the extracted left eye features and right eye features are fused through CNN. For the key point feature L, use a stacked fully connected layer to extract its features. The above process can be expressed as:

[0048] F f = CNN(I f ),

[0049] F e = CNN(I e ),

[0050] F L = FC(L),

[0051] where CNN represents the convolutional neural network, and FC represents the stacked fully connected layer. F f is the extracted face feature, F e is the extracted human eye feature, F L is the extracted key point feature. And stack the above features to get:

[0052] {F f ,F e ,F L}= concat(F f ,F e ,F L ),

[0053] where concat represents the stacking operation. Then, fuse the stacked features {F f ,F e ,F L} through a fully connected layer to get:

[0054] F a = FC({F f ,F e ,F L}),

[0055] where F a represents the fused feature.

[0056] Step 4, Model the input uncertainty using probabilistic embedding

[0057] Map the fused feature F a to a latent variable z, and the conditional distribution of the latent variable z is a Gaussian distribution in the probability space:

[0058]

[0059] where I represents the set of all inputs {I f , I eyel , I eyer , L}. The mean μ of the Gaussian distribution z represents the most likely feature representation, and the covariance ∑ of the Gaussian distribution z represents the uncertainty measure of the input.

[0060] Step 5, Regress the fixation points

[0061] Perform Monte Carlo sampling on the conditional distribution p(z|I) of the latent variable z to obtain a sequence of features:

[0062]

[0063] where T1 is the number of samples for sampling, and the sequence of features is independently and identically distributed. Then, through a regressor R that stacks fully connected layers φ regress the sequence of features respectively:

[0064]

[0065] where g i represents the predicted fixation point after regression, and obtain the predicted fixation points

[0066] Step 6, Model the label uncertainty using probabilistic labels and train the model

[0067] The original label of the training sample in the dataset is a deterministic point, and this representation method cannot model the label uncertainty. Probabilize the label, that is, transform the original deterministic point into a two-dimensional Gaussian distribution p(g):

[0068]

[0069] where represents the Gaussian distribution, the mean μ of the two-dimensional Gaussian distribution g is the original deterministic label The covariance matrix ∑ of the two-dimensional Gaussian distribution g is the size of the label uncertainty, which is used to measure the label uncertainty and thus reflect the label quality.

[0070] To obtain the analytical solution of the final maximum likelihood estimate, Monte Carlo sampling is performed on the probability label p(g) to obtain the label sequence:

[0071]

[0072] where T2 is the number of sampled samples, and the feature sequence is independently and identically distributed.

[0073] The probability embedding and the probability label are introduced into the loss function to obtain the Probabilistic-weight Normalized Mean Square Error (PNMSE) loss function:

[0074]

[0075] In addition, the feature smoothing module is used to smooth the probability embedding feature z and make the embedding probability distribution more consistent with the label distribution. Specifically, the triplet loss is used to constrain the embedding probability distribution and the label distribution, that is, for any given triplet (i, j, k), the following formula is satisfied:

[0076]

[0077] where D(·) represents the probability distribution distance, such as KL divergence, symmetric KL divergence, etc., and d(·) represents the Euclidean distance. However, since the number of traversing all triplets explodes exponentially with the scale of the dataset, a hard example mining method is adopted to reduce the computational amount. Specifically, for each batch of data with a batch size, each index is traversed and selected as the anchor element i of the triplet, then the second element j is randomly selected from the remaining batch, and for the third element k, the index closest to the label is selected from the remaining elements. The formula is as follows:

[0078]

[0079] where, is the index set of the dataset. After the above steps, the number of triplets is equal to the batch size, and the set of all formed triplets is denoted as and the following constraints are used to smooth the probability embedding feature z and make the embedding probability distribution more consistent with the label distribution:

[0080]

[0081] where |·| represents the cardinality of the set, and η is the relaxation parameter. In addition, to prevent the mean of the embedding distribution from deviating or the covariance matrix from being too large, a prior loss function is introduced:

[0082]

[0083] Among them, E represents the identity matrix.

[0084] In summary, the total loss function is as follows:

[0085]

[0086] Among them, λ t and λ p are balance coefficients. With the goal of minimizing the total loss function, the Adam (Adaptive Moment Estimation) algorithm is used to train the fixation estimation model based on uncertainty modeling.

[0087] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A gaze estimation method based on uncertainty modeling, characterized in that Including: S1: Construct a fixation estimation dataset; S2: Detect human faces and eyes in the image; specifically: for the image frames recorded by the mobile phone, first use existing detection methods to detect the left and right eyes and the human face image respectively, and obtain the human face image and the left and right eye images {I f ,I eyel ,I eyer}. Resize the sizes of the left and right eye images {I eyel ,I eyer} to the size of h1×w1 uniformly, where h1 represents the height of the human eye image and w1 represents the width of the human eye image; resize the sizes of all the human face images I f to the size of h2×w2 uniformly, where h2 represents the height of the human face image and w2 represents the width of the human face image; and obtain the coordinate positions of the detected left and right eye image frames and the human face image frames in the original image, and add the mobile phone placement mode information to form the key point feature L: L = {B e ,B f ,O}, where B e represents the position of the left and right eye image detection frame, B f represents the position of the human face image detection frame, and O represents the mobile phone placement orientation information; mark the images without detected human faces and do not enter the subsequent processes; S3: Extract features from the detected human face and human eyes respectively; specifically: Use the CNN model as the feature extraction network to extract features from the detected human face and left and right eyes respectively. In order to extract left and right eye features equally, share the network parameters of the left and right eye feature extraction networks, and fuse the left and right eye features through the CNN model; for the key point feature L, use a stacked fully connected layer to extract its features; In summary, obtain the human face feature F f , human eye feature F e , key point feature F L ; And fuse the above features {F f , F e , F L} to obtain the feature F a ; S4: Model the input uncertainty using probabilistic embedding; specifically: Map the fused feature F a to a latent variable z, where the latent variable z is a Gaussian distribution in the probability space and the mean μ of the Gaussian distribution z represents the most likely feature representation, and the covariance ∑ of the Gaussian distribution z represents the uncertainty measure of the input; S5: Regress the fixation point; specifically: perform Monte Carlo sampling on the latent variable z to obtain a feature sequence: where T1 is the number of sampled samples, and the feature sequence is independently and identically distributed; And through the regressor R that stacks fully connected layers φ Perform regression on the feature sequences respectively: where g i represents the predicted fixation point after regression, obtaining the predicted fixation point S6: Model the label uncertainty using probability labels; specifically: for an original label that is a deterministic point, the label uncertainty cannot be modeled. Probabilize the label, that is, convert the original deterministic point into a two-dimensional Gaussian distribution p(g): Among them represents a Gaussian distribution, and the mean μ of the two-dimensional Gaussian distribution g is the original deterministic label The covariance matrix ∑ of the two-dimensional Gaussian distribution g is the magnitude of label uncertainty, used to measure the label uncertainty and thus reflect the label quality; The training model is specifically: introduce the probability embedding and probability label into the loss function to obtain the PNMSE loss function: Use the triplet loss to constrain the embedding probability distribution and the label distribution, that is, given any triplet (i, j, k) satisfying the following formula: where D(·) represents the probability distribution distance, and d(·) represents the Euclidean distance; for each batch of data with size batch, traverse and select each index as the anchor element i of the triplet, then randomly select the second element j from the remaining batch, and for the third element k, select the index closest to the label from the remaining elements. The formula is as follows: Among them, is the index set of the data set; after the above steps, the number of triples is equal to the batch size, and the set of all formed triples is denoted as And the following constraints are used to smooth the latent variable z and make the embedding probability distribution more consistent with the label distribution: where |·| represents the cardinality of the set, and η is the relaxation parameter; in addition, to prevent the mean of the embedding distribution from deviating or the covariance matrix from being too large, introduce a prior loss function: where E represents the identity matrix; In summary, the total loss function is as follows: where λ t and λ p are balance coefficients; Taking the minimization of the total loss function as the optimization objective, use the Adam algorithm to train the fixation estimation model based on uncertainty modeling.

2. The gaze estimation method based on uncertainty modeling according to claim 1, wherein S1 is specifically: Recruit subjects to observe a mobile phone, and require the subjects to fixate on randomly displayed dots on the mobile phone screen through instructions; record the video of the subjects fixating on the dots through the mobile phone camera, and at the same time record the position coordinates of the randomly displayed dots; align the recorded video and the position coordinates of the dots in time to obtain the image frames and their corresponding fixation ground truth coordinates.

3. The gaze estimation method based on uncertainty modeling according to claim 1, wherein The described CNN model includes VGG, ResNet, and MobileNet.

Citation Information

Patent Citations

  • Uncertain sample detection method and device in defect detection and medium

    CN114092472A

  • Fuzzy 3D skeleton action recognition method and device based on self-supervised learning

    CN114373224A