Head posture correction staring direction estimation method based on multiple modes
By fusing multimodal information of eye movement and head posture, a multi-scale feature-enhanced convolutional neural network and graph convolutional attention fusion module is constructed, which solves the problem of insufficient head posture correction, improves the accuracy and stability of gaze direction estimation, and optimizes the user experience.
Patent Information
- Application Number
- CN202510245210.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-25
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, in multi-view or dynamic environment, the correction of head posture is insufficient, resulting in limited accuracy and robustness of gaze direction estimation.
A multimodal-based head attitude correction method is adopted to construct a multi-scale feature-enhanced convolutional neural network and graph convolutional attention fusion module by fusing the information of eye movement and head attitude, and combined with an adaptive optimization algorithm, the gaze direction deviation caused by changes in head attitude is corrected.
It significantly improves the accuracy and stability of gaze direction estimation, optimizes model performance, and improves user experience, especially in complex scenarios or multi-view situations.
Smart Images

Figure CN120340083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing and pattern recognition, and specifically relates to a multi-modal based head pose correction and gaze direction estimation method. Background Art
[0002] Gaze direction estimation is a human-computer interaction technology aimed at understanding the user's attention and intention by monitoring the user's eye movements and fixation points. This technology is based on eye tracking technology, which uses cameras or sensors to capture the movement information of the user's eyes, thereby inferring the user's focus of attention on the screen or in the environment. Gaze direction estimation technology has a wide range of applications in the fields of virtual reality, augmented reality, human-computer interaction, user experience design, etc., and can help improve the design of the user interface, enhance the user experience, and achieve a more intelligent human-computer interaction method.
[0003] In recent years, deep learning-based gaze direction estimation technology has made significant progress, capable of estimating the user's gaze direction from face images and eye images, and achieving high accuracy and stability in fixation point prediction. However, many existing methods often do not fully consider the influence of head pose when dealing with complex scenes. Although some models can incorporate head pose information, there are still deficiencies in head pose correction in multi-view or dynamic environments, which may limit the accuracy and robustness of the models. Summary of the Invention
[0004] Aiming at the problem that there are still deficiencies in head pose correction in multi-view or dynamic environments in the existing technology, the present invention proposes a multi-modal based head pose correction and gaze direction estimation method. By fusing multi-modal information of eye movement and head pose, the gaze direction deviation caused by head pose changes is corrected, thereby improving the accuracy and stability of gaze direction estimation. Especially in complex scenes or multi-view situations, the method of the present invention can significantly optimize the model performance, provide more accurate fixation point prediction results, and thus improve the user experience of the eye tracking system.
[0005] The present invention adopts the following technical solutions to achieve the above invention objectives:
[0006] A multi-modal based head pose correction and gaze direction estimation method, comprising the following steps:
[0007] Step 1, preprocess the gaze estimation dataset and divide it into a training set and a test set;
[0008] Step 2, construct a multi-modal based head pose correction and gaze direction estimation model;
[0009] Step 3: Use the binocular images and face images in the training set as inputs to train a multi-modal head pose correction and gaze direction estimation model;
[0010] Step 4: Use the trained gaze direction estimation model to predict the test set data.
[0011] In the said Step 1, the method for preprocessing the gaze estimation data set is as follows:
[0012] Sub-step 1-1: Use the Dlib library for face detection and face key point localization. Detect 8 facial key points, match the obtained 2D face key points with the corresponding points on the standard 3D face model through an optimized projection point formula to obtain the key points of the 3D model. The center of the inner corners of the two eyes on the 3D face is used as the origin g0 of the 3D line-of-sight direction. The optimized projection point formula is:
[0013]
[0014] In the formula, P represents the orthogonal projection matrix, R represents the rotation matrix, s represents the scaling coefficient, σ represents the activation function, α i represents the shape vector coefficient of this face model, β i represents the texture vector coefficient of this face model, f(S i ) represents the non-linear transformation function of the shape vector, g(T i ) represents the non-linear transformation function of the texture vector, t 2d represents the displacement matrix, and m represents the number of shape vectors and texture vectors.
[0015] Sub-step 1-2: Normalize the images in the training set and their corresponding three-dimensional gaze direction vectors. Introduce the weight coefficient ω i Considering the importance of each key point, calculate the coordinates of the face center in the pixel coordinate system as C f , and the formula is:
[0016]
[0017] In the formula, α1 is a non-linear adjustment factor, which is adjusted according to the importance of different key points.
[0018] Next, solve the rotation matrix R' so that after the original image is rotated, the face center C in the camera coordinate system f can be aligned with the set reference coordinate C′ in the image. In the process of optimizing the rotation matrix, the inventor proposed an optimization method based on position matching, which minimizes by considering multi-angle rotation,
[0019] ensuring that the model can adapt to diverse rotation scenarios to reduce the error caused by rotation. The rotation matrix
[0020] It is defined as:
[0021] where r j ′ k represents the element in the j-th row and k-th column of the obtained optimal rotation matrix R', and N represents the number of reference points and matching points. represents the k-th component of the i-th reference point. represents the j-th component of the i-th matching point.
[0022] Based on the optimization of the rotation matrix, a perspective affine transformation matrix M' is further introduced to optimize the final spatial transformation effect for processing face images in different poses. The perspective affine transformation matrix M' is defined as:
[0023] M′ = K′S′R′ΛK′ -1
[0024] where K' represents the internal parameter matrix of the virtual camera C', S′ represents the dynamic distance scaling matrix, Λ represents the adjustment matrix, and K' -1 represents the inverse matrix of the K' matrix.
[0025] The three-dimensional gaze direction vectors corresponding to the images in the training set are also normalized, and the 3D gaze vectors are further converted into 2D gaze vectors. The specific formula is as follows:
[0026] g 2D = P·R′·g 3D
[0027] where g 2D represents the normalized gaze direction vector, and g 3D represents the original gaze direction vector.
[0028] Sub-step 1-3: For the normalized pictures, the inventor applies a regional gamma correction method to the detected face and eye regions to adapt to different lighting conditions. The corresponding gamma correction formula is as follows:
[0029]
[0030] where I corrected (x,y) is the corrected pixel value at the position (x,y), I original (x,y) is the original pixel value at the position (x,y), γ is the global gamma value, A face (x,y), B face (x,y), C face (x,y) are the adjustment parameters for the face and eyes.
[0031] Sub-step 1-4. To improve the robustness of the system against camera blur conditions, the inventors proposed an improved multi-scale Gaussian blur method. By introducing multi-scale Gaussian blur with adaptive weights, it can dynamically adjust the blur degree at different scales according to the texture features and variance information of the local regions of the image, thereby more flexibly processing different regions of the image, enhancing the ability to retain details and effectively reducing noise interference. It is defined as follows:
[0032]
[0033] In the formula, O(x, y) is the pixel value of the output image at position (x, y), I(x, y) represents the pixel value of the input image at position (x, y), G(i, j, σ) represents the value of the Gaussian kernel at position (x, y), σ = f(V(x + i, y + j)) represents the standard deviation adjusted by the locality of the pixel, V(x + i, y + j) represents the local characteristics of the pixel, n represents the size of the filter, Z represents the normalization factor, Var k (x, y) represents the local variance at position (x, y) at the k-th scale, α2 represents the attenuation factor, and K represents the number of scales of the Gaussian kernel.
[0034] In the said step 2, the method for constructing the human eye gaze direction estimation model is as follows:
[0035] Sub-step 2-1: Design a multi-scale feature enhancement convolutional neural network module to extract the facial feature f f_face from the preprocessed face image. This module takes the face image as the input and introduces a multi-scale convolutional neural network to enhance the extraction of facial information at different scales. Different convolutional kernel sizes enable the module to capture information from details to the global, forming a richer facial feature representation and outputting the final facial feature vector f f_face which is defined as:
[0036] f f_face = Multi-scaleCNN(Image face )
[0037] where, CNN represents the Convolutional Neural Network.
[0038] To extract the head pose feature vector f f_face from the facial feature vector f f_head, the inventors proposed a graph convolutional attention fusion module, which can adaptively extract head pose-related features and effectively suppress noise. The key points of facial features learn weights through a two-layer graph convolutional network. The graph convolutional network (Graph Convolutional Networks, GCN) performs convolution operations on the input facial key point map to generate an enhanced local spatial information representation. The output graph convolutional features are combined with the facial residual network features to further extract more expressive pose features. Finally, a multi-layer perceptron (Multi-Layer Perceptron, MLP) is used to take f f_head as the input to estimate the head pose, and the estimated head pose α i ′ is defined as:
[0039] α i ′ = MLP(σ(A2f f_face W2 + σ(A1f f_face W1)) + f res )
[0040] In the formula, α i ′ represents the estimated head pose, σ represents the sigmoid activation function, A1 and A2 are the adjacency matrices of the graph convolutional network, W1 and W2 are the weight matrices of each layer of the graph convolutional network, and f res is the facial global feature extracted by the residual network.
[0041] In the head pose regression task, a loss function combining mean squared error and L2 regularization is designed to ensure the accuracy and stability of the model. The specific formula is defined as:
[0042]
[0043] Among them, represents the true value of the head pose, λ is the regularization coefficient, W k is the k-th parameter in the learnable parameters, θ j is the j-th parameter in the MLP, m is the number of weight parameters W k , and n is the number of parameters in the MLP.
[0044] Sub-step 2-2: Design a spatial shallow multi-stream network module based on the attention mechanism to extract fused eye features from the preprocessed binocular images This module takes the binocular images as the input and passes through the SW-GazeNet architecture respectively. This architecture contains multiple convolutional layers, where the first two layers are standard convolutional layers, and the subsequent layers are depthwise separable convolutional layers. Batch normalization and ReLU activation are performed after each convolutional layer, and finally, it is further processed by the spatial mechanism to obtain the left-eye feature representation fl_eye and the right-eye feature representation f r_eye , then calculate the self-correlation weight of the feature through a parallel attention branch network, use the Sigmoid function to adjust the weight range, multiply the feature vectors obtained from both eyes by the corresponding weight vectors, so as to obtain the weighted feature f l ′ _eye and f r ′ _eye , use the Softmax function to calculate the fusion weights ω l_eye and ω r_eye , and then perform weighted fusion to obtain the final eye feature representation f eye ; finally, use an MLP with f eye as the input to obtain the initial gaze direction vector g b ′, the estimated initial gaze direction vector g b ′ is defined as:
[0045] g′ b = MLP(Softmax((f l_eye ⊙A(f l_eye )),(f r_eye ⊙A(f r_eye ))))
[0046] In the formula, g b ′ represents the initial gaze direction vector estimated by the network, MLP represents the multi-layer perceptron, Softmax is a normalization function, ⊙ represents element-wise multiplication, and A represents the attention mechanism network.
[0047] The inventor designed a comprehensive loss function. By introducing directional loss, smooth loss, and weight regularization loss, it not only ensures the accuracy of the predicted line-of-sight direction, but also improves the prediction stability in time-series data and the generalization ability of the model, thereby optimizing the overall prediction effect. The specific formula is defined as:
[0048]
[0049] In the formula, g * represents the true value of the line-of-sight direction of the input image, g t ′ represents the predicted line-of-sight direction vector at the t-th iteration, ω l_eye represents the attention weight vector of the left eye, ω r_eye represents the attention weight vector of the right eye, α represents the weight coefficient of the directional loss, β represents the weight coefficient of the smooth loss, γ represents the weight coefficient of the weight regularization loss, and T max represents the total number of iterations.
[0050] Sub-step 2-3: Design an eye feature vector and head pose feature vector fusion module based on the multi-head attention mechanism. First, the eye feature representation f eye is concatenated with the head pose feature f f_head to obtain the concatenated feature representation f contact . Then, for each attention head of the concatenated feature representation f contact , calculate its corresponding Q i , K i and V i , and use the scaled dot-product attention mechanism to calculate the attention weights of each head. The weighted sum of these weights gives the output of each head. Secondly, the outputs of all heads are concatenated together and passed through a linear transformation to obtain the final fused feature. Finally, use an MLP with f fusion as the input to obtain the gaze direction vector g b ′ _adjust of the fused head pose. The specific formula is defined as follows:
[0051] f contact =Concat(f eye ,f f_head )
[0052]
[0053] g′ b_adjust =MLP(f fusion )
[0054] where d k is the dimensionality of Q i , K i , V i , and W o is the weight parameter of the linear projection.
[0055] The loss function in the gaze direction estimation task of the fused head pose is specifically defined as:
[0056]
[0057] In step 3, to train the multi-modal head pose correction gaze direction estimation model, the inventor designed a full-model adaptive optimization algorithm to optimize the loss function. Based on step 2, the total loss function is designed, and the loss weights of each task are dynamically and adaptively adjusted by introducing a dynamic weight adjustment mechanism, which is defined as:
[0058] Loss total =λ1(t)·Loss eye +λ2(t)·Loss head +λ3(t)·Loss adjust
[0059] Wherein, λ1(t), λ2(t), and λ3(t) are the dynamic weights of each task at the t-th iteration, which are automatically adjusted according to the loss value of each task and satisfy the condition λ1(t) + λ2(t) + λ3(t) = 1. The formula for the dynamic weight adjustment mechanism is:
[0060]
[0061] Wherein, Loss i (t) represents the loss value of the i-th task at time step t, and T represents the temperature parameter.
[0062] By introducing an adaptive weight adjustment mechanism, the loss weights of each task are dynamically adjusted, the weight matrix is corrected, and the AdamW optimizer is combined for adaptive learning rate update. In each iteration, the weight matrix and bias term are corrected through backpropagation, and combined with the adaptive weight adjustment mechanism, the minimum value of each loss function is solved by the gradient descent method. The optimization formula is:
[0063]
[0064] Wherein, θ t is the weight matrix of the model, η is the adaptive learning rate, and are the momentum and squared momentum respectively, λ is the weight decay coefficient, and ε is a small constant for numerical stability.
[0065] Beneficial Effects
[0066] (1) The present invention provides a multi-modal based head pose correction and gaze direction estimation model, which uses a multi-scale feature enhanced convolutional neural network and a GCN network to improve the ability to extract features from face images, effectively correcting the influence of different head poses on gaze estimation; eye features are extracted through the SW-GazeNet architecture, and the weights are adaptively adjusted in combination with the attention mechanism to ensure the accuracy of binocular feature fusion; the eye and head pose features are fused through the multi-head attention mechanism, enabling the model to have high precision and robustness in complex scenarios and significantly improving the performance of gaze direction estimation.
[0067] (2) The present invention provides a full model adaptive optimization algorithm that uses a dynamic weight adjustment mechanism to automatically optimize weights according to task losses, improving the stability and generalization ability of the model; the adaptive optimization algorithm is combined with the multi-modal data fusion technology, not only enhancing the real-time processing ability of the system, but also improving its robustness and gaze direction prediction performance in complex scenarios, significantly improving the user experience in the fields of virtual reality and human-computer interaction, and providing a more intelligent and natural human-computer interaction method. Description of the Drawings
[0068] Figure 1 This is the implementation flowchart of the gaze direction estimation method proposed by the present invention;
[0069] Figure 2 This is the block diagram of the multi-scale feature enhanced convolutional neural network module;
[0070] Figure 3 This is the block diagram of the spatial shallow multi-stream network based on the attention mechanism, where FC is the fully connected layer and Attention-branch is the attention branch network;
[0071] Figure 4 This is the block diagram of the SW-GazeNet architecture;
[0072] Figure 5 This is the comparison chart of the average angular error between the improved model and other methods on the test set, where MHPC-GazeNet (Multi-Head Pose Corrected Gaze Network) is the model of the present invention. Specific implementation method
[0074] The present invention will be described in detail below through specific examples in combination with the accompanying drawings of the specification. It should be noted that the following implementation is only to help those skilled in the art further understand the present invention, but does not limit the present invention in any form. The technical solution of the present invention is as follows:
[0075] See the accompanying drawing of the specification Figure 1 , a multi-modal based head pose corrected gaze direction estimation method, the steps of which include:
[0076] Step 1, preprocess the gaze estimation data set and divide it into a training set and a test set;
[0077] Step 2, construct a multi-modal based head pose corrected gaze direction estimation model;
[0078] Step 3, use the binocular images and face images in the training set as inputs to train the multi-modal based head pose corrected gaze direction estimation model;
[0079] Step 4, use the trained gaze direction estimation model to predict the test set data.
[0080] In the said Step 1, the method for preprocessing the gaze estimation data set is:
[0081] Sub-step 1-1 uses the Dlib library for face detection and facial landmark localization, detecting 8 facial landmarks. The obtained 2D facial landmarks are matched with the corresponding points on the standard 3D face model by optimizing the projection point formula to obtain the landmarks of the 3D model. The center of the inner corners of the two eyes on the 3D face is used as the origin g0 of the 3D line-of-sight direction. The optimized projection point formula is as follows:
[0082]
[0083] In the formula, P represents the orthogonal projection matrix, R represents the rotation matrix, s represents the scaling coefficient, σ represents the activation function, α i represents the shape vector coefficient of this face model, β i represents the texture vector coefficient of this face model, f(S i ) represents the non-linear transformation function of the shape vector, g(T i ) represents the non-linear transformation function of the texture vector, t 2d represents the displacement matrix, and m represents the number of shape vectors and texture vectors.
[0084] Sub-step 1-2 normalizes the images in the training set and their corresponding three-dimensional gaze direction vectors, introduces the weight coefficient ω i considering the importance of each landmark, and calculates the coordinates of the face center in the pixel coordinate system as C f , and the formula is:
[0085]
[0086] In the formula, α1 is a non-linear adjustment factor, which is adjusted according to the importance of different landmarks.
[0087] Next, solve the rotation matrix R' so that after the original image is rotated, the face center C f in the camera coordinate system can be aligned with the set reference coordinate C′ in the image. In the process of optimizing the rotation matrix, the inventor proposed an optimization method based on position matching, which minimizes by considering multi-angle rotation,
[0088] ensuring that the model can adapt to diverse rotation scenarios to reduce the error caused by rotation. The rotation matrix
[0089] is defined as:
[0090] In the formula, r j ′ k represents the element in the j-th row and k-th column of the obtained optimal rotation matrix R', N represents the number of reference points and matching points, represents the k-th component of the i-th reference point, Represents the j-th component of the i-th matching point.
[0091] Based on the optimization of the rotation matrix, a perspective affine transformation matrix M' is further introduced to optimize the final spatial transformation effect for processing face images in different poses. The perspective affine transformation matrix M' is defined as:
[0092] M′ = K′S′R′ΛK′ -1
[0093] In the formula, K' represents the internal parameter matrix of the virtual camera C', S′ represents the dynamic distance scaling matrix, Λ represents the adjustment matrix, and K' -1 represents the inverse matrix of the K' matrix.
[0094] Normalize the three-dimensional gaze direction vectors corresponding to the images in the training set, and further convert the 3D gaze vectors into 2D gaze vectors. The specific formula is as follows:
[0095] g 2D = P·R′·g 3D
[0096] In the formula, g 2D represents the normalized gaze direction vector, and g 3D represents the original gaze direction vector.
[0097] Sub-step 1-3: For the normalized pictures, the inventor applies a regional gamma correction method to the detected face and eye regions to adapt to different lighting conditions. The corresponding gamma correction formula is as follows:
[0098]
[0099] In the formula, I corrected (x,y) is the corrected pixel value at position (x,y), and I original (x,y) is the original pixel value at position (x,y), γ is the global gamma value, A face (x,y), B face (x,y), C face (x,y) are the adjustment parameters for the face and eyes.
[0100] Sub-step 1-4: To improve the robustness of the system to camera blur conditions, the inventor proposes an improved multi-scale Gaussian blur method. By introducing multi-scale Gaussian blur with adaptive weights, the blur degree of different scales can be dynamically adjusted according to the texture features and variance information of the local regions of the image, so as to more flexibly process different regions of the image, enhance the ability to retain details and effectively reduce noise interference. It is defined as follows:
[0101]
[0102] In the formula, O(x, y) is the pixel value of the output image at position (x, y), I(x, y) represents the pixel value of the input image at position (x, y), G(i, j, σ) represents the value of the Gaussian kernel at position (x, y), σ = f(V(x + i, y + j)) represents the standard deviation adjusted by the locality of the pixels, V(x + i, y + j) represents the local characteristics of the pixels, n represents the size of the filter, Z represents the normalization factor, Var k (x, y) represents the local variance at position (x, y) at the k-th scale, α2 represents the decay factor, and K represents the number of scales of the Gaussian kernel.
[0103] In the said step 2, the method for constructing the human eye gaze direction estimation model is as follows:
[0104] Sub-step 2-1: Design a multi-scale feature enhancement convolutional neural network module to extract the facial feature f from the preprocessed face image f_face . This module takes the face image as the input, introduces a multi-scale convolutional neural network to enhance the extraction of facial information at different scales. Different convolutional kernel sizes enable the module to capture information from details to the global, forming a richer facial feature representation, and output the final facial feature vector f f_face which is defined as:
[0105] f f_face = Multi-scaleCNN(Image face )
[0106] where CNN represents Convolutional Neural Network.
[0107] To extract the head pose feature vector f f_face from the facial feature vector f f_head , the inventor proposed a graph convolutional attention fusion module, which can adaptively extract head pose-related features and effectively suppress noise. The key points of the facial features learn weights through a two-layer graph convolutional network. The graph convolutional network (Graph Convolutional Networks, GCN) performs convolutional operations on the input facial key point graph to generate an enhanced local spatial information representation. Combine the output graph convolutional features with the facial residual network features to further extract more expressive pose features. Finally, use a multi-layer perceptron (Multi-Layer Perceptron, MLP) to take f f_head as the input to estimate the head pose. The estimated head pose α i ′ is defined as:
[0108] α i ′ = MLP(σ(A2f f_face W2 + σ(A1f f_face W1)) + f res )
[0109] where α i ′ represents the estimated head pose, σ represents the sigmoid activation function, A1 and A2 are the adjacency matrices of the graph convolutional network, W1 and W2 are the weight matrices of each layer of the graph convolutional network, and f res is the global facial feature extracted by the residual network.
[0110] In the head pose regression task, a loss function combining mean squared error and L2 regularization is designed to ensure the accuracy and stability of the model. The specific formula is defined as:
[0111]
[0112] where represents the true value of the head pose, λ is the regularization coefficient, W k is the k-th parameter in the learnable parameters, θ j is the j-th parameter in the MLP, m is the number of weight parameters W k and n is the number of parameters in the MLP.
[0113] Sub-step 2-2: Design a spatial shallow multi-stream network module based on the attention mechanism to extract fused eye features from the preprocessed binocular images This module takes the binocular images as inputs and passes them through the SW-GazeNet architecture respectively. This architecture contains multiple convolutional layers, where the first two layers are standard convolutional layers and the subsequent layers are depthwise separable convolutional layers. Batch normalization and ReLU activation are performed after each convolutional layer. Finally, it is further processed by the spatial mechanism to obtain the left-eye feature representation f l_eye and the right-eye feature representation f r_eye . Then, the self-correlation weights of the features are calculated through the parallel attention branch network, and the Sigmoid function is used to adjust the weight range. The feature vectors obtained from both eyes are multiplied by the corresponding weight vectors to obtain the weighted features f l ′ _eye and f r ′ _eye . The Softmax function is used to calculate the fusion weights ω l_eye and ω r_eye , and then weighted fusion is performed to obtain the final eye feature representation f eye ; finally, an MLP takes f eye as input to obtain the initial gaze direction vector g b′, the estimated initial gaze direction vector g b ′ is defined as:
[0114] g b ′ = MLP(Softmax((f l_eye ⊙A(f l_eye ))), (f r_eye ⊙A(f r_eye ))))
[0115] In the formula, g b ′ represents the initial gaze direction vector estimated by the network, MLP represents the multi - layer perceptron, Softmax is a normalization function, ⊙ represents element - wise multiplication, and A represents the attention mechanism network.
[0116] The inventor designed a comprehensive loss function. By introducing directional loss, smooth loss, and weight regularization loss, it not only ensures the accuracy of the predicted line - of - sight direction but also improves the prediction stability in time - series data and the generalization ability of the model, thus optimizing the overall prediction effect. The specific formula is defined as:
[0117]
[0118] In the formula, g * represents the true value of the line - of - sight direction of the input image, g t ′ represents the predicted line - of - sight direction vector at the t - th iteration, ω l_eye represents the attention weight vector of the left eye, ω r_eye represents the attention weight vector of the right eye, α represents the weight coefficient of the directional loss, β represents the weight coefficient of the smooth loss, γ represents the weight coefficient of the weight regularization loss, T max represents the total number of iterations.
[0119] Sub - step 2 - 3: Design a fusion module for eye feature vectors and head pose feature vectors based on the multi - head attention mechanism. First, concatenate the eye feature representation f eye and the head pose feature f f_head to obtain the concatenated feature representation f contact . Then, for each attention head of the concatenated feature representation f contact , calculate its corresponding Q i , K i and V i , and use the scaled dot - product attention mechanism to calculate the attention weights of each head. The weighted sum of these weights gives the output of each head. Secondly, concatenate the outputs of all heads together and obtain the final fusion feature through a linear transformation. Finally, use an MLP with f fusion as the input to obtain the gaze direction vector g b ′_adjust , the specific formula is defined as follows:
[0120] f contact = Concat(f eye , f f_head )
[0121]
[0122] g' b_adjust = MLP(f fusion )
[0123] where d k is the dimension size of Q i , K i , V i , and W o is the weight parameter of the linear projection.
[0124] In the gaze direction estimation task of fusing head pose, the loss function is specifically defined as:
[0125]
[0126] In step 3, to train the multi-modal based head pose corrected gaze direction estimation model, the inventor designed a full model adaptive optimization algorithm to optimize the loss function. Based on step 2, the total loss function is designed, and the loss weights of each task are dynamically and adaptively adjusted by introducing a dynamic weight adjustment mechanism, which is defined as:
[0127] Loss total = λ1(t)·Loss eye + λ2(t)·Loss head + λ3(t)·Loss adjust
[0128] In the formula, λ1(t), λ2(t), λ3(t) are the dynamic weights of each task at the t-th iteration, which are automatically adjusted according to the loss values of each task, and satisfy the condition λ1(t)+λ2(t)+λ3(t)=1. The formula of the dynamic weight adjustment mechanism is:
[0129]
[0130] In the formula, Loss i (t) represents the loss value of the i-th task at time step t, and T represents the temperature parameter.
[0131] By introducing an adaptive weight adjustment mechanism, the loss weights of each task are dynamically adjusted, the weight matrix is corrected, and the AdamW optimizer is combined for adaptive learning rate update. In each iteration, the weight matrix and bias terms are corrected through backpropagation, and combined with the adaptive weight adjustment mechanism, the minimum value of each loss function is solved by the gradient descent method. The optimization formula is as follows:
[0132]
[0133] In the formula, θ t is the weight matrix of the model, η is the adaptive learning rate, and are the momentum and squared momentum respectively, λ is the weight decay coefficient, and ε is a small constant used for numerical stability. Embodiment
[0134] An embodiment of the present invention proposes a multi-modal based head pose correction and gaze direction estimation method, including: Step 1, preprocessing the gaze estimation dataset and dividing it into a training set and a test set;
[0135] Use Dlib for face detection and key point localization. The detection includes 8 facial key points, namely 4 eye corners, 1 nose tip, 2 mouth corners, and 1 chin. The 2D and 3D face key points are matched through the projection formula, and weight coefficients are assigned to each key point. The preset weight of the eye corner is 0.2, the nose tip is preset to 0.3, the mouth corner is preset to 0.15, and the chin is preset to 0.2.
[0136] Normalize the images in the dataset and their corresponding three-dimensional gaze direction vectors. During the preprocessing of the face image and the eye image, the empirical value of d x is selected as 800mm. During the preprocessing of the face image, the focal lengths f x and f y are both set to 1500mm, and the principal points c x and c y are both set to 120. While during the preprocessing of the eye image, the focal lengths f x and f y are both set to 950mm, and the principal points c x and c y are set to 32 and 20 respectively.
[0137] The dataset used in the present invention is the publicly available dataset MPIIFaceGaze in the field of fixation point estimation, which contains approximately 37,667 images, divided into 30,133 training samples and 7,534 test samples.
[0138] Step 2: Construct a multi-modal based head pose correction and gaze direction estimation model;
[0139] The CNN-based residual face network module in the model takes a face image as input. First, it passes through a 5×5 convolutional layer (output channels: 64), and then through a 2×2 max pooling layer. Next, it goes through two 3×3 convolutional layers and a 1×1 convolutional layer. Subsequently, it passes through four dilated convolutional layers (kernel size 3×3, output channels 64, 64, 128, 256 respectively, dilation rates r1, r2, r3, r4). Batch normalization and ReLU activation are performed after each convolutional layer and dilated convolutional layer. Finally, it passes through a fully connected layer with 256 units to output the facial feature vector f f_face 。
[0140] The graph convolutional attention fusion module in the model adaptively extracts head pose features, combines the local spatial information generated by convolution and the residual network features, further enhances the pose feature representation, and finally uses an MLP for head pose estimation to obtain the estimated head pose α i ′. The MLP adopted has a three-layer structure, with the number of hidden units set to 128, and the initial weight matrix is set using the random initialization method.
[0141] The spatial shallow multi-stream network module based on the attention mechanism in the model takes the binocular image as input, and respectively obtains the left and right eye feature representations f l_eye and f r_eye through the SW-GazeNet architecture. This architecture contains 5 convolutional layers. The first two are standard convolutional layers (number of convolutional kernels: 32), and the latter three are depthwise separable convolutional layers (number of convolutional kernels 64 and 256 respectively). Batch normalization and ReLU activation are performed after each convolutional layer, and the output of the 5th convolutional layer is further processed by the spatial mechanism. The spatial mechanism part contains 3 convolutional layers, with the number of convolutional kernels all 256, the kernel size 3×3, and the stride set to 1 and 2. Finally, dimensionality reduction is performed through a 4×4 max pooling layer.
[0142] The parallel attention branch network in the model calculates the self-correlation weights of the features and uses the Sigmoid function to adjust the weight range. The feature vectors obtained from the binoculars are multiplied by the corresponding weight vectors to obtain the weighted feature sum, and the Softmax function is used to calculate the fusion weights ω l_eye and ω r_eye , and then weighted fusion is performed to obtain the final eye feature representation f eye . Finally, an MLP is used with f eye as input to obtain the estimated initial gaze direction vector g b ′. The number of hidden units of the adopted MLP is set to 128.
[0143] The eye feature vector and head pose feature vector fusion module based on the multi-head attention mechanism in the model fuses the eye feature representation f eye with the head pose feature f f_head to obtain the concatenated feature representation f contact . Then, the multi-head attention mechanism is used to calculate and weight-fuse the features. Finally, the outputs of all heads are concatenated to obtain the fusion result, and the fused gaze direction vector g b ′ _adjust is obtained through the MLP. The number of hidden units in the MLP is set to 128, the number of heads in the attention mechanism is 8, and the dimension size of each head is 64.
[0144] Step 3: Use the binocular images and face images in the training set as inputs to train the multi-modal head pose correction gaze direction estimation model;
[0145] The optimization algorithm selects the Adam algorithm with default parameter settings, and the initial learning rate is set to 0.001. The network is trained for a total of 50 epochs, the size of each batch is set to 64, and the dropout value of all fully connected layers is set to 0.5
[0146] Step 4, see the appendix of the specification Figure 4 , use the trained gaze direction estimation model to predict the test set data: input the test set data into the gaze direction estimation model trained offline to obtain the line-of-sight angle in the standardized space. Obtain the gaze direction vector in the standardized space Then convert it into a three-dimensional direction vector Then use the inverse matrix of the previous normalization formula to obtain the gaze direction vector in the original camera coordinate system During the evaluation process, the divided test set dataset is used for cross-validation. The test results show that the improved model has reached the current advanced level in gaze estimation performance. The comparison results of the average angle error of the improved model on the test set with other methods are shown in the appendix of the specification Figure 5 as shown.
[0147] From the appendix of the specification Figure 5 it can be seen that compared with the existing technology, the method of the present invention greatly reduces the angle error of the gaze direction and achieves superior technical effects, thus indicating that the method of the present invention is significantly superior to various existing technologies.
[0148] The above has elaborated on the content of the present invention through embodiments. It should be noted that the embodiments of the present invention are only used to explain the present invention, but do not constitute any limitation to the protection scope of the present invention. The protection scope of the present invention is determined by the claims in the application documents. Those skilled in the art, without departing from the spirit and essence of the present invention, by equivalently replacing the technical features in the present invention or changing other technical contents of the present invention, the technical solutions or embodiments obtained all fall within the protection scope of the present invention.
Claims
1. A multi-modal based head pose correction and gaze direction estimation method, characterized in that: The method includes the following steps: Step 1, preprocess the gaze estimation dataset and divide it into a training set and a test set; Step 2, construct a multi-modal head pose correction gaze direction estimation model; Step 3, use the binocular images and face images in the training set as inputs to train the multi-modal head pose correction gaze direction estimation model; Step 4, use the trained gaze direction estimation model to predict the test set data.
2. The multi-modal based head pose correction and gaze direction estimation method according to claim 1, wherein: Step 1 includes the following sub-steps: Sub-step 1-1, use the Dlib library for face detection and face key point localization, detect 8 facial key points, match the obtained 2D face key points with the corresponding points on the standard 3D face model through optimizing the projection point formula to obtain the key points of the 3D model, and use the center of the inner corner points of the two eyes on the 3D face as the origin g0 of the 3D line-of-sight direction. The optimized projection point formula is: where \(P\) represents the orthogonal projection matrix, \(R\) represents the rotation matrix, \(s\) represents the scaling factor, \(\sigma\) represents the activation function, \(\alpha\) i represents the shape vector coefficient of this face model, \(\beta\) i represents the texture vector coefficient of this face model, \(f(S\) i ) represents the non - linear transformation function of the shape vector, \(g(T\) i ) represents the non - linear transformation function of the texture vector, \(t\) 2d is the displacement matrix, \(m\) represents the number of shape vectors and texture vectors; Sub - step 1 - 2: Normalize the images in the training set and their corresponding three - dimensional gaze direction vectors. By introducing the weight coefficient \(\omega\) i , considering the importance of each key point, calculate the coordinates \(C\) of the face center in the pixel coordinate system f , defined as: where α is a non - linear adjustment factor, solve the rotation matrix R' such that after the original image is rotated by this rotation matrix, the face center C in the camera coordinate system f is aligned with the reference coordinate C' set in the image. The rotation matrix is defined as: where r j ′ k represents the element in the j - th row and k - th column of the obtained optimal rotation matrix R'. N represents the number of source points and target points. is the k - th component of the i - th source point, is the j - th component of the i - th target point. On the basis of optimizing the rotation matrix, further introduce the perspective transformation matrix M'. The affine transformation matrix M' is defined as: M′ = K'S'R'ΛK -1 In the formula, K' represents the internal parameter matrix of the virtual camera C', S′ represents the dynamic distance scaling matrix, and Λ represents the adjustment matrix; normalize the three-dimensional gaze direction vectors corresponding to the images in the training set, and further convert the 3D gaze vectors into 2D gaze vectors. The specific formula is as follows: g 2D = P·R'·g 3D where g 2D represents the normalized gaze direction vector, and g 3D represents the original gaze direction vector; Sub-step 1-3 proposes a regional gamma correction method for the normalized image, which is applied to the detected face and eye regions; the applied gamma correction formula is defined as: Where, I corrected (x, y) is the corrected pixel value at the position (x, y), and I original (x, y) is the original pixel value at the position (x, y), γ is the global gamma value, A face (x, y), B face (x, y), C face (x, y) are the adjustment parameters for the face and eyes; Sub-step 1-4, introducing multi-scale Gaussian blur with adaptive weights, dynamically adjusts the blur degree of different scales according to the texture features and variance information of the local area of the image, and its definition is as follows: Where, O(x, y) is the pixel value of the output image represented at the position, I(x, y) is the pixel value of the input image represented at the position, G(i, j, σ) represents the value of the Gaussian kernel at the position (x, y), σ = f(V(x + i, y + j)) is the standard deviation adjusted by the locality of the pixel, V(x + i, y + j) is the local characteristic of the pixel, n is the size of the filter, Z represents the normalization factor, Var k (x, y) represents the local variance at the position (x, y) at the k-th scale, and K represents the number of scales of the Gaussian kernel.
3. A multi-modal based head pose correction and gaze direction estimation method according to claim 2, characterized in that: The said step 2 includes the following sub-steps: sub-step 2-1, designing a multi-scale feature enhancement convolutional neural network module to extract facial feature f from the preprocessed face image f_face , taking the face image as the input and outputting the final facial feature vector f f_face which is defined as: f f_face = Multi-scaleCNN(Image face ) Among them, CNN represents Convolutional Neural Network, and adaptively extracts the head pose feature vector f from the facial feature vector f through a graph convolutional attention fusion module f_face and suppresses noise. The weights are learned through a two-layer graph convolutional network. The graph convolutional network performs convolution operations on the input facial key point map to generate an enhanced local spatial information representation, combines the output graph convolutional features with the facial residual network features, further extracts more expressive pose features, and finally uses a multi-layer perceptron MLP to use f f_head as the input to estimate the head pose, and the estimated head pose α f_head ' is defined as: i ' is defined as: α i ′ = MLP(σ(A2f f_face W2 + σ(A1f f_face W1)) + f res ) where α i ' represents the estimated head pose, MLP represents the multi-layer perceptron, σ represents the sigmoid activation function, A1 and A2 are the adjacency matrices of the graph convolutional network, W1 and W2 are the weight matrices of each layer of the graph convolutional network, and f res is the global facial feature extracted by the residual network; A loss function combining mean square error and L2 regularization is designed to ensure the accuracy and stability of the model, and the specific formula is defined as: wherein, represents the true value of the head pose, λ is the regularization coefficient, and W k is the k-th parameter among the learnable parameters, and θ j is the j-th parameter in the MLP. m is the number of weight parameters W k ; Sub-step 2-2: Design a spatial shallow multi-stream network module based on the attention mechanism to extract fused eye features from the processed binocular images Taking the binocular images as the input, the left and right eye feature representations f l_eye and f r_eye are obtained through the SW-GazeNet architecture respectively; the autocorrelation weights of the features are calculated through the parallel attention branch network, and the sigmoid activation function is used to adjust the weight range; the obtained feature vectors of the eyes are multiplied by the corresponding weight vectors to obtain the weighted features f l ′ _eye and f r ′ _eye . The fusion weights ω l_eye and ω r_eye are calculated using Softmax, and then the weighted fusion is performed to obtain the final eye feature representation f eye ; Then, an MLP is used with f eye as the input to obtain the initial gaze direction vector. The estimated initial gaze direction vector g b ′ is defined as: g b ′ = MLP(Softmax((f l_eye ⊙ A(f l_eye ))), (f r_eye ⊙ A(f r_eye )))) where, g b ' represents the initial gaze direction vector estimated by the network, MLP represents the multi-layer perceptron, Softmax is a normalization function, ⊙ represents element-wise multiplication, and A represents the attention mechanism network; a comprehensive loss function is designed to optimize the overall prediction effect by introducing directional loss, smooth loss, and weight regularization loss. The specific formula is defined as: where g * represents the true value of the line-of-sight direction of the input image, and g t ' represents the predicted line-of-sight direction vector at the t-th iteration, ω l_eye represents the attention weight vector of the left eye, ω r_eye represents the attention weight vector of the right eye, α represents the weight coefficient of the directional loss, β represents the weight coefficient of the smooth loss, γ represents the weight coefficient of the weight regularization loss, and T max represents the total number of iterations; Sub-step 2-3, design an eye feature vector and head pose feature vector fusion module based on the multi-head attention mechanism, and splice the eye feature representation f eye and the head pose feature f f_head to obtain the spliced feature representation f contact . Then, for each attention head of the spliced feature representation f contact , calculate its corresponding Q i , K i and V i , and use the scaled dot-product attention mechanism to calculate the attention weights of each head. The weighted sum of these weights gives the output of each head. Secondly, splice the outputs of all heads together and obtain the final fused feature through a linear transformation. Finally, use an MLP with f fusion as the input to obtain the gaze direction vector g b ' _adjust of the fused head pose. The loss function in the gaze direction estimation task of the fused head pose is specifically defined by the following formula:
4. A multi-modal based head pose correction and gaze direction estimation method according to claim 1, characterized in that: The specific content of Step 3 is as follows: Based on Step 2, design the total loss function, and dynamically and adaptively adjust the loss weights of each task by introducing a dynamic weight adjustment mechanism, which is defined as: Loss total = λ1(t)·Loss eye + λ2(t)·Loss head + λ3(t)·Loss adjust In the formula, λ1(t), λ2(t), and λ3(t) are the dynamic weights of each task at the t-th iteration, which are automatically adjusted according to the loss values of each task, and satisfy the condition λ1(t)+λ2(t)+λ3(t)=1. The formula of the dynamic weight adjustment mechanism is: where Loss i (t) represents the loss value of the i-th task at time step t, and T represents the temperature parameter Dynamically adjust the loss weights of each task by introducing an adaptive weight adjustment mechanism, and update the adaptive learning rate in combination with the AdamW optimizer; In each iteration, correct the weight matrix and bias term through backpropagation, and in combination with the adaptive weight adjustment mechanism, solve the minimum value of each loss function by the gradient descent method. The optimization formula is: Where, θ t is the weight matrix of the model, η is the adaptive learning rate, and are the momentum and squared momentum respectively, λ is the weight decay coefficient, and ε is a small constant used for numerical stability.
Citation Information
Cited By
Eye movement tracking method and system based on multi-modal fusion
CN121392947A
An eye movement tracking method and system based on multi-modal fusion
CN121392947B