Cross-view-angle sight line estimation method based on feature decoupling and attention mechanism

By using a cross-view angle method of feature decoupling and attention mechanism in line of sight estimation, and using two-view face images for feature separation and fusion, the problem of limited vision field of single-view image is solved, and the accuracy and robustness of line of sight estimation are significantly improved.

CN120088839AActive Publication Date: 2025-06-03BEIJING INST OF TECH

Patent Information

Application Number
CN202510085404.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-03
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

Existing appearance-based vision estimation methods mainly rely on single-view images, resulting in limited field of view and difficulty in capturing facial information in full, which may lead to depth blur and ambiguity in line of sight estimation.

Method used

The cross-view line of sight estimation method based on feature decoupling and attention mechanism is adopted. Through the dual-view face image, the feature decoupling module is used to separate the appearance features, head posture features and line of sight features, and adaptive feature fusion is carried out through the cross-attention feature fusion module to achieve the supplement and enhancement of line of sight information.

Benefits of technology

It significantly improves the accuracy and robustness of line of sight estimation, effectively eliminates blind spots in single-view perspectives, alleviates the ambiguity caused by insufficient information, and ensures high performance in diverse scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088839A_ABST
    Figure CN120088839A_ABST
Patent Text Reader

Abstract

The invention provides a cross-view-angle sight line estimation method based on feature decoupling and an attention mechanism, double-view-angle face images are collected through double cameras, a view angle blind area under a single view angle is effectively eliminated, and the ambiguity problem caused by insufficient single-view-angle information is relieved; face features are decoupled into personal appearance features, head posture features and sight line related features, and accurate separation of different types of features is achieved; the adaptive capacity of the model in diversified scenes is enhanced through a feature decoupling strategy, and it is ensured that the model still keeps high-performance under the conditions of different individuals and large-amplitude head posture changes; then, adaptive weighting processing is carried out on the three types of features obtained through decoupling by applying a cross attention mechanism, and supplement and enhancement of sight line information are realized through fusion of double-view-angle sight line features; meanwhile, the influence of personal factors and head movement is compensated by combining the processed appearance features and head posture features, so that the precision and robustness of sight line estimation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a cross-view gaze estimation method based on feature decoupling and attention mechanism. Background Technique

[0002] With the development of technologies such as image processing and computer vision, gaze estimation has gradually become an important research direction in the field of human-computer interaction and has been widely applied in fields such as virtual / augmented reality, assisted driving, and psychological research. The core goal of gaze estimation technology is to infer the user's gaze direction by analyzing the user's eye or face image, so as to achieve natural interaction between the user and the system.

[0003] Traditional gaze estimation methods are mainly divided into model-based methods and appearance-based methods. The model-based method can provide high accuracy by constructing and fitting an eyeball model, but usually requires the user to wear specific devices, which not only increases the hardware cost but also may cause interference to the user, limiting its practical application. In contrast, the appearance-based method can complete gaze estimation only relying on an ordinary network camera by learning the relationship between the image and the gaze direction, without additional hardware configuration, with low cost and broader application prospects, and has gradually become the mainstream method of gaze estimation.

[0004] Currently, most appearance-based gaze estimation methods rely on single-view images for gaze direction estimation. However, due to the limited field of view of single-view images, it is usually difficult to comprehensively capture the complete information of the face, which may lead to problems such as depth ambiguity and ambiguity in gaze estimation. Summary of the Invention

[0005] To solve the above problems, the present invention provides a cross-view gaze estimation method based on feature decoupling and attention mechanism. By combining the rich information provided by dual-view face images, different types of features are effectively separated through feature decoupling, and the attention mechanism is used to learn the relationship between each feature to achieve adaptive feature fusion, thereby significantly improving the accuracy of gaze estimation.

[0006] A cross-view gaze estimation method based on feature decoupling and attention mechanism inputs two sequences of normalized face images under dual views into a cross-view gaze estimation network, and the cross-view gaze estimation network outputs a gaze direction vector of the face; wherein, the cross-view gaze estimation network includes a backbone network for face feature extraction, a feature decoupling module, first to third cross-attention feature fusion modules, a feature splicing module, and a multi-layer perceptron;

[0007] The backbone network for face feature extraction is used to extract features from the sequences of normalized face images under dual views, and respectively obtain the overall face features f 1 and f2 ;

[0008] The feature decoupling module decouples the features of f 1 and f 2 respectively, and correspondingly obtains the appearance features and head pose features and gaze features and

[0009] The first cross-attention feature fusion module is used to fuse the gaze features and The fusion result is then added to the original gaze features and for enhancement to obtain the final fused gaze features and

[0010] The second cross-attention feature fusion module is used to combine the gaze features and to perform dynamic weighting on the appearance features and respectively, to obtain the weighted appearance features and

[0011] The third cross-attention feature fusion module is used to combine the gaze features and to perform dynamic weighting on the head pose features and respectively, to obtain the weighted head pose features and

[0012] The feature splicing module is used to splice the fused gaze features and appearance features and head pose features and under each camera view respectively, to obtain the spliced features and

[0013] The multi-layer perceptron is used to learn the spliced features and to predict the gaze direction vector of the human face and

[0014] Furthermore, the loss function L used when training the cross-view gaze estimation network model total is as follows:

[0015] L total = α·L gt + β·L app_sim + ε·L head + γ·L ga

[0016] where L gt is the angular deviation loss function, α is the weight corresponding to L gt , L app_sim is the appearance similarity loss function, β is the weight corresponding to L app_sim , L head is the head angle loss function, ε is the weight corresponding to L head , and L ga is the cross-view gaze alignment loss function, γ is the weight corresponding to L ga .

[0017] Furthermore, the calculation method of the appearance similarity loss function L app_sim is as follows:

[0018]

[0019] where B is the total number of human faces used as training samples, is the cosine similarity between the appearance features of the i-th human face under dual views and , and the calculation method of is as follows:

[0020]

[0021] where ||·|| represents taking the modulus.

[0022] Furthermore, the calculation method of the head angle loss function L head is as follows:

[0023]

[0024] where B is the total number of human faces used as training samples, is the true value of the head direction vector of the i-th human face, is the estimated value of the head direction vector obtained by inputting the head pose feature of the i-th human face into the head pose feature multi-layer perceptron, and ||·|| represents taking the modulus.

[0025] Furthermore, the method for obtaining the true value of the head direction vector of any human face is:

[0026] Detect a face image through Dlib, extract 68 facial key points, combine with a pre-trained 3D average head model, use the EPnP algorithm to perform 3D fitting on the face image, and calculate the rotation matrix of the head in the face image relative to the standard camera coordinate system; use the rotation matrix to convert the three-axis direction vectors of the head coordinate system to the standard camera coordinate system, and perform unitization processing on the head z-axis direction vector converted to the standard camera coordinate system, and use the unitized head z-axis direction vector as the true value of the head direction vector of the face.

[0027] Further, the cross-view gaze alignment loss function L ga is calculated as follows:

[0028]

[0029] where B is the total number of faces used as training samples, is the initial 3D gaze direction obtained by inputting the gaze feature f of the i-th face at the first view into the multi-layer perceptron, 1(i) gaze the initial 3D gaze direction obtained by inputting the gaze feature f of the i-th face at the second view into the multi-layer perceptron, is the gaze feature f of the i-th face at the second view, 2(i) gaze the initial 3D gaze direction obtained by inputting the gaze feature f of the i-th face at the second view into the multi-layer perceptron, is the denormalization rotation matrix of the i-th face from the first camera normalized coordinate system to the first camera original coordinate system, is the denormalization rotation matrix of the i-th face from the second camera normalized coordinate system to the second camera original coordinate system, is the rotation matrix of the first camera coordinate system relative to the global screen coordinate system, is the rotation matrix of the second camera coordinate system relative to the global screen coordinate system.

[0030] Further, the angular deviation loss function L gt is calculated as follows:

[0031]

[0032] where B is the total number of faces used as training samples, is the true value of the gaze direction vector of the i-th face, is the concatenated feature f corresponding to the i-th face, fused the estimated value of the gaze direction vector obtained by inputting the concatenated feature f corresponding to the i-th face into the multi-layer perceptron, ||·|| is to find the modulus.

[0033] Further, the first cross-attention feature fusion module is used to fuse the gaze features and perform feature fusion, and then fuse the result with the original gaze feature and Add them to enhance and obtain the final fused gaze feature and The method is as follows:

[0034] Convert the gaze feature in the original coordinate system of the first camera to the original coordinate system of the second camera to obtain the gaze feature after perspective transformation

[0035] Use the multi-head cross-attention mechanism to adaptively fuse the gaze feature with the gaze feature to obtain the initial fused feature

[0036] Add the initial fused feature to the gaze feature to obtain the final fused gaze feature

[0037] Convert the gaze feature in the original coordinate system of the second camera to the original coordinate system of the first camera to obtain the gaze feature after perspective transformation

[0038] Use the multi-head cross-attention mechanism to adaptively fuse the gaze feature with the gaze feature to obtain the initial fused feature

[0039] Add the initial fused feature to the gaze feature to obtain the final fused gaze feature

[0040] Furthermore, the method for the second cross-attention feature fusion module to obtain the weighted appearance feature and is as follows:

[0041] Let k = 1, 2, map the gaze feature to the same dimension as the appearance feature to obtain the mapped gaze feature Use the cross-attention component to perform dynamic weighting on the appearance feature and the mapped gaze feature to obtain the weighted appearance feature

[0042] Furthermore, the third cross-attention feature fusion module obtains the weighted head pose feature and The method is specifically as follows:

[0043] Let k = 1, 2, and map the gaze feature to the same dimension as the head pose feature to obtain the mapped gaze feature Use the cross-attention component to perform dynamic weighting on the head pose feature and the mapped gaze feature to obtain the weighted head pose feature

[0044] Beneficial effects:

[0045] The present invention provides a cross-view gaze estimation method based on feature decoupling and attention mechanism. By collecting dual-view face images with dual cameras, it effectively eliminates the view blind area in a single view and alleviates the ambiguity problem caused by insufficient single-view information. By setting different loss functions, the face features are decoupled into personal appearance features, head pose features, and gaze-related features, realizing the precise separation of different types of features. The feature decoupling strategy enhances the adaptability of the model in diverse scenarios, ensuring its high-performance performance under different individuals and head pose changes. Then, the cross-attention mechanism is respectively applied to the three types of decoupled features for adaptive weighting processing. Through the fusion of dual-view gaze features, the supplementation and enhancement of gaze information are realized. At the same time, by combining the processed appearance features and head pose features, the influence of personal factors and head movement is compensated, thereby significantly improving the accuracy and robustness of gaze estimation. Brief Description of the Drawings

[0046] Figure 1 is the overall network architecture diagram of the present invention;

[0047] Figure 2 is the schematic diagram of the DenseNet feature extraction module used in the present invention;

[0048] Figure 3 is the conversion relationship diagram between the coordinate systems of the present invention;

[0049] Figure 4(a) is a visualization diagram showing the distribution of the three types of features extracted by the present invention's embodiment using t-SNE when the model is iterated 100 times;

[0050] Figure 4(b) is a visualization diagram showing the distribution of the three types of features extracted by the present invention's embodiment using t-SNE when the model is iterated 1500 times;

[0051] Figure 5 is the visualization diagram of the gaze estimation result of the present invention's embodiment. Detailed Embodiments

[0052] To enable those skilled in the art to better understand the solution of this application, the technical solution in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.

[0053] With the continuous development of camera devices, the cost of cameras has gradually decreased, and dual-camera systems have gradually become a viable option. Compared with a single view, a dual-view gaze estimation system can provide a broader field of view, effectively making up for the information loss caused by facial occlusion in single-view images, thereby reducing the estimation error brought by view limitations. This dual-view gaze estimation method shows great potential in improving the accuracy and robustness of gaze estimation and has become an important research direction in appearance-based gaze estimation methods.

[0054] With the increase in view information, how to effectively fuse features from different views has become a new important challenge. Feature decoupling technology can effectively reduce the interference between features by decomposing complex feature vectors into multiple independent sub-features, thereby improving the clarity of feature expression. In the process of multi-view information fusion, feature decoupling helps to independently model each type of feature, avoid information confusion, and thus improve the accuracy and robustness of gaze estimation. Combining the decoupled features with the attention mechanism can further strengthen the correlation between different features and enhance the accuracy and robustness of gaze estimation.

[0055] Based on this, the present invention provides a cross-view gaze estimation method based on feature decoupling and attention mechanism, which inputs two sequences of normalized face images under dual views into a cross-view gaze estimation network, and the cross-view gaze estimation network outputs a gaze direction vector of the face; wherein, the cross-view gaze estimation network includes a face feature extraction backbone network, a feature decoupling module, first to third cross-attention feature fusion modules, a feature splicing module, and a multi-layer perceptron; as Figure 1 shown, the data processing process of the cross-view gaze estimation network is as follows:

[0056] S1: The face feature extraction backbone network is used to extract features from the sequences of normalized face images under dual views, and respectively obtain the overall face features f 1 and f 2 ;

[0057] It should be noted that the present invention can obtain any two sequences of normalized face videos V 1 , V 2 synchronously collected by cameras from a publicly available gaze estimation dataset. Resample the videos from the two views at the same frame rate, such as a frame rate of 10 Hz, and adjust the size of the sampled images to 256×256 pixels, so as to generate corresponding N-frame sequences of normalized face images, denoted as and The normalized face image sequences under the dual perspectives are and input into the DenseNet face feature extraction backbone network in sequence, and the face features f 1 and f 2 .

[0058] For example, different labels are assigned to the normalized RGB face image sequences under the two perspectives, and they are respectively input into the DenseNet feature extraction network with a BatchSize of 128, and the input dimension is (128, 3, 256, 256). After the feature extraction by the backbone network, the overall face features f 1 and f 2 under the dual perspectives are obtained respectively.

[0059] Among them, the feature extraction backbone network adopts the DenseNet dense connection network, and the network structure is as Figure 2 shown. The DenseNetEncoder extracts low-level features through the initial convolutional layer, and then extracts higher-level features layer by layer through 3 dense blocks. Each dense block concatenates the output of the previous layer with the features of the current layer, the growth rate is 32, LeakyReLU is used as the activation function, and InstanceNorm2d is used as the normalization method.

[0060] S2: The feature decoupling module decouples f 1 and f 2 respectively, and correspondingly obtains the appearance features and the head pose features and the gaze features and

[0061] It should be noted that it is assumed that the face feature vector f consists of three parts: gaze-related features, head pose features, and personal appearance features. Let f[1:a] be the a-dimensional personal appearance feature f app , f[a + 1:a + b] be the b-dimensional head pose feature f head , and f[a + b + 1:a + b + c] be the c-dimensional gaze-related feature f gaze . For example, the first 42 dimensions of the feature vector f are set as the personal appearance feature f app , the middle 36 dimensions of the feature are set as the head pose feature f head , and the last 132 dimensions are set as the gaze-related feature f gaze .

[0062] S3: The first cross-attention feature fusion module is used to fuse the gaze features and Perform feature fusion, and then add the fusion result to the original line-of-sight feature and respectively for enhancement to obtain the final fused line-of-sight feature and Specifically as follows:

[0063] S31: Transform the line-of-sight feature in the original coordinate system of the first camera to the original coordinate system of the second camera to obtain the line-of-sight feature after perspective transformation

[0064]

[0065] That is to say, in the present invention, the line-of-sight feature in the original coordinate system of the first camera C 1 is successively left-multiplied by the inverse normalization rotation matrix M of the first camera 1 -1 , the rotation matrix between the x c1 -y c1 -z c1 of the first camera coordinate system and the x s -y s -z s of the global screen coordinate system the rotation matrix between the x s -y s -z s of the global screen coordinate system and the x c2 -y c2 -z c2 of the second camera coordinate system the normalization rotation matrix M from the original coordinate system of the second camera to the x m2 -y m2 -z m2 of the normalized coordinate system of the second camera 2 , and complete its conversion to the normalized coordinate system of the second camera C 2 . The overall conversion process is as Figure 3 shown. It should be noted that M 1 , M 2 are 3x3 normalization matrices obtained from the EVE dataset to convert the two camera coordinate systems to the corresponding standardized camera coordinate systems respectively; is a 3x3 rotation matrix for converting from the two camera coordinate systems to the global screen coordinate system.

[0066] S32: Use the multi-head cross-attention mechanism to adaptively fuse the line-of-sight feature with the line-of-sight feature to obtain the initial fusion feature from the second perspective

[0067] Among them, the multi - head cross - attention mechanism is implemented through the nn.MultiheadAttention layer, with a 132 - dimensional feature input and 4 attention heads for calculation. This module uses the gaze feature after perspective transformation as the query, and the gaze feature of the target perspective as the key and value for cross - perspective attention calculation. The weighted gaze feature f 2 gaze_att The implementation principle is as follows:

[0068]

[0069] In the above formula, Q 1 is the gaze feature after perspective transformation K 1 and V 1 are the gaze features under the second camera perspective The feature dimension d k1 = 132 / 4 = 33, corresponding to the feature dimension of the key vector of each attention head.

[0070] S33: Add the initial fusion feature to the gaze feature to achieve the supplementary enhancement of the gaze feature and obtain the final fusion gaze feature under the second perspective

[0071]

[0072] S34: Convert the gaze feature in the original coordinate system of the second camera to the original coordinate system of the first camera to obtain the gaze feature after perspective transformation

[0073] S35: Use the multi - head cross - attention mechanism to adaptively fuse the gaze feature with the gaze feature to obtain the initial fusion feature

[0074] S36: Add the initial fusion feature to the gaze feature to obtain the final fusion gaze feature

[0075] S4: The second cross - attention feature fusion module is used to combine the gaze feature and to perform dynamic weighting on the appearance features and respectively, to obtain the weighted appearance features and Specifically as follows:

[0076] Let k = 1, 2, map the gaze feature to the same dimension as the appearance feature to obtain the mapped gaze feature Use the cross-attention component of 2 attention heads to combine the mapped gaze feature to perform dynamic weighting on the appearance feature to obtain the weighted appearance feature The calculation formula is as follows:

[0077]

[0078] In the above formula, Q 2 is the mapped gaze feature K 2 and V 2 are the appearance feature f app , and the feature dimension d k2 = 42 / 2 = 21.

[0079] S5: The third cross-attention feature fusion module is used to combine the gaze feature and to perform dynamic weighting on the head pose feature and respectively to obtain the weighted head pose feature and Specifically as follows:

[0080] Let k = 1, 2, map the gaze feature to the same dimension as the head pose feature to obtain the mapped gaze feature Use the cross-attention component of 2 attention heads to combine the mapped gaze feature to perform dynamic weighting on the head pose feature to obtain the weighted head pose feature The calculation formula is specifically as follows:

[0081]

[0082] In the above formula, Q 3 is the mapped gaze feature K 3 and V 3 are the head pose features The feature dimension d k3 = 36 / 4 = 9.

[0083] S6: The feature splicing module is used to splice the fused gaze feature and the appearance feature and head pose feature and Perform feature stitching according to the viewing angle, and let k = 1, 2 to obtain the stitched feature f k fused :

[0084]

[0085] S7: The multi-layer perceptron is used to process the stitched feature f k fused for learning to predict the gaze direction vector of the face under dual viewing angles

[0086]

[0087] That is to say, the present invention uses the method of feature stitching to combine the fused line-of-sight features processed by the attention mechanism and appearance features and head pose features and are fused according to the camera viewing angle to obtain 210-dimensional fused line-of-sight features The fused features are processed by a three-layer MLP, and the hidden layer sizes are 105 and 52 in sequence to obtain the three-dimensional gaze direction vector under the dual camera viewing angles

[0088] It should be noted that the present invention uses the corresponding training set and validation set in the public multi-view gaze estimation dataset (EVE, ETH-XGaze) to train and test the cross-view gaze estimation network model. The ratio of the training set to the validation set is 10:1. At the same time, the loss function L used during the training of the cross-view gaze estimation network model total is:

[0089] L total = α·L gt + β·L app_sim + ε·L head + γ·L ga

[0090] where L gt is the angular deviation loss function, α is the weight corresponding to L gt , L app_sim is the appearance similarity loss function, β is the weight corresponding to L app_sim , L head is the head angle loss function, ε is the weight corresponding to L head , L ga is the cross-view gaze alignment loss function, γ is the weight corresponding to Lga The corresponding weight. It should be noted that through multiple experimental verifications, the angle deviation loss L gt has a weight α = 1, the appearance similarity loss L app_sim has a weight β = 0.5, the cross-view gaze alignment loss L ga has a weight γ = 0.3, and the head pose loss L head has a weight ε = 0.5.

[0091] Furthermore, assuming that at the same moment, the personal appearance features of the same person under different camera views are the same, the dual-view appearance features should have a similar distribution in the feature space. The appearance similarity loss function is used to constrain the dual-view appearance features and to make them tend to be similar. The cosine similarity and between the appearance features of the i-th face under the dual views is calculated as follows:

[0092]

[0093] where ||·|| is to find the modulus.

[0094] The arccosine function is used to convert the cosine similarity into an angle, and the average value of the angle losses of all samples is calculated to obtain the appearance similarity loss function L app_sim as follows:

[0095]

[0096] where B is the total number of faces used as training samples.

[0097] Furthermore, the present invention inputs the head pose features and into a three-layer head feature multi-layer perceptron (MLP) with hidden layer sizes of 18 and 9 respectively. ReLU is used as the activation function for each layer, and the feature extraction results of the first two layers are normalized using BatchNorm1d to obtain the estimated value of the three-dimensional head direction vector, and the true value h 3d of the head direction vector is used as the true value label for supervision, then the head angle loss function L head is obtained as follows:

[0098]

[0099] where B is the total number of faces used as training samples, is the true value of the head direction vector of the i-th face, is the estimated value of the head direction vector obtained by inputting the head pose feature of the i-th person into the head pose feature multi-layer perceptron, and ||·|| is for calculating the modulus. Among them, the method for obtaining the true value of the head direction vector of any face is as follows:

[0100] Detect the normalized face image through Dlib and extract 68 facial key points. Combine with the pre-trained 3D average head model, and use the EPnP algorithm to perform three-dimensional fitting on the face image to calculate the rotation matrix of the head in the face image relative to the standard camera coordinate system; use the rotation matrix to convert the three-axis direction vectors of the head coordinate system to the standard camera coordinate system, and perform unitization processing on the head z-axis direction vector converted to the standard camera coordinate system, and use the unitized head z-axis direction vector as the true value of the head direction vector of the face.

[0101] Furthermore, the cross-view line-of-sight alignment loss function L ga is calculated as follows:

[0102]

[0103] Among them, B is the total number of faces used as training samples, is the line-of-sight feature of the i-th face in the first view and is the initial three-dimensional gaze direction obtained by inputting into a three-layer line-of-sight feature multi-layer perceptron with hidden layer sizes of 66 and 33 respectively, is the line-of-sight feature of the i-th face in the second view and is the initial three-dimensional gaze direction obtained by inputting into a three-layer line-of-sight feature multi-layer perceptron with hidden layer sizes of 66 and 33 respectively, is the denormalization rotation matrix of the i-th face from the first camera normalized coordinate system to the first camera original coordinate system, is the denormalization rotation matrix of the i-th face from the second camera normalized coordinate system to the second camera original coordinate system, is the rotation matrix of the first camera coordinate system relative to the global screen coordinate system, is the rotation matrix of the second camera coordinate system relative to the global screen coordinate system.

[0104] That is to say, the present invention uses the image normalization matrices M 1 and M 2 provided by the dataset to left-multiply the initial estimation vectors in the dual-view by M 1 -1 and M 2 -1 respectively for denormalization processing, and convert them to the original coordinate systems x c1 -y c1 -z c1 of the two cameras, xc2 -y c2 -z c2 down; then, multiply each of them by the rotation matrix between the original coordinate systems of the two cameras and the global screen coordinate system x s -y s -z s to achieve the transformation of the dual-view gaze direction to the same coordinate system. The transformation of the dual-view gaze direction to the same coordinate system is realized.

[0105] Furthermore, the present invention uses the ground-truth label value of the gaze direction in the dataset to supervise the three-dimensional gaze direction estimated by the network, and the angular deviation loss function L gt is calculated as follows:

[0106]

[0107] where B is the total number of human faces used as training samples, is the true value of the gaze direction vector of the i-th human face, is the estimated value of the gaze direction vector obtained by inputting the concatenated feature f corresponding to the i-th human face into the multi-layer perceptron, and ||·|| is the modulus calculation. fused The estimated value of the gaze direction vector obtained by inputting the concatenated feature f corresponding to the i-th human face into the multi-layer perceptron, and ||·|| is the modulus calculation.

[0108] It should be noted that the overall network model proposed in the present invention is implemented based on the PyTorch framework, and the parameter initialization adopts the Xavier method. During the training process, the batch size is set to 128, the optimizer is selected as Adam, the initial learning rate is 0.001, and the weight decay value is 0.0005. The learning rate adopts the CosineAnnealingLR scheduling strategy, the total number of training iterations is 4000 times, and the model parameters are saved every 400 iterations for subsequent verification and testing.

[0109] The present invention uses the Mean Angular Error (MAE) to measure the estimation effect of the model on the validation set. The calculation formula of MAE is as follows:

[0110]

[0111] where N is the total number of samples in the test set, and g are the predicted gaze vector and the true gaze vector respectively. The smaller the MAE value, the higher the estimation accuracy of the model.

[0112] The performance results of the trained cross-view gaze estimation model on the validation sets of two gaze estimation datasets are shown in Table 1. The gaze angle estimation errors of the network model proposed in the present invention on the EVE and ETH-XGaze validation sets are 2.57° and 2.69° respectively. Compared with the multi-view gaze estimation models MMGE and MGT, and the full-face gaze estimation models FullFace and GazeCLR, it has higher gaze estimation accuracy, which proves that the network model proposed in the present invention has good performance.

[0113] Table 1 Gaze Estimation Angle Error Results (unit: °)

[0114] Figures 4(a) and 4(b) respectively use t-SNE to visualize the distributions of three types of features extracted by this model at the 100th and 1500th training times. Among them, the personal appearance feature f app , the head pose feature f head and the gaze-related feature f gaze are represented by red, yellow, and blue dots respectively. By comparing the distribution of feature points in Figures 4(a) and 4(b), it can be seen that as the training progresses, different types of features are gradually separated more clearly and compactly in the feature space, which proves that the feature decoupling effect of this model is good and can effectively improve the expression ability of different types of features.

[0115] Figure 5 Shows the convergence curves of each loss function of the network on the training set. As shown in Figure 4(b), the losses of network training quickly tend to the minimum value after 400 iterations on the training set, and then slowly decrease and finally converge. Among them, the angle deviation loss L gt converges to 1.36°, the appearance similarity loss L app_sim converges to 0.057, the head pose loss converges to 2.05°, and the cross-view gaze alignment loss L ga converges to 1.11°, which fully verifies the effectiveness of the network architecture and its good convergence performance.

[0116] To verify the effectiveness of the design of each module in the cross-view gaze estimation network, ablation experiments were carried out on two gaze estimation datasets, EVE and ETH-XGaze. Table 2 shows the angle errors of 6 different ablation experiments. The experimental results show that each module in the network of the present invention has played a positive role in improving the model performance.

[0117] Table 2 Ablation Experiment Results (unit: °)

[0118]

[0119] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can certainly make various corresponding changes and modifications according to the present invention. However, these corresponding changes and modifications should all fall within the protection scope of the appended claims of the present invention.

Claims

1. A cross-view line of sight estimation method based on feature decoupling and attention mechanism, characterized in that: Input two normalized face image sequences under dual perspectives into a cross-perspective sight line estimation network, and the cross-perspective sight line estimation network outputs a gaze direction vector of the face; wherein the cross-perspective sight line estimation network includes a face feature extraction backbone network, a feature decoupling module, a first to third cross-attention feature fusion module, a feature splicing module, and a multi-layer perceptron; The face feature extraction backbone network is used to extract features from the normalized face image sequence under dual viewing angles, and obtain the overall face features f under dual viewing angles respectively. 1 and f 2 ; The feature decoupling module respectively 1 and f 2 Perform feature decoupling to obtain the appearance features under dual perspectives and Head posture characteristics and Sight characteristics and The first cross-attention feature fusion module is used to combine the sight line features and Perform feature fusion, and then combine the fusion result with the original sight feature and Add and enhance to obtain the final fusion sight feature and The second cross-attention feature fusion module is used to combine the line of sight features and Appearance characteristics and Perform dynamic weighting processing to obtain weighted appearance features and The third cross-attention feature fusion module is used to combine the line of sight features and Head pose features and Perform dynamic weighted processing to obtain weighted head posture features and The feature stitching module is used to merge the sight features under each camera perspective and Appearance characteristics and Head posture characteristics and Perform feature splicing separately to obtain splicing features and The multi-layer perceptron is used to concatenate features and Learn to predict the gaze direction vector of a face and 2. The cross-view line of sight estimation method based on feature decoupling and attention mechanism according to claim 1, characterized in that: The loss function L used when training the cross-view line of sight estimation network model is total for: L total =α·L gt +β·L app_sim +e·L head +γ·L ga Among them, L gt is the angle deviation loss function, α is L gt The corresponding weight, L app_sim is the appearance similarity loss function, β is L app_sim The corresponding weight, L head is the head angle loss function, ε is L head The corresponding weight, L ga is the cross-view alignment loss function, γ is L ga The corresponding weight.

3. The cross-view line of sight estimation method based on feature decoupling and attention mechanism as claimed in claim 2, characterized in that: Appearance similarity loss function L app_sim The calculation method is: Among them, B is the total number of faces used as training samples, is the appearance feature of the i-th face under dual view and The cosine similarity between The calculation method is: Among them, ||·|| is the modulus.

4. The cross-view line of sight estimation method based on feature decoupling and attention mechanism as claimed in claim 2, characterized in that: Head angle loss function L head The calculation method is: Among them, B is the total number of faces used as training samples, is the true value of the head direction vector of the i-th face, The head direction vector estimate obtained by inputting the head pose feature multilayer perceptron into the i-th face head pose feature, and ||·|| is the modulus.

5. The cross-view line of sight estimation method based on feature decoupling and attention mechanism according to claim 4, characterized in that: The method for obtaining the true value of the head direction vector of any face is: Use Dlib to detect face images and extract 68 facial key points. Combined with the pre-trained 3D average head model, use the EPnP algorithm to perform 3D fitting on the face image and calculate the rotation matrix of the head in the face image relative to the standard camera coordinate system. The rotation matrix is ​​used to transform the three-axis direction vectors of the head coordinate system into the standard camera coordinate system, and the head z-axis direction vector transformed into the standard camera coordinate system is normalized, and the normalized head z-axis direction vector is used as the true value of the head direction vector of the face.

6. The cross-view line of sight estimation method based on feature decoupling and attention mechanism according to claim 2, characterized in that: Cross-view line of sight alignment loss function L ga The calculation method is: Among them, B is the total number of faces used as training samples, is the sight feature f of the i-th face in the first viewing angle 1 (i) gaze Input the line of sight feature multi-layer perceptron to obtain the initial three-dimensional gaze direction, is the sight feature f of the i-th face in the second perspective 2(i) gaze Input the line of sight feature multi-layer perceptron to obtain the initial three-dimensional gaze direction, is the denormalized rotation matrix of the ith face from the normalized coordinate system of the first camera to the original coordinate system of the first camera, is the denormalized rotation matrix of the i-th face from the normalized coordinate system of the second camera to the original coordinate system of the second camera, is the rotation matrix of the first camera coordinate system relative to the global screen coordinate system, is the rotation matrix of the second camera coordinate system relative to the global screen coordinate system.

7. The cross-view line of sight estimation method based on feature decoupling and attention mechanism according to claim 2, characterized in that: Angular deviation loss function L gt The calculation method is: Among them, B is the total number of faces used as training samples, is the true value of the gaze direction vector of the i-th face, is the splicing feature f corresponding to the i-th face fused The estimated gaze direction vector is input into the multilayer perceptron, and ||·|| is the modulus.

8. The cross-view line of sight estimation method based on feature decoupling and attention mechanism according to claim 1, characterized in that: The first cross-attention feature fusion module is used to combine the sight line features and Perform feature fusion, and then combine the fusion result with the original sight feature and Add and enhance to obtain the final fusion sight feature and The specific method is: The sight line feature in the original coordinate system of the first camera Convert to the original coordinate system of the second camera to obtain the sight line features after the perspective conversion Use multi-head cross attention mechanism to integrate line of sight features With sight characteristics Perform adaptive fusion to obtain the initial fusion features The initial fusion features With sight characteristics Add together to get the final fusion sight feature The sight line feature of the second camera's original coordinate system Convert to the original coordinate system of the first camera to obtain the sight line features after the perspective conversion Use multi-head cross attention mechanism to integrate line of sight features With sight characteristics Perform adaptive fusion to obtain the initial fusion features The initial fusion features With sight characteristics Add together to get the final fusion sight feature 9. The cross-view line of sight estimation method based on feature decoupling and attention mechanism according to claim 1, characterized in that: The second cross-attention feature fusion module obtains the weighted appearance features and The specific method is: Let k = 1, 2, and transform the sight line feature Mapping to appearance features The same dimension, get the mapping line of sight features Using cross-attention components to focus on appearance features and mapping sight line features Perform dynamic weighting processing to obtain weighted appearance features 10. The cross-view line of sight estimation method based on feature decoupling and attention mechanism according to claim 1, characterized in that: The third cross-attention feature fusion module obtains the weighted head posture features and The specific method is: Let k = 1, 2, and transform the sight line feature Mapped to head pose features The same dimension, get the mapping line of sight features Head pose features using cross-attention components and mapping sight line features Perform dynamic weighted processing to obtain weighted head posture features

Citation Information

Patent Citations

  • Line-of-sight direction determination method for double-camera system

    CN116311485A

  • Double-view-angle fixation estimation method and system based on far eye

    CN118865476A

  • Student fixation point estimation method based on double cameras in classroom scene

    CN119007269A

Cited By

  • Differential guidance sight line estimation method and device based on 6D rotation matrix characterization

    CN120496153A

  • User generated content detection method and device, equipment and medium

    CN120980274A