Cross-domain sight line estimation method and device based on attention mechanism

By embedding the channel attention module in the deep learning network and combining the line of sight prediction and image reconstruction task branches, the problem of poor performance caused by the data distribution differences in the line of sight estimation is solved, and the generalization ability of the model is improved.

CN120014689APending Publication Date: 2025-05-16XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510050505.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Deep learning methods perform poorly in practical applications due to differences in data distribution between training and test data sets in line of sight estimation tasks.

Method used

The deep learning network model based on attention mechanism is adopted, and the ability to extract features related to the line of sight estimation task is improved by embedding two channel attention modules in the backbone network and constraining the two task branches of the line of sight prediction module and the image reconstruction module.

Benefits of technology

It effectively alleviates the problem of poor algorithm performance caused by data distribution differences, improves the generalization ability of network models, and makes it more stable in cross-domain line of sight estimation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014689A_ABST
    Figure CN120014689A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain sight line estimation method and device based on an attention mechanism. The cross-domain sight line estimation method comprises the following steps: acquiring a source domain sample data set and a target domain sample data set; constructing a deep learning network model based on an attention mechanism; inputting the source domain sample data set into a deep learning network model, and performing iterative training on the deep learning network model to obtain a trained deep learning network model; and inputting the target domain sample data set into the trained deep learning network model to obtain a sight line estimation result. On the basis of the backbone network, two channel attention modules are embedded, constraint is carried out through two task branches of the sight prediction module and the image reconstruction module, the network is helped to improve the ability of extracting features related to a sight estimation task, and thus the generalization ability of a network model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning and relates to a cross-domain sight line estimation method and device based on an attention mechanism. Background Art

[0002] Gaze estimation is to predict the gaze direction of a person by analyzing the person's eye or facial images. This technology has been widely used in many fields, especially in smart car cockpits, virtual reality (VR) / augmented reality (AR), medical health, user behavior analysis and other fields, playing an important role. With the development of computer vision and deep learning technology, more and more deep learning-based methods have been proposed for gaze estimation tasks. These methods automatically learn to extract features from images by training convolutional neural networks (CNNs) and predict gaze direction based on these features.

[0003] However, these deep learning methods are often trained and tested on the same dataset, so they can achieve good performance in experiments. But in practical applications, training and testing are often not on the same dataset. Only the training dataset is visible, while the test dataset is completely different from the training dataset and invisible. Due to the difference in data distribution between the training and test datasets, the algorithm usually performs poorly in cross-domain line of sight estimation. Summary of the invention

[0004] In order to solve the problem that the algorithm performs poorly in practical applications due to the difference in data distribution between the training data set and the test data set, the present invention provides a cross-domain sight line estimation method and device based on the attention mechanism. The technical solution adopted is:

[0005] A cross-domain sight line estimation method based on an attention mechanism comprises the following steps:

[0006] S1. Obtain source domain sample dataset and target domain sample dataset;

[0007] S2. Construct a deep learning network model based on attention mechanism;

[0008] S3, inputting the source domain sample data set into the deep learning network model, iteratively training the deep learning network model, and obtaining a trained deep learning network model;

[0009] S4. Input the target domain sample data set into the trained deep learning network model to obtain the line of sight estimation result.

[0010] In one embodiment of the present invention, step S1 comprises:

[0011] S11: Obtain the source domain sample dataset, the source domain sample dataset includes K1 face images of N1 different persons, each person corresponds to J1 different RGB images, where N1≥50, K1≥100000, and J1≥2000;

[0012] S12: Obtain the target domain sample data set, wherein the target domain sample data set includes K2 face images of N2 different persons, each person corresponds to J2 different RGB images, wherein N2≥20, K2≥10000, and J2≥500;

[0013] S13: performing normalization preprocessing on the source domain sample dataset and the target domain sample dataset.

[0014] In one embodiment of the present invention, the step S13 includes: adjusting the size of the RGB image to 224*224, and adjusting the pixel value of the RGB image from [0, 255] to [0, 1].

[0015] In one embodiment of the present invention, step S2 comprises:

[0016] S21: Connecting the backbone network to the first channel attention module and the second channel attention module respectively, so as to extract features from the input source domain sample data set or the target domain sample data set;

[0017] S22: Connecting the first channel attention module to the sight line prediction module to obtain an estimated sight line direction;

[0018] S23: Connect the second channel attention module to the image reconstruction module to obtain a reconstructed image.

[0019] In one embodiment of the present invention, step S3 comprises:

[0020] S31: Initialize the number of iterations to i, and the maximum number of iterations to i max ,i max ≥50, use the dataset loader to load the source domain sample dataset to obtain training samples, and let i=1;

[0021] S32: using the training sample as the input of the deep learning network model, and extracting features of the input image by the backbone network, the first channel attention module, and the second channel attention module, respectively inputting the extracted features into the sight line prediction module and the image reconstruction module to obtain an estimated sight line direction and a reconstructed image;

[0022] S33: The sight prediction module uses the mean absolute error function to construct a sight prediction loss function for the estimated sight direction and the real sight direction, and the image reconstruction module uses the mean square error function at the pixel level to construct an image reconstruction loss function for the reconstructed image and the original input image. The sight prediction loss function and the image reconstruction loss function constitute the overall loss function of the deep learning network model. Based on the overall loss function, back propagation is used to perform gradient descent to update the parameters of the deep learning network model.

[0023] S34: Determine i≥i max Is it true? If so, the trained deep learning network model is obtained. Otherwise, let i=i+1 and execute the step S32.

[0024] In one embodiment of the present invention, step S32 includes:

[0025] S321: The backbone network first extracts features from the input facial image to obtain original features F of the image;

[0026] S322: Input the original feature F into the first channel attention module for squeezing excitation operation, obtain the weight relationship W1 of each channel of the original feature F, expand the weight relationship W1 to the same size as the original feature F, and obtain W1 e , and perform a bitwise multiplication operation with the original feature F to obtain a processed feature F1 related to line of sight estimation, and the feature F1 related to line of sight estimation is used as the line of sight prediction branch M g The input is expressed as: Among them, F1 represents the features related to line of sight estimation, F represents the original features, and W1 e Represents the first channel weight relationship with the same size as the original feature F, It is Hadamard;

[0027] S323: Input the original feature F into the second channel attention module for squeezing excitation operation, obtain the weight relationship W2 of each channel of the original feature F, expand the weight relationship W2 to the same size as the original feature F, and obtain And perform negative and plus one operations to obtain the weight relationship after processing Then the weight relationship after processing The original feature F is subjected to a bitwise product operation to obtain a processed feature F2 related to image reconstruction, and the feature F2 related to image reconstruction is used as the image reconstruction branch M. r The input is expressed as: Among them, F2 represents the features related to image reconstruction, F represents the original features, represents the weight relationship after processing, J represents the all-1 matrix with the same size as the original feature F, Represents the second channel weight relationship with the same size as the original feature F, It is Hadamard.

[0028] In one embodiment of the present invention, in step S33, the overall loss function L total It is expressed as:

[0029] L total =L gaze +L re

[0030]

[0031] Among them, L total represents the overall loss function, L gaze represents the line of sight prediction loss function, L re represents the image reconstruction loss function, n is the number of images, g is the true line of sight direction, and g p is the predicted sight direction, h is the height pixel number of the image, w is the width pixel number of the image, and I ijk is the pixel value of the jth row and kth column in the i-th real input image, is the pixel value of the jth row and kth column in the i-th reconstructed image.

[0032] In one embodiment of the present invention, step S4 includes: inputting the target domain sample data set into the trained deep learning network model, and forward propagating through the backbone network, the first channel attention module and the line of sight prediction module in sequence to obtain a line of sight prediction result.

[0033] A cross-domain sight line estimation device based on an attention mechanism, comprising: a deep learning network model based on the attention mechanism;

[0034] The deep learning network model includes a backbone network, a first channel attention module, a second channel attention module, a sight line prediction module and an image reconstruction module;

[0035] The backbone network is connected to input ends of the first channel attention module and the second channel attention module respectively, the output end of the first channel attention module is connected to the sight line prediction module, and the output end of the second channel attention module is connected to the image reconstruction module;

[0036] The backbone network, the first channel attention module and the second channel attention module are used to extract features from input data;

[0037] The sight line prediction module is used to obtain the predicted sight line direction;

[0038] The image reconstruction module is used to obtain a reconstructed image.

[0039] In one embodiment of the present invention, the cross-domain sight line estimation device based on the attention mechanism further includes an acquisition module and a training module, and both the acquisition module and the training module are connected to the deep learning network model;

[0040] The acquisition module is used to acquire a source domain sample data set and a target domain sample data set, wherein the source domain sample data set is used as training data and the target domain sample data set is used as test data;

[0041] The training module receives the training data for training the deep learning network model;

[0042] The test data is input into the trained deep learning network model to obtain the line of sight estimation result.

[0043] Beneficial effects of the present invention:

[0044] The cross-domain line of sight estimation method and device based on the attention mechanism of the present invention embeds two channel attention modules on the basis of the backbone network, and constrains the two task branches of the line of sight prediction module and the image reconstruction module to help the network improve the ability to extract features related to the line of sight estimation task, thereby effectively alleviating the problem of poor performance of the algorithm in practical applications due to the difference in data distribution between the training data set and the test data set, and improving the generalization ability of the network model. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flow chart of a cross-domain sight line estimation method based on an attention mechanism provided by an embodiment of the present invention;

[0046] Figure 2 It is a structural diagram of a deep learning network model based on the attention mechanism provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0048] The present invention provides a cross-domain sight line estimation method based on attention mechanism. Figure 1 , including the following steps:

[0049] S1. Obtain source domain sample dataset and target domain sample dataset;

[0050] S2. Construct a deep learning network model based on attention mechanism;

[0051] S3, inputting the source domain sample data set into the deep learning network model, iteratively training the deep learning network model, and obtaining a trained deep learning network model;

[0052] S4. Input the target domain sample data set into the trained deep learning network model to obtain the line of sight estimation result.

[0053] The source domain and target domain datasets are two different datasets. In actual production, the source domain dataset is the dataset we have accumulated through historical collection, while the target domain dataset is the unforeseen unknown data encountered in actual production. The deep learning network model of the present invention is trained on the source domain dataset, but tested on the target domain dataset.

[0054] In one embodiment of the present invention, step S1 comprises:

[0055] S11: Obtain a source domain sample dataset, which contains K1 face images of N1 different people, each person corresponds to J1 different RGB images, where N1≥50, K1≥100000, and J1≥2000;

[0056] S12: Obtain a target domain sample data set, which contains K2 face images of N2 different persons, each person corresponds to J2 different RGB images, where N2≥20, K2≥10000, and J2≥500;

[0057] S13: Perform normalization preprocessing on the source domain sample dataset and the target domain sample dataset. Resize the RGB image to 224*224, and adjust the pixel value of the RGB image from [0,255] to [0,1].

[0058] In one embodiment of the present invention, step S2 includes:

[0059] S21: connecting the backbone network to the first channel attention module and the second channel attention module respectively, so as to extract features of the input source domain sample data set or the target domain sample data set;

[0060] S22: Connecting the first channel attention module to the sight line prediction module to obtain an estimated sight line direction;

[0061] S23: Connect the second channel attention module to the image reconstruction module to obtain a reconstructed image.

[0062] The feature extraction part of the Resnet network is used as the backbone network M b , followed by two independent SE channel attention modules and and Connect the sight line prediction branch M respectively g and image reconstruction branch M r The Resnet network is selected as the Resnet50 network, and the convolutional layer and maximum pooling layer of Resnet50 are set as the backbone network M b , the remaining last average pooling layer and fully connected layer of Resnet50 are set as the sight prediction branch M g Image reconstruction branch M r It consists of multiple upsampling modules, Resnet's BasicBlock basic residual block, and convolutional layers.

[0063] Step S3 of the present invention comprises:

[0064] S31: Initialize the number of iterations to i, and the maximum number of iterations to i max ,i max ≥50, the network model obtained in the i-th iteration is M i , use the dataset loader to load the source domain sample dataset to obtain training samples, let i = 1;

[0065] S32: The training sample is used as the input of the deep learning network model M, and the backbone network M b , First channel attention module and the second channel attention module The input image is feature extracted, and the extracted features F1 and F2 are respectively input into the sight prediction module M g and image reconstruction module M r , get the estimated sight direction g p and the reconstructed image I p ;

[0066] S33: Line of sight prediction module M g The estimated sight direction g is calculated using the mean absolute error function. p And the real line of sight direction g constructs the line of sight prediction loss function, the image reconstruction module M r The reconstructed image I is reconstructed using the mean square error function at the pixel level. p The image reconstruction loss function is constructed with the original input image I. The sight prediction loss function and the image reconstruction loss function constitute the overall loss function L of the deep learning network model M. total , based on the overall loss function L total , use back-propagation to perform gradient descent to update the parameters of the deep learning network model M;

[0067] The overall loss function L total It is expressed as:

[0068] L total =Lgaze +L re

[0069]

[0070]

[0071] Among them, L total represents the overall loss function, L gaze represents the line of sight prediction loss function, L re represents the image reconstruction loss function, n is the number of images, g is the true line of sight direction, and g p is the predicted sight direction, h is the height pixel number of the image, w is the width pixel number of the image, and I ijk is the pixel value of the jth row and kth column in the i-th real input image, is the pixel value of the jth row and kth column in the i-th reconstructed image.

[0072] S34: Determine i≥i max Is it true? If so, the trained deep learning network model is obtained. Otherwise, let i=i+1 and execute step S32.

[0073] Step S32 of the present invention includes:

[0074] S321: Backbone network M b First, extract the features of the input facial image I to obtain the original features F of the image;

[0075] S322: Input the original feature F into the first channel attention module Perform a squeeze excitation operation in the original feature F to obtain the weight relationship W1 of each channel, and expand the weight relationship W1 to the same size as the original feature F to obtain W1 e , and perform a bitwise multiplication operation with the original feature F to obtain a processed feature F1 related to line of sight estimation, and the feature F1 related to line of sight estimation is used as the line of sight prediction branch M g The input is expressed as: Among them, F1 represents the features related to line of sight estimation, F represents the original features, and W1 e Represents the first channel weight relationship with the same size as the original feature F, It is Hadamard;

[0076] S323: Input the original feature F into the second channel attention module Perform a squeeze excitation operation in the original feature F to obtain the weight relationship W2 of each channel, and expand the weight relationship W2 to the same size as the original feature F to obtain And perform negative and plus one operations to obtain the weight relationship after processing Then the weight relationship after processing The original feature F is multiplied bit by bit to obtain the processed feature F2 related to image reconstruction. The feature F2 related to image reconstruction is used as the image reconstruction branch M. r The input is expressed as: Among them, F2 represents the features related to image reconstruction, F represents the original features, represents the weight relationship after processing, J represents the all-1 matrix with the same size as the original feature F, Represents the second channel weight relationship with the same size as the original feature F, It is Hadamard.

[0077] Step S4 includes: inputting the target domain sample data set into the trained deep learning network model, and forward propagating through the backbone network, the first channel attention module and the sight line prediction module in sequence to obtain the sight line prediction result.

[0078] The present invention also provides a cross-domain sight line estimation device based on the attention mechanism, referring to the attached Figure 2 The cross-domain sight line estimation device includes a deep learning network model based on the attention mechanism. The deep learning network model of the present invention includes a backbone network, a first channel attention module, a second channel attention module, a sight line prediction module and an image reconstruction module.

[0079] The backbone network is connected to the input ends of the first channel attention module and the second channel attention module respectively, the output end of the first channel attention module is connected to the line of sight prediction module, and the output end of the second channel attention module is connected to the image reconstruction module; the backbone network, the first channel attention module and the second channel attention module are used to extract features of the input data; the line of sight prediction module is used to obtain the predicted line of sight direction; the image reconstruction module is used to obtain the reconstructed image.

[0080] The cross-domain sight line estimation device of the present invention also includes an acquisition module and a training module, both of which are connected to the deep learning network model. The acquisition module is used to acquire a source domain sample data set and a target domain sample data set, the source domain sample data set is used as training data, and the target domain sample data set is used as test data. The training module receives the training data for training the deep learning network model. The test data is input into the trained deep learning network model to obtain the sight line estimation result.

[0081] The cross-domain line of sight estimation device of the present invention includes a deep learning network model based on the attention mechanism. On the basis of the backbone network, a channel attention module is embedded, and constraints are performed through two task branches, namely the line of sight prediction module and the image reconstruction module, to help the network improve its ability to extract features related to the line of sight estimation task, thereby improving the generalization ability of the network model.

[0082] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.

Claims

1. A cross-domain sight line estimation method based on attention mechanism, characterized in that: Includes steps: S1. Obtain source domain sample dataset and target domain sample dataset; S2. Construct a deep learning network model based on attention mechanism; S3, inputting the source domain sample data set into the deep learning network model, iteratively training the deep learning network model, and obtaining a trained deep learning network model; S4. Input the target domain sample data set into the trained deep learning network model to obtain the line of sight estimation result.

2. The cross-domain sight line estimation method based on the attention mechanism according to claim 1, characterized in that: The step S1 comprises: S11: Obtain the source domain sample dataset, the source domain sample dataset includes K1 face images of N1 different persons, each person corresponds to J1 different RGB images, where N1≥50, K1≥100000, and J1≥2000; S12: Obtain the target domain sample data set, wherein the target domain sample data set includes K2 face images of N2 different persons, each person corresponds to J2 different RGB images, wherein N2≥20, K2≥10000, and J2≥500; S13: performing normalization preprocessing on the source domain sample dataset and the target domain sample dataset.

3. The cross-domain sight line estimation method based on the attention mechanism according to claim 2 is characterized in that: The step S13 includes: adjusting the size of the RGB image to 224*224, and adjusting the pixel value of the RGB image from [0, 255] to [0, 1].

4. The cross-domain sight line estimation method based on the attention mechanism according to claim 1, characterized in that: The step S2 comprises: S21: Connecting the backbone network to the first channel attention module and the second channel attention module respectively, so as to extract features from the input source domain sample data set or the target domain sample data set; S22: Connecting the first channel attention module to the sight line prediction module to obtain an estimated sight line direction; S23: Connect the second channel attention module to the image reconstruction module to obtain a reconstructed image.

5. The cross-domain sight line estimation method based on the attention mechanism according to claim 4 is characterized in that: The step S3 comprises: S31: Initialize the number of iterations to i, and the maximum number of iterations to i max ,i max ≥50, use the dataset loader to load the source domain sample dataset to obtain training samples, and let i=1; S32: using the training sample as the input of the deep learning network model, and extracting features of the input image by the backbone network, the first channel attention module, and the second channel attention module, respectively inputting the extracted features into the sight line prediction module and the image reconstruction module to obtain an estimated sight line direction and a reconstructed image; S33: The sight prediction module uses the mean absolute error function to construct a sight prediction loss function for the estimated sight direction and the real sight direction, and the image reconstruction module uses the mean square error function at the pixel level to construct an image reconstruction loss function for the reconstructed image and the original input image. The sight prediction loss function and the image reconstruction loss function constitute the overall loss function of the deep learning network model. Based on the overall loss function, back propagation is used to perform gradient descent to update the parameters of the deep learning network model. S34: Determine i≥i max Is it true? If so, the trained deep learning network model is obtained. Otherwise, let i=i+1 and execute the step S32.

6. The cross-domain sight line estimation method based on attention mechanism according to claim 5, characterized in that: The step S32 comprises: S321: The backbone network first extracts features from the input facial image to obtain original features F of the image; S322: Input the original feature F into the first channel attention module for squeezing excitation operation, obtain the weight relationship W1 of each channel of the original feature F, expand the weight relationship W1 to the same size as the original feature F, and obtain W1 e , and perform a bitwise multiplication operation with the original feature F to obtain a processed feature F1 related to line of sight estimation, and the feature F1 related to line of sight estimation is used as the line of sight prediction branch M g The input is expressed as: Among them, F1 represents the features related to line of sight estimation, F represents the original features, and W1 e Represents the first channel weight relationship with the same size as the original feature F, It is Hadamard; S323: Input the original feature F into the second channel attention module for squeezing excitation operation, obtain the weight relationship W2 of each channel of the original feature F, expand the weight relationship W2 to the same size as the original feature F, and obtain And perform negative and plus one operations to obtain the weight relationship after processing Then the weight relationship after processing The original feature F is subjected to a bitwise product operation to obtain a processed feature F2 related to image reconstruction, and the feature F2 related to image reconstruction is used as the image reconstruction branch M. r The input is expressed as: Among them, F2 represents the features related to image reconstruction, F represents the original features, represents the weight relationship after processing, J represents the all-1 matrix with the same size as the original feature F, Represents the second channel weight relationship with the same size as the original feature F, It is Hadamard.

7. The cross-domain sight line estimation method based on attention mechanism according to claim 5, characterized in that: In step S33, the overall loss function L total It is expressed as: L total =L gaze +L re Among them, L total represents the overall loss function, L gaze represents the line of sight prediction loss function, L re represents the image reconstruction loss function, n is the number of images, g is the true line of sight direction, and g p is the predicted sight direction, h is the height pixel number of the image, w is the width pixel number of the image, and I ijk is the pixel value of the jth row and kth column in the i-th real input image, is the pixel value of the jth row and kth column in the i-th reconstructed image.

8. The cross-domain sight line estimation method based on attention mechanism according to claim 4, characterized in that: The step S4 includes: inputting the target domain sample data set into the trained deep learning network model, and performing forward propagation through the backbone network, the first channel attention module and the sight line prediction module in sequence to obtain the sight line prediction result.

9. A cross-domain sight line estimation device based on attention mechanism, characterized in that: include: Deep learning network model based on attention mechanism; The deep learning network model includes a backbone network, a first channel attention module, a second channel attention module, a sight line prediction module and an image reconstruction module; The backbone network is connected to input ends of the first channel attention module and the second channel attention module respectively, the output end of the first channel attention module is connected to the sight line prediction module, and the output end of the second channel attention module is connected to the image reconstruction module; The backbone network, the first channel attention module and the second channel attention module are used to extract features from input data; The sight line prediction module is used to obtain the predicted sight line direction; The image reconstruction module is used to obtain a reconstructed image.

10. The cross-domain sight line estimation device based on attention mechanism according to claim 9, characterized in that: It also includes an acquisition module and a training module, both of which are connected to the deep learning network model; The acquisition module is used to acquire a source domain sample data set and a target domain sample data set, wherein the source domain sample data set is used as training data and the target domain sample data set is used as test data; The training module receives the training data for training the deep learning network model; The test data is input into the trained deep learning network model to obtain the line of sight estimation result.