E2StyleGAN network-based sight line estimation method

Through the E2StyleGAN network-based line of sight estimation method, the source domain conversion and line of sight perception operation are used to extract and fuse the line of sight related features, the problem of the reduction in accuracy in the real world environment is solved, and better cross-domain performance and generalization capabilities are achieved.

CN119919976APending Publication Date: 2025-05-02HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998126.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Existing line-of-sight estimation techniques reduce accuracy when dealing with complex real-world environments, making it difficult to deal with various unexpected factors, especially in the absence of target data, domain generalization methods are more challenging.

Method used

Using the E2StyleGAN network-based line of sight estimation method, the line of sight related features are extracted and fused through source domain conversion and line of sight perception operations, the line of sight is reduced, and the advantageous line of sight features are selected in the latent code.

Benefits of technology

The cross-domain performance and generalization ability of line of sight estimation are improved, the generalization performance of the model is enhanced, the cross-domain effect is better, and the direction of sight can be more accurately estimated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919976A_ABST
    Figure CN119919976A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of line-of-sight estimation, and discloses an E2StyleGAN network-based line-of-sight estimation method, which comprises the following steps of: through source domain conversion of an E2StyleGAN network, mapping an input image into a source domain to obtain a potential code of the image, and adding a line-of-sight distortion loss training encoder; the potential code processing module finds features related to sight through sight perception operation and sends the features to the generator to reconstruct a source domain image; the sight line estimation model carries out feature fusion by using features learned by a source domain image and sight line related features obtained by potential code processing, and a sight line vector is calculated. According to the line-of-sight estimation method based on the E2StyleGAN, the input image is processed, so that the model has better generalization performance and better cross-domain effect; potential information of the image is obtained through inversion of the GAN, and sight line related information is selected through sight line perception operation to assist in sight line estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sight line estimation, and in particular to a sight line estimation method based on an E2StyleGAN network. Background Art

[0002] Gaze estimation is one of the important research directions in the field of computer vision. It contains rich information about human intentions and can provide a deep understanding of human cognition and behavior. It has been widely used in medicine, assisted driving, marketing, and human-computer interaction. With the development of deep learning technology using convolutional neural networks (CNNs), appearance-based gaze estimation technology has made significant progress.

[0003] In recent years, various unconstrained datasets have been proposed with wide ranges of gaze and head pose, but their improvements are limited to fixed environments. The accuracy of gaze estimation is reduced due to the fact that real-world datasets contain various invisible conditions, such as different personal appearances, lighting, and background environments. This shows that factors unrelated to the line of sight can lead to unexpected mapping relationships, i.e., overfitting. This makes it difficult for the model to handle various unexpected factors in a constantly changing environment. To address this issue, domain adaptation methods that use a small number of target samples in training and domain generalization methods that only use source domain data perform better than traditional methods.

[0004] In the real world, domain adaptation methods trained with a small number of target samples are not always applicable, since target data for adaptation is usually not available. Therefore, domain generalization methods that only use source domain data are more effective in practical applications. In particular, these problems are more challenging in line of sight estimation, since the dimensionality of line of sight-independent features is much larger than that of line of sight-dependent features, which are crucial for line of sight estimation.

[0005] In order to solve the problem of difficult line of sight feature extraction, people have proposed a variety of operation methods to extract and control image attributes. Miraz et al. trained to create conditional images to control specific attributes of the image and obtain related images. Subsequently, Abdal et al. analyzed three semantic editing operations that can be applied to latent space vectors. Shen et al. adopted a data-driven approach and used principal component analysis to learn the most important directions. Park et al. proposed a simple and effective method to conditionally continuously normalize the flow in the GAN latent space based on attribute features, providing a basis for operating semantic features in the latent space to extract line of sight related features. Summary of the invention

[0006] The purpose of the present invention is to provide a line of sight estimation method based on the E2StyleGAN network, which processes the input image to make the model more generalized and have better cross-domain effect; the potential information of the image is obtained through GAN inversion, and the line of sight perception operation is used to select information related to the line of sight to assist in line of sight estimation.

[0007] To achieve the above object, the present invention provides a line of sight estimation method based on an E2StyleGAN network, comprising the following steps:

[0008] Step S1, through the source domain conversion of the E2StyleGAN network, the input image is mapped to the source domain to obtain the potential code of the image; in the process of domain conversion, the line of sight distortion loss is added to train the encoder;

[0009] Step S2: The latent code processing module finds the features related to the line of sight through the line of sight perception operation, and feeds them into the generator of E2StyleGAN to reconstruct the source domain image;

[0010] Step S3: The sight estimation model uses the features learned from the source domain image and the sight-related features obtained from the latent code processing to perform feature fusion and calculate the sight vector (θ prd , );

[0011] Among them, the overall framework of a sight line estimation method based on the E2StyleGAN network consists of three parts: latent code processing module, source domain conversion module, and sight line estimation module.

[0012] Preferably, in step S1, the input image is mapped to the source domain to obtain the potential code of the image through source domain conversion of the E2StyleGAN network; in the domain conversion process, the line of sight distortion loss is added to train the encoder, and the specific process is as follows:

[0013] Step S11, source domain conversion of E2StyleGAN;

[0014] Given a labeled source domain D S , the goal is to learn a translation model f, which only accesses the source domain data D S , the image is transferred from the target domain D T Mapping to source domain D S ; The E2StyleGAN inversion method is used to invert the real image using its out-of-distribution generalization ability. The source domain conversion process is as follows:

[0015]

[0016] Among them, F represents the feature extractor, S * Represents a domain converter;

[0017] Step S12, the E2StyleGAN encoder training process includes two stages. First, the gaze-related extractor is trained to estimate the gaze direction from the facial image; second, the gaze distortion loss is used to train the gaze-related extractor to estimate the gaze direction from the facial image. To train the encoder, is the difference between the result of generating the image F(G(E(I))) and the line of sight extractor g. The specific process is as follows:

[0018] Step S121: train the extractor F combined with ResNet-18 and MLP to estimate the gaze direction from the face image, and use the result of the image F(G(E(l))) generated from F to calculate the gaze distortion loss As shown below:

[0019]

[0020] Among them, E is the encoder, G is the generator, I is the input image, and g is the three-dimensional real line of sight angle;

[0021] Step S122: The encoder uses three types of losses, namely, the least square error L2, the perceptual loss L LPIPS , similarity loss L sim Based on this, add the line of sight distortion loss For training, the expression of total loss is as follows:

[0022] L total (I) = λ l2 L2(I)+λ LPIPS L LPIPS (I)+λ sim L sim (I)+λ GD L GD (I) (3);

[0023] Among them, λ l2 ,λ LPIPS ,λ sim ,λ GD is the weight of each loss;

[0024] The encoder and extractor F are trained in an end-to-end manner so that F is optimized to generate images that preserve gaze-related features.

[0025] Preferably, RetinaFace is used to capture facial images and crop the images to reduce interference inputs caused by different environments.

[0026] Preferably, in step S2, the latent code is processed, and the specific process is as follows:

[0027] The inversion of E2StyleGAN uses the line-of-sight related attributes in the input image to invert the target image into the latent space. By discovering the code of the line-of-sight related attributes, the manipulated latent vector is ensured to be highly correlated with the line-of-sight vector. The calculation formula of the inversion is as follows:

[0028]

[0029] Where f represents the optimal gaze estimator, X represents the latent vector, x represents the element of the latent vector X, and d represents the dimension of the latent space.

[0030] Preferably, in step S2, the line of sight perception operation is performed, and the specific process is as follows:

[0031] Step S211: Divide the data set into two groups G L and G R ; G L represents a group of sight vectors with yaw axis directions from 30 degrees to 90 degrees, G R A group representing a sight line vector having a yaw axis direction of -30 degrees to -90 degrees;

[0032] Step S212, calculating the mean of each group of overall potential codes;

[0033] Step S213: Since adjacent tensors represent similar features, the 16 tensors are regarded as a single block and defined as the unit of statistical operation. The calculation process is as follows:

[0034]

[0035] Among them, C i represents the i-th data block, k represents the number of data blocks in the potential code, and j represents the number of elements in the data block; C i and C j Respectively represent G L Group and G R blocks in a group;

[0036] The blocks were sorted in descending order, the block most associated with sight was selected, and then a paired t-test was performed on the two expectations for each block from the two groups;

[0037] Step S214: Paired t-test of two expected values. The specific process is as follows:

[0038] The sample sizes of the two groups are large enough. According to the central limit theorem, the distribution of the sample mean difference is approximately Gaussian, and the variance of each sample is approximately the population variance, as shown below:

[0039]

[0040]

[0041] Among them, H0 represents the null hypothesis, H1 represents the alternative hypothesis, and Represent the overall average values ​​of the two groups on the i-th block, d represents the dimension of the latent space, T represents the test statistic, and Respectively and The sample statistic, n L and n R Respectively represent G L and G R The number of, σ represents the standard deviation;

[0042] Among them, since the indices of sight-related blocks in different datasets are different, the channel attention layer CA-layer is used to improve the cross-domain generalization performance.

[0043] Therefore, the present invention adopts the above-mentioned sight line estimation method based on the E2StyleGAN network, and the beneficial effects are as follows:

[0044] (1) To address the problem of cross-domain performance degradation of the model, the E2StyleGAN network structure is introduced to convert images from the target domain to the source domain, and then the source domain image is reconstructed using line of sight related features, thereby improving the performance of line of sight estimation across data sets.

[0045] (2) To address the difficulty in extracting gaze features, a gaze perception operation based on a data statistics-driven method is performed to select favorable gaze features for gaze estimation in the latent code.

[0046] (3) Aiming at the possible line of sight distortion that may occur during the domain conversion process, the line of sight distortion loss is introduced, and the encoder of E2StyleGAN is retrained to reduce the line of sight distortion between the input image and the generated image.

[0047] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is the overall framework of a sight line estimation method based on the E2StyleGAN network of the present invention;

[0049] Figure 2 is the training process of the encoder of the present invention;

[0050] Figure 3 This is the line of sight perception operation process of the present invention. DETAILED DESCRIPTION

[0051] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.

[0052] like Figure 1 As shown in the figure, the overall framework of a sight line estimation method based on the E2StyleGAN network consists of three parts: a potential code processing module, a source domain conversion module, and a sight line estimation module.

[0053] Based on this, the present invention provides a sight line estimation method based on the E2StyleGAN network, and the specific implementation process is as follows:

[0054] Step S1, through the source domain conversion of the E2StyleGAN network, the input image is mapped to the source domain to obtain the potential code of the image; in the process of domain conversion, the line of sight distortion loss is added to train the encoder so that the input image does not lose line of sight related information during the conversion process.

[0055] Step S11: Source domain conversion of E2StyleGAN.

[0056] The goal of source domain conversion is to generalize the distribution of E2StyleGAN inversion, use the encoder to correctly map the target image to the latent space, and obtain the latent code of the image, so as to extract the line of sight related features through line of sight perception operation, and then use the obtained line of sight related features to reconstruct the source domain image through the generator.

[0057] It is very difficult to extract line of sight features from the entire image because it is necessary to extract line of sight related features from the complex image space of the real scene. In order to minimize the interference input of the image, RetinaFace is used to capture the face image and crop the image to significantly reduce the interference input caused by different environments. Since the generative space of each attribute trained in the source domain is larger, the generator can reconstruct the image in a semantically aligned manner.

[0058] Given a labeled source domain D S , the goal is to learn a translation model f, which only accesses the source domain data D S , the image is transferred from the target domain D T Mapping to source domain D S The E2StyleGAN inversion method is used to invert real images using its out-of-distribution generalization ability. These images are not generated in the same way as the source domain training data. The source domain conversion process is as follows:

[0059]

[0060] Among them, F represents the feature extractor, S * Represents a domain converter.

[0061] Step S12: E2StyleGAN encoder training.

[0062] like Figure 2 As shown in Figure 2, the encoder training process consists of two stages: first, training the gaze-related extractor to estimate the gaze direction from the face image; second, using the gaze distortion loss To train the encoder, is the difference between the result of generating the image F(G(E(I))) and the line of sight extractor g. The specific process is as follows:

[0063] Step S121: train the extractor F combined with ResNet-18 and MLP to estimate the gaze direction from the face image, and use the result of the image F(G(E(l))) generated from F to calculate the gaze distortion loss As shown below:

[0064]

[0065] Among them, E is the encoder, G is the generator, I is the input image, and g is the 3D real view angle.

[0066] Step S122: The encoder uses three types of losses, namely, the least square error L2, the perceptual loss L LPIPS , similarity loss L sim Based on this, add the line of sight distortion loss For training, the expression of total loss is as follows:

[0067] L total (I) = λ l2 L2(I)+λ LPIPS L LPIPS (I)+λ sim L sim (I)+λ GD L GD (I) (3);

[0068] Among them, λ l2 ,λ LPIPS ,λ sim ,λ GD is the weight of each loss.

[0069] The present invention trains the encoder and the extractor F in an end-to-end manner during training, so that F is optimized to generate an image that retains sight-related features.

[0070] Step S2: The latent code processing module finds the features related to the line of sight through the line of sight perception operation, and feeds them into the generator of E2StyleGAN to reconstruct the source domain image.

[0071] Step S21: Potential code processing.

[0072] The inversion of E2StyleGAN can selectively exploit the gaze-related attributes in the input image and invert the target image to the latent space, so that the code related to the gaze-related attributes can be discovered. The goal of inversion is to ensure that the manipulated latent vector is highly correlated with the gaze vector. Therefore, it is necessary to find the manipulation operator H to determine which elements are related to the gaze information in the latent vector.

[0073] The inversion calculation formula is as follows:

[0074]

[0075] Where f represents the optimal gaze estimator, X represents the latent vector, x represents the element of the latent vector X, and d represents the dimension of the latent space.

[0076] Step S21: Line of sight perception operation.

[0077] Step S211: Divide the data set into two groups G L and G R . G L Represents a group of sight vectors with yaw axis directions between 30 and 90 degrees. R Represents a group of sight vectors with yaw axis directions from -30 degrees to -90 degrees.

[0078] Step S212: Calculate the mean of the overall latent code for each group. Since attributes unrelated to line of sight (such as lighting, character appearance, background environment) are averaged out, the grouped latent codes have the same value except for the elements related to line of sight. It can be inferred that elements with large differences in values ​​between groups are significantly associated with line of sight features.

[0079] Step S213: Since adjacent tensors represent similar features, the 16 tensors are regarded as a single block and defined as the unit of statistical operation. The calculation process is as follows:

[0080]

[0081] Among them, C i represents the i-th data block, k represents the number of data blocks in the potential code, and j represents the number of elements in the data block; C i and C j Respectively represent G L Group and G R The blocks were sorted in descending order to select the blocks most relevant to the sight line, and then a paired t-test was performed on the two expectations for each block from the two groups.

[0082] Step S214: Paired t-test of two expected values. The specific process is as follows:

[0083] The sample sizes of the two groups are large enough. According to the central limit theorem, the distribution of the sample mean difference is approximately Gaussian, and the variance of each sample is approximately the population variance, as shown below:

[0084]

[0085]

[0086] Among them, H0 represents the null hypothesis, H1 represents the alternative hypothesis, and Represent the overall average values ​​of the two groups on the i-th block, d represents the dimension of the latent space, T represents the test statistic, and Respectively and The sample statistic, n L and n R Respectively represent G L and G R is the number of , and σ represents the standard deviation.

[0087] The above sight perception operation process is as follows Figure 3 As shown in the figure, some data blocks belonging to the critical area are located in channels 4 and 5. However, since the indices of the sight-related blocks in different datasets are slightly different, a channel attention layer (CA-layer) is used to improve the cross-domain generalization performance.

[0088] Step S3: The sight estimation model uses the features learned from the source domain image and the sight-related features obtained from the latent code processing to perform feature fusion and calculate the sight vector

[0089] Example

[0090] 1. Dataset and experimental environment.

[0091] This example uses the public Gaze360 and MPIIFaceGaze datasets for model training and testing. The MPIIFaceGaze dataset provides 45,000 facial images from 15 research subjects, and the Gaze360 dataset contains 172,000 images of 238 subjects in indoor and outdoor environments.

[0092] The experimental environment uses the Windows 10 operating system, the processor is Intel (R) Xeon (R) Silver 4210R CPU @ 2.40GHz, the computer running memory is 32G, and the graphics card is RTX3080.

[0093] 2. Experimental preprocessing.

[0094] This embodiment first preprocesses the dataset, detects the face area in the image through RetinaFace, and performs consistent cropping of the two datasets based on the face area, reducing the input except the face area in different environments to induce the generator to generate faces of the same size.

[0095] 3. Evaluation indicators.

[0096] This embodiment uses angle error as the evaluation index of model performance. After the model estimates the pitch angle and yaw angle, it calculates the three-dimensional vector representing the gaze direction. The angle between this vector and the true direction vector (ground truth) is the most commonly used evaluation index in the gaze field. The calculation formula of angle error is:

[0097]

[0098] Among them, θ represents the angle between the true value and the estimated value of the line of sight vector, a and b represent the true line of sight direction vector and the estimated line of sight direction vector in three-dimensional space, respectively.

[0099] 4. Comparative experiments and result analysis.

[0100] In the experiment, the Gaze360 dataset is divided into a training set and a test set. The data recorded by 15 volunteers in the MPIIFaceGaze dataset are divided into 36,000 images of 12 volunteers as the training set, and 9,000 images of the remaining 3 volunteers as the test set.

[0101] 4.1 Comparison of the performance of the method proposed in this invention and the line of sight estimation model.

[0102] In order to evaluate the performance of the proposed method compared with other gaze estimation methods, comparative experiments were carried out with ResNet-18, CA-Net, RT-Gene, Dilated-Net, GazeNAS-ETH, FullFace and other methods. The experimental results are shown in Table 1.

[0103] Table 1 Performance comparison of different methods on Gaze360 and MPIIFaceGaze datasets

[0104]

[0105] From the experimental results in Table 1, it can be seen that the proposed method has higher line of sight estimation accuracy than other methods, reaching 10.18° on the Gaze360 dataset and 3.84° on the MPIIFaceGaze dataset. Since the proposed method extracts line of sight related features in the image latent space and generates source domain images, it not only shows excellent performance in a single domain, but also has verified its generalization ability in cross-domain line of sight estimation tasks.

[0106] 4.2 Comparison of cross-domain performance between the method proposed in this invention and the cross-domain line of sight estimation model.

[0107] Generally speaking, in a cross-dataset environment, gaze estimation methods suffer from significant performance degradation due to overfitting to the source domain. In the model of the present invention, the latent code is selectively utilized through gaze-aware operations to solve the overfitting problem of gaze-independent features. In addition, the domain transfer module from the target domain to the source domain maps the input image to the source domain space that the gaze estimator can fully perceive.

[0108] In order to evaluate the cross-domain performance of the proposed method compared with other line of sight estimation methods, comparative experiments were conducted with Gaze360, GazeAdv, DAGEN, ADDA, UMA, PnP-GA, PureGaze, RT-Gene, RUDA and other methods. The model trained on the Gaze360 dataset is used to test the MPIIFaceGaze dataset. In order to prove that the proposed method is suitable for practical application scenarios, no more than 100 target samples are randomly selected in this experiment.

[0109] The experimental results are shown in Table 2, where G represents the Gaze360 dataset and M represents the MPIIFaceGaze dataset. The method of the present invention shows favorable sight angle errors in a cross-domain environment.

[0110] Table 2 Cross-dataset evaluation

[0111] method Target sample G→M Gaze360 >100 7.00° GazeAdv ~500 7.88° DAGEN ~500 8.02° ADDA ~1000 7.18° UMA ~100 8.51° PnP-GA <100 6.18° PureGaze <100 9.28° RT-Gene <100 21.81° RUDA <100 6.20° The present invention <100 5.37°

[0112] Experimental results show that the present invention can transfer the target domain to the source domain while retaining the line of sight related features. Due to the large difference in resolution between the images in the Gaze360 dataset and the MPIIFaceGaze dataset, the method of the present invention adds super-resolution and denoising preprocessing in the generator stage of reconstructing the source domain image to solve this problem.

[0113] 5. Ablation experiment.

[0114] The present invention obtains the potential code of the image based on the E2StyleGAN network and a pre-trained encoder, obtains the sight-related features through the sight-line perception operation, and sends the image reconstructed by the sight-related features into the ResNet-18 network to obtain the sight direction.

[0115] To verify the effectiveness of the gaze-aware operation, images and processed latent codes from the MPIIFaceGaze dataset were used, using 448 (all), 256 (56%), and 64 (14%) blocks in the latent code, respectively. Through the gaze-aware operation, blocks that are not related to the gaze are removed. As shown in Table 3, the quantitative results of evaluating the effect of the gaze-aware operation are presented.

[0116] Table 3 Quantitative results of evaluating the effect of line of sight perception analysis operation

[0117]

[0118] Experimental results show that the angular error using gaze-aware operations is lower than that when using a large number of blocks, indicating that the present invention can effectively and accurately separate gaze-related features from other closely intertwined features.

[0119] Therefore, the present invention adopts the above-mentioned line of sight estimation method based on the E2StyleGAN network, and converts the image from the target domain to the source domain and retains the line of sight related information by adding the line of sight distortion loss to train the encoder of E2StyleGAN, obtains the line of sight related features in the potential code through the line of sight perception operation, and then reconstructs the source domain image using the line of sight related features, sends them to the line of sight estimation network to learn features, and finally integrates the line of sight related features obtained in the potential code to predict the line of sight direction; the present invention processes the input image to make the model more generalized and have better cross-domain effect; obtains the potential information of the image through the inversion of GAN, and uses the line of sight perception operation to select the information related to the line of sight to assist the line of sight estimation.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A line of sight estimation method based on E2StyleGAN network, characterized in that: The following steps are involved: Step S1, through the source domain conversion of the E2StyleGAN network, the input image is mapped to the source domain to obtain the potential code of the image; in the process of domain conversion, the line of sight distortion loss is added to train the encoder; Step S2: The latent code processing module finds the features related to the line of sight through the line of sight perception operation, and feeds them into the generator of E2StyleGAN to reconstruct the source domain image; Step S3: The sight estimation model uses the features learned from the source domain image and the sight-related features obtained from the latent code processing to perform feature fusion and calculate the sight vector Among them, the overall framework of a sight line estimation method based on the E2StyleGAN network consists of three parts: latent code processing module, source domain conversion module, and sight line estimation module.

2. The line of sight estimation method based on the E2StyleGAN network according to claim 1, characterized in that: In step S1, the input image is mapped to the source domain to obtain the latent code of the image through the source domain conversion of the E2StyleGAN network; in the process of domain conversion, the line of sight distortion loss is added to train the encoder. The specific process is as follows: Step S11, source domain conversion of E2StyleGAN; Given a labeled source domain D S , the goal is to learn a translation model f, which only accesses the source domain data D S , the image is transferred from the target domain D T Mapping to source domain D S ; The E2StyleGAN inversion method is used to invert the real image using its out-of-distribution generalization ability. The source domain conversion process is as follows: Among them, F represents the feature extractor, S * Represents a domain converter; Step S12, the E2StyleGAN encoder training process includes two stages. First, the gaze-related extractor is trained to estimate the gaze direction from the facial image; second, the gaze distortion loss is used to train the gaze-related extractor to estimate the gaze direction from the facial image. To train the encoder, is the difference between the result of generating the image F(G(E(I))) and the line of sight extractor g. The specific process is as follows: Step S121: train the extractor F combined with ResNet-18 and MLP to estimate the gaze direction from the face image, and use the result of the image F(G(E(l))) generated from F to calculate the gaze distortion loss As shown below: Among them, E is the encoder, G is the generator, I is the input image, and g is the three-dimensional real line of sight angle; Step S122: The encoder uses three types of losses, namely, the least square error L2, the perceptual loss L LPIPS , similarity loss L sim Based on this, add the line of sight distortion loss For training, the expression of total loss is as follows: L total (I)=λ l2 L2(I)+λ LPIPS L LPIPS (I)+λ sim L sim (I)+λ GD L GD (I) (3); Among them, λ l2 ,λ LPIPS ,λ sim ,λ GD is the weight of each loss; The encoder and extractor F are trained in an end-to-end manner so that F is optimized to generate images that preserve gaze-related features.

3. The sight line estimation method based on the E2StyleGAN network according to claim 2, characterized in that: RetinaFace is used to capture facial images and crop them to reduce interference input caused by different environments.

4. The sight line estimation method based on the E2StyleGAN network according to claim 1, characterized in that: In step S2, the potential code is processed, and the specific process is as follows; The inversion of E2StyleGAN uses the line-of-sight related attributes in the input image to invert the target image into the latent space. By discovering the code of the line-of-sight related attributes, the manipulated latent vector is ensured to be highly correlated with the line-of-sight vector. The calculation formula of the inversion is as follows: Where f represents the optimal gaze estimator, X represents the latent vector, x represents the element of the latent vector X, and d represents the dimension of the latent space.

5. The sight line estimation method based on the E2StyleGAN network according to claim 4, characterized in that: In step S2, the line of sight sensing operation is performed, and the specific process is as follows: Step S211: Divide the data set into two groups G L and G R ; G L represents a group of sight vectors with yaw axis directions from 30 degrees to 90 degrees, G R A group representing a sight line vector having a yaw axis direction of -30 degrees to -90 degrees; Step S212, calculating the mean of each group of overall potential codes; Step S213: Since adjacent tensors represent similar features, the 16 tensors are regarded as a single block and defined as the unit of statistical operation. The calculation process is as follows: Among them, C i represents the i-th data block, k represents the number of data blocks in the potential code, and j represents the number of elements in the data block; C i and C j Respectively represent G L Group and G R blocks in a group; The blocks were sorted in descending order, the block most associated with sight was selected, and then a paired t-test was performed on the two expectations for each block from the two groups; Step S214: Paired t-test of two expected values. The specific process is as follows: The sample sizes of the two groups are large enough. According to the central limit theorem, the distribution of the sample mean difference is approximately Gaussian, and the variance of each sample is approximately the population variance, as shown below: Among them, H0 represents the null hypothesis, H1 represents the alternative hypothesis, and Represent the overall average values ​​of the two groups on the i-th block, d represents the dimension of the latent space, T represents the test statistic, and Respectively and The sample statistic, n L and n R Respectively represent G L and G R The number of, σ represents the standard deviation; Among them, since the indices of sight-related blocks in different datasets are different, the channel attention layer CA-layer is used to improve the cross-domain generalization performance.