Model training method, sight line estimation method, electronic equipment and storage medium

By introducing meta-learning and personalized correction technology into the three-dimensional line of sight estimation model, the problem of insufficient line of sight estimation accuracy among different groups of people is solved, and the accuracy and stability of line of sight estimation are significantly improved.

CN120147768APending Publication Date: 2025-06-13HEFEI DILUSENSE TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311693018.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing three-dimensional line of sight estimation method has insufficient accuracy among different populations, mainly due to the difference in the kappa angle of the human eye.

Method used

Based on the training of the general line of sight estimation model, the meta-learning idea is used to retrain the line of sight prediction network, and a small number of samples are used for personalized correction during the test phase to reduce the impact of kappa angle difference on the line of sight estimation results.

Benefits of technology

Through meta-learning retraining and personalized correction, the accuracy of line of sight estimation is significantly improved, the line of sight estimation error is reduced, and the stability and accuracy are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147768A_ABST
    Figure CN120147768A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of computer vision, and discloses a model training method, a sight line estimation method, electronic equipment and a storage medium, a trained general sight line estimation model is obtained, and the general sight line estimation model comprises a sight line feature network and a sight line prediction network; retraining a sight line prediction network in the general sight line estimation model by adopting meta-learning based on the feature training set to obtain a sight line prediction network after meta-learning; performing an iterative test on the sight line prediction network after the meta-learning based on the feature test set, and correcting the sight line prediction network after the meta-learning based on an iterative test result to obtain a target sight line estimation model; wherein the target line-of-sight estimation model comprises a line-of-sight feature network and a corrected line-of-sight prediction network after meta learning, the target line-of-sight estimation model can reduce the influence of the difference of kappa angles on a line-of-sight estimation result, and the line-of-sight estimation precision of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a model training, a gaze estimation method, an electronic device, and a storage medium. Background Art

[0002] Gaze estimation is a technology that uses mechanical, electronic, optical and other detection means to obtain the current gaze direction of a user, and has wide applications in fields such as game interaction, virtual reality (VR), medical treatment, assisted driving, and web page analysis.

[0003] In the field of computer vision, gaze estimation mainly takes an eye image or a face image as a processing object to estimate the gaze direction of a person or the position of a fixation point. At present, the research in this field can be roughly divided into three categories according to different scenarios and applications: fixation target estimation, fixation point estimation, and three-dimensional gaze estimation. Specifically, fixation target estimation is to detect the target object gazed by a given person; fixation point estimation is to estimate the two-dimensional landing point where the binocular gaze of a person focuses; three-dimensional gaze estimation is to deduce the three-dimensional gaze direction of a person. It is worth mentioning that after obtaining the three-dimensional gaze direction of the human eye, the calibration relationship between the camera screen, and the three-dimensional position of the human eye, the three-dimensional gaze direction can be converted into a two-dimensional landing point.

[0004] Generally speaking, the existing three-dimensional gaze estimation methods can be divided into two major categories: geometric-based methods and appearance-based methods. Among them, the geometric-based method calculates the gaze by detecting key features such as the corners and pupils of the human eye, while the appearance-based method directly learns a mapping model from appearance to gaze.

[0005] For the appearance-based gaze estimation method based on deep learning, when the training data and the test data come from different people, the accuracy of the estimated three-dimensional gaze often fluctuates between 4 and 5 degrees. The essential reason restricting the further improvement of this accuracy lies in the internal structure of the human eye, specifically the human eye kappa angle. The human eye kappa angle of each person is unique, which in turn leads to a certain deviation between the estimated gaze and the real gaze. That is to say, for two different people, even if the eye rotation angles are exactly the same, their gazes may still be 2-3 degrees different. This phenomenon is reflected in the training and testing of deep learning, resulting in different posterior probability distributions of the training data and the test data. Summary of the Invention

[0006] The objective of the embodiments of the present invention is to provide a model training, gaze estimation method, electronic device, and storage medium. On the basis of training a general gaze estimation model, the idea of meta-learning is further introduced to retrain the model, and during the model testing phase, personalized correction of the human eye gaze can be achieved using a small number of samples, thereby reducing the influence of the difference in kappa angle on the gaze estimation result and improving the accuracy of gaze estimation of the model.

[0007] To solve the above technical problems, an embodiment of the present invention provides a model training method, including:

[0008] Obtain a trained general gaze estimation model, where the general gaze estimation model includes: a gaze feature network for extracting gaze features from an image containing a human eye and a gaze prediction network for predicting the gaze from the gaze features;

[0009] Based on a feature training set, retrain the gaze prediction network in the general gaze estimation model using meta-learning to obtain a gaze prediction network after meta-learning;

[0010] Iteratively test the gaze prediction network after meta-learning based on a feature test set, and correct the gaze prediction network after meta-learning based on the iterative test results to obtain a target gaze estimation model;

[0011] Wherein, the target gaze estimation model includes: the gaze feature network and the gaze prediction network after meta-learning that has been corrected.

[0012] An embodiment of the present invention also provides a gaze estimation method, including:

[0013] Obtain a face image for which gaze estimation is to be performed; crop the face image into an image that can be received by the target gaze estimation model;

[0014] Input the cropped image into the target gaze estimation model to obtain predicted left and right eye gazes;

[0015] Wherein, the target gaze estimation model is trained by the model training method as described above.

[0016] An embodiment of the present invention also provides an electronic device, including:

[0017] At least one processor; and,

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the model training method as described above, or the gaze estimation method as described above.

[0020] An embodiment of the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the model training method as described above, or the gaze estimation method as described above.

[0021] Compared with the prior art, an embodiment of the present invention obtains a trained general gaze estimation model, which includes: a gaze feature network for extracting gaze features from an image containing a human eye and a gaze prediction network for predicting a gaze based on the gaze features; retrains the gaze prediction network in the general gaze estimation model using meta-learning based on a feature training set to obtain a gaze prediction network after meta-learning; iteratively tests the gaze prediction network after meta-learning based on a feature test set, and corrects the gaze prediction network after meta-learning based on the iterative test results to obtain a target gaze estimation model; wherein, the target gaze estimation model includes: a gaze feature network and a corrected gaze prediction network after meta-learning. When training the target gaze estimation model, this solution further incorporates the idea of meta-learning to retrain the gaze prediction network in the general gaze estimation model on the basis of the general gaze estimation model, and can achieve personalized correction of the human eye gaze using a small number of samples during the test stage for the gaze prediction network, thereby reducing the influence of the difference in kappa angle on the gaze estimation result and improving the accuracy of gaze estimation of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a specific flowchart of the model training method according to an embodiment of the present invention;

[0023] Figure 2 is a schematic structural diagram of a general gaze estimation model according to an embodiment of the present invention;

[0024] Figure 3 is a specific flowchart of the gaze estimation method according to an embodiment of the present invention;

[0025] Figure 4 is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will elaborate on the various embodiments of the present invention in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in the various embodiments of the present invention, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.

[0027] One embodiment of the present invention relates to a model training method, which can train a target line-of-sight estimation model for estimating the line of sight of a human eye in a face image. As Figure 1 shown, the model training method provided in this embodiment includes the following steps.

[0028] Step 101: Obtain a trained general line-of-sight estimation model, where the general line-of-sight estimation model includes: a line-of-sight feature network for extracting line-of-sight features from an image containing a human eye, and a line-of-sight prediction network for performing line-of-sight prediction on the line-of-sight features.

[0029] Among them, the function of the trained general line-of-sight estimation model is to extract line-of-sight features from an image containing a human eye, such as an image of a face region, an image of a left-eye region, and an image of a right-eye region. In this embodiment, neither the types nor the quantities of the images containing a human eye that can be input into the general line-of-sight estimation model are limited. For example, it can be set that the input images of the general line-of-sight estimation model include the above three types of images, that is, an image of a face region, an image of a left-eye region, and an image of a right-eye region.

[0030] According to the functional division, the general line-of-sight estimation model can include two networks: a line-of-sight feature network and a line-of-sight prediction network. Among them:

[0031] The function of the line-of-sight feature network is to extract line-of-sight features from an image containing a human eye. When the types and quantities of the images of the human eyes being extracted are relatively large, the line-of-sight features extracted by the line-of-sight feature network can be the line-of-sight features formed by separately extracting the line-of-sight features of these images and then fusing the separately extracted features.

[0032] The function of the line-of-sight prediction network is to perform line-of-sight prediction on the line-of-sight features extracted by the line-of-sight feature network, that is, to obtain the predicted left and right eye lines of sight. In this embodiment, the left and right eye lines of sight can be represented by a three-dimensional spatial vector.

[0033] It should be noted that in this embodiment, neither the network structures of the line-of-sight feature network and the line-of-sight prediction network nor the methods for training the general line-of-sight estimation model are limited.

[0034] In some embodiments, the method for obtaining a trained general gaze estimation model, that is, the method for training a general gaze estimation model, may include the following sub-steps 1011 to 1013:

[0035] Sub-step 1011: Construct a general training set. Each training sample in the general training set includes: an image sample composed of a face region image, a left eye region image, and a right eye region image, and a true gaze label composed of true left and right eye gazes.

[0036] As Figure 2 shown, construct a general training set containing images of human eyes. Each training sample (input image) in this general training set includes: an image sample composed of three types of images, namely, a face region image (face), a left eye region image (left eye), and a right eye region image (right eye), and a true gaze label composed of true left and right eye gazes.

[0037] Among them, the process of obtaining the images in the training sample can be to perform regional cropping on the originally collected human images to obtain images of the face region, left eye region, and right eye region. Among them, the true gaze label can be represented by the spatial three-dimensional vector of the true gaze.

[0038] Sub-step 1012: Input the image sample into the general gaze estimation model to be trained to obtain the predicted left and right eye gazes.

[0039] Specifically, when inputting the image sample into the general gaze estimation model to be trained, first extract gaze features through a gaze feature network, and then perform prediction on the extracted gaze features through a gaze prediction network to obtain the predicted left and right eye gazes.

[0040] Sub-step 1013: Construct a first loss function based on the true left and right eye gazes in the true gaze label and the predicted left and right eye gazes, and perform iterative training on the general gaze estimation model based on the first loss function to obtain a trained general gaze estimation model.

[0041] In this embodiment, the loss function for training the general gaze estimation model, that is, the first loss function, is constructed for the difference between the true left and right eye gazes in the true gaze label and the predicted left and right eye gazes. After obtaining the predicted left and right eye gazes, use the loss value calculated by the first loss function to iteratively optimize the network parameters of the general gaze estimation model, and finally obtain a trained general gaze estimation model.

[0042] Among them, the type of the first loss function in this embodiment is not limited. For example, it can be a loss function that calculates squared loss, binary classification loss, cross-entropy loss, etc.

[0043] In addition, regarding the model structure of the general line-of-sight estimation model, in some embodiments, as Figure 2 shown, the line-of-sight feature network may include three branch network layers and one feature fusion layer. The three branch network layers are used to receive the face region image, left eye region image, and right eye region image in the image sample respectively, and the feature fusion layer is used to fuse the feature maps output by the three branch network layers to obtain the line-of-sight feature. Among them, preferably, the network structures of the two branch network layers corresponding to receiving the left eye region image and the right eye region image in the training sample (the two branch network layers at the bottom of the image) are the same, and the same weights are shared during the iterative training of the general line-of-sight estimation model. These two branch network layers are set to be the same, so that the process of feature extraction for the left eye region image and the right eye region image is indistinguishable, which also conforms to the conventional symmetry of the human left and right eyes.

[0044] In some embodiments, the line-of-sight prediction network may include two output layers, which are respectively used to output the predicted left and right eye lines of sight. As Figure 2 shown, the line-of-sight prediction network includes two parallel output layers, which respectively output the predicted left and right eye lines of sight.

[0045] In this embodiment, the network structures adopted by the above three branch network layers may not be completely the same or may be completely the same. For example, the network structure of the branch network layer corresponding to receiving the face image in the training sample may be the same as or different from the network structures of the two branch network layers corresponding to receiving the left eye region image and the right eye region image in the training sample, and the network structures of the two branch network layers corresponding to receiving the left eye region image and the right eye region image in the training sample may be the same or different. For example, all three branch network layers may adopt the ResNet structure.

[0046] Step 102: Retrain the line-of-sight prediction network in the general line-of-sight estimation model using meta-learning based on the feature training set to obtain the line-of-sight prediction network after meta-learning.

[0047] Specifically, the input of the line-of-sight prediction network in the general line-of-sight estimation model is the line-of-sight feature output by the line-of-sight feature network, and the output is the predicted left and right eye lines of sight. Therefore, when retraining the line-of-sight prediction network using meta-learning, each training sample (also called the retraining sample) in the training set used will include the line-of-sight feature output by the line-of-sight feature network and the true line-of-sight label composed of the true left and right eye lines of sight corresponding to this line-of-sight feature. To distinguish it from the above general training set, the training set used in the meta-learning process in this embodiment is denoted as the feature training set.

[0048] During the retraining process of the gaze prediction network in the general gaze estimation model, the gaze feature network can be kept fixed. Using the gaze features output by the gaze feature network and their corresponding true gaze labels as retraining samples, a feature training set is constructed. Based on the feature training set, with the idea of meta-learning (calibrating on one sample set and validating on another sample set), the gaze prediction network is retrained to obtain the gaze prediction network after meta-learning.

[0049] In some embodiments, this step may specifically include the following sub-steps 1021 to 1023.

[0050] Sub-step 1021: Using the gaze features output by the gaze feature network in the trained general gaze estimation model and the true left and right eye gazes as retraining samples, multiple retraining samples belonging to the same person are used as a feature training set, and this feature training set is divided into a calibration set and a validation set.

[0051] Specifically, an image sample containing a face region image, a left eye region image, and a right eye region image can be input into the gaze feature network to obtain gaze features. It should be noted that the image samples here are not completely the same or are completely different from the image samples in the general training set used to train the general gaze estimation model. The obtained gaze features and the true left and right eye gazes corresponding to the image samples are used as retraining samples. Then, for these retraining samples, multiple retraining samples belonging to the same person are used as a feature training set, and this feature training set is divided into a calibration set and a validation set.

[0052] For example, a sample set containing a large number of retraining samples (in each retraining sample, the gaze feature is denoted as z, and the true left and right eye gazes are denoted as g) can be denoted as the training set S train , from S train a random identity id is selected, and then a sample set is randomly sampled from the large number of retraining samples corresponding to this identity id as a feature training set P train , P train = {D C train , D V train}, where D C train = {(z i , g i )|i = 1,..., k} is the calibration set, Dv train = {(z j , g j )|j = 1,..., l} is the validation set, and it can be set that k, l ≤ 20; z i 、z j are the i-th and j-th gaze features in sequence, gi , g j are the i-th and j-th true left and right eye lines of sight in sequence.

[0053] Sub-step 1022: For each feature training set, use the calibration set in this feature training set to obtain the network parameter α of the line-of-sight prediction network layer to be updated through n rounds of stochastic gradient descent n The intermediate parameter α n ’.

[0054] For example, for each feature training set P train perform one iteration of training. First, use the line-of-sight features in the retraining samples in the calibration set D train in this feature training set P C train input into the line-of-sight prediction network to obtain the predicted left and right eye lines of sight; calculate the loss between the predicted left and right eye lines of sight and the corresponding true left and right eye lines of sight, and obtain the updated network parameter of the line-of-sight prediction network layer based on the loss. Since each calibration set D C train contains multiple retraining samples, in this way, through n rounds of stochastic gradient descent, the network parameter α n of the line-of-sight prediction network layer before this iteration of training can be updated to obtain the intermediate parameter α n ’.

[0055] Sub-step 1023: Use the intermediate parameter α n ’, calculate the second loss function on the validation set in this feature training set, calculate the gradient of the second loss function with respect to the network parameter α n , and update the network parameter α n based on this gradient descent.

[0056] For example, after obtaining the intermediate parameter α n ’, based on the line-of-sight prediction network layer with this intermediate parameter α n ’, use the line-of-sight features in the retraining samples in the validation set D train in the above-mentioned feature training set P V train input into the line-of-sight prediction network to obtain the predicted left and right eye lines of sight; calculate the loss between the predicted left and right eye lines of sight and the corresponding true left and right eye lines of sight, and calculate the gradient of this loss with respect to the network parameter α n before this iteration of training, and then update the network parameter α n based on this gradient descent to obtain the line-of-sight prediction network after this iteration of training.

[0057] Similarly, when using the feature training set P of the current identity id trainAfter one iteration of training the gaze prediction network, the training set P of features with other identity IDs can be continuously used for subsequent iterations of training. Eventually, the training phase is completed, and after retraining, that is, after meta-learning, the gaze prediction network is obtained. train Proceed with subsequent iterations of training until the training phase is finally completed, obtaining the gaze prediction network after retraining, that is, after meta-learning.

[0058] Step 103: Iteratively test the meta-learned gaze prediction network based on the feature test set, and correct the meta-learned gaze prediction network based on the iterative test results to obtain the target gaze estimation model.

[0059] Among them, the target gaze estimation model includes: a gaze feature network and the meta-learned gaze prediction network after correction.

[0060] Specifically, the input of the gaze prediction network in the general gaze estimation model is the gaze feature output by the gaze feature network, and the output is the predicted left and right eye gazes. Therefore, when testing the gaze prediction network, each test sample in the test set used will include the gaze feature output by the gaze feature network and the true gaze label composed of the true left and right eye gazes corresponding to this gaze feature. In this embodiment, the test set used in the testing process of the gaze prediction network is denoted as the feature test set.

[0061] During the testing process of the gaze prediction network in the general gaze estimation model after meta-learning, the gaze feature network can be kept fixed. Using the gaze feature output by the gaze feature network and its corresponding true gaze label as test samples, a feature test set is constructed. Based on the feature test set, the gaze prediction network is tested. After obtaining the test results, the model can be evaluated according to the test results, and then the meta-learned gaze prediction network is corrected to obtain the target gaze estimation model composed of the gaze feature network and the meta-learned gaze prediction network after correction.

[0062] In some embodiments, this step may specifically include the following sub-steps 1031 to 1033.

[0063] Sub-step 1031: Use the gaze feature output by the gaze feature network in the trained general gaze estimation model and the true left and right eye gazes as test samples. Multiple test samples of the same person are used as a feature test set, and this feature test set is divided into a calibration set and a validation set.

[0064] Specifically, an image sample composed of a face region image, a left eye region image, and a right eye region image can be input into the gaze feature network to obtain gaze features. It should be noted that the image samples here are not exactly the same or are completely different from the image samples in the general training set used to train the general gaze estimation model and the image samples that formed the feature training set before. The obtained gaze features and the true left and right eye gazes corresponding to the image samples are used as test samples. Then, for these test samples, multiple test samples belonging to the same person are used as a feature test set, and the feature test set is divided into a calibration set and a validation set.

[0065] For example, a sample set containing a large number of test samples (in each test sample, the gaze feature is denoted as z', and the true left and right eye gazes are denoted as g') can be denoted as test set S test , from S test a random identity id is selected, and then a sample set is randomly sampled from the large number of test samples corresponding to this identity id as a feature test set P test , P test ={D C test , D V test}}, where D C test ={(z i ’, g i ’)|i = 1,..., k} is the calibration set, Dv test ={(z j ’, g j ’)|j = 1,..., l} is the validation set, and it can be set that k, l ≤ 20; z i ’, z j ’ are the i-th and j-th gaze features in turn, and g i ’, g j ’ are the i-th and j-th true left and right eye gazes in turn.

[0066] In practical applications, the sample size of the feature test set can be much smaller than that of the feature training set.

[0067] Sub-step 1032: For each feature test set, the network parameters θ n to be updated for the gaze prediction network layer are obtained through n rounds of stochastic gradient descent using the calibration set in this feature test set, and the target parameter θ n ’;

[0068] For example, for each feature test set P test a test is conducted. First, using the calibration set D test in this feature test set P C testThe line-of-sight features in the test samples are input into the line-of-sight prediction network to obtain the predicted left and right eye lines of sight; the loss between the predicted left and right eye lines of sight and the corresponding true left and right eye lines of sight is calculated, and the network parameters of the updated line-of-sight prediction network layer are obtained based on the loss. Since each calibration set D C test contains multiple test samples, in this way, the network parameters θ n of the line-of-sight prediction network layer before this test can be updated through n rounds of stochastic gradient descent to obtain the target parameters θ n ’.

[0069] Sub-step 1033: Using the target parameter θ n ’, the test results are obtained on the validation set in this feature test set, and the target parameter θ n ’ is corrected based on this test result to obtain the target line-of-sight estimation model.

[0070] For example, after obtaining the target parameter θ n ’, based on the line-of-sight prediction network layer of this target parameter θ n ’, the validation set D test in the above-mentioned feature test set P V test The line-of-sight features in the test samples are input into the line-of-sight prediction network to obtain the predicted left and right eye lines of sight as the test results, and then the target parameter θ n ’ is corrected based on this test result, so that the corrected target parameter θ n ’ is used as the final network parameter to obtain the target line-of-sight estimation model.

[0071] The correction method may include: continuing to calculate the loss between the predicted left and right eye lines of sight in the test results and the corresponding true left and right eye lines of sight, and correcting the target parameter θ n ’ according to the magnitude of this loss to reduce the loss.

[0072] Similarly, after completing a test on the line-of-sight prediction network using the feature test set P test of the current identity id, the feature test set P test of other identity ids can be continued to be used for subsequent iterative tests, and finally the test phase is completed to obtain the line-of-sight prediction network after the test. The line-of-sight feature network and the line-of-sight prediction network obtained after correcting the target parameter θ n ’ constitute the target line-of-sight estimation model that this embodiment finally intends to train and complete.

[0073] This application introduces the idea of meta-learning into gaze estimation. By guiding the model to focus more on the "learning" ability rather than overfitting the current data, the generalization ability of the model for different human eyes is improved. During the prediction stage, through fine-tuning with a small number of samples, the influence of the human eye kappa can be reduced, and the accuracy of gaze estimation is enhanced.

[0074] After actual comparison, when testing a newly trained general gaze estimation model, the gaze angle error is greater than 4.5 degrees. After introducing the meta-learning method in this application, the minimum error of the target gaze estimation model obtained through training comes to about 2.86 degrees, and the result of each fine-tuning is stable.

[0075] Compared with the related technology, in this embodiment, when training the target gaze estimation model, on the basis of the general gaze estimation model, the idea of meta-learning is further introduced to retrain the gaze prediction network in the general gaze estimation model, and during the test stage for the gaze prediction network, personalized correction of the human eye gaze can be achieved using a small number of samples, thereby reducing the influence of the difference in the kappa angle on the gaze estimation result and enhancing the accuracy of the gaze estimation of the model.

[0076] Another embodiment of the present invention relates to a gaze estimation method, which is implemented based on the target gaze estimation model trained according to the above method embodiment. As Figure 3 shown, the gaze estimation method includes the following steps.

[0077] Step 201: Obtain a face image for which gaze estimation is to be performed.

[0078] Among them, the face image for which gaze estimation is to be performed should include a clear human eye image area.

[0079] Step 202: Crop the face image into an image that can be received by the target gaze estimation model.

[0080] Among them, the target gaze estimation model is the model obtained through the model training method in the above method embodiment.

[0081] For example, the face area image, left eye area image, and right eye area image can be cropped from the face image to meet the image requirements that can be received by the target gaze estimation model.

[0082] Step 203: Input the cropped image into the target gaze estimation model to obtain the predicted left and right eye gazes.

[0083] For example, input the face area image, left eye area image, and right eye area image cropped from the face image into the target gaze estimation model to obtain the predicted left and right eye gazes.

[0084] Compared with related technologies, in this embodiment, the target line-of-sight estimation model is used to estimate the line of sight of the human eye. Since the target line-of-sight estimation model is obtained by further adding the idea of meta-learning to retrain the line-of-sight prediction network in the general line-of-sight estimation model and realizing the personalized correction of the human eye line of sight with a small number of samples during the test stage of the line-of-sight prediction network, the influence of the difference in kappa angle on the line-of-sight estimation result is reduced, and the accuracy of the line-of-sight estimation of the model is improved.

[0085] Another embodiment of the present invention relates to an electronic device, such as Figure 4 shown, including at least one processor 302; and a memory 301 communicatively connected to the at least one processor 302; wherein, the memory 301 stores instructions executable by the at least one processor 302, and the instructions are executed by the at least one processor 302 to enable the at least one processor 302 to execute any of the above method embodiments.

[0086] Among them, the memory 301 and the processor 302 are connected by a bus. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 302 and the memory 301 together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore, they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 302 is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 302.

[0087] The processor 302 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 301 can be used to store the data used by the processor 302 when executing operations.

[0088] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, any of the above method embodiments is implemented.

[0089] That is, those skilled in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0090] Those of ordinary skill in the art can understand that the above-described embodiments are specific embodiments for implementing the present invention. In actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present invention.

Claims

1. A model training method, characterized in that, it includes: obtaining a trained general line-of-sight estimation model, where the general line-of-sight estimation model includes: a line-of-sight feature network for extracting line-of-sight features from an image containing a human eye and a line-of-sight prediction network for predicting the line of sight of the line-of-sight features; retraining the line-of-sight prediction network in the general line-of-sight estimation model based on a feature training set using meta-learning to obtain a line-of-sight prediction network after meta-learning; iteratively testing the line-of-sight prediction network after meta-learning based on a feature test set, and correcting the line-of-sight prediction network after meta-learning based on the iterative test results to obtain a target line-of-sight estimation model; wherein, the target line-of-sight estimation model includes: the line-of-sight feature network and the line-of-sight prediction network after meta-learning that has been corrected.

2. The method according to claim 1, characterized in that, the obtaining of the trained general line-of-sight estimation model includes: constructing a general training set, where each training sample in the general training set includes: an image sample composed of a face region image, a left-eye region image, and a right-eye region image, and a true line-of-sight label composed of true left and right eye lines of sight; inputting the image sample into the general line-of-sight estimation model to be trained to obtain predicted left and right eye lines of sight; constructing a first loss function based on the true left and right eye lines of sight in the true line-of-sight label and the predicted left and right eye lines of sight, and iteratively training the general line-of-sight estimation model based on the first loss function to obtain a trained general line-of-sight estimation model.

3. The method according to claim 2, characterized in that, the line-of-sight feature network includes three branch network layers and a feature fusion layer. The three branch network layers are used to respectively receive the face region image, the left-eye region image, and the right-eye region image in the image sample, and the feature fusion layer is used to fuse the feature maps output by the three branch network layers to obtain the line-of-sight features; the line-of-sight prediction network includes two output layers, which are respectively used to output the predicted left and right eye lines of sight.

4. The method according to claim 3, characterized in that, the network structures of the two branch network layers corresponding to receiving the left-eye region image and the right-eye region image in the training sample are the same, and the same weights are shared during the iterative training of the general line-of-sight estimation model.

5. The method according to claim 2, characterized in that, the retraining of the line-of-sight prediction network in the trained general line-of-sight estimation model based on a feature training set using meta-learning to obtain a line-of-sight prediction network after meta-learning includes: using the line-of-sight features output by the line-of-sight feature network in the trained general line-of-sight estimation model and the true left and right eye lines of sight as retraining samples, using multiple retraining samples of the same person as a feature training set, and dividing the feature training set into a calibration set and a validation set; For each feature training set, the calibration set in the feature training set is used to obtain the network parameter α to be updated for the line-of-sight prediction network layer through n rounds of stochastic gradient descent n as the intermediate parameter α n '; Using the intermediate parameter α n ’, calculate a second loss function on the validation set in the feature training set, calculate the gradient of the second loss function with respect to the network parameter α n and update the network parameter α based on gradient descent n .

6. The method according to claim 2, characterized in that, Iteratively testing the meta-learned line-of-sight prediction network based on the feature test set, and correcting the meta-learned line-of-sight prediction network based on the iterative test results to obtain a target line-of-sight estimation model, including: Using the line-of-sight features output by the line-of-sight feature network in the trained general line-of-sight estimation model, and the real left and right eye lines of sight as test samples, taking multiple test samples of the same person as one feature test set, and dividing this feature test set into a calibration set and a validation set; For each feature test set, the network parameters θ to be updated for the line-of-sight prediction network layer are obtained by performing n rounds of stochastic gradient descent on the calibration set in the feature test set n of the target parameter θ n ’; Using the target parameter θ n ’, obtaining a test result on the validation set in the feature test set, and correcting the target parameter θ based on the test result n ’ to obtain a target line-of-sight estimation model.

7. The method according to claim 1, wherein, the sample size of the feature test set is much smaller than the sample size of the feature training set.

8. A line-of-sight estimation method, wherein, including: Obtaining a face image for which line-of-sight estimation is to be performed; Cropping the face image into an image that can be received by the target line-of-sight estimation model; Inputting the cropped image into the target line-of-sight estimation model to obtain predicted left and right eye lines of sight; wherein, the target line-of-sight estimation model is trained by the model training method according to any one of claims 1-7.

9. An electronic device, wherein, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the model training method according to any one of claims 1 to 7, or the line-of-sight estimation method according to claim 8.

10. A computer-readable storage medium storing a computer program, wherein, the computer program, when executed by a processor, implements the model training method according to any one of claims 1 to 7, or the line-of-sight estimation method according to claim 8.