Network model construction method, line-of-sight estimation method, device, equipment and medium

By introducing weights determined by the theoretical value of line-of-sight difference into the network model training, and training the model based on weighted loss, the problem of insufficient accuracy in line-of-sight estimation is solved, and more accurate line-of-sight difference and direction estimation are achieved.

CN119339429BActive Publication Date: 2026-01-02BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310899892.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-20
Publication Date
2026-01-02
Estimated Expiration
2043-07-20

AI Technical Summary

Technical Problem

Existing network model training methods result in poor accuracy in gaze estimation, failing to effectively consider the impact of gaze differences between eye images on model training.

Method used

By acquiring multiple training image groups, the first weight is determined based on the theoretical value of the line-of-sight difference, and the initial network model is trained based on the weighted loss to construct the target network model, fully considering the impact of the line-of-sight difference on the model output.

Benefits of technology

It improves the accuracy of line-of-sight difference estimation results, ensures the accuracy of line-of-sight direction estimation, and meets the application needs of model testing or inference stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339429B_ABST
    Figure CN119339429B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a network model construction method, a line-of-sight estimation method, an apparatus, a device and a medium. The network model construction method comprises: obtaining a plurality of training image groups; obtaining a line-of-sight difference theoretical value corresponding to a training image group according to an angle difference between theoretical line-of-sight directions labeled by two eye sample images in the training image group, and determining a first weight corresponding to the training image group according to the line-of-sight difference theoretical value; obtaining line-of-sight difference estimation values respectively output by a preset initial network model for the plurality of training image groups; performing weighted processing on a loss between the line-of-sight difference theoretical value and the line-of-sight difference estimation value corresponding to each of the plurality of training image groups according to the first weight corresponding to each of the plurality of training image groups, to obtain a weighted loss; and training the initial network model based on the weighted loss, to construct a target network model based on the trained initial network model. The embodiments of the present disclosure can effectively guarantee the accuracy of line-of-sight direction estimation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to a network model construction method, a line-of-sight estimation method, device, equipment and medium. BACKGROUND

[0002] In many fields such as game field, medical field, intelligent control field and the like, the line-of-sight direction of human eyes needs to be recognized, so as to take corresponding strategies based on the line-of-sight direction estimation result of human eyes. In general, the line-of-sight direction needs to be estimated by means of a network model. However, the existing network model training manner is not good, which leads to poor accuracy of line-of-sight estimation by means of the output result of the network model. SUMMARY

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a network model construction method, a line-of-sight estimation method, device, equipment and medium.

[0004] In a first aspect, an embodiment of the present disclosure provides a network model construction method, comprising: obtaining a plurality of training image groups; wherein the training image group contains two eye sample images, and the eye sample images are labeled with theoretical line-of-sight directions; obtaining a line-of-sight difference theoretical value corresponding to the training sample group according to the angle difference between the theoretical line-of-sight directions labeled in the two eye sample images in the training image group, and determining a first weight corresponding to the training image group according to the line-of-sight difference theoretical value; obtaining a line-of-sight difference estimation value output by a preset initial network model for the plurality of training image groups respectively; wherein the line-of-sight difference estimation value is obtained by the initial network model estimating the angle difference between the line-of-sight directions of the two eye sample images in the training image group; performing weighted processing on the loss between the line-of-sight difference theoretical value and the line-of-sight difference estimation value corresponding to each of the plurality of training image groups according to the first weight corresponding to each of the plurality of training image groups, to obtain a weighted loss; training the initial network model based on the weighted loss, to construct a target network model based on the trained initial network model; wherein the target network model is used to generate a line-of-sight difference estimation value between two target eye images.

[0005] In a second aspect, an embodiment of the present disclosure provides a line-of-sight estimation method, comprising: obtaining a calibration eye image labeled with a theoretical line-of-sight direction and a target eye image to be estimated; obtaining a line-of-sight difference estimation value between the calibration eye image and the target eye image by a pre-constructed target network model; wherein the target network model is obtained based on the network model construction method provided in the first aspect; estimating the line-of-sight direction corresponding to the target eye image according to the theoretical line-of-sight direction of the calibration eye image and the line-of-sight difference estimation value between the calibration eye image and the target eye image.

[0006] In a third aspect, the embodiments of the present disclosure provide a network model construction device, comprising: an image group acquisition module configured to acquire a plurality of training image groups; wherein the training image groups each comprise two eye sample images, and the eye sample images are labeled with theoretical gaze directions; a first weight determination module configured to determine a first weight corresponding to each of the training image groups according to an angle difference between the theoretical gaze directions of the two eye sample images in the training image group; a network model output module configured to acquire gaze difference estimation values respectively output by a preset initial network model for the plurality of training image groups; wherein the gaze difference estimation value is an estimation value of an angle difference between the theoretical gaze directions of the two eye sample images in the training image group by the initial network model; a weighted loss acquisition module configured to acquire a weighted loss by performing weighted processing on a loss between the gaze difference theoretical value and the gaze difference estimation value corresponding to each of the plurality of training image groups according to the first weight corresponding to each of the plurality of training image groups; and a network model training module configured to train the initial network model based on the weighted loss, and to construct a target network model based on the trained initial network model; wherein the target network model is configured to generate a gaze difference estimation value between two target eye images.

[0007] In a fourth aspect, the embodiments of the present disclosure provide a gaze estimation device, comprising: an image acquisition module configured to acquire a calibration eye image labeled with a theoretical gaze direction and a target eye image with a to-be-estimated gaze direction; a gaze difference estimation module configured to acquire a gaze difference estimation value between the calibration eye image and the target eye image by using a target network model; wherein the target network model is obtained based on the network model construction method of the first aspect; and a gaze direction estimation module configured to estimate a gaze direction corresponding to the target eye image according to the theoretical gaze direction of the calibration eye image and the gaze difference estimation value between the calibration eye image and the target eye image.

[0008] In a fifth aspect, the embodiments of the present disclosure provide an electronic device, comprising: a processor; a memory configured to store executable instructions of the processor; and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the network model construction method of the first aspect, or implement the gaze estimation method of the second aspect.

[0009] In a sixth aspect, the embodiments of the present disclosure provide a computer readable storage medium, wherein the storage medium stores a computer program, and the computer program is configured to execute the network model construction method of the first aspect, or implement the gaze estimation method of the second aspect.

[0010] The above technical solution provided by the embodiments of the present disclosure is based on the obtained line-of-sight difference estimation values respectively output by the preset initial network model for the plurality of training image groups and the line-of-sight difference theoretical values corresponding to the training image groups, and further determines the first weight corresponding to the training image groups according to the line-of-sight difference theoretical values corresponding to the training image groups, so as to weight the loss between the line-of-sight difference theoretical values and the line-of-sight difference estimation values corresponding to the plurality of training image groups according to the first weights corresponding to the plurality of training image groups, and then perform model training based on the weighted loss. This way can fully consider the influence of the line-of-sight difference between the two eye images on the model output result, and in the model training process, the corresponding first weight is determined for the line-of-sight difference theoretical value corresponding to the training image group. By setting the first weight, the influence of the line-of-sight difference on the model training is introduced, which helps to make the line-of-sight difference estimation result output by the finally trained target network model more accurate, and on this basis, the result accuracy of the line-of-sight direction estimation based on the line-of-sight difference estimation value between the calibration eye image and the target eye image output by the target network model and the theoretical line-of-sight direction labeled by the calibration eye image is further ensured.

[0011] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0014] Figure 1 A schematic diagram of an angle panel is provided for the embodiments of the present disclosure;

[0015] Figure 2 A schematic diagram of an angle panel is provided for the embodiments of the present disclosure;

[0016] Figure 3 A flowchart of a network model construction method is provided for the embodiments of the present disclosure;

[0017] Figure 4 A flowchart of a line-of-sight estimation method is provided for the embodiments of the present disclosure;

[0018] Figure 5 A line-of-sight estimation process schematic diagram provided by an embodiment of the present disclosure;

[0019] Figure 6 A structure schematic diagram of a network model construction device provided by an embodiment of the present disclosure;

[0020] Figure 7 A structure schematic diagram of a line-of-sight estimation device provided by an embodiment of the present disclosure;

[0021] Figure 8 A structure schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0023] In the following description, many specific details are set forth in order to provide a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the description are only some of the embodiments of the present disclosure, not all the embodiments.

[0024] In the field of line-of-sight estimation, it is usually necessary to obtain multiple eye sample images and label the theoretical line-of-sight direction of the eye sample images, so as to form multiple training image groups through the multiple eye sample images, and train the network model for estimating the line-of-sight difference between two eye images. For the convenience of research, the theoretical line-of-sight direction can be represented by a corresponding angle, and refer to a kind of angle panel schematic diagram as shown in Figure 1 , which illustrates multiple concentric circles, and the angles corresponding to different circles are different. In some embodiments, the obtained eye sample images can be analyzed through the angle panel as shown in Figure 1 , such as, the angles corresponding to each eye sample image collected can be labeled in Figure 1 , so as to analyze the angle distribution of the eye sample images; and such as, in the model test stage, multiple calibration points and a test point can be set on the basis of Figure 1 , and the model prediction is performed by combining the test point with each calibration point respectively, and the accuracy of the model prediction result is measured. For the convenience of understanding, the following will be described in detail:

[0025] In actual application, multiple calibration points can be set, and the user gazes at the specified calibration points to collect the eye images of the user as calibration eye images, and a test point is determined and the eye image of the user gazing at the test point is obtained as a test eye image. For each calibration eye image, the line-of-sight difference between the test eye image and the calibration eye image is estimated by the network model to obtain a line-of-sight difference estimation value, and then the line-of-sight estimation result of the test eye image is obtained by combining the theoretical line-of-sight direction of the calibration eye image. Then, the line-of-sight estimation results corresponding to the multiple calibration eye images are weighted to obtain the line-of-sight estimation result of the test eye image by the model. For ease of understanding, refer to a schematic diagram of an angle panel shown in FIG. 7. The black heart dots shown in the schematic diagram are calibration points. For each calibration point, the angle corresponding to the calibration point in FIG. 7 is the theoretical line-of-sight direction corresponding to the calibration eye image collected when the user gazes at the calibration point. The black heart square is the actual test point. The angle corresponding to the actual test point in FIG. 7 is the theoretical line-of-sight direction corresponding to the test eye image collected when the user gazes at the test point. The gray heart dot is the prediction result point of each calibration point relative to the test point. The angle corresponding to the prediction result point in FIG. 7 is the angle obtained by the network model based on each calibration point and the test point (that is, the model estimation angle corresponding to the test eye image). The black heart triangle is the comprehensive prediction result point obtained by combining multiple calibration points relative to the test point. The angle corresponding to the comprehensive prediction result point in FIG. 7 is the angle obtained by weighting and fusing multiple model estimation angles corresponding to the test eye image. For ease of intuitive analysis, the positions of the points can be directly observed below. From FIG. 7, it can be intuitively known that the accuracy of the angle prediction result of the calibration point adjacent to the test point relative to the test point is relatively high (that is, the gray heart dot is closer to the black heart square). The accuracy of the comprehensive prediction result point obtained based on multiple calibration points relative to the test point is probably dependent on the calibration point adjacent to the test point, that is, the closeness of the black heart triangle to the black heart square mainly depends on the gray heart dot close to the black heart square. Figure 2 Figure 2 Figure 2 Figure 2 Figure 2 Figure 2

[0026] ​​​​​​However, the inventors have found that there is a big difference between the application of the model test stage and the model training stage. The main reason is that, as analyzed above, the model test stage is affected by the size of the line of sight difference between the two eye images (the closer the calibration point and the test point, the smaller the line of sight difference, and the greater the impact on the final line of sight prediction result of the model). In the model training stage, the two eye sample images in the training image group input to the model are random, and the impact of the line of sight difference between the eye sample images on the model training is not considered. The impact of training image groups with different line of sight differences on model training is consistent, resulting in low accuracy of the model trained in the test stage or subsequent inference stage to predict the line of sight direction of the eye image.

[0027] To at least improve the above problems, a method for constructing a network model can be provided as shown in Figure 3 The method can be executed by a network model construction device. The device can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in Figure 3 The method mainly includes the following steps S302-S310:

[0028] In step S302, a plurality of training image groups are obtained. Each training image group includes two eye sample images, and each eye sample image is labeled with a theoretical line of sight direction.

[0029] The eye sample image can be an eye image taken when the target eye gazes at a preset gaze point. The target eye can be an eye of a target object (such as a human). In some embodiments, a plurality of gaze points can be preset, and an eye image can be taken when the target eye gazes at each gaze point, thereby obtaining an eye sample image corresponding to each gaze point. In actual applications, a user can wear a VR glasses, smart glasses, an eye tracker, or other electronic devices that need to estimate the visual line direction of the user's eye. The electronic device can sense the eye state of the user, thereby determining the visual line direction when the eye gazes at a gaze point, and taking the visual line direction obtained by the electronic device as the theoretical visual line direction of the collected eye sample image, thereby serving as the label of the eye sample image. In actual applications, a collection prompt information can be initiated to the user before collecting the eye image of the user, and the user can be further informed of the purpose of image collection in the prompt information. In other embodiments, the eye image can be obtained through a network or the like, and then a theoretical visual line direction can be labeled for the eye image by using an existing model. In actual applications, the theoretical visual line direction can be represented by a spatial angle, which can be understood as the included angle between the visual line direction when the eye gazes at a gaze point and a reference direction. In some specific embodiments, the theoretical visual line direction can be directly represented by the yaw and pitch of the eyeball; in other embodiments, the theoretical visual line direction can also be represented by a target angle calculated based on the yaw and pitch, which can be used to represent the included angle between the visual line direction when the eye gazes at a gaze point and the visual line direction when the eye horizontally gazes at the gaze point. The above are only two exemplary representations of the visual line direction, and the representation of the visual line direction can be flexibly set in actual applications, which is not limited herein. In order to ensure the reliability of model training, a plurality of training image groups can be obtained, and the number of training image groups is not limited by the embodiments of the present disclosure. Generally, the more the number of training image groups, the stronger the robustness of the trained model.

[0030] In step S304, the angle difference between the theoretical visual line directions labeled by the two eye sample images in the training image group is obtained to obtain the theoretical visual line difference value corresponding to the training image group, and the first weight corresponding to the training image group is determined according to the theoretical visual line difference value. Each of the plurality of training image groups obtained corresponds to a first weight, and the first weights corresponding to different training image groups can be the same or different, mainly depending on the theoretical visual line difference value of the training image group.

[0031] In actual applications, the angle difference between the theoretical visual line directions can be directly represented by the yaw angle difference and the pitch angle difference, or by the target angle difference determined based on the yaw angle difference and the pitch angle difference, which is not limited herein.

[0032] In the embodiments of the present disclosure, the influence of the gaze difference size of the training image set on the model training is fully considered, and compared with the related art, the embodiments of the present disclosure additionally determine the first weight corresponding to the training image set according to the theoretical value of the gaze difference corresponding to the training image set, so as to control the influence degree of different training image sets on the model training based on the first weight. In some implementation examples, for any two training image sets, the first weight corresponding to the training image set with a smaller theoretical value of the gaze difference is not less than the first weight corresponding to the training image set with a larger theoretical value of the gaze difference. That is, the training image set with a smaller theoretical value of the gaze difference has a greater influence degree on the model training, so that the model obtained by training can focus more on the gaze difference estimation result of the two eye images with smaller gaze difference, thereby matching the model application manner in the subsequent test stage or inference stage.

[0033] In step S306, the gaze difference estimation value output by the preset initial network model for the plurality of training image sets is obtained; wherein the gaze difference estimation value is obtained by the initial network model estimating the angle difference between the gaze directions of the two eye sample images in the training image set.

[0034] The initial network model can be a neural network model, and the structure of the initial network model is not limited in the embodiments of the present disclosure. For example, the initial network model can be a differential model, which can specifically include a feature extraction network and a gaze difference prediction network. The feature extraction network is used to extract the feature vector of each eye sample image in the training image set, and the gaze difference prediction network can analyze and process the feature vectors of the two eye sample images in the training image set, thereby generating the gaze difference estimation value (also referred to as the gaze difference value).

[0035] In step S308, the loss between the theoretical value of the gaze difference and the gaze difference estimation value corresponding to each of the plurality of training image sets is weighted according to the first weight corresponding to each of the plurality of training image sets, to obtain a weighted loss.

[0036] In actual application, each training image group corresponds to a first weight, and the first weight can be directly used to perform weighted average processing on the loss between the line-of-sight difference theoretical value and the line-of-sight difference estimated value corresponding to each of the plurality of training image groups, to obtain a weighted loss. Other weights can also be introduced, and a comprehensive weight is obtained on the basis of the first weight and the other weights, and then the comprehensive weight is used to perform weighted average processing on the loss between the line-of-sight difference theoretical value and the line-of-sight difference estimated value corresponding to each of the plurality of training image groups, to obtain a weighted loss. The specific setting can be flexible, and is not limited herein. For each training image group, the loss between the line-of-sight difference theoretical value and the line-of-sight difference estimated value corresponding to the training image group can be obtained in the following manner: based on the difference between the line-of-sight difference theoretical value and the line-of-sight difference estimated value corresponding to the training image group, a preset loss function is used to determine a loss function value, and the loss function value is the loss between the line-of-sight difference theoretical value and the line-of-sight difference estimated value corresponding to the training image group. The present disclosure does not limit the loss function.

[0037] In step S310, the initial network model is trained based on the weighted loss, to construct a target network model based on the trained initial network model; wherein the target network model is used to generate a line-of-sight difference estimated value between two target eye images.

[0038] Training the initial network model based on the weighted loss can adjust the parameters of the network model towards the goal of reducing the weighted loss, so as to reduce the difference as much as possible, and stop training when a preset training end condition is reached (such as the total loss converging to a preset threshold range). Since the weighted loss sufficiently integrates the first weight determined by each of the different training image groups based on the line-of-sight difference theoretical value, the target network model obtained by model training based on the weighted loss can more accurately estimate the line-of-sight difference, and can be better matched with the application mode in the model test or inference stage.

[0039] In summary, the above-mentioned manner can fully consider the influence of the line-of-sight difference between two eye images on the model output result, and in the model training process, the corresponding first weight is determined for the line-of-sight difference theoretical value corresponding to the training image group. By introducing the influence of the line-of-sight difference on the model through the setting of the first weight, it is helpful to make the line-of-sight difference estimated result output by the finally obtained target network model more accurate.

[0040] In order to guarantee the reliability of the first weight, the present disclosure provides an implementation manner for determining the first weight corresponding to the training image group according to the line-of-sight difference theoretical value, which can be executed by referring to the following step one and step two:

[0041] Step one: obtain a preset line-of-sight difference threshold; wherein the number of line-of-sight difference thresholds is one or more, and the specific setting can be flexible.

[0042] In some embodiments, the line-of-sight difference threshold value includes a first line-of-sight difference threshold value and a second line-of-sight difference threshold value, where the second line-of-sight difference threshold value is greater than the first line-of-sight difference threshold value. For example, the first line-of-sight difference threshold value can be determined based on the maximum angle difference between the plurality of calibration points set in the test stage, such as the first line-of-sight difference threshold value can be set as range = 20. The second line-of-sight difference threshold value can be determined based on the effective angle FOV of the human eye looking up, down, left and right, such as the FOV = 40 can be set.

[0043] Step two, comparing the line-of-sight difference theoretical value of the training image set with the line-of-sight difference threshold value to determine the first weight corresponding to the training image set based on the comparison result. That is, through the comparison result between the line-of-sight difference theoretical value and the line-of-sight difference threshold value, the size of the line-of-sight difference theoretical value can be objectively measured, and different comparison results can set different first weight determination methods to reasonably determine the first weight.

[0044] In some embodiments, the step of determining the first weight corresponding to the training image set based on the comparison result can be performed according to the following cases one to three:

[0045] Case one, in the case that the line-of-sight difference theoretical value of the training image set is not higher than the first line-of-sight difference threshold value, the first weight corresponding to the training image set is determined as a preset first value. For example, the first value can be 1. That is, in the case that the line-of-sight difference theoretical value of the training image set is not higher than the first line-of-sight difference threshold value, it is indicated that the line-of-sight difference theoretical value of the training image set is small, and the influence degree on the model is large, so the second weight weight = 1 can be directly set. It should be noted that 1 is only an example, and other values can be set, which is not limited here.

[0046] Case two, in the case that the line-of-sight difference theoretical value of the training image set is between the first line-of-sight difference threshold value and the second line-of-sight difference threshold value, the second weight corresponding to the training image set is determined according to a preset algorithm; where the second weight is between a preset second value and the first value, the second value is less than the first value, and the second weight is negatively correlated with the line-of-sight difference theoretical value.

[0047] For example, the preset algorithm includes a linear algorithm, when the line-of-sight difference theoretical value is between the first line-of-sight difference threshold value and the second line-of-sight difference threshold value, the smaller the line-of-sight difference theoretical value, the greater the second weight, and the relationship between them can be represented by a linear function, and the present disclosure does not limit the linear function, such as, based on the first line-of-sight difference threshold value range and the second line-of-sight difference threshold value FOV have been set, the second weight weight = (1.2FOV-Δangel) / (1.2FOV-range) can be set. Wherein, Δangel is used to represent the line-of-sight difference theoretical value.

[0048] Scenario 3: If the theoretical value of the line-of-sight difference in the training image group is not lower than the second line-of-sight difference threshold, then the first weight corresponding to the training image group is determined as the second value. For example, the second value can be 0.2. That is, if the theoretical value of the line-of-sight difference in the training image group is not lower than the second line-of-sight difference threshold, it indicates that the theoretical value of the line-of-sight difference in the training image group is relatively large, and its impact on the model is relatively small, so the second weight can be directly set to weight = 0.2. It should be noted that 0.2 is only an example, and other values ​​can be set; there are no restrictions here.

[0049] Using the above method, the corresponding first weight can be reasonably determined based on the magnitude of the theoretical value of the line-of-sight difference.

[0050] Based on determining the first weights corresponding to each of the multiple training image groups, this disclosure further provides a specific implementation method for weighting the loss between the theoretical value and the estimated value of the line-of-sight difference corresponding to each of the multiple training image groups according to the first weights corresponding to each of the multiple training image groups. This can be performed with reference to steps A through C below:

[0051] Step A: Determine the second weight corresponding to the training image group based on the target angle range corresponding to the theoretical gaze direction of each of the two eye sample images in the training image group.

[0052] In this embodiment, the impact of the angle difference between two eye sample images on model training is considered not only, but also the impact of the angle range of each eye sample image on model training. Specifically, the inventors have discovered that the currently collected eye sample images suffer from an imbalance in angle; the amount of data for eye sample images at different angles varies. For example, among the collected eye sample images, the number of eye sample images with smaller angle ranges is usually greater, while the number of eye sample images with larger angle ranges is usually less. In practical applications, this can be addressed by utilizing... Figure 1 The angle panel diagram shown is used to mark the angles corresponding to each eye sample image. The marking results reveal that although the angle distribution of the collected eye sample images is wide-ranging, primarily concentrated within the 0-35 degree range, the number of eye sample images corresponding to small angle ranges is very large, while the number of eye sample images corresponding to large angle ranges is relatively small. Therefore, the imbalance in the number of eye sample images at different angles will affect the model accuracy to some extent, and the prediction accuracy of the trained model for eye sample images at large angles is relatively poor. To address this, this embodiment of the disclosure additionally calculates the angle range (i.e., the target angle range) to which each eye sample image belongs, thereby further determining the second weight corresponding to the training image group.

[0053] For ease of understanding, the embodiment of the present disclosure provides a way to determine the target angle interval corresponding to the eye sample image:

[0054] Exemplarily, assuming that the theoretical line-of-sight direction labeled by the eye sample image is represented by the yaw and pitch of the eyeball, the target angle angle corresponding to the eye sample image can be calculated based on this, which can be obtained with reference to the following formula:

[0055] angle = int(arctan(np.sqrt(tan(yaw)*tan(yaw)+tan(pitch)*tan(pitch))*180 / np.pi+0.5) where np.sqrt represents the square root operation, and np.pi represents the circular constant π. The above formula is only an example, which is based on the Pythagorean theorem in three-dimensional space and calculates the target angle angle based on the known yaw and pitch of the eyeball. The target angle can be used to represent the included angle between the line-of-sight direction when the eye is staring at the fixation point and the line-of-sight direction when the eye is horizontally staring at the fixation point. The above is only an example, and the calculated angle can also not be rounded, but can be directly represented in decimal form, which is not limited here.

[0056] In addition, based on the statistical analysis of the angle distribution of a large number of eye sample images obtained, a plurality of angle intervals can be set, such as setting the angle interval to 0-1 degrees, 1-2 degrees, 3-4 degrees, …, 30-31 degrees, or also can be set to 0-0.5 degrees, 0.5-1.5 degrees, 1.5-2.5 degrees, …, 30.5-31.5 degrees, or set to 0-5 degrees, 5-10 degrees, …, 30-35 degrees, etc. The specific angle interval can be flexibly set according to the needs, and different angle intervals can be equally spaced or non-equally spaced, which is not limited here. The angle interval in which the target angle corresponding to each eye sample image obtained according to the above method is located is the target angle interval corresponding to the eye sample image.

[0057] In some specific embodiments, steps A1-A2 can be performed with reference to:

[0058] Step A1, according to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group, determine the image weight corresponding to each of the two eye sample images in the training image group.

[0059] As mentioned above, the embodiments of the present disclosure fully consider the current problem of the uneven quantity of eye sample images of different angles, which also affects the model accuracy to some extent. Therefore, the embodiments of the present disclosure can further acquire a target angle interval corresponding to the eye sample images, and then set a corresponding image weight based on the target angle interval, so as to improve the problem of the uneven quantity of eye sample images of different angles in the model training process in a manner of setting the image weight. Exemplarily, the embodiments of the present disclosure further provide a determination manner of the image weight, which can be performed by referring to steps A1.1-A1.3 as follows:

[0060] In step A1.1, the quantity of eye sample images corresponding to each of the preset angle intervals is acquired. It can be understood that each eye sample image is labeled with a theoretical line of sight direction, and the corresponding angle interval can be determined in the foregoing manner. At this time, the quantity of eye sample images corresponding to each angle interval can be counted, and the quantity of eye sample images corresponding to the angle interval i can be Xi.

[0061] In step A1.2, the weight corresponding to the target angle interval of the theoretical line of sight direction of each of the two eye sample images in the training image group is determined according to the quantity of eye sample images corresponding to each of the preset angle intervals.

[0062] According to the quantity of eye sample images corresponding to each of the angle intervals obtained by counting, it can be known that the quantity of eye sample images corresponding to the target angle interval to which the eye sample images in the training image group belong is high or low, so that the corresponding image weight can be reasonably determined.

[0063] In some specific implementation examples, step A1.2 can be performed by referring to steps A1.2.1-A1.2.2 as follows:

[0064] In step A1.2.1, the maximum quantity is determined from the quantity of eye sample images corresponding to each of the preset angle intervals, and the quantity of eye sample images corresponding to the target angle interval of the theoretical line of sight direction of each of the two eye sample images in the training image group is determined. The maximum quantity is the maximum value in the quantity of eye sample images corresponding to each of the angle intervals. By acquiring the maximum quantity and the quantity of eye sample images corresponding to the target angle interval of the eye sample images, the quantity of images of the target angle interval can be reasonably and objectively measured based on the maximum quantity.

[0065] In step A1.2.2, the weight corresponding to the target angle interval is determined according to the maximum quantity, the quantity of eye sample images corresponding to the target angle interval, and a preset maximum image weight threshold. Specifically, step A1.2.2 can be performed by referring to steps 1)-3) as follows:

[0066] Step 1), obtain the ratio between the maximum number and the number of eye sample images corresponding to the target angle interval. Assuming that the maximum number is XB, and the number of eye sample images corresponding to the ith angle interval is Xi, then the ratio = XB / Xi. The smaller the ratio, the greater the number of eye sample images corresponding to the angle interval. The greater the ratio, the smaller the number of eye sample images corresponding to the angle interval.

[0067] Step 2), in the case where the ratio is not lower than the maximum image weight threshold, determining that the weight corresponding to the target angle interval is the maximum image weight threshold.

[0068] The embodiments of the present disclosure prevent the weight of eye sample images in a certain angle interval from being too large by limiting the maximum image weight threshold, which can weaken the influence of eye sample images in other angle intervals on model training, resulting in low accuracy of the model finally trained. When the ratio is not lower than the maximum image weight threshold, it means that the ratio is high and the number of eye sample images corresponding to the angle interval is small. At this time, the weight corresponding to the target angle interval can be directly set to the maximum image weight threshold, and the influence of eye sample images in the angle interval with a small number of images on model training can be effectively improved by setting the maximum image weight threshold.

[0069] Step 3), in the case where the ratio is lower than the maximum image weight threshold, determining that the weight corresponding to the target angle interval is the ratio. When the ratio is lower than the maximum image weight threshold, it means that the ratio is low and the number of eye sample images corresponding to the angle interval is large. At this time, the ratio can be directly used as the weight corresponding to the target angle interval. Moreover, the number of images corresponding to the angle interval is negatively correlated with the corresponding weight. The greater the number of images, the lower the corresponding weight. By increasing the weight of the angle interval with a small number of images, the influence of the imbalance of the number of images on model training can be effectively improved, which helps to further improve the prediction accuracy of the model for a large number of eye images.

[0070] Step A1.3, taking the weight corresponding to the target angle interval of the theoretical line of sight direction of each of the two eye sample images in the training image group as the image weight corresponding to each of the two eye sample images in the training image group.

[0071] In summary, by the above step A1, the image weight corresponding to each of the eye sample images in the training image group can be reasonably and objectively determined.

[0072] Step A2, obtaining a second weight corresponding to the training image group according to the image weight corresponding to each of the two eye sample images in the training image group.

[0073] For each training image group, on the basis of knowing the image weights of the two eye sample images in the training image group, the image weights corresponding to the two eye sample images in the training image group are averaged respectively, and the average processing result is taken as the second weight corresponding to the training image group.

[0074] Step B: obtaining the comprehensive weight corresponding to each of the plurality of training image groups based on the first weight and the second weight corresponding to each of the plurality of training image groups.

[0075] It can be understood that the first weight is set on the basis of the size of the line-of-sight difference between the two eye sample images in the training image group, and is used to alleviate the problem that the processing manner in the model training stage does not match the processing manner in the model testing or inference stage. In the model training stage, the influence of the eye sample image with smaller line-of-sight difference on the accuracy of the model prediction result is considered more sufficiently, and then the influence of the size of the line-of-sight difference on the model training is adjusted by setting the first weight. The second weight is set on the basis of the angle of the eye sample image in the training image group, and is used to alleviate the problem of imbalance in the number of eye sample images of different angles. Then the influence of the number of eye sample images of different angles on the model training is adjusted by setting the second weight.

[0076] The disclosure embodiments do not limit the manner of determining the comprehensive weight based on the first weight and the second weight. It can be a weighted manner or a product manner. In some specific implementation examples, the comprehensive weight corresponding to each of the plurality of training image groups can be obtained based on the product result of the first weight and the second weight corresponding to each of the plurality of training image groups.

[0077] Step C: weighting the loss between the line-of-sight difference theoretical value and the line-of-sight difference estimated value corresponding to each of the plurality of training image groups according to the comprehensive weight corresponding to each of the plurality of training image groups.

[0078] By taking into account the comprehensive weight obtained by the first weight and the second weight, the model training can be more objective and effective. The prediction accuracy of the line-of-sight difference estimated value between the eye images with smaller line-of-sight difference and the accuracy of the final prediction result of the eye images with larger angle can be effectively improved.

[0079] The disclosure embodiments also provide a line-of-sight estimation method, Figure 4 A flowchart of a line-of-sight estimation method provided by the disclosure embodiments is shown in the figure. The method can be executed by a line-of-sight estimation device, which can be implemented by software and / or hardware, and can be integrated in an electronic device. As shown in the figure, the method mainly includes the following steps S402-S406: Figure 4

[0080] ​At step S402, a calibration eye image labeled with a theoretical gaze direction and a target eye image to be estimated for the gaze direction are acquired. In actual applications, a prompt information for collecting the eye image can be initiated to the user before the eye image is collected, and the eye image is collected after the user authorization is acquired.

[0081] At step S404, a gaze difference estimation value between the calibration eye image and the target eye image is acquired by a pre-constructed target network model. The target network model is obtained based on the construction method of the network model, and details can be referred to the foregoing related content, which will not be repeated here.

[0082] At step S406, the gaze direction corresponding to the target eye image is estimated according to the theoretical gaze direction of the calibration eye image and the gaze difference estimation value between the calibration eye image and the target eye image.

[0083] Since the target network model trained by the embodiment of the present disclosure can output a reliable gaze difference estimation value, the accuracy of the gaze direction estimation for the target image is further ensured on this basis.

[0084] Exemplarily, referring to a gaze estimation flowchart shown in FIG. 1, Figure 5 It is shown that the calibration eye image labeled with the theoretical gaze direction and the target eye image are simultaneously input into the target network model, the target network model can output the gaze difference estimation value between the two, and the gaze direction corresponding to the target eye image can be obtained by adding the gaze difference estimation value and the gaze direction labeled by the calibration eye image. Further, the target network model can be a difference model, Figure 5 It is also shown that the target network model includes a feature extraction network and a gaze difference prediction network, the feature extraction network is used to extract the feature vectors of the calibration eye image and the target eye image respectively, and the gaze difference prediction network can analyze and process the feature vectors of the calibration eye image and the target eye image respectively, so as to generate the gaze difference estimation value. In specific implementation, the feature vectors of the calibration eye image and the target eye image can be spliced to obtain spliced features, and then the gaze difference prediction network is used to analyze and process the spliced features to obtain the gaze difference estimation value. The above is only an implementation example of the target network model, and should not be regarded as a limitation.

[0085] By the above method provided by the embodiment of the present disclosure, more accurate and reliable gaze direction estimation results can be obtained.

[0086] Corresponding to the construction method of the network model, Figure 6 A structure diagram of a network model construction device provided by the embodiment of the present disclosure is shown in FIG. 2, which can be realized by software and / or hardware, and can be integrated in an electronic device, such as a server. Figure 6As shown, the network model construction apparatus comprises:

[0087] An image group acquisition module 602 is configured to acquire a plurality of training image groups; wherein each of the training image groups comprises two eye sample images, and the eye sample images are labeled with theoretical gaze directions;

[0088] A first weight determination module 604 is configured to obtain a theoretical value of a gaze difference corresponding to each of the training image groups according to an angle difference between the theoretical gaze directions labeled by the two eye sample images in the training image group, and determine a first weight corresponding to the training image group according to the theoretical value of the gaze difference.

[0089] A network model output module 606 is configured to acquire gaze difference estimation values respectively output by a preset initial network model for the plurality of training image groups; wherein the gaze difference estimation values are obtained by the initial network model estimating an angle difference between the gaze directions of the two eye sample images in the training image groups.

[0090] A weighted loss acquisition module 608 is configured to perform weighted processing on a loss between the theoretical value of the gaze difference and the gaze difference estimation value corresponding to each of the plurality of training image groups according to the first weight corresponding to each of the plurality of training image groups, to obtain a weighted loss.

[0091] A network model training module 610 is configured to train the initial network model based on the weighted loss, to construct a target network model based on the trained initial network model; wherein the target network model is configured to generate a gaze difference estimation value between two target eye images.

[0092] The above apparatus can fully consider the influence of the size of the gaze difference between two eye images on the output result of the model, and determine a corresponding first weight for the theoretical value of the gaze difference corresponding to each of the training image groups in the model training process. By introducing the influence of the size of the gaze difference on the model through the first weight, the gaze difference estimation result output by the finally obtained target network model can be more accurate.

[0093] In some embodiments, for any two of the training image groups, the first weight corresponding to the training image group with a smaller theoretical value of the gaze difference is not lower than the first weight corresponding to the training image group with a larger theoretical value of the gaze difference.

[0094] In some embodiments, the first weight determination module 604 is specifically configured to: acquire a preset gaze difference threshold; wherein the number of the gaze difference thresholds is one or more; compare the theoretical value of the gaze difference of the training image group with the gaze difference threshold, to determine the first weight corresponding to the training image group based on a comparison result.

[0095] In some embodiments, the first weight determination module 604 is specifically configured to: in a case where the theoretical value of the line-of-sight difference of the training image set is not higher than a first line-of-sight difference threshold, determine the first weight corresponding to the training image set as a preset first numerical value; in a case where the theoretical value of the line-of-sight difference of the training image set is between the first line-of-sight difference threshold and a second line-of-sight difference threshold, determine the second weight corresponding to the training image set according to a preset algorithm; wherein the second weight is between a preset second numerical value and the first numerical value, the second numerical value is smaller than the first numerical value, and the second weight is negatively correlated with the theoretical value of the line-of-sight difference; and in a case where the theoretical value of the line-of-sight difference of the training image set is not lower than the second line-of-sight difference threshold, determine the first weight corresponding to the training image set as the second numerical value.

[0096] In some embodiments, the preset algorithm includes a linear algorithm.

[0097] In some embodiments, the weighted loss obtaining module 608 is specifically configured to: determine the second weight corresponding to the training image set according to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image set; obtain the comprehensive weight corresponding to each of the plurality of training image sets based on the first weight and the second weight corresponding to each of the plurality of training image sets; and perform weighted processing on the loss between the theoretical value of the line-of-sight difference and the estimated value of the line-of-sight difference corresponding to each of the plurality of training image sets according to the comprehensive weight corresponding to each of the plurality of training image sets.

[0098] In some embodiments, the weighted loss obtaining module 608 is specifically configured to: determine the image weight corresponding to each of the two eye sample images in the training image set according to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image set; and obtain the second weight corresponding to the training image set according to the image weight corresponding to each of the two eye sample images in the training image set.

[0099] In some embodiments, the weighted loss obtaining module 608 is specifically configured to: obtain the number of eye sample images corresponding to each of a plurality of preset angle intervals; determine the weight corresponding to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image set according to the number of eye sample images corresponding to each of the plurality of preset angle intervals; and take the weight corresponding to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image set as the image weight corresponding to each of the two eye sample images in the training image set.

[0100] In some embodiments, the weighted loss obtaining module 608 is specifically configured to: determine a maximum number from the preset numbers of eye sample images corresponding to the respective angle intervals, and determine numbers of eye sample images corresponding to the target angle interval to which the respective theoretical gaze directions of the two eye sample images in the training image group correspond; and determine the weight corresponding to the target angle interval according to the maximum number, the number of eye sample images corresponding to the target angle interval, and a preset maximum image weight threshold.

[0101] In some embodiments, the weighted loss obtaining module 608 is specifically configured to: obtain a ratio between the maximum number and the number of eye sample images corresponding to the target angle interval; in a case where the ratio is not lower than the maximum image weight threshold, determine the weight corresponding to the target angle interval as the maximum image weight threshold; and in a case where the ratio is lower than the maximum image weight threshold, determine the weight corresponding to the target angle interval as the ratio.

[0102] In some embodiments, the weighted loss obtaining module 608 is specifically configured to: average the respective image weights corresponding to the two eye sample images in the training image group to obtain an average processing result as a second weight corresponding to the training image group.

[0103] In some embodiments, the weighted loss obtaining module 608 is specifically configured to: obtain a comprehensive weight corresponding to each of the plurality of training image groups based on a product result of the first weight and the second weight corresponding to each of the plurality of training image groups.

[0104] The network model construction apparatus provided by the embodiments of the present disclosure can execute the network model construction method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.

[0105] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the apparatus embodiments described above can refer to the corresponding process in the method embodiments, which will not be described here.

[0106] Corresponding to the aforementioned gaze estimation method, Figure 7 A structural diagram of a gaze estimation apparatus provided by the embodiments of the present disclosure is shown in the figure. The apparatus can be implemented by software and / or hardware, and can be integrated in an electronic device, such as a mobile phone. Figure 7 As shown in the figure, the gaze estimation apparatus includes:

[0107] An image obtaining module 702 is configured to obtain a calibration eye image labeled with a theoretical gaze direction and a target eye image whose gaze direction is to be estimated.

[0108] The line-of-sight difference estimation module 704 is configured to obtain a line-of-sight difference estimation value between the calibration eye image and the target eye image by using a pre-constructed target network model, wherein the target network model is obtained based on the network model construction method described above.

[0109] The line-of-sight direction estimation module 706 is configured to estimate the line-of-sight direction corresponding to the target eye image according to the theoretical line-of-sight direction of the calibration eye image and the line-of-sight difference estimation value between the calibration eye image and the target eye image.

[0110] Since the target network model obtained by training in the embodiments of the present disclosure can output more accurate line-of-sight difference estimation values, the accuracy of line-of-sight direction estimation for target images is further ensured on this basis.

[0111] The line-of-sight direction estimation apparatus provided in the embodiments of the present disclosure can perform the line-of-sight direction estimation method provided in any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of performing the method.

[0112] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the apparatus embodiments described above can refer to the corresponding process in the method embodiments, which will not be described here.

[0113] The embodiments of the present disclosure also provide an electronic device, which includes a processor, a memory for storing processor-executable instructions, and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the network model construction method or the line-of-sight estimation method described above.

[0114] Figure 8 A structural schematic diagram of an electronic device provided in the embodiments of the present disclosure is shown in FIG. 8. Figure 8 As shown in FIG. 8, the electronic device 800 includes one or more processors 801 and a memory 802.

[0115] The processor 801 can be a central processing unit (CPU) or other forms of processing units having data processing and / or instruction execution capabilities, and can control other components in the electronic device 800 to perform desired functions.

[0116] The memory 802 can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read-only memory (ROM), hard disk drives, solid-state drives, and / or the like. The computer-readable storage media can store one or more computer program instructions that, when executed by the processor 801, implement the above-described method of constructing a network model or the above-described method of line-of-sight estimation provided by embodiments of the present disclosure and / or other desired functions. Various contents such as input signals, signal components, noise components, and the like can also be stored in the computer-readable storage media.

[0117] In one example, the electronic device 800 can further include an input device 803 and an output device 804, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0118] In addition, the input device 803 can further include, for example, a keyboard, a mouse, and the like.

[0119] The output device 804 can output various information to the outside, including determined distance information, direction information, and the like. The output device 804 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.

[0120] Of course, in order to simplify, Figure 8 Only some of the components of the electronic device 800 related to the present disclosure are shown in the figure, and components such as buses, input / output interfaces, and the like are omitted. In addition, the electronic device 800 can further include any other appropriate components according to specific application cases.

[0121] In addition to the above-described method and device, embodiments of the present disclosure can also be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform the above-described method of constructing a network model or the above-described method of line-of-sight estimation provided by embodiments of the present disclosure.

[0122] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0123] In addition, the embodiments of the present disclosure can also be a computer readable storage medium, which stores computer program instructions, and the computer program instructions make the processor execute the above-mentioned network model construction method or the above-mentioned line-of-sight estimation method provided by the embodiments of the present disclosure when the processor runs.

[0124] The computer readable storage medium can take any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (non-exhaustive list) of readable storage medium include: electrical connections having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0125] The embodiments of the present disclosure also provide a computer program product, which includes computer programs / instructions, and the computer programs / instructions are executed by a processor to realize the XYZ method in the embodiments of the present disclosure.

[0126] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario and the like should be informed to the user and the authorization of the user should be obtained according to relevant laws and regulations through appropriate means.

[0127] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using personal information of the user. For example, in collecting an eye image of the user, a collection prompt information of the eye image is sent to the user, and the user can be further informed of the image collection purpose in the prompt information. Thus, the user can autonomously select whether to provide personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information, and the operation of the technical solution of the present disclosure is performed after obtaining the authorization of the user.

[0128] As an optional but non-limiting implementation, in response to receiving an active request of a user, the prompt information can be sent to the user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select “agree” or “disagree” to provide personal information to the electronic device.

[0129] It can be understood that the above notification and obtaining of user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0130] It should be noted that in this paper, the relationship terms such as “first” and “second” are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement “including a…” does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0131] The above is only a specific embodiment of the present disclosure, which enables those skilled in the art to understand or implement the present disclosure. Various modifications of these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of constructing a network model, characterized by, The method comprises: obtaining a plurality of training image groups; wherein each of the training image groups comprises two eye sample images, and the two eye sample images are labeled with theoretical gaze directions; obtaining a theoretical value of a gaze difference corresponding to each of the training image groups according to an angle difference between the theoretical gaze directions labeled by the two eye sample images in the training image group, and determining a first weight corresponding to the training image group according to the theoretical value of the gaze difference; obtaining a gaze difference estimation value output by a preset initial network model for each of the plurality of training image groups; wherein the gaze difference estimation value is obtained by the initial network model estimating an angle difference between the theoretical gaze directions of the two eye sample images in the training image group; performing weighted processing on a loss between the theoretical value of the gaze difference and the gaze difference estimation value corresponding to each of the plurality of training image groups according to the first weight corresponding to each of the plurality of training image groups, to obtain a weighted loss; training the initial network model based on the weighted loss, and constructing a target network model based on the trained initial network model; wherein the target network model is used to generate a gaze difference estimation value between two target eye images.

2. The method of claim 1, wherein, For any two of the training image groups, the first weight corresponding to the training image group with a smaller theoretical value of the gaze difference is not less than the first weight corresponding to the training image group with a larger theoretical value of the gaze difference.

3. The method of claim 1, wherein, The method further comprises: obtaining a preset threshold value of the gaze difference; wherein the number of the threshold values of the gaze difference is one or more; comparing the theoretical value of the gaze difference of the training image group with the threshold value of the gaze difference, to determine the first weight corresponding to the training image group based on a comparison result.

4. The method of claim 3, wherein, The method further comprises: in a case where the theoretical value of the gaze difference of the training image group is not higher than a first threshold value of the gaze difference, determining the first weight corresponding to the training image group as a preset first value; in a case where the theoretical value of the gaze difference of the training image group is between the first threshold value of the gaze difference and a second threshold value of the gaze difference, determining a second weight corresponding to the training image group according to a preset algorithm; wherein the second weight is between a preset second value and the first value, the second value is smaller than the first value, and the second weight is negatively correlated with the theoretical value of the gaze difference; in a case where the theoretical value of the gaze difference of the training image group is not lower than the second threshold value of the gaze difference, determining the first weight corresponding to the training image group as the second value.

5. The method of claim 4, wherein, The preset algorithm comprises a linear algorithm.

6. The method according to any one of claims 1 to 5, characterized in that, The method further comprises: determining the second weight corresponding to the training image group according to a target angle interval corresponding to the theoretical gaze direction of each of the two eye sample images in the training image group; obtaining a comprehensive weight corresponding to each of the plurality of training image groups based on the first weight and the second weight corresponding to each of the plurality of training image groups. The loss between the theoretical value of the line-of-sight difference and the estimated value of the line-of-sight difference of each of the plurality of training image groups is weighted according to the comprehensive weight corresponding to each of the plurality of training image groups.

7. The method of claim 6, wherein, The second weight corresponding to the training image group is determined according to a target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group. The image weight corresponding to each of the two eye sample images in the training image group is determined according to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group. The second weight corresponding to the training image group is obtained according to the image weight corresponding to each of the two eye sample images in the training image group.

8. The method of claim 7, wherein, The image weight corresponding to each of the two eye sample images in the training image group is determined according to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group. The number of eye sample images corresponding to each of the plurality of preset angle intervals is obtained. The weight corresponding to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group is determined according to the number of eye sample images corresponding to each of the plurality of preset angle intervals. The weight corresponding to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group is taken as the image weight corresponding to each of the two eye sample images in the training image group.

9. The method of claim 8, wherein, The weight corresponding to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group is determined according to the number of eye sample images corresponding to each of the plurality of preset angle intervals. The maximum number is determined from the number of eye sample images corresponding to each of the plurality of preset angle intervals, and the number of eye sample images corresponding to the target angle interval corresponding to the theoretical line-of-sight direction of each of the two eye sample images in the training image group is determined. The weight corresponding to the target angle interval is determined according to the maximum number, the number of eye sample images corresponding to the target angle interval, and a preset maximum image weight threshold.

10. The method of claim 9, wherein, The weight corresponding to the target angle interval is determined according to the maximum number, the number of eye sample images corresponding to the target angle interval, and a preset maximum image weight threshold. The ratio between the maximum number and the number of eye sample images corresponding to the target angle interval is obtained. In a case where the ratio is not lower than the maximum image weight threshold, the weight corresponding to the target angle interval is determined as the maximum image weight threshold. In a case where the ratio is lower than the maximum image weight threshold, the weight corresponding to the target angle interval is determined as the ratio.

11. The method of claim 7, wherein, The second weight corresponding to the training image group is obtained according to the image weight corresponding to each of the two eye sample images in the training image group. The image weight corresponding to each of the two eye sample images in the training image group is averaged to take the average processing result as the second weight corresponding to the training image group.

12. The method of claim 6, wherein, The obtaining of the comprehensive weight corresponding to each of the plurality of training image groups based on the first weight and the second weight corresponding to each of the plurality of training image groups comprises: The comprehensive weight corresponding to each of the plurality of training image groups is obtained based on a product result of the first weight and the second weight corresponding to each of the plurality of training image groups.

13. A line of sight estimation method, characterized by, Comprise: Obtaining a calibration eye image labeled with a theoretical gaze direction and a target eye image whose gaze direction is to be estimated; Obtaining a gaze difference estimation value between the calibration eye image and the target eye image through a pre-constructed target network model; wherein the target network model is obtained based on the network model construction method of any one of claims 1 to 12; Estimating the gaze direction corresponding to the target eye image according to the theoretical gaze direction of the calibration eye image and the gaze difference estimation value between the calibration eye image and the target eye image.

14. A network model construction apparatus characterized by comprising: Comprise: An image group acquisition module is configured to acquire a plurality of training image groups; wherein each of the training image groups contains two eye sample images, and each of the eye sample images is labeled with a theoretical gaze direction; A first weight determination module is configured to obtain a theoretical gaze difference value corresponding to each of the training image groups according to an angle difference between the theoretical gaze directions labeled by the two eye sample images in the training image group, and determine a first weight corresponding to each of the training image groups according to the theoretical gaze difference value; A network model output module is configured to obtain gaze difference estimation values respectively output by a preset initial network model for the plurality of training image groups; wherein the gaze difference estimation values are obtained by the initial network model for estimating an angle difference between the gaze directions of the two eye sample images in the training image group; A weighted loss acquisition module is configured to perform weighted processing on a loss between a theoretical gaze difference value and a gaze difference estimation value corresponding to each of the plurality of training image groups according to a first weight corresponding to each of the plurality of training image groups, to obtain a weighted loss; A network model training module is configured to train the initial network model based on the weighted loss, to construct a target network model based on the trained initial network model; wherein the target network model is used to generate a gaze difference estimation value between two target eye images.

15. A line of sight estimation apparatus characterized by comprising: Comprise: An image acquisition module is configured to acquire a calibration eye image labeled with a theoretical gaze direction and a target eye image whose gaze direction is to be estimated; A gaze difference estimation module is configured to obtain a gaze difference estimation value between the calibration eye image and the target eye image through a pre-constructed target network model; wherein the target network model is obtained based on the network model construction method of any one of claims 1 to 12; A gaze direction estimation module is configured to estimate a gaze direction corresponding to the target eye image according to the theoretical gaze direction of the calibration eye image and the gaze difference estimation value between the calibration eye image and the target eye image.

16. An electronic device, comprising: The electronic device comprises: A processor; A memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the network model construction method according to any one of claims 1-12 or the line-of-sight estimation method according to claim 13.

17. A computer-readable storage medium, characterized in that, The storage medium stores a computer program configured to implement the network model construction method according to any one of claims 1-12 or the line-of-sight estimation method according to claim 13.

Citation Information

Patent Citations

  • Sight line prediction method, device and system and readable storage medium

    CN110008835A

  • Training method of image processing network, and image processing method and device

    CN115359547A