Gaze direction estimation method and device, equipment and storage medium

By combining a main-view camera and an auxiliary camera, and using multi-view images to correct the gaze direction, the problem of high cost and insufficient accuracy of existing eye-tracking technologies is solved, and low-cost, high-precision gaze direction estimation is achieved.

CN122044337APending Publication Date: 2026-05-15GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU SHIYUAN ELECTRONICS CO LTD
Filing Date
2024-11-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing infrared eye-tracking technologies are costly, while three-dimensional geometry-based methods are difficult to guarantee accuracy due to the diversity of human eyeball shapes. Feature-based learning methods are affected by environmental and facial expression interference and are difficult to obtain high-precision training data, resulting in high cost and insufficient accuracy in gaze direction estimation.

Method used

The initial gaze direction is estimated by capturing facial images of the user with the main camera, and the position of the gaze object is estimated by capturing environmental images within the user's field of vision with the auxiliary camera. The initial gaze direction is corrected by using multi-view images, and the feature extraction and facial landmarks are trained using a pre-set network model to achieve accurate correction of the gaze direction.

Benefits of technology

It achieves reduced gaze direction estimation costs without adding extra sensors, and improves the accuracy and precision of gaze direction through multi-view data, making it suitable for multiple application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122044337A_ABST
    Figure CN122044337A_ABST
Patent Text Reader

Abstract

The invention provides a gazing direction estimation method and device, equipment and a storage medium, applied to target equipment, the target equipment is connected with a main view camera and at least one auxiliary camera, the main view camera is arranged facing a user, and the auxiliary camera is arranged facing the gazing direction of the user; the method comprises the steps that a main view image and an auxiliary image are acquired, the main view image is a target face image shot by a main view camera, the target face image is an image of the face of a target user, and the auxiliary image is an image of a target environment shot by an auxiliary camera; performing user gazing direction estimation on the front view image to obtain an initial gazing direction of a target user in the front view image; performing gaze object position estimation on the auxiliary image to obtain an object position of a user gaze object in the auxiliary image, the user gaze object being an object gazed by the target user; and correcting the initial gazing direction according to the object position to obtain a target gazing direction of the target user. According to the technical scheme, fixation point tracking is realized through a common camera, and the application cost is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gaze tracking technology, and more particularly to gaze direction estimation methods, apparatus, devices, and storage media. Background Technology

[0002] Foveation tracking, also known as eye-tracking technology, is a technique used to detect and analyze the direction of human eye gaze. It has been widely applied in various fields, such as virtual reality (VR), human-computer interaction, advertising design, and finance.

[0003] The mainstream eye-tracking technology typically uses infrared technology to determine gaze direction. This method involves emitting infrared light through a device integrated with an infrared light source. The reflected light is then captured by an infrared camera, and the pupil center is determined by analyzing the image generated from the reflected light. However, the infrared method requires an infrared camera, making it relatively expensive. Summary of the Invention

[0004] This application provides a gaze direction estimation method, apparatus, device, and storage medium for determining gaze direction.

[0005] In a first aspect, a gaze direction estimation method is provided, applied to a target device, the target device being connected to a main gaze camera and at least one auxiliary camera, the main gaze camera being oriented towards a user, and the auxiliary camera being oriented towards the user's gaze direction; the method includes:

[0006] Acquire a front view image and an auxiliary image. The front view image is a target face image captured by the front view camera, which is an image of the target user's face. The auxiliary image is an image of the target environment captured by the auxiliary camera, which is the environment within the target user's field of vision.

[0007] The user gaze direction is estimated by performing a user gaze direction estimation on the front view image to obtain the initial gaze direction of the target user in the front view image;

[0008] The position of the gaze object is estimated by performing gaze object position estimation on the auxiliary image to obtain the position of the user gaze object in the auxiliary image, wherein the user gaze object is an object in the target environment that is gazed at by the target user;

[0009] Based on the object's position, the initial gaze direction is corrected to obtain the target user's target gaze direction.

[0010] In this technical solution, after acquiring the main view image and the auxiliary image, the user's gaze direction is estimated on the main view image to obtain the initial gaze direction of the target user in the main view image. Then, the position of the object being gazed upon is estimated on the auxiliary image to obtain the position of the object being gazed upon by the user in the auxiliary image. Finally, based on the position of the object being gazed upon, the initial gaze direction is corrected to obtain the target gaze direction of the target user in the main view image, thus realizing the detection and correction of the user's gaze direction. Since the main view image is a facial image of the user captured by the main view camera, which is set to face the user, estimating the user's gaze direction on the main view image can achieve a rough estimate of the user's gaze in the image. Since the auxiliary image is a view of the user's field of vision captured by the auxiliary camera... The system uses an image of the environment and an auxiliary camera oriented towards the user's gaze direction. By estimating the position of the object being gazed upon in the auxiliary image, the system obtains the position of the object being gazed upon in the auxiliary image. The user's gaze direction is then corrected using the position of the object being gazed upon. The object being gazed upon is the object in the environment that the user is looking at. The position of the object being gazed upon can indirectly reflect the user's gaze direction. This is equivalent to using multi-view images to help determine the user's gaze direction, which can correct the recognition results of the main view image, thereby ensuring the accuracy of the user's gaze direction. Since the image is obtained and the user's gaze direction is estimated by combining the main view camera and the auxiliary camera, gaze point tracking can be achieved with ordinary cameras without the need for additional sensors, thus saving application costs.

[0011] In conjunction with the first aspect, the number of object positions is multiple; the step of correcting the initial gaze direction based on the object positions to obtain the target user's target gaze direction includes: determining multiple object position matching values ​​based on the initial gaze direction and the multiple object positions, wherein the multiple object position matching values ​​are object position matching values ​​corresponding to the multiple object positions, and the object position matching values ​​are used to reflect the degree of proximity between the object positions and the initial gaze direction; and determining the target user's target gaze direction based on the multiple object positions and the multiple object position matching values.

[0012] When there are multiple estimated object positions, determining the user's target gaze direction based on the object position and the object position matching value, which reflects the proximity of the object position to the initial gaze direction, helps to accurately determine the user's target gaze direction.

[0013] In conjunction with the first aspect, in one possible implementation, determining multiple object position matching values ​​based on the initial gaze direction and multiple object positions includes: calculating the projection distance of the target position vector relative to the initial gaze position vector to obtain the target projection distance, wherein the target position vector is the position vector of the target object position, the target object position is any one of the multiple object positions, and the initial gaze position vector is the position vector of the initial gaze direction; and determining the object position matching value corresponding to the target object position based on the target projection distance.

[0014] Determining the object position matching value based on the projection distance of the object's position vector relative to the initial direction position vector ensures that the object position matching value is reasonable and sufficient to reflect the closeness between the object's position and the initial gaze direction.

[0015] In conjunction with the first aspect, in one possible implementation, determining the target user's target gaze direction based on the plurality of object positions and the plurality of object position matching values ​​includes: performing position clustering on the plurality of object positions to obtain at least one object position set; determining a set matching value corresponding to each object position set based on the object position matching value corresponding to the object position in each object position set; and determining the target user's target gaze direction based on the object position in the object position set corresponding to the largest set matching value.

[0016] By clustering object positions and determining the target user's gaze direction based on the object positions in the set with the largest matching value, the accuracy of the target user's gaze direction can be guaranteed.

[0017] In conjunction with the first aspect, in one possible implementation, determining the set matching value corresponding to each object location set based on the object location matching value corresponding to the object location in each object location set includes: calculating the mean of the object location matching values ​​corresponding to the object locations in the first object location set and / or the sum of the object location matching values ​​corresponding to the object locations in the first object location set to obtain the mean of the target matching values ​​and / or the sum of the target matching values, wherein the first object location set is any one of the at least one object location sets; and determining the set matching value corresponding to the first object location set based on the mean of the target matching values ​​and / or the sum of the target matching values.

[0018] The set matching value corresponding to the set of object locations is determined by the mean and / or sum of the matching values ​​of the object locations in the set of object locations, which can reasonably determine the set matching value of the set of object locations.

[0019] In conjunction with the first aspect, in one possible implementation, determining the target user's target gaze direction based on the object positions in the object position set corresponding to the maximum set matching value includes: determining the position center of the object positions in the second object position set, where the second object position set is the object position set corresponding to the maximum set matching value, and the spatial position coordinates of the position center are the average of the spatial position coordinates of the object positions in the second object position set; and determining the position vector of the position center as the position vector of the target user's target gaze direction.

[0020] The position vector of the center of the object in the set of object positions with the largest matching value is determined as the position vector of the target user's gaze direction, which can accurately determine the user's gaze direction.

[0021] In conjunction with the first aspect, in one possible implementation, before performing position clustering on the plurality of object positions to obtain at least one set of object positions, the method further includes: filtering out noisy object positions from the plurality of object positions, wherein the noisy object positions are object positions whose object position matching values ​​are less than a preset threshold.

[0022] By filtering out the positions of noisy objects, interference from noise on the user's gaze direction can be avoided, thus improving the accuracy of gaze tracking.

[0023] In conjunction with the first aspect, in one possible implementation, the step of estimating the user gaze direction of the main view image to obtain the initial gaze direction of the target user in the main view image includes: inputting the main view image into a gaze direction estimation model to perform user gaze direction estimation to obtain the initial gaze direction of the target user in the main view image.

[0024] By using a gaze direction estimation model to estimate the user's gaze direction, a preliminary estimate of the user's gaze direction can be achieved.

[0025] In conjunction with the first aspect, in one possible implementation, the gaze direction estimation model is trained by a preset network model, which includes a feature extraction network, a first output branch, and a second output branch. The first and second output branches are respectively connected to the feature extraction network. The feature extraction network is used to extract features from a sample facial image to obtain image features, and the sample facial image is a facial image used as a training sample. The first output branch is used to output the user's gaze direction in the sample facial image based on the image features. The second output branch is used to output facial landmarks of the user in the sample facial image based on the image features. The facial landmarks are used to adjust the parameters of the feature extraction network during the training process of the preset network model. The gaze direction estimation model includes the trained feature extraction network and the first output branch.

[0026] During the training of the gaze direction estimation model, a second output branch that outputs facial landmarks is introduced as an auxiliary training output branch. Facial landmarks provide information about the eye region and facial expressions. These landmarks are used to adjust the parameters of the feature extraction network during training, enabling the feature extraction network to learn eye region features and facial expression features in the image. In other words, the feature extraction network can learn and extract richer and more accurate image features, thus allowing the gaze direction estimation model to output a more accurate gaze direction based on the image features extracted by the feature extraction network, ensuring the accuracy of the gaze direction estimation model. Introducing a second output branch that outputs facial landmarks as an auxiliary training output branch also accelerates model convergence.

[0027] In conjunction with the first aspect, in one possible implementation, estimating the position of the gaze object in the auxiliary image to obtain the object position of the user's gaze object in the auxiliary image includes: extracting feature points from the auxiliary image to obtain an auxiliary feature point set, the auxiliary feature point set including multiple feature points in the auxiliary image; determining feature points in the auxiliary feature point set that match a preset gaze environment image to obtain gaze object feature points; and determining the spatial coordinates corresponding to the gaze object feature points to obtain the object position.

[0028] By matching the feature points obtained from feature point extraction of the auxiliary image with the preset gaze environment image to determine the gaze object feature points, it is helpful to accurately determine the gaze object feature points, thereby accurately determining the object position.

[0029] In conjunction with the first aspect, in one possible implementation, determining the spatial coordinates corresponding to the gaze feature point to obtain the object position includes: acquiring projection transformation parameters between the auxiliary camera and the main camera, the projection transformation parameters reflecting the pose of the auxiliary camera relative to the main camera; determining the pixel position of the gaze feature point in the auxiliary image to obtain a first pixel position; determining the spatial coordinates that minimize the target difference as the spatial coordinates corresponding to the gaze feature point, the target difference being the difference between a second pixel position and the first pixel position, the second pixel position being the pixel position obtained by transforming the spatial coordinates using the projection transformation parameters.

[0030] Determining the spatial coordinates that minimize the target difference as the spatial coordinates corresponding to the feature point of the object being gazed upon can improve the accuracy of the spatial coordinates corresponding to the feature point of the object being gazed upon.

[0031] In a second aspect, a gaze direction estimation device is provided, applied to a target device, the target device being connected to a main gaze camera and at least one auxiliary camera, the main gaze camera being oriented towards a user, and the auxiliary camera being oriented towards the user's gaze direction; the device includes:

[0032] The image acquisition module is used to acquire a front view image and an auxiliary image. The front view image is a target face image captured by the front view camera, which is an image of the target user's face. The auxiliary image is an image of the target environment captured by the auxiliary camera, which is the environment within the target user's field of vision.

[0033] A gaze direction estimation module is used to estimate the user's gaze direction in the main view image to obtain the initial gaze direction of the target user in the main view image;

[0034] The gaze object position estimation module is used to estimate the gaze object position of the auxiliary image to obtain the object position of the user gaze object in the auxiliary image, wherein the user gaze object is an object in the target environment that is gazed at by the target user;

[0035] The gaze direction correction module is used to correct the initial gaze direction based on the object's position to obtain the target gaze direction of the target user.

[0036] Thirdly, a computer device is provided, including a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, wherein when the processor executes the one or more computer programs, the computer device enables the gaze direction estimation method of the first aspect described above.

[0037] Fourthly, a computer-readable storage medium is provided, which stores a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the gaze direction estimation method of the first aspect.

[0038] This application achieves the following technical effects: it enables the detection and correction of the user's gaze direction; since the main view image is a facial image of the user captured by the main view camera, which is oriented towards the user, by estimating the user's gaze direction in the main view image, a rough estimate of the user's gaze in the image can be achieved; since the auxiliary image is an image of the user's gaze environment captured by the auxiliary camera, which is oriented towards the user's gaze direction, by estimating the position of the object being gazed upon in the auxiliary image, the position of the object being gazed upon in the auxiliary image can be obtained, and the user's gaze direction can be corrected using the position of the object being gazed upon. The object being gazed upon is an object in the environment that the user is gazing upon, and the position of the object being gazed upon can indirectly reflect the user's gaze direction. This is equivalent to using multi-view images to assist in determining the user's gaze direction, which can correct the recognition result of the main view image, thereby ensuring the accuracy of the user's gaze direction; since the image is obtained and the user's gaze direction is estimated by combining the main view camera and the auxiliary camera, gaze point tracking can be achieved with ordinary cameras without the need for additional sensors, thus saving application costs. Attached Figure Description

[0039] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 A schematic diagram showing the positions of the main camera and the auxiliary camera provided in an embodiment of this application;

[0041] Figure 2 A schematic flowchart illustrating a gaze direction estimation method provided in an embodiment of this application;

[0042] Figure 3 A schematic diagram of a preset network model and a gaze direction estimation model provided for embodiments of this application;

[0043] Figure 4 A process step for correcting the initial gaze direction is provided in an embodiment of this application;

[0044] Figure 5 This is a schematic diagram of the structure of a gaze direction estimation device provided in an embodiment of this application;

[0045] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0047] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0048] The technical solution of this application is applicable to gaze tracking scenarios.

[0049] In gaze tracking scenarios, gaze tracking can be achieved through the following methods:

[0050] (1) Infrared method: Infrared light is emitted by an emitting device that integrates an infrared light source. The infrared light is reflected at the cornea of ​​the eye and then captured by an infrared camera to form an image. The direction of the gaze point is determined by analyzing the image captured by the infrared camera.

[0051] (2) Three-dimensional geometry-based method: By fitting a three-dimensional eyeball model, the eyeball center, radius and other eye features are determined, and then the line of sight is estimated by combining geometric relationships.

[0052] (3) Feature-based learning method: Learn the mapping relationship between high-dimensional features and low-dimensional gaze in eye images through machine learning, obtain gaze estimation model, and perform gaze estimation through gaze estimation model.

[0053] Among the methods mentioned above, the infrared method relies on an infrared camera, which requires additional sensors and has a high application cost; the three-dimensional geometry-based method is affected by the diversity of human eyeball shapes, making it difficult to guarantee the accuracy of gaze estimation; the feature learning-based method is affected by various factors such as environment and facial expressions, making it difficult to guarantee the accuracy of the gaze estimation model. In addition, a large amount of high-precision gaze point data is required when training the gaze estimation model, and high-precision gaze point data itself is not easy to obtain, so it is difficult to truly implement and apply.

[0054] In view of this, this application proposes a gaze direction estimation scheme. A front-view camera, positioned facing the user, captures an image of the user's face to obtain a front-view image. The user's gaze direction is estimated on the front-view image to obtain the user's initial gaze direction, thus achieving a preliminary estimation of the user's gaze direction. At least one auxiliary camera, positioned facing the user's gaze direction, captures an image of the environment within the user's field of vision to obtain an auxiliary image. The position of the object being gazed upon is estimated on the auxiliary image to obtain the position of the object being gazed upon. Finally, the user's initial gaze direction is corrected based on the position of the object being gazed upon to obtain the user's target gaze direction. The position of the object being gazed upon in the auxiliary image can indirectly reflect the user's gaze direction. By combining the position of the object being gazed upon to correct the user's initial gaze direction, accurate estimation of the user's gaze direction can be achieved. Since the user's gaze direction is determined through image recognition, ordinary cameras can be used to estimate the user's gaze direction without the need for special sensors such as infrared cameras, thus saving application costs. By using the position of the object being gazed upon in the auxiliary image to correct the user's initial gaze direction, and combining multi-view data to comprehensively determine the user's gaze direction, the accuracy of single-view image recognition results can be improved. Since the user's gaze direction is determined by combining multi-view data, the initial gaze direction can be a rough and imprecise one. The accuracy requirements for the gaze point data used for training are low, making it easy to implement and apply.

[0055] The technical solution of this application is described in detail below. The technical solution of this application is applied to a target device, which is connected to a main camera and at least one auxiliary camera. The positional relationship between the main camera and at least one auxiliary camera can be referred to... Figure 1The main camera 101 is positioned facing the user and is used to capture images of the user's face. The auxiliary camera 102 is positioned facing the user's gaze direction and is used to capture images of the environment the user is looking at. The target device can be independent of the main camera 101 and the auxiliary camera 102. The target device can be a computer device or control module connected to the main camera 101 and the auxiliary camera 102, such as a control chip in a car cockpit system. Alternatively, the target device can be a device that includes both the main camera 101 and the auxiliary camera 102, such as a car. The connection between the target device and the main camera and the auxiliary camera can be wired or wireless. This application does not limit the specific implementation of the target device or the specific connection relationship with the main camera and the auxiliary camera.

[0056] See Figure 2 , Figure 2 A flowchart illustrating a gaze direction estimation method provided in this application embodiment is shown below. Figure 2 As shown, the method includes the following steps:

[0057] S201, acquire the main view image and auxiliary image.

[0058] Here, the main view image is the target face image captured by the main view camera; the target face image is the image of the target user's face, and the target user is the subject captured by the main view camera. The auxiliary image is the target environment image captured by the auxiliary camera; the target environment is the environment within the target user's field of view. The meaning and positional relationship of the main view camera and auxiliary cameras can be found in the preceding description and will not be repeated here. It is understood that when there are multiple auxiliary cameras, the number of acquired auxiliary images is multiple, and these multiple auxiliary images are images of the environment within the user's field of view captured by the multiple auxiliary cameras.

[0059] S202, perform user gaze direction estimation on the front view image to obtain the initial gaze direction of the target user in the front view image.

[0060] Here, estimating the user's gaze direction in the front view image to obtain the initial gaze direction of the target user in the front view image refers to identifying the direction of the target user's eye gaze in the front view image through image recognition. The target user in the front view image can be understood as the user for whom gaze point tracking is required, such as a driver. The initial gaze direction of the target user can also be called the initial gaze point of the target user, which can be represented as spatial coordinates or a position vector. The initial gaze direction of the target user can be represented as dc = (x... dc y dc , z dc ), x dcy dc , z dc It can represent either the three components of the position vector of the initial gaze direction, or the coordinates of the target user's initial gaze point on the three coordinate axes.

[0061] In this method, the user's gaze direction can be estimated using any image recognition method capable of estimating the user's gaze direction, thereby obtaining the initial gaze direction of the target user in the front view image.

[0062] In some cases, user gaze direction estimation can be achieved through image recognition models to obtain the initial gaze direction of the target user in the front view image. The front view image can be input into a gaze direction estimation model to estimate the user's gaze direction and obtain the initial gaze direction of the target user in the front view image.

[0063] The gaze direction estimation model can be any model that takes a user's facial image as input and outputs the user's gaze direction.

[0064] In one specific implementation, the gaze direction estimation model can be trained based on a pre-defined network model. The pre-defined network model can be as follows: Figure 3 As shown in M1, it includes a feature extraction network T, a first output branch L1 and a second output branch L2, and the first output branch L1 and the second output branch L2 are respectively connected to the feature extraction network T.

[0065] The feature extraction network T is used to extract features from sample facial images to obtain image features. The sample facial images refer to user facial images used as training samples. The feature extraction network can include a convolutional neural network (CNN), which includes multiple convolutional layers and pooling layers. The first output branch L1 is used to output the user's gaze direction in the sample facial image based on the image features extracted by the feature extraction network T. The first output branch L1 can be a fully connected layer connected to the feature extraction network T. The fully connected layer in the first output branch L1 will output different gaze directions and different confidence levels for the gaze directions. The confidence level of the gaze direction is used to represent the credibility of the gaze direction. The higher the confidence level, the more credible the gaze direction. The gaze direction corresponding to the maximum confidence level output by the fully connected layer in the first output branch L1 is the user's gaze direction output by the first output branch L1. The second output branch L2 is used to output facial landmarks of the user in the facial image of the image feature sample extracted by the feature extraction network T. Facial landmarks can also be called facial feature points. Facial landmarks can include feature points reflecting the user's eyebrows, eyes, nose and mouth. Facial landmarks can also include feature points reflecting the user's contour. The second output branch L2 can include multiple fully connected layers.

[0066] During training, multiple sample facial images can be collected as a sample dataset. These sample facial images are then input into a pre-defined network model M1. The first output branch L1 of the pre-defined network model M1 outputs the user gaze direction prediction result, and the second output branch L2 outputs the facial landmark prediction result. The difference between the user gaze direction prediction result and the actual gaze direction corresponding to the sample facial image is used as the gaze direction prediction loss of the pre-defined network model M1, and the difference between the facial landmark prediction result and the actual facial landmark corresponding to the sample facial image is used as the landmark prediction loss of the pre-defined network model M1. The model parameters in the pre-defined network model M1 are adjusted based on these two losses until the pre-defined network model converges. Adjusting the model parameters in the pre-defined network model M1 includes adjusting the parameters in the feature extraction network T, the first output branch L1, and the second output branch L2.

[0067] Among them, the gaze direction estimation model can be as follows: Figure 3 As shown in M11, it includes the feature extraction network T and the first output branch L1.

[0068] During the training of the gaze direction estimation model, a second output branch that outputs facial landmarks is introduced as an auxiliary training output branch. Facial landmarks provide information about the eye region and facial expressions. These landmarks are used to adjust the parameters of the feature extraction network during training, enabling the feature extraction network to learn eye region features and facial expression features in the image. In other words, the feature extraction network can learn and extract richer and more accurate image features, thus allowing the gaze direction estimation model to output a more accurate gaze direction based on the image features extracted by the feature extraction network, ensuring the accuracy of the gaze direction estimation model. Introducing a second output branch that outputs facial landmarks as an auxiliary training output branch also accelerates model convergence.

[0069] Optionally, during training, data augmentation techniques can be used to expand the sample dataset, i.e., increase the number of sample facial images in the dataset. This can be achieved by processing the sample facial images in the dataset using one or more image processing methods such as image rotation, random flipping, image blurring, and image scaling, resulting in new sample facial images and thus increasing the number of sample facial images in the dataset. Expanding the sample dataset using data augmentation techniques can improve the generalization ability of the finally trained model.

[0070] S203, perform gaze position estimation on the auxiliary image to obtain the object position of the user's gaze in the auxiliary image.

[0071] Here, identifying the gaze location of the user in the auxiliary image means identifying the object location of the user's gaze in the auxiliary image through image matching. A user gaze object is an object in the target environment that is being gazed upon by the user. There are multiple object locations for user gaze objects, and each object location represents its spatial position. The object location of the user gaze object can indirectly reflect the direction of the target user's gaze in the main view image. The object location of the user gaze object can be represented as P. j =(x j y j , z j ), x j y j , z j It can represent either the three components of the position vector of the j-th object, or the coordinates of the spatial position of the j-th object on the three coordinate axes.

[0072] In one feasible implementation, the position of the object being gazed upon in the auxiliary image can be estimated through the following steps A1-A3 to obtain the position of the object being gazed upon in the auxiliary image:

[0073] A1. Extract feature points from the auxiliary image to obtain an auxiliary feature point set.

[0074] Here, the auxiliary feature point set includes multiple feature points in the auxiliary image; the feature points in the auxiliary image refer to points in the auxiliary image where the gray value changes drastically or points in the auxiliary image where the edge curvature is large. Feature points usually appear at inflection points or where the texture changes drastically, such as corner points, edge points, etc.

[0075] In this process, feature points can be extracted from the auxiliary image using any one or more feature point extraction algorithms to obtain an auxiliary feature point set. Feature point extraction algorithms include, but are not limited to, the scale-invariant feature transform (SIFT) algorithm and the oriented fast and rotated BRIEF (ORB) algorithm.

[0076] A2. In the set of auxiliary feature points, identify the feature points that match the preset gaze environment image to obtain the gaze object feature points.

[0077] Here, the preset gaze environment image is a pre-set image of objects / environments that the user may gaze at. The preset gaze environment image can be set based on specific application scenarios; different application scenarios will result in different preset gaze environment images. For example, in a driving scenario, objects that the user may gaze at include roads, pedestrians, traffic lights, etc., so the preset gaze environment image could be an image of roads, pedestrians, traffic lights, etc. Similarly, in a meeting scenario, the environment that the user may gaze at is a meeting room, so the preset gaze environment image could be an image of the meeting room. The preset gaze environment image is not limited to these examples.

[0078] In this method, feature points in the auxiliary feature point set can be matched with feature points in the preset gaze environment image based on any feature point matching algorithm, thereby determining the feature points that match the preset gaze environment image and obtaining the gaze object feature points.

[0079] In one specific implementation, for any feature point in the auxiliary feature point set (hereinafter referred to as the target feature point), a feature descriptor can be obtained. The feature descriptor is an attribute parameter used to describe the feature point, usually presented as a multi-dimensional position vector. The feature descriptor is extracted by a feature point extraction algorithm. The Euclidean distance between the feature descriptor of the target feature point and the feature descriptors of feature points in the preset gaze environment image is calculated. If the Euclidean distance between the feature descriptor of the target feature point and the feature descriptors of feature points in the preset gaze environment image is less than a preset distance threshold, then the target feature point is determined to match the preset gaze environment image, and the target feature point is identified as a gaze object feature point. If the Euclidean distance between the target feature point and the feature descriptors of every feature point in the preset gaze environment image is greater than or equal to the preset distance threshold, then the target feature point is determined to not match the preset gaze environment, i.e., the target feature point is not a gaze object feature point. For each feature point in the auxiliary feature point set, it is matched with the preset gaze environment image in the same way, thus identifying the feature points that match the preset gaze environment image, thereby identifying all gaze object feature points in the auxiliary feature point set.

[0080] It is understood that other methods can also be used to match feature points in the auxiliary feature point set with the preset gaze environment image, and this application does not impose any restrictions.

[0081] A3. Determine the spatial coordinates of the feature points of the object being gazed upon, and obtain the object position of the user's gazed upon in the auxiliary image.

[0082] In one specific implementation, the spatial location corresponding to the feature point of the object being viewed can be determined through the following steps A31-A33:

[0083] A31. Obtain the projection conversion parameters between the auxiliary camera and the main camera.

[0084] Here, the projection transformation parameters between the auxiliary camera and the main camera are used to reflect the pose of the auxiliary camera relative to the main camera. The projection transformation parameters between the auxiliary camera and the main camera can be expressed as follows: This represents the projection conversion parameters between the i-th auxiliary camera and the main camera.

[0085] The projection conversion parameters between each auxiliary camera and the main camera are obtained through pre-calibration.

[0086] The process of pre-calibrating and determining the projection conversion parameters between each auxiliary camera and the main camera can be as follows:

[0087] First, use the auxiliary camera and the main camera to photograph the same calibration aid, which can be a QR code, a checkerboard image, etc.

[0088] Then, feature point detection is performed on the images captured by the auxiliary camera and the main camera. Based on the detected feature points, the poses of the auxiliary camera and the main camera relative to the calibration aid are determined. The pose of the auxiliary camera relative to the calibration aid can be expressed as follows: Let represent the projection transformation parameters of the i-th auxiliary camera relative to the calibration aid, and let represent the pose of the main camera relative to the calibration aid as follows: Indicates the projection conversion parameters of the main view camera relative to the calibration aid;

[0089] Finally, based on the poses of the auxiliary camera and the main camera relative to the calibration aid, the projection transformation parameters between the auxiliary camera and the main camera are determined. The projection transformation parameters between the auxiliary camera and the main camera are as follows:

[0090]

[0091] A32. Determine the pixel position of the feature point of the object being gazed upon in the auxiliary image to obtain the first pixel position.

[0092] Here, the position of the first pixel can be represented as p. k =(u k v k ), (u k v k ) represents the pixel position of the k-th gaze feature point in the auxiliary image.

[0093] A33. The spatial coordinates that minimize the target difference are determined as the spatial coordinates corresponding to the feature points of the object being gazed upon.

[0094] Here, the target difference is the difference between the second pixel position and the first pixel position. The second pixel position is the pixel position obtained by transforming the spatial coordinates using projection transformation parameters. The target difference can be expressed as: This represents the pixel position obtained by transforming the spatial coordinates using projection transformation parameters. and p k Given all quantities, find the P that minimizes Cd. k , which is the spatial coordinate of the feature point of the object being viewed.

[0095] Among them, the spatial coordinates that minimize the target difference can be obtained by using the least squares method, thus obtaining the spatial coordinates corresponding to the feature points of the object being gazed upon.

[0096] In steps A31-A33 above, the spatial coordinates that minimize the target difference are determined as the spatial coordinates corresponding to the feature points of the object being gazed upon, which can improve the accuracy of the spatial coordinates corresponding to the feature points of the object being gazed upon.

[0097] In another specific implementation, after determining the pixel position of the object's feature point in the auxiliary image and obtaining the first pixel position, the spatial coordinates corresponding to the first pixel position are determined based on the mapping relationship between the pixel coordinate system and the spatial coordinate system. These spatial coordinates are then used as the spatial coordinates corresponding to the object's feature point. The mapping relationship between the pixel coordinate system and the spatial coordinate system is obtained based on the projection transformation parameters between the auxiliary camera and the main camera.

[0098] In steps A1-A3 above, the feature points obtained by extracting feature points from the auxiliary image are matched with the preset gaze environment image to determine the gaze object feature points, which helps to accurately determine the gaze object feature points, thereby accurately determining the object position.

[0099] In another feasible implementation, the position of the object being gazed upon in the auxiliary image can be estimated using an object position estimation model to obtain the position of the object being gazed upon in the auxiliary image; alternatively, after extracting feature points from the auxiliary image to obtain an auxiliary feature point set, the feature points in the auxiliary feature point set can be clustered, and the cluster centers can be determined as the positions of the objects being gazed upon in the auxiliary image. This application does not limit the specific implementation method for estimating the position of the object being gazed upon in the auxiliary image.

[0100] It should be noted that when there are multiple auxiliary cameras, the same method can be used to estimate the position of the object being gazed at in the auxiliary image corresponding to each auxiliary camera, so as to obtain the position of the object being gazed at by the user in the auxiliary image corresponding to each auxiliary camera.

[0101] S204, based on the object position of the user's gaze in the auxiliary image, correct the initial gaze direction of the target user in the main view image to obtain the target gaze direction of the target user in the main view image.

[0102] In some possible cases, if the number of object positions of the user's gaze obtained through step S203 above is only one, the mean of the initial gaze direction position vector (hereinafter referred to as the initial gaze position vector) of the target user in the main view image and the position vector of the object position of the user's gaze in the auxiliary image (hereinafter referred to as the object position vector) can be calculated. Alternatively, the initial gaze position vector and the object position vector can be weighted and summed to obtain the target gaze direction position vector (hereinafter referred to as the target gaze position vector) of the target user in the main view image. The target gaze position vector can be expressed as: dd = (x dd y dd , z dd ), where x dd =w11*x dc +w12*x p y dd =w11*y dc +w12*y p , z dd =w11*z dc +w12*z p x dd y dd , z dd The three components of the target gaze position vector are x dc y dc , z dc The three components of the initial gaze position vector are x p y p , z p The three components represent the object position vector; w11 and w12 are the weights of the initial gaze position vector and the object position vector, respectively. When calculating the mean of the initial gaze position vector and the object position vector, w11 = w12 = 1 / 2.

[0103] In other possible cases, when there are multiple object positions of the user gaze obtained through the above step S203, the vector angle between the position vector of each object position and the initial gaze position vector can be calculated. Among the multiple object positions, candidate object positions with vector angles smaller than a preset angle are determined, and the average of the position vectors of the candidate object positions is determined as the target position vector.

[0104] The formula for calculating the angle between the object's position vector and the initial gaze position vector is:

[0105]

[0106] x k y k , z kThe three components of the position vector represent the position of the k-th object, and α represents the angle between the position vector of the object and the initial gaze position vector.

[0107] The target gaze position vector can be represented as: dd = (x dd y dd , z dd ),in:

[0108]

[0109] x i y i , z i Let m represent the three components of the position vector of the i-th candidate object, where m is the number of candidate object positions.

[0110] In some other possible cases, where the number of object positions obtained through step S203 is multiple, it can also be achieved through... Figure 4 The process steps shown correct the initial orientation of the target user in the main view image. For specific implementation details, please refer to the subsequent sections. Figure 4 The corresponding description will not be elaborated upon here.

[0111] It is understandable that the specific implementation of correcting the initial gaze direction of the target user in the main view image based on the object position of the user's gaze in the auxiliary image to obtain the target gaze direction of the target user in the main view image can be varied, and this application does not limit it.

[0112] In the above Figure 2In the corresponding technical solution, after acquiring the main view image and the auxiliary image, the user's gaze direction is estimated on the main view image to obtain the initial gaze direction of the target user in the main view image. Then, the position of the object being gazed upon is estimated on the auxiliary image to obtain the position of the object being gazed upon in the auxiliary image. Finally, based on the position of the object being gazed upon, the initial gaze direction is corrected to obtain the target gaze direction of the target user in the main view image, thus realizing the detection and correction of the user's gaze direction. Since the main view image is a user's facial image captured by the main view camera, which is set to face the user, estimating the user's gaze direction on the main view image can achieve a rough estimate of the user's gaze in the image. Since the auxiliary image is captured by the auxiliary camera... The system obtains images of the environment within the user's field of view. An auxiliary camera is positioned facing the user's gaze direction. By estimating the position of the object being gazed upon in the auxiliary image, the position of the object being gazed upon in the auxiliary image is obtained. The user's gaze direction is then corrected using the position of the object being gazed upon. The object being gazed upon is the object in the environment that the user is looking at. The position of the object being gazed upon can indirectly reflect the user's gaze direction. This is equivalent to using multi-view images to help determine the user's gaze direction, thus ensuring the accuracy of the user's gaze direction. Since the image is obtained and the user's gaze direction is estimated by combining the main camera and the auxiliary camera, gaze point tracking can be achieved with ordinary cameras without the need for additional sensors, which can save application costs.

[0113] See Figure 4 , Figure 4 The present application provides a process for correcting the initial gaze direction, such as... Figure 4 As shown, it includes the following steps:

[0114] S2041, determine multiple object position matching values ​​based on the target user's initial gaze direction and multiple object positions in the main view image.

[0115] Here, multiple object position matching values ​​correspond to multiple object positions, with one object position corresponding to one object position matching value. The number of object position matching values ​​is equal to the number of object positions of the user's gaze obtained through the aforementioned step S203. The object position matching value is used to reflect the proximity of the object position to the initial gaze direction; the higher the object position matching value, the closer the object position is to the initial gaze direction; the lower the object position matching value, the farther the object position is from the initial gaze direction.

[0116] In one feasible implementation, multiple object position matching values ​​can be determined through the following steps B1-B2:

[0117] B1. Calculate the projection distance of the target position vector relative to the initial gaze position vector to obtain the target projection distance.

[0118] Here, the target position vector is the position vector of the target object, where the target object position is any one of multiple object positions. The target position vector can be represented as P mentioned above. j =(x j y j , z j The initial gaze position vector is the position vector of the target user's initial gaze direction in the main view image. The initial gaze position vector can be expressed as dc = (x dc y dc , z dc ).

[0119] The projected distance between the target position vector and the initial gaze position vector can be calculated using the vector projection distance calculation formula, which is as follows:

[0120]

[0121] Dt represents the target projection distance.

[0122] B2. Determine the object position matching value corresponding to the target object position based on the target projection distance.

[0123] In one specific implementation, the target projection distance can be determined as the object position matching value corresponding to the target object position.

[0124] In another specific implementation, the object position matching value corresponding to the target object position can also be determined through the following steps B21-B22:

[0125] B21. Obtain the confidence level of the initial gaze direction.

[0126] Here, the confidence level of the initial gaze direction is used to reflect the reliability of the initial gaze direction. The higher the confidence level of the initial gaze direction, the more reliable the initial gaze direction is; the lower the confidence level of the initial gaze direction, the less reliable the initial gaze direction is.

[0127] The confidence level of the initial gaze direction can be obtained through the process of estimating the initial gaze direction in step S202.

[0128] B22. The confidence level of the initial gaze direction and the target projection distance are weighted and summed to obtain the object position matching value corresponding to the target object position.

[0129] The formula for calculating the object position matching value of the target object is: r j= w21*Dt+w22*Dz,r jThe object position matching value is represented by w21 and w22, which represent the weights of the confidence of the target projection distance and the initial gaze direction, respectively. Dz represents the confidence of the initial gaze direction.

[0130] In steps B21-B22 above, the object position matching value is obtained by weighted summing the projection distance of the object's position vector relative to the position vector of the initial gaze direction and the confidence level reflecting the credibility of the initial gaze direction. This makes the object position matching value reasonable and effective.

[0131] In steps B1-B2 above, the object position matching value is determined based on the projection distance of the object's position vector relative to the initial direction's position vector. This ensures that the object position matching value is reasonable and sufficient to reflect the closeness between the object's position and the initial gaze direction.

[0132] In another feasible implementation, the angle between the target position vector and the initial gaze position vector can be calculated to obtain the target vector angle; based on the target vector angle, the object position matching value corresponding to the target object position can be determined. For example, the target vector angle can be determined as the object position matching value corresponding to the target object position; or, the product of the confidence level of the initial gaze direction and the target vector angle can be determined as the object position matching value corresponding to the target object position. The formula for calculating the vector angle refers to the aforementioned formula for calculating cosα.

[0133] This application does not impose any restrictions on the specific implementation method for determining the object position matching value.

[0134] For each of the multiple object locations, the object location matching value is determined using the same method, thus obtaining the object location matching value for each object location, and consequently, multiple object location matching values.

[0135] S2042, determine the target user's gaze direction based on multiple object positions and multiple object position matching values.

[0136] In one feasible implementation, the target user's gaze direction can be determined through the following steps C1-C3:

[0137] C1. Perform location clustering on multiple object locations to obtain at least one object location set.

[0138] In this process, multiple object locations can be clustered using any location clustering algorithm to obtain at least one set of object locations. Location clustering algorithms include, but are not limited to, K-means clustering and density-based spatial clustering of applications with noise (DBSCAN).

[0139] C2. Based on the object position matching value corresponding to the object position in each object position set, determine the set matching value corresponding to each object position set.

[0140] Here, the set matching value corresponding to the object position set is used to reflect how close the object position in the object position set is to the initial gaze direction. The larger the set matching value, the closer the object position in the object position set is to the initial gaze direction; the smaller the set matching value, the farther the object position in the object position set is from the initial gaze direction.

[0141] In one specific implementation, the set matching value corresponding to each object location set can be determined through the following steps C21-C22:

[0142] C21. Calculate the mean of the object position matching values ​​corresponding to the object positions in the first object position set and / or the sum of the object position matching values ​​corresponding to the object positions in the first object position set, to obtain the mean of the target matching values ​​and / or the sum of the target matching values.

[0143] Here, the first object position set is any one of the object position sets in at least one object position set.

[0144] C22. Determine the set matching value corresponding to the first object position set based on the mean of the target matching values ​​and / or the sum of the target matching values.

[0145] Specifically, the average or sum of the target matching values ​​can be used to determine the set matching value corresponding to the first object position set; alternatively, the ratio of the average target matching value to the maximum object position matching value corresponding to the object position in the first object position set can be calculated to obtain the target ratio, and the product of the target ratio and the sum of the target matching values ​​can be used to determine the set matching value corresponding to the first object position set. The maximum object position matching value corresponding to the object position in the first object position set refers to the maximum value among the object position matching values ​​corresponding to the object position in the first object position set. This application does not limit the method of determining the set matching value corresponding to the first object position set.

[0146] In steps C21-C22 above, the set matching value corresponding to the object position set is determined based on the mean and / or sum of the object position matching values ​​corresponding to the object positions in the object position set, which can reasonably determine the set matching value of the object position set.

[0147] By determining the set matching value in the same way for each set of object locations, we can obtain the set matching value corresponding to each set of object locations, thus obtaining at least one set matching value.

[0148] C3. Determine the target user's gaze direction based on the object positions in the object position set corresponding to the maximum set matching value.

[0149] Here, the maximum set matching value refers to the maximum value among the set matching values ​​corresponding to each object location set.

[0150] In one specific implementation, the target user's gaze direction can be determined through the following steps C31-C32:

[0151] C31. Determine the center of position of the objects in the second object position set.

[0152] Here, the second object position set is the object position set corresponding to the maximum set matching value, and the spatial position coordinates of the position center of the object position in the second object position set are the average of the spatial position coordinates of the object position in the second object position set.

[0153] The spatial coordinates of the location center can be represented as (x od y od , z od ),in:

[0154]

[0155] (x u y u , z u ) represents the spatial coordinates of the u-th object position in the second object position set, and n represents the number of object positions in the second object position set.

[0156] C32. Determine the position vector of the position center as the position vector of the target user's gaze direction.

[0157] The position vector of the location center can be represented as od = (x od y od , z od The position vector of the target user's gaze direction can be represented as dd = (x dd y dd , z dd ), x dd=x od y dd =y od , z dd =z od .

[0158] In steps C31-C32 above, the position vector of the center of the object position in the set of object positions with the largest matching value is determined as the position vector of the target user's gaze direction, which can accurately determine the user's gaze direction.

[0159] In some possible cases, before performing location clustering on multiple object locations to obtain at least one set of object locations, noisy object locations can be filtered out from the multiple object locations. Noisy object locations are those whose object location matching values ​​are less than a preset threshold.

[0160] Here, the preset threshold is a threshold used to measure the degree of matching between the object position and the initial gaze direction; the preset threshold will be different depending on the method of calculating the object position matching value, and this application does not impose any restrictions.

[0161] In this process, after determining multiple object position matching values, a target object position matching value less than a preset threshold can be identified from the multiple object position matching values. The object position corresponding to the target object position matching value is then removed from the multiple object positions, thereby filtering out noisy object positions from the multiple object positions.

[0162] By filtering out the positions of noisy objects, interference from noise on the user's gaze direction can be avoided, thus improving the accuracy of gaze tracking.

[0163] In another feasible implementation, the maximum object position matching value can be determined among multiple object position matching values, and the position vector of the object position corresponding to the maximum object position matching value can be determined as the position vector of the target user's target gaze direction.

[0164] In another feasible implementation, multiple object position matching values ​​can be used as weight values ​​for multiple object positions, and the weighted average of the spatial position coordinates of multiple object positions can be calculated. The position vector corresponding to the weighted average is then determined as the position vector of the target user's gaze direction.

[0165] This application does not limit the specific implementation of determining the target user's gaze direction based on multiple object positions and multiple object position matching values.

[0166] The method of this application has been described above; the apparatus of this application will be described below.

[0167] See Figure 5 , Figure 5This is a schematic diagram of a gaze direction estimation device provided in an embodiment of this application, applied to a target device. The target device is connected to a main viewing camera and at least one auxiliary camera. The main viewing camera is oriented towards the user, and the auxiliary cameras are oriented towards the user's gaze direction. Figure 5 As shown, the gaze direction estimation device 30 includes:

[0168] The image acquisition module 301 is used to acquire a front view image and an auxiliary image. The front view image is a target face image captured by the front view camera. The target face image is an image of the face of the target user. The auxiliary image is an image of the target environment captured by the auxiliary camera. The target environment is the environment within the field of view of the target user.

[0169] The gaze direction estimation module 302 is used to estimate the user gaze direction of the main view image to obtain the initial gaze direction of the target user in the main view image;

[0170] The gaze object position estimation module 303 is used to estimate the gaze object position in the auxiliary image to obtain the object position of the user's gaze object in the auxiliary image;

[0171] The gaze direction correction module 304 is used to correct the initial gaze direction according to the position of the object to obtain the target gaze direction of the target user.

[0172] In one possible design, the number of object positions is multiple; the gaze direction correction module 304 is specifically used to: determine multiple object position matching values ​​based on the initial gaze direction and multiple object positions, wherein the multiple object position matching values ​​are object position matching values ​​corresponding to the multiple object positions, and the object position matching values ​​are used to reflect the degree of proximity between the object positions and the initial gaze direction; and determine the target gaze direction of the target user based on the multiple object positions and the multiple object position matching values.

[0173] In one possible design, the gaze direction correction module 304 is specifically used to: calculate the projection distance of the target position vector relative to the initial gaze position vector to obtain the target projection distance, wherein the target position vector is the position vector of the target object position, the target object position is any one of the plurality of object positions, and the initial gaze position vector is the position vector of the initial gaze direction; and determine the object position matching value corresponding to the target object position based on the target projection distance.

[0174] In one possible design, the gaze direction correction module 304 is specifically used to: cluster the positions of the plurality of objects to obtain at least one set of object positions; determine the set matching value corresponding to each set of object positions based on the object position matching value corresponding to the object position in each set of object positions; and determine the target gaze direction of the target user based on the object position in the set of object positions corresponding to the maximum set matching value.

[0175] In one possible design, the gaze direction correction module 304 is specifically used to: calculate the mean of the object position matching values ​​corresponding to the object positions in the first object position set and / or the sum of the object position matching values ​​corresponding to the object positions in the first object position set, to obtain the mean of the target matching values ​​and / or the sum of the target matching values, wherein the first object position set is any one of the at least one object position sets; and determine the set matching value corresponding to the first object position set based on the mean of the target matching values ​​and / or the sum of the target matching values.

[0176] In one possible design, the gaze direction correction module 304 is specifically used to: determine the position center of the object position in the second object position set, wherein the second object position set is the object position set corresponding to the maximum set matching value, and the spatial position coordinates of the position center are the average of the spatial position coordinates of the object positions in the second object position set; and determine the position vector of the position center as the position vector of the target gaze direction of the target user.

[0177] In one possible design, the gaze direction correction module 304 is further configured to: filter out noisy object positions from the plurality of object positions, wherein the noisy object positions are object positions whose object position matching values ​​are less than a preset threshold.

[0178] In one possible design, the image acquisition module 301 is specifically used to: input the front view image into the gaze direction estimation model to estimate the user's gaze direction, and obtain the initial gaze direction of the target user in the front view image.

[0179] In one possible design, the gaze direction estimation model is trained by a preset network model, which includes a feature extraction network, a first output branch, and a second output branch. The first and second output branches are respectively connected to the feature extraction network. The feature extraction network is used to extract features from a sample facial image to obtain image features. The sample facial image is a facial image used as a training sample. The first output branch is used to output the user's gaze direction in the sample facial image based on the image features. The second output branch is used to output facial landmarks of the user in the sample facial image based on the image features. The facial landmarks are used to adjust the parameters of the feature extraction network during the training process of the preset network model. The gaze direction estimation model includes the trained feature extraction network and the first output branch.

[0180] In one possible design, the aforementioned object position estimation module 303 is specifically used to: extract feature points from the auxiliary image to obtain an auxiliary feature point set, the auxiliary feature point set including multiple feature points in the auxiliary image; determine feature points in the auxiliary feature point set that match a preset gaze environment image to obtain object feature points; and determine the spatial coordinates corresponding to the object feature points to obtain the object position.

[0181] In one possible design, the gaze object position estimation module 303 is specifically used to: obtain projection transformation parameters between the auxiliary camera and the main camera, the projection transformation parameters reflecting the pose of the auxiliary camera relative to the main camera; determine the pixel position of the gaze object feature point in the auxiliary image to obtain a first pixel position; determine the spatial position coordinates that minimize the target difference as the spatial position coordinates corresponding to the gaze object feature point, the target difference being the difference between the second pixel position and the first pixel position, the second pixel position being the pixel position obtained by transforming the spatial position coordinates through the projection transformation parameters.

[0182] It should be noted that, Figure 5 For any content not mentioned in the corresponding embodiments, please refer to the description of the foregoing method embodiments, which will not be repeated here.

[0183] The aforementioned device, after acquiring a front view image and an auxiliary image, estimates the user's gaze direction in the front view image to obtain the initial gaze direction of the target user in the front view image. Then, it estimates the position of the object being gazed upon in the auxiliary image to obtain the position of the object being gazed upon in the auxiliary image. Finally, based on the position of the object being gazed upon, it corrects the initial gaze direction to obtain the target gaze direction of the target user in the front view image, thus achieving the detection and correction of the user's gaze direction. Since the front view image is a facial image of the user captured by the main camera, which is oriented towards the user, estimating the user's gaze direction in the front view image allows for a rough estimation of the user's gaze in the image. Because the auxiliary image... The auxiliary image is an image of the environment within the user's field of view captured by the auxiliary camera, which is set to face the user's gaze direction. By estimating the position of the object being gazed upon in the auxiliary image, the position of the object being gazed upon in the auxiliary image is obtained, and the user's gaze direction is corrected using the position of the object being gazed upon. The object being gazed upon is the object in the environment that the user is gazing upon. This is equivalent to using multi-view images to help determine the user's gaze direction, thereby ensuring the accuracy of the user's gaze direction. Since the image is obtained and the user's gaze direction is estimated by combining the main view camera and the auxiliary camera, gaze tracking can be achieved with ordinary cameras without the need for additional sensors, which can save application costs.

[0184] See Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device 40 provided in an embodiment of this application. The computer device 40 includes a processor 401 and a memory 402. The memory 402 is connected to the processor 401, for example, via a bus. The computer device is connected to a main camera and an auxiliary camera.

[0185] Processor 401 is configured to support the computer device 40 in performing the corresponding functions in the methods described in the above method embodiments. Processor 401 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0186] Memory 402 is used to store program code, etc. Memory 402 may include volatile memory (VM), such as random access memory (RAM); memory 402 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 402 may also include combinations of the above types of memory.

[0187] The memory 402 is used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the gaze direction estimation method in the embodiments of this application. The core processor and the graphics processor cooperate to execute various functional applications and data processing of the gaze direction estimation method by running the non-volatile software programs, instructions, and modules stored in the memory, thereby realizing the function of the gaze direction estimation method provided in the above method embodiments.

[0188] Memory 402 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function. The data storage area may store data created based on the use of the gaze direction estimation device. In some embodiments, the memory may include memory remotely located relative to the processor, which can be connected to the gaze direction estimation device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0189] The one or more modules are stored in the memory. When executed by the one or more processors, they perform the gaze direction estimation method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.

[0190] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the method described in the foregoing embodiments.

[0191] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0192] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A method for estimating gaze direction, characterized in that, The method is applied to a target device connected to a main-view camera and at least one auxiliary camera, the main-view camera being oriented towards the user and the auxiliary cameras being oriented towards the user's gaze direction; the method includes: Acquire a front view image and an auxiliary image. The front view image is a target face image captured by the front view camera, which is an image of the target user's face. The auxiliary image is an image of the target environment captured by the auxiliary camera, which is the environment within the target user's field of vision. The user gaze direction is estimated by performing a user gaze direction estimation on the front view image to obtain the initial gaze direction of the target user in the front view image; The position of the gaze object is estimated by performing gaze object position estimation on the auxiliary image to obtain the position of the user gaze object in the auxiliary image, wherein the user gaze object is an object in the target environment that is gazed at by the target user; Based on the object's position, the initial gaze direction is corrected to obtain the target user's target gaze direction.

2. The method according to claim 1, characterized in that, The number of the object's locations is multiple; The step of correcting the initial gaze direction based on the object's position to obtain the target user's target gaze direction includes: Based on the initial gaze direction and multiple object positions, multiple object position matching values ​​are determined. The multiple object position matching values ​​are the object position matching values ​​corresponding to the multiple object positions, and the object position matching values ​​are used to reflect the degree of proximity between the object positions and the initial gaze direction. The target user's gaze direction is determined based on the multiple object positions and the multiple object position matching values.

3. The method according to claim 2, characterized in that, The step of determining multiple object position matching values ​​based on the initial gaze direction and multiple object positions includes: Calculate the projection distance of the target position vector relative to the initial gaze position vector to obtain the target projection distance. The target position vector is the position vector of the target object, and the target object position is any one of the multiple object positions. The initial gaze position vector is the position vector of the initial gaze direction. Based on the target projection distance, determine the object position matching value corresponding to the target object position.

4. The method according to claim 2, characterized in that, Determining the target user's gaze direction based on the multiple object positions and the multiple object position matching values ​​includes: The locations of the multiple objects are clustered to obtain at least one set of object locations; Based on the object position matching value corresponding to the object position in each object position set, determine the set matching value corresponding to each object position set; The target user's gaze direction is determined based on the object positions in the object position set corresponding to the maximum set matching value.

5. The method according to claim 4, characterized in that, The step of determining the set matching value corresponding to each object location set based on the object location matching value corresponding to the object location in each object location set includes: Calculate the mean of the object position matching values ​​corresponding to the object positions in the first object position set and / or the sum of the object position matching values ​​corresponding to the object positions in the first object position set to obtain the mean of the target matching values ​​and / or the sum of the target matching values, wherein the first object position set is any one of the at least one object position sets; The set matching value corresponding to the first object location set is determined based on the mean of the target matching values ​​and / or the sum of the target matching values.

6. The method according to claim 4, characterized in that, Determining the target user's gaze direction based on the object positions in the object position set corresponding to the maximum set matching value includes: Determine the location center of the object locations in the second object location set, where the second object location set is the object location set corresponding to the maximum set matching value, and the spatial location coordinates of the location center are the average of the spatial location coordinates of the object locations in the second object location set. The position vector of the location center is determined as the position vector of the target user's target gaze direction.

7. The method according to claim 4, characterized in that, Before performing location clustering on the multiple object locations to obtain at least one object location set, the method further includes: Noisy object locations are filtered out from the plurality of object locations, wherein the noisy object locations are those whose object location matching values ​​are less than a preset threshold.

8. The method according to any one of claims 1-7, characterized in that, The step of estimating the user's gaze direction in the main view image to obtain the initial gaze direction of the target user in the main view image includes: The front view image is input into the gaze direction estimation model to estimate the user's gaze direction, thereby obtaining the initial gaze direction of the target user in the front view image.

9. The method according to claim 8, characterized in that, The gaze direction estimation model is trained by a preset network model, which includes a feature extraction network, a first output branch, and a second output branch. The first and second output branches are respectively connected to the feature extraction network. The feature extraction network is used to extract features from a sample facial image to obtain image features. The sample facial image is a facial image used as a training sample. The first output branch is used to output the user's gaze direction in the sample facial image based on the image features. The second output branch is used to output facial landmarks of the user in the sample facial image based on the image features. The facial landmarks are used to adjust the parameters of the feature extraction network during the training process of the preset network model. The gaze direction estimation model includes the trained feature extraction network and the first output branch.

10. The method according to any one of claims 1-7, characterized in that, The step of estimating the gaze location of the auxiliary image to obtain the object location of the user's gaze in the auxiliary image includes: Feature points are extracted from the auxiliary image to obtain an auxiliary feature point set, which includes multiple feature points in the auxiliary image. In the set of auxiliary feature points, feature points that match the preset gaze environment image are determined to obtain the gaze object feature points; The spatial coordinates corresponding to the feature points of the object being gazed upon are determined to obtain the position of the object.

11. The method according to claim 10, characterized in that, Determining the spatial coordinates corresponding to the feature points of the object being gazed upon, and obtaining the position of the object, includes: Obtain projection transformation parameters between the auxiliary camera and the main camera, wherein the projection transformation parameters are used to reflect the pose of the auxiliary camera relative to the main camera; Determine the pixel position of the gaze object feature point in the auxiliary image to obtain the first pixel position; The spatial coordinates that minimize the target difference are determined as the spatial coordinates corresponding to the gaze object feature point. The target difference is the difference between the second pixel position and the first pixel position. The second pixel position is the pixel position obtained by transforming the spatial coordinates through the projection transformation parameters.

12. A gaze direction estimation device, characterized in that, An apparatus applicable to a target device connected to a main-view camera and at least one auxiliary camera, the main-view camera being oriented towards a user and the auxiliary camera being oriented towards the user's gaze direction; the apparatus includes: The image acquisition module is used to acquire a front view image and an auxiliary image. The front view image is a target face image captured by the front view camera, which is an image of the target user's face. The auxiliary image is an image of the target environment captured by the auxiliary camera, which is the environment within the target user's field of vision. A gaze direction estimation module is used to estimate the user's gaze direction in the main view image to obtain the initial gaze direction of the target user in the main view image; The gaze object position estimation module is used to estimate the gaze object position of the auxiliary image to obtain the object position of the user gaze object in the auxiliary image, wherein the user gaze object is an object in the target environment that is gazed at by the target user; The gaze direction correction module is used to correct the initial gaze direction based on the object's position to obtain the target gaze direction of the target user.

13. A computer device, characterized in that, The device includes a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, the processor causing the computer device to perform the method as described in any one of claims 1-11 when executing the one or more computer programs.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-11.