Method and related device for determining gaze position

Through image feature extraction and neural network model processing, the problems of poor universality and low efficiency of gaze position determination are solved, and fast and accurate gaze position detection on ordinary devices are achieved.

CN114299598BActive Publication Date: 2025-05-27BEIJING DAJIA INTERNET INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111533438.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-05-27
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

In the prior art, the method of determining the gaze position is poor in versatility and cumbersome and low efficiency, especially in ordinary equipment, it is difficult to accurately determine the gaze position.

Method used

By obtaining the image of the target object, the facial area, left eye area and right eye area are parsed out, and facial features, comprehensive features, left eye feature expression and right eye feature expression are obtained, and gaze position information is obtained based on these features.

Benefits of technology

It realizes the quick and accurate determination of the gaze position on ordinary devices, avoids tedious correction steps, and is suitable for a variety of lighting environments and head postures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299598B_ABST
    Figure CN114299598B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and related device for determining a gaze position, which are used to solve the problems of poor generality, cumbersome process and low efficiency in the related art for determining the gaze position. Based on the image captured by a camera, the present disclosure decomposes the left eye region, the right eye region and the face region therefrom, then analyzes these three regions to obtain a comprehensive feature, analyzes the left and right eye region images based on the comprehensive feature to obtain a left eye feature expression and a right eye feature expression, and finally combines the comprehensive feature and the face feature to obtain the gaze position. Only important features, including the face feature, the comprehensive feature, the left eye feature expression and the right eye feature expression, need to be extracted throughout the process, and then based on these features, the gaze position of the human eye can be classified. The user does not need to fixate on a fixed point to collect calibration data, and the accuracy of determining the gaze position can be ensured through feature descriptions at multiple levels.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Eye gaze usually contains a lot of information, which can reflect interest points, concentration level, and even psychological state. Automatic and real-time estimation of eye gaze is of great value to various researchers and daily production and life. However, in the past, in order to achieve more accurate eye gaze estimation, it was usually necessary to purchase professional equipment.

[0003] Many different technical approaches have been proposed in the field of gaze estimation recently, which can be summarized into three categories: the technical approach of estimating the three-dimensional direction of gaze based on eye model reconstruction; the technical approach of estimating the screen gaze point based on regression of two-dimensional eye features; and the technical approach based on facial features. Among them:

[0004] Three-dimensional eye model reconstruction is to build a three-dimensional geometric model of the eye and use it to estimate the line of sight. The eye model of each object is different, so this type of technical route requires capturing a lot of object information to reconstruct the object's eye model, such as measuring the iris radius, etc., and this method also requires the use of professional equipment to collect a lot of current object information.

[0005] The equipment requirements of the two-dimensional technology route are basically the same as those of the three-dimensional technology route. The two-dimensional technology route directly uses the measured information such as the pupil center and eyelid to regress the position of the gaze point on the screen, so professional equipment is also required.

[0006] The most important difference between the face-based technology route and the first two categories is that the requirements for hardware equipment are very low. It collects facial information through an ordinary webcam and directly regresses the gaze direction or gaze point based on the collected image.

[0007] Although the face-based technology route has lower hardware requirements, its process is also relatively more complicated. For example, this method usually requires each subject to gaze at some fixed points on the screen before use to collect the correction data of the object. This process has poor versatility in determining the gaze position, and the process is cumbersome and inefficient. Therefore, how to determine the gaze position on ordinary devices remains to be studied. Summary of the invention

[0008] The embodiments of the present disclosure provide a method for determining a gaze position and a related device, which are used to solve the problems in the related art that the method for determining a gaze position has poor versatility, a complicated process and low efficiency.

[0009] In a first aspect, the present disclosure provides a method for determining a gaze position, the method comprising:

[0010] Acquire an image of the target object;

[0011] Parsing the face area, left eye area and right eye area of ​​the target object from the image;

[0012] Feature extraction is performed on the facial region to obtain facial features, and feature extraction is performed on the facial region, the left eye region, and the right eye region to obtain comprehensive features;

[0013] Feature extraction is performed on the left eye region, the facial features, and the comprehensive features to obtain a left eye feature representation; and feature extraction is performed on the right eye region, the facial features, and the comprehensive features to obtain a right eye feature representation;

[0014] Based on the left eye feature representation, the right eye feature representation, the facial features, and the comprehensive features, gaze position information of the target object is obtained, where the gaze position information includes gaze point coordinates and / or the region where the gaze point is located.

[0015] Optionally, the feature extraction of the left eye region, the facial features, and the comprehensive features to obtain a left eye feature representation includes:

[0016] Feature extraction is performed on the left eye region, the facial features, and the comprehensive features to obtain a first left eye feature map;

[0017] Encoding operation is performed on the first left eye feature map to obtain the left eye encoded feature of the first left eye feature map; and context information of each feature point of the first left eye feature map is extracted to obtain left eye context features;

[0018] Based on the left eye encoded feature and the left eye context features, a second left eye feature map is obtained;

[0019] Based on the first left eye feature map and the second left eye feature map, the left eye feature representation is extracted.

[0020] Optionally, the obtaining of the second left eye feature map based on the left eye encoded feature and the left eye context features includes:

[0021] The first left eye feature map and the left eye context features are concatenated to obtain left eye concatenated features;

[0022] Convolution operations are sequentially performed on the left eye concatenated features to obtain left eye convolution features;

[0023] Based on the left eye convolution features and the left eye encoded feature, left eye fusion features are obtained;

[0024] The Fusion module is used to process the left eye fusion features and the left eye context features to obtain the second left eye feature map.

[0025] Optionally, the feature extraction of the right eye region, the facial features, and the comprehensive features to obtain the right eye feature representation includes:

[0026] Performing feature extraction on the right eye region, the facial features, and the comprehensive features to obtain a first right eye feature map;

[0027] Performing an encoding operation on the first right eye feature map to obtain the right eye encoded features of the first right eye feature map; and extracting the context information of each feature point of the first right eye feature map to obtain right eye context features;

[0028] Based on the right eye encoded features and the right eye context features, obtaining a second right eye feature map;

[0029] Based on the first right eye feature map and the second right eye feature map, extracting the right eye feature representation.

[0030] Optionally, the obtaining of the second right eye feature map based on the right eye encoded features and the right eye context features includes:

[0031] Concatenating the first right eye feature map and the right eye context features to obtain right eye concatenated features;

[0032] Performing a convolution operation on the right eye concatenated features in sequence to obtain right eye convolution features;

[0033] According to the right eye convolution features and the right eye encoded features, obtaining right eye fusion features;

[0034] Using a fusion module to process the right eye fusion features and the right eye context features to obtain the second right eye feature map.

[0035] Optionally, the neural network layer used for the encoding operation is a convolutional layer with a convolution kernel of 1*1.

[0036] Optionally, the neural network layer for extracting context information is a convolutional layer with a convolution kernel of n*n, where n is greater than 1 and less than a specified value, and n is a positive integer.

[0037] Optionally, the obtaining of the gaze position information of the target object based on the left eye feature representation, the right eye feature representation, the facial features, and the comprehensive features includes:

[0038] Performing a concatenation process on the left eye feature representation, the right eye feature representation, the facial features, and the comprehensive features to obtain global concatenated features;

[0039] Performing a normalization process on the global concatenated features to obtain a normalized feature map;

[0040] Process each channel feature of the normalized feature map using a multi-layer perceptron network module to obtain the feature to be recognized;

[0041] Perform channel mixing on the global concatenated feature to obtain a channel mixed feature;

[0042] Process the feature to be recognized and the channel mixed feature using a first fully connected layer to obtain the gaze position information of the target object.

[0043] Optionally, if the gaze position information includes the region where the gaze point of the target object is located, determining the region includes:

[0044] Perform a classification operation on the feature to be recognized to obtain a region classification result, and the region classification result is used to indicate the region where the gaze point of the target object is located.

[0045] Optionally, the left-eye feature extraction module for extracting the left-eye feature expression and the right-eye feature extraction module for extracting the right-eye feature expression have the same structure, and at least one of the following neural network layers at the same position in the left-eye feature extraction module and the right-eye feature extraction module share network parameters:

[0046] Convolutional layer, pooling layer, squeeze-and-excitation layer.

[0047] Optionally, the left-eye additional module for extracting the second left-eye feature map and the right-eye additional module for extracting the right-eye feature map adopt the same network structure, and the convolutional layers at the same position in the network structure share network parameters.

[0048] In a second aspect, a device for determining a gaze position, the device includes:

[0049] An image acquisition module configured to acquire an image of a target object;

[0050] A region recognition module configured to parse out the facial region, left-eye region, and right-eye region of the target object from the image;

[0051] A comprehensive feature extraction module configured to extract facial features from the facial region and extract comprehensive features from the facial region, the left-eye region, and the right-eye region;

[0052] A binocular feature extraction module configured to extract a left-eye feature expression by extracting features from the left-eye region, the facial features, and the comprehensive features; and extract a right-eye feature expression by extracting features from the right-eye region, the facial features, and the comprehensive features;

[0053] A fixation position determination module, configured to obtain fixation position information of the target object based on the left-eye feature representation, the right-eye feature representation, the facial feature, and the comprehensive feature, where the fixation position information includes fixation point coordinates and / or the region where the fixation point is located.

[0054] Optionally, when performing feature extraction on the left-eye region, the facial feature, and the comprehensive feature to obtain a left-eye feature representation, the binocular feature extraction module is specifically configured to perform:

[0055] Perform feature extraction on the left-eye region, the facial feature, and the comprehensive feature to obtain a first left-eye feature map;

[0056] Perform an encoding operation on the first left-eye feature map to obtain the left-eye encoded feature of the first left-eye feature map; and extract the context information of each feature point of the first left-eye feature map to obtain the left-eye context feature;

[0057] Based on the left-eye encoded feature and the left-eye context feature, obtain a second left-eye feature map;

[0058] Based on the first left-eye feature map and the second left-eye feature map, extract the left-eye feature representation.

[0059] Optionally, when performing the operation of obtaining the second left-eye feature map based on the left-eye encoded feature and the left-eye context feature, the binocular feature extraction module is specifically configured to perform:

[0060] Concatenate the first left-eye feature map and the left-eye context feature to obtain a left-eye concatenated feature;

[0061] Perform a convolution operation on the left-eye concatenated feature in sequence to obtain a left-eye convolution feature;

[0062] According to the left-eye convolution feature and the left-eye encoded feature, obtain a left-eye fusion feature;

[0063] Use a fusion module to process the left-eye fusion feature and the left-eye context feature to obtain the second left-eye feature map.

[0064] Optionally, when performing feature extraction on the right-eye region, the facial feature, and the comprehensive feature to obtain a right-eye feature representation, the binocular feature extraction module is specifically configured to perform:

[0065] Perform feature extraction on the right-eye region, the facial feature, and the comprehensive feature to obtain a first right-eye feature map;

[0066] Perform an encoding operation on the first right-eye feature map to obtain the right-eye encoded feature of the first right-eye feature map; and extract the context information of each feature point of the first right-eye feature map to obtain the right-eye context feature;

[0067] Based on the right-eye encoded feature and the right-eye context feature, obtain a second right-eye feature map;

[0068] Based on the first right-eye feature map and the second right-eye feature map, extract the right-eye feature expression.

[0069] Optionally, when performing the operation of obtaining the second right-eye feature map based on the right-eye encoded feature and the right-eye context feature, the binocular feature extraction module is specifically configured to perform:

[0070] Concatenate the first left-and-right eye feature map and the right-eye context feature to obtain a right-eye concatenated feature;

[0071] Perform a convolution operation on the right-eye concatenated feature in sequence to obtain a right-eye convolution feature;

[0072] According to the right-eye convolution feature and the right-eye encoded feature, obtain a right-eye fusion feature;

[0073] Use a fusion module to process the right-eye fusion feature and the right-eye context feature to obtain the second right-eye feature map.

[0074] Optionally, the neural network layer used for the encoding operation is a convolutional layer with a convolution kernel of 1*1.

[0075] Optionally, the neural network layer for extracting context information is a convolutional layer with a convolution kernel of n*n, where n is greater than 1 and less than a specified value, and n is a positive integer.

[0076] Optionally, when performing the operation of obtaining the gaze position information of the target object based on the left-eye feature expression, the right-eye feature expression, the facial feature, and the comprehensive feature, the gaze position determination module is specifically configured to perform:

[0077] Perform a concatenation process on the left-eye feature expression, the right-eye feature expression, the facial feature, and the comprehensive feature to obtain a global concatenated feature;

[0078] Normalize the global concatenated feature to obtain a normalized feature map;

[0079] Process the channel features of the normalized feature map using a multi-layer perceptron MLP network module to obtain a feature to be recognized;

[0080] Perform channel mixing on the global concatenated feature to obtain a channel mixed feature;

[0081] The first fully connected layer is used to process the feature to be recognized and the channel mixing feature, so as to obtain the gaze position information of the target object.

[0082] Optionally, if the gaze position information includes the area where the gaze point of the target object is located, determining the area includes:

[0083] A classification module, configured to perform a classification operation on the feature to be recognized to obtain a region classification result, where the region classification result is used to indicate the region where the gaze position of the target object is located.

[0084] Optionally, the left-eye feature extraction module for extracting the left-eye feature expression and the right-eye feature extraction module for extracting the right-eye feature expression have the same structure, and at least one of the following neural network layers at the same position in the left-eye feature extraction module and the right-eye feature extraction module share network parameters:

[0085] Convolutional layer, pooling layer, squeeze-and-excitation layer.

[0086] Optionally, the left-eye additional module for extracting the second left-eye feature map and the right-eye additional module for extracting the right-eye feature map adopt the same network structure, and the convolutional layers at the same position in the network structure share network parameters.

[0087] In a third aspect, the present disclosure also provides an electronic device, including:

[0088] A processor;

[0089] A memory for storing instructions executable by the processor;

[0090] Wherein, the processor is configured to execute the instructions to implement any method provided in the first aspect and the second aspect of the present disclosure.

[0091] In a fourth aspect, an embodiment of the present disclosure also provides a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute any method provided in the first aspect and the second aspect of the present disclosure.

[0092] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements any method provided in the first aspect and the second aspect of the present disclosure.

[0093] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0094] In the method for determining the gaze position provided by the embodiments of the present disclosure, by constructing a neural network model, a context feature model, and a normalization model, the neural network model can extract the features of both eyes and the entire face in the image, and cumbersome correction steps are avoided. Thus, the accuracy of estimating the gaze position is ensured, and relatively stable prediction can be achieved under various lighting environments and head postures.

[0095] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments of the present disclosure. Obviously, the following introduced drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0097] Figure 1 Schematic diagram of the application scenario of the neural network model training method provided by the embodiments of the present disclosure;

[0098] FIG. 2(a) is a schematic diagram of starting the front camera provided by the embodiments of the present disclosure;

[0099] FIG. 2(b) is a schematic diagram of collecting facial images provided by the embodiments of the present disclosure;

[0100] Figure 3 Main flowchart provided by an embodiment of the present disclosure;

[0101] Figure 4 One of the gaze position acquisition models provided by an embodiment of the present disclosure;

[0102] Figure 5 Another gaze position acquisition model provided by the embodiments of the present disclosure;

[0103] FIG. 6(a) is the third gaze position acquisition model provided by the embodiments of the present disclosure;

[0104] FIG. 6(b) is a schematic diagram of the bounding boxes of the face and the left and right eyes provided by the embodiments of the present disclosure;

[0105] FIG. 6(c) is the left and right eye feature extraction module provided by the embodiments of the present disclosure;

[0106] FIG. 6(d) is the correction module provided by the embodiments of the present disclosure;

[0107] FIG. 6(e) is the multi-layer perceptron network module provided by the embodiments of the present disclosure;

[0108] Figure 7 Flow chart of obtaining fixation position provided by an embodiment of the present disclosure;

[0109] Figure 8 Flow chart of left-eye feature expression provided by an embodiment of the present disclosure;

[0110] Figure 9 Flow chart of right-eye feature expression provided by an embodiment of the present disclosure;

[0111] Figure 10 Flow chart of normalization provided by an embodiment of the present disclosure;

[0112] Figure 11 Block diagram of device for determining fixation position provided by an embodiment of the present disclosure;

[0113] Figure 12 Schematic structural diagram of an electronic device shown according to an exemplary embodiment provided by an embodiment of the present disclosure. Detailed implementation manners

[0114] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0115] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0116] Hereinafter, some terms in the embodiments of the present disclosure will be explained to facilitate the understanding of those skilled in the art.

[0117] (1) In the embodiments of the present disclosure, the term "a plurality of" means two or more, and other quantifiers are similar thereto.

[0118] (2) "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0119] (3) The server serves the terminal, and the services include providing resources to the terminal and storing terminal data; the server corresponds to the application installed on the terminal and runs in cooperation with the application on the terminal.

[0120] (4) The terminal device can refer to either a software-based APP (Application) or a client. It has a visible display interface and can interact with users; it corresponds to the server and provides local services to customers. For software-based applications, except for some applications that only run locally, they are generally installed on ordinary customer terminals and need to run in cooperation with the server. After the development of the Internet, commonly used applications include short video applications, email clients when sending and receiving emails, and instant messaging clients, etc. For this type of application, corresponding servers and service programs in the network are required to provide corresponding services, such as database services, configuration parameter services, etc. In this way, specific communication connections need to be established between the customer terminal and the server to ensure the normal operation of the application.

[0121] (5) Multilayer Perceptron (MLP): A feedforward artificial neural network model that maps multiple input data sets to a single output data set.

[0122] (6) Pooling layer: Reduces the number of output values by reducing the size of the input, generally accomplished through simple maximum, minimum, or average operations.

[0123] (7) Adaptive Group Normalization (adagn) layer: Performs group normalization in the channel direction of the input to facilitate variable control.

[0124] (8) Squeeze-and-Excitation layer (SElayer): Used to enhance the sensitivity of the model to channel mixed features.

[0125] (9) Concatenate layer: Used to fuse the features extracted by multiple convolutional features or fuse the output information.

[0126] Eye sight usually contains a lot of information. It can reflect points of interest, concentration, and even psychological state. Automatic real-time estimation of eye sight is of great value to various researchers and daily production and life. However, in the past, in order to achieve more accurate sight estimation, it is usually necessary to purchase professional equipment, and those estimation methods based on simple cameras are often very unreliable in complex daily life. In recent years, mobile devices have been popularized rapidly, and the hardware level of devices has also been gradually improved, which provides a strong guarantee for the collection of images of a certain quality. Therefore, how to enable most mobile devices to have the function of more accurately estimating the user's sight direction has become an emerging research direction in the fields of computer vision, virtual reality, and deep learning.

[0127] Many different technical routes have been proposed in the field of gaze estimation recently, which can be summarized into three categories: the technical route of estimating the three-dimensional direction of gaze based on eye model reconstruction; the technical route of estimating the screen gaze point based on the regression of two-dimensional eye features; and the technical route based on face. Among them, the three-dimensional eye model reconstruction is to build a three-dimensional geometric model of the eye and use it to estimate the gaze. The eye model of each object is different, so this technical route requires capturing a lot of object information to reconstruct the eye model of the object, such as measuring the iris radius, etc., and this method also requires the use of professional equipment to collect a lot of current object information, so the accuracy of the three-dimensional eye model reconstruction is still relatively satisfactory. The requirements of the two-dimensional technical route for equipment are basically the same as those of the three-dimensional technical route. The two-dimensional technical route directly uses the measured information such as the pupil center and eyelid to regress the position of the gaze point on the screen. The most important difference between the face-based technical route and the first two categories is that the requirements for hardware equipment are very low. It collects facial information through an ordinary webcam and directly regresses the gaze direction or gaze point based on the collected image. Although the face-based technical route has lower hardware requirements, its process is also relatively more complicated. First, a feature extractor needs to be designed to effectively extract useful features from complex raw high-dimensional data; second, a robust regression function is needed to map the raw features to the coordinates of the gaze point or the gaze reverse direction; finally, a large amount of labeled data is needed to train the neural network to fit this objective function. This method usually requires each subject to gaze at some fixed points on the screen before use, and collect the corrected data of the object. This process has poor versatility in determining the gaze position, and the process is cumbersome and inefficient. Therefore, how to determine the gaze position on ordinary devices remains to be studied.

[0128] In view of this, in order to solve the above problems, the embodiments of the present disclosure provide a method for determining a gaze position and a related device.

[0129] In the embodiments of the present disclosure, in order to determine the gaze position on ordinary devices, another method based on the technical route of facial features is proposed. In this method, the image captured by the camera is decomposed into the left eye region, the right eye region, and the facial region, and then the comprehensive features are obtained by analyzing these three regions. Based on the comprehensive features, the left eye region image and the right eye region image are further analyzed to obtain the left eye feature expression and the right eye feature expression. Finally, the gaze position is obtained by further combining the comprehensive features and the facial features. Only important features, including facial features, comprehensive features, left eye feature expression, and right eye feature expression, need to be extracted during the whole process, and then the gaze position of the human eye can be classified based on these features. The user does not need to fixate on a fixed point to collect calibration data, and the accuracy of determining the gaze position can be ensured through feature descriptions at multiple levels.

[0130] Reference Figure 1 , which is a schematic diagram of the application scenario of the method for determining the gaze position provided by the embodiments of the present disclosure. The application scenario includes multiple terminal devices 101 (including terminal devices 101-1, terminal devices 101-2,..., terminal devices 101-n), and also includes a server 102. Among them, the terminal devices 101 and the server 102 are connected through a wireless or wired network. The terminal devices 101 include, but are not limited to, electronic devices such as desktop computers, mobile phones, mobile computers, tablets, media players, smart wearable devices, and smart TVs. The server 102 can be a single server, a server cluster composed of several servers, or a cloud computing center. The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0131] Of course, the method provided by the embodiments of the present disclosure is not limited to Figure 1 the application scenario shown, and can also be used in other possible application scenarios, which are not restricted by the embodiments of the present disclosure. The functions that can be realized by each device in Figure 1 the application scenario shown will be described together in the subsequent method embodiments, and will not be elaborated here too much.

[0132] In the page shown in FIG. 2(a), the user can enter the front camera mode by clicking the camera icon on FIG. 2(a) through the camera function provided by the terminal device 101, and collect the user image based on the front camera. The collected image is shown in FIG. 2(b). In the front camera mode, the camera collects the user's facial image in real time, and the terminal device 101 analyzes the collected user's facial image to obtain the user's gaze position. Then, the analyzed gaze position and its corresponding interface information can be notified to the server 102. Of course, it should be noted that any information about the user in the embodiments of the present disclosure can be obtained after the user's authorization.

[0133] To further illustrate the technical solutions provided by the embodiments of the present disclosure, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Although the embodiments of the present disclosure provide the method operation steps as shown in the following embodiments or drawings, more or fewer operation steps may be included in the method based on routine or non-creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present disclosure.

[0134] For ease of understanding, first, the main processes involved in the embodiments of the present disclosure will be described, as Figure 3 shown:

[0135] In step 301, an image collected by the camera is obtained.

[0136] In step 302, the facial region, left eye region, and right eye region are extracted from the image.

[0137] In step 303, a neural network model is used to process the extracted facial region, left eye region, and right eye region to obtain the human eye gaze position information, where the gaze position information includes gaze point coordinates and / or the region where the gaze point is located.

[0138] One of the gaze position acquisition models is as Figure 4 shown. This model includes a label network Label-Net model (also referred to as the Label-Net model hereinafter), a face network Face-net model (also referred to as the Face-Net model hereinafter), and an eye network EyeNet model (also referred to as the Eye Net model hereinafter).

[0139] Among them, the input of the Label-Net model is the left eye region, right eye region, and facial region images, and the output is the extracted comprehensive features.

[0140] The input of the Face-net model is the facial region image, and the output is the facial features.

[0141] The Eye Net model includes two modules, one is the left-eye feature extraction module, and the other is the right-eye feature extraction module.

[0142] The input of the left-eye feature extraction module is the comprehensive feature, the facial feature, and the left-eye region image, and the output is the left-eye feature expression.

[0143] The input of the right-eye feature extraction module is the comprehensive feature, the facial feature, and the right-eye region image, and the output is the right-eye feature expression.

[0144] Finally, the neural network model integrates and classifies the left-eye feature expression, the right-eye feature expression, the facial feature, and the comprehensive feature to obtain the final human eye fixation position.

[0145] In some embodiments, as Figure 5 shown, in the embodiments of the present disclosure, in order to improve the accuracy of determining the fixation position, a correction (CCB, Context Correlation Block) module (also referred to as the CCB module hereinafter) is proposed. The CCB module is built into the Eye Net model and can have multiple ones to improve the accuracy of extracting the left and right eye feature expressions.

[0146] In addition, in some other embodiments, in the embodiments of the present disclosure, an MLP module (i.e., a multi-layer perceptron) and a channel mixing module are also connected to the backend of the Eye Net model, which are used to further process various features to mix the human eye features and the facial features to improve the accuracy of the feature expression finally used to determine the fixation point position.

[0147] For ease of understanding, the neural network model structure in the embodiments of the present disclosure will be further explained below. As shown in FIG. 6(a), it is the third fixation position acquisition model proposed by the present disclosure. After obtaining the facial image captured by the terminal, a facial feature point detection algorithm is used to obtain a series of key points of the face, and according to the key point information, the schematic diagrams of the bounding boxes of the face and the left and right eyes shown in FIG. 6(b) are extracted. The bounding box can be composed of two coordinate points, the lower left and the upper right. The face region, the left-eye region, and the right-eye region are cropped according to the bounding box.

[0148] In Fig. 6(a), it includes Lable-Net, Face-Net and Eye-Net. The neural network model in Eye-Net is shown in Fig. 6(c). The CCB module in Fig. 6(c) is shown in Fig. 6(d). The MLP network module in Eye-Net is shown in Fig. 6(e). After the facial region, left eye region and right eye region are processed through n fully connected layers in Lable-Net, comprehensive features are obtained. After the facial region is processed through n convolutional layers and n SElayer layers in Face-Net, facial features are obtained. The comprehensive features and facial features are processed through the left and right eye feature extraction modules in Eye-Net to obtain the left eye feature expression and the right eye feature expression respectively. The left eye feature expression, right eye feature expression, comprehensive features and facial features are processed in the MLP network module in Eye-Net to obtain the features to be recognized. The features to be recognized are further processed through the fully connected layer m and the loss function to obtain the gaze position information of the target object.

[0149] The structures of the left eye feature extraction module and the right eye feature extraction module in Fig. 6(a) are shown in Fig. 6(c). In Fig. 6(c), for the convenience of understanding, the facial features are represented by circles and the comprehensive features are represented by hexagons. Taking the left eye as an example, after the image of the left eye region and the comprehensive features pass through the first convolutional layer, they are input together with the facial features and comprehensive features into the first adaptive group normalization (adagn) layer. Then, the data output by the first adagn layer is sequentially passed through a convolutional layer, a CCB module, a pooling layer, and a squeeze-and-excitation layer (SElayer) for feature extraction to obtain the first intermediate features. To make the training results more accurate, the first intermediate features, facial features, and comprehensive features are first input into the second adagn layer, and then the data output by the second adagn layer is sequentially processed through a convolutional layer, a CCB module, and a pooling layer to obtain the second intermediate result. Then, the second intermediate result, facial features, and comprehensive features are again input into the third adagn layer, and then the data output by the third adagn layer is sequentially processed through a convolutional layer, a CCB module, and an SElayer layer to obtain the third intermediate result. Then, the third intermediate result, facial features, and comprehensive features are first input into the fourth adagn layer, and then the data output by the fourth adagn layer is sequentially processed through a convolutional layer and a CCB module to obtain the left eye feature expression. Similarly, the right eye feature expression is extracted in the same way. The left eye feature expression, right eye feature expression, comprehensive features, and facial features will be processed in the concatenate layer.

[0150] In one embodiment, the left-eye feature extraction module for extracting the left-eye feature representation and the right-eye feature extraction module for extracting the right-eye feature representation have the same structure, and at least one of the following neural network layers at the same position in the left-eye feature extraction module and the right-eye feature extraction module share network parameters: convolutional layer, pooling layer, and SElayer. In this way, the number of parameters can be reduced during training, and the training complexity can be lowered.

[0151] As shown in FIG. 6(c), the upper row identifies the architecture of the left-eye feature extraction module, and the lower row identifies the structure of the right-eye feature extraction module. The structures of the two modules are similar, and the neural network sides of the two modules connected by a dashed line identify the shared neural network layers, where the CCB module is also shared.

[0152] The CCB module is shown in FIG. 6(d). After encoding, context information extraction, convolutional operation, and feature fusion of the first left-eye feature map and the first right-eye feature map in sequence in the CCB module, a second left-eye feature map and a second right-eye feature map are obtained. Taking the left eye as an example, the left-eye region image X L is input, and is respectively used as the value map Value Map to input the 1×1 convolutional kernel V to obtain the left-eye encoded feature, as the key map Key Map to input the 3×3 convolutional kernel K to obtain the left-eye context feature, and as the query value Query to input the concatenation (concat) layer. The Query is concatenated with the left-eye context feature in the concat layer to obtain the left-eye concatenated feature. The left-eye concatenated feature passes through the 1×1 convolutional kernel α and the convolutional kernel β in sequence to obtain the left-eye convolutional feature, and the left-eye fusion feature is obtained based on the left-eye convolutional feature and the left-eye encoded feature. Taking FIG. 6(d) as an example, in the embodiment of the present disclosure, the left-eye convolutional feature and the left-eye encoded feature are subjected to matrix multiplication operation to obtain the left-eye fusion feature. In another embodiment of the present disclosure, the left-eye convolutional feature and the left-eye encoded feature can also be subjected to matrix addition operation to obtain the left-eye fusion feature, and the present disclosure does not limit this.

[0153] Finally, the Fusion module processes the left-eye fusion feature and the left-eye context feature to obtain the second left-eye feature map Y L .

[0154] Similarly, the left-eye additional module for extracting the second left-eye feature map and the right-eye additional module for extracting the right-eye feature map also adopt the same network structure, and the convolutional layers at the same position in the network structure share network parameters. In this way, the number of parameters can be reduced during training, and the training complexity can be lowered.

[0155] In another embodiment, the processing of the Fusion module specifically includes the following process. First, two feature features are added together, then global information is obtained through a global pooling layer, and a fully connected layer feature is obtained through a fully connected layer. This fully connected layer feature is multiplied by the two feature features respectively and then added together to obtain the output result.

[0156] X L The sizes of / Xr after convolution by convolution kernel K and convolution kernel V are both H*W*C. After splicing, the size of the regional image grows to H*W*2C, then becomes H*W*D after convolution kernel α, becomes H*W*(3*3*ch) after convolution kernel β, and the size of the fused feature of the left or right eye after matrix multiplication is H*W*C. At this time, after being processed by the Fusion module, the size of the second left or right eye feature map is still H*W*C. Among them, D < 2C, and 3*3*ch = H*W.

[0157] In the MLP network module in FIG. 6(e), the left-eye feature expression, right-eye feature expression, comprehensive feature, and facial feature output in FIG. 6(c) are processed through a concatenate layer to obtain a global concatenated feature. Then, it is normalized through a channel fusion Layer-Norm layer to obtain a normalized feature map. Finally, it is processed in the n-layer MLP module to obtain the feature to be recognized. Since some data will be lost when the left-eye feature expression, right-eye feature expression, comprehensive feature, and facial feature are processed through the concatenate layer, it is necessary to perform channel mixing layer processing on the global concatenated feature to obtain the channel mixed feature. After the feature to be recognized and the channel mixed feature are processed through the fully connected layer m, and then processed through the loss function, the gaze position information of the target object is obtained.

[0158] After introducing the neural network model used in this disclosure, the solution of this disclosure will be further described below in combination with the flowchart.

[0159] As Figure 7 shown, it is the flowchart for obtaining the gaze position information of this disclosure, and the specific steps are as follows:

[0160] Step 701, obtain an image of the target object.

[0161] Step 702, parse the facial area, left-eye area, and right-eye area of the target object from the image.

[0162] Step 703, extract facial features from the facial area, and extract comprehensive features from the facial area, left-eye area, and right-eye area.

[0163] Step 704: Extract features from the left eye region, facial features, and comprehensive features to obtain the left eye feature representation.

[0164] Among them, as Figure 8 shown in the left eye feature representation flowchart provided by the embodiments of the present disclosure, the left eye feature representation needs to be obtained according to FIGS. 6(c) and 6(d), and specifically includes the following steps:

[0165] Step 801: Extract features from the left eye region, facial features, and comprehensive features to obtain the first left eye feature map.

[0166] Step 802: Perform an encoding operation on the first left eye feature map to obtain the left eye encoded feature of the first left eye feature map.

[0167] Step 803: Extract the context information of each feature point of the first left eye feature map to obtain the left eye context feature.

[0168] Step 804: Based on the first left eye feature map and the left eye context feature, splice the first left eye feature map and the left eye context feature to obtain the left eye spliced feature.

[0169] Step 805: Sequentially perform convolution operations on the left eye spliced feature to obtain the left eye convolution feature.

[0170] Step 806: Obtain the left eye fusion feature according to the left eye convolution feature and the left eye encoded feature.

[0171] Step 807: Use the fusion module to process the left eye fusion feature and the left eye context feature to obtain the second left eye feature map.

[0172] Step 808: Based on the first left eye feature map and the second left eye feature map, extract the left eye feature representation.

[0173] Step 705: Extract features from the right eye region, facial features, and the comprehensive features to obtain the right eye feature representation.

[0174] The left eye feature representation is extracted four times in total. The CCB module proposed by the present disclosure based on the neural network improves the extraction accuracy of the left eye feature representation. After mirroring the right eye region, the right eye region can share the same CCB module with the left eye region, and the convolutional layers at the same positions in the model structure share network parameters, reducing the training complexity.

[0175] Among them, as Figure 9 shown in the right eye feature representation flowchart provided by the embodiments of the present disclosure, the right eye feature representation also needs to be obtained according to FIGS. 6(c) and 6(d), and specifically includes the following steps:

[0176] Step 901: Extract features from the right eye region, facial features, and comprehensive features to obtain the first right eye feature map.

[0177] Step 902: Perform an encoding operation on the first right eye feature map to obtain the right eye encoding feature of the first right eye feature map.

[0178] Step 903: Extract the context information of each feature point of the first right eye feature map to obtain the right eye context feature.

[0179] Step 904: Based on the first right eye feature map and the right eye context feature, splice the first right eye feature map and the right eye context feature to obtain the right eye spliced feature.

[0180] Step 905: Sequentially perform convolution operations on the right eye spliced feature to obtain the right eye convolution feature.

[0181] Step 906: Based on the right eye convolution feature and the right eye encoding feature, obtain the right eye fusion feature.

[0182] Step 907: Use a fusion module to process the right eye fusion feature and the right eye context feature to obtain the second right eye feature map.

[0183] Step 908: Based on the first right eye feature map and the second right eye feature map, extract the right eye feature expression.

[0184] Step 706: Based on the left eye feature expression, right eye feature expression, facial features, and comprehensive features, obtain the gaze position information of the target object.

[0185] After the left eye and right eye feature extraction models perform convolution, pooling, and context feature operations on the first left eye feature map and the first right eye feature map, the output results of the left eye feature expression and the right eye feature expression are more accurate, reducing the processing error in the MLP network model.

[0186] Among them, as Figure 10 shown, for the MLP network module and the channel mixing layer provided in the embodiments of the present disclosure, after being processed by the left eye and right eye feature extraction models, the gaze position of the target object still needs to be obtained according to the MLP network module and the channel mixing layer shown in FIG. 6(e), specifically including the following steps:

[0187] Step 1001: Perform splicing processing on the left eye feature expression, right eye feature expression, facial features, and comprehensive features to obtain a global spliced feature;

[0188] Step 1002: Perform normalization processing on the global spliced feature to obtain a normalized feature map;

[0189] In step 1003, the features of each channel of the normalized feature map are processed using a multi-layer perceptron (MLP) network module to obtain the features to be recognized.

[0190] In step 1004, channel mixing is performed on the global concatenated features to obtain channel-mixed features.

[0191] In step 1005, a first fully connected layer is used to process the features to be recognized to obtain the fixation position of the target object.

[0192] In one embodiment, the neural network layer used for the encoding operation is a convolutional layer with a convolution kernel of 1×1. Compared with a convolution kernel of n×n (n>1), using this convolution kernel can improve the running speed of the present disclosure. In the present disclosure, when the number of convolution kernels is two, the effect of the convolution kernel is the best and the operation speed is not affected.

[0193] In another embodiment, the neural network layer for extracting context information is a convolutional layer with a convolution kernel of n×n, that is, the K: 3×3 layer in FIG. 6(d). Where n is greater than 1 and less than a specified value, and n is a positive integer. Using a convolutional layer of n×n can sense the information of the line of sight, which is beneficial to extracting context information. To improve the accuracy, it is recommended that n be set to 3 in the embodiments of the present disclosure. In the present disclosure, there are two representation methods for the fixation position information. One is represented by a four-grid area, and the other is represented by a coordinate system. Combining the two methods for positioning makes the fixation position information more accurate.

[0194] Based on the same inventive concept, the application also proposes a neural network model training device for determining the fixation position. Figure 11 It is a block diagram showing the device according to an exemplary embodiment. Referring to Figure 11 , the device 1100 includes:

[0195] An image processing module 1101, configured to execute acquiring an image of a target object;

[0196] A region recognition module 1102 is configured to execute parsing out the facial region, left eye region, and right eye region of the target object from the image;

[0197] A comprehensive feature extraction module 1103, configured to execute extracting facial features from the facial region, and extracting comprehensive features from the facial region, the left eye region, and the right eye region;

[0198] A binocular feature extraction module 1104, configured to execute extracting features from the left eye region, the facial features, and the comprehensive features to obtain a left eye feature expression; and extracting features from the right eye region, the facial features, and the comprehensive features to obtain a right eye feature expression;

[0199] A fixation position determination module 1105, configured to obtain fixation position information of the target object based on the left-eye feature representation, the right-eye feature representation, the facial feature, and the comprehensive feature, where the fixation position information includes fixation point coordinates and / or the region where the fixation point is located.

[0200] Optionally, when performing feature extraction on the left-eye region, the facial feature, and the comprehensive feature to obtain the left-eye feature representation, the binocular feature extraction module 1104 is specifically configured to perform:

[0201] Perform feature extraction on the left-eye region, the facial feature, and the comprehensive feature to obtain a first left-eye feature map;

[0202] Perform an encoding operation on the first left-eye feature map to obtain the left-eye encoded feature of the first left-eye feature map; and extract the context information of each feature point of the first left-eye feature map to obtain the left-eye context feature;

[0203] Based on the left-eye encoded feature and the left-eye context feature, obtain a second left-eye feature map;

[0204] Based on the first left-eye feature map and the second left-eye feature map, extract the left-eye feature representation.

[0205] Optionally, when performing the operation of obtaining the second left-eye feature map based on the left-eye encoded feature and the left-eye context feature, the binocular feature extraction module 1104 is specifically configured to perform:

[0206] Concatenate the first left-eye feature map and the left-eye context feature to obtain a left-eye concatenated feature;

[0207] Perform a convolution operation on the left-eye concatenated feature in sequence to obtain a left-eye convolution feature;

[0208] According to the left-eye convolution feature and the left-eye encoded feature, obtain a left-eye fusion feature;

[0209] Use a fusion module to process the left-eye fusion feature and the left-eye context feature to obtain the second left-eye feature map.

[0210] Optionally, when performing feature extraction on the right-eye region, the facial feature, and the comprehensive feature to obtain the right-eye feature representation, the binocular feature extraction module 1104 is specifically configured to perform:

[0211] Perform feature extraction on the right-eye region, the facial feature, and the comprehensive feature to obtain a first right-eye feature map;

[0212] Perform an encoding operation on the first right-eye feature map to obtain the right-eye encoded feature of the first right-eye feature map; and extract the context information of each feature point of the first right-eye feature map to obtain the right-eye context feature;

[0213] Based on the right-eye encoded feature and the right-eye context feature, obtain a second right-eye feature map;

[0214] Based on the first right-eye feature map and the second right-eye feature map, extract the right-eye feature expression.

[0215] Optionally, when performing the operation of obtaining the second right-eye feature map based on the right-eye encoded feature and the right-eye context feature, the binocular feature extraction module 1104 is specifically configured to perform:

[0216] Concatenate the first left and right eye feature map and the right-eye context feature to obtain a right-eye concatenated feature;

[0217] Perform a convolution operation on the right-eye concatenated feature in sequence to obtain a right-eye convolution feature;

[0218] Based on the right-eye convolution feature and the right-eye encoded feature, obtain a right-eye fusion feature;

[0219] Use a fusion module to process the right-eye fusion feature and the right-eye context feature to obtain the second right-eye feature map.

[0220] Optionally, the neural network layer used for the encoding operation is a convolution layer with a convolution kernel of 1*1.

[0221] Optionally, the neural network layer for extracting context information is a convolution layer with a convolution kernel of n*n, where n is greater than 1 and less than a specified value, and n is a positive integer.

[0222] Optionally, when performing the operation of obtaining the gaze position information of the target object based on the left-eye feature expression, the right-eye feature expression, the facial feature, and the comprehensive feature, the gaze position determination module 1105 is specifically configured to perform:

[0223] Perform a concatenation process on the left-eye feature expression, the right-eye feature expression, the facial feature, and the comprehensive feature to obtain a global concatenated feature;

[0224] Perform a normalization process on the global concatenated feature to obtain a normalized feature map;

[0225] Process the channel features of the normalized feature map using a multi-layer perceptron MLP network module to obtain a feature to be recognized;

[0226] Perform channel mixing on the global stitching feature to obtain a channel-mixed feature;

[0227] Use a first fully connected layer to process the feature to be recognized and the channel-mixed feature to obtain the gaze position information of the target object.

[0228] Optionally, if the gaze position information includes the area where the gaze point of the target object is located, determining the area includes:

[0229] A classification module 1106, configured to perform a classification operation on the feature to be recognized to obtain the area classification result, where the area classification result is used to indicate the area where the gaze position of the target object is located.

[0230] Optionally, the left-eye feature extraction module for extracting the left-eye feature expression and the right-eye feature extraction module for extracting the right-eye feature expression have the same structure, and at least one of the following neural network layers at the same position in the left-eye feature extraction module and the right-eye feature extraction module share network parameters:

[0231] Convolutional layer, pooling layer, squeeze-and-excitation layer.

[0232] Optionally, the left-eye additional module 1107 for extracting the second left-eye feature map and the right-eye additional module 1108 for extracting the right-eye feature map adopt the same network structure, and the convolutional layers at the same position in the network structure share network parameters.

[0233] After introducing the method and of the exemplary embodiments of the present disclosure, next, an electronic device according to another exemplary embodiment of the present disclosure is introduced.

[0234] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0235] In some possible implementation manners, the electronic device according to the present disclosure may at least include at least one processor and at least one memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor executes the task scheduling method according to various exemplary embodiments of the present disclosure described above in this specification. For example, the processor may execute the steps in the task scheduling method.

[0236] Next, refer to Figure 12 to describe the electronic device according to this embodiment of the present disclosure.Figure 12 The electronic device shown is only an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.

[0237] As Figure 12 shown, the electronic device 130 is presented in the form of a general electronic device. The components of the electronic device 130 may include, but are not limited to: at least one of the above-mentioned processors 131, at least one of the above-mentioned memories 132, and a bus 133 connecting different system components (including the memory 132 and the processor 131).

[0238] The bus 133 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a processor, or a local bus using any bus structure in a variety of bus structures.

[0239] The memory 132 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 1321 and / or a cache memory 1322, and may further include a read-only memory (ROM) 1323.

[0240] The memory 132 may further include a program / utilities 1325 having a set (at least one) of program modules 1324. Such program modules 1324 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0241] The electronic device 130 may also communicate with one or more external devices 134 (such as a keyboard, a pointing device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 130, and / or communicate with any device that enables the electronic device 130 to communicate with one or more other electronic devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 135. And, the electronic device 130 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 136. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0242] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as the memory 132 including instructions. Optionally, the storage medium may be a non-transitory computer-readable storage medium. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, magnetic tape, a floppy disk, and an optical data storage device, etc.

[0243] In an exemplary embodiment, a computer program product is also provided, including a computer program, which when executed by the processor 131 implements any of the task scheduling methods provided in this disclosure.

[0244] In an exemplary embodiment, various aspects of the task scheduling method provided in this disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps in the task scheduling method according to various exemplary embodiments described above in this specification.

[0245] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0246] The program product for the task scheduling method according to the embodiments of this disclosure can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can run on an electronic device. However, the program product of this disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0247] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0248] The program code contained on the readable medium can be transmitted by any suitable medium, including - but not limited to - wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the above.

[0249] The program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, C++, etc., and also including conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's electronic device, partially on the user's device, executed as a stand-alone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external electronic device (e.g., by connecting through the Internet using an Internet service provider).

[0250] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that the operations must be performed in that specific order, or that all of the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0251] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0252] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable image scaling devices to produce a machine, such that the instructions executed by the processor of the computer or other programmable image scaling devices produce means for realizing the Figure 1 one or more flows or multiple flows and / or blocks Figure 1means for the functions specified in one or more blocks.

[0253] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the functions specified in one or more of the procedures Figure 1 one or more of the procedures and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0254] These computer program instructions may also be loaded onto a computer or other programmable apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the procedures Figure 1 one or more of the procedures and / or blocks Figure 1 means for the functions specified in one or more blocks.

[0255] Although embodiments of the present disclosure have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present disclosure.

[0256] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these modifications and variations.

Claims

1. A method for determining a fixation position, characterized in that, the method comprises: acquiring an image of a target object; parsing the facial region, left eye region and right eye region of the target object from the image; extracting facial features from the facial region and extracting comprehensive features from the facial region, the left eye region and the right eye region; extracting features from the left eye region, the facial features and the comprehensive features to obtain a left eye feature representation; and extracting features from the right eye region, the facial features and the comprehensive features to obtain a right eye feature representation; performing splicing processing on the left eye feature representation, the right eye feature representation, the facial features and the comprehensive features to obtain a global splicing feature, and performing normalization processing on the global splicing feature to obtain a normalized feature map; processing each channel feature of the normalized feature map by using a multi-layer perceptron MLP network module to obtain a feature to be recognized; performing channel mixing on the global splicing feature to obtain a channel mixing feature; processing the feature to be recognized and the channel mixing feature by using a first fully connected layer to obtain the fixation position information of the target object; wherein the fixation position information includes fixation point coordinates and / or the region where the fixation point is located.

2. The method according to claim 1, characterized in that, the extracting features from the left eye region, the facial features and the comprehensive features to obtain a left eye feature representation includes: extracting features from the left eye region, the facial features and the comprehensive features to obtain a first left eye feature map; performing an encoding operation on the first left eye feature map to obtain an encoded left eye feature of the first left eye feature map; and extracting context information of each feature point of the first left eye feature map to obtain a left eye context feature; obtaining a second left eye feature map based on the encoded left eye feature and the left eye context feature; extracting the left eye feature representation based on the first left eye feature map and the second left eye feature map.

3. The method according to claim 2, characterized in that, the obtaining a second left eye feature map based on the encoded left eye feature and the left eye context feature includes: splicing the first left eye feature map and the left eye context feature to obtain a left eye splicing feature; performing a convolution operation on the left eye splicing feature in sequence to obtain a left eye convolution feature; obtaining a left eye fusion feature according to the left eye convolution feature and the encoded left eye feature; processing the left eye fusion feature and the left eye context feature by using a fusion module to obtain the second left eye feature map.

4. The method according to claim 1, characterized in that, the extracting features from the right eye region, the facial features and the comprehensive features to obtain a right eye feature representation includes: extracting features from the right eye region, the facial features and the comprehensive features to obtain a first right eye feature map; Perform an encoding operation on the first right-eye feature map to obtain the right-eye encoded feature of the first right-eye feature map; and extract the context information of each feature point of the first right-eye feature map to obtain the right-eye context feature; Based on the right-eye encoded feature and the right-eye context feature, obtain a second right-eye feature map; Based on the first right-eye feature map and the second right-eye feature map, extract the right-eye feature expression.

5. The method according to claim 4, wherein, the obtaining of the second right-eye feature map based on the right-eye encoded feature and the right-eye context feature includes: Concatenate the first right-eye feature map and the right-eye context feature to obtain a right-eye concatenated feature; Perform a convolution operation on the right-eye concatenated feature in sequence to obtain a right-eye convolution feature; Based on the right-eye convolution feature and the right-eye encoded feature, obtain a right-eye fusion feature; Use a fusion module to process the right-eye fusion feature and the right-eye context feature to obtain the second right-eye feature map.

6. The method according to claim 2 or 4, wherein, The neural network layer used for the encoding operation is a convolution layer with a convolution kernel of 1*1.

7. The method according to claim 2 or 4, wherein, The neural network layer for extracting context information is a convolution layer with a convolution kernel of n*n, where n is greater than 1 and less than a specified value, and n is a positive integer.

8. The method according to claim 1, wherein, If the gaze position information includes the area where the gaze point of the target object is located, then determining the area includes: Perform a classification operation on the feature to be recognized to obtain an area classification result, and the area classification result is used to indicate the area where the gaze point of the target object is located.

9. The method according to claim 1, wherein, The left-eye feature extraction module for extracting the left-eye feature expression and the right-eye feature extraction module for extracting the right-eye feature expression have the same structure, and at least one of the following neural network layers at the same position in the left-eye feature extraction module and the right-eye feature extraction module share network parameters: Convolution layer, pooling layer, squeeze and excitation layer.

10. The method according to claim 4, wherein, The left-eye additional module for extracting the second left-eye feature map and the right-eye additional module for extracting the right-eye feature map adopt the same network structure, and the convolution layers at the same position in the network structure share network parameters.

11. A device for determining a gaze position, wherein, the device includes: An image acquisition module configured to acquire an image of a target object; A region recognition module configured to parse out the facial region, left-eye region, and right-eye region of the target object from the image; A comprehensive feature extraction module configured to perform feature extraction on the facial region to obtain facial features, and perform feature extraction on the facial region, the left-eye region, and the right-eye region to obtain comprehensive features; The binocular feature extraction module is configured to perform feature extraction on the left eye region, the facial features, and the comprehensive features to obtain a left eye feature representation; and perform feature extraction on the right eye region, the facial features, and the comprehensive features to obtain a right eye feature representation; The gaze position determination module is configured to perform splicing processing on the left eye feature representation, the right eye feature representation, the facial features, and the comprehensive features to obtain a global spliced feature, and perform normalization processing on the global spliced feature to obtain a normalized feature map; Use the multi-layer perceptron MLP network module to process the channel features of the normalized feature map to obtain the feature to be recognized; Perform channel mixing on the global spliced feature to obtain a channel mixed feature; Use the first fully connected layer to process the feature to be recognized and the channel mixed feature to obtain the gaze position information of the target object; where the gaze position information includes gaze point coordinates and / or the region where the gaze point is located.

12. The apparatus according to claim 11, wherein, When performing the feature extraction on the left eye region, the facial features, and the comprehensive features to obtain the left eye feature representation, the binocular feature extraction module is specifically configured to perform: Perform feature extraction on the left eye region, the facial features, and the comprehensive features to obtain a first left eye feature map; Perform an encoding operation on the first left eye feature map to obtain the left eye encoded feature of the first left eye feature map; And extract the context information of each feature point of the first left eye feature map to obtain the left eye context feature; Based on the left eye encoded feature and the left eye context feature, obtain a second left eye feature map; Based on the first left eye feature map and the second left eye feature map, extract the left eye feature representation.

13. The apparatus according to claim 12, wherein, When performing the operation of obtaining the second left eye feature map based on the left eye encoded feature and the left eye context feature, the binocular feature extraction module is specifically configured to perform: Splice the first left eye feature map and the left eye context feature to obtain a left eye spliced feature; Perform a convolution operation on the left eye spliced feature in sequence to obtain a left eye convolution feature; Based on the left eye convolution feature and the left eye encoded feature, obtain a left eye fusion feature; Use a fusion module to process the left eye fusion feature and the left eye context feature to obtain the second left eye feature map.

14. The apparatus according to claim 11, wherein, When performing the feature extraction on the right eye region, the facial features, and the comprehensive features to obtain the right eye feature representation, the binocular feature extraction module is specifically configured to perform: Perform feature extraction on the right eye region, the facial features, and the comprehensive features to obtain a first right eye feature map; Perform an encoding operation on the first right eye feature map to obtain the right eye encoded feature of the first right eye feature map; And extract the context information of each feature point of the first right eye feature map to obtain the right eye context feature; Based on the right-eye coding feature and the right-eye context feature, a second right-eye feature map is obtained; Based on the first right-eye feature map and the second right-eye feature map, the right-eye feature expression is extracted.

15. The apparatus according to claim 14, wherein, When performing the operation of obtaining the second right-eye feature map based on the right-eye coding feature and the right-eye context feature, the binocular feature extraction module is specifically configured to perform: Concatenate the first right-eye feature map and the right-eye context feature to obtain a right-eye concatenated feature; Perform a convolution operation on the right-eye concatenated feature in sequence to obtain a right-eye convolution feature; Based on the right-eye convolution feature and the right-eye coding feature, obtain a right-eye fusion feature; Use a fusion module to process the right-eye fusion feature and the right-eye context feature to obtain the second right-eye feature map.

16. The apparatus according to claim 12 or 14, wherein, The neural network layer used for performing the encoding operation is a convolution layer with a convolution kernel of 1*1.

17. The apparatus according to claim 12 or 14, wherein, The neural network layer for extracting context information is a convolution layer with a convolution kernel of n*n, where n is greater than 1 and less than a specified value, and n is a positive integer.

18. The apparatus according to claim 17, wherein, If the gaze position information includes the area where the gaze point of the target object is located, then determining the area includes: A classification module, configured to perform a classification operation on the feature to be recognized to obtain an area classification result, and the area classification result is used to indicate the area where the gaze position of the target object is located.

19. The apparatus according to claim 11, wherein, The left-eye feature extraction module for extracting the left-eye feature expression and the right-eye feature extraction module for extracting the right-eye feature expression have the same structure, and at least one of the following neural network layers at the same position in the left-eye feature extraction module and the right-eye feature extraction module share network parameters: Convolution layer, pooling layer, squeeze and excitation layer.

20. The apparatus according to claim 14, wherein, The left-eye additional module for extracting the second left-eye feature map and the right-eye additional module for extracting the right-eye feature map adopt the same network structure, and the convolution layers at the same position in the network structure share network parameters.

21. An electronic device, wherein, comprising: A processor; A memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the method for determining the gaze position according to any one of claims 1-10.

22. A computer-readable storage medium, wherein, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the method for determining the gaze position according to any one of claims 1-10.

23. A computer program product, including a computer program, wherein, When the computer program is executed by a processor, it implements the method for determining a gaze position according to any one of claims 1-10.