A gaze point prediction model training method and device, and electronic equipment
By constructing a gaze point hypergraph learning model and adjusting parameters, the problems of lighting environment and eye tremor in traditional gaze point prediction methods are solved, achieving more efficient and accurate user gaze point prediction, which is suitable for smart terminal devices.
Patent Information
- Application Number
- CN202211000132.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-08-19
AI Technical Summary
In existing technologies, information push methods based on user click behavior suffer from problems such as false touch noise and inaccurate discovery of the user's true intent. Traditional gaze point prediction methods cannot accurately predict the user's gaze point under the influence of factors such as lighting environment, inherent eye tremors, and changes in head posture.
Using a pre-defined dynamic eye-tracking dataset, multiple sets of sample image data and ground truth coordinates of gaze points are obtained. A gaze hypergraph is constructed using a gaze hypergraph learning model, and the accuracy of gaze point prediction is improved by adjusting the model parameters.
It achieves more accurate capture of the temporal relationship between the user's eye movement trajectory and frame images, improving the efficiency and accuracy of gaze point prediction, and is suitable for smart terminal devices.
Smart Images

Figure CN115359092B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the field of eye-tracking technology. Background Technology
[0002] With the widespread use of portable mobile devices (such as smartphones and tablets), users have become accustomed to browsing information pushed by various platforms on their mobile devices. Therefore, how to accurately target information to user needs is one of the important research directions for various platforms.
[0003] Currently, determining user needs is often achieved by mining user behavioral intentions. For example, by observing users' click-through behavior on information and the length of time their eyes linger on the information, we can determine the user's level of interest in it and identify their needs accordingly. Summary of the Invention
[0004] This disclosure provides a training method, apparatus, and electronic device for a gaze prediction model.
[0005] According to one aspect of this disclosure, a method for training a gaze prediction model is provided, comprising:
[0006] Based on a preset dynamic eye-tracking dataset, multiple sets of first sample image data and gaze point ground truth coordinates are obtained. Each set of first sample image data is an image group data obtained frame by frame from each first sample video data collected when the sample object gazes at preset information.
[0007] The data of each of the first sample image groups are input into the gaze prediction model. A gaze hypergraph is constructed using the gaze hypergraph learning model in the gaze prediction model, and the predicted gaze coordinates are obtained based on the gaze hypergraph.
[0008] The parameters of the gaze prediction model are adjusted based on the distance between the predicted gaze coordinates and the ground truth gaze coordinates.
[0009] According to another aspect of this disclosure, a fixation prediction method is provided, comprising:
[0010] Acquire target image data, wherein the target image data includes the target object for which the gaze point to be predicted is located;
[0011] The target image data is input into a pre-trained gaze prediction model to obtain the predicted gaze coordinates of the target object, wherein the gaze prediction model is trained using any of the gaze prediction model training methods described above.
[0012] According to another aspect of this disclosure, a training apparatus for a gaze prediction model is provided, comprising:
[0013] The first data acquisition module is used to acquire multiple sets of first sample image data and gaze point ground truth coordinates based on a preset dynamic eye tracking dataset. Each set of first sample image data is an image group data obtained frame by frame from each set of first sample video data collected when the sample object gazes at preset information.
[0014] The first coordinate prediction module is used to input the data of each of the first sample image groups into the gaze prediction model, construct a gaze hypergraph using the gaze hypergraph learning model in the gaze prediction model, and obtain the predicted gaze coordinates based on the gaze hypergraph.
[0015] The first parameter adjustment module is used to adjust the parameters of the gaze point prediction model based on the distance between the predicted gaze point coordinates and the true gaze point coordinates.
[0016] According to another aspect of this disclosure, a gaze point prediction device is provided, comprising:
[0017] An image acquisition module is used to acquire target image data, wherein the target image data includes a target object for which the gaze point to be predicted is located;
[0018] A gaze point coordinate prediction module is used to input the target image data into a pre-trained gaze point prediction model to obtain the predicted gaze point coordinates of the target object, wherein the gaze point prediction model is trained by any of the gaze point prediction model training devices described above.
[0019] This disclosure uses a pre-defined dynamic eye-tracking dataset to acquire multiple sets of first sample image data and ground truth coordinates of the gaze point. Each first sample image set is a frame-by-frame capture of sample video data collected based on preset gaze information of the sample object. Then, each first sample image set is input into a gaze point prediction model. A gaze point hypergraph is constructed using the gaze point hypergraph learning model within the gaze point prediction model, and the predicted gaze point coordinates are obtained based on the gaze point hypergraph. Finally, the parameters of the gaze point prediction model are adjusted according to the distance between the predicted gaze point coordinates and the ground truth gaze point coordinates. This achieves the training of the gaze point prediction model.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0021] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0022] Figure 1 This is a flowchart illustrating the training method of the first gaze prediction model provided in this disclosure;
[0023] Figure 2 This is one possible implementation of step S11 provided in this disclosure;
[0024] Figure 3 This is one possible implementation of step S12 provided in this disclosure;
[0025] Figure 4a This is a flowchart illustrating the training method for the second gaze prediction model provided in this disclosure;
[0026] Figure 4b This is an example diagram illustrating the selection of camera positions based on a sample object provided in this disclosure;
[0027] Figure 4c This is an example diagram comparing the prediction accuracy of the dynamic eye-tracking dataset and fixation prediction model provided in this disclosure with that of datasets and models in the prior art;
[0028] Figure 5 This is one possible implementation of step S32 provided in this disclosure;
[0029] Figure 6 This is a flowchart illustrating a fixation prediction method provided in this disclosure;
[0030] Figure 7 This is a schematic diagram of the structure of a training device for a gaze point prediction model provided in this disclosure;
[0031] Figure 8 This is a schematic diagram of a gaze point prediction device provided in this disclosure;
[0032] Figure 9 This is a block diagram of an electronic device used to implement the gaze prediction model training method and the gaze prediction method according to the embodiments of this disclosure. Detailed Implementation
[0033] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0034] In existing technologies, when platforms try to understand user behavior and push information that matches their needs, they mostly rely on click data from users' historical browsing history, treating user clicks as evidence of interest. However, this approach has certain problems in practical applications. For example, the click data may contain a large amount of false click noise, or users may be attracted by misleading titles or thumbnails and click on content they are not actually interested in.
[0035] To more accurately uncover users' true behavioral intentions, researchers have proposed eye-tracking technology. This technology uses facial eye movement information, analyzing how long the user's gaze lingers on the information, combined with user click data, to determine the user's level of interest in the information. The content being pushed to users also reflects their level of interest to some extent. For example, if a message is both viewed and clicked by a user for a certain period, it is considered that the user's interest in that message is higher than in other messages that were only viewed or clicked by the user.
[0036] The goal of eye tracking is to predict the coordinates of a person's gaze point in the prediction space with the highest possible accuracy, based on features abstracted from images. These gaze point coordinates reflect the user's true behavioral intentions. Therefore, the key to eye tracking lies in the prediction of gaze points. Traditional gaze point prediction methods can be divided into two types: model-based methods and representation-based methods. Model-based methods rely on eye features detected by reflections from an external infrared light source on the outermost layer of the eye (cornea), which requires high-resolution images and uniform lighting conditions. Representation-based methods, on the other hand, infer gaze directly by detecting the shape of the eyeball, but therefore require large-scale eye-tracking data.
[0037] To address the shortcomings of traditional methods, existing technologies have proposed combining facial images, eye images, and the original image to predict gaze points. However, this approach also presents several challenges: Firstly, most public eye-tracking datasets are collected using devices like laptops, making gaze prediction models trained on such datasets unsuitable for various smart devices. Furthermore, image datasets are not necessarily sequential, making it difficult to capture eye movement trajectories. Factors such as lighting conditions, inherent eye tremors, head posture changes, and facial expressions of the volunteers providing the images can also affect eye feature extraction, potentially leading to a mismatch between training labels and actual gaze points. Secondly, existing eye-tracking models cannot model higher-order temporal relationships between frames. These issues all contribute to inaccurate gaze point predictions.
[0038] To address at least one of the aforementioned problems, this disclosure provides a method for training a fixation prediction model, comprising:
[0039] Based on a preset dynamic eye-tracking dataset, multiple sets of first sample image data and gaze point ground truth coordinates are obtained. Each set of first sample image data is a set of image data captured frame by frame from sample video data collected when the sample object gazes based on preset information.
[0040] The data of each of the first sample image groups are input into the gaze prediction model. A gaze hypergraph is constructed using the gaze hypergraph learning model in the gaze prediction model, and the predicted gaze coordinates are obtained based on the gaze hypergraph.
[0041] The parameters of the gaze prediction model are adjusted based on the distance between the predicted gaze coordinates and the ground truth gaze coordinates.
[0042] As can be seen from the above, the training method of the gaze prediction model provided in this disclosure is based on a preset dynamic eye tracking dataset, which obtains multiple sets of first sample image data and gaze point ground truth coordinates. Each set of first sample image data is a set of image data extracted frame by frame from sample video data collected when the sample object gazes at preset information. Therefore, each set of first sample image data is continuous in time. Based on such sample image data, the prediction of gaze point coordinates can more accurately capture the movement trajectory of the sample object's eyes.
[0043] Secondly, the data from each first sample image group are input into the gaze prediction model. A gaze hypergraph is constructed using the gaze hypergraph learning model within the model. Based on this hypergraph, the relationships between the first sample image groups can be determined. Since each sample image group is obtained frame-by-frame from the sample video data, determining the predicted gaze coordinates based on this hypergraph effectively considers the temporal and spatial relationships between frames in the sample video data, thus modeling higher-order temporal relationships between frames. The resulting predicted gaze coordinates have higher accuracy. Adjusting the parameters of the gaze prediction model based on the distance between the more accurate predicted gaze coordinates and the ground truth gaze coordinates improves the training efficiency of the model. The trained gaze prediction model effectively improves both the efficiency and accuracy of gaze prediction.
[0044] The training method of the gaze prediction model provided in this disclosure will be described in detail below through specific embodiments.
[0045] The method of this disclosure is applied to a smart terminal and can be implemented through a smart terminal. In actual use, the smart terminal can be a computer, server, data center, etc.
[0046] See Figure 1 , Figure 1 A flowchart illustrating the training method for the first gaze prediction model provided in this disclosure includes:
[0047] Step S11: Based on the preset dynamic eye-tracking dataset, acquire multiple sets of first sample image data and ground truth coordinates of the gaze point.
[0048] Among them, each of the first sample image group data is the image group data obtained frame by frame from the first sample video data collected when the sample object is looking at the preset information.
[0049] The aforementioned preset dynamic eye-tracking dataset is a pre-collected dataset that includes first sample video data of sample objects gazing at preset information. Specifically, the preset information includes multiple pre-defined pieces of information. In one example, the preset information could be words, phrases, graphics, images, etc., as long as it can make the sample objects focus their gaze.
[0050] In one example, the preset information is displayed sequentially in a preset order. Therefore, the sample object also gazes at each piece of preset information in the preset order. The sample video data collected when the sample object gazes at the preset information is acquired continuously, meaning the collected sample video data is sequential in time. After obtaining the sample video data of the sample object gazing at all the preset information, the sample video data is divided into video segments corresponding to each piece of preset information, thus obtaining each first sample video data. In other words, each first sample video data corresponds to a different piece of preset information.
[0051] In one example, the first sample video data includes 90-150 frames. For each first sample video data, each frame is cropped to obtain the first sample image group data corresponding to each frame. Specifically, the first sample image group data includes the left eye image, right eye image, facial image, and overall image of the sample object. That is, each frame of the first sample video data is cropped into a first sample image group data including the left eye image, right eye image, facial image, and overall image of the sample object. If the first sample video data includes 90-150 frames, 90-150 first sample image group data can be obtained after cropping.
[0052] In addition, after obtaining the data of each first sample image group, the coordinates of preset information can also be obtained, and these coordinates can be used as the true coordinates of the gaze point corresponding to the preset information, that is, the true coordinates of the gaze point corresponding to the first sample image group data corresponding to the preset information.
[0053] Step S12: Input the data of each of the first sample image groups into the gaze prediction model, construct a gaze hypergraph using the gaze hypergraph learning model in the gaze prediction model, and obtain the predicted gaze coordinates based on the gaze hypergraph.
[0054] After obtaining the data for each first sample image group, it is input into the gaze point prediction model. The gaze point prediction model includes a gaze point hypergraph learning model. The data for each first sample image group is used as the input required by the gaze point hypergraph learning model to construct the gaze point hypergraph. The nodes and hyperedges in the constructed gaze point hypergraph can represent the relationship between the images included in each first sample image group. Based on this, the gaze point prediction model can further predict the gaze point coordinates.
[0055] Step S13: Adjust the parameters of the gaze point prediction model based on the distance between the predicted gaze point coordinates and the true gaze point coordinates.
[0056] After obtaining the predicted gaze coordinates, the loss and error of the gaze prediction model are calculated based on the distance between the predicted gaze coordinates and the ground truth gaze coordinates. The parameters of the gaze prediction model are then adjusted based on the calculated loss and error. Following this, the parameter-adjusted gaze prediction model is repeatedly trained using a pre-defined dynamic eye-tracking dataset until the distance between the predicted gaze coordinates and the ground truth gaze coordinates output by the gaze prediction model meets the requirements. At this point, the gaze prediction model can be considered successfully trained.
[0057] As can be seen from the above, the training method of the gaze prediction model provided in this disclosure is based on a preset dynamic eye tracking dataset, which obtains multiple sets of first sample image data and gaze point ground truth coordinates. Each set of first sample image data is a set of image data extracted frame by frame from sample video data collected when the sample object gazes at preset information. Therefore, each set of first sample image data is continuous in time. Based on such sample image data, the prediction of gaze point coordinates can more accurately capture the movement trajectory of the sample object's eyes.
[0058] Secondly, the data from each first sample image group are input into the gaze prediction model. A gaze hypergraph is constructed using the gaze hypergraph learning model within the model. Based on this hypergraph, the relationships between the first sample image groups can be determined. Since each sample image group is obtained frame-by-frame from the sample video data, determining the predicted gaze coordinates based on this hypergraph effectively considers the temporal and spatial relationships between frames in the sample video data, thus modeling higher-order temporal relationships between frames. The resulting predicted gaze coordinates have higher accuracy. Adjusting the parameters of the gaze prediction model based on the distance between the more accurate predicted gaze coordinates and the ground truth gaze coordinates improves the training efficiency of the model. The trained gaze prediction model effectively improves both the efficiency and accuracy of gaze prediction.
[0059] In one possible implementation, such as Figure 2 As shown, step S11 above, based on a preset dynamic eye-tracking dataset, acquires multiple sets of first sample image data and ground truth coordinates of the gaze point, including:
[0060] Step S21: Based on the preset dynamic eye-tracking dataset, acquire the first sample video data and the second sample video data.
[0061] The first sample video data is a video recording taken when the sample object is looking at the preset information, and the second sample video data is a screen recording taken when the preset information is displayed on a smart device.
[0062] Step S22: Obtain the true coordinates of the gaze point of the sample object based on the second sample video data.
[0063] As mentioned above, the preset information includes multiple pre-defined pieces of information. The collected sample video data is a continuous video obtained by the sample object sequentially viewing multiple preset pieces of information. The video segments obtained by segmenting such continuous video according to different preset pieces of information are used as the first sample video data corresponding to each preset piece of information. In this embodiment, each first sample video data is a video recording collected when the sample object views different preset pieces of information. Each frame of the first sample video data includes the sample object; therefore, each frame is partially cropped to obtain the first sample image group data. The second sample video data is a screen recording video collected when the preset information is displayed on the smart device. That is, it is a screen recording video obtained by segmenting the screen recording video of the preset information displayed sequentially on the smart device screen according to different preset pieces of information. Therefore, each second sample video data has a one-to-one correspondence with each first sample video data.
[0064] Therefore, the true coordinates of the gaze point of the sample object for each preset information can be obtained based on each second sample video data.
[0065] Step S23: For each frame of the first sample video data, perform partial cropping to obtain the first sample image group data corresponding to each frame.
[0066] The data of each first sample image group includes the left eye image, right eye image, facial image, and overall portrait of the sample object.
[0067] As can be seen from the above, the training method of the gaze prediction model provided in this disclosure includes video recordings of sample objects gazing at preset information and screen recordings of preset information displayed on smart devices. These videos are segmented according to different preset information, so the gaze points of each segmented video segment are consistent. Each frame of the segmented video segment is extracted into the left eye image, right eye image, facial image, and overall portrait of the sample object. The amount of data of the extracted image group is very large, which can cover as many gaze point features of the sample object as possible. Based on this, the gaze prediction model can be trained to achieve the highest possible accuracy in gaze prediction.
[0068] In one possible implementation, such as Figure 3 As shown, the above-mentioned gaze prediction model further includes: a feature extraction network and a multilayer perceptron network; step S12 above inputs the data of each of the first sample image groups into the gaze prediction model, constructs a gaze hypergraph using the gaze hypergraph learning model in the gaze prediction model, and obtains the predicted gaze coordinates based on the gaze hypergraph, including:
[0069] Step S31: Input the data of each of the first sample image groups into the feature extraction network to obtain the sample image features of each of the sample image groups;
[0070] Step S32: Input the features of each sample image into the gaze hypergraph learning model to construct the gaze hypergraph, and obtain the gaze features based on the gaze hypergraph;
[0071] Step S33: Input the gaze point features of each of the above into the multilayer perceptron network to obtain the predicted gaze point coordinates.
[0072] Each set of sample images is input into a feature extraction network to extract features, thus obtaining the individual sample image features of each set. These features are then used as input to a gaze hypergraph learning model to construct a gaze hypergraph that effectively represents the relationships between the features of each sample image. Based on the constructed gaze hypergraph, gaze features are obtained. These gaze features are then input into a multilayer perceptron network to obtain the predicted gaze coordinates.
[0073] As can be seen from the above, the training method of the gaze prediction model provided in this disclosure includes a feature extraction network, a gaze hypergraph learning model, and a multilayer perceptron network. It utilizes a multi-structured model to decompose and analyze a large number of sample image data sets. At the same time, it uses the gaze hypergraph learning model to establish the correlation between images in each sample image data set, thereby realizing the construction of temporal and spatial correlations between frames and determining the higher-order temporal relationships between frame images. The gaze prediction performed in this way can effectively improve the accuracy.
[0074] In one embodiment of this disclosure, such as Figure 4a The diagram shows a flowchart illustrating a training method for a second gaze prediction model, wherein the gaze prediction model further includes a multi-layer neural network, and the method further includes:
[0075] Step S41: Based on the preset dynamic eye-tracking dataset, obtain sample boundary image data and ground truth coordinates of boundary points.
[0076] The sample boundary image data includes images captured when the sample object gazes at a preset boundary point.
[0077] The positions of the preset boundary points are pre-defined, so the true coordinates of the boundary points can be obtained directly from the pre-defined positions.
[0078] In one embodiment of this disclosure, the above-mentioned preset dynamic eye-tracking dataset is constructed as follows:
[0079] Send a command to the sample object indicating the selection of a camera location, so that the sample object selects a camera location on the screen of the smart device.
[0080] Specifically, considering that gaze point prediction is often used in push platforms on portable smart devices, the smart devices used to construct the preset dynamic eye-tracking dataset in this embodiment are portable smart devices, including smartphones and tablets. It is understood that the front-facing camera positions may differ between different smart devices. Therefore, before constructing the preset dynamic eye-tracking dataset, a command to select a camera position is first sent to the sample object, allowing the sample object to select the camera position on the smart device's screen. Specifically, selectable positions include the upper left corner, the top, the upper right corner, and other different locations on the screen. In one example, when the user selects the camera position, a ruler can be displayed on the edge of the smart device's screen. Specifically, the ruler coordinates can range from 0 to 15, where 0 represents the left side of the screen and 15 represents the right side, allowing the user to select specific camera position coordinates when identifying the camera position. Figure 4b As shown, the front-facing camera at the user's selected location is the acquisition device used to collect the first sample video data. Incorporating the different camera positions into the subsequent prediction of gaze coordinates facilitates the normalization of the data from the acquisition device and eliminates errors caused by the camera position.
[0081] Send a command to the sample object to correct the face pose, so that the sample object corrects the face pose according to the preset facial and head poses.
[0082] After the user selects the camera position, a command to correct the facial pose is sent to the sample object. This command causes the sample object to correct its facial and head poses according to preset parameters. These preset facial and head poses are pre-defined to facilitate data collection. Specifically, it may require the sample object to keep its face directly facing the screen of the smart device and its face within the green circular frame generated by the head pose correction algorithm, with both eyes looking directly at the screen. In one example, during facial pose correction, certain preset requirements are also made for the stability of the light source and the smart device. Specifically, the sample object is required to keep its face directly facing the light source, and the stability of the smart device must remain within a certain preset range during the acquisition of sample video data. In another example, if there are any deviations in the pose, light source, or stability of the smart device during the correction process, the smart device will issue a reminder, such as issuing a prompt tone like "Please adjust your head pose."
[0083] The preset boundary points are displayed on the screen of the smart device, so that the sample objects gaze at the preset boundary points in a preset order, and images of the sample objects gazing at the preset boundary points are captured.
[0084] After correcting the facial pose, data collection can begin. First, preset boundary points are displayed. Specifically, these are the boundary points of the smart device's screen. In one example, three dots (a total of 24 dots) could appear sequentially at the top left, top center, top right, center left, center right, bottom left, bottom right, and bottom right corners of the screen. The sample subject is then asked to look at and click on these dots. Each time the sample subject clicks a dot, the front-facing camera captures an image of the sample subject, collecting a total of 24 images.
[0085] The preset information is displayed on the screen of the smart device, so that the sample object looks at the preset information in the order of its display, and video recordings of the sample object looking at the preset information and screen recordings of the preset information being displayed on the screen of the smart device are captured.
[0086] After the image data corresponding to the preset boundary points is collected, the video data corresponding to the preset information is collected. At this time, the front-facing camera starts recording, and the smart device starts screen recording. Specifically, the start timestamp of the screen recording is also recorded. Multiple preset information items are displayed to the sample object in sequence, causing the sample object to look at each preset information item in turn. In one example, to ensure that the sample object's gaze point is stably on the preset information, the sample object is prompted to read the preset information aloud. Specifically, the initial color of the preset information is one color. When the sample object is prompted to look at a certain preset information item, the color of the preset information item is changed to the prompt color. After detecting that the sample object is reading aloud, it is verified whether the information read by the sample object matches the preset information. If they match, the user is prompted to start looking at the next preset information item.
[0087] In one example, the preset information is idioms. The smart device screen can display 60 idioms in 20 rows and 3 columns simultaneously, initially in black. After data collection begins, an idiom is randomly selected, its color changes to red, and a prompt sound instructs the user to look at and read it aloud. Specifically, the user's voice data of reading the idioms is verified using Automatic Speech Recognition (ASR). If the verification fails (the user's reading does not match the preset information), the user is prompted to continue reading the idiom. If the verification passes (the user's reading matches the preset information), the next idiom is randomly selected, prompting the user to look at and read it aloud. ASR supervision ensures that the user's gaze point matches the selected idiom, guaranteeing data accuracy. When the 30th idiom read by the user passes ASR verification, both video recording and screen recording stop simultaneously, indicating to the user that data collection is complete.
[0088] In one example, the dynamic eye-tracking dataset constructed in this embodiment is defined as dataset DGazeCap. The gaze prediction model provided in this disclosure, after being trained, is defined as model HGMSGaze (HGaze for short). This model achieves a prediction accuracy of 1.01 on the DGazeCap dataset, which is approximately 62% higher than existing technologies. Figure 4c As shown. Figure 4c AFF-Net and iTracker are both existing gaze prediction models, while GazeCapture and MPIIFaceGaze are existing eye-tracking datasets. The data represent the prediction accuracy of the two existing gaze prediction models on the two existing eye-tracking datasets and the DgazeCap dataset provided in this disclosure, respectively, and the prediction accuracy of the HGMSGaze model provided in this disclosure on the two existing eye-tracking datasets and the DgazeCap dataset provided in this disclosure. Specifically, the smaller the data value, the smaller the difference between the predicted result and the true value, i.e., the higher the prediction accuracy. This demonstrates the correctness of the dataset construction process and gaze prediction model proposed in this disclosure.
[0089] This constructed dynamic eye-tracking dataset contains richer facial information corresponding to gaze points, which can fully support the training of gaze point prediction models. It can also collect information on eye movement between different targets in the video. Eye movement trajectories are also of great significance for gaze point prediction model research. Furthermore, the data collection for this dataset is not limited by the model of smart device; theoretically, it can be collected from any smart device. The diversity of collection terminals and the resulting data diversity are significantly improved compared to previous datasets.
[0090] Step S42: Perform partial cropping on each of the sample boundary image data to obtain the second sample image group data corresponding to each of the sample boundary image data.
[0091] Each of the second sample image groups includes the left eye image, right eye image, facial image, and overall portrait of the sample object;
[0092] Step S43: Input the data of each second sample image group into the multilayer neural network to obtain the features of each boundary point;
[0093] Step S44: Input the features of each boundary point into the multilayer perceptron network to obtain the predicted boundary point coordinates.
[0094] Step S45: Adjust the parameters of the gaze point prediction model based on the distance between the predicted boundary point coordinates and the true boundary point coordinates.
[0095] The sample boundary image data is also partially cropped to obtain a second sample image group for each sample boundary image data. This second sample image group is then input into a multilayer neural network (specifically, a fully connected neural network) to obtain the features of each boundary point. These features are then input into a multilayer perceptron network to obtain the predicted boundary point coordinates. The loss and error of the gaze prediction model are calculated based on the distance between the predicted and ground truth boundary point coordinates, and the parameters of the gaze prediction model are adjusted accordingly. This process is repeated using the sample boundary image data until the distance between the predicted and ground truth boundary point coordinates output by the gaze prediction model meets the requirements, at which point the gaze prediction model training is considered complete.
[0096] In one example, the aforementioned gaze point features and boundary point features can be input into the multilayer perceptron network in parallel to predict coordinates. Based on this, the loss of the gaze point prediction model and the adjustment of parameters are also performed in parallel.
[0097] As can be seen from the above, the training method for the gaze prediction model provided in this disclosure, in addition to using the predicted gaze coordinates to adjust the gaze prediction model, also uses pre-acquired image data to predict the boundary point coordinates to assist in predicting the gaze prediction model. By using two methods, dual supervision of the gaze prediction model is achieved, which improves the training efficiency and accuracy of the gaze prediction model.
[0098] In one possible implementation, such as Figure 5 As shown, the gaze hypergraph learning model includes a hypergraph convolutional layer. Step S32 above inputs the features of each sample image into the gaze hypergraph learning model to construct the gaze hypergraph, and obtains the gaze features based on the gaze hypergraph, including:
[0099] Step S51: Input the features of each sample image into the gaze hypergraph learning model to obtain each initial node of the gaze hypergraph;
[0100] Step S52: For each initial node, use the K-nearest neighbor algorithm to determine a preset number of neighboring target nodes in the high-dimensional space;
[0101] Step S53: Establish hyperedges based on each target node to obtain the gaze point hypergraph;
[0102] Step S54: Update and iterate each node in the gaze hypergraph using the hypergraph convolutional layer to obtain each gaze feature.
[0103] The features of each sample image are input into the gaze hypergraph learning model as the initial nodes of the hypergraph. Then, multiple neighboring target nodes are determined in a high-dimensional space; in one example, the preset number of target nodes could be 10. Based on this, hyperedges are constructed to obtain the gaze hypergraph, which can represent the correlation between the features of each sample image. Then, the features of each node in the gaze hypergraph are updated iteratively using hypergraph convolutional layers. That is, the high-order features of nodes connected by the same hyperedge are first aggregated onto the corresponding hyperedge, and then the hyperedge distributes the obtained high-order features to each connected node to complete the feature update, finally obtaining the gaze features.
[0104] As can be seen from the above, the training method of the gaze prediction model provided in this disclosure uses the features of each sample image as the initial nodes of the gaze hypergraph learning model to construct the gaze hypergraph. This allows the gaze hypergraph to effectively reflect the correlation between the features of each sample image, and uses the edges between nodes to transmit and share information, thereby achieving the effect of feature enhancement and improving the accuracy of gaze prediction.
[0105] See Figure 6 This disclosure also provides a flowchart of a fixation prediction method, including:
[0106] Step S61: Obtain target image data, wherein the target image data includes the target object of the gaze point to be predicted;
[0107] Step S62: Input the target image data into a pre-trained gaze prediction model to obtain the predicted gaze coordinates of the target object, wherein the gaze prediction model is trained by any of the gaze prediction model training methods described above.
[0108] As can be seen from the above, the gaze prediction method provided in this disclosure, which uses the gaze prediction model trained by any of the above-described gaze prediction models to predict target image data, can accurately and efficiently obtain the predicted gaze of the target object, providing the possibility for subsequent operations.
[0109] See Figure 7 This disclosure also provides a schematic diagram of the structure of a training device for a gaze prediction model, including:
[0110] The first data acquisition module 701 is used to acquire multiple sets of first sample image data and gaze point ground truth coordinates based on a preset dynamic eye tracking dataset. Each set of first sample image data is a set of image data captured frame by frame from sample video data collected when the sample object gazes based on preset information.
[0111] The first coordinate prediction module 702 is used to input the data of each of the first sample image groups into the gaze prediction model, construct a gaze hypergraph using the gaze hypergraph learning model in the gaze prediction model, and obtain the predicted gaze coordinates based on the gaze hypergraph.
[0112] The first parameter adjustment module 703 is used to adjust the parameters of the gaze point prediction model based on the distance between the predicted gaze point coordinates and the true gaze point coordinates.
[0113] As can be seen from the above, the training device for the gaze prediction model provided in this disclosure acquires multiple sets of first sample image data and gaze point ground truth coordinates based on a preset dynamic eye tracking dataset. Each set of first sample image data is a set of image data extracted frame by frame from sample video data collected when the sample object gazes at preset information. Therefore, each set of first sample image data is continuous in time. Based on such sample image data, the prediction of gaze point coordinates can more accurately capture the movement trajectory of the sample object's eyes.
[0114] Secondly, the data from each first sample image group are input into the gaze prediction model. A gaze hypergraph is constructed using the gaze hypergraph learning model within the model. Based on this hypergraph, the relationships between the first sample image groups can be determined. Since each sample image group is obtained frame-by-frame from the sample video data, determining the predicted gaze coordinates based on this hypergraph effectively considers the temporal and spatial relationships between frames in the sample video data, thus modeling higher-order temporal relationships between frames. The resulting predicted gaze coordinates have higher accuracy. Adjusting the parameters of the gaze prediction model based on the distance between the more accurate predicted gaze coordinates and the ground truth gaze coordinates improves the training efficiency of the model. The trained gaze prediction model effectively improves both the efficiency and accuracy of gaze prediction.
[0115] In one embodiment of this disclosure, the first data acquisition module 701 is specifically used for:
[0116] Based on a preset dynamic eye-tracking dataset, first sample video data and second sample video data are obtained, wherein the first sample video data is a video recording taken when the sample object gazes at the preset information, and the second sample video data is a screen recording taken when the preset information is displayed on a smart device.
[0117] The true coordinates of the gaze point of the sample object are obtained based on the second sample video data;
[0118] Each frame of the first sample video data is partially cropped to obtain the first sample image group data corresponding to each frame. Each first sample image group data includes the left eye image, right eye image, facial image, and overall portrait of the sample object.
[0119] As can be seen from the above, the training device for the gaze prediction model provided in this disclosure includes a preset dynamic eye tracking dataset containing video recordings of sample objects gazing at preset information and screen recordings of the preset information displayed on a smart device. These videos are segmented according to different preset information, so the gaze points of each segmented video segment are consistent. Each frame of the segmented video segment is extracted into the left eye image, right eye image, facial image, and overall portrait of the sample object. The amount of data in the extracted image group data is very large, which can cover as many gaze point features of the sample object as possible. Based on this, the gaze prediction model can be trained to achieve the highest possible accuracy in gaze prediction.
[0120] In one embodiment of this disclosure, the gaze prediction model further includes: a feature extraction network and a multilayer perceptron network;
[0121] The first coordinate prediction module 702 includes:
[0122] The feature extraction submodule is used to input the data of each of the first sample image groups into the feature extraction network to obtain the sample image features of each of the sample image groups.
[0123] The feature acquisition submodule is used to input the features of each sample image into the gaze hypergraph learning model to construct the gaze hypergraph, and to obtain the gaze features based on the gaze hypergraph;
[0124] The coordinate prediction submodule is used to input the features of each gaze point into the multilayer perceptron network to obtain the predicted gaze point coordinates.
[0125] As can be seen from the above, the training device for the gaze prediction model provided in this disclosure includes a feature extraction network, a gaze hypergraph learning model, and a multilayer perceptron network. It utilizes a multi-structured model to decompose and analyze a large number of sample image data sets. At the same time, it uses the gaze hypergraph learning model to establish the correlation between images in each sample image data set, thereby realizing the construction of temporal and spatial correlations between frames and determining the higher-order temporal relationships between frame images. The gaze prediction performed in this way can effectively improve the accuracy.
[0126] In one embodiment of this disclosure, the gaze point prediction model further includes a multi-layer neural network, and the apparatus further includes:
[0127] The second data acquisition module is used to acquire sample boundary image data and ground truth coordinates of boundary points based on a preset dynamic eye-tracking dataset, wherein the sample boundary image data includes images captured when the sample object gazes at a preset boundary point;
[0128] The image cropping module is used to partially crop each of the sample boundary image data to obtain a second sample image group data corresponding to each of the sample boundary image data, wherein each second sample image group data includes the left eye image, right eye image, facial image and overall portrait of the sample object.
[0129] The feature acquisition module is used to input the data of each second sample image group into the multilayer neural network to obtain the features of each boundary point;
[0130] The second coordinate prediction module is used to input the features of each boundary point into the multilayer perceptron network to obtain the predicted boundary point coordinates.
[0131] The second parameter adjustment module is used to adjust the parameters of the gaze point prediction model based on the distance between the predicted boundary point coordinates and the true boundary point coordinates.
[0132] As can be seen from the above, the training device for the gaze prediction model provided in this disclosure, in addition to using the predicted gaze coordinates to adjust the gaze prediction model, also uses pre-acquired image data to predict the boundary point coordinates to assist in predicting the gaze prediction model. By using two methods, dual supervision of the gaze prediction model is achieved, which improves the training efficiency and accuracy of the gaze prediction model.
[0133] In one embodiment of this disclosure, the preset dynamic eye-tracking dataset is constructed as follows:
[0134] Send a command to the sample object indicating the selection of a camera position, so that the sample object selects a camera position on the screen of the smart device;
[0135] Send a command to the sample object indicating that the face pose should be corrected, so that the sample object corrects its face pose according to a preset face pose and head pose;
[0136] The preset boundary points are displayed on the screen of the smart device, so that the sample object looks at the preset boundary points in a preset order, and images of the sample object looking at the preset boundary points are captured.
[0137] The preset information is displayed on the screen of the smart device, so that the sample object looks at the preset information in the order of its display, and video recordings of the sample object looking at the preset information and screen recordings of the preset information being displayed on the screen of the smart device are captured.
[0138] In one embodiment of this disclosure, the feature acquisition submodule is specifically used for:
[0139] The features of each sample image are input into the gaze hypergraph learning model to obtain each initial node of the gaze hypergraph;
[0140] For each initial node, the K-nearest neighbor algorithm is used to determine a predetermined number of neighboring target nodes in the high-dimensional space;
[0141] Hyperedges are established based on each of the target nodes to obtain the gaze point hypergraph;
[0142] The nodes in the gaze hypergraph are updated and iterated using the hypergraph convolutional layer to obtain the gaze features of each gaze point.
[0143] As can be seen from the above, the training device for the gaze prediction model provided in this disclosure uses the features of each sample image as the initial nodes of the gaze hypergraph learning model to construct the gaze hypergraph. This allows the gaze hypergraph to effectively reflect the correlation between the features of each sample image, and uses the edges between nodes to transmit and share information, thereby achieving the effect of feature enhancement and improving the accuracy of gaze prediction.
[0144] See Figure 8 This disclosure also provides a schematic diagram of a gaze point prediction device, including:
[0145] Image acquisition module 801 is used to acquire target image data, wherein the target image data includes the target object of the gaze point to be predicted;
[0146] The gaze point coordinate prediction module 802 is used to input the target image data into a pre-trained gaze point prediction model to obtain the predicted gaze point coordinates of the target object, wherein the gaze point prediction model is trained by any of the gaze point prediction model training devices described above.
[0147] As can be seen from the above, the gaze prediction device provided in this disclosure can accurately and efficiently obtain the predicted gaze point of the target object by using the gaze prediction model trained by any of the above-described gaze prediction model training methods to predict the target image data, thus providing the possibility for subsequent operations.
[0148] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0149] It should be noted that the two-dimensional face images in this embodiment are from a publicly available dataset.
[0150] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0151] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0152] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0153] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0154] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the methods for training and predicting gaze points. For example, in some embodiments, the methods for training and predicting gaze points can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the methods for training and predicting gaze points described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured in any other suitable manner (e.g., by means of firmware) to perform a training method for the gaze prediction model and a gaze prediction method.
[0155] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0156] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0157] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0158] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0159] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0160] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0161] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0162] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for training a fixation prediction model, comprising: Based on a preset dynamic eye-tracking dataset, multiple sets of first sample image data and gaze point ground truth coordinates are obtained, as well as sample boundary image data and boundary point ground truth coordinates. Each set of first sample image data is an image group data obtained frame by frame from each set of first sample video data collected when the sample object gazes at preset information. Each set of first sample image data includes the left eye image, right eye image, facial image, and overall portrait of the sample object. The gaze point ground truth coordinates are the coordinates of the preset information. The sample boundary image data includes the image collected when the sample object gazes at the preset boundary point. For each of the sample boundary image data, a portion is cropped to obtain the second sample image group data corresponding to each of the sample boundary image data. Each of the second sample image group data includes the left eye image, right eye image, facial image and overall portrait of the sample object. The data of each of the first sample image groups are input into the gaze prediction model. A gaze hypergraph is constructed using the gaze hypergraph learning model in the gaze prediction model. Based on the gaze hypergraph and the multilayer perceptron network of the gaze prediction model, the predicted gaze coordinates are obtained. The data of each second sample image group are input into the multi-layer neural network of the gaze point prediction model to obtain the features of each boundary point; The features of each boundary point are input into the multilayer perceptron network of the gaze point prediction model to obtain the predicted boundary point coordinates. The parameters of the gaze prediction model are adjusted based on the distance between the predicted gaze coordinates and the ground truth gaze coordinates, and based on the distance between the predicted boundary point coordinates and the ground truth boundary point coordinates.
2. The method according to claim 1, wherein, The method, based on a preset dynamic eye-tracking dataset, acquires multiple sets of first sample image data and ground-value coordinates of the gaze point, including: Based on a preset dynamic eye-tracking dataset, first sample video data and second sample video data are obtained, wherein the first sample video data is a video recording taken when the sample object gazes at the preset information, and the second sample video data is a screen recording taken when the preset information is displayed on a smart device. The true coordinates of the gaze point of the sample object are obtained based on the second sample video data; Each frame of the first sample video data is partially cropped to obtain the first sample image group data corresponding to each frame.
3. The method according to claim 1, wherein, The gaze prediction model also includes: a feature extraction network and a multilayer perceptron network; The step of inputting the data of each of the first sample image groups into the gaze prediction model, constructing a gaze hypergraph using the gaze hypergraph learning model in the gaze prediction model, and obtaining the predicted gaze coordinates based on the gaze hypergraph includes: Each of the first sample image groups is input into the feature extraction network to obtain the sample image features of each sample image group. The features of each sample image are input into the gaze hypergraph learning model to construct the gaze hypergraph, and the features of each gaze point are obtained based on the gaze hypergraph. The gaze point features are input into the multilayer perceptron network to obtain the predicted gaze point coordinates.
4. The method according to any one of claims 1-3, wherein, The preset dynamic eye-tracking dataset is constructed as follows: Send a command to the sample object indicating the selection of a camera position, so that the sample object selects the camera position on the screen of the smart device; Send a command to the sample object indicating that the face pose should be corrected, so that the sample object corrects its face pose according to a preset face pose and head pose; The preset boundary points are displayed on the screen of the smart device, so that the sample object looks at the preset boundary points in a preset order, and images of the sample object looking at the preset boundary points are captured. The preset information is displayed on the screen of the smart device, so that the sample object looks at the preset information in the order of its display, and video recordings of the sample object looking at the preset information and screen recordings of the preset information being displayed on the screen of the smart device are captured.
5. The method according to claim 3, wherein, The gaze hypergraph learning model includes a hypergraph convolutional layer. The step of inputting the features of each sample image into the gaze hypergraph learning model to construct the gaze hypergraph, and obtaining each gaze feature based on the gaze hypergraph, includes: The features of each sample image are input into the gaze hypergraph learning model to obtain each initial node of the gaze hypergraph; For each initial node, the K-nearest neighbor algorithm is used to determine a predetermined number of neighboring target nodes in the high-dimensional space; Hyperedges are established based on each of the target nodes to obtain the gaze point hypergraph; The nodes in the gaze hypergraph are updated and iterated using the hypergraph convolutional layer to obtain the gaze features of each gaze point.
6. A fixation prediction method, comprising: Acquire target image data, wherein the target image data includes the target object for which the gaze point to be predicted is located; The target image data is input into a pre-trained gaze prediction model to obtain the predicted gaze coordinates of the target object, wherein the gaze prediction model is trained by the gaze prediction model training method of any one of claims 1-5.
7. A training device for a fixation prediction model, comprising: The first data acquisition module is used to acquire multiple sets of first sample image data and gaze point ground truth coordinates based on a preset dynamic eye tracking dataset, and to acquire sample boundary image data and boundary point ground truth coordinates. Each set of first sample image data is an image group data obtained frame by frame from each set of first sample video data collected when the sample object gazes at preset information, including the left eye image, right eye image, facial image, and overall image of the sample object. The gaze point ground truth coordinates are the coordinates of the preset information. The sample boundary image data includes the image collected when the sample object gazes at the preset boundary point. The image cropping module is used to partially crop each of the sample boundary image data to obtain a second sample image group data corresponding to each of the sample boundary image data, wherein each second sample image group data includes the left eye image, right eye image, facial image and overall portrait of the sample object. The first coordinate prediction module is used to input the data of each of the first sample image groups into the gaze prediction model, construct a gaze hypergraph using the gaze hypergraph learning model in the gaze prediction model, and obtain the predicted gaze coordinates based on the gaze hypergraph and the multilayer perceptron network of the gaze prediction model. The feature acquisition module is used to input the data of each second sample image group into a multi-layer neural network to obtain the features of each boundary point; The second coordinate prediction module is used to input the features of each boundary point into the multilayer perceptron network to obtain the predicted boundary point coordinates. The first parameter adjustment module is used to adjust the parameters of the gaze point prediction model based on the distance between the predicted gaze point coordinates and the ground truth gaze point coordinates, and based on the distance between the predicted boundary point coordinates and the ground truth boundary point coordinates.
8. The apparatus according to claim 7, wherein, The first data acquisition module is specifically used for: Based on a preset dynamic eye-tracking dataset, first sample video data and second sample video data are obtained, wherein the first sample video data is a video recording taken when the sample object gazes at the preset information, and the second sample video data is a screen recording taken when the preset information is displayed on a smart device. The true coordinates of the gaze point of the sample object are obtained based on the second sample video data; Each frame of the first sample video data is partially cropped to obtain the first sample image group data corresponding to each frame.
9. The apparatus according to claim 7, wherein, The gaze prediction model also includes: a feature extraction network and a multilayer perceptron network; The first coordinate prediction module includes: The feature extraction submodule is used to input the data of each of the first sample image groups into the feature extraction network to obtain the sample image features of each of the sample image groups. The feature acquisition submodule is used to input the features of each sample image into the gaze hypergraph learning model to construct the gaze hypergraph, and to obtain the gaze features based on the gaze hypergraph; The coordinate prediction submodule is used to input the features of each gaze point into the multilayer perceptron network to obtain the predicted gaze point coordinates.
10. The apparatus according to any one of claims 7-9, wherein, The preset dynamic eye-tracking dataset is constructed as follows: Send a command to the sample object indicating the selection of a camera position, so that the sample object selects the camera position on the screen of the smart device; Send a command to the sample object indicating that the face pose should be corrected, so that the sample object corrects its face pose according to a preset face pose and head pose; The preset boundary points are displayed on the screen of the smart device, so that the sample object looks at the preset boundary points in a preset order, and images of the sample object looking at the preset boundary points are captured. The preset information is displayed on the screen of the smart device, so that the sample object looks at the preset information in the order of its display, and video recordings of the sample object looking at the preset information and screen recordings of the preset information being displayed on the screen of the smart device are captured.
11. The apparatus according to claim 9, wherein, The feature acquisition submodule is specifically used for: The features of each sample image are input into the gaze hypergraph learning model to obtain each initial node of the gaze hypergraph; For each initial node, the K-nearest neighbor algorithm is used to determine a predetermined number of neighboring target nodes in the high-dimensional space; Hyperedges are established based on each of the target nodes to obtain the gaze point hypergraph; The nodes in the gaze hypergraph are updated iteratively using a hypergraph convolutional layer to obtain the gaze features of each gaze point.
12. A fixation point prediction device, comprising: An image acquisition module is used to acquire target image data, wherein the target image data includes a target object for which the gaze point to be predicted is located; A gaze point coordinate prediction module is used to input the target image data into a pre-trained gaze point prediction model to obtain the predicted gaze point coordinates of the target object, wherein the gaze point prediction model is trained by the training device of the gaze point prediction model according to any one of claims 7-11.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Eye movement interaction method and device based on head time sequence signal correction
CN113419624A