A method and device for estimating consumer focus points based on AI

By combining head detection and gaze constraints, the difficulty of traditional gaze estimation methods in inferring gaze in non-frontal situations is solved, and the accuracy and robustness of gaze estimation are improved when the face is deviated or occluded.

CN120318864BActive Publication Date: 2025-09-19SHENZHEN AIMALL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510793476.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-19
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Traditional gaze estimation methods have difficulty accurately inferring the gaze target when the face is not facing forward. Especially when the face is away from the camera or severely occluded, the face detection algorithm may fail, resulting in the interruption of the gaze estimation process.

Method used

A head detection algorithm is used to obtain head posture information, and the gaze target is predicted based on the head posture information. Gaze constraints are introduced to screen gaze targets that meet the orientation. The head posture information is used to simulate the natural human visual range and filter out incorrect predictions.

Benefits of technology

When the face is not facing forward or is blocked, it can accurately provide gaze direction cues, improve the accuracy and robustness of gaze estimation, and adapt to multiple people and complex background scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318864B_ABST
    Figure CN120318864B_ABST
Patent Text Reader

Abstract

The present invention discloses an AI-based method and device for estimating consumer focus points, which relates to the technical field of computer vision and solves the technical problem that traditional gaze estimation methods are difficult to accurately infer gaze targets when the face is not facing forward. The method includes performing head detection on an input image using a head detection algorithm to obtain head posture information; adding the head posture information to a feature map of the input image to obtain a head condition feature map; predicting gaze targets corresponding to all heads in the input image based on the head condition feature map; and using the head posture information to filter each predicted gaze target, retaining the gaze targets that meet the gaze constraint conditions. The present invention can also accurately infer gaze targets when the face is not facing forward.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to an AI-based consumer focus estimation method and device. Background Art

[0002] Gaze estimation is a key technology that analyzes the direction of a person's visual attention to infer the target of attention. It has widespread applications in fields such as human-computer interaction, smart retail, driver monitoring, and virtual reality. Traditional gaze estimation methods rely primarily on facial feature detection, particularly the location of key points in the eye region (such as the pupil center and eye corners), combined with geometric models or machine learning algorithms to predict gaze direction. These methods typically require the subject's face to face the camera, ensuring that facial features (such as the eyes) are clearly visible and symmetrical.

[0003] However, in real-world applications (such as analyzing customer behavior in supermarkets and multi-person interactions in conferences), faces often present non-frontal orientations (e.g., sideways, looking down, or turned), are partially or completely occluded (e.g., by wearing a mask, glare from glasses, or hand occlusion), and encounter complex background interference. These issues lead to the following problems with traditional gaze estimation methods: when the face is facing away from the camera or is severely occluded, the face detection algorithm may fail, interrupting the gaze estimation process. While existing methods (such as Gaze-LLE) can extract features of the facial region through face detection, they ignore auxiliary information such as head pose, body orientation, and scene context, making it difficult to accurately infer the gaze target in non-frontal situations.

[0004] In the process of implementing the present invention, the inventors discovered that the prior art has at least the following problems:

[0005] Traditional gaze estimation methods have difficulty in accurately inferring gaze targets when the face is not frontal. Summary of the Invention

[0006] The purpose of the present invention is to provide an AI-based consumer focus point estimation method and device to solve the technical problem in the prior art that traditional gaze estimation methods are difficult to accurately infer the gaze target when the face is not frontal.

[0007] The various technical effects that can be produced by the preferred technical solutions among the various technical solutions provided by the present invention are described in detail below.

[0008] To achieve the above objectives, the present invention provides the following technical solutions:

[0009] The present invention provides an AI-based method for estimating consumer focus points, comprising the following steps: performing head detection on an input image using a head detection algorithm to obtain head posture information; adding the head posture information to a feature map of the input image to obtain a head condition feature map; predicting gaze targets corresponding to all heads in the input image based on the head condition feature map; and using the head posture information to screen each predicted gaze target, retaining the gaze targets that meet gaze constraint conditions.

[0010] Optionally, the head detection is performed on the input image through a head detection algorithm to obtain head posture information, including: generating a head bounding box of the head through a pre-trained multi-person head posture estimation model; obtaining the three-dimensional posture angle of the head through the multi-person head posture estimation model; and calculating the frontal face orientation vector of the head based on the three-dimensional posture angle; wherein the head posture information includes the head bounding box and the frontal face orientation vector.

[0011] Optionally, calculate the face orientation vector for:

[0012] ;

[0013] ;

[0014]

[0015] in, is the three-dimensional posture angle; The rotation matrix of the three-dimensional posture angle; for The rotation matrix of for The rotation matrix of for The rotation matrix of .

[0016] Optionally, adding the head posture information to the feature map of the input image to obtain a head conditional feature map includes: using a visual feature extractor to extract features from the input image to generate the feature map; generating a head position embedding by converting the head bounding box into normalized coordinates; converting the head bounding box into a downsampled binary mask and aligning it with the feature map in the spatial dimension; and adding the head position embedding element-by-element to the part of the binary mask corresponding to the head position to generate the head conditional feature map.

[0017] Optionally, predicting the gaze target corresponding to each head in the input image based on the head condition feature map includes: predicting the position of the gaze target corresponding to each head in the input image based on the head condition feature map through a gaze heat map decoder to obtain a gaze heat map of each gaze target; wherein the gaze heat map represents the probability of being the gaze target in the input image.

[0018] Optionally, a binary cross entropy loss function is used to train and optimize the gaze heat map; the binary cross entropy loss function for:

[0019]

[0020] ;

[0021] in, is the pixel-level binary cross entropy loss; is the height of the gaze heat map; is the width of the gaze heat map; is the index of the gaze heat map, representing the two-dimensional coordinate of the gaze heat map, i represents the row, and j represents the column; For a real gaze heat map; is the predicted gaze heat map; is the regularization parameter; is a loss function for training whether the gaze target is within the input image.

[0022] Optionally, in the step of using the head posture information to filter each predicted gaze target and retaining the gaze targets that meet the gaze constraint condition, the gaze constraint condition is: calculating the direction vector of the line connecting the center position of the head in the input image and the gaze target, and then calculating the angle between the direction vector and the frontal face direction vector; if the angle is less than a gaze threshold, determining that the gaze target with the angle less than the gaze threshold meets the gaze constraint condition.

[0023] Optionally, after screening each of the predicted gaze targets using the head posture information and retaining the gaze targets that meet the gaze constraint condition, the method further includes:

[0024] A target tracking algorithm is adopted, with the head bounding box of the first frame image in the video to be estimated as the initial tracking target, the head bounding box of each frame image in the video to be estimated is tracked, and the gaze trajectory corresponding to the head in each head bounding box is recorded; wherein, each frame image in the video to be estimated is the input image.

[0025] A device for estimating consumer focus points based on AI comprises: a head detection module, configured to perform head detection on an input image using a head detection algorithm to obtain head posture information; a feature extraction module, configured to add the head posture information to a feature map of the input image to obtain a head condition feature map; a gaze estimation module, configured to predict the gaze target corresponding to each head in the input image based on the head condition feature map; and a gaze constraint module, configured to use the head posture information to filter each predicted gaze target and retain the gaze targets that meet the gaze constraint conditions.

[0026] Optionally, the device also includes a target tracking module, which is used to adopt a target tracking algorithm, take the head bounding box of the first frame image in the video to be estimated as the initial tracking target, track the head bounding box of each frame image in the video to be estimated, and record the gaze trajectory corresponding to the head in each head bounding box; wherein, each frame image in the video to be estimated is the input image.

[0027] Implementing one of the above technical solutions of the present invention has the following advantages or beneficial effects:

[0028] The present invention uses head detection to obtain the posture information of the human head in the image. It does not rely on the detection of the frontal face and can estimate the line of sight direction from the human head posture. Therefore, even when the face is not facing the camera (such as the side face or the back of the head is photographed by the camera) or is facing the camera or there is occlusion of the face, it can also accurately provide position prompt information to improve the accuracy of line of sight estimation. It also introduces gaze constraints and uses head posture information to simulate the natural visual range of humans, filter out erroneous predictions that do not conform to the direction, and can significantly improve robustness in scenes with multiple people, occlusion or complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work. In the drawings:

[0030] Figure 1 This is a flowchart of a method for estimating consumer focus points based on AI according to a first embodiment of the present invention;

[0031] Figure 2 This is a flowchart of step S1 of the AI-based consumer focus estimation method according to the first embodiment of the present invention;

[0032] Figure 3This is a flowchart of step S2 of the AI-based consumer focus estimation method according to the first embodiment of the present invention;

[0033] Figure 4 This is a structural block diagram of an AI-based consumer focus estimation device according to a second embodiment of the present invention. DETAILED DESCRIPTION

[0034] In order to make the objects, technical solutions and advantages of the present invention clearer, the various exemplary embodiments to be described below will refer to the corresponding drawings, which constitute a part of the exemplary embodiments, in which various exemplary embodiments that may be used to implement the present invention are described. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation methods described in the following exemplary embodiments do not represent all implementation methods consistent with the present disclosure. It should be understood that they are only examples of processes, methods and devices that are consistent with some aspects of the present disclosure as detailed in the appended claims, and other embodiments may also be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and essence of the present invention.

[0035] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", etc. indicate the orientation or position relationship based on the drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the elements referred to must have a specific orientation, be constructed and operate in a specific orientation. The terms "first", "second", etc. are only used for descriptive purposes and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. The term "plurality" means two or more. The terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.

[0036] In order to illustrate the technical solution of the present invention, a specific embodiment is provided below, in which only the parts related to the embodiment of the present invention are shown.

[0037] Example 1:

[0038] like Figure 1As shown, the present invention provides an AI-based method for estimating consumer focus points, comprising the following steps: S1. Detecting the head of an input image using a head detection algorithm to obtain head posture information; S2. Adding the head posture information to a feature map of the input image to obtain a head conditional feature map; S3. Predicting gaze targets corresponding to all heads in the input image based on the head conditional feature map; S4. Using the head posture information, filtering each predicted gaze target and retaining those that meet the gaze constraint conditions. It should be noted that a consumer's focus point refers to the object seen in the direction of the consumer's line of sight, or the object in the direction of the consumer's gaze or attention.

[0039] This embodiment uses head detection to obtain the posture information of the human head in the image. It does not rely on the detection of the frontal face and can estimate the gaze direction from the human head posture. Therefore, even when the face is not facing the camera (such as the camera captures the side face or the back of the head), or is facing the camera or there is occlusion of the face, it can also accurately provide position prompt information to improve the accuracy of gaze estimation. It also introduces gaze constraints, uses head posture information to simulate the natural visual range of humans, and filters out erroneous predictions that do not conform to the direction. In scenes with multiple people, occlusions, or complex backgrounds, it can significantly improve robustness.

[0040] Next, combine Figures 1 to 3 The specific implementation steps of the AI-based consumer focus estimation method provided in this embodiment are described in detail:

[0041] First, step S1 is performed to detect the head of the input image using a head detection algorithm to obtain head pose information. Using a head detection algorithm to obtain head pose information from an image does not rely on images of a person's face facing the camera directly. Even images of a person's head from the side or back can accurately estimate the gaze target.

[0042] Specifically, such as Figure 2As shown, step S1 includes: S11, generating a head bounding box of the head in the input image through a pre-trained multi-person head pose estimation model; compared with face detection, head detection is more robust to changes in face orientation, which helps to achieve accurate prediction of the line of sight of non-frontal faces. S12, obtaining the three-dimensional posture angle of the head through the multi-person head pose estimation model; for example, using the Euler angle of the head as the three-dimensional posture angle. S13, calculating the frontal face orientation vector of the head based on the three-dimensional posture angle; wherein the head pose information includes the head bounding box and the frontal face orientation vector. The obtained head pose information provides accurate position information for subsequent line of sight estimation. The multi-person head pose estimation model of this embodiment adopts the DirectMHP model, which is an end-to-end multi-person head pose estimation model that breaks through the viewing angle limitation of traditional head pose estimation and provides an implementation basis for multi-person interaction in complex scenes. The DirectMHP model can be pre-trained using an existing data set and can be used directly after training. For example, the DirectMHP model can be trained using the AGORA dataset (Avatars in Geography Optimized for Regression Analysis, a synthetic dataset designed specifically for 3D human pose estimation tasks) and the CMU Panoptic dataset (a large-scale multimodal 3D pose dataset for multi-person social interaction scenarios, designed to provide high-precision annotation support for body, hand, and facial pose estimation in complex environments). Labels can also be added to the dataset, including head bounding boxes and Euler angles, to provide clear optimization directions for the model.

[0043] Furthermore, generating the head bounding box of the head in the input image specifically includes: the backbone network uses YOLOv5 as a feature extractor, processes the input image I to generate a multi-scale feature map F O , the feature scale is O∈{8,16,32,64}; then the backbone network outputs four prediction grids, each of which contains dense prediction results, which include the parameters of the head bounding box and the 3D pose angle. Assume that the output vector of each grid unit is ,in:

[0044] , represents the center of the bounding box, Represents the width and height of the bounding box, Indicates confidence; represents the 3D posture angle (in this embodiment, the Euler angle); Represents a vector and vector Perform splicing operations.

[0045] In each network unit of the multi-person head pose estimation model, the network unit predicts the relative position of the predefined anchor box. The coordinates of the head bounding box are calculated as follows:

[0046]

[0047]

[0048]

[0049]

[0050] in, is the sigmoid function; ( , ) are the coordinates of the upper left corner of the network unit, representing the x-axis coordinate and y-axis coordinate relative to the entire image respectively; and The offset of the center of the bounding box predicted by the network relative to the top left corner of the grid cell; and represents the logarithmic ratio of the width and height of the bounding box predicted by the network relative to the anchor box; the final coordinates are:

[0051]

[0052] .

[0053] Further, calculate the face orientation vector for:

[0054] ;

[0055] ;

[0056]

[0057] in, is the three-dimensional posture angle; The rotation matrix of the three-dimensional posture angle; for The rotation matrix of for The rotation matrix of for The rotation matrix of .

[0058] Then, step S2 is performed to add the head posture information to the feature map of the input image to obtain a head conditional feature map. Incorporating head posture information into the feature map can significantly improve the accuracy of gaze direction prediction.

[0059] Specifically, such as Figure 3 As shown, step S2 includes: S21, using a visual feature extractor to extract features from the input image and generate a feature map; for example, DINOv2 (DINOv2 is a self-supervised visual representation learning model developed by Meta AI Research (formerly Facebook AI Research), which aims to train highly versatile visual features through large-scale unlabeled image data, and directly apply to a variety of downstream tasks without fine-tuning) can be used as a visual feature extractor. S22, generating a head position embedding by converting the head bounding box into normalized coordinates; S23, converting the head bounding box into a downsampled binary mask and aligning it with the feature map in the spatial dimension; S24, adding the head position embedding element-by-element to the part of the binary mask corresponding to the head position to generate a head conditional feature map. Integrating the information of the head detection box into the feature map can enhance the model's perception of the head position and improve the accuracy of line of sight estimation.

[0060] For example, first generate the feature map: transform the RGB format image I∈R 3×H×W Input visual feature extractor, where RGB format is a common color image representation, consisting of three color channels: red (R), green (G), and blue (B). Image I∈R 3×H×W Indicates that image I belongs to a three-dimensional real space; 3 represents the RGB channels, and H and W represent the height and width of the image, respectively. For example, if we input an image with H=448 and W=448, it specifies that both the height and width of the input RGB image are 448 pixels.

[0061] Then the 448×448 image is divided into multiple 14×14 sized small regions (ie patches) to generate the feature map F∈R d×H×W . It is calculated that it can be divided into 448÷14=32 patches in the height direction and 448÷14=32 patches in the width direction, and there will be a total of 32×32=1024 14×14 patches. After the encoder processes the image, a feature map F∈R is generated. d×H×W The feature map F is also a three-dimensional tensor, where d represents the number of channels, and H and W are the height and width of the feature map. The feature map F has the same H and W values ​​as the input image, i.e., both height and width are 448 pixels. Each element in the feature map contains the characteristic information of the image at that location and channel, which can be used for subsequent line of sight estimation tasks.

[0062] Then, generate the head position embedding: obtain the coordinates of the head bounding box and divide them by the width and height of the image to obtain the normalized coordinates. For example, if the image width is W and the height is H, the coordinates of the upper left corner of the bounding box are (x1, y1), and the coordinates of the lower right corner are (x2, y2), then the normalized coordinates are (x1 / W, y1 / H, x2 / W, 2 / H). Use the normalized coordinates to generate a vector, which is the head position embedding. This can be achieved through an embedding layer (Embedding Layer). The embedding layer can be a simple fully connected layer that takes the normalized coordinates as input and outputs a fixed-length vector. This vector contains the head position information. By converting the head position information into an embedding vector, it can be better integrated with other features (such as facial features and posture features), thereby improving the performance of the model.

[0063] Subsequently, a linear projection layer can be used to reduce the feature dimension d to the model-specific dimension d m =256, generate compressed feature map x F ∈R 256×32×32 The weights of the scene encoder remain fixed during training and inference, leveraging robust, general features learned from large-scale datasets. Linear projection layers effectively reduce feature dimensionality while preserving the data's primary structure and information. This helps reduce the number of parameters in the model, speeds up training, and improves the model's generalization capabilities.

[0064] Then, the head bounding box is converted into a downsampled binary mask and aligned with the feature map F in the spatial dimension: box ∈R 4 (Assuming X box =[x min ,y min ,x max ,y max ]) is converted to a downsampled binary mask M∈{0,1} 32×32 , and then spatially aligns it with the generated feature map F. Downsampling refers to reducing the size of an image or data. Here, the head bounding box is converted into a 32×32 binary mask. A binary mask means that the elements in this matrix can only take on two values: 0 and 1. In this mask, the area corresponding to the head position has a value of 1, and the rest of the area has a value of 0. By generating a binary mask and performing feature alignment, the information in the feature map can be more effectively utilized, thereby improving model performance and accuracy.

[0065] Finally, generate the head conditional feature map: embed the head position (assuming the head position is embedded as p head ∈R 256 ) is added element-by-element to the scene tag of M=1 to generate the head conditional feature map S:

[0066]

[0067] in, represents an element-wise broadcast multiplication along the channel dimension. This late injection of head position information preserves the integrity of the frozen encoder features while making the output conditioned on a specific individual.

[0068] Then, step S3 is executed to predict the gaze targets corresponding to all heads in the input image based on the head condition feature map. Predicting the gaze targets based on the feature map with added head posture information can greatly improve the accuracy of the prediction.

[0069] Specifically, step S3 includes: using a gaze heatmap decoder to predict the location of the gaze target corresponding to each head in the input image based on the head conditional feature map, thereby obtaining a gaze heatmap for each gaze target; wherein the gaze heatmap represents the probability of being the gaze target in the input image. For example, a Lightweight Transformer (LWT) can be used as the gaze decoder. This model is suitable for processing large-scale datasets and can reduce model size and computational complexity.

[0070] Furthermore, the binary cross entropy loss function is used to train and optimize the gaze heat map; the binary cross entropy loss function for:

[0071]

[0072] ;

[0073] in, is the pixel-level binary cross entropy loss; is the height of the gaze heat map; is the width of the gaze heat map; is the index of the gaze heat map, representing the two-dimensional coordinate of the gaze heat map, i represents the row, and j represents the column; For a real gaze heat map; is the predicted gaze heat map; is the regularization parameter; is a loss function for training whether the gaze target is within the input image.

[0074] Finally, step S4 is executed to filter each predicted gaze target using head posture information, retaining those that meet the gaze constraints. These gaze targets are the consumer's focus points. Using head posture information, we limit the gaze angle range, simulating the natural human visual range, filtering out incorrect predictions that do not conform to the orientation, and improving accuracy in complex scenarios.

[0075] Specifically, the gaze constraint condition is as follows: calculate the direction vector of the line connecting the center position of the head in the input image and the gaze target, and then calculate the angle between the direction vector and the face heading vector; if the angle is less than the gaze threshold, it is determined that the gaze target with an angle less than the gaze threshold meets the gaze constraint condition.

[0076] For example, the gaze direction vector in the image plane is calculated based on the center position coordinates of the head (xh, yh) and the coordinates of the gaze target (xg, yg) :

[0077] ;

[0078] Then calculate the face orientation vector in the projected head posture information and gaze direction vector The angle θ between them:

[0079] ;

[0080] ;

[0081] in:

[0082] ;

[0083] ;

[0084] Set the gaze threshold as θthreshold. If θ≤θthreshold, the prediction (xg, yg) of the gaze target is accepted; otherwise, it is considered an invalid prediction, thereby further improving the accuracy of the prediction.

[0085] Furthermore, after step S4, the method further includes employing a target tracking algorithm, using the head bounding box of the first frame of the video to be estimated as the initial tracking target, tracking the head bounding box of each frame of the video to be estimated, and recording the gaze trajectory of the head within each head bounding box. Each frame of the video to be estimated serves as the input image. Examples of target tracking algorithms include KCF (Kernelized Correlation Filters) and CSRT (Discriminative Correlation Filter with Channel and Spatial Reliability). These algorithms can continuously locate the target in subsequent frames based on the target's features in the initial frame.

[0086] In the video to be estimated, the head bounding box in the first frame is selected as the initial tracking target. The head bounding box is a rectangular box that defines the area of ​​the person's head in the video. This bounding box can be used to determine information such as the target's position and size, which serves as the starting point for subsequent tracking. A gaze heatmap is a visual representation that shows where a person's eyes are fixated while viewing an image or video. Brighter areas are more likely to be fixated. In each frame of the video, an object tracking algorithm, combined with the initial head bounding box information, continuously tracks the corresponding gaze heatmap, ensuring accurate capture of changes in the gaze area corresponding to each head during video playback. During the tracking process, the gaze trajectory corresponding to each head is recorded. A gaze trajectory refers to the path of the eye gaze point on the image plane during video playback. By recording these trajectories, further information such as the person's attention distribution and focus can be analyzed. By tracking the heads of the characters in the video and recording the gaze trajectory corresponding to each head, the gaze trajectory can reflect the changes in the location of the person's eye gaze in the video and support multi-scene information mining.

[0087] The embodiment is only a special example and does not represent only one way of implementing the present invention.

[0088] Example 2:

[0089] Based on the same inventive concept, Figure 4 As shown, the second embodiment of the present invention further provides an AI-based consumer focus point estimation device, including: a head detection module 1, used to perform head detection on an input image through a head detection algorithm to obtain head posture information; a feature extraction module 2, used to add the head posture information to the feature map of the input image to obtain a head condition feature map; a gaze estimation module 3, used to predict the gaze target corresponding to each head in the input image according to the head condition feature map; a gaze constraint module 4, used to use the head posture information to filter each predicted gaze target and retain the gaze targets that meet the gaze constraint conditions.

[0090] The device also includes a target tracking module 5, which is used to adopt a target tracking algorithm, use the head bounding box of the first frame image in the video to be estimated as the initial tracking target, track the head bounding box of each frame image in the video to be estimated, and record the gaze trajectory corresponding to the head in each head bounding box; wherein, each frame image in the video to be estimated is an input image.

[0091] The device of this embodiment is used to execute the AI-based consumer focus estimation method described in Example 1. It does not rely on the detection of the frontal face, can estimate the gaze direction of the characters in the video with high precision, and can also realize the target tracking function.

[0092] The foregoing is merely a preferred embodiment of the present invention. Those skilled in the art will appreciate that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. Furthermore, under the guidance of the present invention, these features and embodiments may be modified to suit specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are intended to be within the scope of the present invention.

Claims

1. A method for estimating consumer focus points based on AI, characterized in that: The following steps are involved: Perform head detection on the input image through the head detection algorithm to obtain head posture information; Adding the head posture information to the feature map of the input image to obtain a head condition feature map; predicting gaze targets corresponding to all heads in the input image according to the head condition feature map; Using the head posture information to filter each of the predicted gaze targets, and retaining the gaze targets that meet the gaze constraint condition; The head detection algorithm is used to perform head detection on the input image to obtain head posture information, including: Generate a head bounding box of the head through a pre-trained multi-person head pose estimation model; Acquiring a three-dimensional posture angle of the head through the multi-person head posture estimation model; Calculating a frontal face orientation vector of the head according to the three-dimensional posture angle; wherein the head posture information includes the head bounding box and the frontal face orientation vector; Calculate the face orientation vector for: ; ; in, is the three-dimensional posture angle; is the rotation matrix of the three-dimensional posture angle; for The rotation matrix of for The rotation matrix of for The rotation matrix of Adding the head posture information to the feature map of the input image to obtain a head condition feature map includes: Using a visual feature extractor to extract features from the input image to generate the feature map; Generate a head position embedding by converting the head bounding box into normalized coordinates; Convert the head bounding box into a downsampled binary mask and align it with the feature map in the spatial dimension; Embedding the head position element-by-element into the portion of the binary mask corresponding to the head position to generate the head conditional feature map; In the step of screening each predicted gaze target by using the head posture information and retaining the gaze targets that meet a gaze constraint condition, the gaze constraint condition is: Calculating a direction vector of a line connecting the center position of the head in the input image and the gaze target, and then calculating an angle between the direction vector and the frontal face orientation vector; If the included angle is smaller than the gaze threshold, it is determined that the gaze target with the included angle smaller than the gaze threshold meets the gaze constraint condition.

2. The AI-based consumer focus estimation method according to claim 1, characterized in that: 、 and The calculation formula is:

3. The AI-based consumer focus estimation method according to claim 1, characterized in that: The predicting, according to the head condition feature map, a gaze target corresponding to each head in the input image comprises: The gaze heat map decoder predicts the position of the gaze target corresponding to each head in the input image according to the head condition feature map, and obtains a gaze heat map of each gaze target; wherein the gaze heat map represents the probability of being the gaze target in the input image.

4. The AI-based consumer focus estimation method according to claim 3, characterized in that: The gaze heat map is trained and optimized using a binary cross entropy loss function; the binary cross entropy loss function for: ; in, is the pixel-level binary cross entropy loss; is the height of the gaze heat map; is the width of the gaze heat map; is the pixel index of the gaze heat map, representing the two-dimensional coordinate of the gaze heat map, i represents the row, and j represents the column; For a real gaze heat map; is the predicted gaze heat map; is the regularization parameter; is a loss function for training whether the gaze target is within the input image.

5. The AI-based consumer focus estimation method according to claim 1, characterized in that: After screening each of the predicted gaze targets using the head posture information and retaining the gaze targets that meet the gaze constraint condition, the method further includes: A target tracking algorithm is used, with the head bounding box of the first frame image in the video to be estimated as the initial tracking target, the head bounding box of each frame image in the video to be estimated is tracked, and the gaze trajectory corresponding to each head is recorded; wherein, each frame image in the video to be estimated is the input image.

6. An AI-based consumer focus estimation device, characterized in that: include: The head detection module is used to perform head detection on the input image using a head detection algorithm to obtain head posture information; A feature extraction module, configured to add the head posture information to the feature map of the input image to obtain a head condition feature map; a gaze estimation module, configured to predict a gaze target corresponding to each head in the input image based on the head condition feature map; a gaze constraint module, configured to filter each of the predicted gaze targets using the head posture information, and retain the gaze targets that meet the gaze constraint conditions; The head detection algorithm is used to perform head detection on the input image to obtain head posture information, including: Generate a head bounding box of the head through a pre-trained multi-person head pose estimation model; Acquiring a three-dimensional posture angle of the head through the multi-person head posture estimation model; Calculating a frontal face orientation vector of the head according to the three-dimensional posture angle; wherein the head posture information includes the head bounding box and the frontal face orientation vector; Calculate the face orientation vector for: ; ; in, is the three-dimensional posture angle; is the rotation matrix of the three-dimensional posture angle; for The rotation matrix of for The rotation matrix of for The rotation matrix of Adding the head posture information to the feature map of the input image to obtain a head condition feature map includes: Using a visual feature extractor to extract features from the input image to generate the feature map; Generate a head position embedding by converting the head bounding box into normalized coordinates; Convert the head bounding box into a downsampled binary mask and align it with the feature map in the spatial dimension; Embedding the head position element-by-element into the portion of the binary mask corresponding to the head position to generate the head conditional feature map; In the step of screening each predicted gaze target by using the head posture information and retaining the gaze targets that meet a gaze constraint condition, the gaze constraint condition is: Calculating a direction vector of a line connecting the center position of the head in the input image and the gaze target, and then calculating an angle between the direction vector and the frontal face orientation vector; If the included angle is smaller than the gaze threshold, it is determined that the gaze target with the included angle smaller than the gaze threshold meets the gaze constraint condition.

7. The AI-based consumer focus estimation device according to claim 6, characterized in that: The device also includes a target tracking module, which is used to adopt a target tracking algorithm, use the head bounding box of the first frame image in the video to be estimated as the initial tracking target, track the head bounding box of each frame image in the video to be estimated, and record the gaze trajectory corresponding to the head in each head bounding box; wherein, each frame image in the video to be estimated is the input image.

Citation Information

Patent Citations

  • Human-computer interaction method and system , equipment and storage medium

    CN113093907A

  • Gaze target detection method based on visual and semantic clues

    CN116402991A