A single-camera multi-mode line-of-sight tracking method and device, and an electronic device

Through the single-camera multi-mode gaze tracing method, adaptive three-dimensional face alignment and Hough transformation technology are used to solve the real-time and device dependence problems of gaze tracing, and efficient gaze tracing in different environments.

CN114494347BActive Publication Date: 2025-07-18HUMMINGBIRD INNOVATION (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210073470.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-07-18
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

The existing line-of-view tracing technology has problems such as insufficient real-time and high dependence on equipment, especially in complex environments, which are difficult to effectively apply.

Method used

The single-camera multi-mode gaze tracing method is adopted, and through the attention-enhanced adaptive three-dimensional face alignment technology, combined with Hough transformation and three-dimensional face reference vectors, specific information is established to realize user feature calibration and eye movement reference vector determination, and reduce device dependence.

Benefits of technology

It improves the accuracy and real-timeness of gaze tracing, reduces costs, and can be applied in simple or complex environments, especially for partial obstruction, distortion, deformation and other situations of human faces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494347B_ABST
    Figure CN114494347B_ABST
Patent Text Reader

Abstract

The present invention discloses a single-camera multi-mode gaze tracking method and device, and an electronic device, belonging to the technical field of human-computer interaction. The method includes: aligning the face information collected by the camera through an attention-enhanced adaptive three-dimensional face alignment method; calibrating the face reference features and solving the face reference vectors during the gaze tracking process; using the three-dimensional face alignment result to calculate the normal vectors of multiple key-point triangular faces on the facial mesh; using the three-dimensional face center as the starting point to represent the reference projection vector, fitting the gaze transfer matrix, and obtaining the dynamic reference vector of the flexible head in the three-dimensional space; establishing the specific information of the user to be tracked; during the continuous gaze interaction process, using the Hough transform to obtain the iris data of the human eye in the denoised image with prominent edges; and determining the eye movement reference vector based on the iris data of the human eye and the face base vector to complete the gaze tracking, which can improve the timeliness of gaze tracking and has weak dependence on hardware devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and particularly to a single-camera multi-mode gaze tracking method and device, and an electronic device. Background Art

[0002] Gaze tracking has always been a key concern in the field of interaction. The algorithm research and interaction applications for gaze tracking have been quite mature. The face alignment features in three-dimensional space intuitively represent the gaze intention and are the basic information source for most gaze tracking methods. The research on three-dimensional face alignment algorithms has had a huge impact on gaze tracking technology. The methods of gaze tracking can be mainly divided into two categories: hardware device-based gaze tracking methods and computer vision-based gaze tracking methods. The computer vision-based methods can be further divided into end-to-end gaze tracking algorithms for specific data and gaze tracking algorithms based on fixed-point calibration schemes. With the continuous development of gaze tracking, the hardware-based gaze tracking solutions are approaching maturity. However, due to the cost of hardware devices and environmental requirements, it is difficult to promote them comprehensively. Most of them are tested in industrial-level scenarios. The computer vision-based solutions have received more and more attention. More end-to-end algorithms have constructed datasets and gaze tracking methods for proprietary environments, and the fixed-point calibration schemes increasingly rely on multiple cameras for three-dimensional space positioning, increasing the hardware dependence.

[0003] Regarding the related problems of face alignment, a large number of researchers have designed many face alignment methods that conform to the characteristics of the human face structure through the curvature features of the three-dimensional human face and the reference face structure features. Some researchers segmented a convex region within the image based on the features of mean and Gaussian curvature, and created an Extended Gaussian Image (EGI) for each convex region. The EGI describes the shape of the object through the surface normal distribution on the object surface. By correlating the EGI, the regions of the target image are matched with the regions of the gallery image. However, the EGI is insensitive to changes in the size of the object. Therefore, in such a representation form, it is impossible to distinguish two faces with similar shapes but different sizes. Regarding the related problems of gaze tracking, some researchers have proposed a method based on a generalized regression neural network, which uses pupil parameters, pupil flicker displacement, the direction and ratio of the major and minor axes of the pupil ellipse, and the flicker coordinates to map the screen coordinates. This method does not require calibration after initial training, but can only moderately improve head movement.

[0004] There is an urgent need for those skilled in the art to provide a gaze tracking method with good real-time performance and low device dependence. Summary of the Invention

[0005] The objective of the embodiments of the present invention is to provide a single-camera multi-mode gaze tracking method, device, and electronic device, which can improve real-time performance and reduce dependence on devices when performing gaze tracking on users.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] A single-camera multi-mode gaze tracking method, wherein the method includes:

[0008] Align the face information collected by the camera through an attention-enhanced adaptive three-dimensional face alignment method to obtain a three-dimensional face alignment result;

[0009] Calibrate the face reference features and solve the face reference vector during the gaze tracking process;

[0010] Use the three-dimensional face alignment result to calculate the normal vectors of multiple key point triangular faces on the facial mesh;

[0011] Characterize the reference projection vector starting from the three-dimensional face center, fit the gaze transfer matrix, and obtain the dynamic reference vector of the flexible head in the three-dimensional space;

[0012] Calibrate the features of different users according to their interaction habits and establish the specific information of the user to be tracked;

[0013] During the continuous gaze interaction process, use the Hough transform to obtain the human eye iris data in the denoised image of the prominent edge; determine the eye movement reference vector based on the human eye iris data and the face reference vector to complete gaze tracking.

[0014] Among them, the step of aligning the face information collected by the camera through an attention-enhanced adaptive three-dimensional face alignment method to obtain a three-dimensional face alignment result includes:

[0015] Determine an adaptive face detection method and use the determined adaptive face detection method to collect face information;

[0016] For the face information collected by the camera, extract multi-scale image receptive fields through multi-scale convolutional kernels, and use deformable convolution to extract the face features in the face information to determine the face features with deformations;

[0017] Introduce an attention mechanism during the three-dimensional face key point recognition process of the face information collected by the camera to obtain a three-dimensional face alignment result.

[0018] Among them, the step of determining an adaptive face detection method includes:

[0019] Two pooling operations and cascaded rectified linear units are used to eliminate the irrelevant feature points in the face information collected by the camera;

[0020] The feature modules of Inception are used as the backbone network to obtain feature information of different structures;

[0021] Shape adaptation is performed, and a single mapping multi-box prediction method is used to extract features of different scales through several different convolutional kernels after the backbone network.

[0022] Among them, during the three-dimensional face key point recognition process of the face information collected by the camera, an attention mechanism is introduced, and the three-dimensional face alignment result includes:

[0023] Three-dimensional face feature extraction based on heat maps, using the first two layers of the DenseNet structure to extract the feature maps of two-dimensional key points, and integrating the features into a heat map of 69 key points, where 68 are key points and 1 is the background;

[0024] A stackable residual structure based on the attention mechanism is used for feature decoding.

[0025] Among them, during the continuous gaze interaction process, the iris data of the human eye in the denoised image with prominent edges is obtained by using the Hough transform; based on the iris data of the human eye and the face reference vector, the eye movement reference vector is determined to complete gaze tracking, including:

[0026] During the continuous gaze interaction process, based on the face reference vector, an iris recognition method based on scale-fixed Hough features is used to locate the dynamic iris center of the eye movement vector;

[0027] Select the key point coordinates with stable features among the face key points as the reference coordinates, and calculate the key features of the iris center coordinates of both eyes and the three-dimensional space information of the face reference vector;

[0028] According to the key features of the three-dimensional space information of the face reference vector, a gaze tracking regression function is used to complete the gaze tracking of the eye movement vector.

[0029] A single-camera multi-mode gaze tracking device, where the device includes:

[0030] An alignment module, which is used to align the face information collected by the camera through an attention-enhanced adaptive three-dimensional face alignment method to obtain a three-dimensional face alignment result;

[0031] A calibration module, which is used to calibrate the face reference features and solve the face reference vector during the gaze tracking process;

[0032] A calculation module, which is used to calculate the normal vectors of multiple key point triangular faces on the facial mesh by using the three-dimensional face alignment result;

[0033] A fitting module, configured to project a reference vector starting from the center of a three-dimensional face as a characterization reference, fit a gaze transfer matrix, and obtain a dynamic reference vector of a flexible head in three-dimensional space;

[0034] An establishment module, configured to calibrate the features of a user according to the interaction habits of different users, and establish specific information of the user to be tracked;

[0035] A tracking module, configured to obtain eye iris data in a denoised image with prominent edges by using the Hough transform during a continuous gaze interaction process; and determine an eye movement reference vector according to the eye iris data and the face reference vector to complete gaze tracking.

[0036] Wherein, the alignment module includes:

[0037] A first sub-module, configured to determine an adaptive face detection method, and collect face information by using the determined adaptive face detection method;

[0038] A second sub-module, configured to extract multi-scale image receptive fields through multi-scale convolutional kernels for the face information collected by a camera, and extract face features in the face information by using deformable convolution to determine face features with deformations;

[0039] A third sub-module, configured to introduce an attention mechanism during the process of performing three-dimensional face key point recognition on the face information collected by the camera to obtain a three-dimensional face alignment result.

[0040] Wherein, when the first sub-module determines the adaptive face detection method, it is specifically configured to:

[0041] Use two pooling operations and a cascaded rectified linear unit to eliminate irrelevant feature points in the face information collected by the camera;

[0042] Obtain feature information of different structures through the feature module of Inception as a backbone network;

[0043] Perform shape adaptation, and use a single mapping multi-box prediction method to extract features of different scales through several different convolutional kernels after the backbone network.

[0044] Wherein, the third sub-module is specifically configured to:

[0045] Extract a feature map of two-dimensional key points by using the first two layers of the DenseNet structure for three-dimensional face feature extraction based on a heat map, and integrate the features into a heat map of 69 key points, where 68 are key points and 1 is the background;

[0046] Perform feature decoding by using a stackable residual structure based on an attention mechanism.

[0047] Among them, the tracking module includes:

[0048] A fourth sub-module, configured to, during a continuous line-of-sight interaction process, based on the face reference vector, use an iris recognition method based on scale-fixed Hough features to locate the dynamic iris center of the eye movement vector;

[0049] A fifth sub-module, configured to select the key point coordinates with stable features among the face key points as reference coordinates, and calculate the key features of the iris center coordinates of both eyes and the three-dimensional space information of the face reference vector;

[0050] A sixth sub-module, configured to complete the eye movement vector line-of-sight tracking by using a line-of-sight tracking regression function according to the key features of the three-dimensional space information of the face reference vector.

[0051] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of any one of the above single-camera multi-mode line-of-sight tracking methods are implemented.

[0052] An embodiment of the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of any one of the above single-camera multi-mode line-of-sight tracking methods are implemented.

[0053] The single-camera multi-mode line-of-sight tracking method provided by the embodiment of the present invention aligns the face information collected by the camera through an attention-enhanced adaptive three-dimensional face alignment method to obtain a three-dimensional face alignment result; calibrates the face reference features and solves the face reference vector in the line-of-sight tracking process; uses the three-dimensional face alignment result to calculate the normal vectors of multiple key point triangular faces on the face mesh; represents the reference projection vector starting from the three-dimensional face center, fits the line-of-sight transfer matrix, and obtains the dynamic reference vector of the flexible head in the three-dimensional space; calibrates the features of the user according to the interaction habits of different users, and establishes the specific information of the user to be tracked; during the continuous line-of-sight interaction process, uses the Hough transform to obtain the eye iris data in the denoised image with prominent edges; determines the eye movement reference vector based on the eye iris data and the face reference vector to complete the line-of-sight tracking. The single-camera multi-mode line-of-sight tracking method provided by the embodiment of the present invention, on the one hand, solves the problems of device dependence and environmental dependence through an adaptive scale face detection method and a single-camera multi-mode line-of-sight tracking method, improves the accuracy and real-time performance of line-of-sight tracking, and reduces the cost of the line-of-sight tracking interaction method; on the other hand, it has a wide range of applications, can be applied to relatively simple environments and environments with higher complexity, and is applicable to problems such as partial occlusion, distortion, and deformation of the face. Description of the Drawings

[0054] Figure 1 It is a flowchart showing the steps of a single - camera multi - mode gaze tracking method according to an embodiment of the present application;

[0055] Figure 2 It is a flowchart showing the steps of the gaze tracking method provided by an embodiment of the present invention;

[0056] Figure 3 It is a network diagram of an attention - enhanced adaptive three - dimensional face alignment method provided by an embodiment of the present invention;

[0057] Figure 4 It is a framework diagram of a face - benchmark fast gaze tracking method provided by an embodiment of the present invention;

[0058] Figure 5 It is a framework diagram of a space - limited eye movement vector gaze tracking method provided by an embodiment of the present invention;

[0059] Figure 6 It is a flowchart of a scale Hough transform method provided by an embodiment of the present invention;

[0060] Figure 7 It is a process diagram of determining an eye movement vector provided by an embodiment of the present invention;

[0061] Figure 8 It is a diagram of iris localization results provided by an embodiment of the present invention;

[0062] Figure 9 It is a structural block diagram of a single - camera multi - mode gaze tracking device according to an embodiment of the present application;

[0063] Figure 10 It is a structural block diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0064] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0065] The following will, with reference to the accompanying drawings, through specific embodiments and their application scenarios, elaborate in detail on the single - camera multi - mode gaze tracking solution provided by an embodiment of the present application.

[0066] Figure 1 It is a flowchart showing the steps of a single - camera multi - mode gaze tracking method according to an embodiment of the present application.

[0067] The single - camera multi - mode gaze tracking method according to an embodiment of the present application includes the following steps:

[0068] Step 101: Align the face information collected by the camera through an attention-enhanced adaptive 3D face alignment method to obtain a 3D face alignment result.

[0069] An optionally attention-enhanced adaptive 3D face alignment method for aligning the face information collected by the camera to obtain a 3D face alignment result includes the following sub-steps:

[0070] Sub-step one: Determine an adaptive face detection method and use the determined adaptive face detection method to collect face information;

[0071] Specifically, the process of determining an adaptive face detection method is as follows:

[0072] First, use two pooling operations and a cascaded rectified linear unit to eliminate irrelevant feature points in the face information collected by the camera;

[0073] Second, use the feature module of Inception as the backbone network to obtain feature information of different structures;

[0074] Finally, perform shape adaptation, and use the single mapping multi-box prediction method to extract features of different scales through several different convolutional kernels after the backbone network.

[0075] Sub-step two: For the face information collected by the camera, extract multi-scale image receptive fields through multi-scale convolutional kernels, and use deformable convolution to extract face features in the face information to determine face features with deformations;

[0076] While extracting multi-scale image receptive fields through multi-scale convolutional kernels, use deformable convolution to further extract face features and regress face features that may have deformations to improve the accuracy of face detection.

[0077] Sub-step three: Introduce an attention mechanism during the 3D face key point recognition process for the face information collected by the camera to obtain a 3D face alignment result.

[0078] For the recognition of 3D face key points, while using the UV position map feature to ensure a low number of calculation parameters, in the case of weak image feature extraction, an attention mechanism is introduced during the process of decoding features to strengthen the role of image texture features in the UV map.

[0079] Specifically, the way to introduce an attention mechanism during the 3D face key point recognition process for the face information collected by the camera to obtain a 3D face alignment result can include the following process:

[0080] First, for three-dimensional face feature extraction based on a heat map, the first two layers of the DenseNet structure are used to extract the feature map of two-dimensional key points, and the features are integrated into a heat map of 69 key points, where 68 are key points and 1 is the background;

[0081] Secondly, a stackable residual structure based on an attention mechanism is used for feature decoding.

[0082] Step 102: Calibrate the face reference features and solve the face reference vector during the gaze tracking process.

[0083] Step 103: Use the three-dimensional face alignment result to calculate the normal vectors of multiple key point triangular faces on the facial mesh.

[0084] Step 104: Represent the reference projection vector starting from the center of the three-dimensional face, fit the gaze transfer matrix, and obtain the dynamic reference vector of the flexible head in the three-dimensional space.

[0085] Step 105: Calibrate the features of the user according to the interaction habits of different users, and establish the specific information of the user to be tracked.

[0086] When establishing the specific information of the gaze tracking user, the user can be set to calibrate at a distance of 30 cm from the computer screen (the face is parallel to the screen), and the calibration point is the center of the screen. The initial calibration data is obtained by the proposed key point detection method, and the remaining calibration information can be gradually deduced through the imaging method, and the calibration result is obtained by calculating the spatial matrix to achieve real-time alignment.

[0087] Steps 102 to 105 are the specific processes of constructing a fast gaze tracking method based on the face reference and using the three-dimensional face alignment result and the single-point feature calibration method to establish the feature mapping relationship. This method can ensure short-time and frequent gaze tracking operations.

[0088] In an optional implementation manner, after establishing the specific information of the user to be tracked, the existing errors can also be corrected.

[0089] When correcting the errors, the deviations caused by the movement of the human head parallel to the screen, perpendicular to the screen, and left and right rotation can be deduced. A control factor α is introduced for the offset problem caused by head rotation. α is taken as 0.1, and the intersection coordinates of the user's head orientation vector on the screen are controlled by α, that is, D(X_d) becomes αD(X_d), and the offset parameter is set in the vertical direction.

[0090] Step 106: During the continuous gaze interaction process, use the Hough transform to obtain the human eye iris data in the denoised image of the prominent edge; determine the eye movement reference vector based on the human eye iris data and the face reference vector to complete the gaze tracking.

[0091] In an alternative embodiment, during a continuous line-of-sight interaction, the human eye iris data in the denoised image with prominent edges is obtained using the Hough transform; based on the human eye iris data and the face reference vector, when determining the eye movement reference vector and completing the line-of-sight tracking, the following sub-steps may be included:

[0092] Sub-step 1: During a continuous line-of-sight interaction, based on the face reference vector, use an iris recognition method based on scale-fixed Hough features to locate the dynamic iris center of the eye movement vector;

[0093] The iris recognition method based on scale-fixed Hough features mainly uses a three-dimensional face alignment method to obtain the human eye frame and cut out the positions of both eyes, and processes each human eye image separately, including the following specific steps:

[0094] First, judge the human behavior noise, that is, judge whether the human eye is in the process of blinking, using the aspect ratio threshold judgment method.

[0095] Secondly, apply a Gaussian filter to the filtered pictures to remove interference.

[0096] Again, use the result of Gaussian filtering in the form of a normal distribution to enhance the edge features of the image, calculate the gradient change direction of the image, and enhance the gradient change. Further filter out the non-maximum gradient features using the non-maximum suppression method, keep the edges greater than the gradient value, and set the edge grayscales less than the gradient value to 0.

[0097] Finally, use the double-threshold method to detect edges.

[0098] Sub-step 2: Select the key point coordinates with stable features among the face key points as the reference coordinates, and calculate the key features of the iris center coordinates of both eyes and the three-dimensional space information of the face reference vector;

[0099] Specifically, calculating the key features of the iris center coordinates of both eyes and the three-dimensional space information of the face reference vector includes the following specific steps:

[0100] First, use 33 dense points (including the corners of the eyes) with obvious and fixed structural features in the human face as the reference point set of the eye movement vector, and the 33 three-dimensional points are obtained by the method proposed in Chapter 3 for recognizing the images collected by a single camera.

[0101] Secondly, due to the loss of features caused by occlusion problems, the information of one eye is abnormal or the information symmetry of part of the face is abnormal. The final eye movement vector is the result of integrating two eye movement vectors, and a defined threshold method is used to limit the integration vector of the binocular eye movement vectors.

[0102] Among them, the average depth coordinate of the eye contour is used as the iris depth coordinate.

[0103] Sub-step 3: According to the key features of the three-dimensional space information of the face reference vector, use the gaze tracking regression function to complete the gaze tracking of the eye movement vector.

[0104] In sub-step 3, according to the key features of the three-dimensional space information of the face reference vector, use the gaze tracking regression function to complete the gaze tracking of the eye movement vector. Calibrate the initial features using the 9-point calibration method, use the regression function that mixes the first- and second-order eye movement vectors and the head reference, and adopt the regression space influence of the depth vector of the iris.

[0105] The single-camera multi-mode gaze tracking method provided by the embodiment of the present application, on the one hand, solves the problems of device dependence and environmental dependence through the adaptive scale face detection method and the single-camera multi-mode gaze tracking method, improves the accuracy and real-time performance of gaze tracking, and reduces the cost of the gaze tracking interaction method; on the other hand, it has a wide range of applications and can be applied to relatively simple environments and environments with higher complexity, and is applicable to problems such as partial occlusion, distortion, and deformation of the face.

[0106] Figure 2 It is the step flow chart of a single-camera multi-mode gaze tracking method provided by the embodiment of the present invention; the single-camera multi-mode gaze tracking method provided by the embodiment of the present application includes the following steps:

[0107] Step (1): Align the face information collected by the single camera through the attention-enhanced adaptive three-dimensional face alignment method.

[0108] Step (2): Adopt the three-dimensional face alignment result and the single-point feature calibration method to construct a fast gaze tracking method based on the face reference, and establish a feature mapping relationship to ensure short-term and frequent gaze tracking operations.

[0109] Step (3): For continuous gaze interaction behaviors, construct a gaze tracking method for the eye movement vector of the face reference features, use the Hough transform to obtain the human eye iris data in the denoised image with prominent edges, cooperate with the face reference vector to form an eye movement reference vector, and complete gaze tracking.

[0110] Figure 3 It is the network diagram of the attention-enhanced adaptive three-dimensional face alignment method provided by the embodiment of the present invention.

[0111] Further, as Figure 3 shown, step (1): The attention-enhanced adaptive three-dimensional face alignment method aligns the face information collected by the single camera, including the following specific steps (1-1) to step (1-3).

[0112] Step (1-1): Reduce the face detection box recognition time through a single-stage object detection method and propose an adaptive face detection method.

[0113] Step (1-2): While extracting multi-scale image receptive fields using multi-scale convolutional kernels, use deformable convolutions to further extract face features, regress the face features that may be deformed, and improve the accuracy of face detection.

[0114] Step (1-3): For the recognition of 3D face key points, use the UV position map feature to ensure a low number of calculation parameters. In the case of weak image feature extraction, introduce an attention mechanism during the decoding of features to strengthen the role of image texture features in the UV map.

[0115] Furthermore, step (1-1) reduces the face detection box recognition time through a single-stage object detection method and proposes an adaptive face detection method, including the following specific steps (1-1-1) to step (1-1-2):

[0116] Step (1-1-1): Feature extraction: Initially, the network uses a 7×7 and two 5×5 large convolutional kernels to obtain a large receptive field. Each convolution uses a sliding step of length 2 to shrink the feature space. Among them, two pooling operations are used to eliminate irrelevant feature points. Use the Inception feature module as the backbone network to obtain feature information of different structures, set several convolution scales such as 1×1, 3×3, 5×5, 7×7, etc., add a path that can represent the 7×7 module on the basis model of Inception v2, and obtain a richer receptive field by fusing features of different layers.

[0117] The following Table 1 is the feature extraction network structure table:

[0118]

[0119] Step (1-1-2): Shape adaptation: One is to reduce the training difficulty through prior boxes of different scales, use different default scales in different feature layers, then set different aspect ratios to form feature maps of different scales, and train by selecting the prediction box closest to the fixed scale. The default box size of each feature map is calculated as follows:

[0120]

[0121] Another prediction at a different scale enhances the network's adaptability to shapes. First, 1×1 convolutions are used to integrate features on the feature maps output by the backbone network. Then, by adding an offset to each convolutional sample to represent the relationship between each sample point and its surrounding points, the feature effects at the deformation scale are collected. First, the input feature map passes through a common convolution to obtain the offset of the pixels. The offset is added to the index of the original feature map, causing the point P0 on the original feature map to move to P l position, and the position coordinates of P l are obtained from the convolution result, so 4 pairs of integers can be used to represent the existing actual coordinates P l :

[0122] [[f(x), f(y)], [f(x), e(y)], [e(x), f(y)], [e(x), f(y)]]

[0123] Furthermore, in step (1-3), the UV position map features are adopted. For the case where the image feature extraction is weak, an attention mechanism is introduced during the process of decoding features, including the following specific steps (1-3-1) to (1-3-2):

[0124] Step (1-3-1): 3D face feature extraction based on the heat map: The first two layers of the DenseNet structure are used to extract the feature maps of 2D key points, and the features are integrated into a heat map of 69 key points (where 68 are key points and 1 is the background).

[0125] The following Table 2 shows the feature extraction network:

[0126]

[0127] Step (1-3-2): Face alignment based on attention decoding: A stackable residual structure based on the attention mechanism is used for feature decoding.

[0128] The following Table 3 shows the PRNet decoder structure:

[0129]

[0130]

[0131] Figure 4 This is the framework diagram of the fast gaze tracking method based on face landmarks provided by the embodiments of the present invention;

[0132] Furthermore, as Figure 4 shown, step (2) constructs a fast gaze tracking method based on face landmarks by using the 3D face alignment result and the single-point feature calibration method, and establishes a feature mapping relationship to ensure short-term and frequent gaze tracking operations, including the following specific steps:

[0133] Step (2-1): First, perform face fiducial feature calibration, solve the specific face fiducial vector during the gaze tracking process, calculate the sum of the normal vectors of multiple key-point triangular faces on the facial mesh using the result of 3D face alignment, represent the fiducial projection vector with the 3D face center as the starting point, fit the gaze transfer matrix, and deduce the dynamic fiducial vector of the flexible head in 3D space. For different users' interaction habits, it is necessary to calibrate the users' features and establish the specific information of the gaze-tracking users.

[0134] Step (2-2): Secondly, correct the existing errors.

[0135] Furthermore, in step (2-1) for establishing the specific information of the gaze-tracking users, it is set that the user calibrates at a distance of 30 cm from the computer screen (the face parallel to the screen), and the calibration point is the center of the screen. The initial calibration data is obtained by the key-point detection method proposed in step (1), and the remaining calibration information can be gradually deduced by the imaging method, and the calibration result is obtained by calculating the spatial matrix to achieve real-time alignment.

[0136] Furthermore, in step (2-2) for correcting the errors, the deviations caused by the movement of the human head parallel to the screen, perpendicular to the screen, and rotation left and right are deduced. For the offset problem caused by the head rotation, a control factor α is introduced, α takes 0.1, and the intersection coordinates of the user's head orientation vector on the screen are controlled by α, that is, D(X_d) becomes αD(X_d), and the offset parameter is set in the vertical direction.

[0137] Figure 5 This is the framework diagram of the spatially defined eye movement vector gaze tracking method provided by the embodiment of the present invention.

[0138] Furthermore, as Figure 5 shown, in step (3) for continuous gaze interaction behavior, a gaze tracking method of the face fiducial feature eye movement vector is constructed. The iris data of the human eye in the denoised image with prominent edges is obtained by using the Hough transform, and the gaze reference vector is formed in cooperation with the face fiducial vector to complete the gaze tracking, including the following specific steps:

[0139] Step (3-1): According to the face fiducial vector obtained in the previous step, use the iris recognition method based on the scale-fixed Hough feature to locate the dynamic iris center of the eye movement vector.

[0140] The flow chart of the scale Hough transform method is as Figure 6 shown.

[0141] Step (3-2): Then, select the key point coordinates of the face contour, the tip of the nose, the eye contour, etc. with stable features among the face key points as the reference coordinates, and calculate the iris center coordinates of both eyes (using the average depth coordinate of the eye contour as the iris depth coordinate) and the key features of the three-dimensional space information of the face reference vector.

[0142] Step (3-3): Complete the eye movement vector gaze tracking according to the key features of the three-dimensional space information of the face reference vector by using the gaze tracking regression function.

[0143] Furthermore, the iris recognition method based on the scale-fixed Hough feature in step (3-1) mainly uses the three-dimensional face alignment method in step (1) to obtain the eye frame and cut out the positions of both eyes, and processes each eye image separately, including the following specific steps:

[0144] Step (3-1-1): Judgment of human behavior noise: The human eye contour can be represented by P1, P2, P3, P4, P5, P6, and the aspect ratio of the human eye edge, that is, the Eye Aspect Ratio (EAR), can be calculated through these 6 points:

[0145]

[0146] In the embodiment of the present invention, the set threshold is 0.2, and if it is less than 0.2, it is considered to be in the blinking stage.

[0147] Step (3-1-2): Removal of interference by Gaussian filter: For any point P(x, y) on the picture, assuming its gray value is f(x, y), after filtering, it becomes:

[0148]

[0149] Step (3-1-3): Enhancement of edge features: Utilize the result of Gaussian filtering in the form of normal distribution to enhance the edge features of the image, calculate the gradient change direction of the image, and enhance the gradient change.

[0150] Step (3-1-4): Filtering of non-maximum gradient features: The degree of gray value change, that is, the result of gradient change, is obtained by multiplying each point by the sobel operator to represent the edge. However, there is still a situation of edge amplification during the filtering process. The non-maximum gradient features are further filtered out by using the non-maximum suppression method, and the edges greater than the gradient value are retained, and the gray values of the edges less than the gradient value are set to 0.

[0151] Step (3-1-5): Detection of edges: Perform double-threshold detection on the edge feature map, remove the features outside the lower threshold and outside the upper threshold, and update the intermediate data to 1 to emphasize the eye features.

[0152] Further, step (3-2) calculates the key features of the iris center coordinates of both eyes and the three-dimensional spatial information of the face reference vector, including the following specific steps:

[0153] Step (3-2-1): Reference point set: Use 33 dense points with obvious and fixed structural features in the face

[0154] (including the corners of the eyes) as the reference point set of the eye movement vector. The 33 three-dimensional points are obtained by identifying the images collected by a single camera through the method proposed in Chapter 3. The coordinate point calculation formula of the final reference point is as follows:

[0155]

[0156] Step (3-2-2): Integrate vectors: For the loss of features due to occlusion problems, resulting in abnormal information of one eye or abnormal symmetry of part of the face information. The final eye movement vector is the result of integrating two eye movement vectors. The integration vector of the binocular eye movement vector is restricted by defining a threshold. The final eye movement vector can be expressed as:

[0157]

[0158] Further, in step (3-3), the eye movement vector line-of-sight tracking is completed according to the key features of the three-dimensional spatial information of the face reference vector. The initial features are calibrated using the 9-point calibration method, and a regression function that combines the first-order and second-order eye movement vectors and the head reference is used to adopt the depth vector regression space influence of the iris.

[0159] Among them, the process diagram of determining the eye movement vector is shown in Figure 7, and the iris positioning result diagram is as Figure 8 shown.

[0160] The single-camera multi-mode line-of-sight tracking method provided by the embodiment of the present invention solves the problems of device dependence and environmental dependence through the adaptive-scale face detection method and the single-camera multi-mode line-of-sight tracking method, improves the accuracy and real-time performance of line-of-sight tracking, and reduces the cost of the line-of-sight tracking interaction method.

[0161] Figure 9 It is a structural block diagram of a single-camera multi-mode line-of-sight tracking device for implementing an embodiment of the present application.

[0162] The single-camera multi-mode line-of-sight tracking device of the embodiment of the present application includes the following functional modules:

[0163] An alignment module 201, configured to align the face information collected by the camera through an attention-enhanced adaptive three-dimensional face alignment method to obtain a three-dimensional face alignment result;

[0164] The calibration module 202 is used to calibrate the face reference features and solve the face reference vector during the gaze tracking process;

[0165] The calculation module 203 is used to calculate the normal vectors of multiple key-point triangular faces on the facial mesh by using the three-dimensional face alignment result;

[0166] The fitting module 204 is used to represent the reference projection vector starting from the center of the three-dimensional face, fit the gaze transfer matrix, and obtain the dynamic reference vector of the flexible head in the three-dimensional space;

[0167] The establishment module 205 is used to calibrate the features of the user according to the interaction habits of different users and establish the specific information of the user to be tracked;

[0168] The tracking module 206 is used to obtain the human eye iris data in the denoised image of the prominent edge by using the Hough transform during the continuous gaze interaction process; and determine the eye movement reference vector according to the human eye iris data and the face reference vector to complete the gaze tracking.

[0169] Optionally, the alignment module includes:

[0170] The first sub-module is used to determine an adaptive face detection method and collect face information by using the determined adaptive face detection method;

[0171] The second sub-module is used to extract the multi-scale image receptive fields through multi-scale convolution kernels for the face information collected by the camera, and use deformable convolution to extract the face features in the face information to determine the face features with deformations;

[0172] The third sub-module is used to introduce an attention mechanism during the three-dimensional face key-point recognition process for the face information collected by the camera to obtain the three-dimensional face alignment result.

[0173] Optionally, when the first sub-module determines the adaptive face detection method, it is specifically used for:

[0174] Using two pooling operations and a cascaded rectified linear unit to remove the irrelevant feature points in the face information collected by the camera;

[0175] Obtaining the feature information of different structures through the feature module of Inception as the backbone network;

[0176] Performing shape adaptation, and using the single-shot multibox prediction method to extract features of different scales through several different convolution kernels after the backbone network.

[0177] Optionally, the third sub-module is specifically used for:

[0178] 3D face feature extraction based on heat map, using the first two layers of DenseNet structure to extract the feature map of 2D key points, and integrating the features into a heat map of 69 key points, where 68 are key points and 1 is the background;

[0179] A stackable residual structure based on the attention mechanism is used for feature decoding.

[0180] Optionally, the tracking module includes:

[0181] A fourth sub-module for locating the dynamic iris center of the eye movement vector based on the face reference vector using an iris recognition method based on scale-fixed Hough features during continuous line-of-sight interaction;

[0182] A fifth sub-module for selecting the key point coordinates with stable features among the face key points as the reference coordinates, and calculating the key features of the iris center coordinates of both eyes and the three-dimensional spatial information of the face reference vector;

[0183] A sixth sub-module for completing the line-of-sight tracking of the eye movement vector using a line-of-sight tracking regression function based on the key features of the three-dimensional spatial information of the face reference vector.

[0184] The single-camera multi-mode line-of-sight tracking device provided by the embodiments of the present application.

[0185] In the embodiments of the present application Figure 9 The single-camera multi-mode line-of-sight tracking device shown can be a device, or a component, integrated circuit, or chip in a server. In the embodiments of the present application Figure 2 The single-camera multi-mode line-of-sight tracking device shown can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0186] The single-camera multi-mode line-of-sight tracking device provided by the embodiments of the present application Figure 9 shown can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0187] Optionally, as Figure 10 shown, the embodiments of the present application also provide an electronic device 300, including a processor 301, a memory 302, a program or instruction stored on the memory 302 and executable on the processor 301. When the program or instruction is executed by the processor 301, it implements each process of the above-mentioned single-camera multi-mode line-of-sight tracking method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0188] It should be noted that the electronic device in the embodiments of the present application includes the server described above.

[0189] The embodiments of the present application further provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned embodiment of the single-camera multi-mode gaze tracking method is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0190] Wherein, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc.

[0191] The embodiments of the present application further provide a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run a program or instruction to implement each process of the above-mentioned embodiment of the single-camera multi-mode gaze tracking method, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.

[0192] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0193] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including that element.

[0194] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A single-camera multi-mode gaze tracking method, characterized in that, The method includes: Aligning the face information collected by the camera through an attention-enhanced adaptive 3D face alignment method to obtain a 3D face alignment result; Calibrating the face reference features and solving the face reference vector during the gaze tracking process; Using the 3D face alignment result to calculate the normal vectors of multiple key-point triangular faces on the facial mesh; Characterizing the reference projection vector starting from the 3D face center, fitting the gaze transfer matrix, and obtaining the dynamic reference vector of the flexible head in 3D space; Calibrating the features of different users according to their interaction habits and establishing the specific information of the user to be tracked; During the continuous gaze interaction process, using the Hough transform to obtain the human eye iris data in the denoised image with prominent edges; determining the eye movement reference vector based on the human eye iris data and the face reference vector to complete gaze tracking; Among them, aligning the face information collected by the camera through an attention-enhanced adaptive 3D face alignment method to obtain a 3D face alignment result includes: Determining an adaptive face detection method and using the determined adaptive face detection method to collect face information; For the face information collected by the camera, extracting multi-scale image receptive fields through multi-scale convolutional kernels, and using deformable convolution to extract the face features in the face information to determine the deformed face features; Introducing an attention mechanism during the 3D face key-point recognition process for the face information collected by the camera to obtain a 3D face alignment result; Among them, determining an adaptive face detection method includes: Using two pooling operations and a cascaded rectified linear unit to eliminate the irrelevant feature points in the face information collected by the camera; Using the feature module of Inception as the backbone network to obtain feature information of different structures; Performing shape adaptation, and using a single mapping multi-box prediction method to extract features of different scales through several different convolutional kernels after the backbone network; Among them, introducing an attention mechanism during the 3D face key-point recognition process for the face information collected by the camera to obtain a 3D face alignment result includes: 3D face feature extraction based on a heat map, using the first two layers of the DenseNet structure to extract the feature map of 2D key points, and integrating the features into a heat map of 69 key points, where 68 are key points and 1 is the background; Performing feature decoding using a stackable residual structure based on the attention mechanism.

2. The method according to claim 1, wherein During the continuous gaze interaction process, using the Hough transform to obtain the human eye iris data in the denoised image with prominent edges; determining the eye movement reference vector based on the human eye iris data and the face reference vector to complete gaze tracking includes: During the continuous gaze interaction process, based on the face reference vector, using an iris recognition method based on scale-fixed Hough features to locate the dynamic iris center of the eye movement vector; Selecting the key-point coordinates with stable features among the face key points as the reference coordinates, and calculating the key features of the iris center coordinates of both eyes and the 3D space information of the face reference vector; According to the key features of the 3D space information of the face reference vector, using a gaze tracking regression function to complete the gaze tracking of the eye movement vector.

3. A single-camera multi-mode line-of-sight tracking device, characterized in that, The device includes: An alignment module, which is used to align the face information collected by the camera through an attention-enhanced adaptive 3D face alignment method to obtain a 3D face alignment result; A calibration module, which is used to calibrate the face reference features and solve the face reference vectors during the gaze tracking process; A calculation module, which is used to calculate the normal vectors of multiple key-point triangular faces on the facial mesh by using the 3D face alignment result; A fitting module, which is used to represent the reference projection vector starting from the 3D face center, fit the gaze transfer matrix, and obtain the dynamic reference vector of the flexible head in the 3D space; A building module, which is used to calibrate the features of the user according to the interaction habits of different users and establish the specific information of the user to be tracked; A tracking module, which is used to obtain the eye iris data in the denoised image with prominent edges by using the Hough transform during the continuous gaze interaction process; and determine the eye movement reference vector based on the eye iris data and the face reference vector to complete the gaze tracking; Among them, the alignment module includes: A first sub-module, which is used to determine an adaptive face detection method and collect face information by using the determined adaptive face detection method; A second sub-module, which is used to extract multi-scale image receptive fields from the face information collected by the camera through multi-scale convolutional kernels, and use deformable convolution to extract the face features in the face information to determine the face features with deformations; A third sub-module, which is used to introduce an attention mechanism during the 3D face key-point recognition of the face information collected by the camera to obtain a 3D face alignment result; Among them, when the first sub-module determines the adaptive face detection method, it is specifically used for: Using two pooling operations and a cascaded rectified linear unit to eliminate the irrelevant feature points in the face information collected by the camera; Obtaining feature information of different structures through the feature module of Inception as the backbone network; Performing shape adaptation, and using a single mapping multi-box prediction method to extract features of different scales through several different convolutional kernels after the backbone network; Among them, the third sub-module is specifically used for: 3D face feature extraction based on a heat map, using the first two layers of the DenseNet structure to extract the feature map of 2D key points, and integrating the features into a heat map of 69 key points, where 68 are key points and 1 is the background; Performing feature decoding by using a stackable residual structure based on the attention mechanism.

4. An electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the single-camera multi-mode gaze tracking method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Second-level sight line tracing method based on face orientation constraint

    CN107193383A

  • Sight tracking-based care demand identification method for old and disabled people

    CN112232128A