Personalized sight line estimation method and system based on causal inference
Through the personalized line of sight estimation method based on causal inference, the influence of personalized information is eliminated and the relevant features of sight is purified, and the prediction deviation problem caused by personalized differences in the prior art is solved, thereby achieving a low-cost and portable line of sight estimation effect.
Patent Information
- Application Number
- CN202510116713.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The existing line of sight estimation technology has prediction deviations caused by personalized differences when facing different users, and the equipment is costly and complex, making it difficult to popularize in ordinary families.
The personalized sight estimation method based on causal inference is adopted to analyze the causal relationship between images, sight and user individual information through a structured causal model, eliminate the influence of personalized information, and purify the sight-related characteristics.
The adaptability and generalization ability of the line of sight estimation model to new users is improved, and unconstrained personalized line of sight estimation is achieved, with low cost, simple operation and strong portability.
Smart Images

Figure CN119992636A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a personalized sight line estimation method and system based on causal inference. Background Art
[0002] At present, the devices used for line of sight estimation on the market are mainly eye trackers, including desktop eye trackers and wearable eye trackers. Desktop devices are generally fixed in a laboratory environment and require the subjects to maintain a stable head posture for a certain period of time. Compared with desktop devices, head-mounted devices are more flexible to use, but they require the subjects to wear the collection equipment with them, which will bring certain constraints on the senses. Therefore, both types of devices are not suitable for certain usage scenarios that require the collection of the user's unconstrained line of sight direction. For example, when collecting and analyzing the eye movement data of children with autism spectrum disorders, due to their low level of empathy and poor obedience to instructions, it is difficult to maintain a fixed head and body posture for a long period of time. The use of head-mounted devices has a strong sense of constraint, which may have a negative impact on their condition. In addition, the prices of existing eye tracker devices are relatively expensive, and their operation and use require the guidance and training of professionals, making it difficult to popularize them in ordinary families in real life.
[0003] In contrast, desktop computers have been integrated into people's daily lives, and people are familiar with their operation methods. Therefore, the present invention develops an intelligent sight line estimation system based on desktop computers, aiming to accurately estimate the sight line direction information of users under unconstrained conditions.
[0004] At present, the gaze estimation method based on deep learning is the mainstream method in the field of gaze estimation. This kind of method relies on a large amount of labeled training data, and learns the mapping from eye or face images to gaze direction by designing effective deep neural networks, and achieves good accuracy. However, in complex and diverse practical application scenarios, gaze estimation models are often oriented to open test environments, that is, the user to be tested is different from the model training user. There are personalized differences between different users, including the following aspects. 1) Physiological differences: The physiological structure of the eyeballs of different users is different, that is, the kappa angle between the visual axis (used to define the gaze direction) and the optical axis (defined by the iris center) varies from person to person. Therefore, when different people look at the same gaze target, their eyeball appearance also has certain differences. 2) There are also inevitable differences in facial appearance, head posture, gaze habits, background environment, light intensity, etc. of different users. This leads to large differences in the acquired facial images. Due to the personalized differences between users, the gaze estimation model that only uses fixed user optimization is prone to overfitting to the personalized information of the training user, and when it is applied to new users in unseen scenes, it will produce large prediction deviations. Therefore, the performance of existing line of sight estimation models cannot meet the high-precision requirements of practical applications.
[0005] As early as the ancient Greek period, people began to study eye movements. In the Middle Ages, researchers tried to combine mathematical analysis with anatomy to study the eyeball, which pioneered the introduction of instruments into this research. Over the past hundreds of years, with the development of science and technology and the continuous efforts of researchers, the line of sight estimation technology has made rapid progress and has become increasingly mature. Early traditional line of sight estimation methods were all invasive, requiring the physical equipment for the experiment to be in direct contact with the user, mainly including the contact lens method and the electro-oculogram method. The basic principle of the contact lens method is to place a special reflector directly on the eye, and obtain eye movement signals about the line of sight based on the direction of the fixed light beam reflected when the eyeball moves. This method has high accuracy, but the equipment used for the experiment is very complex and interferes with the user, so it cannot be promoted and applied. The basic principle of the electrooculogram method is to use the weak potential difference generated by the movement of the eyeballs. Electrodes are placed on the upper and lower sides of the two eyes, the potential difference is amplified and output, and the information of eye movement is obtained by detecting the change in the potential difference. This method is relatively simple in terms of equipment, but the error is large and it will also cause great interference to the user.
[0006] With the rapid development of optical imaging technology and computer technology, modern gaze estimation methods have gradually shifted from signal-based research methods to video image technology-based research methods. The core idea of existing gaze estimation methods is to design an effective gaze estimation model to improve model accuracy, which is mainly divided into model-based methods and appearance-based methods. The model-based method is a traditional gaze estimation method, which mainly estimates the corresponding gaze information based on the precise geometric features of the human eye. The appearance-based gaze estimation method adopts a learning strategy to learn the mapping relationship between face or eye images and gaze to perform gaze estimation.
[0007] (1) Model-based line of sight estimation method
[0008] As a traditional method, model-based line of sight estimation mainly uses the principle of human eye imaging, and uses a camera to capture features such as the cornea, iris, and pupil of the eye to construct a geometric model of the eyeball for line of sight estimation. This type of method can be further divided into shape-based methods and pupil-corneal reflection-based methods.
[0009] Shape-based methods mainly use prior knowledge such as the external structure of the eyeball, and obtain local features of the eye, such as the position of the eye corners and the pupil, and then establish a line of sight estimation model based on the coordinate relationship in three-dimensional space to obtain the line of sight direction. Since this method is overly dependent on the inherent structure of the eyeball and the geometric model of the eyeball varies greatly between individuals, it cannot be flexibly applied to different users and the performance of the method is unstable. In addition, this method has high requirements on image quality, and is prone to large estimation errors when processing low-quality images, and has poor robustness.
[0010] The method based on pupil-corneal reflection mainly uses the relationship between pupil size and the distance of light source to realize line of sight estimation. When the human pupil faces a close light source, the distance between the reflected light spot and the center of the pupil is closer, otherwise, the distance between the two will become farther. This method uses an infrared light source to illuminate the user's eyes, and the cornea will reflect the infrared light source. The corresponding line of sight direction is estimated by calculating the position between the center of the user's pupil and the light spot. Early methods based on pupil-corneal reflection require users to maintain a fixed head posture in a single light source environment, and later gradually expanded to handle the situation where the user's head posture can change freely in a multi-light source environment. For example, Ohno et al. proposed a line of sight estimation method to solve the problem of free head movement using three cameras and an infrared light source. Yoo et al. used the characteristic of the invariant cross ratio of the feature space to allow a wide range of head movement while ensuring the accuracy of line of sight estimation. The method based on pupil-corneal reflection is the mainstream direction of early line of sight estimation methods. Many built-in algorithms in commercial eye trackers are such methods, such as Tobii, SMI, etc. However, such methods generally rely on high-definition cameras, multiple active light sources, complex calibration processes, strict experimental conditions, etc., and are difficult to promote on a large scale.
[0011] (2) Line of sight estimation based on appearance
[0012] The appearance-based line of sight estimation method adopts a machine learning strategy to learn appearance features from face or eye images, and establish a mapping relationship between images and line of sight to perform line of sight estimation. Compared with the model-based method, it has the following advantages: 1) The input samples are flexible and do not require complex preprocessing. This type of method directly takes the eye or face image as input, and the output is the corresponding line of sight direction, without the need to detect and extract features such as reflected light spots, pupils, and irises; 2) Low hardware requirements and easy to promote. This type of method does not rely on additional active light sources and high-definition cameras, and only requires an ordinary webcam to complete line of sight estimation. Therefore, it has great application potential on mobile devices and is easier to promote to daily life; 3) Low requirements on image quality. This type of method is not limited to high image resolution and fixed lighting conditions, and has great potential in processing low-quality images. Therefore, the appearance-based method has gradually become the mainstream method for line of sight estimation. Traditional appearance-based gaze estimation methods first extract manual features from the image (such as gradient direction histograms, etc.), and then design a regression model to establish a mapping relationship from image features to gaze directions (such as local linear interpolation, adaptive linear regression, and Gaussian process regression).
[17] wait).
[0013] In recent years, with the development of deep learning technology, apparent line of sight estimation methods using deep neural networks are rapidly emerging. Such methods use deep learning networks to extract multi-level abstract line of sight direction features in images, and at the same time establish a nonlinear mapping relationship from image to line of sight direction. Through a large number of labeled samples, the model is optimized and trained to learn an effective line of sight estimation model.
[0014] 1) Gaze estimation methods based on eye images. Zhang et al. used deep neural networks for gaze estimation for the first time and proposed a deep gaze estimation model that cascades head posture and eye feature information: GazeNet. The core network used in this model is derived from the LeNet network. Cheng et al. proposed an asymmetric gaze estimation method based on binocular gaze. On the basis of the monocular gaze estimation model, the complementary relationship between the two eyes was considered, further improving the accuracy of the gaze estimation model. Liu et al. proposed a differential gaze estimation network, which eliminates the influence of personalized information by estimating the gaze difference between two eye images of the same person to achieve high-precision gaze estimation. Guo et al. proposed a method based on prediction consistency to learn gaze representation. This method uses the gaze direction of the training sample to linearly represent the predicted value of the test sample, and generalizes the linear combination to the embedding space to improve the generalization ability of the model for the test sample. Yu et al. proposed a gaze estimation method based on gaze redirection sample augmentation. The core of this method is to synthesize a large number of gaze redirection samples through existing reference samples to provide additional training data for the model to improve the accuracy of the model. Sun et al. proposed an encoding framework based on unsupervised learning. This method separates the gaze-related features from the eye appearance features in the eye image through feature-level cross-reconstruction, and can accurately predict the gaze direction. The above methods all use eye images as the input of the model and have achieved relatively ideal prediction results.
[0015] 2) Gaze estimation method based on full-face image. In order to make full use of the additional information of other facial regions, researchers have proposed a gaze estimation method based on the whole face. Zhang et al. proposed a deep gaze estimation model based on spatial weighting, which can flexibly suppress or enhance the information of different facial regions, thereby improving the estimation accuracy. Park et al. proposed a meta-learning-based gaze estimation framework: FAZE, which separates feature representations representing changes in appearance, gaze, and head posture from face images, and then uses a small number of calibration samples to adaptively adjust the gaze estimation model for specific objects. Abdelrahman et al. improved the accuracy of gaze estimation under unconstrained conditions by constraining the angles of each gaze direction.
[0016] 3) Gaze estimation methods based on eye and face images. There are also many methods that use both face and eye images as model inputs. Krafka et al. proposed a four-branch gaze estimation model, including: face image branch, left eye image branch, right eye image branch, and grid image branch. After multiple layers of convolution and feature fusion, the four branches finally output a two-dimensional gaze position. Cheng et al. proposed a coarse-grained to fine-grained strategy to estimate the gaze direction. The core of this method is to refine the basic gaze direction estimated by the face image through the gaze difference predicted by the eye image. Chen et al. proposed a gaze estimator containing void convolution: GEDDNet, which aims to decompose the gaze direction into the sum of the gaze estimate independent of the subject and the deviation dependent on the subject, thereby improving the accuracy of the model between different subjects.
[0017] 4) Gaze estimation methods with the help of attention mechanisms. In order to reduce the impact of gaze-irrelevant factors on model estimation, researchers have proposed using different forms of attention mechanisms to make the model pay more attention to features related to gaze estimation. Murthy et al. proposed an attention-based gaze estimation network AGE-Net (Attention-based Gaze Estimation Network), adding an attention branch to the feature extraction branch to perform feature operations to predict the gaze direction, and achieved advanced results. Hu et al. proposed a multi-task gaze focus network that uses an attention layer to guide the formation of facial features and improves the accuracy of gaze point estimation. Wang et al. proposed a capsule network GazeCaps with a self-attention routing mechanism for gaze estimation, representing various facial features as different capsules, and using self-attention routing to dynamically assign attention to different capsules containing important information, reducing the interference of features irrelevant to gaze, and the model achieved state-of-the-art performance.
[0018] First, we analyze the existing technical solutions from the perspective of theoretical research. As real-life scenarios become more complex, some personalized line of sight estimation methods have been proposed to meet the requirements of high accuracy and high robustness in practical applications. As the research deepens, researchers have roughly three directions of thinking to solve the problem of model personalization:
[0019] 1) Modeling personalized information
[0020] This idea is to learn the user's personalized information from existing samples to reduce the line of sight deviation. For example, Zhao et al.
[33] In 2022, a dual-branch personalized gaze estimation network (episode-based personalization network, EbPN) based on subtask learning was proposed to extract common features and personalized features respectively. The training strategy based on subtask learning enables the network to fully extract both aspects of individual user characteristics, and the gaze direction of new individual users can be estimated more accurately without the need for calibration samples.
[0021] In addition to extracting and integrating personalized information or personalized differences, some researchers have also reduced the impact of personalized information by separating and discarding personalized information from visual features and only describing the independent components of the subject. Murthy et al.
[30] A network based on difference layering is proposed to obtain the feature differences between the left and right eyes, and all human-related features are removed based on these differences to help the model obtain better generalization on user individuals that have never been seen. Cheng et al. regard the personalized gaze estimation task as a cross-domain problem, and design a plug-and-play self-adversarial framework to eliminate personalized information such as lighting and identity, thereby improving the domain generalization ability of the model. Zhang et al. introduced consistency constraints on the gaze space and feature embedding space between and within domains, decoupled the features extracted by the model from the personalized difference information, eliminated the influence of gaze-irrelevant factors, and proved the feasibility of the model using cross-dataset experiments.
[0022] 2) Fine-tune using calibration samples
[0023] Fine-tuning with calibration samples requires obtaining calibration samples of new users to update some parameter information of the pre-trained model, and then train a gaze estimation model specific to the user. Krafka et al. collected 13 calibration samples of new users for model fine-tuning, and obtained a model specific to the new user to predict the location of their gaze points. Since fine-tuning with a very small number of calibration samples may lead to model overfitting, this method has a large estimation error when using a small number of calibration samples. The experimental effect will be significantly improved after the number of calibrations increases. In recent studies, Jindal et al. proposed an unsupervised gaze estimation pre-training method using contrastive representation learning, which relies on the invariance of gaze direction under certain appearance changes and the equivariance of camera perspectives. It can show excellent performance when fine-tuning with only a small number of samples in the data set and in cross-dataset evaluation. There is sometimes no strict boundary between the method of fine-tuning with calibration samples and the method of modeling personalized information. Some methods achieve gaze estimation using only a small number of calibration samples or no calibration samples by modeling personalized information or eliminating personalized information.
[0024] 3) Expand the size of the dataset
[0025] Learning from large amounts of data is an important means to improve the generalization performance of deep learning models. However, in real-world gaze estimation applications, it is not practical to obtain hundreds or thousands of calibration samples for new users, so researchers try to solve this problem from a technical level. One main approach is to use gaze redirection technology to generate a large number of calibration images of a certain user, and the model learns from the large number of generated images to improve accuracy. Yin et al. used Neural Radiance Field (NeRF) for head-eye redirection. The model can decouple the face and eyes, and control the attributes of the face, user identity, lighting, and gaze separately. It can edit eye features separately without affecting the overall face image, accurately manipulate head posture and gaze, and increase the coverage of training samples for application scenario data. The results of using the generated data for the training of several basic models prove the effectiveness of new data in improving the generalization performance of the model domain.
[0026] Secondly, the existing technical solutions are analyzed from the perspective of practical application. The devices currently available on the market for line of sight estimation are called eye trackers, which can be divided into desktop eye trackers and wearable eye trackers according to different usage scenarios. Among them, the desktop eye tracker consists of an eye tracker and a computer display screen as a complete set of equipment. The device is usually placed at a certain distance from the subject to monitor his or her eye movements. The computer screen is used to display visual stimulus materials in the form of images, videos, etc., and the eye tracker is used to record the subject's eye movements, including screen gaze point coordinates, gaze duration, gaze heat map, etc. The wearable eye tracker collects the subject's eye movement behavior in a real environment by integrating the eye tracking system and the scene camera on a lightweight frame, such as glasses or helmets, and the subject can move freely. The wearable eye tracker can record the scene seen by the subject and record the eye movement behavior during the observation process.
[0027] Both of the above-mentioned eye tracker devices are based on pupil-corneal reflection technology, which is the most mainstream technology on the market. The basic principle of this technology is: (1) First, use an infrared light source to illuminate the eye. As a non-visible light, infrared light can penetrate the cornea and intraocular media, providing the necessary light source conditions for subsequent image acquisition; (2) With the help of a high-sensitivity camera, the infrared light reflected from the cornea and retina is accurately captured. These reflected lights contain rich information about the structure of the eyeball, providing a data basis for subsequent eye movement analysis; (3) Due to the unique physiological structure and physical properties of the eyeball, the position of the light spot formed by corneal reflection is relatively stable while keeping the relative position of the light source and the head unchanged. This feature allows us to accurately distinguish between corneal reflected light and retinal reflected light, thereby further analyzing the eye movement state; (4) The direction of the light reflected on the retina can clearly indicate the direction of the pupil. When the light from the light source enters the pupil and is refracted by the intraocular medium, the reflected light on the retina will be emitted from the pupil, and its direction directly reflects the orientation of the pupil; (5) The direction of sight is accurately calculated by calculating the angle between the light reflected from the cornea and the light reflected from the pupil. This calculation process is based on the principle of geometric optics, ensuring the accuracy and reliability of the eye movement direction measurement. Summary of the invention
[0028] The embodiments of the present invention provide a personalized line of sight estimation method and system based on causal inference, which are used to solve the problems existing in the prior art.
[0029] In order to achieve the above object, the present invention adopts the following technical scheme.
[0030] Personalized sight line estimation method based on causal inference, including:
[0031] Obtain the internal parameters of the camera through calibration;
[0032] Perform face detection on the image acquired by the camera to obtain the key points of the face;
[0033] Based on the internal parameters, the external parameters of the camera used to characterize the position and posture of the head in the camera coordinate system are obtained by calculation;
[0034] Normalizing the image through internal and external parameters, including perspective transformation and unified coordinate system processing;
[0035] The normalized image is input into a personalized sight line estimation model based on causal inference for processing to obtain the three-dimensional sight line direction in the normalized camera coordinate system; the three-dimensional sight line direction is represented by a pitch angle and a yaw angle;
[0036] Convert the three-dimensional coordinates of the sight direction into two-dimensional coordinates in the image coordinate system.
[0037] Preferably, obtaining the internal parameters of the camera through calibration includes:
[0038] Mount the checkerboard on a flat, non-reflective surface; the checkerboard has a black and white checkerboard pattern with clear boundaries;
[0039] Images of the checkerboard are captured by a camera from multiple angles and distances;
[0040] Preprocessing the checkerboard image, and then extracting the coordinates of the corner points in the checkerboard image;
[0041] Based on the corner point coordinates, the internal parameters are calculated using Zhang’s camera calibration algorithm.
[0042] Preferably, performing face detection on the image acquired by the camera to obtain key points of the face includes:
[0043] Determine whether the image acquired by the camera contains a human face. If so, locate the area of the human face in the image. Otherwise, discard the image. Perform this substep multiple times to complete the processing of all images.
[0044] Set the initial positions of all facial key points in the image;
[0045] Extract local features of all faces in the image;
[0046] The facial key points are obtained by learning the mapping from the current key point position to the actual mark position through cascaded regression trees; each regression tree can correct the prediction error of the previous step and gradually approach the position of the real key point.
[0047] Preferably, the process of obtaining the head posture through internal parameters includes:
[0048] Matching the feature points in the two-dimensional image corresponding to the coordinates of multiple facial key points with the corresponding points in the known three-dimensional facial model;
[0049] Select four non-coplanar feature points from the matched feature points as control points;
[0050] A virtual camera coordinate system is constructed through the control points, and four virtual control points are placed in the virtual camera coordinate system;
[0051] The EPnP algorithm is used to solve the camera's rotation matrix R and translation vector t by minimizing the reprojection error between the projection of the four virtual control points on the image plane and the real detected feature points.
[0052] Preferably, the process of using the EPnP algorithm to solve the rotation matrix R and the translation vector t of the camera includes:
[0053] Based on the four virtual control points, a virtual camera coordinate system is constructed and
[0054]
[0055] Calculate the weighted linear combination of four virtual control points; where X i is the coordinate of point i in the three-dimensional space where the four virtual control points are located, C j are the coordinates of the four control points, α ij is the weighting coefficient used to represent point i;
[0056] Pass-through
[0057]
[0058] Calculate the rotation matrix R; where x c ,y c and z c are the three axis vectors of the camera coordinate system after rotation, ‖x c ‖,‖y c ‖ and ‖z c ‖ are the moduli of the three vectors, used to normalize the vectors, z c is the z-axis of the camera coordinate system after rotation, aligned with the z-axis of the head coordinate system;
[0059] By setting the standard distance d of the camera n Make all images have the same scale and use the formula
[0060]
[0061] Calculate the translation vector t=(t x ,t y ,t z ).
[0062] Preferably, normalizing the image by using internal parameters and external parameters specifically includes:
[0063] By rotating the virtual camera coordinate system, the z-axis of the virtual camera coordinate system is aligned with the z-axis of the head coordinate system, and the rotation matrix R = [x c / ‖x c ‖;y c / ‖y c ‖; z c / ‖z c ‖]; where z c is the z-axis of the camera after rotation, x r is the x-axis of the head coordinate system, y c is the y-axis after rotation, x c is the x-axis of the rotation camera;
[0064] By adjusting the camera so that the distance between the camera and the chessboard is d n , and make all images have the same scale;
[0065] By transforming the matrix
[0066]
[0067] All images are standardized; where M = SR is the overall transformation matrix including rotation and scaling, C r is the projection matrix of the original camera, C n is the normalized camera projection matrix.
[0068] Preferably, the personalized sight line estimation model based on causal inference has a structured causal sub-model and a personalized sight line estimation sub-model, and the structured causal sub-model has: a facial image I variable, a user S variable, a potential user information C variable, and a sight line estimation result G variable;
[0069] I→G. characterizes the causal relationship between sample image features and gaze estimation results. S→I. The user's gaze habits, eye anatomical structure, and surrounding environment lead to specific appearance features in the facial image recorded when the user is gazing. S→C←IC represents potential user information, including general user information and personalized user information. I→G←C. The traditional gaze estimation model aims to estimate the gaze estimation result G variable as accurately as possible.
[0070] The normalized image is input into the personalized sight line estimation model based on causal inference for processing, and the three-dimensional sight line direction in the normalized camera coordinate system is obtained, including:
[0071] Pass-through
[0072]
[0073] Calculate the causal effect of each user; where S = {s i |i=1,…,N s}, f(·) is the given sample image I and the specific user s i represents the function of potential information C;
[0074] The personalized gaze estimation sub-model includes:
[0075] The feature extraction module based on the residual neural network includes five convolutional layers arranged in sequence along the data flow direction. The output of the last convolutional layer is average-pooled and used as the extracted line-of-sight feature.
[0076] The personalized information decoupling module is used to eliminate personalized information irrelevant to the line of sight in the user's apparent features, purify the line of sight related features, and improve the accuracy of the model's line of sight estimation; the processing process of the personalized information decoupling module includes:
[0077] Pass-through
[0078]
[0079] Calculate the prototype of each user; where S is a user dictionary S = {s i |i=1,…,N s},N s is the number of users included in the training data, each is the prototype of user i, N i is the number of samples of user i;
[0080] Pass-through
[0081]
[0082] Calculate the intermediate variable C; where P(s i ) is the user probability, is a learnable mapping matrix;
[0083] Pass-through
[0084]
[0085] Calculate the apparent features after decoupling personalized information; where W1, W2, and W3 are the learnable weights of the model. is the image feature extracted by the backbone network from the sample image, c j It is the potential user information expression in a specific sample image;
[0086] Pass-through
[0087] P(G|do(I))=W3f concat (I) (10) Calculate and obtain the sight line estimation result G variable;
[0088] By minimizing Loss
[0089]
[0090] Update the output of the personalized gaze estimation sub-model.
[0091] In a second aspect, the present invention provides a personalized sight line estimation system based on causal inference, comprising:
[0092] A camera, for capturing images;
[0093] The camera calibration module is used to obtain the internal parameters of the camera through calibration;
[0094] The face and key point detection module is used to: perform face detection on the image acquired by the camera and obtain the key points of the face;
[0095] The head pose estimation module is used to calculate the external parameters of the camera through the head pose obtained by the internal parameters;
[0096] The image normalization module is used to: normalize the image through internal parameters and external parameters, including perspective transformation and unified coordinate system processing of the image;
[0097] The personalized sight line estimation module based on causal inference is used to: input the normalized image into the personalized sight line estimation model based on causal inference for processing, and obtain the three-dimensional sight line direction in the normalized camera coordinate system; the three-dimensional sight line direction is represented by the pitch angle and the yaw angle; and convert the coordinates of the three-dimensional sight line direction into two-dimensional coordinates in the image coordinate system;
[0098] The visualization output module is used to visualize and output the two-dimensional coordinates in the image coordinate system after the transformation.
[0099] It can be seen from the technical solutions provided by the above-mentioned embodiments of the present invention that the present invention provides a personalized line of sight estimation method and system based on causal inference, wherein the method includes: obtaining the internal parameters of the camera through calibration; performing face detection on the image acquired by the camera to obtain face key points; calculating the external parameters of the camera through the head posture obtained by the internal parameters; normalizing the image through the internal parameters and the external parameters, including performing perspective transformation processing and unified coordinate system processing on the image; inputting the normalized image into a personalized line of sight estimation model based on causal inference for processing to obtain a three-dimensional line of sight direction in a normalized camera coordinate system; the three-dimensional line of sight direction is represented by a pitch angle and a yaw angle; and converting the coordinates of the three-dimensional line of sight direction into two-dimensional coordinates in the image coordinate system. The method and system provided by the present invention have the following advantages: in view of the prediction bias caused by personalized information in the line of sight estimation task, a novel causal inference-based method is proposed, which extracts the independent appearance features of the subject of each image through causal intervention, eliminates irrelevant interference personalized information, and improves the adaptability of the line of sight estimation model to new users; realizes unconstrained personalized line of sight estimation, the acquisition process is comfortable, and the equipment is highly portable; the cost is low, and it can be deployed only by relying on a desktop computer; the operation is simple and convenient; the line of sight estimation algorithm based on this platform has high accuracy, strong robustness, and strong generalization. It can accurately predict the line of sight direction of different users in a variety of environments. And it does not require high head and body posture of the subject.
[0100] Additional aspects and advantages of the present invention will be given in part in the following description, which will become obvious from the following description, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0102] Figure 1 A processing flow chart of the personalized sight line estimation method based on causal inference provided by the present invention;
[0103] Figure 2 A schematic diagram of a system interface of a personalized sight line estimation method based on causal inference provided by the present invention;
[0104] Figure 3 A schematic diagram of camera calibration for the personalized sight line estimation method based on causal inference provided by the present invention;
[0105] Figure 4 A schematic diagram of the image normalization process of the personalized sight line estimation method based on causal inference provided by the present invention;
[0106] Figure 5 A schematic diagram of a personalized sight line estimation model based on causal inference of the personalized sight line estimation method based on causal inference provided by the present invention;
[0107] Figure 6 A schematic diagram of real-time visual estimation of the personalized line of sight estimation method based on causal inference provided by the present invention. DETAILED DESCRIPTION
[0108] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.
[0109] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.
[0110] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless defined as herein.
[0111] To facilitate understanding of the embodiments of the present invention, several specific embodiments will be further explained below with reference to the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.
[0112] The present invention provides a personalized line of sight estimation method and system based on causal inference, which are used to solve the following technical problems existing in the prior art:
[0113] 1: Disadvantages of existing eye tracker equipment.
[0114] (1) Ambient light has a significant impact on the performance of the eye tracker. In certain scenarios, if the light is too strong or too weak, the eye tracker's measurement results may be inaccurate. In addition, the eye tracker is also sensitive to the subject's head movement, and even a slight movement of the head may cause data deviation.
[0115] (2) The comfort and portability of eye trackers are also disadvantages. Fixed desktop eye trackers require subjects to maintain consistent head and body postures for a period of time, and no large head or body movements are allowed. This usage condition limits its scope of application. Head-mounted eye trackers can even cause discomfort to subjects when worn. Therefore, the above disadvantages severely limit the use of existing eye tracker equipment in a variety of application scenarios.
[0116] (3) Cost is also an important factor limiting the popularity of eye trackers. At present, high-precision eye tracker equipment is expensive, ranging from thousands to hundreds of thousands of yuan, which makes it difficult for many ordinary consumers and research institutions to afford it, thus limiting the widespread application of eye trackers.
[0117] (4) The operation and use of existing eye trackers require guidance and training from professionals, and are therefore not suitable for widespread use in daily life.
[0118] Compared with the shortcomings of the above-mentioned eye trackers, the present invention provides an intelligent line of sight acquisition platform based on an unconstrained apparent line of sight estimation method. That is, through the built-in camera of the device (laptop), the user's facial image is collected, and the user's line of sight direction is automatically predicted in real time based on the image. The algorithm proposed in the invention can be well adapted to users in different environments, and is also relatively robust to head postures. The invention can be used after being installed on a laptop equipped with a webcam, without the user having to wear any additional equipment, so its comfort and cost price are far superior to eye tracker equipment.
[0119] 2. Disadvantages of existing line of sight estimation algorithms
[0120] As described in the above problem 1, the existing sight line estimation method has a low accuracy when used for new users in unseen scenarios. To address this problem, the present invention proposes a personalized sight line estimation method based on causal inference. Analyzing the technical points, the shortcomings of the existing sight line estimation method can be summarized as follows.
[0121] The influence of personalized factors has not been eliminated. There are personalized differences between individual users. The model learned by training users does not eliminate the influence of personalized factors. The extracted line of sight related features often embed the personalized information of these users, which makes the line of sight estimation model easily overfit to the personalized features of the training users, resulting in deviations when applied to new users.
[0122] In view of the above shortcomings, the present invention first proposes a personalized line of sight estimation algorithm based on causal inference, then designs an intelligent line of sight estimation platform, and deploys the algorithm on the platform, thus completely realizing the entire line of sight estimation process. Users can collect face images based on computer cameras and track the line of sight direction in real time. The main contents of the present invention are as follows:
[0123] In order to solve the influence of personalized information on the model, thereby purifying the line of sight related features and preventing the model from overfitting with the personalized information in the training samples. The present invention proposes a personalized feature decoupling method based on causal inference, adopts the structured causal model (SCM) in the causal inference theory, establishes a directed acyclic graph to analyze the complex causal relationship between the image, line of sight and the user's individual personalized information, and regards the personalized information as a confusing factor that affects the causal relationship between the image features and the line of sight features. On the basis of the causal model, the backdoor adjustment criterion is used to remove the influence of personalized information. Specifically, the model considers the features of all users on average through user weights and user dictionaries, extracts the main independent appearance features of each image, eliminates irrelevant interfering personalized information, and improves the adaptability of the line of sight estimation model to new users.
[0124] The present invention aims to capture a face image based on a network camera and predict the three-dimensional sight direction thereof. The system interface is as follows Figure 3 As shown. Since the user is not subject to any constraints when collecting images, the distance between the user and the screen, the user's head posture, the light intensity, etc. in each image may be different. In order to eliminate the influence of these factors on the prediction accuracy of the line of sight estimation model, the present invention provides a personalized line of sight estimation method based on causal inference. Figure 1 The flow chart is shown as follows:
[0125] The first step is to obtain the camera intrinsic parameters through camera calibration, which are used to estimate the head posture later;
[0126] The second step is to perform face detection on the collected image, and then perform face key point detection when a face can be detected from the image;
[0127] The third step is to obtain the external parameters of the camera for characterizing the position and posture of the head in the camera coordinate system through calculation based on the internal parameters;
[0128] The fourth step is to normalize the image using the internal and external parameters of the camera, perform perspective transformation on the image, and convert it to a unified camera coordinate system;
[0129] The fifth step is to input the normalized image into a personalized line of sight estimation model based on causal inference to obtain the three-dimensional line of sight direction in the normalized camera coordinate system, which is represented by the pitch angle and the yaw angle.
[0130] The sixth step is to convert the coordinates of the three-dimensional sight line direction in the camera coordinate system into two-dimensional coordinates in the image coordinate system, and then display the coordinates of the three-dimensional sight line direction and / or the two-dimensional coordinates on a visualization device such as a screen. The displayed result is the attention direction of the subject or the monitored person.
[0131] The application environment (purpose) of the obtained two-dimensional coordinates of the line of sight includes: improving and optimizing the architectural design and layout; obtaining user preferences to improve the recommendation algorithm or optimize the product layout; monitoring the driver's attention to reduce traffic accidents; monitoring the behavior of production personnel to improve production efficiency and reduce safety accidents.
[0132] In the preferred embodiment provided by the present invention, the specific execution process of each step is as follows.
[0133] Considering that the camera parameters of different computers may be different, this system first uses the chessboard calibration algorithm to obtain the camera internal parameters of the computer camera. The schematic diagram of the camera calibration chessboard is as follows Figure 3 As shown, the specific steps are as follows:
[0134] Prepare a checkerboard calibration board: Select a checkerboard with clear boundaries and alternating black and white as the calibration object. Ensure the calibration board is made with high accuracy and its size and flatness meet the calibration requirements. Common checkerboard specifications include 8×8 or 9×6 (e.g. Figure 3 shown).
[0135] Fix the checkerboard calibration board: Fix the checkerboard on a flat, non-reflective surface to ensure that the calibration board does not move or deform during the calibration process.
[0136] Take chessboard images: Use a camera to take images containing chessboards from multiple different angles and distances to ensure that the chessboards are clearly visible in the image and cover the entire image area. These images will be used for subsequent corner point extraction and parameter calculation.
[0137] Corner point extraction: Preprocess each chessboard image, including grayscale, denoising and other operations, to improve the accuracy of corner point extraction. Then, use the Harris corner point detection algorithm of the cv2.findChessboardCorners function to extract the coordinates of the corner points in the chessboard image.
[0138] Parameter calculation: Using the extracted corner coordinates, the camera's intrinsic parameters are calculated through Zhang's camera calibration algorithm, which can be implemented by calling the cv2.calibrateCamera function. In this process, the algorithm uses checkerboard images from multiple perspectives to optimize the estimation of intrinsic parameters by minimizing the reprojection error.
[0139] When performing face detection on the captured image, and then performing face key point detection when a face can be detected from the image, first determine whether the continuous image frames captured by the camera / camera contain a face, and accurately locate the face area in the image. Image frames where no face is detected are directly discarded, and then the next frame of the image is determined. Before face detection, the image is first dedistorted by calling the undistort() function in the cv2 library combined with the internal parameters of the camera and the camera's distortion matrix, so as to calibrate the image to restore the authenticity of the image. When performing face detection on an image, the face detection function get_frontal_face_detector() in the dlib library is used for implementation. This function uses a face detection algorithm based on HOG (Histogram of Oriented Gradients) and SVM (Support Vector Machine). First, the image is converted into a grayscale image; then the HOG method is used to extract the gradient information of the image as image features; then an SVM classifier is trained to distinguish between faces and non-faces; secondly, a window is moved on the image, and the SVM model is used to detect whether there is a face in the window; finally, overlapping detection results are merged through non-maximum suppression, and only the best detection window is retained.
[0140] After ensuring that the image contains a face, perform face key point detection. The shape_predictor (shape_predictor_68_face_landmarks.dat) function in the dlib library is used to detect face key points. This function can obtain the positions of 68 face key points. Finally, we only retain the coordinates of six of the key points, which are the two eye corners of the left and right eyes and the two mouth corners. This function uses a face key point detection algorithm based on a cascade regression tree. The detailed steps are as follows: First, set an initial estimated position for each key point to be detected, that is, the average position of all key points; then for each face in the image, extract local features; next, use a series of regression trees to learn the mapping from the current key point position to the actual landmark position. Each tree corrects the prediction error of the previous step. The structure of the cascade regression tree continuously optimizes the key point position through multiple iterations, so that the final detection result gradually approaches the true key point position. During the training process, each tree learns a regression function that predicts an update vector (i.e., position correction) based on the predicted position of the current key point and local image features; finally, in the detection phase, these trees are organized into a cascade framework. The cascade regression tree model applies the predictions of each tree in sequence to gradually update the position of each key point. After each tree, the predicted position of the key point is updated to be close to the true position. In this way, the model refines the position of the key point at each cascade level. After the predictions of all trees, the final key point position obtained is the final prediction of each key point of the face.
[0141] In the third step, the EPnP (Efficient Perspective-n-Point) algorithm is used to determine the rotation and translation of the head relative to the camera, thereby obtaining the head pose and providing parameters for subsequent image normalization. Specifically, the camera pose is calculated by calling the solvePnP() function in the cv2 library. The main goal of the EPnP algorithm is to estimate the camera pose by matching points in three-dimensional space (such as facial feature points) with their projections in two-dimensional images. The following are the basic steps for estimating head pose using the EPnP algorithm:
[0142] Feature point detection and matching: First, the coordinates of the six key points detected in the face image are input, and then the feature points in these two-dimensional images are matched with the corresponding points in the known three-dimensional face model.
[0143] Initialization: Select four non-coplanar points from the matched feature points as control points. These four points will be used for subsequent EPnP calculations.
[0144] Construct virtual control points: Use the four control points in the real three-dimensional space to construct a virtual camera coordinate system and place four virtual control points in it.
[0145] Solution: By minimizing the reprojection error between the projection of the four virtual control points on the image plane and the real detected feature points, the camera's rotation matrix R and translation vector t are solved. This process involves complex mathematical optimization. The EPnP algorithm reduces the computational complexity through clever methods, making it very efficient in practical applications.
[0146] Head pose estimation: Once the rotation matrix R and translation vector t are obtained, the pose of the head can be estimated. The rotation matrix R represents the rotation direction of the head relative to the camera, while the translation vector t represents the position of the head relative to the camera.
[0147] After selecting four non-coplanar control points, the EPnP algorithm uses these four real control points in three-dimensional space to construct a virtual camera coordinate system. The points in this three-dimensional virtual coordinate system are weighted linear combinations of these four control points, namely:
[0148]
[0149] Where X i is the coordinate of point i in three-dimensional space, C j are the coordinates of the four control points, α ij is the weighting coefficient used to represent point i, which is uniquely determined by the set position of the point. Specifically, the weight α of the control point ijIt is calculated by analytical geometry. These weights reflect how each point is distributed in the virtual coordinate system. Through this representation, all points in the three-dimensional space can be based on these four control points, that is, these four control points can be used to replace all three-dimensional points for optimization calculation, which simplifies the calculation of complex three-dimensional point sets.
[0150] In actual calculation, these four control points are mapped to the image plane, and the corresponding two-dimensional points are obtained through the camera projection model. In this process, the rotation matrix R and translation vector t are solved by minimizing the reprojection error between the two-dimensional projections of the four control points and the real feature points.
[0151] The fourth step above aims to adjust all images to a consistent perspective and scale before entering the gaze estimation model, so that the model can focus more on learning the changes in gaze direction and greatly reduce the impact of head posture on the model prediction accuracy. Figure 4 As shown in Figure 2, it mainly includes two key steps: camera alignment and image normalization.
[0152] The following are the detailed steps and calculation formulas involved:
[0153] 1. Camera alignment: This step mainly rotates and scales the camera coordinate system to align it with the head coordinate system, thereby reducing the impact of different head postures on the line of sight estimation model.
[0154] Rotation:
[0155] The rotation matrix R is used to rotate the original camera coordinate system to align with the head coordinate system, unify the perspective, and solve the visual differences caused by head rotation. The rotation matrix R describes how to rotate from one coordinate system to another. Each column of the matrix represents the new direction of an axis of the original coordinate system in the target coordinate system.
[0156] First, the camera coordinate system needs to be rotated so that its z-axis is aligned with the z-axis of the head coordinate system (usually the axis facing backwards, perpendicular to the plane of the face). The rotation matrix R obtained in the head pose estimation module is:
[0157]
[0158] where z c is the z-axis of the rotated camera, which needs to be aligned with the z-axis of the head coordinate system. r is the x-axis of the head coordinate system, y c is the rotated y-axis, defined as y c =z c × r .x c is the x-axis of the rotation camera, defined as x c=y c × c .
[0159] The rotation matrix R maps the axes of the original coordinate system to the new coordinate system so that the original camera coordinate system is aligned with the target head coordinate system, as Figure 4 (b) This means that we use the head coordinate system as a unified reference coordinate system. When processing images with different head postures, the rotation matrix R can "rotate" the facial feature points back to a unified frontal head position. After this rotation operation, no matter how the head posture in the original image changes, the result is always a frontal standardized perspective.
[0160] Zoom:
[0161] After rotating the camera to see the origin of the head coordinate system, it is also necessary to adjust the camera to a fixed distance d n , so that all images have the same scale. The scaling matrix S is usually a diagonal matrix with diagonal elements d n / ‖t‖, where ‖t‖ is the distance from the center of the eye to the original position of the camera.
[0162] Translation vector t=(t x ,t y ,t z ) represents the translation position of the head in the camera coordinate system, which is used to adjust the viewing distance of the image to ensure that the relative position and proportion of the head in all images are consistent in order to unify the scale. Specifically, the camera needs to be adjusted to a fixed standard distance d n , so that all images have the same scale, that is, Figure 4 (c) The image is adjusted using the following calculation formula of the scaling matrix S:
[0163]
[0164] in, is the distance from the camera to the center of the eye.
[0165] 2. Image Standardization
[0166] After the cameras are aligned, the next task is to adjust the images to eliminate as much variation in image appearance as possible due to different head poses. We process the images by perspective warping, that is, regardless of the head pose in the original image, all images are processed as if they were taken from a fixed distance d. n It can be obtained by observing from different angles.
[0167] Transformation Matrix
[0168]
[0169] Where M = SR is the overall transformation matrix including rotation and scaling, representing the camera alignment process, C r is the projection matrix of the original camera, describing the focal length, principal point and other internal parameters of the original camera, C n It is the normalized camera projection matrix and the camera's internal parameter, which projects the three-dimensional space onto a two-dimensional plane.
[0170] In summary, the role of camera extrinsics is mainly reflected in the perspective transformation and coordinate system process, which is used to control the rotation of viewing angle and the unification of distance. After image standardization, each image enters the line of sight estimation model with the same observation conditions.
[0171] In the preferred embodiment provided by the present invention, the present invention proposes a personalized line of sight estimation method based on causal inference, which mines the true causal relationship between users, images and line of sight, thereby eliminating the influence of model priors caused by training user personalized information, purifying line of sight related features, and improving the estimation accuracy of the model when facing new users. The framework diagram of this method is shown in Figure 5 As shown, it mainly consists of two parts: structural causal sub-model ( Figure 5 The left sub-image a) and the personalized gaze estimation sub-model based on causal inference ( Figure 5 The right sub-figure b).
[0172] 1. Structural Causal Submodel
[0173] In order to explain the cause of the personalized estimation bias problem, the present invention constructs a structured causal model to illustrate the causal relationship between variables in the line of sight estimation model. Figure 6 As shown on the left, the gaze estimation causal model in this paper contains four variables (processing factors), namely facial image I (image), user S (subject), potential user information C (customized information) and gaze estimation result G (gaze). The directed line represents the causal relationship between two nodes, that is: cause → effect.
[0174] I→G. The causal relationship between sample image features and gaze estimation results. Ideally, the model makes decisions based on the causal relationship between the two.
[0175] S→I. The user's gaze habits, eye anatomy, surrounding environment, etc. result in specific appearance features in the facial images recorded when the user is looking. Even if the gaze direction is the same or the gaze position is the same, the features presented in the image by different users are different.
[0176] S→C←IC represents potential user information, including general user information and personalized user information. General user information is determined by basic features common to all users, such as the basic structure of the eye. However, due to differences in personalized information such as the user's environment, gaze habits, appearance, and κ angle, the apparent features displayed in a single sample of each user include not only general gaze features, but also gaze features specific to the current user and the current image. Such a causal relationship causes similar sample images to extract different sight features because they belong to different individual users. Different sample images belonging to the same individual user may also result in differences in the extracted personalized features. The difference in personalized features ultimately leads to deviations in the prediction results. This type of deviation can be effectively eliminated by the above formula.
[0177] I→G←C. The traditional line of sight estimation model aims to estimate the line of sight G as accurately as possible. Figure 4 As can be seen from the causal graph in (a), the gaze G is the result of two causal paths, namely: (1) I→G, which means that the gaze estimation model estimates the gaze G based on the appearance features extracted from the input facial image; (2) C→G, which means that the potential user information embedded in the model uses the prior influence in the training data to estimate the gaze. In other words, the traditional gaze estimation model approximates P(Y|I), and the potential user information learned from the training data will implicitly affect the gaze estimation result. Although when the appearance features obtained from the image are not sufficient to determine the gaze direction or gaze point, the existence of C allows the model to use the prior information in the training data to make better estimates, it may also be confused by the training data and erroneously associate or separate certain appearance features with the final estimation result.
[0178] The present invention regards the user S of the sight estimation task as a confounding factor, which confuses the relationship between I and G through the path I←S→C→G. In this path, even if there is no causal relationship between I and G, they are still connected through S and C, resulting in a false mapping between the image and the sight. In order to eliminate the adverse effects of the confounding factor S, according to the theory of causal inference, intervention is performed by applying the do(·) operation to the variable I. As reflected in the structured causal model graph, the do(·) operation deletes all arrows entering I, and the causal relationship from S to I is cut off, thus obtaining a sight estimation model that approximates P(Y|do(X)) rather than P(Y|X).
[0179] The most direct way to intervene in I is to collect any facial image of any user for a randomized controlled experiment. However, since the number of users and facial images is infinite in the gaze estimation task, it is impossible to implement a randomized controlled experiment. Therefore, the present invention uses a back-door adjustment formula and uses statistical language to describe this intervention, as shown in formula (5). Specifically, the influence of S is decoupled, and the causal effect of each user in the training data is first estimated. Then, the weighted average is calculated based on the sample proportion of each user in the training data to estimate the average causal effect.
[0180]
[0181] Where S = {s i |i=1,...,N s}, f(·) is the given sample image I and the specific user s i When , it represents the function of the potential information C. By using the proportion of each user’s sample data in the whole, each user is fairly included in the estimation of the line of sight, so that S no longer affects the causal relationship from I to G.
[0182] However, since this formula needs to traverse every pair of I and s i , which requires a relatively large computational overhead, the present invention uses the normalized weighted edgeometric mean (NWGM) for simplification. NWGM allows approximation of the expectation of formula (5) at the feature level, and the computational target of the model can be expressed by formula (6).
[0183]
[0184] 2. Personalized Gaze Estimation Model
[0185] In order to address the problem of personalized information decoupling in line of sight estimation, the present invention introduces causal inference into line of sight estimation, with the aim of eliminating personalized information irrelevant to line of sight in the user's apparent features extracted by the backbone network, purifying line of sight-related features, and improving the accuracy of line of sight estimation. Specifically, the present invention develops a feature extraction module and a personalized information decoupling module based on a residual neural network (Residual Network, ResNet), which respectively extracts the apparent features of the input user image and decouples the personalized information.
[0186] 2.1 Feature extraction module based on residual neural network
[0187] The original ResNet network was proposed for target detection tasks. Taking ResNet18 as an example, the original model outputs a 1000-dimensional one-hot vector to represent the target detection result after receiving the image input. When this paper uses it to extract the line of sight features of the image, the fully connected layer for feature fusion is removed, and the output of the last residual module is averaged and pooled as the extracted line of sight features. The specific structure is shown in Table 1.
[0188]
[0189] Table 1 Feature extraction module based on residual neural network
[0190] 2.2 Personalized Information Decoupling Module
[0191] The personalized information decoupling module aims to eliminate personalized information irrelevant to the line of sight in the user's apparent features, purify the line of sight related features, and improve the accuracy of the model's line of sight estimation. i After the face image I is input, the output of the hidden layer of the backbone network is used as the extracted image feature This feature is used as the input of the personalized information decoupling module. The causal features output by the module and the apparent features extracted by the backbone network jointly determine the line of sight estimation result. The construction of the personalized information decoupling module includes the following steps:
[0192] 1) Construction of causal dictionary. For the confounding variable S, since it is impossible to collect samples from all users, this paper fixes S as a user dictionary S = {s i |i=1,…,N s} to approximate, and update the model after each iteration, where N s is the number of users included in the training data, each is the prototype of user i. In order to calculate the prototype of each user, a prototype container is maintained for each user, which is composed of the output of the backbone network for the user's training sample images, expressed as formula (7).
[0193]
[0194] Where S is a user dictionary S = {s i |i=1,…,N s},N s is the number of users included in the training data, each is the prototype of user i, N i is the number of samples of user i.
[0195] 2) Calculation of intermediate variable C. C is approximately the weighted geometric mean of all user prototypes, and the weight is expressed as the sample image feature of a specific user In user dictionariesi The attention score in is calculated using scaled dot-product attention (ScaledDot-ProductAttention), as shown in formula (8).
[0196]
[0197] The user probability P(s i ) is the ratio of the number of user i samples to the total number of training samples, is a learnable mapping matrix.
[0198] 3) Apparent features after decoupling of personalized information. Figure 5 The causal model and formula (2) shown in the left sub-figure can be approximated using a linear model to calculate the gaze feature embedding after intervention on the user image I, expressed as formula (9).
[0199]
[0200] Where W1, W2, W3 are the learnable weights of the model. is the image feature extracted by the backbone network from the sample image, c j It is the potential user information expression in a specific sample image. The final sight line G is obtained by embedding the feature using linear regression, that is, the sight line G is calculated by formula (10).
[0201] P(G|do(I))=W3f concat (I) (10)
[0202] In summary, with the help of the theory of causal inference, the model can obtain the appearance feature embedding of a specific facial image after decoupling the user information, reducing the interference of the prior information learned by the model in the training data on the line of sight estimation of new users. After obtaining the final appearance features, the model uses a fully connected layer to fuse the causal features with the image features and regress them into the line of sight vector, as shown in formula (10).
[0203] 3. Model training process
[0204] In order to fully learn the gaze-related features in the image, at each iteration, the personalized information decoupling module purifies the gaze-related personalized features in the current sample image based on the existing user prototype, and uses it to update the user causal dictionary; after each round of iteration, the user prototype is updated according to the causal dictionary calculated during the iteration process. The purpose of this network is to decouple the gaze-irrelevant personalized information. Therefore, given a set of samples, the model uses the learned user prototype to eliminate the gaze-irrelevant features, obtain the gaze-related causal features, and calculate the gaze vector after fusing with the image features to minimize The loss (Formula (11)) updates the entire network.
[0205]
[0206] 4. Visualization of sight direction
[0207] Inputting the normalized face image into the above gaze estimation model can obtain the three-dimensional gaze direction (pitch, yaw) in the camera coordinate system. In order to display the three-dimensional gaze direction on the computer screen, it is converted from the camera coordinate system to the two-dimensional image coordinate system. This conversion can be completed by using the cv2.projectPoints() function. Then, by connecting the end point and the starting point of the gaze (i.e. the center coordinates of the key points of the face), the three-dimensional gaze direction can be visualized on the screen. The final result is as follows: Figure 6 shown.
[0208] In a second aspect, the present invention provides a personalized sight line estimation system based on causal inference, comprising:
[0209] A camera, for capturing images;
[0210] The camera calibration module is used to obtain the internal parameters of the camera through calibration;
[0211] The face and key point detection module is used to: perform face detection on the image acquired by the camera and obtain the key points of the face;
[0212] The head pose estimation module is used to calculate the external parameters of the camera through the head pose obtained by the internal parameters;
[0213] The image normalization module is used to: normalize the image through internal parameters and external parameters, including perspective transformation and unified coordinate system processing of the image;
[0214] The personalized sight line estimation module based on causal inference is used to: input the normalized image into the personalized sight line estimation model based on causal inference for processing, and obtain the three-dimensional sight line direction in the normalized camera coordinate system; the three-dimensional sight line direction is represented by the pitch angle and the yaw angle; and convert the coordinates of the three-dimensional sight line direction into two-dimensional coordinates in the image coordinate system;
[0215] The visualization output module is used to visualize and output the two-dimensional coordinates in the image coordinate system after the transformation.
[0216] In summary, the present invention provides a personalized sight line estimation method and system based on causal inference, wherein the method includes: obtaining the internal parameters of the camera through calibration; performing face detection on the image obtained by the camera to obtain the key points of the face; calculating the external parameters of the camera through the head posture obtained by the internal parameters; normalizing the image through the internal parameters and the external parameters, including performing perspective transformation processing and unified coordinate system processing on the image; inputting the normalized image into a personalized sight line estimation model based on causal inference for processing to obtain the three-dimensional sight line direction in the normalized camera coordinate system; the three-dimensional sight line direction is represented by the pitch angle and the yaw angle; converting the coordinates of the three-dimensional sight line direction into two-dimensional coordinates in the image coordinate system. The method and system provided by the present invention design a personalized sight line estimation model that can be transplanted to different devices. Compared with the eye tracker equipment on the market, the invention has low cost, good scalability, and is more comfortable and portable; a personalized sight line estimation method based on causal inference is proposed, which can eliminate the interference of personalized information of training samples on the model. The Structural Causal Model (SCM) in causal inference theory is used to establish a directed acyclic graph to analyze the complex causal relationship between images, sight lines, and user individual personalized information, and the personalized information is regarded as a confounding factor that affects the causal relationship between image features and sight line features. Based on the causal model, the backdoor adjustment criterion is used to remove the influence of personalized information, improve the model's adaptability and generalization ability to new users, and eliminate personalized bias.
[0217] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.
[0218] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.
[0219] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0220] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A personalized line of sight estimation method based on causal inference, characterized in that: include: Obtain the internal parameters of the camera through calibration; Perform face detection on the image acquired by the camera to obtain the key points of the face; Based on the internal parameters, obtaining the external parameters of the camera for characterizing the position and posture of the head in the camera coordinate system by calculation; Normalizing the image by using the internal parameters and the external parameters, including performing perspective transformation and unified coordinate system processing on the image; Inputting the normalized image into a personalized sight line estimation model based on causal inference for processing to obtain a three-dimensional sight line direction in a normalized camera coordinate system; the three-dimensional sight line direction is represented by a pitch angle and a yaw angle; The coordinates of the three-dimensional sight line direction are converted into two-dimensional coordinates in the image coordinate system.
2. The method according to claim 1, characterized in that: The method of obtaining the internal parameters of the camera through calibration includes: Fixing a checkerboard on a flat and non-reflective surface; the checkerboard has a pattern of black and white checkerboards with clear boundaries; Taking images of the checkerboard from multiple angles and distances by a camera; Preprocessing the image of the checkerboard, and then extracting the coordinates of the corner points in the image of the checkerboard; Based on the corner point coordinates, the internal parameters are calculated using Zhang's camera calibration algorithm.
3. The method according to claim 2, characterized in that The method of performing face detection on an image acquired by a camera to obtain key points of a face includes: Determine whether the image acquired by the camera contains a human face. If so, locate the area of the human face in the image; otherwise, discard the image. Perform this sub-step multiple times to complete the processing of all the images. Setting initial positions of key points of all faces in the image; Extracting local features of all faces in the image; The facial key points are obtained by learning the mapping from the current key point position to the actual marked position through cascaded regression trees; each regression tree can correct the prediction error of the previous step and gradually approach the position of the real key point.
4. The method according to claim 3, characterized in that The process of obtaining the head posture through the internal parameters includes: Matching feature points in the two-dimensional image corresponding to the coordinates of the plurality of facial key points with corresponding points in a known three-dimensional facial model; Selecting four non-coplanar feature points from the matched feature points as control points; Constructing a virtual camera coordinate system through the control points, and placing four virtual control points in the virtual camera coordinate system; The EPnP algorithm is used to solve the rotation matrix R and translation vector t of the camera by minimizing the reprojection error between the projection of the four virtual control points on the image plane and the real detected feature points.
5. The method according to claim 4, characterized in that The process of using the EPnP algorithm to solve the camera's rotation matrix R and translation vector t includes: Based on the four virtual control points, a virtual camera coordinate system is constructed and Calculate and obtain the weighted linear combination of the four virtual control points; where X i is the coordinate of point i in the three-dimensional space where the four virtual control points are located, C j are the coordinates of the four control points, α ij is the weighting coefficient used to represent point i; Pass-through Calculate the rotation matrix R; where x c ,y c and z c are the three axis vectors of the camera coordinate system after rotation, ‖x c ‖,‖y c ‖ and ‖z c ‖ are the moduli of the three vectors, used to normalize the vectors, z c is the z-axis of the camera coordinate system after rotation, aligned with the z-axis of the head coordinate system; By setting the standard distance d of the camera n Make all images have the same scale and use the formula Calculate the translation vector t=(t x ,t y ,t z ).
6. The method according to claim 5, characterized in that The normalization process of the image by using the internal parameters and the external parameters specifically includes: By rotating the virtual camera coordinate system, the z-axis of the virtual camera coordinate system is aligned with the z-axis of the head coordinate system, and the rotation matrix R=[x c / ‖x c ‖;y c / ‖y c ‖; z c / ‖z c ‖]; where z c is the z-axis of the camera after rotation, x r is the x-axis of the head coordinate system, y c is the y-axis after rotation, x c is the x-axis of the rotation camera; By adjusting the camera so that the distance between the camera and the chessboard is d n , and making all the images have the same scale; By transforming the matrix All the images are normalized; where M = SR is the overall transformation matrix including rotation and scaling, C r is the projection matrix of the original camera, C n is the normalized camera projection matrix.
7. The method according to claim 6, characterized in that The personalized sight line estimation model based on causal inference has a structured causal sub-model and a personalized sight line estimation sub-model, and the structured causal sub-model has: a facial image I variable, a user S variable, a potential user information C variable, and a sight line estimation result G variable; I→G. is used to characterize the causal relationship between the sample image features and the gaze estimation results; S→I. is used to characterize the user's gaze habits, eye anatomical structure, and surrounding environment, which lead to specific appearance features in the facial image recorded when the user is gazing; S→C←I. is used to characterize potential user information, including general user information and personalized user information; I→G←C. is used to estimate the gaze estimation result G variable; The step of inputting the normalized image into a personalized sight line estimation model based on causal inference for processing to obtain a three-dimensional sight line direction in a normalized camera coordinate system includes: Pass-through Calculate the causal effect of each user; where S = {s i |i=1,…,N s }, f(·) is the given sample image I and the specific user s i represents the function of potential information C; The personalized sight line estimation sub-model includes: A feature extraction module based on a residual neural network includes five convolutional layers arranged in sequence along the data flow direction, and the output of the last convolutional layer is average-pooled to extract the sight line feature; The personalized information decoupling module is used to eliminate personalized information irrelevant to the line of sight in the user's apparent features, purify line of sight related features, and improve the accuracy of the model line of sight estimation; the processing process of the personalized information decoupling module includes: Pass-through Calculate the prototype of each user; where S is a user dictionary S = {s i |i=1,…,N s },N s is the number of users included in the training data, each is the prototype of user i, N i is the number of samples of user i; Pass-through Calculate the intermediate variable C; where P(s i ) is the user probability, W Q , is a learnable mapping matrix; Pass-through Calculate the apparent features after decoupling personalized information; where W1, W2, and W3 are the learnable weights of the model. is the image feature extracted by the backbone network from the sample image, c j It is the potential user information expression in a specific sample image; Pass-through Calculate and obtain the sight line estimation result G variable; By minimizing Loss Update the output of the personalized gaze estimation sub-model.
8. A personalized sight line estimation system based on causal inference, characterized in that: include: A camera, for capturing images; The camera calibration module is used to obtain the internal parameters of the camera through calibration; The face and key point detection module is used to: perform face detection on the image acquired by the camera and obtain the key points of the face; A head posture estimation module, used to calculate and obtain the external parameters of the camera through the head posture obtained by the internal parameters; An image normalization module, used to: perform normalization processing on the image by using the internal parameters and the external parameters, including performing perspective transformation processing and unified coordinate system processing on the image; A personalized sight line estimation module based on causal inference, used to: input the normalized image into a personalized sight line estimation model based on causal inference for processing, and obtain a three-dimensional sight line direction in a normalized camera coordinate system; the three-dimensional sight line direction is represented by a pitch angle and a yaw angle; Convert the coordinates of the three-dimensional sight line direction into two-dimensional coordinates in an image coordinate system; The visualization output module is used to visualize and output the coordinates of the three-dimensional sight line direction and / or the two-dimensional coordinates in the converted image coordinate system.