Sight line estimation method based on three-dimensional point cloud and surface reconstruction technology
By mapping 2D eye images to 3D space and building a gaze ball model, the problem of invisibility of 3D coordinates in the traditional line of sight estimation method is solved, and the prediction accuracy and adaptability of line of sight direction are improved.
Patent Information
- Application Number
- CN202510053058.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
The traditional appearance-based line of sight estimation method is unable to directly observe 3D coordinates, resulting in limited prediction accuracy of line of sight direction.
Using a line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology, a gaze sphere model is constructed to more accurately predict the line of sight direction by mapping the geometric eye model in a 2D image into a 3D space.
It improves the accuracy and efficiency of line of sight estimation, can better adapt to individual differences and environmental changes, and alleviates the troubles caused by invisibility of the center of the eye.
Smart Images

Figure CN120014160A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of image processing and pattern recognition, and in particular relates to a sight line estimation method based on three-dimensional point cloud and surface reconstruction technology. Background Art
[0002] Gaze estimation is an important technology in the field of computer vision, which aims to predict the gaze direction of an individual by analyzing eye information. Gaze estimation plays a key role in many practical application scenarios, such as intelligent driving and human-computer interaction. Traditional appearance-based gaze estimation methods directly predict 3D gaze direction from 2D eye images. However, due to the limitation that 3D coordinates, such as the position of the pupil and the center of the eyeball, cannot be directly observed in 2D images, this affects the accuracy of gaze estimation. To solve the above problems, we propose a gaze estimation method based on 3D point cloud and surface reconstruction technology. This method uses 3D point cloud to accurately map the geometric eyeball model in the 2D image into 3D space, and constructs a gaze ball model, so as to predict the gaze direction more intuitively and efficiently. Summary of the invention
[0003] The present invention proposes a line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology to solve the problems mentioned in the above background technology. The technical solution steps of the present invention are as follows:
[0004] Step 1: Preprocess the data in the public gaze estimation datasets MPIIGaze, EYEDIAP, and Columbia to obtain the processed eye images and their corresponding real gaze spheres;
[0005] Step 1.1: Convert all monocular images in the dataset into grayscale images, and adjust the converted grayscale images to a fixed size of 72*120 pixels;
[0006] Step 1.2: Use 3D point cloud and surface reconstruction technology to generate a real gaze sphere from the labels of the real sight direction in the dataset;
[0007] Step 1.2.1: The fixation ball is composed of two spheres, one simulating the eyeball and the other iris. The diameter ratio of the two spheres is 2:1. In terms of spatial relationship, the center of the sphere simulating the iris is located at the edge of the sphere simulating the eyeball.
[0008] Step 1.2.2: The real sight direction is represented by a 3-element unit vector in 3D space. The center coordinate of the eyeball is taken as the origin by default, and the position of the entire gaze ball model is determined based on this.
[0009] Step 1.2.3: Use the Open3D tool to generate a real gaze ball, construct a point cloud model of the gaze ball, and use surface reconstruction technology to obtain mesh representations at different angles from the dense point cloud to build an accurate model;
[0010] Step 2: Design and build a deep learning model for gaze estimation, input the eye image and the gaze ball corresponding to the true gaze direction into the initialized model, calculate the optimization loss, and update the model parameters for supervised training;
[0011] Step 2.1: Use the encoder based on Convolutional Neural Network (CNN) as the feature extractor of monocular grayscale image, and use Continuous Normalizing Flow (CNF) to map the eye information and the corresponding real fixation ball to the latent space. Through the forward reversible transformation of CNF, sample and reconstruct multiple candidate fixation balls from the latent space. The optimization target formula in this step is as follows:
[0012] L cnf =L recon +βL ent (1)
[0013]
[0014]
[0015] In formula (1), L recon is the reconstruction loss, L ent is the Gaussian entropy loss, β>0 is the regularization parameter; p in formula (2) θ (y|z,x) represents the probability of generating the fixation ball y under the condition of the latent variable z and the eye image feature x. represents the inverse mapping of the continuous normalized flow (CNF), which is used to map the fixation sphere y to the latent space to obtain the latent variable z, Involving the integral calculation in the CNF process, it is used to measure the reconstruction error; Q in formula (3) φ (ω|x) is the encoder used to learn the eye image condition information, d is the dimension of the latent variable, σ i is the standard deviation of the diagonal Gaussian distribution;
[0016] Step 2.1.1: Use CNN to extract the features of the monocular grayscale image as conditional information;
[0017] Step 2.1.2: Construct a continuous normalized flow CNF to learn the mapping from simple distribution to complex distribution, input the conditional information and the real gaze ball information into the CNF, map the data to the latent space through the forward reversible transformation, and obtain the latent variables;
[0018] Step 2.1.3: Use the inverse transform sampling of CNF to transform the latent variables back to the data space, reconstruct and generate multiple candidate fixation balls, the goal is to learn the representation from the eye image to the possible fixation ball. The relevant calculation formula is as follows:
[0019] y=G θ (z; x) (4)
[0020]
[0021] Among them G θ (·,·) is the CNF transformation formula, y is the candidate fixation ball, z is the latent variable, x is the eye image feature, is the prior distribution, q φ (z|y,x) is the approximate posterior distribution, is the expected term, D KL This is the KL divergence calculation formula.
[0022] Step 2.2: Use the gaze sphere attention fusion module (GSAF) to adaptively fuse multiple candidate gaze spheres to obtain dense gaze spheres to improve the accuracy and stability of gaze estimation;
[0023] Step 2.2.1: For multiple candidate fixation balls, first use the farthest point sampling (FPS) algorithm to downsample each candidate fixation ball to obtain a new candidate fixation ball;
[0024] Step 2.2.2: Calculate the confidence C of each fixation ball through two learnable matrices W1 and W2 and the Sigmoid function i , select the TopK fixation balls with high confidence according to the confidence. The formula is as follows:
[0025] C i =Sigmoid(W2Pool(W1Sample(y i ))) (6)
[0026] Where Sample(·) represents the farthest point sampling (FPS) algorithm, and Pool(·) represents the pooling operation;
[0027] Step 2.2.3: Construct the selected K high-confidence candidate fixation spheres into a dense and more representative fixation sphere. The formula is as follows:
[0028]
[0029] in For the final intensive look at the ball;
[0030] Step 2.2.4: Input the constructed dense gaze ball into the PointMLP network to predict the final gaze direction
[0031]
[0032] Step 2.3: The final gaze direction prediction adopts two strategies: 1) Point cloud fusion (PCF) inputs the fused dense gaze ball into the PointMLP network to obtain the final predicted gaze direction; 2) Gaze direction fusion (GDF) inputs multiple candidate gaze balls into the PointMLP to obtain different prediction results, calculates the weight of each result and then weightedly fuses them to obtain the final predicted gaze direction. The optimization target formula is as follows:
[0033]
[0034] Using L2 loss, where g is the true view direction, is the gaze direction predicted by the model;
[0035] Step 2.3.1: For multiple candidate gaze balls, first input each candidate gaze ball into the PointMLP network to obtain multiple different gaze directions g i ;
[0036] Step 2.3.2: Also use the farthest point sampling (FPS) algorithm to downsample the candidate fixation sphere to obtain a new candidate fixation sphere, and then calculate the weight A of each gaze direction through two learnable matrices M1 and M2 and the SoftMax function. i :
[0037] A i =SoftMax(M2Pool(M1Sample(y i ))) (10)
[0038] Step 2.3.3: Perform weighted fusion based on the calculated weights and gaze directions to obtain the final predicted gaze direction.
[0039]
[0040] Where N is the number of candidate fixation balls; the overall optimization goal of the above steps is:
[0041] L=λL cnf +L g (12)
[0042] Where λ is the regularization hyperparameter;
[0043] Step 2.4: Calculate the loss between the predicted sight direction and the real sight direction after fusion, and use the optimization algorithm to update the model training parameters to minimize the loss and improve the performance of the model. The sight estimation model sets its initial learning rate to 5e-4, and the parameter update uses the Adam optimizer;
[0044] Step 3: Use the trained model as the gaze estimation model, input the eye images of the preprocessed test data in the dataset into the model for prediction, obtain the prediction results of the gaze direction, and calculate the relevant evaluation indicators.
[0045] Step 3.1: Input the pre-processed eye images from the public dataset into the trained model;
[0046] Step 3.2: Use the same convolutional neural network-based encoder, CNF module, and GSAF module as in the training phase to extract the conditional information in the eye image, reconstruct multiple candidate fixation balls, and select different fusion strategies;
[0047] Step 3.3: Use the trained PointMLP network to predict the gaze direction, obtain the final predicted gaze direction, compare it with the true gaze direction, and calculate the performance index to evaluate the prediction performance of the model;
[0048] A line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology, characterized by comprising a preprocessing module, a gaze ball generation module, a gaze attention fusion module, and a line of sight direction prediction module:
[0049] Preprocessing module: Convert the input monocular image into a grayscale image and adjust it to a fixed size of 72*120 pixels; at the same time, use the 3D point cloud and surface reconstruction technology to construct a gaze ball based on the eyeball geometry model, including simulating the 3D structure of the eyeball and iris;
[0050] Fixation sphere generation module: It uses continuous normalized flow CNF to sample from the latent space and convert the original eye image into multiple candidate fixation spheres;
[0051] Gaze Attention Fusion Module: It includes two fusion mechanisms: 1) Point Cloud Fusion (PCF), which evaluates the quality of each candidate gaze ball and selects the K gaze balls with the highest confidence for fusion; 2) Gaze Direction Fusion (GDF), which predicts the gaze of multiple generated candidate gaze balls and performs weighted fusion;
[0052] Gaze direction prediction module: Based on the different sampling mechanisms in the gaze attention fusion module, the PointMLP model is used to predict the final gaze direction.
[0053] The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology of the present invention has the following beneficial effects:
[0054] The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology proposed in the present invention mainly solves the problem of limited accuracy of line of sight estimation due to the invisible center of the eyeball and the complex structure of the eye in the field of line of sight estimation. Traditional methods are difficult to accurately obtain 3D coordinates from 2D eye images, which in turn affects the accurate prediction of the line of sight direction. The method of the present invention introduces the innovative representation of the gaze sphere, uses the three-dimensional point cloud to map the 2D image to the 3D space, and reconstructs the eye geometry model, which effectively alleviates the problem caused by the invisible center of the eyeball, and can well adapt to the influence of individual differences and environmental changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flow chart of the line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology of the present invention;
[0056] Figure 2 This is a model framework diagram of the line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology of the present invention;
[0057] Figure 3 This is a framework diagram of a gaze ball attention fusion module for line of sight estimation based on three-dimensional point cloud and surface reconstruction technology of the present invention; DETAILED DESCRIPTION
[0058] The present invention is further described in detail below in conjunction with preferred embodiments. More details are elaborated in the following description to facilitate a full understanding of the present invention. However, the present invention can obviously be implemented in a variety of other ways different from the description. Those skilled in the art can make similar generalizations and deductions based on actual application situations without violating the connotation of the present invention. Therefore, the protection scope of the present invention should not be limited by the content of this specific embodiment.
[0059] A flow chart of a line of sight estimation method based on 3D point cloud and surface reconstruction technology, such as Figure 1As shown, the following steps are included:
[0060] S1. Preprocess the data in the public dataset to obtain a fixed-size grayscale image of the eye and a real gaze ball;
[0061] S2: Build and initialize the model;
[0062] s2.1: Use a convolutional neural network (CNN) based encoder as a feature extractor for monocular grayscale images; construct a continuous normalization flow (CNF module) to map the eye information and the corresponding real fixation ball into the latent space;
[0063] S2.2: Initialize model parameters and prepare for training;
[0064] S3: fixation ball generation stage;
[0065] S3.1: Sample and reconstruct multiple candidate fixation spheres from the latent space through the forward reversible transformation of CNF. In this process, CNF is used to map the eye image and fixation sphere into the latent space to generate multiple candidate fixation spheres;
[0066] S3.2: Calculate the reconstruction loss and entropy loss to help the model learn the potential structure of data representation. The formula is as follows:
[0067] L cnf =L recon +βL ent (13)
[0068]
[0069]
[0070] In formula (1), L recon is the reconstruction loss, L ent is the Gaussian entropy loss, β>0 is the regularization parameter; p in formula (2) θ (y|z, x) represents the probability of generating the fixation ball y under the condition of the latent variable z and the eye image feature x. represents the inverse mapping of the continuous normalized flow (CNF), which is used to map the fixation sphere y to the latent space to obtain the latent variable z, Involving the integral calculation in the CNF process, it is used to measure the reconstruction error; Q in formula (3) φ (ω|x) is the encoder used to learn the eye image condition information, d is the dimension of the latent variable, σ i is the standard deviation of a diagonal Gaussian distribution
[0071] S4: gaze direction prediction stage;
[0072] S4.1: Use the GSAF module to process candidate fixation balls and select the PCF mechanism or the GDF mechanism
[0073] S4.2: Use PointMLP model to predict the gaze direction;
[0074] S4.3: Calculate the loss between the real sight direction and the predicted sight direction and optimize the model. The formula is as follows:
[0075]
[0076] Using L2 loss, where g is the true view direction, is the line of sight direction predicted by the model, and the overall optimization goal is:
[0077] L=λL cnf +L g (17)
[0078] Where λ is the regularization hyperparameter;
[0079] S5: Input the eye grayscale image of the test data into the trained model;
[0080] S6: Extract image features, and the CNF module generates multiple candidate fixation balls;
[0081] S7: Use different fusion strategies (point cloud fusion or gaze sphere fusion) to get the predicted sight direction;
[0082] S8: Calculate relevant indicators and evaluate model performance.
[0083] A model framework diagram of a line of sight estimation method based on 3D point cloud and surface reconstruction technology, such as Figure 2 As shown:
[0084] Training phase: First, the data of MPIIGaze, EYEDIAP and Columbia datasets are preprocessed. The monocular images are converted into grayscale images and resized to 72*120 pixels. The three-dimensional point cloud and surface reconstruction technology are used to generate a real gaze ball according to the real line of sight direction. The ball is composed of a sphere that simulates the eyeball and iris, and is constructed by setting the diameter ratio and determining the position. Then the model is designed and built. The monocular grayscale image features are extracted using a CNN encoder. The eye information and the real gaze ball are mapped to the latent space using the continuous normalized flow (CNF). Multiple candidate gaze balls are sampled and reconstructed through the forward reversible transformation of the CNF. The sum of the reconstruction loss and the Gaussian entropy loss is calculated as the gaze ball generation loss. At the same time, the point cloud fusion (PCF) and gaze direction fusion (GDF) mechanisms in the gaze ball attention fusion (GSAF) module are used to process the candidate gaze balls. The gaze direction prediction regression loss is calculated, and the weighted sum of the two is used to obtain the total loss.
[0085] Testing phase: First, the eye images of the test data in the public dataset that have been preprocessed (the same image conversion and resizing as in the training phase) are input into the trained model. The CNN encoder and CNF module are used to extract the conditional information in the eye image and reconstruct multiple candidate gaze balls. Then, the point cloud fusion (PCF) or gaze direction fusion (GDF) strategy is selected according to the needs, and the candidate gaze balls are input into the trained PointMLP network for line of sight direction prediction to obtain the final predicted line of sight direction. Finally, the predicted line of sight direction is compared with the true line of sight direction, and the prediction performance of the model on the dataset is evaluated by averaging the errors of all test samples.
[0086] The above description is only a preferred embodiment of the present invention and is not intended to be a further limitation of the present invention. All equivalent changes made using the contents of the present specification and drawings are within the protection scope of the present invention.
Claims
1. A line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology, characterized in that: The specific steps are as follows: Step 1: Preprocess the data in the public gaze estimation datasets MPIIGaze, EYEDIAP, and Columbia to obtain the processed eye images and their corresponding real gaze spheres; Step 2: Design and build a deep learning model for gaze estimation, input the eye image and the gaze ball corresponding to the true gaze direction into the initialized model, calculate the optimization loss, and update the model parameters for supervised training; Step 3: Use the trained model as the gaze estimation model, input the eye images of the preprocessed test data in the dataset into the model for prediction, obtain the prediction results of the gaze direction, and calculate the relevant evaluation indicators.
2. The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology according to claim 1, characterized in that: The sub-steps of preprocessing the data of the public data set in step 1 are: Step 1.1: Convert all monocular images in the dataset into grayscale images, and adjust the converted grayscale images to a fixed size of 72*120 pixels; Step 1.2: Use 3D point cloud and surface reconstruction technology to generate a real gaze sphere from the labels of the real sight direction in the dataset.
3. The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology according to claim 2, characterized in that: The sub-steps of generating a real gaze ball in step 1.2 are: Step 1.2.1: The fixation ball is composed of two spheres, one simulating the eyeball and the other iris. The diameter ratio of the two spheres is 2:
1. In terms of spatial relationship, the center of the sphere simulating the iris is located at the edge of the sphere simulating the eyeball. Step 1.2.2: The real sight direction is represented by a 3-element unit vector in 3D space. The center coordinate of the eyeball is taken as the origin by default, and the position of the entire gaze ball model is determined based on this. Step 1.2.3: Use the Open3D tool to generate a real gaze ball and construct a point cloud model of the gaze ball. Use surface reconstruction technology to obtain mesh representations at different angles from the dense point cloud to build an accurate model.
4. The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology according to claim 1, characterized in that: The sub-steps of building and training the sight line estimation model in step 2 are: Step 2.1: Use the encoder based on Convolutional Neural Network (CNN) as the feature extractor of monocular grayscale image, and use the conditional continuous normalizing flow (conditional CNF) to map the eye information and the corresponding real fixation ball to the latent space. Through the forward reversible transformation of CNF, sample and reconstruct multiple candidate fixation balls from the latent space. The optimization target formula in this step is as follows: L cnf =L recon +βL ent (1) In formula (1), L recon is the reconstruction loss, L ent is the Gaussian entropy loss, β>0 is the regularization parameter; p in formula (2) θ (y|z, x) represents the probability of generating the fixation ball y under the condition of the latent variable z and the eye image feature x. represents the inverse mapping of the continuous normalized flow (CNF), which is used to map the fixation sphere y to the latent space to obtain the latent variable z, Involving the integral calculation in the CNF process, it is used to measure the reconstruction error; Q in formula (3) φ (ω|x) is the encoder used to learn the eye image condition information, d is the dimension of the latent variable, σ i is the standard deviation of the diagonal Gaussian distribution; Step 2.2: Use the gaze sphere attention fusion module (GSAF) to adaptively fuse multiple candidate gaze spheres to obtain dense gaze spheres to improve the accuracy and stability of gaze estimation; Step 2.3: The final gaze direction prediction adopts two strategies: 1) Point cloud fusion (PCF) inputs the fused dense gaze ball into the PointMLP network to obtain the final predicted gaze direction; 2) Gaze direction fusion (GDF) inputs multiple candidate gaze balls into the PointMLP to obtain different prediction results, calculates the weight of each result and then weightedly fuses them to obtain the final predicted gaze direction. The optimization target formula is as follows: Using L2 loss, where g is the true view direction, is the gaze direction predicted by the model; Step 2.4: Calculate the loss between the predicted sight direction and the actual sight direction after fusion, and use the optimization algorithm to update the model training parameters to minimize the loss and improve the performance of the model; the sight estimation model sets its initial learning rate to 5e-4, and uses the Adam optimizer to update the parameters.
5. The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology according to claim 4, characterized in that: The sub-steps of extracting eye image condition information in step 2.1 are: Step 2.1.1: Use CNN to extract the features of the monocular grayscale image as conditional information; Step 2.1.2: Construct a conditional continuous normalized flow conditional CNF to learn the mapping from simple distribution to complex distribution. Input the conditional information and the real gaze ball information into CNF, map the data to the latent space through the forward reversible transformation, and obtain the latent variables. Step 2.1.3: Use the inverse transform sampling of CNF to transform the latent variables back to the data space, reconstruct and generate multiple candidate fixation balls, the goal is to learn the representation from the eye image to the possible fixation ball. The relevant calculation formula is as follows: y=G θ (z;x) (5) Among them G θ (·,·) is the CNF transformation formula, y is the candidate fixation ball, z is the latent variable, x is the eye image feature, is the prior distribution, q φ (z|y,x) is the approximate posterior distribution, is the expected term, D KL This is the KL divergence calculation formula.
6. The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology according to claim 4, characterized in that: The sub-steps of the strategy of using the fixation ball attention fusion module to fuse multiple candidate fixation balls and make predictions in step 2.2 are: Step 2.2.1: For multiple candidate fixation balls, first use the farthest point sampling (FPS) algorithm to downsample each candidate fixation ball to obtain a new candidate fixation ball; Step 2.2.2: Calculate the confidence C of each fixation ball through two learnable matrices W1 and W2 and the Sigmoid function i , select the TopK fixation balls with high confidence according to the confidence. The formula is as follows: C i =Sigmoid(W2Pool(W1Sample(y i ))) (7) Where Sample(·) represents the farthest point sampling (FPS) algorithm, and Pool(·) represents the pooling operation; Step 2.2.3: Construct the selected K high-confidence candidate fixation spheres into a dense and more representative fixation sphere. The formula is as follows: in For the final intensive look at the ball; Step 2.2.4: Input the constructed dense gaze ball into the PointMLP network to predict the final gaze direction 7. The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology according to claim 4, characterized in that: The sub-steps of the strategy for inputting multiple gaze balls into the PointMLP network for line of sight direction in step 2.3 are: Step 2.3.1: For multiple candidate gaze balls, first input each candidate gaze ball into the PointMLP network to obtain multiple different gaze directions g i ; Step 2.3.2: Also use the farthest point sampling (FPS) algorithm to downsample the candidate fixation sphere to obtain a new candidate fixation sphere, and then calculate the weight A of each gaze direction through two learnable matrices M1 and M2 and the SoftMax function. i : TO i =SoftMax(M2Pool(M1Sample(and i ))) (10) Step 2.3.3: Perform weighted fusion based on the calculated weights and gaze directions to obtain the final predicted gaze direction. Where N is the number of candidate fixation balls; The overall optimization goal of the above steps is: L=λL cnf +L g (12) where λ is the regularization hyperparameter.
8. The line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology according to claim 1, characterized in that: The sub-steps of using the trained model to perform line of sight estimation in step 3 are: Step 3.1: Input the pre-processed eye images from the public dataset into the trained model; Step 3.2: Use the same convolutional neural network-based encoder, CNF module, and GSAF module as in the training phase to extract the conditional information in the eye image, reconstruct multiple candidate fixation balls, and select different fusion strategies; Step 3.3: Use the trained PointMLP network to predict the gaze direction, obtain the final predicted gaze direction, compare it with the true gaze direction, and calculate the performance index to evaluate the prediction performance of the model.
9. A line of sight estimation method based on three-dimensional point cloud and surface reconstruction technology, characterized in that: Including data preprocessing module, gaze ball generation module, gaze attention fusion module, and sight direction prediction module: Data preprocessing module: Convert the input monocular image into a grayscale image and adjust it to a fixed size of 72*120 pixels; at the same time, use the 3D point cloud and surface reconstruction technology to construct a gaze ball based on the eyeball geometry model, including simulating the 3D structure of the eyeball and iris; Gaze ball generation module: It uses the conditional continuous normalization flow conditional CNF to sample from the latent space and convert the original eye image into multiple candidate gaze balls; Gaze attention fusion module: It includes two fusion mechanisms: 1) Point Cloud Fusion (PCF), which evaluates the quality of each candidate gaze ball and selects the K gaze balls with the highest confidence for fusion; 2) Gaze Direction Fusion (GDF), which predicts the gaze direction of multiple generated candidate gaze spheres and performs weighted fusion; Gaze direction prediction module: Based on the different sampling mechanisms in the gaze attention fusion module, the PointMLP model is used to predict the final gaze direction.