Gaze estimation method based on double-view explicit geometric modeling
The gaze estimation method based on dual-view explicit geometric modeling solves the problems of unclear spatial correspondence between gaze and gaze point and high system cost in existing technologies. It achieves high stability and low cost of gaze and gaze point estimation, which is suitable for rehabilitation training and psychotherapy scenarios.
Patent Information
- Application Number
- CN202511769762.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-27
AI Technical Summary
Existing gaze estimation methods lack explicit geometric modeling, multi-view fusion relies on precise calibration, cannot achieve bidirectional constraints, and lack cross-view consistency supervision, resulting in unclear spatial correspondence between gaze and gaze point, high system cost, complex deployment, decreased accuracy, and poor stability.
A dual-view explicit geometric modeling approach is adopted. By acquiring face images and panoramic images, deep features are extracted using a pre-set network model. The gaze endpoints are constructed by combining the three-dimensional gaze ray formula and geometric constraints. Iterative training is then performed through joint loss to achieve a unified estimation of gaze direction and gaze point.
It improves the accuracy and interpretability of gaze estimation, ensures high stability and consistency under conditions of lighting changes, pose shifts, or partial occlusion, reduces system costs, and simplifies the deployment process.
Smart Images

Figure CN121582987A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical-assisted visual perception and eye-tracking analysis technology, and in particular to a gaze estimation method based on dual-view explicit geometric modeling. Background Technology
[0002] During rehabilitation training or psychotherapy, doctors need to monitor in real time whether the patient is looking at the task target or specific stimulus on the screen in order to assess their attention, cognitive response and emotional feedback.
[0003] However, existing gaze estimation methods are mostly limited to single-view image analysis or rely on high-cost multi-sensor systems, making it difficult to achieve stable and low-cost spatial gaze analysis in medical scenarios. The main technical problems are as follows: 1) Lack of explicit geometric modeling, resulting in unclear spatial correspondence between gaze direction and gaze point. Existing methods mostly use end-to-end regression from a single camera viewpoint, estimating the gaze direction only in the image plane without explicitly constructing the spatial transformation relationship between different viewpoints. This leads to gaze point prediction relying on empirical projection or linear approximation. When head posture deviates or patient position changes, the geometric consistency between the gaze direction and the gaze target is severely distorted. 2) Multi-view fusion relies on precise calibration, resulting in high system costs and complex deployment. Existing multi-camera fusion solutions often require extrinsic parameter calibration or eye-tracking-assisted calibration, necessitating sophisticated optical systems and multi-step registration processes. This is not only costly but also susceptible to camera micro-motion or ambient light interference, making it difficult to meet the clinical needs of "rapid deployment and plug-and-play" in treatment room scenarios. 3) Decoupling of 2D and 3D information, leading to insufficient constraints and decreased accuracy. Existing algorithms often model 3D spatial gaze estimation and 2D gaze point prediction separately, failing to achieve bidirectional constraints. Errors are amplified during gaze projection, especially with changes in lighting or partial eye occlusion, leading to predictive point drift and hindering spatial consistency and robustness. 4) Lack of cross-viewpoint consistency supervision makes model self-correction difficult. Current deep learning solutions mostly rely on single-viewpoint supervision signals, failing to simultaneously constrain the geometric relationship between the face viewpoint and the scene viewpoint. The feature extraction process is unidirectional and one-sided, preventing the network from automatically correcting biases under complex poses or rapid eye movements. CN120783382A discloses a gaze estimation method based on face and eye features, acquiring face images through a single camera and combining an attention mechanism to regress the gaze direction, achieving high-precision gaze estimation. However, this method relies entirely on manually labeled gaze direction tags, neglecting scene context and gaze target location, and cannot guarantee stability under complex backgrounds or different viewpoints. CN120406751A proposes a pilot eye-tracking method based on multi-camera fusion. In the cockpit environment, it achieves 3D gaze direction estimation through geometric modeling and deep learning fusion, and utilizes an eye tracker for error compensation. This scheme offers high accuracy, but the system is complex, requiring multi-camera calibration and dedicated eye tracker equipment, thus limiting its application scope. In summary, existing technologies either can only estimate gaze direction from a single viewpoint and lack target reconstruction capabilities, or rely on multi-sensor calibration systems, making it difficult to extend to general-purpose scenarios. Summary of the Invention
[0004] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a gaze estimation method based on dual-view explicit geometric modeling, which solves the problems of existing methods lacking explicit geometric modeling, relying on accurate calibration for multi-view fusion, being unable to achieve bidirectional constraints, and lacking cross-view consistency supervision.
[0005] To achieve the above objectives, the present invention provides the following solution:
[0006] A gaze estimation method based on dual-view explicit geometric modeling includes:
[0007] Acquire facial images and panoramic images of the target object;
[0008] The face image and the panoramic image behind the face are input into a pre-deployed gaze estimation model for calculation to obtain a gaze estimation result; the training process of the gaze estimation model includes:
[0009] The face image and the panoramic image behind the research subject are pre-collected to obtain the original acquired image;
[0010] The original acquired image is cropped for the eye area, and the center of the left eye, the center of the right eye, the center of the head, and the center of both eyes are calculated to obtain a local image of the eye and key pixels. The key pixels are then back-projected using the intrinsic parameters of the first camera to obtain projection data.
[0011] The pre-collected face image, the panoramic image behind the face, and the local eye image are used to extract deep features and perform preliminary regression to obtain the network extraction results. The network extraction results include: gaze direction unit vector, target pixel prediction vector, and depth parameters.
[0012] The line-of-sight endpoint is obtained by calculating the unit vector of the line-of-sight direction and the depth parameter using the three-dimensional line-of-sight ray formula;
[0013] The gaze endpoints are sequentially transformed and mapped to the SceneCam coordinate system and projected onto the pixel plane to obtain the mapped gaze point;
[0014] A unit two-dimensional orientation is constructed based on the projection data corresponding to the head center and the mapped gaze point;
[0015] By constructing and integrating 3D line-to-line nearest distance geometric constraints, Scene 2D vector angle consistency constraints, reprojection and pixel-level error constraints, and robustness and regularization constraints, the optimization objective is obtained.
[0016] The optimization objective is calculated based on the network extraction results and the unit two-dimensional direction to obtain the joint loss. The joint loss is then used to iteratively train the preset network model to obtain the gaze estimation model.
[0017] Preferably, the expression for the projection data is: ;in, The projection data; This refers to the intrinsic parameters of the first camera; These are the key pixels.
[0018] Preferably, the expression for the unit vector of the line of sight direction is: The expression for the target pixel prediction vector is: The expression for the depth parameter is: ;in, , , These are the gaze direction unit vector, the target pixel prediction vector, and the depth parameter, respectively. , , These are the first regression network, the second regression network, and the third regression network, respectively. The encoding result of the face image; The encoding result of the panoramic image behind it; This is the encoding result of the local image of the eye.
[0019] Preferably, the expression for the line-of-sight endpoint is: ;in, The endpoint of the line of sight; The center of each eye.
[0020] Preferably, the expression for the mapped gaze point is: ;in, ; The mapped gaze point; This represents the homogeneous coordinate projection divided by the scale factor; The SceneCam projection matrix; This is the internal reference of the second camera; This is the relative rotation matrix between the FaceCam coordinate system and the SceneCam coordinate system; This is the relative translation vector between the FaceCam coordinate system and the SceneCam coordinate system.
[0021] Preferably, the expression for the optimization objective is: ;in, The optimization objective is as described above; , , , , These are the first coefficient, the second coefficient, the third coefficient, the fourth coefficient, and the fifth coefficient, respectively. , , , , These are, respectively, the supervised loss for line-of-sight regression, the reprojection and pixel-level error constraint, the 3D line-to-line nearest distance geometric constraint, the Scene 2D vector angle consistency constraint, and the robustness and regularization constraint.
[0022] Preferably, the expression for the reprojection and pixel-level error constraint is: ;in, These are the actual pixel coordinates of the target in the SceneCam coordinate system; This represents the perspective projection function in the FaceCam coordinate system; This is the projection matrix of the FaceCam coordinate system; The actual pixel coordinates of the gaze point in the FaceCam coordinate system.
[0023] Preferably, the expression for the three-dimensional line-to-line nearest distance geometric constraint is: ;in, These are the coordinates of the camera center in the SceneCam coordinate system; The pixel coordinates of the face center in the SceneCam coordinate system; This is the unit vector of the line of sight in the FaceCam coordinate system; This is the result of backprojecting the target through the intrinsic parameters of the second camera and normalizing it.
[0024] Preferably, the expression for the Scene two-dimensional vector angle consistency constraint is: ;in, The unit is a two-dimensional direction.
[0025] Preferably, the expressions for the robustness and regularization constraints are as follows: ;in, , , For the sixth, seventh, and eighth coefficients; The rotation matrix is initially fixed; It is the Frobenius norm; This is the initial relative translation vector; This is the initial depth.
[0026] The present invention discloses the following technical effects:
[0027] This invention provides a gaze estimation method based on dual-view explicit geometric modeling. By using explicit geometric modeling, it solves the problem of existing methods lacking explicit geometric modeling, thereby improving the accuracy and interpretability of the model. By using two-layer geometric constraints, it solves the problems of existing methods relying on precise calibration for multi-view fusion, being unable to achieve bidirectional constraints, and lacking cross-view consistency supervision. It achieves the unification of the gaze direction and gaze point, as well as high stability and consistency under conditions of illumination changes, posture shifts, or partial occlusion. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 A schematic diagram of the gaze estimation process based on dual-view explicit geometric modeling provided in an embodiment of the present invention;
[0030] Figure 2 A flowchart of gaze estimation based on dual-view explicit geometric modeling is provided for embodiments of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] The purpose of this invention is to provide a gaze estimation method based on dual-view explicit geometric modeling, which solves the problems of existing methods lacking explicit geometric modeling, relying on accurate calibration for multi-view fusion, being unable to achieve bidirectional constraints, and lacking cross-view consistency supervision.
[0033] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] Figure 1 This is a schematic diagram of the gaze estimation process based on dual-view explicit geometric modeling provided in an embodiment of the present invention, as shown below. Figure 1 As shown, this invention provides a gaze estimation method based on dual-view explicit geometric modeling, comprising:
[0035] Step 100: Acquire the face image and panoramic image behind the target object;
[0036] Step 200: Input the face image and the panoramic image behind the face into a pre-deployed gaze estimation model for calculation to obtain the gaze estimation result; the training process of the gaze estimation model includes:
[0037] Step 201: Pre-collect the face image and the panoramic image behind the research subject to obtain the original acquired image;
[0038] Step 202: Crop the eye image and calculate the center of the left eye, the center of the right eye, the center of the head, and the center of both eyes in the original acquired image to obtain the local eye image and key pixels. Then, back-project the key pixels using the intrinsic parameters of the first camera to obtain projection data.
[0039] Step 203: Use a preset network model to extract deep features and perform preliminary regression on the pre-collected face image, the panoramic image behind, and the local eye image to obtain the network extraction results; the network extraction results include: gaze direction unit vector, target pixel prediction vector, and depth parameters;
[0040] Step 204: Calculate the line-of-sight direction unit vector and the depth parameter using the three-dimensional line-of-sight ray formula to obtain the line-of-sight endpoint;
[0041] Step 205: The gaze endpoints are sequentially transformed and mapped to the SceneCam coordinate system and projected onto the pixel plane to obtain the mapped gaze point;
[0042] Step 206: Construct a unit two-dimensional orientation based on the projection data corresponding to the head center and the mapped gaze point;
[0043] Step 207: Construct and integrate the 3D line-to-line nearest distance geometric constraints, the Scene 2D vector angle consistency constraints, the reprojection and pixel-level error constraints, and the robustness and regularization constraints to obtain the optimization objective;
[0044] Step 208: Calculate the optimization objective based on the network extraction results and the unit two-dimensional direction to obtain the joint loss, and use the joint loss to iteratively train the preset network model to obtain the gaze estimation model.
[0045] Specifically, this embodiment proposes a gaze estimation method based on dual-view explicit geometric modeling to simultaneously estimate the patient's gaze point and gaze direction using two ordinary RGB cameras (front and rear) in a treatment room environment. For example... Figure 2 As shown, the specific steps are as follows:
[0046] S1: Camera acquisition and initial preprocessing. Two ordinary RGB cameras were used to simultaneously acquire data in the treatment room; FaceCam acquired face images ( SceneCam acquires the panoramic image behind the scenes. Distortion correction, grayscale / color normalization, and scale standardization are performed on both images; and the approximate intrinsic parameter matrices of the two cameras are recorded. (If unknown, calibration or initial estimation can be performed).
[0047] S2: Face and eye landmark detection and eye-center calculation. (In...) Detect facial landmarks and crop the eye image. ), calculate the center of the left and right eyes ( ) and head center ( ), and the center of both eyes:
[0048]
[0049] These pixels can be back-projected into rays in the camera coordinate system using intrinsic parameters:
[0050]
[0051] in These are pixel coordinates.
[0052] S3: Extracting deep features and preliminary regression. (For...) The encoder (ResNet18) is applied to obtain the features ( ) respectively. ), and extract eye features from facial keypoints detected by Dlib using ResNet18. Based on these features, the gaze direction unit vector of the face branch is regressed. (In the FaceCam coordinate system) and viewpoint pixel prediction of the Scene branch ( ):
[0053]
[0054]
[0055] S4: Explicit Camera Extrinsic Initialization and Robust Estimation. If the extrinsic parameters of the two cameras are unknown, the initial relative rotation and translation are estimated using several pairs of known corresponding 3D points or keypoint pairs via the PnP / five-point method and RANSAC. The projection matrix is initialized using the following formula:
[0056]
[0057]
[0058] in Represents the identity matrix. This indicates that FaceCam's extrinsic parameters in its own coordinate system are a unit rotation matrix and a zero translation vector, meaning the camera center is at the origin of its own coordinate system. To prevent the influence of outliers, RANSAC is used for robust filtering of data correspondences.
[0059] S5: Explicit 3D construction and depth calibration strategy for the gaze endpoint. (The image shows the eye center under the FaceCam.) ) is considered as a point in the camera coordinate system ( ), with the regression line of sight unit vector ( Constructing a three-dimensional line of sight:
[0060]
[0061] Due to depth ( The uncertainty is addressed by introducing deep regression during training. Then the endpoint of the line of sight is:
[0062]
[0063] S6: Explicit coordinate transformation and projection onto the panoramic image. This involves transforming the viewpoint or 3D point in the FaceCam coordinate system... The relative transformation is used to map the viewpoint to the SceneCam coordinate system and project it onto the pixel plane to obtain the mapped gaze point. ):
[0064]
[0065]
[0066] in This represents the homogeneous coordinate projection divided by the scale factor.
[0067] S7: SceneCam 2D gaze vector construction. On the SceneCam image, this can be achieved using the pixels at the center of the head (…). (or projected from the head) and target pixels ( )(reality Or network regression ( ) constitute a two-dimensional vector ( Normalizing it yields the unit two-dimensional direction:
[0068]
[0069] This two-dimensional vector represents the "viewpoint projection direction" on the SceneCam image plane and can be used as a second layer of supervision constraint.
[0070] S8: 3D line-to-line nearest distance geometric constraint. The mapped 3D ray obtained from FaceCam ( (in the Scene coordinate system) and SceneCam pixels ( Three-dimensional rays generated along the back projection direction ( (From the center of the Scene camera through pixels) Treating the two lines as two straight lines in space, calculate their shortest distance and use this as the geometric consistency loss: if the two lines are parameterized as and Then the square of the shortest distance is:
[0071]
[0072] This term is included as part of the loss to constrain the two lines to ideally intersect or be as close as possible.
[0073] S9: Scene 2D Vector Angle Consistency Constraint. On the pixel plane, the SceneCam is a 2D vector pointing from the head center to the network-predicted gaze point (…). ) and the two-dimensional vector obtained by mapping and projection from FaceCam ( Apply angle consistency constraints and define the two-dimensional angle loss as:
[0074]
[0075] This loss provides robust line-of-sight direction supervision even in the absence of precise 3D depth information.
[0076] S10: Reprojection and pixel-level error constraints. For known or labeled targets ( ), calculate its reprojection error under the two cameras for direct supervision:
[0077]
[0078] The true pixel coordinates of the gaze point on FaceCam. The reprojection error and the line-to-line distance of S8 together form a strong geometric constraint.
[0079] S11: Robustness and Regularization Measures. To reduce the risk of parameter estimation degradation, regularization must be applied to the externally involved deep variables, such as... Prior bias penalty, for deep regression ( The range constraints and smoothing regularization terms for the projection scale can be added.
[0080]
[0081] in All values are initialization values or empirical values.
[0082] S12: Joint Loss and Optimization Strategy. The angle / direction loss, pixel regression loss, line-to-line shortest distance loss, 2D vector angle consistency loss, reprojection loss, and regularization term are weighted and combined to obtain the final optimization objective:
[0083]
[0084] in, This represents the supervised loss for commonly used gaze direction regression, based on the true gaze unit vector labeled in the FaceCam coordinate system. , can be obtained During the training phase, a phased optimization approach is adopted: first, RANSAC+PnP is used to fix the extrinsic parameters for estimation, then the deep regression and regression networks are optimized, and finally all parameters are fine-tuned jointly. During deployment, it is possible to choose to perform only forward inference and use the minimum cost method to fine-calibrate the extrinsic parameters online.
[0085] S13: Periodic calibration of extrinsic parameters. To address camera installation micro-motion or long-term drift, an online calibration process is designed: periodically acquire short-sequence synchronization frames, perform short-term PnP estimation using identifiable markers in the scene or key points of the doctor / device, and use a sliding window Bundle Adjustment to minimize short-term reprojection error and update (…). ).
[0086] S14: Output confidence score. The final output of the system is: the gaze vector in the FaceCam coordinate system ( ), estimated depth ( ), and the gaze point in the SceneCam pixel system ( Because it explicitly uses camera matrices and geometric formulas, the output has interpretable 3D geometric meaning, making it easy to integrate into clinical monitoring or interactive modules, and providing additional diagnostic information (such as the shortest distance between two lines). (as a confidence level indicator).
[0087] The beneficial effects of this invention are as follows:
[0088] This invention ensures the geometric consistency between the gaze direction and the gaze point position in three-dimensional space through explicit geometric modeling, without relying on implicit network learning, thereby improving the accuracy and interpretability of the estimation. Through two-layer geometric constraints, it ensures the uniformity between the gaze direction and the gaze point, and can maintain high stability and consistency under conditions of illumination changes, posture shifts, or partial occlusion.
[0089] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0090] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A gaze estimation method based on dual-view explicit geometric modeling, characterized in that, include: Acquire facial images and panoramic images of the target object; The face image and the panoramic image behind the face are input into a pre-deployed gaze estimation model for calculation to obtain the gaze estimation result; The training process of the gaze estimation model includes: The face image and the panoramic image behind the research subject are pre-collected to obtain the original acquired image; The original acquired image is cropped for the eye area, and the center of the left eye, the center of the right eye, the center of the head, and the center of both eyes are calculated to obtain a local image of the eye and key pixels. The key pixels are then back-projected using the intrinsic parameters of the first camera to obtain projection data. The pre-collected face image, the panoramic image behind the face, and the local eye image are used to extract deep features and perform preliminary regression to obtain the network extraction results. The network extraction results include: gaze direction unit vector, target pixel prediction vector, and depth parameters. The line-of-sight endpoint is obtained by calculating the unit vector of the line-of-sight direction and the depth parameter using the three-dimensional line-of-sight ray formula; The gaze endpoints are sequentially transformed and mapped to the SceneCam coordinate system and projected onto the pixel plane to obtain the mapped gaze point; A unit two-dimensional orientation is constructed based on the projection data corresponding to the head center and the mapped gaze point; By constructing and integrating 3D line-to-line nearest distance geometric constraints, Scene 2D vector angle consistency constraints, reprojection and pixel-level error constraints, and robustness and regularization constraints, the optimization objective is obtained. The optimization objective is calculated based on the network extraction results and the unit two-dimensional direction to obtain the joint loss. The joint loss is then used to iteratively train the preset network model to obtain the gaze estimation model.
2. The gaze estimation method based on dual-view explicit geometric modeling according to claim 1, characterized in that, The expression for the projection data is: ;in, The projection data; This refers to the intrinsic parameters of the first camera; These are the key pixels.
3. The gaze estimation method based on dual-view explicit geometric modeling according to claim 1, characterized in that, The expression for the unit vector of the line of sight direction is: The expression for the target pixel prediction vector is: The expression for the depth parameter is: ;in, , , These are the line-of-sight direction unit vector, the target pixel prediction vector, and the depth parameter, respectively. , , These are the first regression network, the second regression network, and the third regression network, respectively. The encoding result of the face image; The encoding result of the panoramic image behind it; This is the encoding result of the local image of the eye.
4. The gaze estimation method based on dual-view explicit geometric modeling according to claim 3, characterized in that, The expression for the line-of-sight endpoint is: ;in, The endpoint of the line of sight; The center of each eye.
5. The gaze estimation method based on dual-view explicit geometric modeling according to claim 4, characterized in that, The expression for the mapped gaze point is: ;in, ; The mapped gaze point; This represents the homogeneous coordinate projection divided by the scale factor; The SceneCam projection matrix; This is the internal reference of the second camera; This is the relative rotation matrix between the FaceCam coordinate system and the SceneCam coordinate system; This is the relative translation vector between the FaceCam coordinate system and the SceneCam coordinate system.
6. The gaze estimation method based on dual-view explicit geometric modeling according to claim 5, characterized in that, The expression for the optimization objective is: ;in, The optimization objective is as described above; , , , , These are the first coefficient, the second coefficient, the third coefficient, the fourth coefficient, and the fifth coefficient, respectively. , , , , These are, respectively, the supervised loss for line-of-sight regression, the reprojection and pixel-level error constraint, the 3D line-to-line nearest distance geometric constraint, the Scene 2D vector angle consistency constraint, and the robustness and regularization constraint.
7. The gaze estimation method based on dual-view explicit geometric modeling according to claim 6, characterized in that, The expression for the reprojection and pixel-level error constraint is: ;in, These are the actual pixel coordinates of the target in the SceneCam coordinate system; Represents the perspective projection function in the FaceCam coordinate system; This is the projection matrix of the FaceCam coordinate system; The actual pixel coordinates of the gaze point in the FaceCam coordinate system.
8. The gaze estimation method based on dual-view explicit geometric modeling according to claim 6, characterized in that, The expression for the three-dimensional line-to-line nearest distance geometric constraint is: ;in, These are the coordinates of the camera center in the SceneCam coordinate system; The pixel coordinates of the face center in the SceneCam coordinate system; This is the unit vector of the line of sight in the FaceCam coordinate system; This is the result of backprojecting the target through the intrinsic parameters of the second camera and normalizing it.
9. A gaze estimation method based on dual-view explicit geometric modeling according to claim 6, characterized in that, The expression for the Scene two-dimensional vector angle consistency constraint is: ;in, The unit is a two-dimensional direction.
10. A gaze estimation method based on dual-view explicit geometric modeling according to claim 6, characterized in that, The expressions for the robustness and regularization constraints are as follows: ;in, , , The sixth, seventh, and eighth coefficients; The rotation matrix is initially fixed; It is the Frobenius norm; This is the initial relative translation vector; This is the initial depth.
Citation Information
Patent Citations
Pilot eye movement tracking method, system and device based on multi-camera fusion and storage medium
CN120406751A
3D sight line target estimation method and device
CN120783382A