Multi-camera eye movement self-calibration fusion method based on ensemble learning and graph neural network
Patent Information
- Application Number
- CN202611030870.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-07-13
AI Technical Summary
[0006]为了解决现有技术中的上述问题,即在多相机条件下通常依赖人工标定或固定标定板,在头部运动范围较大以及低照度成像导致眼部特征不稳定时难以实现跨相机几何关系与视线映射的快速标定,且不同相机观测难以在统一坐标系下稳定融合输出,从而阻碍了多相机眼动系统在真实复杂场景中的实用化问题,本发明提出了基于集成学习与图神经网络的多相机眼动自标定融合方法,包括以下步骤:
1)实现快速免人工的跨相机自标定:利用时间同步的多相机眼动观测,结合集成外参估计、图神经网络一致性聚合以及可微分束调整迭代优化,在无需标定板和人工配合的情况下获得全局一致的相机外参数,实现跨相机几何关系与视线映射的快速自标定;
Smart Images

Figure CN122530290B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and eye tracking technology, specifically involving a multi-camera eye tracking self-calibration fusion method based on ensemble learning and graph neural networks. Background Technology
[0002] Eye tracking and gaze estimation technologies are widely used in scenarios such as driving and flight simulation, human-computer interaction, and behavior analysis.
[0003] With the development of deep learning applications in eye feature extraction and gaze estimation, methods such as gaze direction regression and 3D eye position estimation based on monocular cameras or single eye trackers have become relatively mature. In complex application environments, to expand the field of view, reduce occlusion, and improve output stability, the industry is gradually adopting multi-camera or multi-eye tracker deployment schemes, and achieving gaze output in a unified coordinate system through time synchronization, cross-camera coordinate mapping, and multi-source fusion. For obtaining cross-camera geometric relationships, existing technologies typically employ offline calibration using calibration boards or markers, prior extrinsic parameter setting based on the installation structure, or correction of extrinsic parameters using multi-view geometry and optimization methods under certain conditions. For multi-source gaze fusion, common methods include simple weighted averaging, confidence-based fusion, and time-series filtering smoothing.
[0004] However, the existing technologies still have shortcomings in multi-camera eye-tracking applications, mainly in the following aspects: First, cross-camera calibration often relies on manual intervention, calibration boards, or specific viewpoint guidance, making it difficult to complete quickly in actual sessions. Furthermore, recalibration is required when there are minor changes in camera installation, affecting efficiency. Second, in situations with large head movements, frequent blinking occlusion, and decreased eye texture contrast due to low-light imaging, single-camera eye-tracking observations are prone to degradation, and extrinsic parameter correction and gaze output are easily jittered or drifted, resulting in insufficient robustness. Third, there may be time offsets or alignment errors between multiple cameras, leading to significant differences in observation quality between different cameras. Existing fusion methods lack adaptive modeling for cross-camera consistency constraints and observation quality, making it difficult to obtain stable and reliable fused gaze results in a unified coordinate system. These shortcomings limit the practical application of multi-camera eye-tracking systems in real-world scenarios such as low-light conditions and large head movements.
[0005] Therefore, a multi-camera eye-tracking self-calibration fusion method is needed to address the shortcomings of the existing technologies. Summary of the Invention
[0006] To address the aforementioned problems in existing technologies—namely, the reliance on manual calibration or fixed calibration plates in multi-camera environments, the difficulty in rapidly calibrating cross-camera geometric relationships and line-of-sight mapping when the head motion range is large and low-light imaging leads to unstable eye features, and the challenge in stably fusing observations from different cameras in a unified coordinate system—this invention proposes a multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks, comprising the following steps:
[0007] Acquire image sequences and timestamps from multiple cameras, determine the reference coordinate system and reference plane, and obtain a synchronization dataset; Face eye localization and preprocessing are performed on the image sequence in the synchronous dataset to obtain eye region images and initial camera extrinsic parameters; The eye region image is input into the eye movement estimation model based on the timestamp information in the synchronous dataset to obtain eye movement observation data including gaze direction, three-dimensional eye position, and observation quality parameters; For the eye-tracking observation data, the initial camera extrinsic parameters, and the reference plane, integrated extrinsic parameter estimation is performed: candidate extrinsic parameter increments are calculated using a weighted 3D point set registration sub-model and a weighted line-of-sight alignment sub-model, and candidate confidence is calculated and filtered based on the line-of-sight reprojection residuals to obtain a set of candidate extrinsic parameter increments; Construct an initial graph with the camera as a node and the relative pose prior and time alignment parameters as edges; Graph neural network inference is performed on the initial graph. Cross-camera consistency constraints and aggregation are performed on the candidate extrinsic increments of each node through message passing and node updates to obtain the node fusion weights and globally consistent extrinsic increments of each node. The initial camera extrinsic parameters are updated based on the global consistent extrinsic parameter increment. Under the framework of differentiable beam adjustment, the objective function is used to iteratively optimize the camera extrinsic parameters to be optimized, and the self-calibrated camera extrinsic parameters are obtained. The objective function includes a line-of-sight reprojection error term and a cross-camera line-of-sight consistency error term weighted by the node fusion weight and the observation quality parameter. Based on the self-calibrated camera extrinsic parameters, the eye-tracking observation data is transformed to a reference coordinate system and intersected with the reference plane. The data is then weighted and fused according to the node fusion weights and observation quality parameters to obtain the fused gaze result in the reference coordinate system.
[0008] The beneficial effects of this invention are: 1) Achieve rapid, manual-free cross-camera self-calibration: Utilize time-synchronized multi-camera eye-tracking observations, combined with integrated extrinsic parameter estimation, graph neural network consistency aggregation, and differentiable beam adjustment iterative optimization, to obtain globally consistent camera extrinsic parameters without the need for calibration boards or manual assistance, thus achieving rapid self-calibration of cross-camera geometric relationships and line-of-sight mapping; 2) Improve robustness and stability under conditions of large head movement range and low illumination: By enhancing and denoising preprocessing under low illumination, and introducing observation quality parameters, candidate confidence and node fusion weights, low-quality observations such as blink occlusion and decreased pupil visibility are adaptively downweighted or screened to reduce the impact of degraded observations on external parameter estimation and gaze output. 3) Achieve stable fused gaze output in a unified coordinate system: Transform the line of sight of multiple cameras and the three-dimensional position of the eye to a reference coordinate system and intersect with the reference plane. Through node fusion weight and observation quality parameters, weighted fusion is performed and time series smoothing can be combined to obtain continuous and stable fused gaze results, reducing single camera jitter and line of sight drift. Attached Figure Description
[0009] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the steps of the multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks of the present invention. Detailed Implementation
[0010] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0011] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0012] To more clearly explain the multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks of this invention, the following will be combined with... Figure 1 The steps in the embodiments of the present invention will be described in detail below.
[0013] The first embodiment of this invention is a multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks, see [link to relevant documentation]. Figure 1 This includes the following steps: S1: Obtain image sequences and timestamps from multiple cameras, determine the reference coordinate system and reference plane, and obtain the synchronization dataset; In this embodiment, the method for obtaining the synchronized dataset is as follows: S1.1 At the start of the session, a synchronous acquisition trigger signal is sent to the multiple cameras. Specifically, at the start of the session, the synchronization controller sends a synchronous acquisition trigger signal to multiple cameras (at least two) through the hardware trigger bus. The synchronous acquisition trigger signal adopts a TTL pulse with a single rising edge and the pulse width is set to 10 ms. After detecting the rising edge, each camera takes the exposure start time as the starting point of the first frame acquisition and enters the continuous acquisition mode. S1.2 Generate timestamp information for each frame of image acquired by each camera based on the same common clock. Specifically: The synchronization controller provides the same common clock and synchronizes the clocks of each camera device through the IEEE 1588 PTP method, so that each camera reads the common clock count and generates timestamp information at the start of exposure of each frame of image acquisition. The timestamp information is stored as a 64-bit unsigned integer in microseconds and bound to the frame image to form an "image-timestamp" pair. S1.3 When the timestamp information of different cameras meets the preset synchronization threshold, the images from each camera that meet the preset synchronization threshold are merged into a synchronization frame group. Specifically, the synchronization data processing thread maintains a circular buffer for each camera, sorted by timestamp in ascending order, and performs synchronization frame group merging when there is at least one image to be matched in all camera buffers. The merging criterion is set as follows: the absolute value of the timestamp difference between any two cameras in the same synchronization frame group does not exceed the preset synchronization threshold. ; in, Indicates the first The first camera The timestamps of the images selected in each synchronized frame group. Indicates the first The first camera The timestamps of the images selected in each synchronized frame group. and Index the camera and satisfy For synchronization frame group index, Set the preset synchronization threshold and set it to When multiple frames satisfy the above criteria, the frame with the smallest difference between each camera and the minimum timestamp in the group is taken as the matching frame of that camera in the synchronization frame group, and all matching frames in the synchronization frame group are dequeued from their respective buffers to avoid duplicate merging. S1.4 The origin position and coordinate axis direction of the reference coordinate system are determined according to the pre-stored simulator geometric parameters, and the plane parameters corresponding to the instrument display surface are determined as the reference plane according to the simulator geometric parameters. Specifically: In determining the reference coordinate system and the reference plane, the pre-stored simulator geometric parameters are read from the non-volatile memory. The simulator geometric parameters include the three-dimensional position definition of the origin of the reference coordinate system at the simulator structural reference point and the direction definition of each coordinate axis of the reference coordinate system on the simulator structure. Based on this, a right-handed reference coordinate system is constructed. The X-axis of the reference coordinate system is defined as the unit direction pointing forward of the simulator, the Y-axis is defined as the unit direction pointing to the right of the simulator, and the Z-axis is uniquely determined by the right-hand rule to ensure that the coordinate system is orthogonal and consistent. Simultaneously, based on the geometric parameters of the simulator, the three-dimensional coordinates of three non-collinear structural reference points corresponding to the instrument display surface are read, and a reference plane is calculated accordingly. The reference plane is stored with "three-dimensional coordinates of a point on the plane" and "plane unit normal vector" as plane parameters. The unit normal vector is obtained by normalizing the normal vector obtained by the cross product of two reference edge vectors to ensure the numerical stability of subsequent line intersection calculations. S1.5 Bind the synchronization frame group, the timestamp information, the reference coordinate system, and the reference plane to obtain the synchronization dataset. Specifically, bind each camera image, each camera timestamp information, the reference coordinate system, and the reference plane parameters within each synchronization frame group and write them into the data structure of the synchronization dataset. Each record uses the synchronization frame group index as the primary key and contains the mapping relationship between the camera identifier and the "image-timestamp" pair to obtain the synchronization dataset. S2, perform face eye localization and preprocessing on the image sequence in the synchronous dataset to obtain eye region images and initial camera extrinsic parameters; In this embodiment, the preprocessing includes: S2.1 Perform face detection frame by frame on the image sequence of each camera in the synchronous dataset to determine the face region. Specifically, face detection is implemented using the RetinaFace single-stage detection network. The RetinaFace uses ResNet-50 as the backbone network and outputs face bounding boxes and five face key points on three scale feature layers. The network input is an RGB tensor of 640×640 with the original image scaled proportionally along the long side and padded to the same size, and the pixels are normalized to [0,1]. During inference, the face confidence threshold is set to 0.80 and the IoU threshold for non-maximum suppression is set to 0.40. The face bounding box with the highest confidence is selected as the face region from the output results. S2.2 Extract eye key points within the face region and determine the eye region boundary based on the eye key points, thereby cropping the image sequence to obtain the eye region image. Specifically: perform eye key point extraction within the face region to determine the eye region boundary. The eye key point extraction is implemented using the HRNet-W18 key point regression network. The HRNet-W18 constructs a feature extraction backbone with four-stage parallel multi-resolution branches and inputs a key point heatmap header at the output. The network input is the cropped face region scaled to... RGB tensor and pixels normalized to The network output is the two-dimensional pixel coordinates of 6 key points of the eyes for each eye. The 6 key points of the eyes are the inner corner of the eye, the outer corner of the eye, the midpoint of the upper eyelid, the midpoint of the lower eyelid, the edge of the upper iris, and the edge of the lower iris. Based on the key points of the left and right eyes, the bounding boxes of the eye regions are calculated respectively. The center of the bounding box is the midpoint between the inner and outer corners of the eye, and the width of the bounding box is 2.0 times the pixel distance between the inner and outer corners of the eye. The height of the bounding box is 2.2 times the pixel distance between the midpoint of the upper eyelid and the midpoint of the lower eyelid. The bounding boxes are restricted to the range of the original image and then cropped from the original image to obtain the left and right eye region images. If the face detection or eye key point extraction does not meet the above threshold conditions, the frame is marked as unusable and is not written into the preprocessing dataset. S2.3 Perform low-light enhancement processing based on a deep network on the eye region image to improve the contrast of eye texture, and perform denoising processing on the low-light enhanced eye region image to suppress noise interference on eye features. Specifically, the low-light enhancement is implemented using a U-Net structure enhancement model based on a deep network. The enhancement model consists of a 4-level encoder and a 4-level decoder, with each level containing two layers. Convolutional and ReLU activations are used, and symmetric layer features are fused through skip connections. The first layer has 32 channels, which doubles with each downsampling step and halves with each upsampling step. The model output uses a sigmoid function to limit the pixels of the enhanced eye region image to a certain limit. To improve texture contrast and suppress overexposure; The denoising process is implemented using a DnCNN denoising network, which consists of 17 convolutional layers, with the first layer being... Convolution with ReLU, the last layer is... Convolution output noise residual, the remaining 15 layers are The network employs convolution, batch normalization, and ReLU. The input is a low-light enhanced image of the eye region, and the output is a noise residual map. The denoising result is obtained by subtracting the noise residual map from the enhanced image and then cropping it. It is then quantized into 8-bit grayscale or RGB format for consistent use in subsequent eye-tracking estimation; S2.4 Determine the initial pose of each camera relative to the reference coordinate system based on the camera installation prior, and convert the initial pose into the initial camera extrinsic parameters. Specifically: determine the initial camera extrinsic parameters for each camera based on the camera installation prior. The camera installation prior is given by the installation configuration file and includes the initial translation vector and initial attitude Euler angles of the camera relative to the reference coordinate system. The Euler angles adopt... The initial rotation parameters of the camera are obtained by sequentially converting the parameters in radians to a rotation matrix, thereby representing the initial camera extrinsic parameters as a homogeneous transformation matrix from the camera coordinate system to the reference coordinate system: ; in, Indicates the camera index. Indicates the first The initial rotation matrix of the camera. Indicates the first The initial translation vector of the camera in the reference coordinate system. This represents a row vector consisting of three zeros, where 1 represents the homogeneous coordinate constant. S2.5 The eye region image after low-light enhancement and denoising is bound with the corresponding timestamp information to form a preprocessed dataset. Specifically, the eye region image of each camera after low-light enhancement and denoising in each synchronous frame group is bound with the corresponding timestamp information and written into the data structure of the preprocessed dataset along with the camera identifier. At the same time, the initial camera extrinsic parameters of each camera are output. S3, input the eye region image into the eye movement estimation model according to the timestamp information in the synchronous dataset to obtain eye movement observation data including gaze direction, three-dimensional position of the eye and observation quality parameters; In this embodiment, the eye region image is input into the eye movement estimation model based on the timestamp information in the synchronized dataset to obtain eye movement observation data including gaze direction, three-dimensional eye position, and observation quality parameters. The method is as follows: S3.1 The eye region images are input into the same eye-tracking estimation model in the order of timestamp information in the synchronized dataset; the eye-tracking estimation model outputs the gaze direction and the three-dimensional position of the eye for each frame of the eye region image. The eye-tracking estimation model includes a deep learning model, where the gaze direction is represented by a unit direction vector in the camera coordinate system, and the three-dimensional position of the eye is represented by three-dimensional coordinates in the camera coordinate system. Specifically: the left and right eye region images and their associated timestamp information are read from the preprocessed dataset, and each synchronized frame group is processed in ascending order of timestamp information. The record; For the The camera in the synchronized frame group index is The corresponding left and right eye region images were normalized in size, and each eye region image was bilinearly scaled to [size missing]. It also maintains the RGB three channels and normalizes the pixel values to... This then forms the model input; Eye-tracking estimation employs the same eye-tracking estimation model to share inference parameters across all cameras, achieving consistent observation generation across cameras. This model is a two-branch convolutional regression network, containing two structurally identical and parameter-shared feature extraction branches for the left and right eyes. Each branch uses ResNet-18 as its backbone and outputs a 512-dimensional feature vector after global average pooling. The 512-dimensional feature vectors from the two branches are concatenated by channel to form a 1024-dimensional fused feature, which is then input to a three-way output head. The gaze direction output head consists of two fully connected layers with a hidden layer dimension of 256, outputting a three-dimensional gaze direction raw vector. This vector is then normalized to obtain the unit gaze direction in the camera coordinate system. The eye 3D position output head consists of two fully connected layers with a hidden layer dimension of 256, and outputs the eye 3D position in the camera coordinate system. Furthermore, the three-dimensional component unit is meters, the confidence output head consists of two fully connected layers with a hidden layer dimension of 128, and the output is mapped by the Sigmoid function to obtain the model confidence. and ; The camera coordinate system is a right-handed system with the camera optical center as the origin and the positive direction of the optical axis as the coordinate system. Positive axis, horizontal direction of the image The positive axis direction and the vertical downward direction of the image are The positive axis ensures that the coordinate conventions for subsequent steps of external parameter transformation and line-of-sight intersection are consistent; S3.2 Based on the confidence level output by the eye-tracking estimation model and the usability assessment results of the eye region image, observation quality parameters are generated. The usability assessment results include pupil visibility and blinking status. Specifically, pupil visibility is obtained by performing Otsu adaptive thresholding on the enhanced and denoised grayscale eye region image to obtain binary candidate pupil masks and then performing... The structuring element opening operation and connected component filtering are used to obtain the largest connected component as the pupil region. The pupil visibility is obtained by calculating the ratio of the number of pixels in the pupil region to the total number of pixels in the eye region. and The blinking status is determined by calculating the ratio of the longitudinal to the lateral distance of the eye fissure using key eye points and comparing it with a threshold of 0.18 to obtain the blink marker. and When the conditions for blinking are met otherwise ; Based on the above quantities, construct the observation quality parameters: ; in, Indicates the first The camera is in the synchronized frame group index. Observation quality parameters at that time This represents the model confidence level output by the eye-tracking estimation model. This represents the pupil visibility calculated from pupil segmentation and pixel ratio. This indicates a blink marker determined by the palpebral fissure ratio threshold. This is the camera index and corresponds to the camera identifier in the synchronized dataset. The index is for the synchronization frame group and is consistent with the synchronization frame group obtained by merging in step S1. S3.3 The line-of-sight direction, three-dimensional eye position, and corresponding observation quality parameters of each camera at the same time are bound to the timestamp information to form eye-tracking observation data aligned with the timestamp information. Specifically: each camera is indexed in each synchronization frame group as... unit of sight direction Three-dimensional position of the eyes With observation quality parameters The timestamp information of the camera within the synchronization frame group The data is bound and written into the data structure of the camera observation set, and the same synchronization frame group is indexed as... The binding results of all cameras are aggregated into aligned observation records at the same time according to the camera index, resulting in eye-tracking observation data (i.e., camera observation set) aligned with the timestamp information. S4, perform integrated extrinsic parameter estimation on the eye-tracking observation data, the initial camera extrinsic parameters and the reference plane: calculate candidate extrinsic parameter increments using a weighted 3D point set registration sub-model and a weighted line-of-sight alignment sub-model, calculate candidate confidence based on line-of-sight reprojection residuals and filter them to obtain a set of candidate extrinsic parameter increments; In this embodiment, a weighted 3D point set registration sub-model and a weighted line-of-sight alignment sub-model are used to calculate candidate extrinsic parameter increments. The candidate reliability is calculated and filtered based on the line-of-sight reprojection residuals to obtain a set of candidate extrinsic parameter increments. The method is as follows: S4.1 Select the camera with the largest average observation quality parameter as the reference camera, and use the synchronization frame group with the observation quality parameter greater than the preset quality threshold as the valid frame. Specifically: based on the camera observation set (i.e., the camera observation set obtained from all cameras, as the statistical basis), count the observation quality parameters of each camera in continuous... The average (arithmetic mean) of the observation quality parameters within each synchronization frame group is used to select the camera with the largest average as the reference camera. The reference camera keeps its initial camera extrinsic parameters unchanged and sets its candidate extrinsic parameter increments to zero increments to serve as anchor points for subsequent relative estimations. Here, N is a predefined threshold that can be selected according to calibration requirements: the smaller the value of N, the faster the calibration speed; the larger the value of N, the more stable the calibration results, but the longer the time required. Then, for any non-reference camera Extract the index of the camera and the reference camera in the same synchronization frame group from the camera observation set. The observed quality parameters exist simultaneously greater than the preset quality threshold (i.e., satisfying the following conditions). and The observation pairs are taken as valid frames, where, and The first The synchronization frame group index of the primary camera and the reference camera is Observation quality parameters at that time For reference camera index; S4.2 For each non-reference camera, the weighted 3D point set registration sub-model uses the 3D eye positions of the local camera and the reference camera in the same valid frame. As a 3D point pair, the relative rotation matrix and translation vector are solved through weighted singular value decomposition, using the product of the observation quality parameters of the two cameras as weights, to output the first candidate relative pose. Specifically, the weighted 3D point set registration sub-model uses... and As the three-dimensional point pair corresponding to time, with weights Weighting is performed by calculating the weighted centroids of the two sets of 3D points, constructing the weighted covariance matrix, and solving it using singular value decomposition to obtain the weighted centroids of the two sets of 3D points. 3D position of the eye in the camera coordinate system Rigid transformation to the reference camera coordinate system for approximation The relative rotation matrix and relative translation vector are used to obtain the first candidate relative pose. ,in, For from the first The relative rotation matrix from the camera coordinate system to the reference camera coordinate system satisfies the orthogonality constraint. For from the first The translation vector from the camera coordinate system to the reference camera coordinate system, with units in meters; The weighted gaze orientation alignment sub-model described in S4.3 uses the gaze orientations of the local camera and the reference camera in the same valid frame as unit orientation vector pairs. It obtains the relative rotation matrix by solving a weighted Wahba problem with equal weights, and then calculates the relative translation vector using the weighted centroid of the eye's 3D position, outputting the second candidate relative pose. Specifically: the weighted gaze orientation alignment sub-model uses... and As the unit (line of sight) direction vector pair corresponding to the time, with the same weight We perform weighted calculations and obtain the weighted Wahba problem by solving the weighted Wahba problem. Approximation in the least squares sense after rotation relative rotation matrix And in getting Then, the relative translation vector is calculated using the weighted centroids of two sets of three-dimensional points. To make through Rotate and add Then aligned to the mean. The second candidate relative pose is obtained. ; S4.4 Combine the first candidate relative pose and the second candidate relative pose with the initial camera extrinsic parameters of the reference camera, and then relativize them with the initial camera extrinsic parameters of the local camera to obtain two candidate extrinsic parameter increments. Specifically: using the initial camera extrinsic parameters of the reference camera... The candidate relative poses are combined to obtain the first... Candidate updated external parameters for Taiwanese cameras ,in, Index of the sub-model for extrinsic parameter estimation and For from the first The homogeneous transformation matrix from the camera coordinate system to the reference coordinate system; Further With the camera's initial camera extrinsic parameters Relativization is performed to obtain candidate extrinsic parameter increments, and these increments are decomposed into rotation increments. With displacement increment ,in, Used to update the rotation parameters in the initial camera extrinsic parameters. Used to update the displacement parameters in the initial camera extrinsic parameters; S4.5 For each candidate extrinsic parameter increment, the gaze direction and eye 3D position of each effective frame are transformed to the reference coordinate system using the initial camera extrinsic parameters, and the intersection with the reference plane is obtained to obtain the gaze point. Simultaneously, based on all camera gaze points, the weighted robust estimation error reference plane gaze point is calculated according to the observation quality parameters, and a candidate confidence level is generated using an exponential decay function. Specifically: for each candidate extrinsic parameter increment, an extrinsic parameter update is first performed to obtain the candidate updated extrinsic parameters corresponding to that increment. Then the first The camera indexes in each synchronization frame group Three-dimensional position of the lower eyelid relative to the unit's line of sight pass Transform to the reference coordinate system and intersect with the reference plane determined in step S1 to obtain the reference plane gaze point under the candidate extrinsic parameters. Among them, when the line of sight is parallel to the reference plane or the intersection parameter is non-positive, causing the intersection point to be located in the opposite direction of the line of sight, the frame is removed from the residual statistics; Simultaneously based on all cameras being in the same synchronized frame group index The reference plane gaze point is determined according to the observation quality parameters of each camera. The index of the synchronization frame group is calculated by weighting and iterative reweighted least squares using Huber loss. Error reference plane gaze point Used as the reprojection alignment target; Therefore, the candidate confidence is defined as a line-of-sight reprojection residual decay function weighted by observation quality parameters: ; in, Indicates the first The camera was from the first The candidate confidence level of the output of each extrinsic parameter estimation submodel, with a value range of [value missing]. , Represents the residual scaling parameter and is set to... , Indicates the first The camera is in the synchronized frame group index. Observation quality parameters at that time Indicates the first The camera is in the synchronized frame group index. Time-based candidate-based extrinsic parameter update The calculated reference plane gaze point, Indicates the synchronization frame group index is The error reference plane gaze point is obtained by multi-camera weighted robust estimation. Denotes the Euclidean norm. The index is for the synchronization frame group participating in the statistics and only includes valid frames that have not been removed. S4.6 Multiple candidate extrinsic increments from the same camera are filtered according to their candidate confidence level. Candidate extrinsic increments with a confidence level greater than a preset confidence threshold are retained and bound to their candidate confidence levels to form a set of candidate extrinsic increments. Specifically: multiple candidate extrinsic increments from the same camera are filtered according to their candidate confidence levels. Filter and retain The candidate extrinsic increments are obtained, and each retained candidate extrinsic increment is bound to its corresponding candidate credibility and aggregated to form a candidate extrinsic increment set. S5. Construct an initial graph with the camera as a node and the relative pose prior and time alignment parameters as edges. In this embodiment, the method for constructing an initial graph using the camera as a node and the relative pose prior and time alignment parameters as edges is as follows: S5.1 A node is created for each camera in the candidate extrinsic parameter increment set, and node features are configured for the node. The node features include the camera's candidate extrinsic parameter increments, the candidate confidence level bound to the candidate extrinsic parameter increments, the camera's observation quality parameter statistics, the camera's initial camera extrinsic parameters, and the node's effective observation ratio. Specifically: for each camera in the candidate extrinsic parameter increment set... Establish nodes Configure node characteristics for this node The node features It consists of four types of information pieced together in a fixed order with fixed dimensions. The first type of information is what the camera retains. The increment of the candidate extrinsic parameter and its candidate confidence will be used to determine the first extrinsic parameter increment. Rotation increment of each candidate extrinsic parameter increment Transformed into axis-angle vectors via Lie algebra logarithmic mapping and displacement increment and candidate credibility The concatenation yields a candidate description vector of length 7. If the number of candidates for this camera after filtering is less than... Then, zero-candidate description vectors are used to fill in the blanks to maintain the consistency of node feature dimensions; The second type of information is the statistical quantity of observation quality parameters, which is derived from the index of all valid synchronization frame groups of that camera in the camera observation set that participated in the statistics in step S4. Calculate the observation quality parameters The mean and standard deviation are denoted as and , respectively. and ,in Used to characterize the overall observation reliability of the camera and Used to characterize the fluctuations in the observation quality of this camera; The third type of information is the camera's initial external parameters. The compact representation will The initial rotation matrix is partially converted to a unit quaternion. Let the initial translation vector be denoted as Then and The concatenation ensures that rotation and translation information can be directly used by the graph neural network at the node level and is consistent with the parameterization of subsequent incremental updates of extrinsic parameters; The fourth type of information is the proportion of effective observations at each node. The effective observation ratio of the nodes Defined as the camera satisfying within a session The ratio of the number of synchronized frame groups to the total number of synchronized frame groups in the session is used to reflect the coverage of the camera that is available for observation during the session; S5.2 For any two cameras, determine the relative pose prior between the two cameras based on the camera installation prior, and calculate the time alignment parameters between the two cameras based on the timestamp information. The time alignment parameters include at least a time offset. Specifically, regarding edge establishment, for any two cameras... and First, read the camera's relative pose prior from the installation configuration file and record it as... ,in, For from the camera Coordinate system to camera The homogeneous transformation of the coordinate system is given by the fixed geometric relationship of the installation structure and stored in the configuration file in the form of rotation quaternions and translation vectors; Then, based on the timestamp information, calculate the time alignment parameter, i.e., the time offset, between the two cameras and record it as . The Take the index of the same synchronization frame group from both cameras. The following timestamp difference sequence The median, to suppress occasional jitter, and The unit is unified to seconds to be consistent with the time interpolation module in the subsequent differentiable beam adjustment; S5.3 When the relative pose prior exists and the time alignment parameter meets the preset synchronization threshold, an edge is established between the corresponding two nodes, and edge features are configured for the edge. The edge features include the relative pose prior and the time alignment parameter. Specifically: an edge is established between the corresponding nodes if and only if the camera relative pose prior is available and the time alignment parameter meets the preset synchronization threshold. and Establish edges between The determination rule is written as follows: ; in, Represents the adjacency matrix elements and Indicates the establishment of edges, This indicates that no edge is established. The relative pose prior of the camera is indicated by a tag if and only if If it exists in the installation configuration file and passes verification Indicates an indicator function, Indicates camera With camera The time offset between them This indicates the preset synchronization threshold used in step S1, and its value is [value missing]. And in this step, it is converted to seconds for comparison; For each established edge Configure edge features The edge features Composed of camera relative pose prior and time alignment parameters, The rotational part is converted into an axis-angle vector. And the translation part is denoted as Then and Edge features are formed by splicing them together in a fixed order. The translation amount is normalized to 1 m and the time offset is normalized to 1 s to ensure that different units are on the same numerical scale in the subsequent graph neural network encoding. S5.4 The graph containing the node, the node feature, the edge, and the edge feature is determined as the initial graph. Specifically: defined as ,in, For a set of nodes, Given a set of edges, and configure node features for each node. Configure edge features for each edge. All node features constitute a set All edge features constitute a set ; S6, Perform graph neural network inference on the initial graph, and perform cross-camera consistency constraints and aggregation on the candidate extrinsic increments of each node through message passing and node updates to obtain the node fusion weights and globally consistent extrinsic increments of each node. In this embodiment, graph neural network inference is performed on the initial graph. Cross-camera consistency constraints and aggregation are applied to the candidate extrinsic parameter increments of each node through message passing and node updates to obtain the node fusion weights and globally consistent extrinsic parameter increments for each node. The method is as follows: S6.1 Encodes the node features of each node in the initial graph to generate an initial node representation for each node. Specifically: The graph network inference thread reads the initial graph. And perform graph neural network inference on it to obtain the graph network output, where This represents the initial graph. This represents a set of nodes, where each node corresponds to one camera. This represents a set of edges, where each edge corresponds to an established connection between two cameras. For each node Read its node features And generate the initial node representation through the node encoder. The node encoder is a two-layer multilayer perceptron with fixed parameters. The first layer is a fully connected mapping. Mapping to 128 dimensions and then connecting ReLU activation with LayerNorm, the second layer is a fully connected mapping that maintains the 128-dimensional parallel connection with LayerNorm, thus enabling... ; S6.2 Calculate the message passing weight of each edge based on the edge features in the initial graph. The message passing weight is calculated based on the edge features, the observation quality parameter statistics, the candidate confidence level, and the time offset. Specifically: for each edge... Read its edge features And generate edge representations through an edge encoder. The side encoder is a two-layer multilayer perceptron, and the first layer will... Mapping to 64 dimensions and connecting ReLU and LayerNorm, the second layer also maintains 64 dimensions and connects to LayerNorm, thus... ; Based on edge representation Calculate message passing weight The Output from the edge weight network and constrained by the Sigmoid function. The input to the edge weight network is The concatenated vector is output as a scalar, and the concatenated vector contains... and The first Mean and standard deviation of statistical parameters for observation quality of multiple cameras. For the first The first camera The candidate confidence level corresponding to each increment of candidate extrinsic parameters For camera With camera Time offset between; S6.3 Perform message passing and node representation updates on the initial graph according to a preset number of iterations. In each iteration, the node representations of adjacent nodes are aggregated under the constraint of the message passing weights, and combined with the node representation of the current node to generate an updated node representation. Specifically: Perform a preset number of iterations on the initial graph. Message passing and node representation updates, in the first In the next iteration, messages are aggregated based on neighboring nodes and node representations are updated. The aggregation and update are given by the following formula: ; in, Represents a node In the The aggregated message vector received in the next iteration. Indicates the initial graph In and nodes There exists a set of neighbor node indices connected by an edge. This represents the slave node calculated by the edge weight network. To the node Message passing weights, This represents the linear transformation matrix of the message and serves as the learnable parameters of the graph neural network. Represents a node In the The node representation vector for the next iteration. Represents a node In the The node representation vector updated in the next iteration. Indicates iterative index, This indicates that the gated loop unit updates the operator and that both its input dimension and hidden state dimension are 128 to ensure that the node representation dimension is consistent across different iterations; S6.4 Based on the node representation after completing the preset number of iterations, generate the node fusion weights for each node, and perform consistency correction on the candidate extrinsic parameter increments for each node to generate the corrected extrinsic parameter increments for each node. Specifically: complete... After the iteration, for each node Represent the final node Input the fusion weight header to generate node fusion weights. The fusion weight head is a two-layer fully connected network with a hidden layer dimension of 64, and it outputs a scalar score. Then, Softmax normalization is performed on the scores of all nodes to obtain the final weight. And satisfy This allows the node fusion weights to have an interpretable weight meaning in subsequent extrinsic parameter aggregation and gaze fusion; Simultaneously, consistency correction is performed on the candidate extrinsic parameter increments of each node to generate corrected extrinsic parameter increments, specifically by... Candidate confidence of this node Candidate weights are obtained by inputting a common candidate gating head. And perform Softmax normalization on the candidate set of the same node to obtain Then, the candidate rotation axis angle increments carried in the node features of step S5 are further processed. With candidate displacement increment according to The weighted summation yields the corrected extrinsic parameter increment for that node. ,in and The rotation axis angle vector representation and the three-dimensional translation vector representation are kept consistent with those in steps S4 and S5, respectively. S6.5 Aggregate the correction extrinsic increments based on the node fusion weights of each node to obtain the globally consistent extrinsic increments. Specifically: based on the node fusion weights of each node... Increment of correction extrinsic parameters for each node Perform weighted aggregation to obtain globally consistent extrinsic parameter increments And increment the globally consistent extrinsic parameters. Node fusion weight set Together they serve as the output of the graph neural network; S7. Update the initial camera extrinsic parameters according to the global consistent extrinsic parameter increment. Under the differentiable beam adjustment framework, iteratively optimize the camera extrinsic parameters to be optimized using an objective function to obtain the self-calibrated camera extrinsic parameters. The objective function includes a line-of-sight reprojection error term and a cross-camera line-of-sight consistency error term weighted by the node fusion weight and the observation quality parameter. In this embodiment, the initial camera extrinsic parameters are updated based on the globally consistent extrinsic parameter increment. Within the framework of differentiable beam adjustment, the objective function is used for iterative optimization of the camera extrinsic parameters to be optimized, resulting in self-calibrated camera extrinsic parameters. The method is as follows: S7.1 Update the initial camera extrinsic parameters of each camera based on the globally consistent extrinsic parameter increment, and generate the camera extrinsic parameters to be optimized. Specifically: The beam adjustment optimization thread reads the globally consistent extrinsic parameter increment of each camera. Node fusion weights And read the initial camera extrinsic parameters. With camera observation set; For each camera based on Perform an extrinsic parameter update to generate initial values for the extrinsic parameters of the camera to be optimized. In this case, rotation updates use axis-angle vectors. Incremental rotation obtained through exponential mapping and combined with The rotational part is left-multiplied for synthesis, and the displacement update uses... and The displacement portions are added together to ensure that... It remains a homogeneous transformation matrix from the camera coordinate system to the reference coordinate system; S7.2 During the iterative optimization process, based on the extrinsic parameters of the camera to be optimized, the gaze direction and the three-dimensional position of the eye at multiple moments are mapped to the reference coordinate system to obtain the intermediate gaze direction and the intermediate three-dimensional position of the eye. The intersection of the intermediate gaze direction and the reference plane is calculated to obtain the intermediate plane gaze point corresponding to each camera at each moment. Specifically: in the differentiable beam adjustment framework, each camera... The parameterized variable to be optimized is a rotation axis angle vector. With translation vector , and by The rotation matrix is constructed in real time through exponential mapping, thereby forming the rotation matrix in the iterative process. To participate in forward computation; In the forward computation of each iteration, for the synchronization frame group index First determine the unified alignment time The Take the timestamps of all cameras within the synchronization frame group. The median was used to suppress the impact of single-camera timestamp jitter on alignment; The iterative optimization process also includes joint optimization of the time alignment parameters: Using the time offset of each camera as a variable to be optimized, aligned observations are generated by interpolating at least one of the line-of-sight direction or the three-dimensional position of the eye. Specifically, a time offset is introduced for each camera. As a variable to be optimized and with a reference camera set. and time alignment parameters As The initial values are set to ensure consistent initial alignment during optimization; For each camera During alignment time The alignment observation is generated by locating a point on the timestamp axis of the camera's observation sequence that satisfies the following conditions: Adjacent frame indexes and Linear interpolation is performed on the gaze direction and the three-dimensional position of the eye, respectively. The interpolated gaze direction vector is then normalized to maintain the same semantics as the unit direction vector in step S3. At the same time, the observation quality parameter is subjected to equal-weighted linear interpolation to obtain the aligned... For error weighting; The interpolated 3D eye position and gaze direction are then used as current external parameters. Transform to the reference coordinate system to obtain the three-dimensional position of the central eye and the central line of sight, and generate the camera's index in the synchronized frame group by intersecting with the reference plane parameters. The intermediate plane gaze point below The reference plane is a point on the plane in the reference coordinate system. With unit normal vector Characterization, and when the direction of the intermediate line of sight is... Approximate orthogonality leads to the absolute value of the intersection denominator being less than If the intersection parameter is non-positive, causing the intersection point to be located in the opposite direction of the line of sight, the observation of that camera at that moment will be marked as invalid and removed from the error term at that moment; S7.3 For multiple intermediate plane gaze points at the same time, an error reference plane gaze point is generated at that time based on the node fusion weight, and the distance between the intermediate plane gaze point of each camera and the error reference plane gaze point is determined as the gaze reprojection error term. Specifically: for the same synchronization frame group index All valid Based on node fusion weight Generate error reference plane gaze point The generation method is based on all valid cameras. according to Normalized weighted summation and using the result as To ensure that it is consistent with the semantics of the fusion weights; Line reprojection error term and The Euclidean distance between them is used, and the Huber robust kernel is employed to suppress anomalous frames; S7.4 For multiple line-of-sight directions at the same time, a cross-camera line-of-sight consistency error term is generated based on the angular difference between the line-of-sight directions in the reference coordinate system; the cross-camera observation consistency error term takes the index of the same synchronization frame group. The difference in the angle between the line of sight directions of different cameras in the reference coordinate system; S7.5 Based on the node fusion weights and the observation quality parameters, the line-of-sight reprojection error term and the cross-camera line-of-sight consistency error term are weighted to construct an objective function; specifically: differentiable calculation is achieved in dot product form, and a time offset regularization term is added to minimize the objective function, resulting in an objective function for automatic differentiation optimization: ; in, Indicates the objective function for adjusting the bundle. This indicates the synchronization frame group index, which is consistent with the synchronization frame group in step S1. and Indicates the camera index. Indicates the first The camera nodes are fused with weights that satisfy the condition that the summation over all cameras is equal to 1. Indicates the first The camera is in the synchronized frame group index. The observed quality parameters after time interpolation are consistent with the range of values in step S3. This indicates that the Huber robust kernel has its inflection point parameter set to... Indicates the first The camera is in the synchronized frame group index. The current external reference The gaze point of the intermediate plane obtained by intersecting with the reference plane. Indicates the synchronization frame group index is Each camera according to The weighted error reference plane gaze point, Denotes the Euclidean norm. Represents the weight of the cross-camera observation consistency error term and is set to Indicates the weight of the time offset regularization term and sets it to... Indicates the first The camera is in the synchronized frame group index. After time interpolation and by the current external parameter Rotate to the unit center line-of-sight vector in the reference coordinate system. Indicates the first The time offset of the camera is a variable to be optimized, and its unit is seconds. It will be cropped after each parameter update. To avoid gradient instability caused by interpolation exceeding the limit; S7.6 The objective function is iteratively optimized within the differentiable beam adjustment framework. The derivative of the objective function with respect to the extrinsic parameters of the camera to be optimized is calculated using automatic differentiation. The extrinsic parameters are then updated based on the derivative until a preset convergence condition is met, at which point the self-calibrated camera extrinsic parameters are output. Specifically, the differentiable beam adjustment framework is implemented using a computational graph based on automatic differentiation. Perform iterative updates of the Adam optimizer, where the learning rates for rotation and translation parameters are set to... And the time offset learning rate is set to The maximum number of iterations is set to 80, and the decrease in the objective function is calculated after each iteration. When the decrease in the objective function is less than 80 in two consecutive iterations, the result is considered a positive result. When the preset convergence condition is met, the iteration stops. After the iteration is complete, the final extrinsic parameters of each camera will be obtained. As the output of the self-calibrated camera's extrinsic parameters and the final time offset Write to the time alignment parameter storage area; S8. Based on the self-calibrated camera extrinsic parameters, the eye-tracking observation data is transformed to the reference coordinate system and intersected with the reference plane. The data is then weighted and fused according to the node fusion weight and the observation quality parameters to obtain the fused gaze result in the reference coordinate system. In this embodiment, the method for obtaining the fused gaze result in the reference coordinate system is as follows: First, determine the uniform alignment time of the synchronization frame group. ,in Take the timestamps of each camera within the synchronization frame group. the median of For the first The camera is in the synchronized frame group index. The timestamp information is linked to the image. Then, for each camera Calculate the target alignment time And select on the time axis of the camera observation sequence that satisfy Adjacent frame indexes and Perform linear interpolation to obtain the aligned unit line-of-sight direction. Aligned 3D eye position and the aligned observation quality parameters ,in The unit direction vector in the camera coordinate system. The coordinates are three-dimensional coordinates in the camera coordinate system, with the unit being meters. And it is consistent with the definition of step S3, and the line of sight obtained by interpolation is first linearly interpolated and then normalized to keep the semantics of the unit direction vector unchanged; when When the time range of the camera's observation sequence is exceeded or adjacent frames required for interpolation are missing, the camera is indexed as [missing information] in the synchronization frame group. This section is marked as invalid and will not participate in subsequent fusion. S8.1 Based on the self-calibrated camera extrinsic parameters, transform the line of sight and the three-dimensional position of the eye in the eye-tracking observation data from the camera coordinate system to the reference coordinate system; S8.2 Under the reference coordinate system, calculate the intersection of the line of sight direction and the reference plane at each time moment to generate the reference plane gaze point of each camera at each time moment; Specifically: for each valid camera, and pass Transform to the reference coordinate system to obtain the 3D position of the eye and the unit line-of-sight direction in the reference coordinate system, and intersect with the reference plane to generate the camera's index in the synchronized frame group. Reference plane gaze point The reference plane is defined by a point on the plane of the reference coordinate system. With unit normal vector Characterization, and The parameters are calculated from the simulator geometric parameters in step S1 and remain unchanged in this step; When the line of sight is approximately parallel to the reference plane, the absolute value of the intersection denominator is less than... If the intersection parameter is non-positive, causing the intersection point to be located in the opposite direction of the line of sight, then the index of that synchronization frame group of the camera is... The location is marked as invalid and not used. Write it into the fusion candidate set; S8.3 For multiple reference plane gaze points at the same time, a weighted fusion is performed based on the product of the node fusion weight corresponding to each camera and the observation quality parameter to generate the fused gaze point at that time. S8.4 Perform time-series smoothing on the fused gaze points to obtain the fused gaze result in the reference coordinate system; Specifically: the index for the same synchronization frame group is Download all valid cameras The node fusion weights and observation quality parameters are used for weighted fusion, followed by time-series smoothing to generate the fused gaze result. The fusion and smoothing are implemented using the following formula: ; in, Indicates the synchronization frame group index is The fused gaze point in the reference coordinate system is located on the reference plane. Indicates the synchronization frame group index is The set of valid camera indexes that participate in the fusion at the same time. Indicates the first The camera is in the synchronized frame group index. The fusion weights are defined as follows: Indicates the first The camera nodes are fused with weights, and the normalization constraint is satisfied for all cameras. Indicates the first The camera is in the synchronized frame group index. The observation quality parameters after time interpolation Indicates the first The camera is in the synchronized frame group index. The reference plane is the point of focus obtained by intersecting the line of sight in the reference coordinate system with the reference plane. Indicates the synchronization frame group index is The output shows a smooth, fused gaze point located on the reference plane. This indicates the smooth fusion gaze point of the previous synchronized frame group index. This represents the smoothing coefficient and is set to 0.25. When Initialize to To complete the smoother state initialization; Finally, each synchronization frame group is indexed as Corresponding output smooth fusion gaze point With alignment time Perform the binding, and simultaneously output the associated reference coordinate system identifier and reference plane parameters. To clarify the correspondence between the fused gaze results and the reference coordinate system; Therefore, through a two-layer preprocessing architecture of multi-camera synchronous acquisition and spatiotemporal alignment, multiple image sequences, timestamps, and simulator geometric parameters are obtained to complete face and eye localization, low-light enhancement, and denoising. This achieves accurate alignment of cross-camera eye-tracking observation data, significantly improving the robustness and completeness of eye feature extraction under complex conditions (large head movement range, low light) compared to traditional single-camera or calibration board-dependent methods. Furthermore, by leveraging a multi-hypothesis incremental generation mechanism integrating extrinsic parameter estimation, the system inputs time-aligned gaze direction, 3D eye position, and observation quality parameters, and performs weighted point set registration and weighted gaze alignment. The orientation alignment sub-model outputs candidate extrinsic parameter increments and their confidence levels in parallel. High-confidence increments are filtered using a reprojection residual decay function, breaking the local optimum trap of a single extrinsic parameter solver and significantly improving the reliability of extrinsic parameter initialization. Through graph neural network cross-camera consistency aggregation and node fusion weight generation, candidate extrinsic parameter increments, observation quality parameter statistics, relative pose priors, and time alignment parameters are input. Node representations are iteratively updated using message passing and gated loop units, outputting globally consistent extrinsic parameter increments and node fusion weights. This achieves collaborative correction of multi-camera extrinsic parameters, avoiding the pitfalls of traditional step-by-step methods. This addresses the shortcoming of independent camera calibration, which is prone to error accumulation. By leveraging a differentiable beam adjustment framework and a weighted objective function for joint optimization, and inputting globally consistent extrinsic parameter increments, node fusion weights, and observation quality parameters, an objective function containing line-of-sight reprojection error and cross-camera line-of-sight consistency error is constructed. A time offset regularization term is introduced for automatic differential iteration, outputting the extrinsic parameters of the self-calibrated camera. This precisely adapts to the real-time extrinsic parameter correction requirements in scenarios with large head movement ranges and fluctuating observation quality, overcoming the challenge of offline calibration's inability to dynamically adapt. Furthermore, it utilizes weighted fusion of node fusion weights and observation quality parameters, along with time series smoothing... This method takes as input on self-calibration extrinsic parameters, gaze direction, and 3D eye position, and outputs a stable fused gaze result in a reference coordinate system. It achieves adaptive fusion of gaze points from multiple cameras, avoiding gaze drift caused by single-camera jitter or occlusion. Simultaneously, the cascaded preprocessing of the low-light enhancement network (U-Net) and the denoising network (DnCNN) ensures eye texture contrast in degraded images. An adaptive weighting mechanism for pupil visibility and blink state in the observation quality parameters further reduces the interference of low-quality observations on extrinsic parameter estimation and fusion output, significantly reducing the workload of manual calibration, parameter tuning, and data cleaning. Furthermore, this method requires no dedicated calibration board and does not rely on manual assistance; self-calibration can be completed using only natural eye movements, lowering the barrier to system deployment and maintenance. It provides fully automated and robust technical support for multi-camera eye tracking in scenarios such as driving / flight simulation, human-computer interaction, and behavior analysis, promoting the engineering application of eye tracking systems from controlled laboratory environments to real-world complex scenarios.
[0014] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process and related explanations of the method described above can be found in the corresponding process in the foregoing system embodiments, and will not be repeated here.
[0015] Those skilled in the art will recognize that the modules and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the invention.
[0016] The terms “first”, “second”, etc., are used to distinguish similar objects, not to describe or indicate a specific order or sequence.
[0017] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent in such process, method, article, or apparatus / device.
[0018] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks, characterized in that, Includes the following steps: Acquire image sequences and timestamps from multiple cameras, determine the reference coordinate system and reference plane, and obtain a synchronization dataset; Face eye localization and preprocessing are performed on the image sequence in the synchronous dataset to obtain eye region images and initial camera extrinsic parameters; The eye region image is input into the eye movement estimation model based on the timestamp information in the synchronous dataset to obtain eye movement observation data including gaze direction, three-dimensional eye position, and observation quality parameters; For the eye-tracking observation data, the initial camera extrinsic parameters, and the reference plane, integrated extrinsic parameter estimation is performed: candidate extrinsic parameter increments are calculated using a weighted 3D point set registration sub-model and a weighted line-of-sight alignment sub-model, and candidate confidence is calculated and filtered based on the line-of-sight reprojection residuals to obtain a set of candidate extrinsic parameter increments; Construct an initial graph with the camera as a node and the relative pose prior and time alignment parameters as edges; Graph neural network inference is performed on the initial graph. Cross-camera consistency constraints and aggregation are performed on the candidate extrinsic increments of each node through message passing and node updates to obtain the node fusion weights and globally consistent extrinsic increments of each node. The initial camera extrinsic parameters are updated based on the global consistent extrinsic parameter increment. Under the framework of differentiable beam adjustment, the objective function is used to iteratively optimize the camera extrinsic parameters to be optimized, and the self-calibrated camera extrinsic parameters are obtained. The objective function includes a line-of-sight reprojection error term and a cross-camera line-of-sight consistency error term weighted by the node fusion weight and the observation quality parameter. Based on the self-calibrated camera extrinsic parameters, the eye-tracking observation data is transformed to a reference coordinate system and intersected with the reference plane. The data is then weighted and fused according to the node fusion weights and observation quality parameters to obtain the fused gaze result in the reference coordinate system.
2. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 1, characterized in that, The method for obtaining the synchronized dataset is as follows: At the start of the session, a synchronization acquisition trigger signal is sent to the multiple cameras; Timestamp information is generated for each frame of image captured by each camera based on the same common clock; When the timestamp information of different cameras meets the preset synchronization threshold, the images of each camera that meet the preset synchronization threshold are merged into a synchronization frame group. The origin and coordinate axis directions of the reference coordinate system are determined based on the pre-stored simulator geometric parameters, and the plane parameters corresponding to the instrument display surface are determined as the reference plane based on the simulator geometric parameters. The synchronization dataset is obtained by binding the synchronization frame group, the timestamp information, the reference coordinate system, and the reference plane.
3. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 1, characterized in that, The preprocessing includes: Face detection is performed frame-by-frame on the image sequences of each camera in the synchronized dataset to determine face regions; Key points of the eyes are extracted within the face area, and the boundaries of the eye area are determined based on the key points of the eyes, thereby cropping the image sequence to obtain the image of the eye area; The low-light enhancement processing based on a deep network is performed on the eye region image to improve the contrast of eye texture, and the low-light enhancement processing of the eye region image is then subjected to denoising processing to suppress the interference of noise on eye features. The initial pose of each camera relative to the reference coordinate system is determined based on the prior knowledge of camera installation, and the initial pose is converted into the initial camera extrinsic parameters.
4. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 1, characterized in that, The eye region image is input into the eye movement estimation model based on the timestamp information in the synchronous dataset to obtain eye movement observation data including gaze direction, three-dimensional eye position, and observation quality parameters. The method is as follows: The eye region images are sequentially input into the same eye-tracking estimation model according to the timestamp information in the synchronization dataset; the eye-tracking estimation model outputs the gaze direction and the three-dimensional position of the eye for each frame of the eye region image. The eye-tracking estimation model includes a deep learning model, the gaze direction is represented by a unit direction vector in the camera coordinate system, and the three-dimensional position of the eye is represented by three-dimensional coordinates in the camera coordinate system. Based on the confidence level output by the eye-tracking estimation model and the usability assessment results of the eye region image, observation quality parameters are generated. The usability assessment results include pupil visibility and blinking status. The eye movement observation data is generated by binding the line of sight, the three-dimensional position of the eye, and the corresponding observation quality parameters of each camera at the same time with the timestamp information.
5. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 1, characterized in that, Candidate extrinsic parameter increments are calculated using a weighted 3D point set registration sub-model and a weighted line-of-sight alignment sub-model. The candidate reliability is then calculated and filtered based on the line-of-sight reprojection residuals to obtain the candidate extrinsic parameter increment set. The method is as follows: The camera with the largest average observation quality parameter is selected as the reference camera, and the synchronous frame group with the observation quality parameter greater than the preset quality threshold is the valid frame. For each non-reference camera, the weighted 3D point set registration sub-model uses the 3D eye positions of the local camera and the reference camera in the same effective frame as 3D point pairs, and the product of the observation quality parameters of the two cameras as weights. It solves the relative rotation matrix and translation vector through weighted singular value decomposition and outputs the first candidate relative pose. The weighted gaze direction alignment sub-model uses the gaze directions of the local camera and the reference camera in the same effective frame as the unit direction vector pair. It obtains the relative rotation matrix by solving the weighted Wahba problem with the same weight, and then calculates the relative translation vector by combining the weighted centroid of the three-dimensional position of the eye to output the second candidate relative pose. The first candidate relative pose and the second candidate relative pose are combined with the initial camera extrinsic parameters of the reference camera, and then relativized with the initial camera extrinsic parameters of the local camera to obtain two candidate extrinsic parameter increments. For each candidate extrinsic parameter increment, the gaze direction and eye 3D position of each effective frame are transformed to the reference coordinate system using the initial camera extrinsic parameters and intersected with the reference plane to obtain the gaze point. At the same time, based on all camera gaze points, the weighted robust estimation error reference plane gaze point is calculated according to the observation quality parameters, the weighted gaze reprojection residual is calculated, and the candidate confidence level is generated according to the exponential decay function. Multiple candidate extrinsic increments from the same camera are filtered according to their candidate credibility. Candidate extrinsic increments with a candidate credibility greater than a preset credibility threshold are retained and bound to their candidate credibility to form a set of candidate extrinsic increments.
6. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 1, characterized in that, The method for constructing an initial graph using cameras as nodes and relative pose priors and time alignment parameters as edges is as follows: A node is established for each camera in the candidate extrinsic parameter increment set, and node features are configured for the node. The node features include the candidate extrinsic parameter increment of the camera, the candidate confidence level bound to the candidate extrinsic parameter increment, the observation quality parameter statistics of the camera, the initial camera extrinsic parameters of the camera, and the effective observation ratio of the node. For any two cameras, the relative pose prior between the two cameras is determined based on the camera installation prior, and the time alignment parameters between the two cameras are calculated based on the timestamp information. The time alignment parameters include at least a time offset. When the relative pose prior exists and the time alignment parameter meets the preset synchronization threshold, an edge is established between the corresponding two nodes, and edge features are configured for the edge, the edge features including the relative pose prior and the time alignment parameter; The graph containing the node, the node features, the edge, and the edge features is determined as the initial graph.
7. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 6, characterized in that, Graph neural network inference is performed on the initial graph. Cross-camera consistency constraints and aggregation are applied to the candidate extrinsic parameter increments of each node through message passing and node updates to obtain the node fusion weights and globally consistent extrinsic parameter increments for each node. The method is as follows: The node features of each node in the initial graph are encoded to generate an initial node representation for each node; The message passing weights of each edge are calculated based on the edge features in the initial graph. The message passing weights are calculated based on the edge features, the observation quality parameter statistics, the candidate confidence, and the time offset. The initial graph is processed by message passing and node representation updates according to a preset number of iterations. In each iteration, the node representations of adjacent nodes are aggregated under the constraint of the message passing weights and combined with the node representation of the current node to generate the updated node representation. Based on the node representation after completing the preset number of iterations, the node fusion weight of each node is generated, and the candidate extrinsic parameter increments of each node are consistently corrected to generate the corrected extrinsic parameter increments of each node. The correction extrinsic increments are aggregated based on the node fusion weights of each node to obtain the globally consistent extrinsic increments.
8. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 1, characterized in that, The initial camera extrinsic parameters are updated incrementally based on the globally consistent extrinsic parameters. Within the framework of differentiable beam adjustment, the objective function is used to iteratively optimize the camera extrinsic parameters to obtain the self-calibrated camera extrinsic parameters. The method is as follows: The initial camera extrinsic parameters of each camera are updated once based on the globally consistent extrinsic parameter increment, generating camera extrinsic parameters to be optimized; During the iterative optimization process, the gaze direction and the three-dimensional position of the eye at multiple moments are mapped to the reference coordinate system based on the extrinsic parameters of the camera to be optimized, so as to obtain the intermediate gaze direction and the intermediate three-dimensional position of the eye, and the intersection of the intermediate gaze direction and the reference plane is calculated to obtain the intermediate plane gaze point corresponding to each camera at each moment. For multiple intermediate plane gaze points at the same time, an error reference plane gaze point at that time is generated based on the node fusion weight, and the distance between the intermediate plane gaze point of each camera and the error reference plane gaze point is determined as the line of sight reprojection error term. For multiple line-of-sight directions at the same time, a cross-camera line-of-sight consistency error term is generated based on the angle difference between the line-of-sight directions in the reference coordinate system; The objective function is constructed by weighting the line-of-sight reprojection error term and the cross-camera line-of-sight consistency error term based on the node fusion weights and the observation quality parameters. The objective function is iteratively optimized within the differentiable beam adjustment framework. The derivative of the objective function with respect to the extrinsic parameters of the camera to be optimized is calculated by automatic differentiation, and the extrinsic parameters of the camera to be optimized are updated according to the derivative until the preset convergence condition is met, at which point the self-calibrated extrinsic parameters are output.
9. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 8, characterized in that, The iterative optimization process also includes joint optimization of the time alignment parameters: Using the time offset of each camera as a variable to be optimized, aligned observations are generated by interpolating the time of at least one of the line of sight or the three-dimensional position of the eye. A time offset regularization term is added to the objective function to minimize the objective function.
10. The multi-camera eye-tracking self-calibration fusion method based on ensemble learning and graph neural networks according to claim 1, characterized in that, The method for obtaining the fused gaze result in the reference coordinate system is as follows: Based on the self-calibrated camera extrinsic parameters, the line of sight and the three-dimensional position of the eye in the eye-tracking observation data are transformed from the camera coordinate system to the reference coordinate system; In the reference coordinate system, the intersection of the line of sight direction and the reference plane at each time moment is calculated to generate the reference plane gaze point of each camera at each time moment; For multiple reference plane gaze points at the same time, a weighted fusion is performed based on the product of the node fusion weight corresponding to each camera and the observation quality parameter to generate the fused gaze point at that time. The fused gaze points are subjected to time-series smoothing to obtain the fused gaze result in the reference coordinate system.
Citation Information
Patent Citations
Virtual-real fusion dynamic calibration method based on binocular eye movement tracking
CN121582525A