Eye movement detection method and system with multi-mode data fusion and dynamic visual target adjustment
Patent Information
- Application Number
- CN202610193307.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2026-02-06
- Filing Date
- 2026-02-10
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-02-10
AI Technical Summary
[0004]因此,本发明的目的在于提供一种多模数据融合与动态视靶调整的眼动检测方法及系统,通过端到端人工智能模型来对多模态头部传感数据进行融合处理,并基于处理结果获得高精度头部姿态;据此实时调整视靶位姿以减少非必要的眼动;最后利用头部姿态数据及根据眼部图像数据计算获得的眼球扭转角度进行分析,获得用于评估前庭功能的检测结果,以解决在头部运动不可避免的临床检测(如反向眼扭转试验OCR)中,因头部运动干扰导致眼球位置变化、虹膜图像质量下降及眼球扭转角度测量误差大、结果重复性差的问题
1、本申请取代了传统椭圆拟合法,通过人工智能模型实现对磁力计误差的自适应校正:具备抗噪声、抗离群点能力,可应对时变及非椭球畸变。同时实现磁力计、加速度计、陀螺仪的统一坐标系标定,解决跨传感器对准问题。
Smart Images

Figure CN122074962B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vestibular function assessment technology, and in particular to an eye movement detection method and system for multimodal data fusion and dynamic visual target adjustment. Background Technology
[0002] In vestibular function assessments, clinical testing systems used for vestibular function and oculomotor analysis (such as the reverse oculomotor test, OCR) typically require subjects to maintain continuous fixation on a visual target during the test to enable high-precision visual analysis such as iris feature extraction and oculomotor angle calculation. However, physiological head micro-movements in subjects and operator-guided head posture changes are often unavoidable, introducing additional eye position adjustments and eye image changes related to fixation maintenance during the test, reducing the measurement stability and repeatability of oculomotor parameters such as oculomotor angle. Current technologies have a series of shortcomings in addressing this challenge, as follows: During examinations, a fixed target mode is often used. In routine OCR and other examinations, the visual stimulus target (visual target) is usually fixed at a certain position on the screen or physical plane. When the operator (such as a doctor) manually turns the user's head along the Roll axis (left and right tilt) or in other directions, the subject often needs to make additional eye adjustments to maintain fixation on the fixed visual target. This changes the visible area of the iris, the degree of occlusion, and the distribution of reflected light spots in the eye image. Since the calculation of the eye's torsion angle usually relies on iris information for computer vision analysis, the requirements for iris exposure are higher when the eye position changes, and the interference is much greater than in other tests. In traditional detection methods, the visual target does not change position when the doctor turns the user's head. After the head position changes, the eye position changes significantly, often leading to a significant impact on iris exposure and unstable results. In addition, multiple tests on the same user cannot accurately reproduce the head movement trajectory, resulting in low consistency of results for the same subject across multiple tests, thus affecting the reliability of clinical assessments.
[0003] Furthermore, existing sensors for data acquisition mostly employ attitude tracking based on inertial measurement units (IMUs), but their performance has inherent limitations. For example, magnetometer calibration commonly relies on offline calibration methods such as ellipsoid fitting. This method assumes that magnetic interference is fixed, uniform, and ellipsoidally distributed. In real medical environments (including metal equipment, electric beds, and displays), magnetic field interference is time-varying, non-uniform, and non-ellipsoidally distorted, which traditional methods cannot effectively compensate for. Additionally, alignment is lacking: traditional calibration only addresses the magnetometer itself, failing to resolve the axial misalignment issue between the magnetometer and the accelerometer / gyroscope, leading to a lack of unified fusion foundations. Summary of the Invention
[0004] Therefore, the purpose of this invention is to provide an eye movement detection method and system for multimodal data fusion and dynamic visual target adjustment. This method uses an end-to-end artificial intelligence model to fuse multimodal head sensing data and obtains high-precision head posture based on the processing results. The visual target posture is then adjusted in real time to reduce unnecessary eye movements. Finally, the head posture data and the eye torsion angle calculated from eye image data are analyzed to obtain detection results for evaluating vestibular function. This addresses the problems of eye position changes, decreased iris image quality, large measurement errors in eye torsion angle, and poor repeatability in clinical tests where head movement is unavoidable (such as reverse eye torsion test OCR) due to head movement interference.
[0005] To achieve the above objectives, the present invention provides an eye-tracking detection method for multimodal data fusion and dynamic visual target adjustment, comprising the following steps: Collect multimodal head sensing data from the user, the multimodal head sensing data including data from at least two types of sensors; The multimodal head sensing data is input into an artificial intelligence model for fusion processing, and the user's head posture data is output, including roll angle. Based on the head posture data, the position adjustment data of the visual target on the display device is calculated, and the position of the visual target on the display device is dynamically updated based on the position adjustment data. Synchronously collect eye image data when the user observes the visual target; Based on the eye image data, identify iris features and calculate the eyeball torsion angle; Based on the head posture data and the eyeball torsion angle, the test results for evaluating vestibular function are output.
[0006] Optionally, the multimodal sensing data includes triaxial magnetometer data, triaxial acceleration data, and triaxial gyroscope data; Optionally, the artificial intelligence model includes: A temporal convolutional encoder is used to extract multi-scale temporal features from the multimodal head sensing data, respectively. A self-attention encoder is used to perform global analysis on the multi-scale temporal features, calculate the correlation between features at each time step based on multi-head self-attention, and output the encoded features after weighted fusion through a feedforward network. A regression head is used to output the head pose data or intermediate variables of the head pose data based on the encoded features. The temporal convolutional encoder includes multiple one-dimensional convolutional layers connected by residuals, and different one-dimensional convolutional layers correspond to different time windows to form multi-scale temporal features.
[0007] Optionally, the self-attention encoder is further configured to: In the process of weighted fusion of features from different sensors to generate the encoded features, when the consistency between the features corresponding to the data from the magnetometer and the features corresponding to the data from the accelerometer and gyroscope in the same period is lower than a threshold, the weight of the magnetometer features in the weighted fusion process is reduced to suppress the influence of abnormal magnetic interference on the encoded features.
[0008] Optionally, the artificial intelligence model further includes: The context modulation layer is used to receive external context features, conditionally modulate the encoded features, and output them to the regression head.
[0009] Optionally, the intermediate variables of the head posture data include correction parameters or corrected magnetic field vectors used for magnetometer calibration. The head pose data of the user obtained from the processing results based on the artificial intelligence model specifically includes: The triaxial magnetometer data is calibrated based on the calibration parameters and / or the calibrated magnetic field vector of the magnetometer. The calibrated triaxial magnetometer data, triaxial accelerometer data, and triaxial gyroscope data are fused using an extended Kalman filter algorithm to output a continuous attitude data sequence including three-dimensional spatial position parameters and Euler angle attitude information, which serves as the head attitude data.
[0010] Optionally, the artificial intelligence model includes a cascaded magnetometer correction model and a multi-sensor fusion attitude estimation model; wherein: The magnetometer correction model is used to receive the original magnetometer timing data, acceleration timing data, gyroscope timing data and context information, and output the corrected magnetometer timing data. The multi-sensor fusion attitude estimation model is used to input the corrected magnetochronous timing data, the acceleration timing data, the gyroscope timing data, and context information, and outputs the six-degree-of-freedom attitude of the head. The magnetometer correction model and the multi-sensor fusion attitude estimation model are jointly trained, and the loss function for joint training is set according to minimizing the estimation error of attitude information.
[0011] Optionally, the artificial intelligence model also outputs confidence information corresponding to the head pose data, and adaptively controls the update process of the visual target position based on the confidence information. The adaptive control includes any one or more of the following: Adjust the gain and / or smoothing coefficient for target position updates; Apply amplitude or rate of change constraints to the update of the target position; When the confidence level is below the threshold, freeze the update or use the target position from the previous time step.
[0012] Optionally, the eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment further includes a step of correcting the position of the visual target, the correction step including: The user's actual gaze point location is identified based on the aforementioned eye image data; Calculate the deviation between the actual gaze point position and the visual target position controlled based on the head posture data; The display position of the visual target is dynamically adjusted based on the deviation.
[0013] Another aspect of the present invention discloses an eye-tracking detection system for multimodal data fusion and dynamic visual target adjustment, comprising: The data acquisition module collects multimodal head sensing data of the user's head, which includes data from at least two types of sensors; The eye-tracking acquisition module is used to acquire the user's eye-tracking images; One or more data processors, the data processors being coupled to the data acquisition module and the eye movement, the data processors being configured to perform the steps of the eye movement detection method for multimodal data fusion and dynamic visual target adjustment as described above.
[0014] The eye-tracking detection method and system for multimodal data fusion and dynamic visual target adjustment disclosed in this application have at least one of the following advantages compared with the prior art: 1. This application replaces the traditional ellipse fitting method and achieves adaptive correction of magnetometer errors through an artificial intelligence model: it has the ability to resist noise and outliers, and can cope with time-varying and non-ellipsoidal distortions. At the same time, it realizes the unified coordinate system calibration of magnetometer, accelerometer and gyroscope, and solves the cross-sensor alignment problem.
[0015] 2. This application adjusts the position of the displayed visual target in real time based on changes in head posture. When the user's head rotates / translates, the system automatically adjusts the visual target to maintain a stable line of sight and eye position, avoiding calculation errors caused by changes in iris exposure. It has specific advantages in OCR (reverse eye torsion) testing, mainly in maintaining a constant eye position relative to the visual target when the user's head rotates around the roll axis, reducing changes in iris image illumination. This improves the stability and repeatability of eye torsion angle analysis and reduces interference from human factors.
[0016] This application achieves automatic compensation for head movement interference during testing, significantly reducing result fluctuations and greatly improving the repeatability and comparability of the tests. The system exhibits good robustness, adapting to time-varying, non-ellipsoidal distortion, and dynamic noise environments while maintaining stable output. This application implements automated calibration and synchronization control, reducing manual intervention, simplifying operational procedures, and improving the efficiency and consistency of clinical testing. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an embodiment of the eye-tracking detection method for multimodal data fusion and dynamic visual target adjustment of the present invention. Figure 2 This is a schematic diagram of an artificial intelligence model architecture in one embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an embodiment of the eye-tracking detection system for multimodal data fusion and dynamic visual target adjustment of the present invention. Detailed Implementation
[0018] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown, one embodiment of the present invention provides an eye-tracking detection method for multimodal data fusion and dynamic visual target adjustment, comprising the following steps: S1. Collect multimodal head sensing data of the user, wherein the multimodal head sensing data includes data from at least two types of sensors; Specifically, for example, the multimodal head sensing data collected is nine-axis inertial measurement data, including data from three types of sensors: a three-axis magnetometer, a three-axis accelerometer, and a three-axis gyroscope. In this embodiment, the head attitude data is first collected in real time by an integrated nine-axis inertial measurement unit (including a three-axis accelerometer, a three-axis gyroscope, and a three-axis magnetometer), and synchronized with a unified timestamp.
[0020] S2. Input the multimodal head sensing data into an artificial intelligence model for fusion processing, and obtain the user's head posture data based on the processing result of the artificial intelligence model. Preferably, the head posture data includes at least the roll angle. For example, by using an artificial intelligence model to unify and fuse the coordinates of the aforementioned nine-axis inertial measurement data, the six-degree-of-freedom attitude of the head can be estimated in real time.
[0021] The artificial intelligence model is a trained model that is trained offline and deployed in a terminal device or edge computing unit. The training process includes: collecting multimodal head sensing data and corresponding real head poses as training samples, iteratively updating the model parameters, so that the model can learn the mapping relationship between multimodal sensing data and head poses.
[0022] In another implementation, the AI model employs a convolutional and attention-based fusion structure to automatically compensate for non-ellipsoidal distortion, noise, and time-varying interference, outputting a corrected magnetic field vector, which is then aligned with accelerometer and gyroscope data. The nine-axis data, after being aligned to a unified coordinate system, generates real-time three-dimensional spatial position parameters and Euler angle attitude information, forming a stable six-dimensional motion state data stream.
[0023] This application does not limit the specific form of the artificial intelligence model, and provides the following three embodiments for illustration.
[0024] Example 1: The preferred artificial intelligence model in this example includes a cascaded magnetometer correction model and a multi-sensor fusion attitude estimation model. The magnetometer correction model is used to receive raw magnetometer timing data, acceleration timing data, gyroscope timing data and context information, and output corrected magnetometer timing data. The multi-sensor fusion attitude estimation model is used to input the corrected magnetochronous timing data, the acceleration timing data, the gyroscope timing data, and context information, and outputs the six-degree-of-freedom attitude of the head. In this first embodiment, both the magnetometer correction model and the multi-sensor fusion attitude estimation model are neural network-based artificial intelligence models. These models are used to extract temporal features and perform nonlinear mapping on multimodal head sensing data. Their network structure can be implemented using a combination of temporal convolution, self-attention, and regression layers, but is not limited to the specific structure of the artificial intelligence models described in the second or third embodiments below. In this embodiment, the magnetometer correction model and the multi-sensor fusion attitude estimation model employ a joint training method, using a joint loss function constructed based on the final head attitude estimation error to collaboratively constrain and optimize the parameters of the two sub-models.
[0025] Example 2, Artificial Intelligence Model, including: A temporal convolutional encoder is used to extract multi-scale temporal features from the multimodal head sensing data; Specifically, this temporal convolutional encoder is used to extract temporal features from windowed multimodal head sensing data. Since head movements during eye-tracking detection and vestibular assessment simultaneously include instantaneous rotations (e.g., short-duration rapid swaying) and longer-term motion trends, this embodiment extracts multi-scale features through one-dimensional convolutional layers with different time windows to represent short-term dynamic changes and longer-term motion trends, thereby improving the ability to express head posture change patterns.
[0026] A self-attention encoder is used to perform global analysis on the multi-scale temporal features, calculate the correlation between features at each time step based on multi-head self-attention, and output the encoded features after weighted fusion through a feedforward network. Specifically, the self-attention encoder is used to perform global modeling and cross-sensor consistency analysis of the multi-scale temporal features. Since magnetometers are prone to abnormal disturbances in complex electromagnetic environments, and accelerometers and gyroscopes have a strong constraint effect on attitude changes within short timescales, this embodiment uses multi-head self-attention to calculate the correlation between features at each time step and between different sensors. During the fusion process, the weights of relevant magnetometer features are adaptively adjusted based on the consistency results, thereby suppressing the impact of abnormal magnetic interference on the correction results.
[0027] A context modulation layer is used to receive external context features and conditionally modulate the encoded features; the context features include temperature information, detected duration information, current time and location, etc. Specifically, considering that the eye-tracking detection process may experience zero-bias drift due to increased detection time and sensor bias changes due to changes in ambient temperature, this embodiment utilizes a context modulation layer to modulate the encoded features based on contextual features, thereby enhancing the robustness of the model under different detection stages and environmental conditions.
[0028] The first regression head is used to obtain the magnetometer's correction parameters and / or the corrected magnetic field vector based on the modulated feature regression, thereby achieving adaptive magnetometer calibration. The correction parameters are used to compensate the original magnetometer data to obtain calibrated magnetometer data, which is then aligned with the accelerometer and gyroscope data before being input into the extended Kalman filter algorithm for fusion, so as to output a continuous attitude sequence as head attitude data.
[0029] Specifically, after obtaining the calibrated triaxial magnetometer data through an artificial intelligence model, the data then enters the subsequent sensor fusion processing flow, which includes: Coordinate alignment: Align the calibrated triaxial magnetometer data with the triaxial accelerometer data and triaxial gyroscope data; Extended Kalman Filter (EKF) Fusion: Multimodal head sensing data with unified coordinates are fused using the extended Kalman filter algorithm to output three-dimensional spatial position parameters and Euler angle attitude information in real time, forming a stable six-dimensional motion state data stream as head attitude data.
[0030] In this second embodiment, the artificial intelligence model uses the output magnetometer correction parameters and / or the corrected magnetic field vector as intermediate variables to generate head posture data. During model training, the intermediate variables are indirectly supervised by the error between the head posture data generated by the intermediate variables and the actual head posture data.
[0031] Example 3, as Figure 2 As shown, this embodiment is an improvement on embodiment two, including: A temporal convolutional encoder is used to extract multi-scale temporal features from the multimodal head sensing data; Specifically, since head movements in eye-tracking detection and vestibular function assessment are characterized by both short-term rapid changes and long-term movement trends, this application improves the ability to model dynamic changes in head posture by setting up a multi-scale temporal convolutional encoder to capture instantaneous changes, action segments, and overall movement trends of head movements in different time windows.
[0032] A self-attention encoder is used to perform global analysis on the multi-scale temporal features, calculate the correlation between features at each time step based on multi-head self-attention, and output the encoded features after weighted fusion through a feedforward network. Specifically, in eye-tracking detection applications, the detection equipment may be in a complex electromagnetic environment, and the magnetometer is susceptible to local magnetic interference. This application uses a self-attention mechanism to perform correlation analysis on features from different sensors. When the consistency between the magnetometer features and the features of the accelerometer and gyroscope is low, the weight of the magnetometer features in the fusion process is adaptively reduced to suppress the impact of abnormal magnetic interference on head posture estimation.
[0033] A context modulation layer is used to receive external context features and conditionally modulate the encoded features; During vestibular function testing, changes in equipment operating time, ambient temperature, and testing stage can affect the zero bias and noise characteristics of the sensor. Therefore, this application introduces a context modulation layer to conditionally modulate the encoded features based on external context features, thereby enhancing the stability of the model under different testing stages and environmental conditions.
[0034] The second regression head maps the modulated encoded features to the attitude output of the corresponding time step, thereby obtaining six-dimensional attitude data including three-dimensional spatial position parameters and rotational attitude, and forming a continuous attitude data sequence as head attitude data.
[0035] The network structure and feature fusion method of the artificial intelligence model in this application are not simply a general application of the model. Instead, they are designed to address technical issues in application scenarios such as rapid changes in head movement, inconsistent noise from multiple sensors, and magnetic interference, thereby improving the robustness and accuracy of head pose estimation in eye-tracking detection scenarios.
[0036] It should be noted that the temporal convolutional encoder of the artificial intelligence model in Embodiments 2 and 3 is used to extract multi-scale temporal features from the multimodal head sensing data; the temporal convolutional encoder includes at least three one-dimensional convolutional layers connected by a residual structure, wherein the first convolutional layer captures the instantaneous patterns in the multimodal head sensing data with a preset first time window and outputs the first scale features. The second convolutional layer integrates the first-scale features with a preset second time window and extracts action segments as the second-scale features; The third convolutional layer uses a preset third time window to perform global perception on all measurement data within the test period based on the extracted second-scale features, and obtains motion trends rich in semantic information as the third-scale features.
[0037] A self-attention encoder is used to perform global analysis on the multi-scale temporal features, calculate the correlation between features at each time step based on multi-head self-attention, and output the encoded features after weighted fusion through a feedforward network. The self-attention encoder performs a global analysis of the temporal features and calculates the correlation between features at each time step based on multi-head self-attention. The weighted fused features are then processed by a feedforward network to output the encoded features; this includes the following steps: Multi-scale temporal features are organized into a time series according to time steps (sampling time step size); For each time series, a first feature vector is used to calculate the association weights between it and all second feature vectors in the series. These association weights are dynamically determined based on the similarity between the first feature vector and each second feature vector. Based on these association weights, all second feature vectors in the feature series are weighted and fused to generate an updated feature vector corresponding to the first feature vector. When a magnetometer feature component in a certain second feature vector has a low similarity to accelerometer and gyroscope feature components from the same period, a lower weight value is assigned to the current second feature vector to suppress abnormal magnetometer data. The above steps are iteratively executed to generate a corresponding updated feature vector for each position in the time series, and the calibrated and fused encoded features are output.
[0038] In other words, the self-attention encoder is also configured to: during the process of weighted fusion of features from different sensors to generate the encoded features, when the consistency between the features corresponding to the data from the magnetometer and the features corresponding to the data from the accelerometer and gyroscope at the same time is lower than a threshold, reduce the weight of the magnetometer features in the weighted fusion process to suppress the influence of abnormal magnetic interference on the encoded features.
[0039] In this third embodiment, the artificial intelligence model directly regresses the six-degree-of-freedom pose data of the head in an end-to-end manner.
[0040] Examples and explanations of training schemes for the artificial intelligence models in the above embodiments. The artificial intelligence model is trained before being deployed in the eye-tracking detection system. The purpose of this training is to enable the AI model to extract feature representations related to head posture changes from multimodal head sensing data, thereby obtaining accurate head posture data directly or indirectly through intermediate variables. The following are the steps included in the model training process: (1) Training data acquisition and labeling: Collect multimodal head sensing data of multiple users under different head posture conditions. The multimodal head sensing data includes at least three-axis magnetometer data, three-axis accelerometer data and three-axis gyroscope data. At the same time, obtain the real head posture data at the corresponding time through a high-precision optical motion capture system, mechanical turntable or known posture reference device as supervision label.
[0041] (2) Data preprocessing and synchronization: The collected multimodal head sensing data is time-synchronized, outlier removal, noise filtering and normalization are performed, and the data is divided into time-series samples according to a fixed time window; the real head posture data is aligned with the multimodal head sensing data in the corresponding time window to form training sample pairs.
[0042] (3) Model output format and training objective: During model training, the artificial intelligence models in different embodiments can output results in different forms, including but not limited to: Six-DOF pose data of the head; Intermediate variables used to generate head posture data include at least magnetometer calibration parameters and / or calibrated magnetic field vectors.
[0043] Regardless of the direct output form of the model, its training objective is uniformly to minimize the error between the head pose data obtained based on the model output and the real head pose data, thereby ensuring the accuracy of the final data used for head pose estimation.
[0044] (4) Loss function construction and model optimization Construct the training loss function based on the model's output format: When the artificial intelligence model directly outputs head pose data, the loss function includes at least the pose error loss based on the predicted pose and the actual head pose. When the artificial intelligence model outputs intermediate variables used to generate head pose data, the loss function includes at least: The posture error loss between the head posture data generated based on the intermediate variables and the real head posture data is used as the main supervision signal. In addition, consistency constraints or regularization constraints are set for the intermediate variables themselves as auxiliary constraints. Among them, the consistency constraint loss is used to constrain the rationality of the intermediate variables in terms of physical meaning or time dimension.
[0045] In one implementation, when the intermediate variable is a magnetometer-calibrated magnetic field vector, the consistency constraint loss is a magnetic field vector consistency constraint loss, which is constructed based on at least one of the following: ① Continuity constraint on the magnitude of magnetic field vector change in adjacent time steps; ② Constraints based on the stability of the magnitude of the magnetic field vector after coordinate transformation; ③ Constraints based on the physical consistency between the magnetic field vector and the attitude calculation results.
[0046] Based on the loss function, the parameters of the artificial intelligence model are iteratively updated using the backpropagation algorithm until the loss function converges or reaches the preset training rounds, thus obtaining the trained artificial intelligence model.
[0047] (5) Model deployment and inference: The trained artificial intelligence model is deployed in the eye-tracking detection system. In the actual detection process, only the forward inference process is executed to output head pose data in real time or intermediate variables used to generate head pose data.
[0048] Through the above training methods, the artificial intelligence model can adaptively learn the correlation between multimodal head sensor data under the presence of magnetic interference, noise, and time-varying environmental factors, thereby improving the stability and accuracy of head posture estimation and providing a reliable data foundation for subsequent dynamic visual target adjustment and vestibular function assessment.
[0049] S3. Dynamically update the position of the visual target on the display device according to the head posture data, so that the position of the visual target is adjusted synchronously with the head movement; wherein, the synchronous adjustment is used to keep the visual target in the preset target gaze direction before and after the user's head posture changes. In this invention, the "preset target gaze direction" is used to characterize the directional reference that the user is expected to gaze continuously during eye movement detection. It is used as a target constraint for the dynamic adjustment of the visual target position, so that when the user's head posture changes, the visual target is still in the user's expected gaze direction, thereby reducing non-target eye movement interference generated in order to maintain gaze.
[0050] Conventional display devices are two-dimensional planar devices. Using this solution, the visual target moves synchronously with the user's head. This ensures that the user's gaze on the target after head movement remains consistent with the gaze in the reference state, and the target's gaze direction relative to the user's head remains consistent (within the allowable error range), thus keeping the visual target always directly in front of the user's head. For example, when the user's head deflects to the left, the system accordingly compensates for the adjustment of the visual target's display position, causing a corresponding lateral displacement of the target on the display device.
[0051] The preset target gaze direction can be determined by combining the relationship between the visual target and the user's head / eyes in a reference state (e.g., the initial zero-position posture or calibration posture before posture change). Since the visual target moves with the head, in the head reference coordinate system, in the reference state, the visual target is in the direction of the forward axis of the user's head, so the preset target gaze direction can be set to the direction of the forward axis of the user's head; or in the reference state, the direction from the center of the user's eyeball to the visual target; the present invention does not limit its specific representation.
[0052] It should be noted that the "preset target gaze direction" does not mean that the target remains stationary in the external space, nor does it limit the target to a simple translation that strictly follows head movement. Instead, by using the "preset target gaze direction" as a control constraint, the display position of the target on the display device is dynamically compensated when the user's head posture changes, so that the target is always located in the target gaze direction consistent with the reference state (within the allowable error range) before and after the head posture change.
[0053] Therefore, although the display position of the target on a two-dimensional display device may shift synchronously with changes in the user's head posture, this synchronous shift is merely one possible outcome of maintaining the target's gaze direction, and not the control method itself as defined in this invention. This application uses a "preset target gaze direction" as an evaluation criterion or control target to compensate for the display position of the target; any technical solution that can ensure the target is positioned in the preset target gaze direction before and after changes in head posture, regardless of whether it employs simple translation, direction calculation, or other equivalent methods, should be considered an equivalent implementation of this invention.
[0054] In one embodiment, the system determines a preset target gaze direction in a reference state and continuously acquires user head posture data during the detection process. Based on the posture change relationship between the current head posture and the reference state, the system calculates the position adjustment data of the target on the display device (e.g., 2D screen coordinate offset, projection angle coordinate offset, or rendering coordinate offset), and updates the display position of the target accordingly, so that the target remains located in the preset target gaze direction before and after the head posture change.
[0055] In one embodiment, the preset target gaze direction can be determined by an initial zero position: at the start of detection, the user's head is guided to be in a centered position and gaze at the target, and the direction from the center of the user's eyeball to the target at this time is recorded as the preset target gaze direction; of course, it can also be determined by calibration. It should be noted that the "preset target gaze direction" can be defined with reference to the head coordinate system or the eye coordinate system, or it can be defined as the reference direction using the eye coordinate system, the display device coordinate system, or the world coordinate system. By establishing a calibration mapping relationship with head posture and / or eye posture, its corresponding gaze direction in the head or eye coordinate system is determined, thus achieving consistent directional constraints when head posture changes. This invention does not limit the explicit establishment of a specific coordinate system during calculation; its key lies in achieving the technical effect of satisfying the "target gaze direction constraint" before and after changes in head posture through any equivalent method.
[0056] In one embodiment, dynamically updating the position of the visual target on the display device based on the head posture data includes: A three-dimensional coordinate system is established with the user's head as the origin. The front of the user is defined as the Z-axis, the horizontal direction as the X-axis, and the vertical direction as the Y-axis. The rotational attitude of the head is represented by Euler angles (Pitch, Yaw, Roll) in the six degrees of freedom of attitude, and the spatial displacement of the head is represented by the three-dimensional position (Δx, Δy, Δz). The position of the target is then calculated according to the following formula (1): Here, Xscreen and Yscreen represent the real-time position offset of the target in the screen coordinate system, and d is the distance between the user's head and the screen. When the head rotates around the Roll axis, the system compensates for the target image at the corresponding angle. To avoid instability caused by small-angle jitter, low-pass filtering and proportional constraints are added to the offset signal in the actual implementation (e.g., the maximum X and Y offsets do not exceed 5% of the screen width and height). Through this geometric projection and dynamic mapping, the six-dimensional posture data is smoothly transformed into two-dimensional screen coordinates, realizing synchronous compensation between head movement and the target, so that the target always maintains a relatively stable position visually, regardless of whether the user rotates or slightly translates.
[0057] S4. Synchronously collect eye image data when the user observes the target position; Specifically, eye-tracking devices worn by the user (such as video nystagmography or eye trackers) can simultaneously capture video of the user's eyes, obtaining eye video image frames. Alternatively, cameras positioned around the user can capture video images of the user's eyes. S5. Based on the video image data of the eyes, iris features are identified and the eyeball torsion angle is calculated; During the examination, when a change in head posture (especially along the Roll axis) is detected, the system dynamically adjusts the position and orientation of the visual target on the display end using the posture change. Through spatial projection calculations, the system maintains relative stability of the visual target within the user's field of vision, thereby reducing eye displacement caused by head movement. Simultaneously, the front-end camera module captures images of the user's eyes and calculates the eyeball torsion angle using an iris feature recognition network. It also incorporates head posture data for compensation analysis to achieve synchronous correlation between head movements and eye movements. The entire process runs automatically and continuously, ultimately ensuring that in scenarios such as reverse eye torsion test (OCR), user head movements do not cause additional visual disturbances, significantly improving iris exposure stability and the accuracy and repeatability of torsion angle measurement.
[0058] The process of identifying iris features and calculating the eyeball torsion angle based on video image data of the eye includes: The eye image is preprocessed to segment the iris region; The iris region is transformed from the image coordinate system to the polar coordinate system to obtain the iris unfolding diagram; The horizontal pixel offset between the current frame's iris unfolded image and the reference template unfolded image is calculated by registering them. The initial twist angle is then calculated based on the ratio of the horizontal pixel offset to the total pixel width. The roll axis rotation angle of the head is extracted from the six-degree-of-freedom posture data; the apparent rotation component caused by the roll axis rotation angle to the initial torsion angle is calculated; the apparent rotation component is subtracted from the initial torsion angle to obtain the physiological torsion angle of the eyeball.
[0059] S6. Based on the head posture data and the eyeball torsion angle, output the detection results for evaluating vestibular function.
[0060] Specifically, before the OCR detection begins, the system guides the user to maintain an upright head posture, and in this posture, it acquires eye images and head posture data as initial position data. This initial position data is used to determine the baseline value of the eye torsion angle and the zero reference point of the head posture. Subsequent changes in eye torsion angle and head posture acquired during the detection process are calculated relative to this initial position.
[0061] During the experiment, the user's left and right tilt angles, i.e., the roll angle φ of the head relative to gravity, are combined with the corresponding eye torsion angle data to calculate the OCR gain. Symbol explanation: OCR gain under left head tilt condition; OCR gain under right head tilt condition; : Represents the time average of the head roll angle within the left and right tilt steady-state ranges; , : These represent the time averages of the relative ocular torsion angle within the steady-state intervals of left and right tilt, respectively; OCR gain essentially means the intensity of the ocular torsional response caused by a unit of head roll and tilt (relative to gravity). It mainly reflects the function of the utricle-vestio-ocular reflex pathway. Unlike aVOR (vHIT), which is dominated by the semicircular canals, it is an indicator of otolith organ function. Therefore, the "magnitude" and "left-right symmetry" of the gain are the core of judging vestibular function.
[0062] Left-right asymmetry index: ; For example, if and If the difference is significant and the AI exceeds a preset threshold, the side with the smaller AI is identified as the suspected damaged side. This is because the function of the utricle or its afferent pathway is impaired on one side, resulting in asymmetrical contributions from the two eyes to the same head tilt stimulation. If the gains are balanced and stable on both sides, and the AI is less than the preset asymmetry threshold, it indicates that the bilateral utricle-ocular reflex pathways are intact and the vestibular function is normal.
[0063] In another embodiment of this application, based on the above embodiments, in order to eliminate the error of the prediction results and adapt to individual differences, this embodiment includes a closed-loop correction / correction step in addition to the method steps of the above embodiments: identifying the user's actual gaze point position based on eye image data; Calculate the deviation between the actual gaze point position and the visual target position controlled based on head posture data; The display position of the visual target is dynamically adjusted based on this deviation.
[0064] Preferably, for the multimodal head sensing data stream and the eye image data stream with different sampling frequencies, a timestamp-based interpolation algorithm is used for time axis alignment. In practical implementation, the sampling frequency of IMU sensors (such as accelerometers, gyroscopes, and magnetometers) is usually higher than the video frame rate of the eye-tracking camera module; for example, the IMU is 100Hz, while the video is 60Hz. To achieve time synchronization and fusion of the two data streams, the system uses a timestamp-based multi-source data alignment and interpolation algorithm. Specifically, all data streams are appended with a high-precision system timestamp during the acquisition phase as a unified time reference. During the fusion phase, when an eye-tracking video frame arrives, the system extracts the nearest multi-frame pose data from the IMU data buffer based on the frame's timestamp and calculates the head pose and position at that moment using linear interpolation or spline interpolation. Conversely, when continuous pose trajectories need to be generated, the IMU data can be downsampled and matched according to the time interval of the video frames. For time offsets caused by delay or jitter, the system introduces a sliding window time synchronization mechanism to dynamically align and compensate for errors in the time series of different modalities. In this way, even if the IMU and video sampling frequencies are different, temporally consistent and spatially unified fused data can be formed on the same time axis, ensuring that the head posture and eye image analysis results strictly correspond, and improving the accuracy of subsequent iris feature calculation and torsion angle analysis.
[0065] In another embodiment of this application, based on any of the above embodiments, the artificial intelligence model outputs confidence information to characterize the reliability of the attitude estimation, in addition to outputting head pose data. Based on the confidence information, the system adaptively controls the position update process of the target.
[0066] The confidence information is used to reflect the stability and reliability of the head attitude estimation result at the current moment. It can be a continuous value, a discrete level, or other parameter form that can characterize the reliability of the attitude estimation. For example, the confidence information can be any one of the following: confidence score c∈[0,1], uncertainty / variance σ, covariance matrix P, or discrete level (high / medium / low); of course, the confidence can correspond to the overall attitude or a single component (especially the roll angle).
[0067] Based on the confidence information, the system adaptively controls the target position update process to suppress erroneous target updates caused by attitude estimation errors when attitude estimation is unstable or its reliability is reduced, thereby reducing non-target eye movement components generated to maintain fixation. The following are some examples of adaptive control of the target position update process based on confidence information: (a) Confidence-based update gain adjustment In one embodiment, the adaptive control includes adjusting the update gain of the target position based on confidence information. When the confidence information is high, a larger update gain is used to enable the target position to respond quickly to changes in head posture; when the confidence information decreases, the update gain is reduced to decrease the amplitude of target position changes and suppress the impact of posture estimation jitter on target updates.
[0068] (ii) Confidence-based smoothing and amplitude limiting control In another embodiment, the adaptive control includes applying smoothing and / or limiting processing to the target position data. Specifically, when the confidence information is lower than a preset threshold, the target position sequence and / or target position update amount at consecutive time steps are low-pass filtered, and / or the target position change at adjacent time steps is limited to not exceeding a preset maximum displacement, thereby avoiding abrupt changes in target position due to abnormal attitude estimation.
[0069] (iii) Confidence-based freezing or delayed update strategies In another embodiment, when the confidence level is below a first preset threshold, the update of the target position data is paused or frozen, keeping the target in its previous display position. When the confidence level recovers to above a second preset threshold, the update of the target position data is resumed. In one embodiment, a hysteresis strategy with different first and second preset thresholds can be used to avoid frequent switching caused by fluctuations in confidence level around the threshold. By employing the above methods, in cases where attitude estimation reliability is insufficient, additional eye-position adjustment behavior caused by erroneous position updates is avoided.
[0070] In practical applications, the adaptive control may include a combination of one or more of the above methods, such as simultaneously performing update gain adjustment and amplitude limiting control, or preferentially freezing the update and superimposing smoothing processing when the confidence level is below a threshold. This invention does not limit the specific combination of adaptive control strategies.
[0071] like Figure 3 As shown, the present invention also provides an embodiment of an eye-tracking detection system for multimodal data fusion and dynamic visual target adjustment, comprising: The multimodal data acquisition module 10 is used to acquire multimodal head sensing data of the user's head, including three-axis magnetometer data, three-axis accelerometer data and three-axis gyroscope data. Eye-tracking acquisition module 20 is used to synchronously acquire images of the user's eyes; The pose estimation module 30 is used to perform coordinate unification and fusion of multimodal head sensing data through an artificial intelligence model, and obtain the user's head pose data based on the fusion processing result; preferably, the head pose data includes at least the roll angle. In one implementation, the artificial intelligence model is a time-series model, and the input is a sequence of inertial measurement data within a sliding time window. The architecture design of the artificial intelligence model can be found in the description of the model architecture in the foregoing method embodiments, and will not be repeated here. The target adjustment module 40 is used to dynamically update the position of the target on the display device according to the head posture data, so that the target position is adjusted synchronously with head movement; wherein, the synchronous adjustment is used to keep the target positioned in a preset target gaze direction before and after changes in the user's head posture. In one implementation, the preset target gaze direction is the projection direction of the forward axis of the head reference coordinate system onto the display plane, and the target is displayed at the display position corresponding to the projection direction.
[0072] The eye-tracking analysis module 50 is used to identify iris features and calculate the eye torsion angle based on synchronously acquired eye image data when the user observes the visual target position. The vestibular function assessment module 60 is used to output test results for assessing vestibular function based on the head posture data and the eyeball torsion angle.
[0073] In one implementation, an artificial intelligence model is used during calibration to improve the accuracy of the calibrated magnetometer data, thereby facilitating subsequent attitude assessment.
[0074] In another implementation, 9DOF (three-axis magnetometer data, three-axis accelerometer data, and three-axis gyroscope data) is used as input data to the artificial intelligence model for coordinate unification and fusion processing. This continuously and stably outputs new, unified 6D data consisting of three-dimensional coordinates and Euler angles, automatically achieving complete coordinate unification and synchronization of the magnetometer, gyroscope, and accelerometer data, thus facilitating subsequent feature fusion. Subsequently, the new, precise head posture data output by the model is used to synchronously adjust the position of the displayed visual target, guiding the user to adjust the target's position synchronously when the user's head rotates or translates, ensuring minimal changes in eye position.
[0075] The technical details of each functional module in this embodiment can be found in the corresponding descriptions in the foregoing method embodiments. To avoid repetition, they will not be repeated here.
[0076] The above system embodiments are described using functional modules. Another system embodiment of the present invention is described from the perspective of system devices, and is actually the same scheme as the above system embodiments. Specifically, it is as follows: The eye-tracking detection system with multimodal data fusion and dynamic visual target adjustment provided in this embodiment includes: The data acquisition module is used to collect multimodal head sensing data of the user's head. The multimodal head sensing data includes data from at least two types of sensors; for example, the multimodal head sensing data includes data from a three-axis magnetometer, a three-axis accelerometer, and a three-axis gyroscope. The eye-tracking acquisition module is used to capture images of the user's eyes; One or more data processors are coupled (i.e., communicated) to the data acquisition module, and the one or more data processors are configured to perform the eye-tracking detection method steps of multi-modal data fusion and dynamic visual target adjustment as described in any of the foregoing method embodiments.
[0077] In one implementation, the one or more data processors are configured to: The system receives multimodal head sensing data of the user collected by the data acquisition module, wherein the multimodal head sensing data includes data from at least two types of sensors; The multimodal head sensing data is input into an artificial intelligence model for fusion processing, and the user's head posture data is obtained based on the processing result of the artificial intelligence model. The head posture data includes at least the roll angle. The position of the visual target on the display device is dynamically updated according to the head posture data, so that the position of the visual target is adjusted synchronously with the head movement. The synchronous adjustment is used to keep the visual target in the preset target gaze direction before and after the user's head posture changes. Receive eye image data synchronously acquired by the eye-tracking acquisition module when the user observes the visual target; Based on the eye image data, identify iris features and calculate the eyeball torsion angle; Based on the head posture data and the eyeball torsion angle, the test results for evaluating vestibular function are output.
[0078] The system embodiments of the present invention correspond to the aforementioned method embodiments. The technical details in the method embodiments are also applicable to the system embodiments. To reduce repetition, they will not be repeated here.
[0079] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. An eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment, characterized in that, Includes the following steps: Collect multimodal head sensing data from the user, the multimodal head sensing data including data from at least two types of sensors; The multimodal head sensing data is input into an artificial intelligence model for fusion processing, and the user's head posture data is obtained based on the processing result of the artificial intelligence model; the position of the visual target on the display device is dynamically updated according to the head posture data, so that the position of the visual target is adjusted synchronously with the head movement; wherein, the synchronous adjustment is used to keep the visual target in the preset target gaze direction before and after the user's head posture changes. Simultaneously collect eye image data when the user observes the visual target; based on the eye image data collected after the visual target is adjusted, identify iris features and calculate the eyeball torsion angle; Based on the head posture data collected after the visual target adjustment and the eyeball torsion angle, the detection results used to evaluate vestibular function are obtained.
2. The eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment according to claim 1, characterized in that, The multimodal head sensing data includes three-axis magnetometer data, three-axis accelerometer data, and three-axis gyroscope data.
3. The eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment according to claim 2, characterized in that, The artificial intelligence model includes: A temporal convolutional encoder is used to extract multi-scale temporal features from the multimodal head sensing data, respectively. A self-attention encoder is used to perform global analysis on the multi-scale temporal features, calculate the correlation between features at each time step based on multi-head self-attention, and output the encoded features after weighted fusion through a feedforward network. A regression head is used to output the head pose data or intermediate variables of the head pose data based on the encoded features. The temporal convolutional encoder includes multiple one-dimensional convolutional layers connected by residuals, and different one-dimensional convolutional layers correspond to different time windows to form multi-scale temporal features.
4. The eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment according to claim 3, characterized in that, The self-attention encoder is also configured to: In the process of weighted fusion of features from different sensors to generate the encoded features, when the consistency between the features corresponding to the data from the magnetometer and the features corresponding to the data from the accelerometer and gyroscope in the same period is lower than a threshold, the weight of the magnetometer features in the weighted fusion process is reduced to suppress the influence of abnormal magnetic interference on the encoded features.
5. The eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment according to claim 3, characterized in that, The artificial intelligence model also includes: The context modulation layer is used to receive external context features, conditionally modulate the encoded features, and output them to the regression head.
6. The eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment according to any one of claims 3-5, characterized in that, The intermediate variables of the head posture data include correction parameters or corrected magnetic field vectors used for magnetometer calibration. The head pose data of the user obtained from the processing results based on the artificial intelligence model specifically includes: The triaxial magnetometer data is calibrated based on the calibration parameters and / or the calibrated magnetic field vector of the magnetometer. The calibrated triaxial magnetometer data, triaxial accelerometer data, and triaxial gyroscope data are fused using an extended Kalman filter algorithm to output a continuous attitude data sequence including three-dimensional spatial position parameters and Euler angle attitude information, which serves as the head attitude data.
7. The eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment according to claim 1, characterized in that, The artificial intelligence model includes a cascaded magnetometer correction model and a multi-sensor fusion attitude estimation model; The magnetometer correction model is used to receive the original magnetometer timing data, acceleration timing data, gyroscope timing data and context information, and output the corrected magnetometer timing data. The multi-sensor fusion attitude estimation model is used to input the corrected magnetochronous timing data, the acceleration timing data, the gyroscope timing data, and context information, and outputs the six-degree-of-freedom attitude of the head. The magnetometer correction model and the multi-sensor fusion attitude estimation model are jointly trained, and the loss function for joint training is set according to minimizing the estimation error of attitude information.
8. The eye-tracking detection method for multimodal data fusion and dynamic visual target adjustment according to any one of claims 1-5, characterized in that, The artificial intelligence model also outputs confidence information corresponding to the head pose data, and adaptively controls the update process of the visual target position based on the confidence information. The adaptive control includes any one or more of the following: Adjust the gain and / or smoothing coefficient for target position updates; Apply amplitude or rate of change constraints to the update of the target position; When the confidence level is below the threshold, freeze the update or use the target position from the previous time step.
9. The eye-tracking detection method for multi-modal data fusion and dynamic visual target adjustment according to claim 1, characterized in that, It also includes a step of correcting the position of the target, the correction step including: The user's actual gaze point location is identified based on the aforementioned eye image data; Calculate the deviation between the actual gaze point position and the visual target position controlled based on the head posture data; The display position of the visual target is dynamically adjusted based on the deviation.
10. An eye-tracking detection system for multimodal data fusion and dynamic visual target adjustment, characterized in that, include: The data acquisition module collects multimodal head sensing data of the user's head, which includes data from at least two types of sensors; The eye-tracking acquisition module is used to acquire the user's eye-tracking images; One or more data processors, the data processors being coupled to the data acquisition module and the eye-tracking acquisition module, the data processors being configured to perform the eye-tracking detection method of multimodal data fusion and dynamic visual target adjustment as described in any one of claims 1-9.
Citation Information
Patent Citations
Eye movement detection method and system and storage medium
CN122096685A