Eye movement event detection method and system based on Bayesian optimization
Through the search matrix task based on the dual matrix paradigm and the Bayesian optimization 1DCNN-BLSTM hybrid model, the problem of insufficient subjectivity and accuracy of existing eye movement event detection is solved, efficient and comparable eye movement data acquisition and detection is achieved, and the accuracy of AD-assisted diagnosis is improved.
Patent Information
- Application Number
- CN202510895427.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-12
AI Technical Summary
The existing eye movement event detection technology relies on manual parameter adjustment threshold algorithms with strong subjectivity and low detection accuracy. The traditional model ignores the time series dependence of eye movement data, resulting in insufficient detection accuracy and lack of standardization of data acquisition paradigm, which affects the reliability of AD-assisted diagnosis.
Eye movement data is collected using a search matrix task based on the dual matrix paradigm, and combined with Bayesian optimized 1DCNN-BLSTM hybrid model, the model hyperparameters are automatically tuned through the dual-scale Kappa objective function to balance event and sample hierarchical detection accuracy.
It improves the accuracy and comparability of eye movement event detection, enhances the reliability of AD-assisted diagnosis, and provides more accurate eye movement analysis tools.
Smart Images

Figure CN120458493A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of eye movement event detection, and in particular to an eye movement event detection method and system based on Bayesian optimization. Background Art
[0002] Eye movement analysis plays a crucial role in the diagnosis of neurodegenerative diseases, where accurate detection of eye movement events such as fixations and saccades is crucial. Alzheimer's disease (AD) patients experience significant visual impairments due to degeneration of neural circuits. Automated analysis of eye movement data, a non-invasive biomarker, is crucial for early diagnosis of AD.
[0003] Existing eye movement event detection mostly relies on threshold-based algorithms, which require manual parameter adjustment, resulting in strong subjectivity and low detection accuracy. For example, when dealing with complex eye movement events such as post-saccade oscillations, traditional methods are prone to missed detections or misjudgments due to deviations in threshold setting. At the same time, the 2D convolutional network models used in some studies overly focus on spatial feature extraction, ignore the time series dependence of eye movement data, and have high computational overhead, making it difficult to efficiently capture dynamic eye movement patterns. In addition, the data acquisition paradigm lacks standardized design, and the traditional visual search task scenarios are complex, resulting in poor data comparability between different studies and an inability to specifically highlight the abnormal eye movement characteristics of AD patients, limiting the clinical application of eye movement analysis in AD auxiliary diagnosis. Summary of the Invention
[0004] To address these issues, this paper proposes a method and system for eye movement event detection based on Bayesian optimization. By employing a search matrix to acquire eye movement data, this method accurately captures dynamic eye movement patterns, avoiding the shortcomings of traditional data collection paradigms. Furthermore, Bayesian optimization is introduced to automatically optimize the parameters of the eye movement event detection algorithm, overcoming the subjectivity of manual parameter adjustment. This combination improves detection accuracy, enhances the comparability of data from different studies, and enhances the accuracy of eye movement analysis for AD diagnosis.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides an eye movement event detection method based on Bayesian optimization, comprising: A search matrix task was constructed based on a dual-matrix paradigm consisting of two uniformly distributed arrow matrices, with the target arrow pointing in the opposite direction to the remaining arrows and in random positions. Collecting the original eye movement data of the subjects in the search matrix task and performing preprocessing; The preprocessed eye movement data is input into a pre-trained 1DCNN-BLSTM hybrid model based on Bayesian optimization to obtain eye movement event detection results. The Bayesian optimization automatically tunes the model hyperparameters based on a two-scale Kappa objective function to balance event and sample-level detection accuracy.
[0006] In a second aspect, the present invention provides an eye movement event detection system based on Bayesian optimization, comprising: A visual task construction module is configured to construct a search matrix task based on a dual matrix paradigm; the dual matrix paradigm includes two uniformly distributed arrow matrices, with the target arrow facing in the opposite direction to the remaining arrows and in random positions; a data acquisition module configured to collect the original eye movement data of the subject in the search matrix task and perform preprocessing; The event detection module is configured to input the preprocessed eye movement data into a pretrained 1DCNN-BLSTM hybrid model based on Bayesian optimization to obtain eye movement event detection results; the Bayesian optimization automatically tunes the model hyperparameters based on the two-scale Kappa objective function to balance event and sample-level detection accuracy.
[0007] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the eye movement event detection method based on Bayesian optimization described in the first aspect.
[0008] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the eye movement event detection method based on Bayesian optimization described in the first aspect are implemented.
[0009] Compared with the prior art, the present invention has the following beneficial effects: (1) This paper constructs a search matrix task based on a dual-matrix paradigm, which can accurately capture dynamic eye movement patterns and avoid the defects of the traditional data collection paradigm, such as complex scenarios and incomparable data. Through two evenly distributed arrow matrices (the target arrow is in the opposite direction to the others and the position is random), it can carefully depict complex eye movement events such as post-saccade oscillations and highlight the abnormal eye movements of AD patients. After the data is collected, it is pre-processed and input into the pre-trained 1DCNN-BLSTM hybrid model. Combined with Bayesian optimization, the hyperparameters are automatically tuned based on the dual-scale Kappa objective function to balance the event and sample level detection accuracy, improve the accuracy of eye movement event detection, enhance the comparability of different research data, and provide a more reliable and accurate eye movement analysis solution for AD auxiliary diagnosis.
[0010] (2) The present invention introduces Bayesian optimization into the 1DCNN-BLSTM hybrid model. By constructing a probability model, the model hyperparameters can be intelligently iterated and optimized, replacing manual parameter adjustment and reducing subjectivity. When processing time series eye movement data, the parameters can be adaptively adjusted to improve detection accuracy. The dual-scale Kappa objective function balances the detection accuracy of the event and sample levels, solving the imbalance problem of traditional model indicators.
[0011] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.
[0013] Figure 1 A main flow chart of an eye movement event detection method based on Bayesian optimization provided by an embodiment of the present invention; Figure 2 A schematic diagram of a search array and data acquisition provided by an embodiment of the present invention; Figure 3 Waveform diagram of eye movement data before and after filtering provided by an embodiment of the present invention; Figure 4 Provides a schematic diagram of the overall framework of PSOs-Net for the embodiment of the present invention; Figure 5 A flowchart of extracting eye movement features of a patient provided by an embodiment of the present invention; Figure 6 A schematic diagram comparing hyperparameter search curves provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0015] Example 1 like Figure 1 As shown, this embodiment discloses an eye movement event detection method based on Bayesian optimization, comprising the following steps: S1: Construct a search matrix task based on a dual-matrix paradigm; the dual-matrix paradigm consists of two uniformly distributed arrow matrices, with the target arrow pointing in the opposite direction to the remaining arrows and in random positions; S2: collecting the original eye movement data of the subjects in the search matrix task and performing preprocessing; S3: Input the preprocessed eye movement data into a pre-trained 1DCNN-BLSTM hybrid model based on Bayesian optimization to obtain eye movement event detection results; the Bayesian optimization automatically tunes the model hyperparameters based on a two-scale Kappa objective function to balance event and sample-level detection accuracy.
[0016] Next, combine Figure 1 , a method for detecting eye movement events based on Bayesian optimization disclosed in this embodiment is described in detail.
[0017] (1) Data collection and preprocessing Compared with visual search in natural environment, this embodiment adopts Figure 2 Shown is a controlled paradigm using a simplified search array to examine eye movement patterns in patients with Alzheimer's disease (AD).
[0018] This examination format features two different search matrices (Grid-4 and Grid-6), consisting of 4×4 and 6×6 uniformly distributed arrays of arrows, respectively. On each trial, participants were required to identify a unique target arrow that faced in the opposite direction from all other arrows in the array. The specific orientation (left / right / up / down) and spatial location of the target were randomized on each trial using a uniform distribution to ensure variation across the eight repetitions of each condition. To prevent fatigue, rest intervals were implemented between blocks of Grid-4 and Grid-6 trials.
[0019] The search matrix, displayed as an experimental scenario on a computer screen, is used to collect eye movement coordinate data from patients observing the experimental scenario using an eye tracker and supporting software. This coordinate data constitutes a dataset. The search matrix simulates search behavior in a natural environment within the experimental scenario, making it easy to control interference variables and difficulty, and its arrow shape is simple and easy to recognize.
[0020] As an implementation, visual stimuli were presented on a 23.8-inch monitor (1920 × 1080 resolution) controlled by a high-performance laptop computer. Eye movement data were recorded at 250 Hz using a Tobii Pro Fusion 250 eye tracker.
[0021] Prior to the experiment, all participants completed a standard five-point calibration procedure to establish an accurate mapping of eye movements with eye tracker data collection to ensure data accuracy.
[0022] Data collection was conducted at the East Campus of the Affiliated Hospital of Shandong University of Traditional Chinese Medicine from June to December 2024. Exclusion criteria included individuals under 50 years of age or those with psychiatric disorders (such as anxiety) or vascular disease. Participants were divided into four groups: healthy young controls (YHC, n=39, MoCA ≥26), healthy elderly controls (HC, n=38, MoCA ≤26), mild cognitive impairment (MCI, n=20, 20≤MoCA≤25), and Alzheimer's disease (AD) patients (n=22, MoCA <19). The mean ages of the MCI and AD, HC, and YHC groups were 65, 60, and 26 years, respectively, with a male-to-female ratio of 15:27, 13:26, and 25:13, respectively. MoCA refers to the Montreal Cognitive Assessment score, a tool used to assess cognitive function. In this example, MoCA scores were used to divide and define different groups.
[0023] Raw eye movement data were filtered to ensure quality. Trials with more than 50% missing data or fewer than 101 samples were excluded. Missing values were interpolated using first-order linear interpolation, followed by noise reduction using a moving median filter (window size = 11, equivalent to 44 ms at 250 Hz). Figure 3 The waveform characteristics of eye movement data before and after filtering are shown.
[0024] Taking into account specific experimental scenarios, for example, in the eye movement detection of patients with Alzheimer's disease (AD), the researchers constructed the search matrix. When healthy young people and AD patients completed the task, the search matrix accurately captured the eye movement abnormalities unique to AD patients. For example, healthy young people may only need two saccades to locate the target arrow, while AD patients, due to attention deficits and impaired visual working memory, will frequently rescan (possibly requiring four or more saccades), accompanied by prolonged fixation time and a larger search area.
[0025] The search matrix proposed in this example controls experimental variables (such as target complexity and distractor distribution) through a standardized arrow matrix, making eye movement data comparable across subjects. Compared to traditional complex visual search tasks in natural scenes, the search matrix more intuitively highlights eye movement patterns in AD patients, such as "repeated saccades" and "inability to efficiently locate targets," thereby focusing on specific pathological features. This provides accurate data support for subsequent AD diagnosis using eye movement characteristics (such as saccade count and fixation duration).
[0026] 2. Bayesian Optimization Eye Movement Event Detection Model Figure 4This paper demonstrates the architecture of the Post-Saccadic Oscillation Eye Movement Event Detection Network (PSOs-Net) algorithm, a 1DCNN-BLSTM hybrid model based on Bayesian optimization. This algorithm combines the 1DCNN-BLSTM hybrid model (implemented in PyTorch) with Bayesian optimization to automatically and concurrently detect three different eye movement events: fixations, saccades, and post-saccadic oscillations (PSOs). Post-saccadic oscillations refer to involuntary, additional oscillatory movements that occur after a saccade (an eye movement that rapidly shifts the gaze from one target to another).
[0027] Figure 4 The hyperparameter configurations and their respective search space boundaries are clearly documented. Ten different hyperparameters are included, including the number of convolutional layers, kernel size, and number of recurrent layers, all set within reasonable ranges. Each convolutional layer may have a different number of neurons, while each recurrent layer has the same number of neurons.
[0028] To determine the optimal hyperparameter combination, this example uses the Scikit Optimize library and Gaussian process-based Bayesian optimization to perform 50 optimization iterations.
[0029] The 1DCNN-BLSTM hybrid model architecture includes: 1. A 1D CNN module with multiple convolutional layers sharing the same kernel size but with different numbers of neurons in each layer, thus achieving effective local feature extraction.
[0030] 2. A BLSTM module with uniform neuron counts across layers to capture temporal dependencies, using mean aggregation.
[0031] 3. Final fully connected layer that generates the eye movement event detection output.
[0032] Based on the 1DCNN-BLSTM model, the Bayesian optimization model's hyperparameters, such as the number of convolutional layers, the number of neurons in each layer, the convolution kernel size, the number of LSTM layers and neurons, and the optimizer parameters, are optimized through continuous iteration to find the optimal network hyperparameter configuration.
[0033] Bayesian optimization is a method for optimizing functions by constructing probabilistic models (such as Gaussian processes). It leverages prior knowledge and existing sampling point information to gradually narrow the search space and find optimal parameters. Furthermore, this embodiment proposes a dual-scale Kappa objective function, which evaluates model performance from both the event and sample levels, balancing detection accuracy at different levels. Thus, Bayesian optimization uses the dual-scale Kappa objective function as its optimization goal. By continuously adjusting model hyperparameters to optimize this objective function, the model achieves high accuracy and balance in eye movement event detection, enabling the model to better capture the characteristics of different types of eye movement events when processing eye movement data.
[0034] Among them, a composite objective function is designed based on the dual-scale metric: ; here, represents the event-level Kappa indicator, Denotes the sample-level Kappa indicator, and X defines the hyperparameter search space. This formula ensures a balanced optimization of detection accuracy and global classification performance by leveraging the geometric properties of the L2 norm, which penalizes unilateral dominance of either metric. Global convergence is achieved only when both metrics improve simultaneously.
[0035] The model outputs labeled eye movement events. Each sample point is assigned an eye movement event label, which can be a fixation, a saccade, or a post-saccadic oscillation (PSO). Eye movement features are further calculated based on these labeled events. The inter-group differences in these features are then analyzed (statistically) to produce the final disease diagnosis.
[0036] Traditional eye movement detection models (such as threshold-based algorithms and some deep learning models) often face an imbalance in performance. For example, they may perform well in event-level classification but suffer from large errors in sample-level detail annotation, or vice versa. This imbalance can lead to insufficient accuracy in detecting complex eye movement events (such as post-saccadic oscillations). Using this function as the objective function to guide hyperparameter search during Bayesian optimization, PSOs-Net achieves comprehensive improvements in both event-level and sample-level Kappa scores on datasets such as Lund2013.
[0037] 3. Extracting the patient’s eye movement features like Figure 5 As shown in Figure 2, eye movement events in the AD-VSEMD dataset are automatically annotated using PSOs-Net. The annotated data is preprocessed as follows: Gaze correction: Saccadic events with an amplitude of 1° and subsequent PSOs (if present) were reclassified as fixations to merge adjacent fixations that were fragmented by minor saccadic errors.
[0038] Label adjustment: PSOs without saccade events were relabeled as fixation events to correct the pronunciation errors caused by the algorithm.
[0039] Duration filtering: fixation events lasting 60 ms were excluded to ensure data robustness.
[0040] As shown in Table 1, 13 quantitative metrics were extracted from the annotated events, including 4 behavioral features and 9 eye movement features. Specifically, the 4 behavioral features include accuracy, number of gaze passes over the area of interest, reaction time, and search efficiency. The 9 eye movement features include pupil diameter, number of saccades, saccade speed, number of post-saccade oscillations, post-saccade oscillation frequency, fixation time, fixation rate, search area, and K coefficient. The visual search paradigm used for the AD-VSEMD dataset consists of two arrow matrix configurations (grid-4 and grid-6), each of which includes eight trials with random target locations. For unified analysis, this example aggregates the eight trials in each configuration and uses the average rather than the sum to calculate the composite metric.
[0041] Table 1 Disease-related quantitative indicators and their definitions;
[0042] This analysis approach was necessary because some participants' eye movement data did not meet the inclusion criteria for the trials. Because the number of valid trials thus varied across conditions, summing the data would introduce systematic biases in cross-condition comparisons. Averaging effectively normalized for these variations, ensuring comparability across conditions and robustness of the results.
[0043] (IV) Experimental setup and results analysis To verify the effectiveness of this example, a PSOs-Net eye movement detection model was trained and tested using the Lund2013 dataset (500Hz sampling rate). To prevent overfitting, this example implemented early stopping with validation every 25 training steps. The best-performing model was selected based on the validation score. Training was terminated if there was no improvement after 150 consecutive validation steps (approximately 3.5 epochs). Considering the limited size of the validation set (approximately 1 minute of data) and the potential risk of overfitting, this example used the average of the validation and test set objective functions as the Bayesian optimization objective. Specifically, after each iteration, the validation and test set data were fed into the model to obtain the model's detection metric (kappa score) on both datasets. The average of these scores was then taken as the final objective function score for the model at that iteration.
[0044] The model uses a weighted cross-entropy loss (class weights: gaze = 0.1443, saccade = 0.8955, PSOs = 0.9602, inversely proportional to class proportion) and RMSprop optimization (batch size = 100, maximum iterations = 50). A temporal context window of 100 samples is used, and data augmentation includes: ① Random Gaussian noise (0-0.4°) simulates measurement error.
[0045] ② Channel swapping within a batch to enhance diversity.
[0046] Also, two markup errors in Lund2013 were found and corrected: ①UH27_img_vy_labelled_RA.mat: 37 samples (8676-8748ms) were incorrectly labeled as smooth pursuit (correct: fixation).
[0047] ②UH29 img_Europe_labelled_RA.mat: 5 samples (2868-2876ms) showing the PSOs before the saccade (correct: fixation).
[0048] For cross-dataset validation, an equivalent feature scale mapping (EFSM) was developed to process 250Hz GazeCom data without resampling. EFSM assumes constant velocity between samples and scales coordinate offsets by 0.5 to match the 500Hz feature scale. This approach preserves the original signal characteristics when operating across different sampling frequencies. Essentially, equivalent feature scale mapping achieves resampling at the model input level, rather than at the data level, ensuring that 250Hz data can be correctly input to a model trained on 500Hz data.
[0049] 1. Input feature extraction Since eye movement events exhibit spatiotemporal invariance, absolute coordinates and timestamps do not contribute to classification efficiency. Therefore, the feature extraction of this embodiment focuses on relative kinematic metrics between consecutive samples, including velocity components. , acceleration component , their Euclidean norm and binocular parallax Together these form an 8-dimensional input feature vector x that captures only the motion dynamics necessary for event discrimination.
[0050] 2. Verification To alleviate class imbalance and avoid the accuracy paradox, this example adopts the revised Kappa (κ) metric, which evaluates performance at both the sample and event levels. The (κ) statistic is calculated as , where po represents the observed agreement (accuracy) and pe represents the expected chance agreement. This metric inherently penalizes prediction bias: when the model disproportionately favors the majority class, increasing pe reduces the κ value. This design ensures rigorous evaluation under imbalanced distributions by weighting the scores according to the class proportions.
[0051] 3. Comparison with existing best methods To demonstrate the power of Bayesian optimization, Figure 6 The convergence curves of Bayesian optimization and random search in the process of hyperparameter adjustment are compared. The x-axis represents the iteration count, the y-axis represents the objective function value, and each point represents the historical best value. Figure 6 As shown, the Bayesian optimization objective function decreases rapidly, approaching optimal performance (-1.414) early in the exploration phase and ultimately reaching an optimal value (-1.246) at iteration 40. In contrast, random search converges more slowly, reaching only a suboptimal value (-1.238), indicating a less efficient search. These results confirm that Bayesian optimization effectively navigates the high-dimensional hyperparameter space for local refinement.
[0052] Furthermore, a comparison of the optimal hyperparameter configurations showed that both methods converged to a similar neural network architecture: three convolutional layers and two recurrent layers, with a consistent kernel size of 5 in the convolutional layers. Notably, the distribution of neurons across the convolutional layers follows a U-shaped pattern of neurons in the intermediate layers compared to the first and last layers. This architecture may offer specific advantages for detecting PSOs in eye movement analysis. The deeper initial and final convolutional layers may enhance spatiotemporal feature extraction, while the narrower intermediate layers may act as informative filters, reducing dimensionality and preventing overfitting. These findings highlight the adaptability of the 1DCNN-BLSTM architecture for eye movement event detection.
[0053] Table 2 Experimental results on the Lund2013 test set;
[0054] Table 2 shows the performance of different algorithms on the Lund2013 test set. The Bayesian-optimized PSOs-Net achieves the highest score across all event types, with event-level Kappa values of 0.972 (gaze), 0.968 (saccade), and 0.786 (PSOs), and sample-level Kappa values of 0.898, 0.913, and 0.745, respectively. Compared to the state-of-the-art GazeUNet, PSOs-Net improves PSOs detection by 1.1% (event-level) and 1.8% (sample-level), and achieves additional gains on other metrics.
[0055] Table 3 Experimental results on the GazeCom dataset;
[0056] Table 3 evaluates the generalization capabilities of the models on the 250Hz GazeCom dataset, which lacks PSOs events and focuses only on gaze and saccade detection. PSOs-Net demonstrates strong generalization, especially after EFSM processing, which improves performance without resampling. Although the sample-level Kappa for gaze detection is slightly lower than that of the random search variant, the Bayesian-optimized PSOs-Net achieves the best results in all other metrics, reaffirming its superior adaptability.
[0057] 4. Abnormal visual search behavior in AD patients The Mann-Whitney U test (α = 0.05) was used to assess group differences in eye movement characteristics, which provides robustness to non-normal distribution and enhances statistical power with small sample sizes. Behavioral indicators - including accuracy (Acc), number of regions of interest ( ), reaction time (T), and search efficiency (Eff) - showed no significant differences between the MCI and AD groups (p>0.05), suggesting that behavioral characteristics alone may not be effective in distinguishing disease severity. However, when healthy young controls (YHC) were compared with healthy elderly controls (HC), HC with MCI, and HC with the combined MCI and AD group, significant differences in Nroi, T, and Eff were observed (p<0.05), while Acc remained consistent across these comparisons. Participants demonstrated high accuracy on the Grid-4 task, confirming appropriate task difficulty. YHC typically located the target within two scans, while HC required up to four scans, suggesting that the risk of visual neglect is age-related. AD patients rescanned significantly more frequently than HC (p<0.01), which may be due to attention deficits and impaired visual working memory.
[0058] Eye movement features provide further discriminative power, and saccade counts ( ) and post-saccadic oscillation frequency ( ) showed significant differences between MCI and AD (p = 0.026 and p = 0.043, respectively). Pupil diameter decreased gradually from YHC to HC to AD (p < 0.05), indicating increased cognitive load and autonomic dysfunction, consistent with previous studies. AD patients also showed prolonged fixation duration ( )、 and The increase in the number of visual acuity scores and the expansion of the search area (Asearch) (p < 0.05) reflect the inefficiency of visual processing, attentional disorder and attentional spatial fragmentation. It is worth noting that although the reaction time is comparable (p > 0.05), the difference between MCI and AD is not significant. and In contrast, fixation rate (FR) lost significance, suggesting that Tfix differences may be due in part to generalized slowing rather than disease-specific effects.
[0059] Task difficulty modulated these effects: in the more complex Grid-6 task, , Asearch (HC and MCI) and The significant difference between (MCI and AD) was reduced, likely due to a training effect on the simpler Grid-4 task. However, feature K (fixation-to-saccade ratio) newly differentiated HC from MCI and MCI & AD (p < 0.05), revealing a shift toward longer fixations and shorter saccades—a maladaptive strategy that increases search area while decreasing efficiency. These findings highlight the importance of multimodal eye-tracking analysis, as behavioral and eye movement features capture complementary aspects of cognitive decline. Short-term task exposure may partially mitigate deficits, but further research is needed to assess long-term training effects.
[0060] Table 4 Comparison of eye movement characteristic indicators between groups;
[0061] This paper proposes a novel Bayesian optimization framework for eye movement event detection, PSOs-Net, to aid in the diagnosis of Alzheimer's disease (AD). By combining a 1DCNN-BLSTM architecture with Bayesian hyperparameter tuning, the model of this paper demonstrates superior performance in detecting fixations, saccades, and post-saccadic oscillations (PSOs), compared to existing technologies. The optimized model achieves high event- and sample-level Kappa scores, particularly for fixation and saccade detection. Furthermore, through the proposed equivalent feature scale mapping (EFSM) technique, it demonstrates improved generalization across datasets with varying sampling frequencies. Application of this framework to the AD-VSEMD dataset reveals significant abnormalities in visual processing patterns in AD patients, including prolonged fixation duration, increased saccades, and altered visual search strategies, all of which are associated with cognitive decline and attention deficits.
[0062] Despite these advances, some limitations should be acknowledged. First, while PSOs-Net achieved high accuracy in detecting fixations and saccades, the detection performance of PSOs remained relatively poor for both event types. This suggests that further refinement of the model architecture or feature representation may be needed to improve PSOs’ detection accuracy. Second, the current dataset, while carefully curated, is limited in size, particularly in the MCI and AD patient populations. Larger and more diverse datasets would improve the robustness of the findings and allow for more nuanced subgroup analyses. Furthermore, the study’s controlled laboratory paradigm, while conducive to standardization, may not fully capture the complexity of natural visual search behavior, potentially limiting the generalizability of the results to real-world scenarios.
[0063] These findings highlight the potential of eye movement analysis as a non-invasive biomarker for AD, particularly when combined with advanced machine learning techniques. The discriminative power of eye movement features, such as saccade counts and fixation characteristics, highlights their utility in distinguishing healthy aging, MCI, and AD. Future studies should explore longitudinal designs to evaluate the predictive value of these metrics in tracking disease progression. Furthermore, integrating multimodal data, such as neuroimaging or genetic markers, could provide a more complete understanding of the neural mechanisms underlying these eye movement abnormalities. Ultimately, this work contributes to the growing body of evidence supporting the use of eye tracking in the diagnosis of neurodegenerative diseases and opens new avenues for the development of objective, data-driven diagnostic tools.
[0064] This specific embodiment addresses the shortcomings of existing eye movement event detection, which relies on manually adjusted threshold algorithms, data collection, and model building, and adopts a search matrix and Bayesian optimization. The search matrix can collect data from multiple sources and multiple time periods based on eye movement characteristics, comprehensively record eye movement trajectories, improve the data confusion problem of the traditional paradigm, and specifically present the abnormal eye movements of AD patients. For example, in complex visual search tasks, it can accurately identify the characteristics of AD patients' excessive attention to distractors. Bayesian optimization uses probabilistic models such as Gaussian processes to automatically search for optimal detection algorithm parameters based on a small number of experiments, reducing missed detections and misjudgments caused by manual threshold deviations. At the same time, compared to 2D convolutional networks, it considers the time series characteristics of eye movement data, dynamically optimizes parameters, and efficiently captures continuous eye movement changes. The two work together to significantly improve the accuracy of eye movement event detection, make different research data have a unified and comparable standard, and transform complex eye movement analysis into a practical clinical tool, providing an objective and accurate basis for early auxiliary diagnosis of AD, helping clinicians to identify AD patients earlier and more accurately.
[0065] Example 2 This embodiment provides an eye movement event detection system based on Bayesian optimization, including: A visual task construction module is configured to construct a search matrix task based on a dual matrix paradigm; the dual matrix paradigm includes two uniformly distributed arrow matrices, with the target arrow facing in the opposite direction to the remaining arrows and in random positions; a data acquisition module configured to collect the original eye movement data of the subject in the search matrix task and perform preprocessing; The event detection module is configured to input the preprocessed eye movement data into a pretrained 1DCNN-BLSTM hybrid model based on Bayesian optimization to obtain eye movement event detection results; the Bayesian optimization automatically tunes the model hyperparameters based on the two-scale Kappa objective function to balance event and sample-level detection accuracy.
[0066] Example 3 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the eye movement event detection method based on Bayesian optimization as described in the first embodiment above are implemented.
[0067] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the eye movement event detection method based on Bayesian optimization as described in the first embodiment above are implemented.
[0068] The steps or modules involved in Examples 2 to 4 above correspond to those in Example 1. For detailed implementations, please refer to the relevant description of Example 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media that includes one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to perform any method of the present invention.
[0069] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for detecting eye movement events based on Bayesian optimization, characterized in that: include: Constructing a search matrix task based on the dual-matrix paradigm; The dual-matrix paradigm includes two uniformly distributed arrow matrices, with the target arrow facing in the opposite direction to the remaining arrows and in random positions; Collecting the original eye movement data of the subjects in the search matrix task and performing preprocessing; The pre-processed eye movement data is input into the pre-trained 1DCNN-BLSTM hybrid model based on Bayesian optimization to obtain the eye movement event detection results; The Bayesian optimization automatically tunes the model hyperparameters based on the dual-scale Kappa objective function to balance event and sample-level detection accuracy.
2. The eye movement event detection method based on Bayesian optimization according to claim 1, characterized in that: The arrow directions of the arrow matrix include four types: left, right, up, and down. The target position is randomly generated according to a uniform distribution in each trial; the two arrow matrices have different sizes.
3. The eye movement event detection method based on Bayesian optimization according to claim 1, characterized in that: The hyperparameters targeted by the Bayesian optimization include: the number of convolutional layers, the number of neurons in each layer, the size of the convolution kernel; the number of LSTM layers and neurons; and the optimizer parameters, which achieve global optimization by constraining the search space boundaries.
4. The eye movement event detection method based on Bayesian optimization according to claim 1, characterized in that: After obtaining the eye movement event detection results, perform data correction, including: Reclassify saccades and subsequent PSOs whose amplitude does not exceed a preset angle as fixations; Relabel isolated PSOs events without associated saccades as fixations; Filters gaze events that last less than a preset time.
5. The eye movement event detection method based on Bayesian optimization according to claim 1, characterized in that: The dual-scale Kappa objective function is: ; in, represents the event-level Kappa index, represents the sample-level Kappa indicator, and X defines the hyperparameter search space.
6. The eye movement event detection method based on Bayesian optimization according to claim 1, characterized in that: It also includes the use of equivalent feature scale mapping when processing datasets with different sampling frequencies.
7. An eye movement event detection system based on Bayesian optimization, characterized in that: include: The visual task building module is configured to build a search matrix task based on the dual-matrix paradigm; The dual-matrix paradigm includes two uniformly distributed arrow matrices, with the target arrow facing in the opposite direction to the remaining arrows and in random positions; a data acquisition module configured to collect the original eye movement data of the subject in the search matrix task and perform preprocessing; The event detection module is configured to input the preprocessed eye movement data into a pretrained 1DCNN-BLSTM hybrid model based on Bayesian optimization to obtain eye movement event detection results; the Bayesian optimization automatically tunes the model hyperparameters based on the two-scale Kappa objective function to balance event and sample-level detection accuracy.
8. The eye movement event detection system based on Bayesian optimization according to claim 7, characterized in that: The arrow directions of the arrow matrix include four types: left, right, up, and down. The target position is randomly generated according to a uniform distribution in each trial; the two arrow matrices have different sizes.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the eye movement event detection method based on Bayesian optimization as described in any one of claims 1 to 6 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the eye movement event detection method based on Bayesian optimization as described in any one of claims 1 to 6 are implemented.