Eye Movement Event Detection Method Based on Multi-Scale Convolution

Through the combination of multi-scale convolution and recurrent neural network, the problem that a single-scale convolution kernel cannot extract small sample features is solved, and higher-precision eye movement event detection is achieved.

CN116386124BActive Publication Date: 2025-07-22XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310161594.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-07-22
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

The convolutional neural network of a single-scale convolution kernel cannot effectively extract small sample features in eye movement event detection, resulting in limited detection performance.

Method used

Using a multi-scale convolution method, the UNet model is used to perform multi-scale feature extraction and feature fusion, and combined with recurrent neural networks to simulate eye movement event sequences, using linear fully connected layer and Softmax for classification, and finally the detection performance is evaluated through event-level Cohen’s Kappa.

Benefits of technology

It improves the accuracy of eye movement event detection, can better learn the characteristic information of small samples in the sequence, so that the model has the same attention to large samples and small sample events, and improves the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386124B_ABST
    Figure CN116386124B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting eye movement events based on multi-scale convolution, comprising the following steps: Step 1, preprocessing of the eye movement sequence; Step 2, using a UNet model to perform multi-scale feature extraction and feature fusion on the differential eye movement sequence; Step 3, using a recurrent neural network to simulate the eye movement event sequence; Step 4, using a linear fully connected layer and Softmax to classify the sample points at each moment in the eye movement sequence into fixation, saccade, and post-saccadic oscillation, so as to achieve eye movement event detection; Step 5, using event-level Cohen's Kappa to evaluate the performance of the three classified eye movement events. The method of the present invention solves the problem that the performance of the eye movement event detection method is limited by the fact that a convolutional neural network with a single-scale convolution kernel cannot effectively extract the features of small-sample events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of eye movement event detection, and particularly relates to an eye movement event detection method based on multi-scale convolution. Background Art

[0002] The purpose of eye movement event detection is to accurately and robustly extract eye movement events such as fixation, saccade, and post-saccadic oscillation from the raw eye movement data extracted by an eye tracker. One of the key challenges is to learn the correlation of each eye movement event in the eye movement sequence and capture the temporal and spatial information related to each eye movement event. Traditional model-based detection methods rely on handcrafted features (such as eye position, velocity, acceleration, etc.). The detection effects of these traditional methods are not only limited by the reliability of the feature extraction method, but also difficult to handle the detection of multiple types of eye movement events. At the same time, such methods often have many tunable parameters or hard-coded parameters, and a large amount of empirical knowledge is required for parameter tuning.

[0003] The new development is the emergence of event detection methods based on machine learning techniques. Tafaj et al. (2012) proposed a machine learning method of Bayesian mixture model (BMM), which uses instantaneous velocity to learn the parameters of Gaussian distributions representing fixation and saccade. This method was developed for assistance during driving and was tested using driving data. Santini et al. (2015) used a different method to add the classification of smooth pursuit events to the BMM algorithm. The classification methods of fixation and saccade are the same as those of the BMM algorithm, and the probability of smooth pursuit is calculated through velocity and movement rate. The above methods still use hand-designed features, but only use machine learning methods in classifier design.

[0004] In recent years, the development of deep learning has been increasingly rapid. Its deep and data-driven architecture has led to very significant performance improvements in many tasks. However, in the field of eye movement event detection, the application of deep learning methods is relatively less. The earliest use of deep learning methods in the field of eye movement event detection was the eye movement event detection algorithm proposed by Hoppe and Bulling (2016). This algorithm consists of an end-to-end single-layer convolutional neural network, a max-pooling layer, and a fully connected layer, and is used to detect three eye movement events: fixation, saccade, and smooth pursuit. Startsev et al. (2018) proposed a 1D-CNN-BLSTM network, which combined various features and compared them with several state-of-the-art detection algorithms that only detect fixation and saccade and some detection algorithms for smooth pursuit on the GazeCom dataset. Most of the feature combinations were either competitive or superior to the competitors. Zemblys et al. (2019) proposed a network called gazeNet for sequence-to-sequence classification. This network consists of two convolutional layers with a convolutional kernel size of 2×11, three LSTM layers, and a fully connected layer to perform event classification on three eye movement events: fixation, saccade, and post-saccadic oscillation.

[0005] Currently, deep learning-based eye movement event detection methods generally use convolutional neural networks and LSTM and their variants as the backbone networks. Since the lengths of different events in the eye movement sequence vary, the duration of fixation is long and it contains many sample points, while the durations of saccade and post-saccadic oscillation are short and they contain relatively few sample points compared to fixation. Therefore, using a convolutional neural network with a single-scale convolutional kernel to extract features cannot effectively extract features for small-sample events. Summary of the Invention

[0006] The purpose of the present invention is to provide an eye movement event detection method based on multi-scale convolution, which solves the problem of restricting the performance of the eye movement event detection method caused by the inability of a convolutional neural network with a single-scale convolutional kernel to effectively extract features of small-sample events.

[0007] The technical solution adopted by the present invention is an eye movement event detection method based on multi-scale convolution, including the following steps:

[0008] Step 1, preprocessing of the eye movement sequence;

[0009] Step 2, using the UNet model to perform multi-scale feature extraction and feature fusion on the differential eye movement sequence;

[0010] Step 3, using a recurrent neural network to simulate the eye movement event sequence;

[0011] Step 4: Use the linear fully connected layer and Softmax to classify the sample points at each moment in the eye movement sequence into fixation, saccade, and post-saccade oscillation to achieve eye movement event detection;

[0012] Step 5: Use event-level Cohen's Kappa to evaluate the performance of the three classified eye movement events.

[0013] The present invention is also characterized in that

[0014] Step 1 is implemented as follows:

[0015] Step 1.1, select the public Lund2013 eye movement event detection dataset as the original eye movement sequence. For the original eye movement sequence, it is necessary to remove the smooth pursuit, blink and undefined eye movement events, so that the eye movement sequence only contains three events: fixation, saccade and post-saccade oscillation for training and testing;

[0016] Step 1.2, segment the original eye movement sequence after removing redundant events, so that each eye movement sequence contains only 100 sample points. This is because the original eye movement sequence is very long, which will cause a large amount of calculation for model training. Segmenting the sequence reduces the difficulty of model training. During the segmentation process, the sequence is cut in an overlap manner to obtain segmented eye movement sequences. The end of each segmented eye movement sequence overlaps with the beginning of the next segmented eye movement sequence by 10 sample points.

[0017] Step 1.3, perform a differential operation on the segmented eye movement sequence to obtain a differential eye movement sequence. The differential eye movement sequence is used as the input of the UNet network. The input segmented eye movement sequence first copies the first sample point at the front end of the sequence, so that the segmented eye movement sequence contains 101 sample points; wherein the segmented eye movement sequence is expressed as: [(x s0 ,y s0 ), (x s1 ,y s1 ), (x s2 ,y s2 ),…,(x s(m-1) ,y s(m-1) ), (x sm ,y sm )], the differential eye movement sequence after difference is expressed as: There are 100 sample points in total, and the difference calculation formula is:

[0018]

[0019]

[0020] where x sm and sm represents the coordinate value of the sample point at the mth moment of the segmented eye movement sequence, xs(m-1) and y s(m-1) represent the coordinate values of the sample point at the (m - 1)-th moment of the segmented eye movement sequence, and represent the coordinate values of the sample point at the m-th moment of the differential eye movement sequence.

[0021] In step 2, the UNet model consists of an encoder and a decoder. The encoder module is responsible for feature extraction and consists of 4 downsampling blocks. Each downsampling block is composed of two 3×5 convolutional kernels for convolution and a 3×5 pooling kernel for max pooling. The decoder module is responsible for restoring the original resolution and consists of 4 upsampling blocks. Each upsampling block is composed of a feature fusion operation between the feature vector generated by upsampling and the feature vector generated by the downsampling block at the same level on the left side, and a convolution operation with two 3×5 convolutional kernels. Among them, the feature vector obtained after upsampling in each upsampling block is fused with the feature vector of the same dimension generated by the downsampling block at the same level, so as to achieve the feature fusion of multi-scale convolution, enabling the model to have the same attention to large samples and small samples.

[0022] In step 2, zero padding is used in the convolutional layer of the downsampling block to keep the size of the sequence unchanged before and after convolution. After two layers of convolution in each downsampling block, the receptive fields reach 9, 13, 17, and 21 respectively. The sizes of the feature vectors output by each downsampling block are [2, 100, 32] (where 2 represents two channels in the horizontal and vertical directions of the eye movement sequence, 100 represents 100 sample points in the sequence length, and 32 represents the dimension of the feature vector), [2, 96, 64], [2, 92, 128], and [2, 88, 256].

[0023] In step 3, the recurrent neural network used has 3 layers, and each layer contains 64 neurons.

[0024] In step 4, the linear fully connected layer used has 1 layer, and the output categories are 3 categories, corresponding to three eye movement events: fixation, saccade, and post-saccadic oscillation respectively. The loss function used in the training process is the weighted cross-entropy loss function. The sample weight ratios of fixation, saccade, and post-saccadic oscillation in the training set are [0.8557, 0.1045, 0.0398], and the weights of the weighted cross-entropy loss function are calculated as [0.1443, 0.8955, 0.9602] according to the sample weight ratios in the training set.

[0025] In step 5, event-level Cohen's Kappa is used to evaluate the performance of the detection results during the evaluation process. Cohen's Kappa score is a measure between the ratings of two raters for the same signal, used to compare the degree of agreement between the algorithm and the manual evaluation. The formula is as follows:

[0026]

[0027] where p o is the relative observed agreement among raters, representing the fraction of samples labeled as positive relative to the total number of samples, and p e represents the probability obtained by randomly shuffling the submission results, and its formula is as follows:

[0028]

[0029] where N is the number of samples, k is the number of classes, and n k1 is the number of times rater 1 predicts class k, and n k2 is the number of times rater 2 predicts class k.

[0030] The beneficial effects of the present invention are as follows:

[0031] The method of the present invention uses a multi-scale convolution-based method to detect eye movement events, integrating the advantages of the UNet and recurrent neural networks. The UNet model performs multi-scale deep feature extraction and feature fusion, and the recurrent neural network is used to simulate the eye movement event sequence, responsible for detecting the start and offset of fixation, saccade, and post-saccadic oscillation sample points. Compared with the existing method using a single-scale convolutional neural network, the UNet model can better learn the feature information of small samples in the sequence, enabling the model to have the same attention to large-sample events with long durations and small-sample events with short durations, comprehensively improving the detection accuracy of eye movement event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the overall framework flowchart of the method of the present invention;

[0033] Figure 2 is the network structure diagram of the UNet layer used in the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0035] The present invention provides a method for detecting eye movement events based on multi-scale convolution, as Figure 1 shown, including the following steps:

[0036] Step 1, preprocessing of the original eye movement sequence;

[0037] Step 1 is specifically implemented according to the following steps:

[0038] Step 1.1: Select the publicly available Lund2013 eye movement event detection dataset as the original eye movement sequence. For the original eye movement sequence, smooth pursuits, blinks, and undefined eye movement events need to be removed so that the eye movement sequence only contains three events: fixations, saccades, and post-saccadic oscillations for training and testing;

[0039] Step 1.2: Segment the original eye movement sequence after removing redundant events so that each segment of the eye movement sequence contains only 100 sample points. This is because the long original eye movement sequence will lead to a large computational load in model training. After segmenting the sequence, the model training difficulty is reduced. During the segmentation process, it is cropped in an overlap manner to obtain the segmented eye movement sequence. The end of each segmented eye movement sequence overlaps 10 sample points with the beginning of the next segmented eye movement sequence;

[0040] Step 1.3: Perform a difference operation on the segmented eye movement sequence to obtain the differential eye movement sequence. The differential eye movement sequence is used as the input to the UNet network. The input segmented eye movement sequence first replicates the first and last sample points at the front of the sequence so that the segmented eye movement sequence contains 101 sample points; among them, the segmented eye movement sequence is expressed as: [(x s0 , y s0 ), (x s1 , y s1 ), (x s2 , y s2 ), …, (x s(m-1) , y s(m-1) ), (x sm , y sm )], and the differential eye movement sequence after differentiation is expressed as: A total of 100 sample points, where the difference calculation formula is:

[0041]

[0042]

[0043] Among them, x sm and y sm represent the coordinate values of the sample points at the m-th moment of the segmented eye movement sequence, x s(m-1) and y s(m-1) represent the coordinate values of the sample points at the (m - 1)-th moment of the segmented eye movement sequence, and represent the coordinate values of the sample points at the m-th moment of the differential eye movement sequence.

[0044] Step 2: Use the UNet model to perform multi-scale feature extraction and feature fusion on the differential eye movement sequence.

[0045] Such as Figure 2As shown in the figure, the UNet model consists of an encoder and a decoder. The encoder module is responsible for feature extraction and consists of 4 downsampling blocks. Each downsampling block is composed of two 3*5 convolutional kernels for convolution and a 3*5 pooling kernel for max pooling. Among them, zero padding is used in the convolutional layer of the downsampling block to keep the size of the sequence unchanged before and after convolution. After two layers of convolution in each downsampling block, the receptive fields reach 9, 13, 17, and 21 respectively. The sizes of the feature vectors output by each downsampling block are [2, 100, 32] (where 2 represents two channels in the horizontal and vertical directions of the eye movement sequence, 100 represents 100 sample points in the sequence length, and 32 represents the dimension of the feature vector), [2, 96, 64], [2, 92, 128], and [2, 88, 256]. The decoder module is responsible for restoring the original resolution and consists of 4 upsampling blocks. Each upsampling block is composed of a feature fusion operation between the feature vector generated by upsampling and the feature vector generated by the downsampling block at the same level on the left, and a convolution operation with two 3*5 convolutional kernels. Among them, after upsampling in each upsampling block, the feature vector obtained is fused with the feature vector of the same dimension generated by the downsampling block at the same level, so as to realize the feature fusion of multi-scale convolution, making the model pay the same attention to large samples and small samples.

[0046] Step 3: Use a recurrent neural network to simulate the eye movement event sequence

[0047] A recurrent neural network is used to simulate the eye movement event sequence, which is responsible for detecting the start and offset of fixation, saccade, and post-saccadic oscillation sample points. The recurrent neural network adds a forgetting mechanism on the basis of the RNN, selectively retaining or forgetting some previous data, and no longer using multiplication but addition to avoid the problem of gradient explosion. The recurrent neural network combines the input gate and forgetting gate in the LSTM into an update gate to control whether the previous memory information can continue to be retained to the current moment, and another gate is the reset gate to control how much past information the model should forget. The recurrent neural network used in this method has 3 layers, and each layer contains 64 neurons.

[0048] Step 4: Use a linear fully connected layer and Softmax to classify the sample points at each moment in the eye movement sequence into fixation, saccade, and post-saccadic oscillation, and realize eye movement event detection;

[0049] The linear fully connected layer used in Step 4 has 1 layer, and the output categories are 3 categories, corresponding to the three eye movement events of fixation, saccade, and post-saccadic oscillation respectively. The loss function used in the training process is the weighted cross-entropy loss function. The sample weight ratios of fixation, saccade, and post-saccadic oscillation in the training set are [0.8557, 0.1045, 0.0398], and the weights of the weighted cross-entropy loss function are calculated as [0.1443, 0.8955, 0.9602] according to the sample weight ratios in the training set.

[0050] Step 5: Use event-level Cohen's Kappa to evaluate the performance of the three classified eye movement events.

[0051] During the evaluation process of Step 5, use event-level Cohen's Kappa to evaluate the performance of the detection results. The Cohen's Kappa score is a measure between the ratings of the same signal by two raters, which is used to compare the degree of agreement between the algorithm and the manual evaluation. The formula is as follows:

[0052]

[0053] where p o is the relative observed agreement between raters, which can be considered a kind of accuracy, representing the fraction of positive samples relative to the total number of samples. p e represents the probability obtained by randomly shuffling the submission results. The formula is as follows:

[0054]

[0055] where N is the number of samples, k is the number of categories, n k1 is the number of times rater 1 predicts category k, and n k2 is the number of times rater 2 predicts category k. The event-level Cohen's Kappa evaluation method finds the events with the largest overlap, then matches the events in the test sequence and the true sequence. If the matched events belong to different categories, the matched events are marked as false positives or false negatives. If the matched events belong to the same category, the matched events are marked as true positives, and then the remaining unmatched events are marked as false positives or false negatives.

[0056] Example 1

[0057] The test set used in this Example 1 is 22 eye movement sequences divided from the Lund2013 dataset. Execute the above Step 1 to Step 5. Finally, the event-level Cohen's Kappa scores for fixation, saccade, and post-saccadic oscillation on this test set are 0.960, 0.954, and 0.784 respectively.

[0058] Table 1 Evaluation scores of each eye movement sequence on the test set

[0059]

[0060] Table 1 shows the evaluation scores of each eye movement sequence on the test set. Ke_Fixation represents the event-level Cohen's Kappa score for fixations, Ke_Saccade represents the event-level Cohen's Kappa score for saccades, and Ke_PSOs represents the event-level Cohen's Kappa score for post-saccadic oscillations.

[0061] Corresponding to the current state-of-the-art eye movement event detection method gazeNet, the event-level Cohen's Kappa scores for fixations, saccades, and post-saccadic oscillations are 0.959, 0.947, and 0.776 respectively. From the experimental results, this method can better learn the feature information of small samples in the sequence, solving the problem that the performance of the eye movement event detection method is limited by the inability of the convolutional neural network with a single-scale convolutional kernel to effectively extract the features of small sample events, enabling the model to have the same attention to fixations with long durations and saccades and post-saccadic oscillations with short durations, and comprehensively improving the detection accuracy of eye movement event detection.

Claims

1. An eye movement event detection method based on multi-scale convolution, characterized in that It includes the following steps: Step 1, preprocessing of the eye movement sequence; Step 2, using the UNet model to perform multi-scale feature extraction and feature fusion on the differential eye movement sequence; In step 2, the UNet model consists of an encoder and a decoder. The encoder module is responsible for feature extraction and is composed of 4 downsampling blocks. Each downsampling block consists of convolution with two convolution kernels and max pooling with one pooling kernel. The decoder module is responsible for restoring the original resolution and is composed of 4 upsampling blocks. Each upsampling block consists of a feature fusion operation between the feature vector generated by upsampling and the feature vector generated by the downsampling block at the same level, as well as convolution operations with two convolution kernels. Among them, the feature vector obtained after upsampling in each upsampling block is fused with the feature vector of the same dimension generated by the downsampling block at the same level, so as to achieve feature fusion of multi-scale convolution, enabling the model to have the same attention to large and small samples; Step 3, using a recurrent neural network to simulate the eye movement event sequence; In Step 3, the recurrent neural network used has 3 layers, and each layer contains 64 neurons; Step 4, using a linear fully connected layer and Softmax to classify the sample points at each moment in the eye movement sequence into fixation, saccade, and post-saccadic oscillation, and realizing eye movement event detection; Step 5, using event-level Cohen's Kappa to evaluate the performance of the three classified eye movement events.

2. The eye movement event detection method based on multi-scale convolution according to claim 1, wherein Step 1 is specifically implemented according to the following steps: Step 1.1, select the publicly available Lund2013 eye movement event detection dataset as the original eye movement sequence. For the original eye movement sequence, smooth pursuit, blinks, and undefined eye movement events need to be removed so that the eye movement sequence only contains three events: fixation, saccade, and post-saccadic oscillation for training and testing; Step 1.2, perform segmentation on the original eye movement sequence after removing redundant events so that each segment of the eye movement sequence only contains 100 sample points. The segmentation process is performed in an overlap manner to obtain the segmented eye movement sequence. The end of each segmented eye movement sequence overlaps 10 sample points with the beginning of the next segmented eye movement sequence; Step 1.3, perform a difference operation on the segmented eye movement sequence to obtain a differential eye movement sequence, and the differential eye movement sequence is used as the input of the UNet network. The input segmented eye movement sequence first replicates the first and last sample points at the front end of the sequence, so that the segmented eye movement sequence contains 101 sample points; among them, the segmented eye movement sequence is expressed as: , and the differential eye movement sequence after differentiation is expressed as: There are 100 sample points in total, and the difference calculation formula is: where and represent the coordinate values of the sample points at the th moment of the segmented eye movement sequence, and represent the coordinate values of the sample points at the th moment of the segmented eye movement sequence, and represent the coordinate values of the sample points at the th moment of the differential eye movement sequence.

3. The eye movement event detection method based on multi-scale convolution according to claim 1, characterized in that, In Step 2, zero padding is used in the convolutional layer of the downsampling block so that the size of the sequence remains unchanged before and after convolution; after two layers of convolution in each downsampling block, the receptive fields reach 9, 13, 17, 21 respectively, and the sizes of the feature vectors output by each downsampling block are [2, 100, 32], [2, 96, 64], [2, 92, 128], [2, 88, 256].

4. The eye movement event detection method based on multi-scale convolution according to claim 1, wherein In Step 4, the linear fully connected layer used has 1 layer, and the output categories are 3 categories, corresponding to the three eye movement events of fixation, saccade, and post-saccadic oscillation respectively. The loss function used in the training process is the weighted cross-entropy loss function. The sample weight ratios of fixation, saccade, and post-saccadic oscillation in the training set are [0.8557, 0.1045, 0.0398], and the weights of the weighted cross-entropy loss function are calculated according to the sample weight ratios in the training set as [0.1443, 0.8955, 0.9602].

5. The eye movement event detection method based on multi-scale convolution according to claim 4, characterized in that In Step 5, event-level Cohen's Kappa is used to evaluate the performance of the detection results during the evaluation process. Cohen's Kappa score is a measure between the ratings of the same signal by two raters, used to compare the degree of agreement between the algorithm and manual evaluation. The formula is as follows: Among them is the relative observed agreement among raters, representing the score of positive samples relative to the total number of samples, represents the probability obtained by randomly shuffling the submission results, and its formula is as follows: Among them is the number of samples, is the number of categories, is the number of times rater 1 predicts category k, is the number of times rater 2 predicts category k.

Citation Information

Patent Citations

  • Dynamic balance assessment method and apparatus, and device and medium

    WO2021184792A1

  • Arithmetic question marking system based on mixnet-yolov3 and convolutional recurrent neural network (CRNN)

    WO2022147965A1