A High-Precision Eye Movement Behavior Detection Method Driven by the Fusion of Time Series and Time Frequency

By performing time-frequency conversion and feature fusion of eye movement data, combined with multiple parallel networks for modeling and prediction, the problem of failure to effectively integrate time and frequency information in traditional methods is solved, and the accuracy and robustness of eye movement behavior detection is improved.

CN119830125BActive Publication Date: 2025-07-25KUNSHAN INNOVATION RES INST OF XIAN UNIV OF ELECTRONIC SCI & TECH +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411865593.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-07-25
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The existing eye movement behavior detection technology fails to effectively integrate time and frequency information, making it difficult for complex eye movement behavior to be accurately identified and affects the detection accuracy.

Method used

By performing short-time Fourier transform on the original eye movement data sequence, the time domain data is converted into time-frequency domain information, and combining multiple parallel Mamba networks and convolutional neural networks to extract timing and time-frequency features, a bidirectional gated recurrent unit network is used for sequence modeling and classification prediction.

Benefits of technology

It significantly improves the accuracy and robustness of eye movement behavior detection, especially in the classification and detection effect of behaviors such as gaze, saccade and late tremor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830125B_ABST
    Figure CN119830125B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of eye movement behavior detection, and disclose a high-precision eye movement behavior detection method driven by time series and time-frequency fusion. The method includes: performing a short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected collected, and converting the original eye movement data sequence in the time domain into time-frequency domain information; using multiple parallel Mamba networks to extract features from the time series information in the original eye movement data sequence to obtain time series features, and at the same time using a convolutional neural network to extract features from the time-frequency domain information to obtain time-frequency features; splicing and fusing the time series features and the time-frequency features to obtain fused features; inputting the fused features into a bidirectional gated recurrent unit network for sequence modeling; using a fully connected layer to perform classification prediction on the output features of the bidirectional gated recurrent unit network to obtain the detection result of the eye movement behavior to be detected. This method effectively improves the accuracy and robustness of eye movement behavior detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of eye movement behavior detection, and particularly to a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency. Background Art

[0002] Eye movement behavior detection technology is playing an increasingly important role as a key research tool in multiple fields such as psychology, neuroscience, and human-computer interaction. By tracking and analyzing eye movements, researchers can reveal the attention distribution, cognitive processes, and visual information processing mechanisms of humans in different situations. This technology has been widely applied in multiple fields such as user interface design, advertising effect evaluation, driver fatigue detection, clinical psychology diagnosis, and educational research.

[0003] Most of the currently proposed eye movement behavior detection technologies rely on time series data and use traditional classification algorithms to identify and classify different eye movement events, such as fixation, saccade, and late tremor, etc. These technologies usually only focus on the time information in the sequence, process the structure and characteristics of the time series itself, and have achieved certain results.

[0004] However, in addition to the time dimension information, time series data also contains rich frequency information, and these time-frequency domain features are also of great value for accurately detecting eye movement behavior. Most traditional methods ignore the frequency domain information of time series data and fail to effectively fuse time and frequency information for eye movement behavior detection. The lack of time-frequency domain information will lead to some complex eye movement behaviors being difficult to be accurately identified, thereby affecting the overall detection accuracy of the eye movement behavior detection system. Therefore, how to combine time and frequency information to improve the accuracy and robustness of eye movement behavior detection has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, the present application proposes a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency, which effectively improves the accuracy and robustness of eye movement behavior detection by combining time series information and frequency domain information.

[0006] To achieve the above object, an embodiment of the present application proposes a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency, including the following steps: performing a short-time Fourier transform (Short Time Fourier Transform, abbreviated as STFT) on the original eye movement data sequence of the eye movement behavior to be detected collected, and converting the original eye movement data sequence in the time domain into time-frequency domain information; using multiple parallel Mamba networks to extract features from the time series information in the original eye movement data sequence to obtain time series features, and at the same time using a convolutional neural network to extract features from the time-frequency domain information to obtain time-frequency features; splicing and fusing the time series features and the time-frequency features to obtain fused features; inputting the fused features into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies; using a fully connected layer to classify and predict the output features of the bidirectional gated recurrent unit network to obtain the detection result of the eye movement behavior to be detected.

[0007] To achieve the above object, an embodiment of the present application also proposes a high-precision eye movement behavior detection system driven by the fusion of time series and time-frequency, including: an STFT preprocessing module for performing a short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected and converting the original eye movement data sequence in the time domain into time-frequency domain information; a time series extraction module for using multiple parallel Mamba networks to extract features from the time series information in the original eye movement data sequence to obtain time series features; a time-frequency extraction module for using a convolutional neural network to extract features from the time-frequency domain information to obtain time-frequency features; a splicing and fusing module for splicing and fusing the time series features and the time-frequency features to obtain fused features; a sequence modeling module for inputting the fused features into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies; a classification and prediction module for using a fully connected layer to classify and predict the output features of the bidirectional gated recurrent unit network to obtain the detection result of the eye movement behavior to be detected.

[0008] To achieve the above object, an embodiment of the present application also proposes an electronic device, where the electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency as described above.

[0009] To achieve the above object, an embodiment of the present application also proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it can implement a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency as described above.

[0010] An accurate eye movement behavior detection method driven by time series and time-frequency fusion proposed in an embodiment of the present application performs a short-time Fourier transform on the original eye movement data sequence, converting the original eye movement data sequence in the time domain into time-frequency domain information to extract time-frequency features. This effectively captures the frequency information not fully utilized in traditional methods, and the introduction of time-frequency domain information significantly improves the accuracy of eye movement behavior classification. Subsequently, multiple parallel Mamba networks are used to process the sequence information with long-range dependencies to extract time series features, and at the same time, a convolutional neural network is combined to extract time-frequency features. Through the parallel structure and the design of smaller convolutional kernels, the model can maintain the complex feature expression ability while reducing the model scale, reducing the number of parameters, and effectively improving the computing efficiency. After that, the time series features and time-frequency features are concatenated and fused to obtain fused features, and the fused features are input into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies. Finally, a fully connected layer is used to classify and predict the output features of the bidirectional gated recurrent unit network to obtain the detection result of the eye movement behavior to be detected, which not only solves the technical problem of integrating sequence and frequency information, but also significantly improves the accuracy and robustness of the classification and detection of eye movement behaviors such as fixation, saccade, and late tremor.

[0011] In some alternative embodiments, a short-time Fourier transform is performed on the original eye movement data sequence of the eye movement behavior to be detected collected, and the conversion of the original eye movement data sequence in the time domain into time-frequency domain information includes: performing a short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected collected, expanding the time information of the original eye movement data sequence to the frequency dimension, capturing the frequency change characteristics at different time points, so as to convert the original eye movement data sequence in the time domain into time-frequency domain information; in the process of performing the short-time Fourier transform, by setting the length and overlap degree of the time window, the continuity and integrity of the time-frequency domain information are ensured.

[0012] In some alternative embodiments, the multiple parallel Mamba networks are composed of 3 parallel branches, and each branch is composed of a linear layer and a state space model connected in series, and the number of channels of each branch is 32.

[0013] In some alternative embodiments, for the original eye movement data sequence with an input of X = {x1, x2, …, x t , …, x T}, after being mapped by the linear layer, the output of the linear layer is expressed by the formula:

[0014] h t = W·x t + b;

[0015] where W is the weight matrix of the linear transformation, b is the bias term, x t is the t-th item in the original eye movement data sequence, and ht is the t-th item of the output of the linear layer;

[0016] The output of the linear layer is further processed by a state space model, and the output of the state space model is expressed by the formula:

[0017] y t = C·s t + D·h t ;

[0018] s t = A·s t-1 + B·h t-1 ;

[0019] where A is the state transition matrix, B is the input mapping matrix, C and D are the observation matrices, and s t represents the hidden state at the current time step t, and s t-1 represents the hidden state at the previous time step t-1, and y t is the t-th item of the output of the state space model.

[0020] In some alternative embodiments, the convolutional kernel size used by the convolutional neural network is (3,7), the stride is (1,1), the number of layers is 3, and the number of channels in each layer is 32.

[0021] In some alternative embodiments, the concatenating and fusing the temporal feature and the time-frequency feature to obtain a fused feature includes: concatenating the outputs of the state space models of 3 parallel branches based on the channel dimension, and using a linear layer to map the number of channels after concatenation to 32 to obtain the mapped temporal feature; keeping the size of the length dimension of the time-frequency feature unchanged, fusing the width dimension into the channel dimension, and using a linear layer to map the fused channel dimension to 64 to obtain the mapped time-frequency feature; concatenating the mapped temporal feature and the mapped time-frequency feature based on the channel dimension to obtain the fused feature.

[0022] In some alternative embodiments, the bidirectional gated recurrent unit network is composed of 4 layers of bidirectional gated recurrent units, and the dimension of its hidden layer is 64, and the dimension of the output is also 64. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related art, the following will briefly introduce the drawings required for the description of the embodiments of the present application or the related art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1It is a flowchart of a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency in an embodiment of the present application;

[0025] Figure 2 It is a schematic diagram of the principle of a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency in an embodiment of the present application;

[0026] Figure 3 It is a schematic diagram of the structure of a high-precision eye movement behavior detection model provided in an embodiment of the present application;

[0027] Figure 4 It is a schematic diagram of the structure of the Mamba network provided in an embodiment of the present application;

[0028] Figure 5 It is a schematic diagram of the structure of a high-precision eye movement behavior detection system driven by the fusion of time series and time-frequency provided in another embodiment of the present application;

[0029] Figure 6 It is a schematic diagram of the structure of an electronic device provided in another embodiment of the present application. Detailed implementation manners

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are provided for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can still be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation on the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of not being contradictory.

[0031] An embodiment of the present application proposes a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency, which is applied to an electronic device. Among them, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described by taking the server as an example. The implementation details of a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency proposed in this embodiment are specifically described below. The following content is only relevant implementation details provided for convenient understanding and is not necessary for implementing this solution.

[0032] The specific process of a high-precision eye movement behavior detection method driven by the fusion of time series and time-frequency proposed in this embodiment can be as Figure 1 shown and includes:

[0033] Step 101: Perform a short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected, and convert the original eye movement data sequence in the time domain into time-frequency domain information.

[0034] In a specific implementation, when there is a need for eye movement behavior detection, the server first collects the original eye movement data sequence of the eye movement behavior to be detected, and then performs a short-time Fourier transform on the collected original eye movement data sequence of the eye movement behavior to be detected, thereby converting the original eye movement data sequence in the time domain into time-frequency domain information.

[0035] In an example, the server performs a short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected, expands the time information of the original eye movement data sequence to the frequency dimension, thereby capturing the frequency change characteristics at different time points, and further converting the original eye movement data sequence in the time domain into time-frequency domain information. During the process of performing the short-time Fourier transform, the server reasonably sets the length and overlap degree of the time window to ensure the continuity and integrity of the time-frequency domain information.

[0036] In an example, the server can perform a short-time Fourier transform on the original eye movement data sequence using a rectangular window with a size of 6 and a step size of 1.

[0037] Step 102: Use multiple parallel Mamba networks to extract features from the temporal information in the original eye movement data sequence to obtain temporal features, and at the same time use a convolutional neural network to extract features from the time-frequency domain information to obtain time-frequency features.

[0038] In a specific implementation, after the server obtains the time-frequency domain information, the data preprocessing process is completed. Next, it enters the eye movement behavior recognition process, including feature extraction, feature fusion, feature modeling, and classification prediction. This process is actually implemented by an eye movement behavior detection model. The structure of the eye movement behavior detection model can be as Figure 2 shown, and can be roughly divided into a feature extraction network, a feature fusion network, a feature modeling network, and a classification network. The feature extraction network consists of two parts, namely multiple parallel Mamba networks and a convolutional neural network. Using multiple parallel Mamba networks to extract features from the temporal information in the original eye movement data sequence can obtain temporal features, and using a convolutional neural network to extract features from the time-frequency domain information can obtain time-frequency features.

[0039] In an example, the specific structure of the Mamba network can be as Figure 3As shown, multiple parallel Mamba networks consist of 3 parallel branches, each branch is composed of a linear layer and a state space model in series, and the number of channels of each branch is 32. Multiple parallel Mamba networks can well handle the long-distance dependence characteristics of the original eye movement data sequence, thereby capturing the subtle and complex dynamic changes in the original eye movement data sequence.

[0040] For the original eye movement data sequence with input X = {x1, x2, …, x t , …, x T}, after being mapped by the linear layer, the output of the linear layer is expressed by the formula:

[0041] h t = W·x t + b;

[0042] Among them, W is the weight matrix of the linear transformation, b is the bias term, x t is the t-th item in the original eye movement data sequence, and h t is the t-th item of the output of the linear layer;

[0043] The output of the linear layer is further processed by the state space model, and the output of the state space model is expressed by the formula:

[0044] y t = C·s t + D·h t ;

[0045] s t = A·s t-1 + B·h t-1 ;

[0046] Among them, A is the state transition matrix, B is the input mapping matrix, C and D are the observation matrices, s t represents the hidden state at the current time step t, s t-1 represents the hidden state at the previous time step t - 1, and y t is the t-th item of the output of the state space model.

[0047] In one example, the convolution kernel size used by the convolutional neural network is (3, 7), the stride is (1, 1), the number of layers is 3, and the number of channels of each layer is 32. Through multi-level convolution operations, local patterns and multi-scale information of time-frequency features can be extracted, which is crucial for identifying eye movement changes within a short time.

[0048] Step 103, splice and fuse the temporal features and time-frequency features to obtain fused features.

[0049] In a specific implementation, after feature extraction is completed, the feature fusion process can be entered, which is implemented by the fusion network of the eye movement behavior detection model. The fusion network splices and fuses the temporal features and time-frequency features to obtain the fused features.

[0050] In one example, the fusion network first splices the outputs of the state space models of three parallel branches based on the channel dimension, and uses a linear layer to map the number of channels after splicing to 32 to obtain the mapped temporal features. After that, the time-frequency features keep the size of the length dimension unchanged, fuse the width dimension into the channel dimension, and use a linear layer to map the fused channel dimension to 64 to obtain the mapped time-frequency features. Finally, the mapped temporal features and the mapped time-frequency features are spliced based on the channel dimension to obtain the fused features. The feature fusion process maps the number of channels to a unified dimension, ensuring that the model can utilize the information of these two types of features simultaneously for a more comprehensive eye movement behavior analysis.

[0051] Step 104, input the fused features into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies.

[0052] In a specific implementation, after feature fusion is completed, feature modeling can be carried out, which is implemented by the feature modeling network. The feature modeling network selected is a bidirectional gated recurrent unit network. Inputting the fused features into the bidirectional gated recurrent unit network for sequence modeling can capture forward and backward dependencies and global dependencies.

[0053] In one example, the bidirectional gated recurrent unit network consists of 4 layers of bidirectional gated recurrent units (bi-GRU), the dimension of its hidden layer is 64, and the dimension of the output is also 64.

[0054] Step 105, use a fully connected layer to classify and predict the output features of the bidirectional gated recurrent unit network to obtain the detection result of the eye movement behavior to be detected.

[0055] In a specific implementation, after feature modeling is completed, the final classification and prediction can be carried out. The classification and prediction are implemented by a fully connected layer. Using the fully connected layer to classify and predict the output features of the bidirectional gated recurrent unit network can obtain the detection result of the eye movement behavior to be detected. The fully connected layer maps the high-dimensional features to a specific classification space (such as fixation, saccade, post-saccadic tremor, etc.), and its output is the specific classification probability, ensuring the accuracy and robustness of the eye movement behavior detection.

[0056] In one example, when iteratively training the eye movement behavior detection model, the cross-entropy loss function can be used to calculate the classification error to guide the model training. The cross-entropy loss function can effectively measure the difference between the model prediction result and the actual label, and through the error feedback mechanism, adjust the model parameters to finally obtain the trained eye movement behavior detection model.

[0057] In this embodiment, first, a short-time Fourier transform is performed on the original eye movement data sequence to convert the original eye movement data sequence in the time domain into time-frequency domain information for extracting time-frequency features. This effectively captures the frequency information that is not fully utilized in traditional eye movement behavior detection methods. The introduction of time-frequency domain information significantly improves the accuracy of eye movement behavior classification. Subsequently, multiple parallel Mamba networks are used to process the sequence information with long-range dependencies to extract temporal features, and at the same time, a convolutional neural network is combined to extract time-frequency features. Through the parallel structure and the design of a smaller convolutional kernel, the model can maintain the complex feature expression ability while reducing the model scale, reducing the number of parameters, and effectively improving the computational efficiency. Then, the temporal features and time-frequency features are concatenated and fused to obtain fused features, and the fused features are input into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies. Finally, a fully connected layer is used to classify and predict the output features of the bidirectional gated recurrent unit network to obtain the detection result of the eye movement behavior to be detected, which not only solves the technical problem of integrating sequence and frequency information but also significantly improves the accuracy and robustness of the classification and detection of eye movement behaviors such as fixation, saccade, and late tremor.

[0058] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process, but not changing the core design of its algorithm and process, are all within the protection scope of this application.

[0059] Another embodiment of this application proposes a high-precision eye movement behavior detection system driven by the fusion of time series and time-frequency. The following specifically describes the implementation details of a high-precision eye movement behavior detection system driven by the fusion of time series and time-frequency in this embodiment. The following content is only relevant implementation details provided for easy understanding and is not necessary for implementing this solution.

[0060] The specific structure of a high-precision eye movement behavior detection system driven by the fusion of time series and time-frequency proposed in this embodiment can be as Figure 5 shown, including: an STFT preprocessing module 201, a temporal feature extraction module 202, a time-frequency feature extraction module 203, a concatenation and fusion module 204, a sequence modeling module 205, and a classification and prediction module 206.

[0061] The STFT preprocessing module 201 is used to perform a short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected collected, and convert the original eye movement data sequence in the time domain into time-frequency domain information.

[0062] The timing extraction module 202 is used to extract the timing information features in the original eye movement data sequence by using multiple parallel Mamba networks to obtain timing features.

[0063] The time-frequency extraction module 203 is used to extract the time-frequency domain information features by using a convolutional neural network to obtain time-frequency features.

[0064] The splicing and fusion module 204 is used to splice and fuse the timing features and the time-frequency features to obtain fusion features.

[0065] The sequence modeling module 205 is used to input the fusion features into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies.

[0066] The classification and prediction module 206 is used to classify and predict the output features of the bidirectional gated recurrent unit network by using a fully connected layer to obtain the detection result of the eye movement behavior to be detected.

[0067] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or can be implemented as a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0068] It is not difficult to find that this embodiment is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details and technical effects mentioned in the above method embodiment are still valid in this embodiment, and in order to reduce repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.

[0069] In another embodiment, in order to verify the superiority of the high-precision eye movement behavior detection method driven by the fusion of timing and time-frequency (hereinafter referred to as OURS), we conducted a comparative experiment on OURS and other methods on the same dataset Lund2013, and the experimental results are shown in Table 1.

[0070] Table 1: Comparison of Event-level Cohen’s Kappa between OURS and other methods

[0071] Method Gaze Saccade Post-saccadic oscillation NH2010 0.639 0.798 0.350 MNH 0.837 0.759 0.598 IRF 0.780 0.848 0.616 gazeNet 0.959 0.947 0.776 OURS 0.969 0.950 0.803

[0072] As can be seen from Table 1, the accuracy of OURS in the three-class eye movement behavior detection of fixation, saccade, and post-saccadic oscillation is much higher than that of other methods. Especially in the detection of fixation, it has an improvement of more than 1.5% compared with the traditional method.

[0073] Another embodiment of the present application provides an electronic device, the structure of which may be as Figure 6 shown, including: at least one processor 301; and a memory 302 communicatively connected to the at least one processor 301; wherein the memory 302 stores instructions executable by the at least one processor 301, and the instructions are executed by the at least one processor 301 to enable the at least one processor 301 to execute a high-precision eye movement behavior detection method driven by time sequence and time-frequency fusion as described in the above method embodiments.

[0074] Wherein, the memory and the processor are connected by a bus. The bus may include any number of interconnected buses and bridges, and the bus connects the circuits of one or more processors and the memory together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and thus will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices over a transmission medium. The data processed by the processor is transmitted over a wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0075] The processor is responsible for managing the bus and general processing, and may also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory may be used to store data used by the processor when executing operations.

[0076] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program, which when executed by a processor, can implement a high-precision eye movement behavior detection method driven by time sequence and time-frequency fusion as described in the above method embodiments.

[0077] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions to enable a device (such as a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0078] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made to them in form and details without departing from the spirit and scope of the present application.

Claims

1. A high-precision eye movement behavior detection method driven by the fusion of time series and time frequency, characterized in that, Including: Performing short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected collected, and converting the original eye movement data sequence in the time domain into time-frequency domain information; Using multiple parallel Mamba networks to extract features from the temporal information in the original eye movement data sequence to obtain temporal features, and at the same time using a convolutional neural network to extract features from the time-frequency domain information to obtain time-frequency features; Concatenating and fusing the temporal features and the time-frequency features to obtain fused features; Inputting the fused features into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies; Using a fully connected layer to classify and predict the output features of the bidirectional gated recurrent unit network to obtain the detection result of the eye movement behavior to be detected; The multiple parallel Mamba networks are composed of 3 parallel branches, each branch is composed of a linear layer and a state space model connected in series, and the number of channels of each branch is 32; The concatenating and fusing the temporal features and the time-frequency features to obtain fused features includes: Concatenating the outputs of the state space models of the 3 parallel branches based on the channel dimension, and using a linear layer to map the number of channels after concatenation to 32 to obtain the mapped temporal features; Keeping the size of the length dimension of the time-frequency features unchanged, fusing the width dimension into the channel dimension, and using a linear layer to map the fused channel dimension to 64 to obtain the mapped time-frequency features; Concatenating the mapped temporal features and the mapped time-frequency features based on the channel dimension to obtain fused features.

2. The high-precision eye movement behavior detection method driven by time sequence and time-frequency fusion according to claim 1, characterized in that The performing short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected collected, and converting the original eye movement data sequence in the time domain into time-frequency domain information includes: Performing short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected collected, expanding the time information of the original eye movement data sequence to the frequency dimension, capturing the frequency change features at different time points, so as to convert the original eye movement data sequence in the time domain into time-frequency domain information; During the process of performing short-time Fourier transform, by setting the length and overlap degree of the time window, ensuring the continuity and integrity of the time-frequency domain information.

3. A high-precision eye movement behavior detection method driven by time series and time-frequency fusion according to claim 1, characterized in that For the original eye movement data sequence with the input after being mapped by the linear layer, the output of the linear layer is expressed by the formula: ; Among them, is the weight matrix of the linear transformation, is the bias term, is the -th item in the original eye movement data sequence, is the -th item of the output of the linear layer; The output of the linear layer is further processed by the state space model, and the output of the state space model is expressed by the formula: ; ; Among them, is the state transition matrix, is the input mapping matrix, and are the observation matrices, represents the hidden state at the current time step of, represents the hidden state at the previous time step of, is the -th item of the output of the state space model.

4. A high-precision eye movement behavior detection method driven by time series and time-frequency fusion according to claim 1, characterized in that The convolutional kernel size used in the convolutional neural network is , the stride is , the number of layers is 3, and the number of channels in each layer is 32.

5. A high-precision eye movement behavior detection method driven by time series and time-frequency fusion according to any one of claims 1 to 4, characterized in that The bidirectional gated recurrent unit network is composed of 4 layers of bidirectional gated recurrent units, the dimension of its hidden layer is 64, and the dimension of the output is also 64.

6. A high-precision eye movement behavior detection system driven by the fusion of time series and time frequency, characterized in that, Including: An STFT preprocessing module for performing short-time Fourier transform on the original eye movement data sequence of the eye movement behavior to be detected collected, and converting the original eye movement data sequence in the time domain into time-frequency domain information; A temporal extraction module for using multiple parallel Mamba networks to extract features from the temporal information in the original eye movement data sequence to obtain temporal features; A time-frequency extraction module for using a convolutional neural network to extract features from the time-frequency domain information to obtain time-frequency features; A concatenating and fusing module for concatenating and fusing the temporal features and the time-frequency features to obtain fused features; A sequence modeling module for inputting the fused features into a bidirectional gated recurrent unit network for sequence modeling to capture global dependencies; A classification prediction module, which is used to perform classification prediction on the output features of the bidirectional gated recurrent unit network using a fully connected layer, and obtain the detection result of the eye movement behavior to be detected; Multiple parallel Mamba networks are composed of 3 parallel branches. Each branch is composed of a linear layer and a state space model connected in series, and the number of channels of each branch is 32; The splicing and fusion of the temporal features and the time-frequency features to obtain the fusion features includes: Splicing the outputs of the state space models of the 3 parallel branches based on the channel dimension, and using a linear layer to map the number of channels after splicing to 32 to obtain the mapped temporal features; Keeping the size of the length dimension of the time-frequency features unchanged, fusing the width dimension into the channel dimension, and using a linear layer to map the fused channel dimension to 64 to obtain the mapped time-frequency features; Splicing the mapped temporal features and the mapped time-frequency features based on the channel dimension to obtain the fusion features.

7. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a high-precision eye movement behavior detection method according to any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements a high-precision eye movement behavior detection method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Signal type identification method and system based on fusion feature and group convolution ViT network

    CN117743946A