Neuromorphic eye movement tracking method based on global-local space-time LSTM modeling

Through the global-local spatiotemporal LSTM modeling method, combined with GFMM and MSM, multi-head attention and convolution operation are used to solve the problems of subtle structural changes and sparseness processing in neuromorphic eye movement tracking, and efficient eye movement tracking effect is achieved.

CN120564249APending Publication Date: 2025-08-29DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510548460.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing neuromorphic eye movement tracking methods lack the effective extraction of time cues and ignore the properties of neuromorphic data and eye structure, it is difficult to identify subtle structural changes and capture fine details, adapt to dynamic angle changes, and have poor sparsity in processing.

Method used

The global-local spatiotemporal LSTM modeling method is adopted, combined with the global feature extraction model GFMM and the multi-scale feature extraction model MSM, and the multi-head attention mechanism and convolution operation are used to construct a neuromorphic eye tracking model, extract spatiotemporal features and adapt to the scale and angle changes of the eye area.

Benefits of technology

It realizes effective capture of local gaze mode and global context dependence, improves the robustness of eye tracking, can locally identify subtle structural changes and capture fine details, capture the overall layout of the eye area globally, and extract space-time information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564249A_ABST
    Figure CN120564249A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of computer vision, and relates to a neural morphology eye movement tracking method based on global-local space-time LSTM modeling, which comprises the following steps: constructing a neural morphology eye movement tracking model by using a global feature extraction model GFMM and a multi-scale feature extraction model MSM in combination with a long short-term memory network LSTM; converting the neuromorphic visual data into a frame image, and inputting the frame image into a neuromorphic eye movement tracking model to extract spatio-temporal features; the global feature extraction model captures global features by using a multi-head attention mechanism; the multi-scale feature extraction model identifies local features by using convolution operation; and inputting features output by the global feature extraction model and the multi-scale feature extraction model into a regression layer to realize eye movement tracking. According to the method, multi-scale recognition fine structure changes are provided locally, fine details are captured, adaptability is achieved in the aspects of eye area scale and angle changes, the overall layout of the eye area is captured globally, and features of spatio-temporal information can be further extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling. Background Art

[0002] Eye tracking has enormous application potential in fields such as human-computer interaction, mental health assessment, education, and safe driving, exemplified by virtual reality and augmented reality. Furthermore, because rapid eye movement results in high angular velocity and acceleration, eye tracking systems must provide real-time feedback with low latency and high frequency. Therefore, compared to traditional RGB-based cameras, neuromorphic vision cameras can provide rich spatial and temporal information while offering advantages such as high temporal resolution, high dynamic range, low power consumption, and high pixel bandwidth. These advantages are particularly well-suited for eye tracking tasks, enabling neuromorphic vision cameras to overcome complex eye movement scenarios and providing new possibilities for the further development of eye tracking.

[0003] In order to effectively extract and utilize the spatiotemporal information of neuromorphic visual data, the typical approach is to first convert the neuromorphic data into frame form and then use Transformers or Convolutional Neural Networks (CNNs) network structures to capture features. However, these methods lack the intrinsic ability to extract temporal cues, which limits the performance of tracking models. In addition, some research works have attempted to introduce Convolutional Long Short-Term Memory Networks (ConvLSTMs), Gate Recurrent Units (GRU) and Spiking Neural Networks (SNNs). These methods explore modeling sequence data and establishing long-term dependencies. However, using only these mechanisms ignores the properties of neuromorphic data and eye structure, limiting the representation of the model.

[0004] In neuromorphic eye tracking tasks, given the particularity of the human eye's structure and movement, it is very important to study how to identify subtle structural changes and capture fine details, adapt to dynamic changes in angles, and effectively handle the sparsity of neuromorphic data and capture continuous dependencies. Summary of the Invention

[0005] According to the technical problems raised above, a neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling is provided. The present invention mainly utilizes a multi-scale module (MSM) and a global features modeling module (GFMM) to perform feature extraction in local and global aspects, and then combines a long short-term memory network (LSTM) to extract spatiotemporal information. The present invention provides multi-scale recognition of subtle structural changes and captures fine details locally, is adaptable in terms of scale and angle changes in the eye area, captures the overall layout of the eye area globally, and can further extract features of spatiotemporal information.

[0006] The technical means adopted in the present invention are as follows:

[0007] A neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling, comprising:

[0008] A neuromorphic eye tracking model is constructed using the global feature extraction model GFMM and the multi-scale feature extraction model MSM, combined with the long short-term memory network LSTM;

[0009] Convert neuromorphic visual data into frame images and input them into the neuromorphic eye tracking model to extract spatiotemporal features;

[0010] The global feature extraction model uses a multi-head attention mechanism to capture global features;

[0011] The multi-scale feature extraction model uses convolution operations to identify local features;

[0012] The features output by the global feature extraction model and the multi-scale feature extraction model are input into the regression layer to achieve eye tracking.

[0013] Furthermore, the neuromorphic visual data includes N events, and the set ε of neuromorphic visual data is expressed as:

[0014]

[0015] Among them, (x k ,y k ) represents the kth event e k The pixel position, t k Indicates the timestamp, p k Represents polarity, converting neuromorphic visual data within ΔT into frame images, where ΔT represents the length of the time window, which is used to divide the continuous event stream into multiple segments;

[0016] The context in eye tracking tasks is continuous, and the neuromorphic eye tracking model is used to extract spatiotemporal features:

[0017] Z′ t =[ψ(X t ),H t-1 ]

[0018] z t =α(MSM(Z′ t ))+β(GFMM(Z′ t ))

[0019] Z i,t ,Z f,t ,Z g,t ,Z o,t =Split(ψ(Z t ))

[0020] i t =σ(Z i,t ),f t =σ(Z f,t ),g t =tanh(Z g,t ),o t =σ(Z o,t )

[0021] C t =f t ⊙C t-1 +i t ⊙g t

[0022] H t =o t ⊙tanh(C t )

[0023] Among them, Z′ t represents the input of the neuromorphic eye tracking model, Z t represents the output of the neuromorphic eye tracking model, Z i,t ,Z f,t ,Z g,t ,Z o,t It's Z t The output after the convolution operation is equally divided into four parts in the channel dimension, corresponding to the pre-activation values ​​of the input gate, forget gate, cell gate, and output gate at time step t in LSTM; ψ represents the convolution layer; [·] is the concatenation operation; Split is the segmentation operation; ⊙ is the Hadamard product; σ is the Sigmoid function; tanh is the hyperbolic tangent function; i t ,f t ,g t ,o tThey represent the input gate, forget gate, cell gate, and output gate of time step t in LSTM respectively; C t 、H t , and X t Represent the memory unit state, hidden state output and input event state respectively; α and β are parameters.

[0024] Furthermore, the global feature extraction model uses a multi-head attention mechanism to learn the global feature dependencies of neuromorphic data and assigns the attention value of the nth head at position (i, j) to Expressed as:

[0025]

[0026] in, represents the Softmax function, d is the balance parameter; q n 、k n 、v n They are the query vector, key vector, and numerical vector in attention calculation respectively; represents the pixel area of ​​spatial range k centered at (i, j), is the output of the GFMM at position (i,j).

[0027] Furthermore, the multi-scale feature extraction model uses convolution operation to adapt to the multi-scale information of eye scale changes; m branches with different convolution kernel sizes are combined to simulate the traditional convolution layer with a kernel size of k×k as k 2 different 1×1 convolutional layers to generate k 2 Features:

[0028]

[0029] in, z m,n The nth head representing the intermediate features z1, z2, and z3;

[0030] Apply depthwise separable convolution with fixed kernel size on k 2 Among the features, perform shift and sum operations:

[0031]

[0032] Aggregate the output features along the branch dimension into the output F of the multi-scale feature extraction model M .

[0033] Furthermore, the neuromorphic eye tracking model utilizes a global feature extraction model to extract global features, utilizes a multi-scale feature extraction model to extract local features, and combines a long short-term memory network to extract spatiotemporal information to achieve eye tracking.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] The present invention provides a neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling. This method utilizes a global feature extraction model (GFMM) and a multi-scale feature extraction model (MSM), combined with a long short-term memory (LSTM) network, to construct a neuromorphic eye tracking model. Neuromorphic visual data is converted into frame images and input into the neuromorphic eye tracking model to extract spatiotemporal features. The global feature extraction model utilizes a multi-head attention mechanism to capture global features, while the multi-scale feature extraction model uses convolution operations to identify local features. The features output by the global and multi-scale feature extraction models are input into a regression layer to implement eye tracking. This method effectively captures local gaze patterns and global contextual dependencies, achieving robust eye tracking. Developed based on the structural properties of the human eye and the properties of neuromorphic data, this method provides multi-scale recognition of subtle structural changes and captures fine details locally. It is adaptable to changes in the scale and angle of the eye region, globally captures the overall layout of the eye region, and further extracts features of spatiotemporal information.

[0036] Based on the above reasons, the present invention can be widely promoted in fields such as computer vision. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0038] Figure 1 This is a flow chart of the neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling of the present invention.

[0039] Figure 2 This is the overall framework diagram of the method of the present invention.

[0040] Figure 3 This is a structural diagram of the neuromorphic eye tracking model of the present invention.

[0041] Figure 4 This is a visualization diagram of the eye tracking results in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0043] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The following description of at least one exemplary embodiment is actually only illustrative and is in no way intended to limit the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0044] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0045] Unless otherwise specifically stated, the relative arrangement of the parts and steps, numerical expressions and numerical values ​​described in these embodiments do not limit the scope of the present invention. At the same time, it should be clear that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship. The techniques, methods and equipment known to ordinary technicians in the relevant fields may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific value should be interpreted as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, so once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0046] In the description of the present invention, it should be understood that the directions or positional relationships indicated by directional words such as "front, back, up, down, left, right", "horizontal, vertical, vertical, horizontal" and "top, bottom" are usually based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description. Unless otherwise specified, these directional words do not indicate or imply that the device or element referred to must have a specific direction or be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the scope of protection of the present invention: the directional words "inside and outside" refer to the inside and outside relative to the outline of each component itself.

[0047] For ease of description, spatially relative terms such as "above", "above", "on the upper surface of", "above", etc. may be used herein to describe the spatial positional relationship of a device or feature to other devices or features as shown in the figures. It should be understood that spatially relative terms are intended to include different orientations of the device in use or operation in addition to the orientation described in the figures. For example, if the device in the drawings is inverted, the device described as "above other devices or structures" or "above other devices or structures" will be positioned as "below other devices or structures" or "below their position devices or structures". Thus, the exemplary term "above" can include both "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatially relative descriptions used here are interpreted accordingly.

[0048] In addition, it should be noted that the use of terms such as "first" and "second" to limit components is only for the convenience of distinguishing the corresponding components. Unless otherwise stated, the above terms have no special meaning and therefore cannot be understood as limiting the scope of protection of the present invention.

[0049] like Figure 1 As shown, the present invention provides a neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling, comprising:

[0050] The global feature extraction model GFMM and the multi-scale feature extraction model MSM are used in combination with the long short-term memory network LSTM to build a neuromorphic eye tracking model, such as Figure 3 shown.

[0051] Convert neuromorphic visual data into frame images and input them into the neuromorphic eye tracking model to extract spatiotemporal features;

[0052] In a specific implementation, as a preferred embodiment of the present invention, the neuromorphic visual data includes N events, and the set ε of the neuromorphic visual data is represented as:

[0053]

[0054] Among them, (x k ,y k ) represents the kth event e k The pixel position, t k Indicates the timestamp, p k Represents polarity, converting neuromorphic visual data within ΔT into frame images, where ΔT represents the length of the time window, which is used to divide the continuous event stream into multiple segments;

[0055] The context in eye tracking tasks is continuous. This paper combines two modules that can respectively extract local and global spatial features with LSTM, allowing the network to effectively selectively memorize and forget information in time and space, effectively processing long-distance dependencies, and using the neuromorphic eye tracking model to extract spatiotemporal features:

[0056] Z′ t =[ψ(X t ),H t-1 ]

[0057] Z t =α(MSM(Z′ t ))+β(GFMM(Z′ t ))

[0058] Z i,t ,Z f,t ,Z g,t ,Z o,t =Split(ψ(Z t ))

[0059] i t =σ(Z i,t ),f t =σ(Z f,t ),g t =tanh(Z g,t ),o t =σ(Z o,t )

[0060] C t =f t ⊙C t-1 +i t ⊙g t

[0061] H t =o t ⊙tanh(C t )

[0062] Among them, Z′ t represents the input of the neuromorphic eye tracking model, Z t represents the output of the neuromorphic eye tracking model, Z i,t ,Z f,t ,Z g,t ,Z o,t It's Z t The output after the convolution operation is equally divided into four parts in the channel dimension, corresponding to the pre-activation values ​​of the input gate, forget gate, cell gate, and output gate at time step t in LSTM; ψ represents the convolution layer; [·] is the concatenation operation; Split is the segmentation operation; ⊙ is the Hadamard product; σ is the Sigmoid function; tanh is the hyperbolic tangent function; it ,ft t ,g t ,o t They represent the input gate, forget gate, cell gate, and output gate of time step t in LSTM respectively; C t 、H t , and X t Represent the memory unit state, hidden state output and input event state respectively; α and β are parameters.

[0063] The traditional convolutional layer with kernel size k×k can be divided into k 2 Different 1×1 convolution layers are then used for shift and sum operations. Similarly, the query vector, key vector, and value vector required for attention calculation can all be obtained from 1×1 convolution, and then the attention weight is calculated. Therefore, MSM and GFMM share the intermediate features obtained from three 1×1 convolution operations, namely z1, z2, and z3, and both use a multi-head mechanism.

[0064] The global feature extraction model uses a multi-head attention mechanism to capture global features.

[0065] In specific implementation, as a preferred embodiment of the present invention, the global feature extraction model uses a multi-head attention mechanism to learn the global feature dependency of neuromorphic data, and the attention value of the nth head at position (i, j) is Expressed as:

[0066]

[0067] in, represents the Softmax function, d is the balance parameter; q n 、k n 、v n are the query vector, key vector, and numerical vector in the attention calculation, i.e., the nth head of the three intermediate features z1, z2, and z3. Considering that most of the valuable eye data is concentrated in the center of the image, the attention value of the nth head at position (i, j) is is calculated within a restricted neighborhood. represents the pixel area of ​​spatial range k centered at (i, j), is the output of the GFMM at position (i,j).

[0068] The multi-scale feature extraction model uses convolution operations to identify local features. Because the human eye contains a wealth of local features, the MSM module designed in this paper uses convolution operations. This facilitates the multi-scale information extraction module to adapt to changes in eye scale, identifying subtle structural variations and capturing fine details. This module can also adapt to dynamic changes in the scale of the eye region caused by changes in viewing angle and other variables.

[0069] In specific implementation, as a preferred embodiment of the present invention, the multi-scale feature extraction model uses convolution operation to adapt to the multi-scale information of eye scale changes; m branches with different convolution kernel sizes are combined to simulate the traditional convolution layer with a kernel size of k×k as k 2 different 1×1 convolutional layers to generate k 2 Features:

[0070]

[0071] in, z m,n The nth head representing the intermediate features z1, z2, and z3;

[0072] Apply depthwise separable convolution with fixed kernel size on k 2 Among the features, perform shift and sum operations:

[0073]

[0074] Aggregate the output features along the branch dimension into the output F of the multi-scale feature extraction model M .

[0075] The features output by the global feature extraction model and the multi-scale feature extraction model are input into the regression layer to achieve eye tracking.

[0076] In specific implementation, as a preferred embodiment of the present invention, the neuromorphic eye tracking model uses a global feature extraction model to extract global features, uses a multi-scale feature extraction model to extract local features, and combines a long short-term memory network to extract spatiotemporal information to achieve eye tracking.

[0077] Example

[0078] like Figure 2 As shown in Figure 2, the neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling in this invention uses the PyTorch deep learning framework for network construction, training, and testing. The network is optimized using the Adam algorithm. The training batch size and total number of epochs are set to 8 and 50, respectively. The initial learning rate is set to 1×10 -3After 25 epochs, the learning rate is decreased by a factor of 10. The entire network is trained and tested on one NVIDIA RTX3090 GPU.

[0079] The datasets used in this example include:

[0080] (1) The SEET dataset is a synthetic dataset that converts the LPW dataset (RGB dataset) into neuromorphic data through the DVS simulator v2e, providing neuromorphic visual data and corresponding RGB images.

[0081] (2) The Ini-30 dataset is an eye tracking dataset based on real neuromorphic data collected from 30 volunteers, with variability in recording length and neuromorphic data count. The neuromorphic data count at each time step spans 3,000-5,000 events, and the input dimension is adjusted to 64×64×2.

[0082] (3) The EV-Eye dataset is a multimodal frame-based neuromorphic dataset for high-frequency eye tracking. Here, we use only RGB images as input to evaluate the effectiveness of our invention on a frame-based dataset. 6,344 images were randomly selected for training and 2,677 images for testing.

[0083] The comparison of eye tracking visualization results of our method (Ours) with ConvLSTM and ViT methods on SEET and Ini-30 datasets is shown in the following example: Figure 4 As shown in Figure 1. GT represents the actual eye movement, which is represented by Figure 4 It can be seen that the eye tracking result obtained by the method of the present invention is better than the ConvLSTM method and the ViT method, and can achieve eye tracking more accurately.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling, characterized by: include: A neuromorphic eye tracking model is constructed using the global feature extraction model GFMM and the multi-scale feature extraction model MSM, combined with the long short-term memory network LSTM; Convert neuromorphic visual data into frame images and input them into the neuromorphic eye tracking model to extract spatiotemporal features; The global feature extraction model uses a multi-head attention mechanism to capture global features; The multi-scale feature extraction model uses convolution operations to identify local features; The features output by the global feature extraction model and the multi-scale feature extraction model are input into the regression layer to achieve eye tracking.

2. The neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling according to claim 1, characterized in that The neuromorphic visual data contains N events, and the set ε of neuromorphic visual data is expressed as: Among them, (x k ,y k ) represents the kth event e k The pixel position, t k Indicates the timestamp, p k Represents polarity, converting neuromorphic visual data within ΔT into frame images, where ΔT represents the length of the time window, which is used to divide the continuous event stream into multiple segments; The context in eye tracking tasks is continuous, and the neuromorphic eye tracking model is used to extract spatiotemporal features: Z′ t =[ψ(X t ),H t-1 ] WITH t =α(MSM(Z′ t ))+β(GFMM(Z′ t )) WITH i,t ,WITH f,t ,WITH g,t ,WITH o,t =Split(ψ(Z t )) i t =σ(Z i,t ),f t =σ(Z f,t ),g t =tanh(Z g,t ),o t =σ(Z o,t ) C t =f t ⊙C t-1 +i t ⊙g t H t =o t ⊙tanh(C t ) Among them, Z′ t represents the input of the neuromorphic eye tracking model, Z t represents the output of the neuromorphic eye tracking model, Z i,t ,Z f,t ,Z g,t ,Z o,t It's Z t The output after the convolution operation is equally divided into four parts in the channel dimension, corresponding to the pre-activation values ​​of the input gate, forget gate, cell gate, and output gate at time step t in LSTM; ψ represents the convolution layer; [·] is the concatenation operation; Split is the segmentation operation; ⊙ is the Hadamard product; σ is the Sigmoid function; tanh is the hyperbolic tangent function; i t ,f t ,g t ,o t They represent the input gate, forget gate, cell gate, and output gate of time step t in LSTM respectively; C t 、H t , and X t Represent the memory unit state, hidden state output and input event state respectively; α and β are parameters.

3. The neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling according to claim 1, characterized in that The global feature extraction model uses a multi-head attention mechanism to learn the global feature dependencies of neuromorphic data and assigns the attention value of the nth head at position (i, j) to Expressed as: in, represents the Softmax function, d is the balance parameter; q n 、k n 、v n They are the query vector, key vector, and numerical vector in attention calculation respectively; represents the pixel area of ​​spatial range k centered at (i, j), is the output of the GFMM at position (i,j).

4. The neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling according to claim 1, characterized in that The multi-scale feature extraction model uses convolution operation to adapt to the multi-scale information of eye scale changes; m branches with different convolution kernel sizes are combined to simulate the traditional convolution layer with kernel size k×k as k 2 Different 1×1 convolutional layers generate k 2 Features: in, z m,n The nth head representing the intermediate features z1, z2, and z3; Apply depthwise separable convolution with fixed kernel size on k 2 Among the features, perform shift and sum operations: Aggregate the output features along the branch dimension into the output F of the multi-scale feature extraction model M .

5. The neuromorphic eye tracking method based on global-local spatiotemporal LSTM modeling according to claim 1, characterized in that The neuromorphic eye tracking model uses a global feature extraction model to extract global features, uses a multi-scale feature extraction model to extract local features, and combines a long short-term memory network to extract spatiotemporal information to achieve eye tracking.