Low-energy-consumption long-sequence action recognition method and system based on mamba pulse neural network

Through the low-energy long-sequence action recognition method based on mamba pulse neural network, the RGB image algorithm is solved while achieving efficient and accurate action recognition of event data while achieving a variety of scenarios and motion states.

CN120496181APending Publication Date: 2025-08-15NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510584248.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing RGB image-based human action recognition algorithms have privacy concerns in privacy-sensitive environments, and the challenge of sparsity and high temporal resolution when modeling event data using event cameras, resulting in high energy consumption and difficulty in effectively handling scenes under complex lighting conditions.

Method used

The low-energy long-sequence action recognition method based on mamba pulse neural network is adopted. By obtaining the event pulse data set, the spatiotemporal dimension processing is performed, and combined with spatiotemporal context encoding and mamba pulse neural network model training, the precise identification of event pulse data is achieved.

Benefits of technology

While reducing energy consumption, it improves the accuracy of motion recognition in diverse scenarios and motion states, with higher computing efficiency and fewer model parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496181A_ABST
    Figure CN120496181A_ABST
Patent Text Reader

Abstract

The invention discloses a low-energy-consumption long-sequence motion recognition method and system based on a mamba pulse neural network, and relates to the field of motion recognition, and the method comprises the steps: obtaining a motion recognition data set of an original event pulse, carrying out the space-time dimension processing of the motion recognition data set of the original event pulse, and obtaining an embedded vector sequence, the dimension of the embedded vector sequence is fixed; performing spatio-temporal context coding on the embedded vector sequence to obtain an embedded vector sequence containing the position context information, inputting the embedded vector sequence containing the position context information into a pre-established mamba spiking neural network model for training, and outputting to obtain a trained spiking neural network model; and the action data of the to-be-recognized event pulse is acquired, the action data of the to-be-recognized event pulse is input into the trained pulse neural network model, and a final human body action recognition result is output, so that accurate recognition of human actions in diversified scenes and motion states can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of action recognition, and in particular to a low-energy consumption long-sequence action recognition method and system based on a mamba pulse neural network. Background Art

[0002] Spiking Neural Networks (SNNs) are the next generation of artificial neural networks (ANNs). Spiking neurons are a key component of SNNs. They process input currents with complex neural dynamics and emit pulses as output when the membrane potential reaches a threshold. SNNs use discrete pulses to communicate between layers, which enables an event-driven computing paradigm. Due to their high biological plausibility, SNNs are viewed by neuroscientists as effective tools for analyzing, simulating, and learning from biological systems. As a bridge between neuroscience and computational science, SNNs have attracted increasing research interest in recent years. The low energy consumption of SNNs is particularly important given the current scaling of AI models and the increasing importance of energy consumption.

[0003] Human action recognition (HAR) plays a key role in various applications such as video analytics, surveillance, autonomous driving, robotics, and healthcare. Most HAR algorithms are developed based on RGB images, which capture detailed visual information. However, due to the recording of identifiable features, these algorithms raise concerns in privacy-sensitive environments. Event cameras provide a promising solution by sparsely capturing scene brightness changes at the pixel level without capturing the full image. In addition, event cameras have a high dynamic range and can effectively handle scenes in complex lighting conditions, such as low-light or high-contrast environments. However, using event cameras introduces challenges in modeling event data with spatial sparseness and high temporal resolution. Summary of the Invention

[0004] In order to solve the deficiencies mentioned in the above background technology, the purpose of the present invention is to provide a low-energy long-sequence action recognition method and system based on mamba pulse neural network.

[0005] In a first aspect, the purpose of the present invention can be achieved by the following technical solution: a low-energy long-sequence action recognition method based on a mamba pulse neural network, the method comprising the following steps:

[0006] Acquire an action recognition dataset of original event pulses, perform spatiotemporal dimension processing on the action recognition dataset of original event pulses, and obtain an embedding vector sequence, wherein the dimension of the embedding vector sequence is fixed;

[0007] Performing spatiotemporal context encoding on the embedded vector sequence to obtain an embedded vector sequence containing position context information, inputting the embedded vector sequence containing position context information into a pre-established mamba spiking neural network model for training; and outputting a trained spiking neural network model;

[0008] The motion data of the event pulse to be identified is obtained, the motion data of the event pulse to be identified is input into the trained spiking neural network model, and the final human motion recognition result is output.

[0009] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: performing spatiotemporal processing on the original event pulse action recognition dataset by a Spiking 3D Patch Embedding module, the process including the following steps:

[0010] Input data and spatial segmentation:

[0011] Spatial segmentation divides the time frame sequence in the input data into blocks of specific size to form a Patch format. The input data is a five-dimensional tensor: Where B represents the batch size, C represents the input channel, T represents the number of time frames, H and W represent the height and width of the image frame, and each frame is divided into non-overlapping blocks. When the input frame size is (H, W) and the patch size is patch_size×patch_size, the number of patches N after segmentation is calculated as:

[0012]

[0013] After spatial segmentation, the original data is reorganized into spatiotemporal units Tubelets to form a B×C×T×N structured input;

[0014] 3D convolution spatiotemporal feature extraction Conv3D:

[0015] The 3D convolution operation is used to extract temporal and spatial features of the input data and capture the local spatiotemporal dynamic relationship. The input data format is: The convolution operation uses the convolution kernel K to act on the time dimension T and the space dimension (H, W). The convolution formula is:

[0016] X′=Conv3D(X,K)

[0017] The convolution formula is for each Tubelet x in X ijk Perform convolution operation, the eigenvalue of each output position (i, j, k) is:

[0018]

[0019] Among them, Tk ,H k ,W k is the spatiotemporal size of the convolution kernel, K is the learnable weight, and the output dimension feature map is converted to:

[0020]

[0021] E is the embedding dimension, T′, H′, W′ are reduced due to convolution stride or padding, and the dimension is halved when the stride is 2;

[0022] Batch normalization stabilizes feature distribution:

[0023] Normalized along the batch and time dimensions, the normalization formula is:

[0024]

[0025] where μ e ,σ e is the mean and standard deviation of channel e in the current batch, γ e ,β e is a learnable scaling offset parameter;

[0026] Spiking

[0027] Incorporate the spiking neural network (SNN) and apply the spiking neuron model (LIF) after convolutional projection.

[0028] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: an operation process of the spiking neuron model LIF includes:

[0029] Membrane potential initialization, input excitation accumulation, membrane potential leakage, and threshold V th The pulse triggers and membrane potential resets.

[0030] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the process of performing spatiotemporal context encoding on the embedded vector sequence is performed by adding the spatial position encoding Spatial Position Embedding and the temporal position encoding Temporal Position Embedding to the Patch embedding element by element,

[0031] Among them, the spatial position encoding adds a code to each Patch embedding vector, which represents the position of the Patch in the original spatial field of view, and is combined with the element-by-element addition method of Patch embedding. The spatial embedding converts the two-dimensional coordinate information into a feature vector through encoding; the time position encoding adds a code to each Patch embedding vector, which represents the position of the Patch in the time series. The goal of time position embedding is to encode the relationship between data points changing over time.

[0032] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: the spatial position encoding and the temporal position encoding adopt a rotation encoding method, and the method is as follows:

[0033] Spatial position embedding realizes the encoding of spatial position information in features by generating a rotation vector for each spatial point. For spatial embedding, the processing process is as follows: the rotation angle θ in each embedding dimension k Calculated using the following formula:

[0034]

[0035] Where base is the frequency base, k is the embedding dimension index, and k max To maximize the embedding index, we construct the position label (x, y) of each spatial point by gridding, and apply the rotation angle to generate the embedding feature. For each spatial direction, we use the formula:

[0036] x embed [x,y]=x·θ k ,y embed [x,y]=y·θ k

[0037] Embed the (x,y) coordinates in the grid into the rotation matrix to get the rotated position embedding:

[0038]

[0039] The goal of time position embedding is to encode the relationship between data points over time, helping the model understand the dynamics of input data in the time dimension. By rotating the embedding mechanism, a unique feature expression is generated for each time point. For each time point t, an embedding value is generated: embed [t]=t·θ k , the embedding vectors at different time points contain information in different frequency ranges.

[0040] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: a process of converting a raw event stream in the action recognition dataset of the raw event pulse into an embedded representation with position awareness:

[0041] The original data X undergoes preliminary feature extraction and patch division through 3D convolution. After batch normalization, the pulse layer converts the continuous value features into pulse signals that conform to the characteristics of SNN, and adds position embedding information to give the feature vector the ability to perceive spatiotemporal context, which can be expressed as follows:

[0042] P=SL patch (BN(Conv3d(X)))+PE

[0043] Among them, Conv3d(·) is a 3D convolution layer that processes the original input data X and divides the continuous data stream into discrete 3D spatiotemporal patches. BN(·) is a batch normalization layer that normalizes the feature map after convolution. patch (·) is the spiking layer, which converts continuous values into discrete pulse signals and introduces the discharge mechanism of biological neurons. The spiking layer simulates the behavior of biological neurons and only generates output when the input exceeds the preset threshold. PE is position embedding information, which includes spatial position encoding and temporal position encoding. Through element-by-element addition operations and combined with the generated feature vector, the pre-established Mamba spiking neural network model obtains the relative position and order relationship of each patch in the original spatiotemporal data.

[0044] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: the process of inputting the embedding vector sequence containing the location context information into the pre-established Mamba spiking neural network model for training includes:

[0045] The pre-built mamba spiking neural network model consists of L series-connected SpikMamba Blocks and a final Prediction Layer;

[0046] The internal processing flow of SpikMamba Block includes:

[0047] The SpikMamba Block receives the embedded image block as input data and first enters the SpikeLinear Attention module. The input data is output through two preset paths to obtain the output result. The output result is added element by element to the original input of the SpikMamba Block to obtain the result of the residual connection.

[0048] The result of the residual connection is input into the Spike Mamba module and output through two preset paths. The Hadamard product is performed on the output of path one and path two to obtain the multiplied output result.

[0049] The multiplied output result is input into the feedforward network FFN to obtain the FFN output result, and the FFN output result and the residual connection result are added element by element to obtain the second residual connection result as the final output result.

[0050] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: the attention mechanism of the Linear Attention unit of the Spike LinearAttention module is to optimize the self-attention mechanism using a kernel function and a feature map, and the formula is as follows:

[0051] LinearAttention(Q,K,V)=φ(Q)·(φ(K) T V)

[0052] Where φ(·) is a feature mapping function and · represents matrix multiplication;

[0053] In the Spike Linear Attention module, input data is processed along two paths. In path one, the input is transformed by the linear layer, processed by the convolution layer Conv, and the SiLU activation function σ is applied. It is then divided into two parts and fed into parallel linear layers. Each linear layer is followed by a spike neural network layer SNN, and the output is fed into the Linear Attention calculation unit. The output of LinearAttention is expressed as:

[0054]

[0055] Among them, W o , W2 is the learnable weight matrix, represents the Hadamard product, S(·) is the impulse activation function of SNN, σ(·) is the SiLU activation function, and X is the original input.

[0056] In combination with the first aspect, in some implementations of the first aspect, the method further includes: during use, the Spike Mamba module obtains a multiplied output result by utilizing feature fusion of two paths, a discrete state space model SSM and a pulse neural network SNN.

[0057] In conjunction with the first aspect, in certain implementations of the first aspect, the method further includes: a formula of the discrete state space model SSM structure is as follows:

[0058]

[0059] y t =Ch t +Dx t

[0060] h t represents the hidden state vector at time t, x t Represents the input event stream characteristics at time t, y t As the output prediction result at time t;

[0061] During the state transition, As a state transition matrix, it dynamically controls how historical information is retained and decayed; As the input projection matrix, it determines how the current input affects the state; C acts as the output projection matrix, mapping the processed hidden state to the prediction space; D acts as the skip connection matrix, realizing the path where the input information directly affects the output;

[0062] Parameter Matrix and It is input-dependent and is converted from a continuous representation using the zero-order hold (ZOH) method:

[0063]

[0064] The discrete state space model SSM dynamically adjusts the memory and forgetting mechanisms according to the input content.

[0065] In a second aspect, in order to achieve the above-mentioned purpose, the present invention discloses a low-energy long-sequence action recognition system based on a mamba spiking neural network, comprising:

[0066] a data processing module, configured to obtain a motion recognition dataset of original event pulses, perform spatiotemporal dimension processing on the motion recognition dataset of original event pulses, and obtain an embedded vector sequence, wherein the dimension of the embedded vector sequence is fixed;

[0067] The model training module is used to perform spatiotemporal context encoding on the embedded vector sequence to obtain an embedded vector sequence containing position context information, input the embedded vector sequence containing position context information into a pre-established mamba spiking neural network model for training, and output the trained spiking neural network model;

[0068] The action recognition module is used to obtain the action data of the event pulse to be recognized, input the action data of the event pulse to be recognized into the trained pulse neural network model, and output the final human action recognition result.

[0069] Beneficial effects of the present invention:

[0070] The superior performance of the present invention compared with the most advanced existing algorithms reduces energy consumption and improves efficiency while ensuring accurate recognition of human actions in a variety of scenarios and motion states. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0072] Figure 1 It is a schematic flow chart of the method of the present invention;

[0073] Figure 2 is an example schematic diagram of event data of the present invention;

[0074] Figure 3 It is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0075] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0076] Example 1:

[0077] like Figure 1 As shown, a low-energy long-sequence action recognition method based on a mamba pulse neural network includes the following steps:

[0078] S101: Acquire an action recognition dataset of original event pulses, perform spatiotemporal dimension processing on the action recognition dataset of original event pulses, and obtain an embedded vector sequence, wherein the dimension of the embedded vector sequence is fixed;

[0079] Data Acquisition of Raw Event Pulse Action Recognition Datasets (Data Acquisition) Acquire event / pulse-based action recognition datasets required for subsequent network training. These data come from event cameras or other sensors that generate pulse signals, recording the changes in light intensity in the scene over time, such as Figure 2 shown.

[0080] The public event-based action recognition benchmark datasets HARDVS and E-FAction are used. The HARDVS dataset is designed to evaluate the model's ability to process everyday human activities captured by event cameras. Instead of providing RGB images, the dataset provides asynchronous event stream data for each action sequence. Each event record contains the exact timestamp (t) of its occurrence, pixel coordinates (x, y), and the polarity of the light intensity change (p). The HARDVS dataset is large in scale, containing more than 100,000 event sequences covering 300 different categories of everyday activities. Crucially, the design of the dataset fully reflects the challenging factors in the real world, such as multi-view shooting, significant lighting changes, different motion speeds, and dynamically changing backgrounds, which place high demands on the robustness and generalization ability of the model. The dataset provides standard training and test set partitioning to ensure the comparability of the results of different research works.

[0081] The E-FAction dataset is considered to be the first fine-grained human action dataset based on event cameras. The dataset provides 3,304 pairs of synchronously recorded "event streams and RGB sequences", making it also useful for multimodal research. This experiment mainly focuses on its event stream data part. The action category design of E-FAction is hierarchical, covering a total of 15 coarse-grained action categories, which are further subdivided into 128 fine-grained action categories. This requires the model to not only recognize large action categories, but also distinguish subtle action differences under the same large category, which places higher demands on the model's fine discrimination ability. Similar to HARDVS, the core data of E-FAction is the event stream (t, x, y, p), which captures the fine dynamic changes during the execution of the action. This dataset also provides clearly defined training and test sets for model training and evaluation.

[0082] The process of processing the spatiotemporal dimensions of the original event pulse action recognition dataset:

[0083] Input data and spatial segmentation.

[0084] Spatial segmentation divides the time frame sequence in the input data into blocks of a specific size to form a patch format. The input data is a five-dimensional tensor: Where B represents the batch size. C represents the input channels. T represents the number of time frames. H and W represent the height and width of the image frame. To capture local spatiotemporal features, each frame is first divided into non-overlapping blocks. If the input frame size is (H, W) and the patch size is patch_size × patch_size, the number of patches N after segmentation is calculated as:

[0085]

[0086] After spatial segmentation, each patch will further integrate the time dimension. This step reorganizes the original data into spatiotemporal units (Tubelets), forming a B×C×T×N structured input, laying the foundation for subsequent feature extraction.

[0087] 3D convolutional spatiotemporal feature extraction (Conv3D).

[0088] Through 3D convolution operation, the temporal and spatial features of the input data are extracted to capture the local spatiotemporal dynamic relationship. Input data format: The convolution operation uses the convolution kernel K to act on the time dimension T and the space dimension (H, W). The convolution formula is:

[0089] X'=Conv3D(X,K)

[0090] The convolution formula is for each Tubelet x in X ijk Perform convolution operation, the eigenvalue of each output position (i, j, k) is:

[0091]

[0092] Among them, T k ,H k ,W k is the spatiotemporal size of the convolution kernel, and K is the learnable weight. The output dimension feature map is converted to:

[0093]

[0094] E is the embedding dimension, and T', H', W' are reduced due to the convolution stride or padding. For example, when the stride is 2, the dimension is halved.

[0095] Batch normalization stabilizes feature distribution.

[0096] To alleviate the internal covariate shift during training, normalization is performed along the batch and time dimensions. The normalization formula is:

[0097]

[0098] where μ e ,σ e is the mean and standard deviation of channel e in the current batch, γ e ,β e is a learnable scaling offset parameter. This operation stabilizes the feature distribution and accelerates model convergence.

[0099] The conversion process of the raw event stream in the action recognition dataset of the raw event pulse to the embedding representation with position awareness:

[0100] The original data X undergoes preliminary feature extraction and patch division through 3D convolution, and the stability of the features is improved through batch normalization. Then, the pulse layer converts the continuous-value features into pulse signals that conform to the characteristics of SNN, and adds position embedding information to give the feature vector the ability to perceive spatiotemporal context, which can be expressed as follows:

[0101] P=SL patch (BN(Conv3d(X)))+PE

[0102] Among them, Conv3d(·) is a 3D convolutional layer with a stride of 1×8×8 and a kernel size of 1×8×8. It processes the original input data X and divides the continuous data stream into discrete 3D spatiotemporal patches. BN(·) is a batch normalization layer that standardizes the feature map after convolution to improve training stability and convergence speed and reduce internal covariate shift. patch (·) is the spiking layer, which converts continuous values into discrete pulse signals and introduces the discharge mechanism of biological neurons. The spiking layer simulates the behavior of biological neurons and only generates output when the input exceeds the preset threshold. PE (Position Embedding) is position embedding information, including spatial position encoding and temporal position encoding. By combining it with the generated feature vector through element-by-element addition operations, the pre-established Mamba spiking neural network model can obtain the relative position and order relationship of each patch in the original spatiotemporal data.

[0103] S102: performing spatiotemporal context encoding on the embedded vector sequence to obtain an embedded vector sequence containing position context information, inputting the embedded vector sequence containing the position context information into a pre-established mamba spiking neural network model for training, and outputting a trained spiking neural network model;

[0104] The process of spatiotemporal context encoding of the embedding vector sequence adds the spatial position encoding Spatial Position Embedding and the temporal position encoding Temporal Position Embedding to the patch embedding element by element, so that the model can understand the relative position and order of each patch in the original spatiotemporal data.

[0105] Spatial and temporal position embeddings are key components for fine-grained spatiotemporal feature modeling, helping the model effectively perceive the spatial and temporal characteristics of input data. The module implements this encoding through Rotary Positional Embedding (RoPE), which utilizes frequency representation and rotation matrices to generate efficient embeddings. The following details the principles and embedding formulas for these two components.

[0106] Spatial Position Embedding

[0107] Spatial position embedding encodes spatial position information in features by generating a rotation vector for each spatial point. Its goal is to provide the model with a structured representation of each point in two-dimensional space (x, y). For spatial embedding, the process is as follows: the rotation angle θ in each embedding dimension is k Calculated using the following formula:

[0108]

[0109] Where base is the frequency base (usually set to 10000) and k is the embedding dimension index. max is the maximum embedding index (defined as the embedding dimension d embed Divide by the number of spatial dimensions). Then, the position label (x, y) of each spatial point is constructed by gridding, and the rotation angle is applied to generate the embedded feature. For each spatial direction, the formula is used:

[0110] x embed [x,y]=x·θ k ,y embed [x,y]=y·θ k

[0111] Finally, the (x,y) coordinates in the grid are embedded into the rotation matrix to obtain the rotated position embedding:

[0112]

[0113] Spatial embedding converts two-dimensional coordinate information into feature vectors through this rotation encoding method, realizing the structured expression of spatial dimensions.

[0114] Temporal Position Embedding

[0115] The goal of temporal position embedding is to encode the relationship between data points over time, helping the model understand the dynamics of the input data in the time dimension. By rotating the embedding mechanism, a unique feature expression is generated for each time point. For each time point t, the embedding value is generated: embed [t]=t·θ k ,The embedding vectors at different time points contain information in different frequency ranges, and can capture fine-grained temporal feature changes.

[0116] In this module, spatial and temporal position embeddings are tightly integrated. By generating a joint rotational embedding matrix, the spatial and temporal information of the input data is simultaneously encoded. This joint rotational embedding allows the model to simultaneously encode dynamic features across time and space while preserving the data's structured information and long-range dependencies. This approach not only effectively captures the spatiotemporal correlations of the input data but also provides a more comprehensive feature representation, enhancing the model's ability to understand and recognize complex spatiotemporal patterns.

[0117] The process of inputting the embedding vector sequence containing the location context information into the pre-established mamba spiking neural network model for training includes:

[0118] Core Network Framework Training trains the SpikMamba network framework based on the resulting embedded patches containing positional context. This network framework primarily consists of L cascaded SpikMambaBlocks (core feature extractors) and a final Prediction Layer (for classification or regression). The training process optimizes network parameters (including those of the Spike Linear Attention, Spike Mamba, and FFN components within the SpikMamba Blocks) through backpropagation, with the goal of enabling the network to accurately learn and distinguish different motion patterns from input spike trains.

[0119] The SpikMamba Block receives embedded image patches as input. These inputs typically incorporate spatiotemporal position encodings. The input first enters the Spike Linear Attention module. Within this module, the input data is processed along two paths: in path one, the input first passes through a linear layer, then a convolutional layer (Conv), and then the SiLU activation function (σ) is applied. The result is then split into two parts and fed into two parallel linear layers, each followed by a spiking neural network (SNN) layer. The Spiking Neural Network (SNN) is one of the key steps in Spike Linear Attention. It achieves sparse activation by simulating the spiking mechanism of biological neurons, enhancing the efficiency of feature expression. It uses the Leak-Integrate-and-Fire (LIF) spiking neuron model. In Spike Linear Attention, SNN activation sparsifies the Q, K, and V features respectively. The specific operation process is as follows:

[0120] Membrane potential initialization: Each spiking neuron has a specific membrane potential V(t), and the initial state is usually set to:

[0121] V(t=0)=V rest

[0122] Where V rest is the resting potential of the neuron, which is usually initialized to zero (or a fixed constant) in the simulation, that is:

[0123] V(t=0)=0

[0124] Input excitation accumulation: When the neural network input reaches a neuron, it drives the membrane potential to change in the form of current or voltage. For a certain time step t, the effect of the input on the neuron membrane potential can be expressed as:

[0125] I(t)=w·x(t)

[0126] Where x(t) is the input pulse signal at the current time step, and w is the connection weight input to the current neuron. The membrane potential therefore accumulates this input excitation and is expressed as:

[0127] V(t+1)=V(t)+w·x(t)

[0128] Membrane potential leakage: An important feature of the LIF model is the leakage effect of the membrane potential, which is similar to the natural loss of current in biological membranes. This can be represented by a leakage coefficient τ, and the membrane potential decays based on the leakage coefficient at each time step:

[0129] V(t+1)=α·V(t)+w·x(t)

[0130] in, represents the leakage factor (usually 0 < α < 1), and τ is the leakage time constant. Due to the leakage effect, the membrane potential of the neuron does not grow indefinitely, but gradually decays (similar to the characteristics of biological nervous systems).

[0131] Threshold-based pulse triggering: When the membrane potential V(t) accumulates to a certain level and exceeds the preset threshold When , the neuron will trigger a pulse (i.e., transmit a signal). This process can be expressed by the following formula:

[0132]

[0133] Where S(t) indicates whether the neuron emits a pulse at the current time step (1 for emission, 0 for non-emission).

[0134] Membrane potential reset: After a trigger pulse, the membrane potential of the neuron will be reset (usually to zero or set to a fixed value) to simulate the absolute refractory period of biological neurons and prevent the firing of pulses too frequently. The membrane potential reset formula is:

[0135]

[0136] in, is the threshold of the trigger pulse, V reset is the reset value of the membrane potential (usually zero).

[0137] Putting the above processes together, we get the complete dynamic equation of the LIF neuron:

[0138]

[0139] The corresponding output pulse can be written as:

[0140]

[0141] Among them S t It is a pulse output, a binary signal (discrete value of pulse emission), V th is the membrane potential threshold, which sets the conditions for the neuron to trigger a pulse. After the pulse is triggered, the membrane potential will be reset: V mem (t) = 0. This discretization process achieves sparse activation, retaining only the signal at the salient feature location and ignoring other noise.

[0142] Finally, the pulse output encoding passes the pulse output as the data representation of the feature to the "Linear Attention" calculation unit. The input of path 2 also passes through a linear layer and then directly applies the SiLU activation function (σ). The output of the "LinearAttention" unit (which first passes through an SNN) is Hadamard Product (Hadamard Product) with the output of path 2 (the result after SiLU activation) ), which is element-wise multiplication. Finally, the result of this multiplication is output through a linear layer. The output of the Spike Linear Attention module will be element-wise added (Elementwise Sum⊕) to the original input entering the SpikMamba Block. This is a standard residual connection that retains the original information. The result of the residual connection in the previous step is then sent to the Spike Mamba module. Similarly, the input data will be processed along two paths: in path one, the input first passes through a linear layer, then an SNN layer, followed by a convolutional layer (Conv), a SiLU activation function (σ), and a core state space model (SSM) module. The core of Mamba lies in its innovative discrete state space model (SSM) representation. The present invention applies the Mamba state space model architecture to the field of event stream action recognition, significantly improving the model's ability to capture long sequences and spatiotemporal dependencies. It efficiently simulates continuous-time systems through discretized state space equations, and can capture long-distance dependencies in sequences with linear complexity. The Mamba architecture is described by the following formula:

[0143]

[0144] y t =Ch t +Dx t

[0145] In this architecture, first, h t Represents the hidden state vector at time t, which is responsible for encoding all historical event information; then, x t Represents the input event stream characteristics at time t, that is, the currently observed data; then, y t The output prediction result at time t is the final recognition output of the model.

[0146] During the state transition process, first As a state transfer matrix, it dynamically controls how historical information is retained and decayed; secondly, As the input projection matrix, it determines how the current input affects the state; again, C acts as the output projection matrix, mapping the processed hidden state to the prediction space; finally, D acts as the skip connection matrix, realizing the path by which the input information directly affects the output.

[0147] The key innovation lies in the parameter matrix and It is input-dependent and implements a "selectivity" mechanism that allows the model to dynamically adjust the memory mechanism based on the content of the event stream. These parameters are converted from the continuous representation using the zero-order hold (ZOH) method:

[0148]

[0149] This design enables the model to dynamically adjust its memory and forgetting mechanisms based on the input content. For important information, the model can set a larger B value to strengthen memory; for less important information, a smaller value can be set to achieve intelligent filtering. This input-dependent parameterization method greatly improves the model's ability to handle complex sequences.

[0150] The input of path 2 is then passed through a linear layer and the SiLU activation function (σ) is applied. The output of path 1 after SSM and SNN processing is Hadamard producted with the output of path 2 (the result after SiLU activation). The result of this multiplication is finally output through a linear layer. The output of the Spike Mamba module is then fed into a feed-forward network (FFN). The FFN consists of several linear layers and activation functions for further feature transformation and nonlinear mapping. The output of the FFN is element-wise added (⊕) to the result of the first residual connection (that is, the result of adding the output of the Spike Linear Attention module to its input), which is the second residual connection. The result after the second residual connection is the final output of this SpikMamba Block. This output can be fed into the next repeated SpikMamba Block (as shown by "×L" in the figure, indicating that this Block will stack L layers) and used for the final prediction task after the last layer.

[0151] S103: Acquire motion data of the event pulse to be identified, input the motion data of the event pulse to be identified into the trained spiking neural network model, and output the final human motion recognition result.

[0152] Action Recognition & Prediction uses the trained SpikMamba network framework model to perform forward propagation on new, to-be-recognized event / pulse data sequences. The network extracts spatiotemporal features and outputs the final human action recognition result, i.e., the action category label, through the Prediction Layer.

[0153] The raw event camera data stream is input and first segmented into spatiotemporal 3D patches by the Spiking 3D Patch Embedding module and converted into embedding vectors containing spike information. Next, spatial and temporal position encodings are added to these vectors to preserve their spatiotemporal context. The feature sequence fused with position information is then fed into a SpikMamba Block stacked L times for deep feature extraction and sequence modeling. Finally, the feature representation after L layers of processing is fed into the top Prediction layer to generate the final prediction output, resulting in human action recognition results.

[0154] Table 1

[0155] Model GLOPs #Params. ExACT 1.1 2.13M EVT 0.2 0.48M Ours 0.12 0.18M

[0156] The training results of the embodiment are shown in the following table, in which the framework of the present invention is compared with the most advanced ANN and SNN methods ExACT and EvT in terms of computational efficiency. The framework of the present invention surpasses the previous state-of-the-art methods on the HARDVS and E-FAction datasets, respectively, with significant improvements in recognition performance of 7.22% and 3.92%. At the same time, the framework of the present invention is more lightweight and has fewer learnable parameters. As shown in Table 1, compared with the most advanced ANN and SNN methods ExACT and EvT in terms of computational efficiency, the number of parameters of the framework of the present invention is 0.18M, which is 1.95M and 0.30M less than ExACT and EvT, respectively. At the same time, the number of floating-point operations of this framework is 0.12GLOPs, which is less than the common ExACT (1.1GLOPs) and EvT (0.2GLOPs). The framework of the present invention strikes a balance between computational efficiency and HAR performance. Our method has the fewest parameters and GLOPs, and also outperforms the most advanced ANN and SNN methods in HAR performance.

[0157] Table 2

[0158] Dataset Model SNN Acc(%) HARDVS X3D × 45.82 SlowFast × 46.54 ACTION-Net √ 46.85 R2Plus1D × 49.06 ResNet18 × 49.20 TAM × 50.41 C3D × 50.52 ESTF √ 51.22 Video-SwinTrans × 51.91 TSM × 52.63 ExACT × 90.10 Ours √ 97.32 E-FAction CLIP-L × 61.90 ResNet3D-N × 65.60 ResNet3D-K × 66.30 MASTAF × 67.10 ExACT × 67.93 Ours √ 71.02

[0159] Comparison of event-based action recognition with state-of-the-art models on the HARDVS and E-FAction datasets. Models are evaluated by accuracy (ACC) and also indicate the model type, either ANN or SNN. The method with the highest accuracy is shown in bold.

[0160] By combining the energy efficiency of spiking neural networks (SNNs) and the long sequence modeling capabilities of Mamba, global features can be efficiently captured from spatially sparse and high-temporal-resolution event data. In addition, in order to improve the locality of modeling, a linear attention mechanism based on a pulse window is used. The training results of the embodiment show that the framework of the present invention has achieved the best level in terms of computational efficiency (amount of operations) and model lightweight (number of parameters), and is a framework that can effectively protect user privacy and accurately identify human actions.

[0161] Embodiment 2: In the second aspect, in order to achieve the above-mentioned purpose, the present invention discloses a low-energy long-sequence action recognition system based on a mamba spiking neural network, comprising:

[0162] A data processing module 11 is used to obtain a motion recognition dataset of original event pulses, perform spatiotemporal dimension processing on the motion recognition dataset of original event pulses, and obtain an embedded vector sequence, wherein the dimension of the embedded vector sequence is fixed;

[0163] A model training module 12 is configured to perform spatiotemporal context encoding on the embedded vector sequence to obtain an embedded vector sequence containing position context information, input the embedded vector sequence containing position context information into a pre-established mamba spiking neural network model for training, and output a trained spiking neural network model;

[0164] The motion recognition module 13 is used to obtain the motion data of the event pulse to be recognized, input the motion data of the event pulse to be recognized into the trained pulse neural network model, and output the final human motion recognition result.

[0165] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is used to implement one or more instructions, specifically for loading and executing one or more instructions in a computer storage medium to implement the above method.

[0166] It should be further explained that, based on the same inventive concept, the present invention also provides a computer storage medium having a computer program stored thereon, which executes the above method when executed by a processor. The storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, device or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.

[0167] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present disclosure. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0168] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present disclosure. Various changes and improvements may be made to the present disclosure without departing from the spirit and scope of the present disclosure, and such changes and improvements shall fall within the scope of the present disclosure.

Claims

1. A low-energy long-sequence action recognition method based on mamba pulse neural network, characterized by: The method comprises the following steps: Acquire an action recognition dataset of original event pulses, perform spatiotemporal dimension processing on the action recognition dataset of original event pulses, and obtain an embedding vector sequence, wherein the dimension of the embedding vector sequence is fixed; Performing spatiotemporal context encoding on the embedded vector sequence to obtain an embedded vector sequence containing position context information, inputting the embedded vector sequence containing position context information into a pre-established mamba spiking neural network model for training; and outputting a trained spiking neural network model; The motion data of the event pulse to be identified is obtained, the motion data of the event pulse to be identified is input into the trained spiking neural network model, and the final human motion recognition result is output.

2. The low-energy long-sequence action recognition method based on mamba pulse neural network according to claim 1 is characterized in that: The process of processing the spatiotemporal dimensions of the original event pulse action recognition dataset is performed by the Spiking 3DPatch Embedding module, and the process includes the following steps: Input data and spatial segmentation: Spatial segmentation divides the time frame sequence in the input data into blocks of specific size to form a Patch format. The input data is a five-dimensional tensor: Where B represents the batch size, C represents the input channel, T represents the number of time frames, H and W represent the height and width of the image frame, and each frame is divided into non-overlapping blocks. When the input frame size is (H, W) and the patch size is patch_size×patch_size, the number of patches N after segmentation is calculated as: After spatial segmentation, the original data is reorganized into spatiotemporal units Tubelets to form a B×C×T×N structured input; 3D convolution spatiotemporal feature extraction Conv3D: The 3D convolution operation is used to extract temporal and spatial features of the input data and capture the local spatiotemporal dynamic relationship. The input data format is: The convolution operation uses the convolution kernel K to act on the time dimension T and the space dimension (H, W). The convolution formula is: X′=Conv3D(X,K) The convolution formula is for each Tubelet x in X ijk Perform convolution operation, the eigenvalue of each output position (i, j, k) is: Among them, T k ,H k ,W k is the spatiotemporal size of the convolution kernel, K is the learnable weight, and the output dimension feature map is converted to: E is the embedding dimension, T′, H′, W′ are reduced due to convolution stride or padding, and the dimension is halved when the stride is 2; Batch normalization stabilizes feature distribution: Normalized along the batch and time dimensions, the normalization formula is: where μ e ,σ e is the mean and standard deviation of channel e in the current batch, γ e ,β e is a learnable scaling offset parameter; Spiking Incorporate the spiking neural network (SNN) and apply the spiking neuron model (LIF) after convolutional projection.

3. The low-energy long-sequence action recognition method based on mamba pulse neural network according to claim 2 is characterized in that: The operation process of the spiking neuron model LIF includes: Membrane potential initialization, input excitation accumulation, membrane potential leakage, and threshold V th The pulse triggers and membrane potential resets.

4. The low-energy long-sequence action recognition method based on mamba pulse neural network according to claim 3 is characterized in that: The process of spatiotemporal context encoding of the embedding vector sequence is performed by adding the spatial position encoding SpatialPosition Embedding and the temporal position encoding Temporal Position Embedding to the patch embedding element by element. Among them, the spatial position encoding adds a code to each Patch embedding vector, which represents the position of the Patch in the original spatial field of view, and is combined with the element-by-element addition method of Patch embedding. The spatial embedding converts the two-dimensional coordinate information into a feature vector through encoding; the time position encoding adds a code to each Patch embedding vector, which represents the position of the Patch in the time series. The goal of time position embedding is to encode the relationship between data points changing over time.

5. The low-energy long-sequence action recognition method based on mamba pulse neural network according to claim 4 is characterized in that: The spatial position coding and temporal position coding adopt a rotation coding method, and the method is as follows: Spatial position embedding realizes the encoding of spatial position information in features by generating a rotation vector for each spatial point. For spatial embedding, the processing process is as follows: the rotation angle θ in each embedding dimension k Calculated using the following formula: Where base is the frequency base, k is the embedding dimension index, and k max To maximize the embedding index, we construct the position label (x, y) of each spatial point by gridding, and apply the rotation angle to generate the embedding feature. For each spatial direction, we use the formula: x embed [x,y]=x·θ k ,and embed [x,y]=y·θ k Embed the (x,y) coordinates in the grid into the rotation matrix to get the rotated position embedding: The goal of time position embedding is to encode the relationship between data points over time, helping the model understand the dynamics of input data in the time dimension. By rotating the embedding mechanism, a unique feature expression is generated for each time point. For each time point t, an embedding value is generated: embed [t] = t·θ k , the embedding vectors at different time points contain information in different frequency ranges.

6. The low-energy long-sequence action recognition method based on mamba spiking neural network according to claim 5 is characterized in that: The conversion process of the raw event stream in the action recognition dataset of the raw event pulse to the embedding representation with position awareness: The original data X undergoes preliminary feature extraction and patch division through 3D convolution. After batch normalization, the pulse layer converts the continuous value features into pulse signals that conform to the characteristics of SNN, and adds position embedding information to give the feature vector the ability to perceive spatiotemporal context, which can be expressed as follows: P=SL patch (BN(Conv3d(X)))+PE Among them, Conv3d(·) is a 3D convolution layer that processes the original input data X and divides the continuous data stream into discrete 3D spatiotemporal patches. BN(·) is a batch normalization layer that normalizes the feature map after convolution. patch (·) is the spiking layer, which converts continuous values into discrete pulse signals and introduces the discharge mechanism of biological neurons. The spiking layer simulates the behavior of biological neurons and only generates output when the input exceeds the preset threshold. PE is position embedding information, which includes spatial position encoding and temporal position encoding. Through element-by-element addition operations and combined with the generated feature vector, the pre-established Mamba spiking neural network model obtains the relative position and order relationship of each patch in the original spatiotemporal data.

7. The low-energy long-sequence action recognition method based on mamba pulse neural network according to claim 1 is characterized in that: The process of inputting the embedding vector sequence containing the location context information into the pre-established mamba spiking neural network model for training includes: The pre-built mamba spiking neural network model consists of L series-connected SpikMamba Blocks and a final Prediction Layer; The internal processing flow of SpikMamba Block includes: The SpikMamba Block receives the embedded image block as input data and first enters the Spike LinearAttention module. The input data is output through two preset paths to obtain the output result. The output result is added element by element to the original input of the SpikMamba Block to obtain the result of the residual connection; The result of the residual connection is input into the Spike Mamba module and output through two preset paths. The Hadamard product is performed on the output of path one and path two to obtain the multiplied output result. The multiplied output result is input into the feedforward network FFN to obtain the FFN output result, and the FFN output result and the residual connection result are added element by element to obtain the second residual connection result as the final output result.

8. The low-energy long-sequence action recognition method based on mamba pulse neural network according to claim 7 is characterized in that: The Linear Attention unit attention mechanism of the Spike Linear Attention module optimizes the self-attention mechanism using kernel functions and feature maps. The formula is as follows: LinearAttention(Q,K,V)=φ(Q)·(φ(K) T ·V) Where φ(·) is a feature mapping function and · represents matrix multiplication; In the Spike Linear Attention module, input data is processed along two paths. In path one, the input is transformed by the linear layer, processed by the convolution layer Conv, and the SiLU activation function σ is applied. It is then divided into two parts and fed into parallel linear layers. Each linear layer is followed by a spike neural network layer SNN, and the output is fed into the Linear Attention calculation unit. The output of LinearAttention is expressed as: Among them, W o , W2 is the learnable weight matrix, represents the Hadamard product, S(·) is the impulse activation function of SNN, σ(·) is the SiLU activation function, and X is the original input.

9. The low-energy long-sequence action recognition method based on mamba spiking neural network according to claim 8 is characterized in that: During use, the Spike Mamba module obtains a multiplied output result by fusing the features of two paths, the discrete state space model SSM and the spiking neural network SNN.

10. The low-energy long-sequence action recognition method based on mamba spiking neural network according to claim 9 is characterized in that: The formula of the discrete state space model SSM structure is as follows: y t =Ch t +Dx t h t represents the hidden state vector at time t, x t Represents the input event stream characteristics at time t, y t As the output prediction result at time t; During the state transition, As a state transition matrix, it dynamically controls how historical information is retained and decayed; As the input projection matrix, it determines how the current input affects the state; C acts as the output projection matrix, mapping the processed hidden state to the prediction space; D is the skip connection matrix, which realizes the path where input information directly affects output; Parameter Matrix and It is input-dependent and is converted from a continuous representation using the zero-order hold (ZOH) method: The discrete state space model SSM dynamically adjusts the memory and forgetting mechanisms according to the input content.

11. A low-energy long-sequence action recognition system based on mamba pulse neural network, characterized by: include: a data processing module, configured to obtain a motion recognition dataset of original event pulses, perform spatiotemporal dimension processing on the motion recognition dataset of original event pulses, and obtain an embedded vector sequence, wherein the dimension of the embedded vector sequence is fixed; The model training module is used to perform spatiotemporal context encoding on the embedded vector sequence to obtain an embedded vector sequence containing position context information, input the embedded vector sequence containing position context information into a pre-established mamba spiking neural network model for training, and output the trained spiking neural network model; The action recognition module is used to obtain the action data of the event pulse to be recognized, input the action data of the event pulse to be recognized into the trained pulse neural network model, and output the final human action recognition result.

Citation Information

Cited By

  • Robot control method based on pulse neural network

    CN121670683A