Human body action recognition method based on pulse neural network and self-attention mechanism

By introducing a multi-headed synaptic filtered self-attention mechanism and pulse feedforward network layer into the pulse neural network, combining time and channel fusion attention layer, the performance bottleneck problem of pulse neural networks in the existing technology in complex space-time dependence tasks is solved, and higher action recognition accuracy and training efficiency are achieved.

CN120047995APending Publication Date: 2025-05-27DONGGUAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411993624.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing Transformer model based on pulsed neural networks shows performance bottlenecks when processing large data sets and cannot compete with traditional artificial neural networks. Especially in complex space-time dependency tasks, the learning ability is limited and it is difficult to deal with long-term dependency features.

Method used

The human body movement recognition method based on pulsed neural network and self-attention mechanism is adopted. By introducing a multi-headed synaptic filtered self-attention mechanism and pulse feedforward network layer, the model's feature capture ability in time and space dimensions is enhanced, and the attention layer is fused with the time and channel to process long-term dependency features.

Benefits of technology

It improves the ability of pulsed neural networks to extract and train spatiotemporal features and trains, enhances the accuracy and robustness of complex action recognition, reduces computing resource consumption, is suitable for low-power scenarios, and shortens training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005224132620000031
    Figure BDA0005224132620000031
  • Figure BDA0005224132620000061
    Figure BDA0005224132620000061
  • Figure BDA0005224132620000071
    Figure BDA0005224132620000071
Patent Text Reader

Abstract

The invention discloses a human body action recognition method based on a pulse neural network and a self-attention mechanism, and the method comprises the steps: building a pulse neural network model based on the self-attention mechanism through employing a pulse attention word segmentation device and a pulse converter model coding module, and building an action data set of an event form; and according to different set time steps, segmenting the data into a plurality of frames of sequence data for model training. The invention relates to the technical field of computer vision. According to the human body action recognition method based on the spiking neural network and the self-attention mechanism, while the advantages of low power consumption and few parameters of the spiking neural network are kept, the spatial-temporal feature extraction capability and training efficiency of the spiking neural network are improved, so that the performance of the spiking neural network in an event-driven task is improved; according to the method, the difference between the pulse neural network and the artificial neural network in performance is reduced, fine time sequence changes can be captured, energy consumption can be reduced while efficient calculation is carried out, the real-time performance and applicability of the model are improved, and accurate recognition of human body actions is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and specifically to a human action recognition method based on spiking neural network and self-attention mechanism. Background Art

[0002] Traditional human action recognition mostly performs action recognition based on RGB videos, depth images, and skeleton data. Such methods usually focus on single action recognition or intention inference, and it is difficult to make full use of the continuity of actions in the time and space dimensions, thus limiting the action classification effect. People have provided new ideas for solving this problem through research on brain-inspired intelligence. As an important model of brain-inspired intelligence, the spiking neural network processes discrete spike signals by simulating the dynamic mechanism of biological neurons, showing potential in processing complex spatio-temporal features. Although the spiking neural network has advantages in energy efficiency and biological neural network simulation, its accuracy in complex tasks still lags significantly behind traditional artificial neural networks. Its main disadvantages lie in limited scalability and unstable training efficiency, especially in dealing with video and time series data.

[0003] Existing spiking neural network methods face problems of gradient disappearance and performance bottlenecks when dealing with complex spatio-temporal dependence tasks. In contrast, traditional artificial neural network models such as Transformer have made remarkable progress in image classification and action recognition through self-attention mechanism, but their high energy consumption and the number of model parameters limit their application in low-power fields. Transformer is famous for its self-attention mechanism, which can effectively capture long-range dependencies and complex features in data, especially performing well in the fields of natural language processing and image processing.

[0004] Existing methods that combine the Transformer architecture with spiking neural networks have made some progress, but still have significant drawbacks. Among them, existing Transformer models based on spiking neural networks exhibit obvious performance bottlenecks when dealing with large datasets and cannot compete with traditional artificial neural networks such as Transformer or convolutional neural networks. This is mainly reflected in the limited learning ability of spiking neural networks (SNNs) when facing large-scale data and their difficulty in handling complex spatio-temporal dependencies. In addition, due to the non-differentiability of the spiking neuron model in spiking neural networks, training methods such as backpropagation through time face problems of vanishing gradients and slow convergence, resulting in low training efficiency. Even when introducing residual networks to alleviate the vanishing gradients, the training efficiency of deep networks is still insufficient. Moreover, although some existing spiking neural network models have introduced self-attention mechanisms to improve the ability to extract spatio-temporal features, due to the way of processing spike signals in spiking neural networks themselves, their ability to capture complex features, especially in terms of long-term dependencies, is still insufficient, thus unable to fully exploit the advantages of Transformer. Based on this, a human action recognition method based on spiking neural networks and self-attention mechanism is proposed to improve the spatio-temporal feature extraction ability and training efficiency of spiking neural networks, ultimately enhancing their performance in event-driven tasks and narrowing the performance gap between spiking neural networks and artificial neural networks. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a human action recognition method based on spiking neural networks and self-attention mechanism, which solves the problems raised in the above background technology.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A human action recognition method based on spiking neural networks and self-attention mechanism specifically includes the following steps:

[0007] S1. Use a dynamic event camera to collect the corresponding actions made by the subjects and establish an action dataset in the form of events;

[0008] S2. Cut the action dataset into several frame sequence data according to different set time steps T, and divide the several frame sequence data into a training set and a test set according to a ratio of 8:2;

[0009] S3. Build a spiking neural network model based on self-attention mechanism using a spiking attention tokenizer and a spiking transformer model encoding module;

[0010] S4. Input the training set data into the spiking neural network model with self-attention mechanism for training to obtain a trained spiking neural network model with self-attention mechanism;

[0011] S5. Input the test set into the trained spiking neural network model with self-attention mechanism for testing, output the predicted ranking after testing, and obtain the most likely action recognition result according to the predicted ranking.

[0012] The present invention is further configured that: the establishment of the action data set of event forms in S1 includes:

[0013] The original data output by the dynamic event camera is in the format of asynchronous discrete events, denoted as E(x, y, i, t, p), where x and y are coordinates in the image, i is the event index, t is the time series, p is the event state, and the event state includes ON event and OFF event. When the dynamic event camera records that the subject has a displacement, it is an ON event. On the contrary, when the dynamic event camera records that the subject has no displacement, it is an OFF event.

[0014] The present invention is further configured that: the spiking attention tokenizer in S3 includes a spiking neuron signal conversion layer and a time and channel fusion attention layer;

[0015] The spiking neuron signal conversion layer and time are used to convert the frame sequence data into discrete spiking sequence data;

[0016] The time and channel fusion attention layer is used to extract the time and channel information of the action displacement points from the discrete spiking sequence data and output the preprocessed spiking sequence data.

[0017] The present invention is further configured that: the spiking neuron signal conversion layer is used to receive the frame sequence data through spiking neurons and convert it into spiking sequence data. After the spiking sequence data is sequentially processed by two-dimensional convolution processing, regularization processing, and max-pooling layer processing, spatio-temporal features are extracted and the firing frequency of the spiking sequence data is increased;

[0018] The discrete equation distribution of the spiking neuron includes:

[0019]

[0020] Among them, is the charging discrete equation, is the discharging discrete equation, is the reset discrete equation, Θ(x) is the step function, τ is the time constant, V t-1 is the membrane potential voltage at the previous moment of charging, X t is the external input current, V reset is the reset voltage. The reset voltage is usually set to 0. At this time, the reset method adopts the soft reset method, that is, the membrane potential after charging minus the threshold voltage, rather than directly becoming the reset voltage V reset .

[0021] The present invention is further configured such that: the time and channel fusion attention layer is used to extract spatio-temporal feature information in the pulse sequence data output by the pulse neuron signal conversion layer through time attention and channel attention, and obtain preprocessed pulse sequence data after attention fusion.

[0022] The present invention is further configured such that: the pulse converter model encoding module includes a multi-head synaptic filtering pulse self-attention module and a pulse feed-forward network layer;

[0023] The multi-head synaptic filtering pulse self-attention module is used to receive the preprocessed pulse sequence data, simultaneously process different parts thereof, capture spatio-temporal features, and output weighted pulse sequence data;

[0024] The pulse feed-forward network layer is used to perform non-linear activation processing on the weighted pulse sequence data.

[0025] The present invention is further configured such that: after receiving the preprocessed pulse sequence data, the multi-head synaptic filtering pulse self-attention module sequentially passes through a pulse neuron layer and a synaptic filter to correct the pulse signal, obtains the corrected pulse sequence data, and performs convolution and regularization processing respectively in three attention heads through the multi-head attention mechanism. After neuron extraction, the weight matrices of query, key, and value are obtained and then dot product operations are performed. The attention weights of the three attention heads are weighted and summed, and then sequentially pass through a pulse neuron, convolution, and regularization processing to output the weighted pulse sequence data;

[0026] Among them, the weight matrices of query, key, and value are respectively set in three attention heads.

[0027] The present invention is further configured such that: after receiving the weighted pulse sequence data, the pulse feed-forward network layer performs non-linear activation on the weighted pulse sequence data through pulse neuron, convolution, and regularization processing, combined with linear transformation.

[0028] The present invention provides a human action recognition method based on a spiking neural network and a self-attention mechanism. It has the following beneficial effects:

[0029] (1) By introducing the multi-head synaptic filtering self-attention mechanism, the present invention enhances the model's ability to capture complex features in the time and space dimensions. It can not only allocate attention to different time steps and channels of the input, but also process long-term dependencies through synaptic filtering, ensuring more accurate recognition of key temporal changes in the action sequence, thereby greatly improving the accuracy of action classification. The spatio-temporal feature extraction ability is effectively improved, making it have higher accuracy and robustness in the action recognition task.

[0030] (2) By introducing a pulse feedforward network layer and a pulse neuron signal conversion layer, the present invention greatly reduces the consumption of computing resources while maintaining high precision. The pulse feedforward network layer utilizes the event-driven characteristic of pulse neurons, enabling computation to occur only when the input pulse signal arrives, reducing unnecessary computations and enhancing computational efficiency. This event-driven mechanism enables the model to perform excellently in power-sensitive scenarios such as embedded devices or the Internet of Things, not only reducing the operating cost but also greatly enhancing the real-time response ability and effectively saving energy consumption.

[0031] (3) Through the design of residual connections, the present invention maintains the stable transmission of gradients in the deep pulse network, avoiding the rapid decay of gradients. The introduction of residual connections not only improves the training efficiency of the model but also shortens the training time, enabling the network to achieve higher accuracy in a shorter time. While shortening the training time, it improves the convergence and stability of the model.

[0032] (4) The present invention performs spatio-temporal fusion processing on the input pulse signal through a pulse attention tokenizer, thereby improving the robustness of the model in complex environments. It utilizes a fusion attention mechanism of time and channels, which can effectively separate useful spatio-temporal information and exclude the interference of noise, enabling the model to maintain a high accuracy rate in various complex scenarios and having wide applicability in practical application scenarios such as low light and background interference. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic diagram of the processing flow of the present invention;

[0034] Figure 2 is a schematic diagram of the format of frame sequence data in the present invention;

[0035] Figure 3 is a schematic diagram of the process of splitting the event form data stream into frame sequence data in the present invention;

[0036] Figure 4 is a schematic diagram of the internal structure process model of the pulse attention tokenizer of the present invention;

[0037] Figure 5 is a schematic diagram of the internal structure flow chart of the multi-head synaptic filtering pulse self-attention module of the present invention;

[0038] Figure 6 is a schematic diagram of the internal structure flow chart of the pulse feedforward network layer of the present invention;

[0039] Figure 7 is a schematic diagram of the training flow chart of the action sequence model of the present invention;

[0040] Figure 8 is a schematic diagram of the comparison of the model fitting situation in the embodiment of the present invention;

[0041] Figure 9 This is the abnormal action warning flowchart in the embodiments of the present invention. Specific Embodiments

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0043] Please refer to Figures 1-9 , the embodiments of the present invention provide the following technical solutions: a human action recognition method based on a spiking neural network and a self-attention mechanism, specifically including the following steps:

[0044] S1. Use a dynamic event camera to collect the corresponding actions made by the subject, and establish an action dataset in the form of events.

[0045] The original data output by the dynamic event camera is in the format of asynchronous discrete events, denoted as E(x, y, i, t, p), where x and y are the coordinates in the image, i is the event index, t is the time series, p is the event state, and the event state includes ON event and OFF event. When the dynamic event camera records that the subject has a displacement, it is an ON event. Conversely, when the dynamic event camera records that the subject has no displacement, it is an OFF event.

[0046] S2. Cut the action dataset into several frame sequence data according to the set different time steps T, and divide the several frame sequence data into a training set and a test set according to a ratio of 8:2.

[0047] As a detailed description, for the segmentation of the action dataset, a certain frame data in the segmented frame sequence data is denoted as F(j):

[0048]

[0049] Among them, the index j l and j r are respectively the start and end event indexes of the current frame, and Θ x,y,p (E x,y,i,t,p ) represents an indicator function. As shown in Appendix Figure 2 and Appendix Figure 3 shown, on the left side of the figure is the initial data format in the form of events. After being cut into T frames, it obtains the T×H×W format, where T, H, and W respectively represent the time step, the height of the image, and the width of the image.

[0050] S3. Build a spiking neural network model based on the self-attention mechanism using a spiking attention tokenizer and a spiking converter model encoding module. Among them, the spiking attention tokenizer can extract key feature information from the image, and this feature information will be used for subsequent model training and prediction, and can better capture local and global features in the image, thereby improving the accuracy of image recognition and classification. As shown in the appendix Figure 1 The spiking attention tokenizer includes a spiking neuron signal conversion layer and a time and channel fusion attention layer.

[0051] Specifically, on the one hand, the spiking neuron signal conversion layer and time are used to convert frame sequence data into discrete spiking sequence data;

[0052] The spiking neuron signal conversion layer is used to receive frame sequence data through spiking neurons and convert it into spiking sequence data. As shown in the appendix Figure 4 After the spiking sequence data passes through two-dimensional convolution processing, regularization processing, and max-pooling layer processing in sequence, spatio-temporal features are extracted and the firing frequency of the spiking sequence data is increased;

[0053] The discrete equation distribution of spiking neurons includes:

[0054]

[0055] Among them, is the charging discrete equation, is the discharging discrete equation, is the reset discrete equation, Θ(x) is the step function, τ is the time constant, V t-1 is the membrane potential voltage at the previous moment of charging, X t is the external input current, V reset is the reset voltage, usually set to 0. At this time, the reset method adopts the soft reset method, that is, the membrane potential after charging minus the threshold voltage, rather than directly becoming the reset voltage V reset .

[0056] On the other hand, the time and channel fusion attention layer is used to extract the time and channel information of the action displacement points from the discrete spiking sequence data and output the preprocessed spiking sequence data;

[0057] As shown in the appendix Figure 4 The time and channel fusion attention layer is used to extract spatio-temporal feature information in the spiking sequence data output by the spiking neuron signal conversion layer through time attention and channel attention, and obtain the preprocessed spiking sequence data after attention fusion;

[0058] As a further illustration, the time and channel fusion attention layer is used to analyze the spike signal data extracted in the spiking neuron signal conversion layer and extract the time and channel information of the action displacement points therein. Traditional spiking neural networks using the Transformer structure usually only utilize the step of spiking neuron signal conversion and lack the feature extraction work on the time and channel information of the input signal during the data processing stage. At this time, the spiking neuron signal conversion layer and the time and channel attention fusion layer are jointly used to form the spiking attention tokenizer module, which can effectively make up for the problem of the low recognition rate of complex actions in previous models.

[0059] As a preferred solution, as shown in the appendix Figure 1 The spiking converter model encoding module includes a multi-head synaptic filtering spiking self-attention module and a spiking feed-forward network layer.

[0060] As a detailed description, on the one hand, the multi-head synaptic filtering spiking self-attention module is used to receive the preprocessed spiking sequence data, process different parts of it simultaneously, capture the spatio-temporal features, and output the weighted spiking sequence data.

[0061] As shown in the appendix Figure 5 After receiving the preprocessed spiking sequence data, the multi-head synaptic filtering spiking self-attention module sequentially passes through the spiking neuron layer and the synaptic filter for spiking signal correction, obtains the corrected spiking sequence data, and performs convolution and regularization processing in three attention heads respectively through the multi-head attention mechanism. After neuron extraction, the query Q, key K, and value V matrices, and the weight matrices of the key and value are obtained and then dot product operations are performed. Among them, the weight matrices of the query, key, and value are respectively set in the three attention heads, and the attention weights of the three attention heads are weighted and summed, and then pass through the spiking neuron, convolution, and regularization processing in sequence to output the weighted spiking sequence data. Among them, a synaptic filtering mechanism is added, and by simulating the gradual influence of biological synapses on neuron excitation, the ability of the model to handle long-term dependencies is increased. Combining the synaptic filtering and the multi-head self-attention mechanism can capture the features of the input spiking sequence in space and time simultaneously, greatly improving the performance of the model on complex spatio-temporal dependence tasks. Specifically, the calculation process of the multi-head synaptic filtering spiking self-attention module includes the following steps:

[0062] A1. Linear transformation of the input:

[0063] Q = X · W Q , K = X · W K , V = X · W V

[0064]

[0065] Among them, W Q , WK and W V are the weight matrices for query, key, and value respectively, and d k is the dimension of the key vector;

[0066] A2. Calculation of the synaptic filtering mechanism:

[0067] To introduce the temporal characteristics of the spiking neural network, the attention weight W is corrected by a spike filter σ(t) to simulate the temporal decay of the spike signal by biological synapses. Its calculation method includes:

[0068] W'(t) = W · σ(t

[0069]

[0070] where τ is the time constant and W'(t) is the attention weight after temporal decay;

[0071] A3. Results of the output calculation:

[0072] The adjusted attention weight W'(t) is used to perform weighted summation on the value matrix V to generate the output of the multi-head attention:

[0073] Attention(Q, K, V) = W′(t) · V

[0074] MultiHead(Q, K, V) = Concat(head 1 , …, head h ) · W O

[0075] where W O is the trainable weight matrix used to map the outputs of multiple heads back to the original dimension after concatenation.

[0076] On the other hand, the spiking feedforward network layer is used to perform non-linear activation processing on the weighted spike sequence data. As shown in Figure 6 , after receiving the weighted spike sequence data, the spiking feedforward network layer performs non-linear activation on the weighted spike sequence data through spiking neurons, convolution, and regularization processing, combined with linear transformation. Among them, the spiking feedforward network layer is a key component in the overall network responsible for non-linear mapping of the input spike sequence. It plays a role in further processing and transformation of spatio-temporal features in the network. Similar to the traditional feedforward neural network layer, the spiking feedforward network layer can process the input features through a series of spiking neurons, but it also retains the event-driven mechanism of the spiking neural network, thus having higher energy efficiency and real-time performance.

[0077] Furthermore, the pulsed feedforward network layer further processes the input features by combining linear transformation and the non-linear activation of pulsed neurons. When enhancing the non-linear ability of the model, it also maintains high efficiency and low latency through the pulsed mechanism, especially playing an important role in the processing of spatio-temporal features and the solution of complex tasks.

[0078] S4. To improve the model's recognition ability for processing complex action information and enhance the model's robustness by inputting action sequences with different time steps and different action postures, the training set data is input into the pulsed neural network model with self-attention mechanism for training to obtain the trained pulsed neural network model with self-attention mechanism. As a detailed description, when the training set is input into the pulsed neural network model with self-attention mechanism for training according to different time steps and different action postures, assuming there are N different actions and action sequences divided according to time steps, the steps of learning, processing, and analysis performed by the action sequences input into the pulsed neural network model with self-attention mechanism include:

[0079] B1. Divide a certain action into T equal parts according to the time step T, and then perform feature fusion through methods such as dot product, weighted summation, or taking the maximum mean value in sequence, and perform the conversion and extraction of pulsed signals on the action sequence data divided according to time.

[0080] B2. The data input into the pulsed neural network model with self-attention mechanism also includes different types of action postures, such as actions like shaking hands, playing the guitar, walking, throwing things, pushing and fighting, etc., to improve the generalization ability of the model and further improve the recognition accuracy of the model when processing various complex actions.

[0081] B3. Input the operations of B1 and B2 into a pre-designed learning model for learning and training, and complete the overall feature fusion through pulsed operation by the pulsed neuron LIF layer, effective feature extraction by two-dimensional convolution operation, regularization, and max pooling processing, and finally realize the understanding of pedestrian action intentions input into the pulsed neural network model with self-attention mechanism, as shown in the appendix. Figure 7 as shown.

[0082] S5. Input the test set into the trained pulsed neural network model with self-attention mechanism for testing, output the predicted ranking after testing, and obtain the most likely action recognition result according to the predicted ranking.

[0083] Comparative experiment

[0084] Apply the 3D convolutional neural network C3D of traditional artificial neural networks, ResNet-SNN-18 based on spiking residual networks, and the spiking neural network model with self-attention mechanism to the action recognition of the same video, compare the experimental results, and obtain the evaluation index table shown in the following table, which shows the action classification accuracy under different simulated time steps T:

[0085]

[0086] As can be seen from the above table, the action recognition method corresponding to the spiking neural network model with self-attention mechanism has a high accuracy at different simulated time steps, which further illustrates the accuracy and reliability of this model.

[0087] Compare the fitting situations of the model ResNet-SNN-18 based on spiking residual networks, the point-based end-to-end spiking neural network model SpikePoint, and the spiking neural network model with self-attention mechanism. As shown in the appendix Figure 8 It can be seen that the accuracy of the spiking neural network model with self-attention mechanism rises rapidly within the first 100 rounds and approaches the fitting state, while the other two models show fluctuations in accuracy. Finally, the spiking neural network model with self-attention mechanism reaches the maximum action classification accuracy of 72.19% near 600 rounds, demonstrating the high convergence ability and high action classification accuracy of this model.

[0088] As an extended explanation, after implementing the action recognition and understanding of video data through the human action recognition method based on spiking neural networks and self-attention mechanism, when dangerous actions are recognized in the identified actions, such as actions like fighting and pushing, early warning processing will be carried out in a timely manner, that is, an early warning signal will be issued to avoid further exacerbating the severity of the situation. As shown in the appendix Figure 9 It can be seen that each time the output action recognition result will be analyzed. When a dangerous action is found in the action, an early warning signal will be issued, and when no danger is recognized, the analysis of the next action recognition result will be carried out.

[0089] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A human action recognition method based on a pulse neural network and a self-attention mechanism, characterized by: The specific steps include: S1. Use a dynamic event camera to collect the corresponding actions made by the subject and establish an action dataset in the form of an event; S2, dividing the action data set into a number of frame sequence data according to different set time steps T, and dividing the frame sequence data into a training set and a test set according to a ratio of 8:2; S3, using the pulse attention tokenizer and pulse converter model encoding module to build a pulse neural network model based on the self-attention mechanism; S4, inputting the training set data into the pulse neural network model of the self-attention mechanism for training, and obtaining the trained pulse neural network model of the self-attention mechanism; S5. Input the test set into the trained pulse neural network model of the self-attention mechanism for testing, output the predicted ranking after the test, and obtain the most likely action recognition result based on the predicted ranking.

2. The human motion recognition method based on pulse neural network and self-attention mechanism according to claim 1, characterized in that: The establishment of the action data set of the event form in S1 includes: The raw data format output by the dynamic event camera is an asynchronous discrete event, denoted as E(x, y, i, t, p), where x and y are the coordinates in the image, i is the event index, t is the time series, and p is the event state. The event state includes ON events and OFF events. If the dynamic event camera records the displacement of the subject, it is an ON event. Otherwise, if the dynamic event camera records that the subject does not move, it is an OFF event.

3. The human motion recognition method based on pulse neural network and self-attention mechanism according to claim 1, characterized in that: The pulse attention word segmenter in S3 includes a pulse neuron signal conversion layer and a time and channel fusion attention layer; The pulse neuron signal conversion layer and time are used to convert frame sequence data into discrete pulse sequence data; The time and channel fusion attention layer is used to extract the time and channel information of the action displacement points from the discrete pulse sequence data, and output the preprocessed pulse sequence data.

4. The human motion recognition method based on pulse neural network and self-attention mechanism according to claim 3, characterized in that: The pulse neuron signal conversion layer is used to receive frame sequence data through pulse neurons and convert it into pulse sequence data. After the pulse sequence data is processed by two-dimensional convolution, regularization and maximum pooling layer in sequence, the spatiotemporal features are extracted and the excitation frequency of the pulse sequence data is increased; The discrete equation distribution of the spiking neuron includes: in, is the charging discrete equation, is the discharge discrete equation, is to reset the discrete equation, Θ(x) is the step function, τ is the time constant, V t-1 is the membrane potential voltage before charging, X t is the external input current, V reset is the reset voltage.

5. The human motion recognition method based on pulse neural network and self-attention mechanism according to claim 4, characterized in that: The time and channel fusion attention layer is used to extract the spatiotemporal feature information in the pulse sequence data output by the pulse neuron signal conversion layer through time attention and channel attention, and obtain the preprocessed pulse sequence data after attention fusion.

6. The human motion recognition method based on pulse neural network and self-attention mechanism according to claim 5, characterized in that: The pulse converter model encoding module includes a multi-head synaptic filter pulse self-attention module and a pulse feedforward network layer; The multi-head synaptic filter pulse self-attention module is used to receive the pre-processed pulse train data, and process different parts of it simultaneously, capture the spatiotemporal features, and output weighted pulse train data; The pulse feedforward network layer is used to perform nonlinear activation processing on weighted pulse sequence data.

7. The human motion recognition method based on pulse neural network and self-attention mechanism according to claim 6, characterized in that: After receiving the preprocessed pulse sequence data, the multi-head synaptic filter pulse self-attention module sequentially passes through the pulse neuron layer and the synaptic filter to correct the pulse signal, obtains the corrected pulse sequence data, and performs convolution and regularization processing in the three attention heads respectively through the multi-head attention mechanism, extracts the neurons, obtains the weight matrix of the query, key and value, and then performs a dot multiplication operation, performs a weighted summation on the attention weights of the three attention heads, sequentially passes through the pulse neuron, convolution and regularization processing, and outputs the weighted pulse sequence data; The weight matrices of query, key, and value are set in three attention heads respectively.

8. The human motion recognition method based on pulse neural network and self-attention mechanism according to claim 7, characterized in that: After receiving the weighted pulse sequence data, the pulse feedforward network layer performs nonlinear activation on the weighted pulse sequence data through pulse neurons, convolution and regularization processing combined with linear transformation.

Citation Information

Cited By

  • Low-energy-efficiency crack detection method and device based on spiking neural network

    CN121259538A