Target identification method, terminal, storage medium and computer program product

By using pulsed neural network model in virtual reality scenarios combined with attention mechanism for target recognition, the high power consumption and low accuracy problems of traditional methods are solved, and low latency and high precision target recognition is achieved, which is suitable for VR gesture recognition and eye tracking.

CN120356135APending Publication Date: 2025-07-22BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510520965.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional object recognition methods have problems such as high power consumption, high latency and low recognition accuracy in complex environments in virtual reality scenarios. Conventional SNN models have low recognition accuracy and low efficiency in complex scenarios, and are prone to pulse degradation problems.

Method used

The pulse neural network model is used to combine the attention mechanism to obtain the event data flow, accumulate the event frames, and use the time, space and channel attention mechanism to optimize the membrane potential to generate pulse signals for target recognition.

Benefits of technology

It realizes low power consumption and low latency target recognition, significantly improving recognition accuracy and real-time performance, and is suitable for VR gesture recognition and eye tracking in complex environments, reducing latency and improving response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356135A_ABST
    Figure CN120356135A_ABST
Patent Text Reader

Abstract

The invention provides a target identification method, a terminal, a storage medium and a computer program product. The target identification method comprises the following steps: acquiring a to-be-identified event data stream; performing accumulation processing on the event data flow according to a preset time step number to obtain a plurality of event frames; and inputting the event frame to a pre-trained spiking neural network model for processing, and determining a target recognition result according to a pulse signal output by the spiking neural network model, the spiking neural network model being a leakage integral issuing neuron model combined with an attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology. More specifically, it relates to an object recognition method, a terminal, a storage medium, and a computer program product. Background Art

[0002] In a virtual reality (VR) scenario, object recognition, such as gesture recognition and eye movement tracking, is a core technology for realizing immersive interaction. Traditional object recognition usually relies on visual sensors and artificial neural network (ANN) algorithms. However, traditional object recognition methods have defects such as high power consumption, high latency, and low recognition accuracy in complex environments. Summary of the Invention

[0003] The purpose of the present disclosure is to provide an object recognition method, a terminal, a storage medium, and a computer program product to solve at least one of the above technical problems.

[0004] To achieve the above object, the present disclosure adopts the following technical solutions:

[0005] The first aspect of the present disclosure provides an object recognition method, including the following steps:

[0006] Obtain an event data stream to be recognized;

[0007] Accumulate and process the event data stream according to a preset number of time steps to obtain a plurality of event frames;

[0008] Input the event frames into a pre-trained spiking neural network model for processing, and determine the object recognition result according to the spike signals output by the spiking neural network model. The spiking neural network model is a leaky integrate-and-fire neuron model combined with an attention mechanism.

[0009] Optionally, the attention mechanism includes at least one of a temporal attention mechanism, a spatial attention mechanism, and a channel attention mechanism. The spiking neural network model includes at least one of a temporal attention module, a spatial attention module, and a channel attention module.

[0010] Optionally, the spiking neural network model includes a time-domain network structure with a preset number of time steps. Each time-domain network structure inputs one of the event frames. The time-domain network structure includes a convolutional block, a temporal attention module, an integration module, a spatial attention module, a channel attention module, and a spiking leaky integrate-and-fire module connected in sequence. The step of inputting the event frames into the pre-trained spiking neural network model for processing includes: inputting the multiple event frames into the corresponding time-domain network structures, and each time-domain network structure processes the input event frame and outputs a corresponding spiking signal. For the t-th time-domain network structure, where 1 < t ≤ T and T is the preset number of time steps, the step of this time-domain network structure processing the input event frame and outputting a corresponding spiking signal includes:

[0011] Performing convolutional processing on the input event frame using the convolutional block to obtain the current membrane potential of the event frame;

[0012] Processing the current membrane potential of each event frame using the temporal attention module to generate a first membrane potential corresponding to each event frame one by one;

[0013] Integrating and processing the first membrane potential and the fourth membrane potential output by the previous time-domain network structure using the integration module to form a second membrane potential;

[0014] Processing the second membrane potential using the channel attention module and the spatial attention module to generate a third membrane potential;

[0015] Processing the third membrane potential using the spiking leaky integrate-and-fire module to generate and output the spiking signal and the fourth membrane potential, and inputting the fourth membrane potential into the integration module of the next time-domain network structure.

[0016] Optionally, after the step of performing convolutional processing on the event frame using the convolutional block, the following steps are further included:

[0017] Performing batch normalization on the event frame after convolutional processing;

[0018] Performing average pooling on the event frame after batch normalization to obtain the current membrane potential.

[0019] Optionally, the step of processing the current membrane potential of each event frame using the temporal attention module to generate a first membrane potential corresponding to each event frame one by one includes:

[0020] Performing average pooling and max pooling on the current membrane potential of each event frame;

[0021] Performing non-linear mapping on the average pooling result and the max pooling result respectively to generate a first initial temporal weight corresponding to the average pooling result and a second initial temporal weight corresponding to the max pooling result;

[0022] Perform weighted normalization on the first initial time weight and the second initial time weight to generate a time attention weight;

[0023] Use the time attention weight to recalibrate the current membrane potential to generate the first membrane potential.

[0024] Optionally, the step of using a channel attention module and a spatial attention module to process the second membrane potential to generate a third membrane potential includes:

[0025] Perform average pooling and max pooling on the second membrane potential of the event frame;

[0026] Perform non-linear mapping on the average pooling result and the max pooling result respectively to generate a first initial channel weight corresponding to the average pooling result and a second initial channel weight corresponding to the max pooling result;

[0027] Perform weighted normalization on the first initial channel weight and the second initial channel weight to generate a channel attention weight;

[0028] Use the channel attention weight to recalibrate the second membrane potential to generate an intermediate membrane potential;

[0029] Use the spatial attention module to process the intermediate membrane potential to generate the third membrane potential.

[0030] Optionally, the step of using the spatial attention module to process the intermediate membrane potential to generate the third membrane potential includes:

[0031] Perform average pooling and max pooling on the intermediate membrane potential of the event frame;

[0032] Perform concatenation and convolution processing on the average pooling result and the max pooling result to generate an initial spatial weight;

[0033] Perform normalization on the initial spatial weight to generate a spatial attention weight;

[0034] Use the spatial attention weight to recalibrate the intermediate membrane potential to generate the third membrane potential.

[0035] Optionally, before the step of accumulating the event data stream according to a preset number of time steps to obtain multiple event frames, it further includes:

[0036] Downsample the event data stream to reduce the resolution of the event data stream to a target resolution.

[0037] Optionally, before the step of inputting the event frame into a pre-trained spiking neural network model, it further includes:

[0038] Obtain training data, and use the training data to train an initial spiking neural network model to obtain the pre-trained spiking neural network model. The training data includes first training data and / or second training data. The first training data is an event data stream output by an event camera, and the second training data is an event data stream obtained by simulating image data.

[0039] Optionally, the step of using the training data to train the initial spiking neural network model to obtain the pre-trained spiking neural network model includes:

[0040] Obtain the spike signals output by each time-domain network structure during the current training process, and calculate the mean value of each spike signal;

[0041] Calculate the mean square error between the mean value and the true value;

[0042] End the training process of the initial spiking neural network model when the mean square error reaches a preset error threshold or the number of traversals of the training data reaches the traversal number threshold.

[0043] The second aspect of the present disclosure provides a terminal, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the target recognition method described above are implemented.

[0044] The third aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the target recognition method described above are implemented.

[0045] The fourth aspect of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the target recognition method described above are implemented.

[0046] The beneficial effects of the present disclosure are as follows:

[0047] The target recognition method of the embodiments of the present disclosure uses event data for target recognition, has the advantages of low power consumption, low latency, and small data volume, and accumulates the event data stream to obtain an event frame instead of a traditional image frame, which can effectively reduce the area where most of the event data is zero and improve the processing efficiency. When applied to application scenarios such as VR gesture recognition or eye movement tracking, the latency can be significantly reduced and the response speed can be increased. In addition, the target recognition method combines an attention mechanism in the spiking neural network model, so that key spatio-temporal features can be accurately focused, the accuracy and real-time performance of target recognition can be significantly improved, the spike degradation problem during the training process can be alleviated, and the accuracy and reliability of target recognition in complex environments such as high-speed movement, strong light, or weak light can be ensured. Description of the Drawings

[0048] The following further elaborates on the specific implementation manners of the present disclosure in conjunction with the accompanying drawings.

[0049] Figure 1 It is a flowchart of the target recognition method provided for the embodiments of the present disclosure;

[0050] Figure 2 It is a schematic diagram of the network architecture of the spiking neural network model provided for the embodiments of the present disclosure;

[0051] Figure 3 It is Figure 2 a schematic diagram of the structure of the nth layer of the spiking neural network model in

[0052] Figure 4 It is Figure 3 a schematic diagram of the structure of the integration module and the spiking leaky integrate-and-fire module in network layer n corresponding to the tth time step in

[0053] Figure 5 It is a flowchart of the spiking neural network model for processing the input spikes and outputting the prediction result. Specific Implementation Manners

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions of the embodiments of the present disclosure in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0055] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, the terms such as "a", "an", or "the" do not denote a quantity limitation, but mean that there is at least one. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0056] To better understand the technical solutions of the object recognition method, terminal, storage medium and computer program product of the present disclosure, the design concept of the present disclosure will be briefly introduced first.

[0057] As an asynchronous sensor, the event camera has changed the traditional way of obtaining visual information. Traditional image sensors sample according to the clock, while the event camera samples depending on the light changes in the dynamic scene, which has great similarities with the signal transmission of the retina system. For a certain pixel point, if there is no change in the light intensity at this pixel point in the scene, the event camera will not produce a corresponding output; if the change in the light intensity in the scene exceeds a certain threshold, the event camera will output the corresponding polarity. This output method has no concept of "frame", liberating the speed limit at the source, and events can be output like point clouds. With its microsecond-level time resolution, low latency, high dynamic range and low power consumption characteristics, the event camera has extremely high application potential in challenging scenarios such as high speed and high dynamic range. The data format of the event camera is generally a quadruple (x, y, t, p), where x and y represent the pixel positions where the event occurs, t represents the timestamp of the event occurrence, and p represents the polarity. When p is -1, it means a decrease in brightness, and when p is +1, it means an increase in brightness. The event data can represent: at what time, which pixel point, and whether there is an increase or decrease in brightness.

[0058] Compared with other sensors, the event camera has great advantages in scenarios such as real-time interaction systems. Especially under uncontrolled lighting conditions, the event camera has the advantages of low latency, low power consumption, and sensitivity to light changes, which are lacking in traditional image sensors. Among them, common event cameras include the Dynamic Vision Sensor (DVS for short).

[0059] When using event camera data for object recognition, the traditional method is to convert it into an image frame and then process it through artificial neural networks such as convolutional neural networks and residual neural networks. However, most of the areas in the converted image frame are 0. If the artificial neural network also calculates this part of the area normally, it will cause unnecessary resource consumption, which will offset the low latency and low data volume advantages of the event camera and exacerbate the consumption of computing resources.

[0060] Therefore, an alternative solution using a Spiking Neural Network (SNN for short) has been proposed. SNN simulates the biological neuron communication mechanism with pulse signals, and its event-driven characteristics are naturally compatible with event data: it only triggers calculations for non-zero events, avoiding the multiplication and addition operations of traditional ANNs and significantly reducing power consumption. However, the conventional SNN model still has the following deficiencies:

[0061] (1) Conventional SNNs have low recognition accuracy and efficiency in complex scenarios, such as high-speed movement, complex lighting environments, multi-gesture overlaps, rapid eye movements, etc.;

[0062] (2) In long-term sequence training of conventional SNNs, the sparsity of pulse signals is prone to cause gradient vanishing or explosion, that is, there is a problem of pulse degradation, which limits the depth and generalization ability of the SNN model.

[0063] To solve at least one of the above technical problems, embodiments of the present disclosure provide an object recognition method, a terminal, a storage medium, and a computer program. The specific embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0064] Please refer to Figure 1 , Figure 1 which is a flowchart of the object recognition method provided by the embodiments of the present disclosure. As Figure 1 shown, it includes the following steps:

[0065] Step S101, obtain an event data stream to be recognized.

[0066] In the embodiments of the present disclosure, the event data stream specifically refers to a data stream in the data format of an event camera, which can either be a data stream directly collected by an event camera or a data stream in the data format of an event camera obtained by simulating image data.

[0067] Specifically, the i-th event e i in the event data stream can be expressed as e i= (x i , y i , t’, p i ), where x i , y i represents the pixel position where the event e i occurs, that is, the spatial position of the event e i , t’ represents the timestamp when the event e i occurs, and p i represents the polarity of the event e i , and the value of p i is -1 or +1. For example, the ON event is +1 and the OFF event is -1.

[0068] The set E t’ of events occurring at the time point t’ is expressed as E t’ ={e i |e i =(x i , y i , t’, p i )}, and the event set E t’ can be represented by a pulse tensor. The event set Et’ The impulse tensor S at time t' t’ Expressed as That is, the impulse tensor S t’ is a three-dimensional tensor representing the set of events at a specific time point t', with dimensions Where h0*w0 represents the image resolution of the event camera, 2 represents two channels, corresponding to the polarities of ON and OFF events (such as +1 and -1), and each channel records the number or intensity of events of the corresponding polarity.

[0069] Step S102: performing accumulation processing on the event data stream according to a preset number of time steps to obtain a plurality of event frames.

[0070] In spiking neural networks, time step and number of time steps are two key concepts. Time step refers to the division of continuous time into discrete intervals (such as 1 millisecond), each interval is called a time step. This discretization allows the simulation of continuous dynamic processes of neurons through iterative calculations, such as changes in membrane potential and the transmission of pulses. The number of time steps refers to the number of time steps, which is a fixed parameter that has been determined before the spiking neural network training. Correspondingly, the spiking neural network model will include multiple network structures that correspond one to one to the number of time steps, and each network structure will include multiple network layers. In order to distinguish and represent, multiple network structures in the time dimension are represented as time domain network structures.

[0071] In the embodiment of the present disclosure, the preset time step number T is the number of time steps of the spiking neural network preset by the model designer. Multiple pulse signals can be obtained by accumulating the event data stream using the preset time step number T, each pulse signal representing an event frame, so that the event data stream can be converted into a real-valued event frame with a new frame rate.

[0072] When the event data stream is accumulated, the original time resolution (that is, the original time window) dt' of the event data stream needs to be adjusted by using the scaling factor a to obtain a new time resolution (that is, a new time window) dt. The new time resolution dt = dt'*a, then the pulse tensor S at a continuous time point is t’ can form a subset E t ={S t’}, where t'∈[a*t,a*(t+1)-1]. For example, the original time resolution dt' is 1ms, that is, 1000 frames are generated per second. If a=10, then dt=10ms, that is, the frame rate per second is reduced to 100. For example, when t=0, the time window is [0,a-1], and when t=1, the time window is [a,2a-1].

[0073] Afterwards, for the subset E tthe pulsed tensor S therein t’ Accumulate according to a new time window to generate an event frame for each time step. Exemplarily, the input representation of the first network layer at the t-th time step is S t,0 and can be expressed as where q() represents element-wise accumulation, that is, for the subset E t the pulsed tensor S therein t’ Accumulate according to the time window, t ∈ {1, 2,..., T}. The accumulated event frame is the result of the polarity accumulation of events within the time window, where positive and negative events may cancel each other out. T is the total number of time steps, that is, the preset number of time steps

[0074] Through the above steps, the asynchronous event data stream can be converted into a tensor sequence with a fixed frame rate, enabling the pulsed neural network model to directly process it, which is suitable for dynamic scene analysis. And by adjusting the scaling factor a, the balance between time resolution and computational efficiency can be achieved

[0075] The object recognition method of the embodiments of the present disclosure accumulates the event data stream to obtain an event frame instead of a traditional image frame, which can effectively reduce the regions mostly filled with zeros in the event data, reduce computational redundancy and resource consumption, and improve processing efficiency. This method is applied to high-speed and high-dynamic-range VR environments, such as application scenarios like VR gesture recognition or eye movement tracking, which can significantly reduce latency and improve response speed, and give full play to the advantages of the event camera dataset in real-time interaction

[0076] Step S103: Input the event frame into a pre-trained pulsed neural network model for processing, and determine the object recognition result according to the pulsed signal output by the pulsed neural network model. The pulsed neural network model is a Leaky Integrate-and-Fire Neuron (LIF) model combined with an attention mechanism

[0077] Optionally, the attention mechanism is a multi-dimensional attention mechanism, and the attention mechanism includes at least one of a temporal attention mechanism, a spatial attention mechanism, and a channel attention mechanism. Correspondingly, the pulsed neural network model includes at least one of a temporal attention module, a spatial attention module, and a channel attention module

[0078] The embodiments of the present disclosure are described by taking the attention mechanism including a temporal attention mechanism, a spatial attention mechanism, and a channel attention mechanism as an example. At this time, please refer to Figures 2 to 4 , Figure 2 which is the schematic diagram of the network architecture of the pulsed neural network model provided by the embodiments of the present disclosure Figure 3 is Figure 2 the schematic diagram of the structure of the n-th layer of the pulsed neural network model in Figure 4 isFigure 3 The structural schematic diagram of the integration module and the spiking leakage firing module in the network layer n corresponding to the t-th time step in [description], as shown in Figures 2 to 4 shown. The spiking neural network model includes a time-domain network structure 10 with a preset number of time steps. Each time-domain network structure 10 inputs one of the event frames. The time-domain network structure 10 includes a convolution block (Convolution), a time attention module, an integration module (Integrate), a channel attention module, a spatial attention module, and a spiking leakage firing module (Fire and Leak) connected in sequence.

[0079] Exemplarily, each time-domain network structure 10 includes an input layer → a 4*4 max pooling layer → a 3*3 convolutional layer → a time attention module → an integration module → a 3*3 convolutional layer → a channel attention module → a 2*2 average pooling layer → a 3*3 convolutional layer → a spatial attention module → a spiking leakage firing module → a 2*2 average pooling → a fully connected layer → a softmax layer. Among them, the 4*4 max pooling layer and the 3*3 convolutional layer between the input layer and the time attention module can be understood as the convolution block shown above. The 3*3 convolutional layer between the integration module and the channel attention module can be understood as being included in the channel attention module. The 2*2 average pooling layer and the 3*3 convolutional layer between the channel attention module and the spatial attention module can be understood as being included in the spatial attention module. The 2*2 average pooling, the fully connected layer, and the softmax layer after the spiking leakage firing module belong to the classification at the end of the spiking neural network model and are not shown in Figure 4 [description]. Exemplarily, for the DVS128 Gesture gesture recognition dataset, the output result of the softmax layer is the probability values of 11 types of gestures.

[0080] It can be understood that each time-domain network structure 10 inputs one event frame, or it can also be understood that each time-domain network structure 10 corresponds to an event frame of one time step. When the preset number of time steps is T, the spiking neural network model includes T time-domain network structures 10.

[0081] In addition, the spiking neural network model usually includes multiple network layers, that is, each time-domain network structure 10 includes multiple network layers, and the structure of each network layer is the same. Since the spiking neural network model in the embodiments of the present disclosure uses LIF neurons, each network layer can also be represented as LIF-SNN. For example, the spiking neural network model includes N network layers. The first layer to the last layer of network layers are respectively denoted as network layer 1, network layer 2,..., network layer n,..., network layer N. The input spikes of the n-th network layer corresponding to the t-th time step can be represented as S t,n-1 , where t is greater than or equal to 1 and less than or equal to T, and n is greater than or equal to 1 and less than or equal to N. It can be understood thatFigure 3 Only the network structure of the n-th layer of the spiking neural network model is schematically shown. In fact, the structures of all network layers are the same, that is, each network layer includes a convolutional block, a temporal attention module, an integration module, a channel attention module, a spatial attention module, and a spiking leaky integrate-and-fire module connected in sequence.

[0082] Among them, the spiking signals output by each temporal domain network structure 10, that is, the output spiking signals at each time step t, are specifically the spiking signals output by the last network layer (that is, network layer N) of each temporal domain network structure 10, and their values are 0 or 1. The final prediction result of the spiking neural network model needs to be obtained by integrating the spiking signals output at multiple time steps. For example, the mean value of the spiking signals output at T time steps can be taken as the prediction result, and this prediction result is also the target recognition result.

[0083] The spiking neural network model of the embodiments of the present disclosure combines an attention mechanism, so it can accurately focus on key spatio-temporal features, significantly improve the accuracy and real-time performance of target recognition, reduce recognition errors. When applied to target recognition tasks such as VR gesture recognition and eye movement tracking, the user's rapid gestures and eye movements can be accurately captured, thus improving the interaction experience.

[0084] Compared with the related art, the target recognition method of the embodiments of the present disclosure uses event data for target recognition, has the advantages of low power consumption, low latency, and small data volume, and accumulates the event data stream to obtain an event frame instead of the traditional image frame, which can effectively reduce the regions with mostly zeros in the event data and improve the processing efficiency. When applied to application scenarios such as VR gesture recognition or eye movement tracking, it can significantly reduce latency and improve the response speed. In addition, the target recognition method combines an attention mechanism in the spiking neural network model, so that it can accurately focus on key spatio-temporal features, significantly improve the accuracy and real-time performance of target recognition, alleviate the problem of spiking degradation during the training process, and ensure the accuracy and reliability of target recognition in complex environments such as high-speed movement, strong light, or weak light.

[0085] Based on Figures 2 to 4 For the spiking neural network model shown, the step of inputting the event frame into the pre-trained spiking neural network model for processing includes: inputting the multiple event frames into the corresponding temporal domain network structures, and each temporal domain network structure processes the input event frame and outputs the corresponding spiking signal.

[0086] For the t-th temporal domain network structure, 1 < t ≤ T, where T is the preset number of time steps, the step of this temporal domain network structure processing the input event frame and outputting the corresponding spiking signal is as Figure 5 shown, including:

[0087] Step S201, perform convolution processing on the event frame using a convolution block to obtain the current membrane potential of the event frame.

[0088] For the n-th network layer of the time-domain network structure 10 corresponding to the t-th time step, its event frame, that is, the input pulse is represented as S t,n-1 , then the convolution processing result obtained by processing the event frame S t,n-1 using a convolution block is represented as: Conv(W n , S t,n-1 ), where Conv in this formula represents the convolution operation, and W n represents the weight matrix of the convolution kernel during the convolution processing of the n-th network layer. Through convolution processing, the spatial features of the input pulse can be extracted.

[0089] Optionally, in some embodiments, batch normalization (BatchNormalization) processing and average pooling (Average Pooling) processing can be further performed on the convolution processing result, that is, after the step of performing convolution processing on the input pulse S t ,n-1 , the following steps are further included: performing batch normalization processing on the convolution-processed event frame; performing average pooling processing on the batch-normalized event frame to obtain the current membrane potential X t,n , and the current membrane potential X t,n is represented as:

[0090] X t,n = AvgPool(BN(Conv(W n , S t,n-1 ))), where c n represents the number of channels of the n-th network layer, and h n * w n represents the image resolution of the n-th network layer.

[0091] In this formula, AvgPool(), BN(), and Conv() represent average pooling, batch normalization, and convolution operations respectively. Among them, batch normalization can normalize the convolution processing result so that its mean is 0 and variance is 1, thereby accelerating training and improving the stability of the spiking neural network model; average pooling can perform downsampling on the batch-normalized processing result, reducing the size of the feature map and retaining the main features at the same time.

[0092] After being processed by step S201, for the n-th network layer, the current membrane potential X n of the spiking neural network model can be represented as X n = [X 1,n ,..., X t,n ,..., XT,n ], Where T represents the preset time step, X n That is, the current membrane potential X corresponding to each time step t t,n A collection of .

[0093] Step S202: Use the temporal attention module to process the current membrane potential of each event frame to generate a first membrane potential corresponding to each event frame.

[0094] Among them, the membrane potential of the spiking neuron is optimized through the attention mechanism, thereby adjusting the spiking response of the spiking neuron. This process can be expressed by the formula: Att =g(x)⊙x, where x represents the input of the attention module, x Att represents the output of the attention module, g(x) represents the function of generating the attention weight, and ⊙ represents element-by-element multiplication. In the disclosed embodiment, the attention module includes a time attention module, a channel attention module, and a spatial attention module.

[0095] Optionally, a temporal attention module is used to analyze each event frame S t,n-1 The current membrane potential X t,n Processing is performed to generate each event frame S t,n-1 The one-to-one corresponding first membrane potential step includes steps (11) to (14):

[0096] (11) The current membrane potential X for each event frame t,n Perform average pooling and maximum pooling.

[0097] That is, the current membrane potential X at each time step t,n Perform dual-path pooling, and the average pooling result is expressed as AvgPool(X n ), the maximum pooling result is expressed as MaxPool(X n ), where the average pooling result AvgPool(X n ) and the maximum pooling result MaxPool(X n )∈R T*1*1*1 .

[0098] (12) Nonlinear mapping is performed on the average pooling result and the maximum pooling result respectively to generate a first initial time weight corresponding to the average pooling result and a second initial time weight corresponding to the maximum pooling result.

[0099] Optionally, a fully connected layer and an activation function are used for nonlinear mapping. Exemplarily, a first fully connected layer, a second fully connected layer and an activation function ReLU are used to implement nonlinear mapping.

[0100] Among them, the first fully connected layer can map the pooled statistics to a low-dimensional space, reducing the computational complexity. The output of the average pooling result AvgPool(X n ) after being processed by the first fully connected layer is expressed as The output of the max pooling result MaxPool(X n ) after being processed by the first fully connected layer is expressed as Where represents the weight of the first fully connected layer, where r t represents the reduction factor in the time dimension, which is used to control the amount of computation.

[0101] The second fully connected layer can map the low-dimensional space features output by the first fully connected layer back to the original dimension, generating preliminary weights, that is, the first initial time weight and the second initial time weight. The first initial time weight can be expressed as The second initial time weight can be expressed as Where represents the weight of the second fully connected layer, where r t represents the reduction factor in the time dimension, which is used to control the amount of computation.

[0102] (13) Perform weighted normalization on the first initial time weight and the second initial time weight to generate time attention weights.

[0103] Among them, the time attention weight g t (X n ) is expressed as:

[0104]

[0105] By adding the first initial time weight and the second initial time weight, the robustness can be enhanced; σ represents the Sigmoid activation function, which can normalize the weights to the interval [0,1], indicating the importance of each time step.

[0106] (14) Use the time attention weight to recalibrate the current membrane potential to generate the first membrane potential.

[0107] Assume that the first membrane potential is expressed as Then:

[0108]

[0109] In the embodiments of the present disclosure, for the event data stream, through the temporal attention module, the features of key time steps can be highlighted, such as highlighting the features at the moment of motion mutation, reducing the weights of irrelevant time steps, realizing selective enhancement of temporal information, finely adjusting the input pulse signal, and improving the robustness of the spiking neural network model and the modeling ability for dynamic scenes; and through lightweight pooling operations and fully connected operations, the computational efficiency can be improved, and efficient temporal attention adjustment can be realized.

[0110] It can be understood that in Figure 3 the illustrated embodiment, the current membrane potential of each time step is input to the temporal attention module, and after being processed by the temporal attention module, the first membrane potential corresponding to each time step is output. Its essence is to perform temporal attention adjustment on the current membrane potential of each time step. In the embodiments of the present disclosure, each temporal network structure includes a temporal attention processing module mainly used to express the temporal attention adjustment of the current membrane potential of each time step.

[0111] Step S203: For any event frame, use the integration module to integrate and process the first membrane potential and the fourth membrane potential output by the previous temporal network structure to form a second membrane potential.

[0112] In the embodiments of the present disclosure, the fourth membrane potential output by the previous temporal network structure can also be understood as the membrane potential output by the temporal network structure corresponding to the previous time step of the same network layer. Assuming that the fourth membrane potential output by the nth network layer of the temporal network structure 10 corresponding to the tth time step is represented as H t,n , then the fourth membrane potential output by the previous temporal network structure is represented as H t-1,n . If the second membrane potential of this network layer is represented as U t,n , then:

[0113] wherein, the second membrane potential U t,n is obtained by adding the fourth membrane potential output at the previous moment, that is, H t-1,n to the input pulse signal at the current moment . That is, the integration module is used to calculate the membrane potential at the current moment according to the membrane potential at the previous moment and the current input, and realize the update of the membrane potential at the current moment.

[0114] Step S204: Use the channel attention module and the spatial attention module to process the second membrane potential to generate a third membrane potential.

[0115] Specifically, in implementation, first use the channel attention module to process the second membrane potential U t,n , and then use the spatial attention module to process the output of the channel attention module. The output of the spatial attention module is the third membrane potential.

[0116] Among them, for event camera data, channels may correspond to optical flow features in different directions. For example, a channel attention module can enhance channels sensitive to the motion direction and suppress the static background.

[0117] Optionally, the step of using the channel attention module and the spatial attention module to process the second membrane potential to generate the third membrane potential includes (21) to (25):

[0118] (21) Perform average pooling and max pooling on the second membrane potential of the event frame.

[0119] That is, perform dual-path pooling on the second membrane potential U at the t-th time step t,n , the average pooling result is denoted as AvgPool(U t,n ), and the max pooling result is denoted as MaxPool(U t,n ). Through the two pooling operations, the global average response and significant peaks of the channel features can be captured respectively. Among them,

[0120] (22) Perform non-linear mapping on the average pooling result and the max pooling result respectively to generate the first initial channel weight corresponding to the average pooling result and the second initial channel weight corresponding to the max pooling result.

[0121] Optionally, use a fully connected layer and an activation function to perform non-linear mapping. Exemplarily, use the third fully connected layer, the fourth fully connected layer, and the activation function ReLU to implement non-linear mapping.

[0122] Among them, the third fully connected layer can map the pooled statistics to a low-dimensional space and reduce the computational complexity. The output of the average pooling result AvgPool(U t,n ) after being processed by the third fully connected layer is denoted as The output of the max pooling result MaxPool(U t,n ) after being processed by the third fully connected layer is denoted as Among them represents the weight of the third fully connected layer, Among them r c represents the dimensionality reduction factor of the channel dimension, which is used to control the amount of computation.

[0123] The fourth fully connected layer can map the low-dimensional space features output by the third fully connected layer back to the original dimension to generate preliminary weights, that is, the first initial channel weight and the second initial channel weight. Among them, the first initial channel weight can be denoted as The second initial channel weight can be denoted as Among them represents the weight of the fourth fully connected layer, where r c represents the reduction factor of the channel dimension, which is used to control the amount of computation.

[0124] (23) Perform weighted normalization on the first initial channel weight and the second initial channel weight to generate a channel attention weight.

[0125] where the channel attention weight g c (U t,n ) is expressed as:

[0126]

[0127] where

[0128] By adding the first initial channel weight and the second initial channel weight, the robustness can be enhanced; σ represents the Sigmoid activation function, which can normalize the weight to the interval [0,1], indicating the importance of each channel.

[0129] (24) Use the channel attention weight to recalibrate the second membrane potential to generate an intermediate membrane potential.

[0130] Assume the intermediate membrane potential is expressed as then This intermediate membrane potential is the membrane potential of the nth network layer at the tth time step after channel attention refinement.

[0131] In the embodiments of the present disclosure, the channel attention module can highlight the channels important for the current task (such as gesture recognition, eye movement tracking), such as channels sensitive to motion, improve the effectiveness of spike firing, reduce the weights of background noise channels, that is, selectively enhance key channels, thereby optimizing the efficiency and effectiveness of spike firing, and improving the modeling ability of the spiking neural network model for complex temporal tasks and the robustness in dynamic scenarios; and through channel compression, the amount of computation can be reduced to meet the lightweight requirements of the spiking neural network model.

[0132] (25) Use the spatial attention module to process the intermediate membrane potential to generate the third membrane potential.

[0133] Optionally, the step of using the spatial attention module in step (25) to process the intermediate membrane potential to generate the third membrane potential includes steps (31) to (34):

[0134] (31) Perform average pooling and max pooling on the intermediate membrane potential of the event frame.

[0135] That is, the intermediate membrane potential at the t-th time step is subjected to dual-channel pooling, and the average pooling result is denoted as The max-pooling result is denoted as Through these two pooling operations, the global average response and significant peaks at each spatial position can be captured respectively, where

[0136] (32) Concatenate and perform convolution processing on the average pooling result and the max-pooling result to generate the initial spatial weights.

[0137] Among them, the concatenation process refers to concatenating the results of average pooling and max-pooling along the channel dimension, and the concatenation result is denoted as The convolution process refers to using a convolutional layer to map the concatenated features to a single-channel spatial weight map, and the initial spatial weights obtained after convolution processing are denoted as where f 7*7 represents the convolution operation, and the size of the convolutional kernel in this convolution operation is 7*7. Through concatenation and convolution processing, the dual-channel pooling information can be fused to model the dependence relationship of spatial positions.

[0138] (33) Normalize the initial spatial weights to generate spatial attention weights.

[0139] Among them, the spatial attention weights are denoted as:

[0140]

[0141] σ represents the Sigmoid activation function, which can normalize the weights to the interval [0,1]. The spatial attention weights represent the scaling factors corresponding to each spatial position at each time step, and can be used to enhance the feature responses of important regions and weaken secondary regions.

[0142] (34) Use the spatial attention weights to recalibrate the intermediate membrane potential to generate the third membrane potential.

[0143] Assume that the third membrane potential is denoted as Then This third membrane potential is the membrane potential after spatial attention refinement at the n-th network layer at the t-th time step.

[0144] In dynamic vision tasks, for event camera data, the spatial attention module can focus on event-dense regions, such as the edges of moving objects, highlighting spatial positions important for tasks (such as gesture recognition or eye tracking). Such important spatial positions are, for example, regions of moving objects, thereby enhancing the targeting and transmission efficiency of spike firing, reducing the weights of static background or noise regions, optimizing the response of the spiking neural network model to dynamic events, and through lightweight pooling and convolution operations, it can adapt to the low-power requirements of the spiking neural network model.

[0145] Step S205, use the spike leakage and emission module to process the third membrane potential to generate and output a spike signal and a fourth membrane potential. The fourth membrane potential is input to the integration module of the next time-domain network structure, where the fourth membrane potential is the reset membrane potential.

[0146] In the embodiments of the present disclosure, please refer to Figure 4 , the spiking neural network model selects LIF neurons, and the basic principle of LIF neurons can be expressed as where τ represents the time constant, u(t) and I(t) respectively represent the membrane potential of the postsynaptic neuron and the input collected from the presynaptic neuron.

[0147] The iterative process of LIF neurons mainly includes three parts, namely: membrane potential update, spike generation, and membrane potential reset.

[0148] Among them, LIF neurons mainly include an integration module and a spike leakage and emission module.

[0149] Among them, the input and output expressions of the integration module are: In this formula, the first membrane potential is the membrane potential obtained by recalibrating the current membrane potential using the time attention weight, that is, the current input, H t-1,n is the output membrane potential of the previous time step. This formula corresponds to the membrane potential update process in LIF neurons.

[0150] The input and output expressions of the spike leakage and emission module are:

[0151]

[0152] Among them, the third membrane potential is the membrane potential obtained by performing channel attention and spatial attention processing on the membrane potential updated by the integration module. Hea() represents the Heaviside step function, which outputs 1 when the input is greater than or equal to 0, and otherwise outputs 0. This function determines whether to generate a spike, u th represents the threshold of the membrane potential, that is, when the third membrane potential reaches or exceeds this threshold u thWhen a neuron generates a pulse signal S t,n , otherwise, the membrane potential remains at 0. Exemplarily, the threshold u th = 0.3. Among them, a lower threshold u th can make the neuron more easily activated, may accelerate convergence but increase noise. This process corresponds to the pulse generation process in LIF neurons. By setting the threshold u th , redundant pulses can be reduced and the data processing efficiency can be improved. Especially in the VR eye movement tracking system, the user's eye movement can be detected quickly and accurately, realizing a natural human-computer interaction experience.

[0153] In the above formula, V reset represents the reset potential, that is, when the neuron generates a pulse, the membrane potential will be reset to V reset for the next potential accumulation. β represents the decay exponent, whose value is greater than 0 and less than 1, representing the decay degree of the membrane potential. Exemplarily, β = 0.3. When the β value is small, the membrane potential decays rapidly, and the neuron is more sensitive to recent inputs. This process corresponds to the membrane potential reset process in LIF neurons.

[0154] It can be understood that the principle of the first temporal network structure (i.e., t = 1) processing the input event frame and outputting the corresponding pulse signal is the same as that of Figure 5 the t-th (1 < t ≤ T) temporal network structure shown in processing the input time frame and outputting the pulse signal. The difference is that for the first temporal network structure, the fourth membrane potential H of the previous temporal network structure input by its integration module 0,n represents the residual potential at the end of the previous calculation of the n-th network layer.

[0155] Optionally, before the step of accumulating the event data stream according to the preset number of time steps to obtain multiple event frames, it further includes: downsampling the event data stream to reduce the resolution of the event data stream to the target resolution.

[0156] In the embodiments of the present disclosure, when downsampling the event data stream, sufficient temporal information is retained for the prediction basis of the spiking neural network model, and it is ensured that the spiking neural network model can operate efficiently under hardware constraints. Exemplarily, the resolution of the event data stream to be recognized obtained is 240*180, and this resolution represents the size of the spatial distribution grid of the events, that is, the value range of the event coordinates (x, y) is x ∈ [0, 239] (a total of 240 columns), y ∈ [0, 179] (a total of 180 rows). The resolution can be reduced to the target resolution, for example, 80*60, by using downsampling. The downsampling method may include: spatial merging: dividing the original resolution into coarser grids, for example, every 3*3 original grids are merged into 1 new grid; event aggregation: the events within the same new grid are accumulated according to the polarity (ON or OFF), or the pulse closest in time is taken. Through downsampling, the resolution requirements can be dynamically adjusted according to memory and computing requirements, achieving a balance between efficiency and feature retention in combination with the scenario. This optimization is particularly suitable for embedded devices in VR systems, ensuring efficient real-time interaction can still be achieved under limited hardware resources.

[0157] Optionally, when training the initial spiking neural network model, a spatio-temporal backpropagation function and a loss function are used for convergence.

[0158] Exemplarily, the steps of training the initial spiking neural network model with the training data to obtain the pre-trained spiking neural network model include:

[0159] (41) Obtain the pulse signals output by each temporal network structure in the current training process, and calculate the mean value of each pulse signal.

[0160] (42) Calculate the mean square error between the mean value and the true value, where the true value is also the label value of the training data.

[0161] (43) End the training process of the initial spiking neural network model when the mean square error reaches the preset error threshold or the number of traversals of the training data reaches the traversal number threshold.

[0162] Exemplarily, as Figure 2 shown, when the spiking neural network model includes T time steps, assume that the output of the t-th time step is represented as O t,N , and this output O t,N is also the pulse signal S t,N output by the N-th layer at the t-th time step. Assume that the true value is represented as Y label , then the loss function can be expressed as:

[0163]

[0164] The training ends when the loss function L reaches a preset error threshold, or when the number of traversals of the training data reaches the traversal number threshold. For example, when the traversal number threshold Epoch = 100, the training ends. In addition, during the training process, the learning rate can also be set to 1e-4, and the threshold u th = 0.3, and the decay exponent β = 0.3.

[0165] Optionally, before the step of inputting the event frame into the pre-trained spiking neural network model, it further includes: obtaining training data, and training the initial spiking neural network model with the training data to obtain the pre-trained spiking neural network model. The training data includes first training data and / or second training data. The first training data is the event data stream output by the event camera, and the second training data is the event data stream obtained by simulating the image data.

[0166] Among them, the first training data is a real event data stream, which can be directly collected by an event camera.

[0167] Exemplarily, the first training data is a gesture recognition dataset, such as the DVS128 Gesture gesture recognition dataset. This dataset is captured by a DVS camera and contains 11 different gesture categories. 29 different individuals demonstrate samples under 3 different lighting conditions. The dynamic information of the gesture is directly captured by the DVS camera, and then combined into a gesture recognition dataset after annotation. This dataset has a total of 1464 sample data, including 1176 training sets and 288 test sets. On the one hand, this gesture recognition dataset can be used to verify the performance of the spiking neural network model of the present disclosure embodiment on real event data. On the other hand, it can also test the model's ability to capture the temporal features of gestures.

[0168] Among them, the second training data is the event data stream obtained by simulating the image data.

[0169] Exemplarily, the second training data is an eye movement tracking data set. For example, the second training data is an event data stream obtained by simulating the LPW (Labelled Pupils in the Wild) data set using the simulation tool V2E. Among them, the LPW data set is a novel data set for studying pupil detection in an unconstrained environment. It contains 66 high-quality, high-frame-rate eye region videos provided by 22 participants, generating a total of 130,856 images with a resolution of 640*480. This data set is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. The LPW data set can be used to verify the performance of the spiking neural network model of the embodiments of the present disclosure on simulated event data, as well as the adaptability of the model in the VR scenario, and test its robustness to complex lighting and dynamic environments.

[0170] The embodiments of the present disclosure train the spiking neural network model using different tasks and different types of event data, which can ensure the wide applicability of the spiking neural network model for different tasks, as well as real event data and simulated event data.

[0171] Please refer to Table 1. Table 1 is a comparison table of the accuracy of target recognition results obtained by testing the SNN model and the spiking neural network model of the present disclosure using the test set in the training data.

[0172] Table 1 Comparison table of the SNN model and the SNN model combined with the attention mechanism

[0173] model dt*T accuracy SNN 15×60 90.63% SNN + attention mechanism 15×60 96.53%

[0174] As can be seen from Table 1, compared with the conventional SNN model, the SNN model combined with the attention mechanism provided by the embodiments of the present disclosure has a significant improvement in the target recognition accuracy, which has increased from 90.63% to 96.53%. That is, the spiking neural network model provided by the embodiments of the present disclosure can effectively operate in a high-dynamic, low-latency VR environment, providing strong technical support for natural interaction and immersive experience in virtual reality.

[0175] Based on the same inventive concept, the second aspect of the present disclosure provides a terminal, including a memory, a processor, and a program stored on the memory and executable on the processor. The processor is configured to read the program in the memory to implement the steps in the target recognition method as described above.

[0176] Exemplarily, the terminal can be a display device such as a mobile phone, a tablet, a computer, or an all-in-one machine.

[0177] Based on the same inventive concept, a third aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned target recognition method are implemented. In a specific implementation process, the computer storage medium may include: Universal Serial Bus Flash Drive (USB for short), mobile hard disk, Read Only Memory (ROM for short), Random Access Memory (RAM for short), magnetic disk or optical disc and other storage media that can store program codes.

[0178] Based on the same inventive concept, a fourth aspect of the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned target recognition method are implemented. Since the principle of the above computer program for solving problems is similar to the principle of the target recognition method, the implementation of the above computer program can refer to the implementation of the target recognition method, and the repeated parts will not be elaborated.

[0179] The computer program product may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or flash memory), an optical fiber, a portable compact disk read only memory (CDROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

Claims

1. A target recognition method, characterized in that, Including the following steps: Obtain the event data stream to be recognized; Perform cumulative processing on the event data stream according to a preset number of time steps to obtain multiple event frames; Input the event frames into a pre-trained spiking neural network model for processing, and determine the target recognition result according to the spike signals output by the spiking neural network model. The spiking neural network model is a leaky integrate-and-fire neuron model combined with an attention mechanism.

2. The object recognition method according to claim 1, wherein The attention mechanism includes at least one of a temporal attention mechanism, a spatial attention mechanism, and a channel attention mechanism. The spiking neural network model includes at least one of a temporal attention module, a spatial attention module, and a channel attention module.

3. The object recognition method according to claim 2, characterized in that, The spiking neural network model includes a time-domain network structure with a preset number of time steps. Each time-domain network structure inputs one of the event frames. The time-domain network structure includes a convolutional block, a temporal attention module, an integration module, a spatial attention module, a channel attention module, and a spike leakage firing module connected in sequence. The step of inputting the event frames into the pre-trained spiking neural network model for processing includes: inputting the multiple event frames into the corresponding time-domain network structures. Each time-domain network structure processes the input event frame and outputs the corresponding spike signal. For the t-th time-domain network structure, 1 < t ≤ T, where T is the preset number of time steps, the step of this time-domain network structure processing the input event frame and outputting the corresponding spike signal includes: Performing convolutional processing on the input event frame using the convolutional block to obtain the current membrane potential of the event frame; Processing the current membrane potential of each event frame using the temporal attention module to generate a first membrane potential corresponding to each event frame one by one; Integrating and processing the first membrane potential and the fourth membrane potential output by the previous time-domain network structure using the integration module to form a second membrane potential; Processing the second membrane potential using the channel attention module and the spatial attention module to generate a third membrane potential; Processing the third membrane potential using the spike leakage firing module to generate and output the spike signal and the fourth membrane potential. The fourth membrane potential is input to the integration module of the next time-domain network structure.

4. The object recognition method according to claim 3, wherein After the step of performing convolutional processing on the event frame using the convolutional block, it further includes: Performing batch normalization processing on the event frame after convolutional processing; Performing average pooling processing on the event frame after batch normalization processing to obtain the current membrane potential.

5. The target recognition method according to claim 3, wherein The step of processing the current membrane potential of each event frame using the temporal attention module to generate a first membrane potential corresponding to each event frame one by one includes: Performing average pooling and max pooling on the current membrane potential of each event frame; Performing non-linear mapping on the average pooling result and the max pooling result respectively to generate a first initial time weight corresponding to the average pooling result and a second initial time weight corresponding to the max pooling result; Performing weighted normalization processing on the first initial time weight and the second initial time weight to generate a temporal attention weight; Re-calibrating the current membrane potential using the temporal attention weight to generate the first membrane potential.

6. The target recognition method according to claim 3, wherein The steps of processing the second membrane potential by using the channel attention module and the spatial attention module to generate the third membrane potential include: Performing average pooling and max pooling on the second membrane potential of the event frame; Performing non-linear mapping on the average pooling result and the max pooling result respectively to generate a first initial channel weight corresponding to the average pooling result and a second initial channel weight corresponding to the max pooling result; Performing weighted normalization processing on the first initial channel weight and the second initial channel weight to generate a channel attention weight; Using the channel attention weight to recalibrate the second membrane potential to generate an intermediate membrane potential; Using the spatial attention module to process the intermediate membrane potential to generate the third membrane potential.

7. The object recognition method according to claim 6, characterized in that The steps of using the spatial attention module to process the intermediate membrane potential to generate the third membrane potential include: Performing average pooling and max pooling on the intermediate membrane potential of the event frame; Performing concatenation and convolution processing on the average pooling result and the max pooling result to generate an initial spatial weight; Performing normalization processing on the initial spatial weight to generate a spatial attention weight; Using the spatial attention weight to recalibrate the intermediate membrane potential to generate the third membrane potential.

8. The object recognition method according to claim 1, characterized in that Before the step of accumulating the event data stream according to the preset number of time steps to obtain multiple event frames, it further includes: Downsampling the event data stream to reduce the resolution of the event data stream to the target resolution.

9. The target recognition method according to claim 3, wherein Before the step of inputting the event frame into the pre-trained spiking neural network model, it further includes: Obtaining training data, and using the training data to train an initial spiking neural network model to obtain the pre-trained spiking neural network model. The training data includes first training data and / or second training data. The first training data is the event data stream output by the event camera, and the second training data is the event data stream obtained by simulating the image data.

10. The target recognition method according to claim 9, wherein, The steps of using the training data to train the initial spiking neural network model to obtain the pre-trained spiking neural network model include: Obtaining the spike signals output by each time-domain network structure in the current training process, and calculating the mean value of each spike signal; Calculating the mean square error between the mean value and the true value; Ending the training process of the initial spiking neural network model when the mean square error reaches the preset error threshold or the number of traversals of the training data reaches the traversal number threshold.

11. A terminal, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the target recognition method according to any one of claims 1-10.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the target recognition method according to any one of claims 1-10.

13. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the target recognition method according to any one of claims 1-10.