Event video reconstruction method based on spiking neural network
By designing an event video reconstruction method based on heterogeneous pulse neural network, the problem of failure to fully consider the asynchronous characteristics of event streams in the prior art is solved, and the effect of reducing video artifacts and flickering is achieved, and time consistency and image quality are improved.
Patent Information
- Application Number
- CN202510290957.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
AI Technical Summary
The existing video reconstruction methods fail to fully consider the asynchronous characteristics of the event stream, resulting in the reconstructed video such as screen flickering, artifacts and contrast reduction.
A method of event video reconstruction based on heterogeneous pulse neural network was designed. By constructing a heterogeneous pulse neural network model, the membrane attenuation coefficient is dynamically updated using the pulsed neuron layer to perceive the spatiotemporal distribution differences of event streams, thereby reducing the artifacts and flickering of the reconstructed video.
Effectively reduce video artifacts and flickering, improve the time consistency and image quality of the reconstructed video, and enhance its usability in practical applications.
Smart Images

Figure BDA0005308612660000081
Abstract
Description
Technical Field
[0001] The present invention relates to a video reconstruction method, in particular to an event video reconstruction method based on a spiking neural network. Background Art
[0002] Event cameras output event streams with high temporal resolution, high dynamic range, and low latency through asynchronous sensors, showing great application potential in fields such as privacy protection and industrial inspection. Event cameras capture the brightness changes of pixels instead of traditional images, so the output data stream is unstructured and most existing computer vision algorithms cannot directly process it. To solve this problem, researchers have designed video reconstruction algorithms for event cameras, whose goal is to reconstruct the unstructured event stream into a video that is easy to understand and analyze for downstream vision tasks. Currently, most video reconstruction methods rely on traditional deep learning models to achieve high-quality reconstruction results by consuming high computing resources. In addition, some methods have explored the application of spiking neural networks, attempting to reduce the computational cost using spike coding. However, these methods do not fully consider the asynchronous characteristics of the event stream - that is, in different scenarios, there are significant differences in the spatio-temporal distribution of the event stream. This neglect of the asynchronous characteristics leads to problems such as frame flickering, artifacts, and contrast degradation in the reconstructed video, affecting its usability in practical applications. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide an event video reconstruction method based on a spiking neural network that can significantly reduce video artifacts and flickering.
[0004] The technical solution adopted by the present invention to solve the above technical problems is: an event video reconstruction method based on a spiking neural network, comprising the following steps:
[0005] Step 1): Obtain a training set consisting of N event streams from the original video database, divide the event streams in the training set at a fixed time interval to obtain multiple sub-event streams, and then convert each sub-event stream into a corresponding 3D voxel grid feature, and each 3D voxel grid feature corresponds to a grayscale image and an optical flow map;
[0006] Step 2): Construct a heterogeneous spiking neural network model to be trained. The heterogeneous spiking neural network model to be trained includes an input layer, an encoder, a residual layer, a decoder, and an output layer. Send the 3D voxel grid features into the input layer in sequence according to the time steps. The input layer consists of a first convolutional module, and the first convolutional module updates the features of the 3D voxel grid features to obtain initial features;
[0007] Step 3): Input the initial features into the encoder. The encoder consists of two encoding layers. The first encoding layer is composed of a first spiking neuron layer and a second convolutional module. The first spiking neuron layer perceives and updates the spatio-temporal information of the initial features to obtain the updated initial features, and then the second convolutional module updates the features of the updated initial features to obtain the first encoded features. The second encoding layer is composed of a second spiking neuron layer and a third convolutional module. The second spiking neuron layer perceives and updates the spatio-temporal information of the first encoded features to obtain the updated first encoded features, and then the third convolutional module updates the features of the updated first encoded features to obtain the second encoded features;
[0008] Step 4): Input the second encoded features into the residual layer. The residual layer is composed of a fourth convolutional module and a third spiking neuron layer. The third spiking neuron layer perceives and updates the spatio-temporal information of the second encoded features to obtain the updated second encoded features, and then the fourth convolutional module updates the features of the updated second encoded features to obtain the primary features. Then, add the primary features and the second encoded features to obtain the residual features;
[0009] Step 5): Input the residual features into the decoder. The decoder consists of two decoding layers. The first decoding layer is composed of a first upsampling layer, a fourth spiking neuron layer, and a fifth convolutional module. The first upsampling module upsamples the residual features to increase the spatial dimension and decrease the channel dimension of the residual features, obtaining the first upsampled features. The fourth spiking neuron layer perceives and updates the spatio-temporal information of the first upsampled features to obtain the updated first upsampled features, and then the fifth convolutional module updates the features of the updated first upsampled features to obtain the first decoded features. The second decoding layer is composed of a second upsampling layer, a fifth spiking neuron layer, and a sixth convolutional module. The second upsampling layer upsamples the added features obtained by adding the first encoded features and the first decoded features to increase the spatial dimension and decrease the channel dimension of the added features, obtaining the second upsampled features. The fifth spiking neuron layer perceives and updates the spatio-temporal information of the second upsampled features to obtain the updated second upsampled features, and then the sixth convolutional module updates the features of the updated second upsampled features to obtain the second decoded features;
[0010] Step 6): Input the second decoded features into the output layer. The output layer is composed of a sixth spiking neuron layer and a seventh convolutional module. The sixth spiking neuron layer perceives and updates the spatio-temporal information of the result of adding the second decoded features and the initial features to obtain the updated second decoded features, and then the seventh convolutional module updates the features of the updated second decoded features to obtain the reconstructed features;
[0011] Step 7): Define the total loss function of the heterogeneous spiking neural network model to be trained. The total loss function is obtained by adding a first loss function for improving the image quality of the reconstructed features and a second loss function for improving the temporal consistency of the reconstructed features. Then, obtain the values of the first loss function and the second loss function through the 3D voxel grid features and the corresponding grayscale images, optical flow maps, and reconstructed features, and add them to obtain the value of the total loss function;
[0012] Step 8): Set the maximum number of iterations. According to the total loss function, use the backpropagation algorithm to iteratively optimize the heterogeneous spiking neural network model to be trained until the iterative process stops when the set maximum number of iterations is reached, and obtain the trained heterogeneous spiking neural network model;
[0013] Step 9): Divide the target event video into multiple sub-event videos at fixed time intervals, then convert each sub-event video into the corresponding 3D voxel grid features, and finally sequentially input the 3D voxel grid features corresponding to each sub-event video into the trained heterogeneous spiking neural network model to obtain the corresponding reconstructed features, and finally combine them to obtain the reconstructed video, completing the reconstruction process of the target event video.
[0014] Specifically, in step 3), use the initial feature input at the current time step as the input feature of the first spiking neuron layer. The specific process of the first spiking neuron layer for perceiving and updating the spatio-temporal information of the input feature is as follows:
[0015] Step 3-1: The first spiking neuron layer includes a membrane decay coefficient update module, a membrane potential update module, a spike emission module with a preset initial spike state, and a membrane potential reset module. The membrane decay coefficient update module includes a first spiking convolution module, a second spiking convolution module, and a first normalization layer. After the initial feature at the first time step is input into the membrane decay coefficient update module, the first spiking convolution module updates the preset initial spike state in the spike emission module to obtain a spike update feature, the second spiking convolution module updates the initial feature at the first time step to obtain an input update feature, and the first normalization layer adds the spike update feature and the input update feature and normalizes them to obtain the membrane decay coefficient at the first time step;
[0016] Step 3-2: The membrane potential update module presets an initial membrane potential. The membrane potential update module uses the membrane decay coefficient at the first time step to perform a decay operation on the preset initial membrane potential to obtain a decay membrane potential, and then adds the decay membrane potential and the initial feature at the first time step to obtain the membrane potential at the first time step;
[0017] Step 3-3: Update the pulse state at the first time step through the pulse emission module. The pulse emission module uses the step function to convert the membrane potential at the first time step into the pulse state S1[t] at the first time step, S1[t]=Θ(V1[t] - V th ), where Θ(·) represents the step function, V1[t] represents the membrane potential at the first time step, and V th represents the preset pulse emission threshold, and take the pulse state S[t] at the first time step as the updated initial feature at the first time step;
[0018] Step 3-4: If V1[t] - V th >0, then the membrane potential reset module subtracts the preset pulse emission threshold from the membrane potential at the first time step to update the membrane potential at the first time step to obtain the final membrane potential at the first time step; if V1[t] - V th ≤0, then the membrane potential at the first time step remains unchanged and serves as the final membrane potential at the first time step;
[0019] Step 3-5: After inputting the initial feature of the current time step starting from the second time step into the membrane decay coefficient update module, the first pulse convolution module updates the pulse state of the previous time step to obtain the pulse update feature, and the second pulse convolution module updates the initial feature of the previous time step to obtain the input update feature. The first normalization layer adds and normalizes the pulse update feature and the input update feature to obtain the membrane decay coefficient at the current time step;
[0020] Step 3-6: The membrane potential update module uses the membrane decay coefficient of the previous time step to perform a decay operation on the final membrane potential of the previous time step to obtain the decay membrane potential, and then adds the decay membrane potential and the initial feature of the previous time step to obtain the membrane potential at the current time step;
[0021] Step 3-7: Update the pulse state at the current time step through the pulse emission module. The pulse emission module uses the step function to convert the membrane potential at the current time step into the pulse state S[t] at the current time step, S[t]=Θ(V[t] - V th ), where Θ(·) represents the step function, V[t] represents the membrane potential at the current time step, and take the pulse state S[t] at the current time step as the updated initial feature at the current time step;
[0022] Step 3-8: If V[t] - V th >0, then the membrane potential reset module subtracts the preset pulse emission threshold from the membrane potential at the current time step to update the membrane potential at the current time step to obtain the final membrane potential at the current time step; if V[t] - V th ≤0, then the membrane potential at the current time step remains unchanged and serves as the final membrane potential at the current time step.
[0023] Compared with the prior art, the advantages of the present invention are that a spiking neural network with heterogeneity is specifically designed to efficiently perceive the spatio-temporal distribution differences of the event stream, thereby reducing the artifacts and flicker of the reconstructed video; the core component of the spiking neural network with heterogeneity is the spiking neuron layer with heterogeneity, and this layer dynamically updates the membrane decay coefficient of the spiking neurons through the membrane decay coefficient update module; specifically, at each time step, the membrane decay coefficient update module uses the spike state of the previous time step and the neuron input of the current time step to update the membrane decay coefficient of the current time step. In this way, all the spiking neurons in the layer have heterogeneous membrane decay coefficients, that is, the membrane decay coefficients of all the spiking neurons are different from each other, so that the spiking neurons have the ability to perceive the spatio-temporal distribution differences of the event stream in different scenarios, and improve the temporal consistency of the reconstructed video. Specific embodiments
[0024] The present invention will be further described in detail below in conjunction with embodiments.
[0025] An event video reconstruction method based on a spiking neural network, characterized by comprising the following steps:
[0026] Step 1): Obtain a training set composed of N event streams from the original video database, divide the event streams in the training set at a fixed time interval to obtain multiple sub-event streams, and then convert each sub-event stream into a corresponding 3D voxel grid feature, and each 3D voxel grid feature corresponds to a grayscale image and an optical flow map.
[0027] Step 2): Construct a heterogeneous spiking neural network model to be trained. The heterogeneous spiking neural network model to be trained includes an input layer, an encoder, a residual layer, a decoder, and an output layer. The 3D voxel grid features are sequentially sent into the input layer according to the order of time steps. The input layer is composed of a first convolution module, and the first convolution module updates the features of the 3D voxel grid features to obtain initial features. During this process, the spatial resolution of the 3D voxel grid features remains unchanged, and the number of channels increases.
[0028] Step 3): Input the initial features into the encoder. The encoder consists of two encoding layers. The first encoding layer is composed of a first spiking neuron layer and a second convolutional module. The first spiking neuron layer perceives and updates the spatio-temporal information of the initial features to obtain the updated initial features, and then the second convolutional module updates the features of the updated initial features to obtain the first encoded features. The second encoding layer is composed of a second spiking neuron layer and a third convolutional module. The second spiking neuron layer perceives and updates the spatio-temporal information of the first encoded features to obtain the updated first encoded features, and then the third convolutional module updates the features of the updated first encoded features to obtain the second encoded features. During the encoding process of each encoding layer, the spatial resolution of the encoded features decreases while the number of channels increases.
[0029] Take the initial features input at the current time step as the input features of the first spiking neuron layer. The specific process of the first spiking neuron layer perceiving and updating the spatio-temporal information of the input features is as follows:
[0030] Step 3-1: The first spiking neuron layer includes a membrane decay coefficient update module, a membrane potential update module, a spike emission module with a preset initial spike state, and a membrane potential reset module. The membrane decay coefficient update module includes a first spiking convolution module, a second spiking convolution module, and a first normalization layer. After the initial features at the first time step are input into the membrane decay coefficient update module, the first spiking convolution module updates the preset initial spike state in the spike emission module to obtain spike update features, the second spiking convolution module updates the initial features at the first time step to obtain input update features, and the first normalization layer adds the spike update features and the input update features and normalizes them to obtain the membrane decay coefficient at the first time step.
[0031] Step 3-2: The membrane potential update module presets an initial membrane potential. The membrane potential update module uses the membrane decay coefficient at the first time step to decay the preset initial membrane potential to obtain a decayed membrane potential, and then adds the decayed membrane potential and the initial features at the first time step to obtain the membrane potential at the first time step.
[0032] Step 3-3: Update the spike state at the first time step through the spike emission module. The spike emission module uses the step function to convert the membrane potential at the first time step into the spike state S1[t] at the first time step, S1[t] = Θ(V1[t] - V th ), Θ(·) represents the step function, V1[t] represents the membrane potential at the first time step, V th represents the preset spike emission threshold, and takes the spike state S[t] at the first time step as the updated initial features at the first time step.
[0033] Step 3-4: If V1[t] - V thIf it is > 0, the membrane potential reset module subtracts the preset pulse emission threshold from the membrane potential of the first time step to update the membrane potential of the first time step to obtain the final membrane potential of the first time step; if V1[t] - V th ≤ 0, the membrane potential of the first time step remains unchanged and serves as the final membrane potential of the first time step.
[0034] Step 3-5: After inputting the initial feature of the current time step starting from the second time step into the membrane decay coefficient update module, the first pulse convolution module updates the pulse state of the previous time step to obtain a pulse update feature, the second pulse convolution module updates the initial feature of the previous time step to obtain an input update feature, and the first normalization layer adds and normalizes the pulse update feature and the input update feature to obtain the membrane decay coefficient of the current time step.
[0035] Step 3-6: The membrane potential update module uses the membrane decay coefficient of the previous time step to perform a decay operation on the final membrane potential of the previous time step to obtain a decay membrane potential, and then adds the decay membrane potential and the initial feature of the previous time step to obtain the membrane potential of the current time step.
[0036] Step 3-7: Update the pulse state of the current time step through the pulse emission module. The pulse emission module uses the step function to convert the membrane potential of the current time step into the pulse state S[t] of the current time step, S[t] = Θ(V[t] - V th ), Θ(·) represents the step function, V[t] represents the membrane potential of the current time step, and the pulse state S[t] of the current time step is used as the updated initial feature of the current time step.
[0037] Step 3-8: If V[t] - V th > 0, the membrane potential reset module subtracts the preset pulse emission threshold from the membrane potential of the current time step to update the membrane potential of the current time step to obtain the final membrane potential of the current time step; if V[t] - V th ≤ 0, the membrane potential of the current time step remains unchanged and serves as the final membrane potential of the current time step.
[0038] Step 4): Input the second encoded feature into the residual layer. The residual layer consists of a fourth convolution module and a third pulse neuron layer. The third pulse neuron layer perceives and updates the spatio-temporal information of the second encoded feature to obtain an updated second encoded feature, and then the fourth convolution module updates the feature of the updated second encoded feature to obtain a primary feature. Then, add the primary feature and the second encoded feature to obtain a residual feature; the residual layer uses the residual connection to perform a cascaded update operation on the second encoded feature, which can improve the information utilization rate, alleviate the problems of gradient disappearance and gradient explosion during model training, and improve the reconstruction quality.
[0039] Step 5): Input the residual features into the decoder. The decoder consists of two decoding layers. The first decoding layer is composed of a first upsampling layer, a fourth spiking neuron layer, and a fifth convolutional module. The first upsampling module upsamples the residual features, increasing the spatial dimension and decreasing the channel dimension of the residual features to obtain the first upsampled features. The fourth spiking neuron layer perceives and updates the spatio-temporal information of the first upsampled features to obtain the updated first upsampled features, and then the fifth convolutional module updates the features of the updated first upsampled features to obtain the first decoded features. The second decoding layer is composed of a second upsampling layer, a fifth spiking neuron layer, and a sixth convolutional module. The second upsampling layer upsamples the added features obtained by adding the first encoded features and the first decoded features, increasing the spatial dimension and decreasing the channel dimension of the added features to obtain the second upsampled features. The fifth spiking neuron layer perceives and updates the spatio-temporal information of the second upsampled features to obtain the updated second upsampled features, and then the sixth convolutional module updates the features of the updated second upsampled features to obtain the second decoded features.
[0040] Step 6): Input the second decoded features into the output layer. The output layer is composed of a sixth spiking neuron layer and a seventh convolutional module. The sixth spiking neuron layer perceives and updates the spatio-temporal information of the result of adding the second decoded features and the initial features to obtain the updated second decoded features, and then the seventh convolutional module updates the features of the updated second decoded features to obtain the reconstructed features.
[0041] Step 7): Define the total loss function of the heterogeneous spiking neural network model to be trained. The total loss function is obtained by adding a first loss function for improving the image quality of the reconstructed features and a second loss function for improving the temporal consistency of the reconstructed features. Then, obtain the values of the first loss function and the second loss function through the 3D voxel grid features, the grayscale images, the optical flow maps, and the reconstructed features corresponding to the 3D voxel grid features, and add them to obtain the value of the total loss function. Obtain the optical flow mapping error of the reconstructed features through the optical flow information between two consecutive images. Among them, through the perceptual reconstruction loss, the reconstructed features are made more semantically consistent with the real images, and through the temporal consistency loss, the temporal consistency and content coherence of the reconstructed video are improved, reducing the flicker problem.
[0042] Step 8): Set the maximum number of iterations. According to the total loss function, iteratively optimize the heterogeneous spiking neural network model to be trained through the backpropagation algorithm until the iterative process stops when the set maximum number of iterations is reached, and obtain the trained heterogeneous spiking neural network model.
[0043] Step 9): Divide the target event video into multiple sub-event videos at fixed time intervals, then convert each sub-event video into its corresponding 3D voxel grid feature, and finally send the 3D voxel grid features corresponding to each sub-event video into the trained heterogeneous spiking neural network model in sequence to obtain the corresponding reconstruction features, and finally combine them to obtain the reconstructed video, completing the reconstruction process of the target event video.
[0044] In the above embodiments, the working principles of the remaining spiking neuron layers are the same as those of the first spiking neuron layer, and the only difference lies in their different inputs and outputs.
[0045] The following uses the specific comparison examples shown in Table 1 to compare the reconstruction performance of the event video reconstruction method of this embodiment, where the event video reconstruction method of this embodiment is abbreviated as this method.
[0046] Table 1
[0047]
[0048] Table 1 shows the results of the event video reconstruction method of this embodiment and other methods in terms of MSE, SSIM, and LPIPS metrics on three datasets ECD, MVSEC, and HQF.
[0049] As can be seen from Table 1, compared with the general method, when using the video reconstruction method of this embodiment, the MSE metrics of the video reconstruction method of this embodiment on the three datasets decreased by 7.5%, 3.3%, and 4.4% respectively, the SSIM metrics increased by 2.0%, 5.7%, and 3.2% respectively, and the LPIPS metrics increased by 13.3%, 3.0%, and 3.4% respectively.
[0050] In addition, on the calibration samples in the ECD dataset, the results of the reconstructed videos of the event video reconstruction method of this embodiment and other methods show that, compared with the general method, the results of the video reconstruction method of this embodiment have better temporal consistency, clearer texture details, and fewer artifacts.
Claims
1. An event video reconstruction method based on a pulse neural network, characterized in that The following steps are involved: Step 1): Obtain N event streams from the original video database to form a training set, divide the event stream in the training set according to a fixed time interval to obtain multiple sub-event streams, and then convert each sub-event stream into a corresponding 3D voxel grid feature, each 3D voxel grid feature corresponds to a grayscale image and an optical flow image; Step 2): construct a heterogeneous spiking neural network model to be trained, the heterogeneous spiking neural network model to be trained includes an input layer, an encoder, a residual layer, a decoder and an output layer, the 3D voxel grid features are sequentially sent to the input layer in the order of time steps, the input layer is composed of a first convolution module, and the first convolution module updates the 3D voxel grid features to obtain initial features; Step 3): Input the initial features into the encoder, the encoder includes two encoding layers, the first encoding layer is composed of a first pulse neuron layer and a second convolution module, the first pulse neuron layer performs spatiotemporal information perception on the initial features and updates them to obtain updated initial features, and then the second convolution module performs feature update on the updated initial features to obtain first encoding features, the second encoding layer is composed of a second pulse neuron layer and a third convolution module, the second pulse neuron layer performs spatiotemporal information perception on the first encoding features and updates them to obtain updated first encoding features, and then the third convolution module performs feature update on the updated first encoding features to obtain second encoding features; Step 4): Input the second coding feature into the residual layer, which is composed of a fourth convolution module and a third pulse neuron layer. The third pulse neuron layer perceives the spatiotemporal information of the second coding feature and updates it to obtain an updated second coding feature. The fourth convolution module then updates the updated second coding feature to obtain a primary feature, and then adds the primary feature to the second coding feature to obtain a residual feature. Step 5): input the residual feature into the decoder, the decoder includes two decoding layers, the first decoding layer is composed of a first upsampling layer, a fourth pulse neuron layer and a fifth convolution module, the first upsampling module performs upsampling processing on the residual feature to increase the spatial dimension of the residual feature and reduce the channel dimension to obtain the first upsampling feature, the fourth pulse neuron layer performs spatiotemporal information perception on the first upsampling feature and updates it to obtain the updated first upsampling feature, and then the fifth convolution module performs feature update on the updated first upsampling feature to obtain the first decoding feature; The second decoding layer is composed of a second upsampling layer, a fifth pulse neuron layer and a sixth convolution module. The second upsampling layer performs upsampling processing on the added feature obtained by adding the first encoding feature and the first decoding feature, so that the spatial dimension of the added feature increases and the channel dimension decreases to obtain a second upsampling feature. The fifth pulse neuron layer performs spatiotemporal information perception on the second upsampling feature and updates it to obtain an updated second upsampling feature. The sixth convolution module then performs feature update on the updated second upsampling feature to obtain a second decoding feature. Step 6): the second decoded feature is input into the output layer, the output layer is composed of a sixth pulse neuron layer and a seventh convolution module, the sixth pulse neuron layer performs spatiotemporal information perception on the result of adding the second decoded feature and the initial feature and updates it to obtain an updated second decoded feature, and then the seventh convolution module performs feature update on the updated second decoded feature to obtain a reconstructed feature; Step 7): define a total loss function of the heterogeneous spiking neural network model to be trained, the total loss function is obtained by adding a first loss function for improving the image quality of the reconstructed feature and a second loss function for improving the temporal consistency of the reconstructed feature, and then obtain the values of the first loss function and the second loss function through the 3D voxel grid feature and the grayscale image, optical flow map and reconstruction feature corresponding to the 3D voxel grid feature, and add them to obtain the value of the total loss function; Step 8): Set the maximum number of iterations, and iteratively optimize the heterogeneous spiking neural network model to be trained through the back propagation algorithm according to the total loss function, until the iterative process is stopped when the set maximum number of iterations is reached, and the trained heterogeneous spiking neural network model is obtained; Step 9): Divide the target event video into multiple sub-event videos according to fixed time intervals, then convert each sub-event video into a corresponding 3D voxel grid feature, and finally send the 3D voxel grid features corresponding to each sub-event video into the trained heterogeneous pulse neural network model in turn to obtain the corresponding reconstruction features, and finally combine them to obtain the reconstructed video, completing the reconstruction process of the target event video.
2. The event video reconstruction method based on pulse neural network according to claim 1 is characterized in that In the step 3), the initial features input at the current time step are used as the input features of the first pulse neuron layer. The specific process of the first pulse neuron layer sensing and updating the input features with spatiotemporal information is as follows: Step 3-1: The first pulse neuron layer includes a membrane attenuation coefficient update module, a membrane potential update module, a pulse emission module preset with an initial pulse state, and a membrane potential reset module. The membrane attenuation coefficient update module includes a first pulse convolution module, a second pulse convolution module, and a first normalization layer. After the initial feature of the first time step is input into the membrane attenuation coefficient update module, the first pulse convolution module updates the initial pulse state preset in the pulse emission module to obtain a pulse update feature. The second pulse convolution module updates the initial feature of the first time step to obtain an input update feature. The first normalization layer adds the pulse update feature and the input update feature and normalizes them to obtain the membrane attenuation coefficient of the first time step. Step 3-2: The membrane potential update module is preset with an initial membrane potential. The membrane potential update module uses the membrane attenuation coefficient of the first time step to attenuate the preset initial membrane potential to obtain an attenuated membrane potential, and then adds the attenuated membrane potential to the initial feature of the first time step to obtain the membrane potential of the first time step; Step 3-3: Update the pulse state of the first time step through the pulse emission module. The pulse emission module uses a step function to convert the membrane potential of the first time step into the pulse state S1[t] of the first time step, S1[t] = Θ(V1[t]-V th ), Θ(·) represents the step function, V1[t] represents the membrane potential at the first time step, V th represents the preset pulse emission threshold, and the pulse state S[t] of the first time step is used as the updated initial feature of the first time step; Step 3-4: If V1[t]-V th > 0, the membrane potential reset module subtracts the preset pulse emission threshold from the membrane potential of the first time step, and updates the membrane potential of the first time step to obtain the final membrane potential of the first time step; if V1[t]-V th ≤0, the membrane potential of the first time step remains unchanged and serves as the final membrane potential of the first time step; Step 3-5: After the initial features of the current time step starting from the second time step are input into the membrane attenuation coefficient update module, the first pulse convolution module updates the pulse state of the previous time step to obtain the pulse update feature, the second pulse convolution module updates the initial features of the previous time step to obtain the input update feature, and the first normalization layer adds the pulse update feature and the input update feature and normalizes them to obtain the membrane attenuation coefficient of the current time step; Step 3-6: The membrane potential update module uses the membrane attenuation coefficient of the previous time step to attenuate the final membrane potential of the previous time step to obtain the attenuated membrane potential, and then adds the attenuated membrane potential to the initial feature of the previous time step to obtain the membrane potential of the current time step; Step 3-7: Update the pulse state of the current time step through the pulse emission module. The pulse emission module uses a step function to convert the membrane potential of the current time step into the pulse state S[t] of the current time step, S[t] = Θ(V[t]-V th ), Θ(·) represents the step function, V[t] represents the membrane potential at the current time step, and the pulse state S[t] at the current time step is used as the updated initial feature of the current time step; Step 3-8: If V[t]-V th > 0, the membrane potential reset module subtracts the preset pulse emission threshold from the membrane potential of the current time step, and updates the membrane potential of the current time step to obtain the final membrane potential of the current time step; if V[t]-V th ≤0, the membrane potential at the current time step remains unchanged and serves as the final membrane potential at the current time step.