Pulse neural network training and processing method based on adaptive integer neuron

CN122819348APending Publication Date: 2026-09-25HENAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611028981.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]为了解决上述标准漏积分发放神经元因二值硬阈值机制导致的表征粒度不足、量化损失较大,以及在复杂扰动环境下膜电位与脉冲状态不稳定的的技术问题,本发明的目的在于提供一种基于自适应整数神经元的脉冲神经网络训练及处理方法,所采用的技术方案具体如下:

Benefits of technology

[0016]本发明具有如下有益效果:通过在脉冲神经网络中配置自适应整数漏电流积分发放神经元,利用动态放电阈值将膜电位转换为有界整数脉冲,有效克服了传统二值脉冲硬截断造成的量化损失,显著提升了网络对特征表征的精细粒度;同时,通过结合动态放电阈值和有界整数脉冲对残余状态进行软复位更新,平滑了神经元放电后的状态过渡,维持了膜电位的时间动态稳定性;进而,利用时空注意力机制对具备高信息密度的有界整数脉冲进行精准加权调制,有效增强了网络捕捉关键时空依赖的能力,从根本上缓解了复杂扰动环境下的脉冲跳变与误差放大问题,最终通过端到端优化显著提升了模型在复杂环境下的抗干扰鲁棒性与图像分类准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819348A_ABST
    Figure CN122819348A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of neural network, in particular to a kind of pulse neural network training and processing method based on adaptive integer neuron.The sample image and real label are input into the pulse neural network configured with adaptive integer leakage current integral release neuron during training;At each time step, membrane potential is updated according to synaptic input and last time step residual state, and bounded integer pulse is generated based on dynamic firing threshold;Residual state is reset in combination with membrane potential, bounded integer pulse and dynamic firing threshold soft reset residual state;Then, bounded integer pulse is modulated using spatiotemporal attention map and task loss is constructed to perform end-to-end optimization.Processing time, bounded integer pulse is unfolded into binary pulse sequence for sparse accumulation reasoning.The present application can improve training representation granularity and maintain pulse inference efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of neural network technology, and specifically to a method for training and processing a spiking neural network based on adaptive integer neurons. Background Technology

[0002] Spiking Neural Networks (SNNs) achieve energy-efficient and low-latency computation through sparse, event-driven spike processing, making them highly attractive for neuromorphic computing deployments and real-time edge sensing. By encoding information into time-precise spike sequences and explicitly modeling the dynamic characteristics of neurons, SNNs are highly similar to biological nervous systems, naturally capturing temporal information in sequential data. Compared to traditional Artificial Neural Networks (ANNs), SNNs exhibit significant advantages in deployment on neuromorphic hardware, particularly in energy efficiency and low-latency inference. These characteristics make SNNs particularly suitable for resource-constrained applications requiring real-time responses, such as autonomous sensing and edge intelligence systems. In SNNs, Leaked Integral Discharge (LIF) neurons are commonly used basic processing units. Furthermore, to enhance the network's ability to capture spatiotemporal dependencies in sequential data, existing techniques often introduce spatiotemporal attention mechanisms, which utilize continuous attention weights to multiplicatively weight and modulate discrete spikes, thereby improving the model's performance on ideal, clean data.

[0003] However, existing attention-based spiking neural networks have the following obvious limitations when faced with complex real-world environments: First, standard LIF neurons employ a hard-threshold firing mechanism, whose output is strictly limited to discrete binary (0 or 1) pulses. This single binarization process restricts the fine granularity of feature representation and incurs a non-negligible quantization loss during the conversion of continuous membrane potentials into binary pulses, hindering the efficient transmission of complex information in deep networks.

[0004] Secondly, in order to compensate for the aforementioned binary quantization loss, if the neuron output is directly modified into a continuous value or a high-bit-width integer, dense multiplication operations must be introduced at the synaptic connections during the inference phase. This would completely destroy the event-driven sparsity of SNN, leading to a sharp increase in hardware inference power consumption and completely losing the energy efficiency advantage of the underlying SNN.

[0005] Finally, attention-based SNNs exhibit extreme vulnerability to noise or minute perturbations. The root cause of this vulnerability lies in the interaction between hard impulse quantization and continuous attention modulation in existing networks: on the one hand, hard binary quantization makes neuronal responses highly sensitive to minute membrane potential perturbations, easily triggering abrupt impulse changes; on the other hand, existing continuous attention mechanisms, when performing multiplicative modulation on discrete impulses, further amplify such perturbation errors, accumulating and propagating these errors through cyclic time dynamics. Existing techniques often focus only on performance improvements under noise-free conditions, lacking effective mechanisms to suppress these error propagation paths, resulting in highly unstable outputs when faced with input perturbations. Summary of the Invention

[0006] To address the technical problems of insufficient representation granularity, significant quantization loss, and unstable membrane potential and pulse state under complex perturbation environments caused by the binary hard threshold mechanism in standard under-integral firing neurons, the present invention aims to provide a training and processing method for spiking neural networks based on adaptive integer neurons. The specific technical solution adopted is as follows: This invention provides a training method for a spiking neural network based on adaptive integer neurons, the method comprising: Sample images and their true labels are acquired and input into a spiking neural network to be trained. The spiking neural network is configured with adaptive integer leakage current integral firing neurons. At each time step, the adaptive integer leakage current integral firing neuron updates the membrane potential of the current time step based on the received synaptic input of the current time step and the residual state of the previous time step; it obtains the residual state of the previous time step and calculates the dynamic discharge threshold of the current time step, and converts the membrane potential to a bounded range based on the dynamic discharge threshold to generate a bounded integer pulse; based on the membrane potential of the current time step and the bounded integer pulse, and in combination with the dynamic discharge threshold, it performs a soft reset update on the residual state of the current time step. A spatial attention map for the current time step is generated based on a spatiotemporal attention mechanism. The bounded integer pulse is modulated using the spatial attention map to obtain the modulation output characteristics of the current time step. The modulation output features at each time step are mapped to corresponding logical values. The task loss is constructed based on the logical values ​​at each time step and the true labels. The spiking neural network is then optimized and updated end-to-end based on the task loss.

[0007] Furthermore, the membrane potential at the current time step is updated, and the corresponding calculation formula is as follows: in, Indicates the current time step The membrane potential; Indicates the previous time step The residual state; Indicates the current time step received. Synaptic input; Indicates from the previous level The first neuron is connected to the current layer. Synaptic weights of individual neurons; Indicates the current layer number Bias of each neuron; Indicates the current time step No. The pulse input to a presynaptic neuron.

[0008] Furthermore, the residual state of the previous time step is obtained and the dynamic discharge threshold of the current time step is calculated. The corresponding calculation formula is as follows: = in, Indicates the current time step The dynamic discharge threshold; This indicates the preset initial reference discharge threshold. This represents the adaptive adjustment coefficient; Indicates the previous time step The residual state.

[0009] Furthermore, the bounded integer pulse is generated, and the corresponding calculation formula is: in, Indicates the current time step Bounded integer pulses; Indicates the current time step The membrane potential; Indicates the current time step The dynamic discharge threshold; Indicates the preset maximum discharge capacity; This represents the floor function; This represents a truncation function used to restrict the input value to a certain range. Within the range.

[0010] Furthermore, a soft reset update is performed on the residual state at the current time step, and the corresponding calculation formula is as follows: in, Indicates the current time step The residual state; Indicates the current time step The membrane potential; Indicates the leakage factor. ; Indicates the current time step Bounded integer pulses; Indicates the current time step The dynamic discharge threshold.

[0011] Furthermore, the spiking neural network is optimized and updated end-to-end based on the task loss, including: Determine the spatial regularization term used to constrain the spatial attention map in the spatial dimension Jacobian gradient, and the temporal regularization term used to constrain the difference between spatial attention maps in adjacent time steps, respectively. The spatial regularization term and the temporal regularization term are weighted based on preset weights and fused with the task loss to construct an overall loss function. Based on the overall loss function, the spiking neural network is optimized and updated end-to-end.

[0012] Furthermore, the spatial regularization term used to constrain the spatial attention map in the Jacobian gradient of the spatial dimension is determined, and the corresponding calculation formula is as follows: in, Represents the space regularization term; Indicates input At the current time step Spatial attention map; This indicates that the attention map is applied to the input. The Jacobian matrix; Denotes the Frobenius norm; This represents the shear threshold used to prevent gradient explosion; This represents a function that takes the minimum value. This represents the total number of time steps.

[0013] Furthermore, the time regularization term used to constrain the spatial attention map differences between adjacent time steps is determined, and the corresponding calculation formula is as follows: in, Represents the time regularization term; and These represent the input samples respectively. At the current time step and the next time step Spatial attention map; Represents the L2 norm; This represents the total number of time steps.

[0014] Furthermore, the spiking neural network is optimized and updated end-to-end, including: When calculating the gradient of the network loss with respect to the membrane potential based on the task loss during backpropagation, a preset surrogate gradient function is used to replace the gradient of the membrane potential rounding operation. When the membrane potential at the current time step is greater than or equal to zero and less than or equal to the product of the maximum discharge capacity and the dynamic discharge threshold at the current time step, the gradient of the bounded integer pulse with respect to the membrane potential at the current time step is replaced with the reciprocal of the discharge threshold. Otherwise, the gradient of the bounded integer pulse with respect to the membrane potential at the current time step is replaced with zero.

[0015] The present invention also provides a data processing method for a spiking neural network based on adaptive integer neurons, the method comprising: The image to be classified is acquired and input into a pre-trained spiking neural network; the spiking neural network is configured with adaptive integer leakage current integral firing neurons; At each time step, the adaptive integer leakage current integral firing neuron updates the membrane potential of the current time step based on the received synaptic input of the current time step and the residual state of the previous time step; it obtains the residual state of the previous time step and calculates the dynamic discharge threshold of the current time step, and converts the membrane potential to a bounded range based on the dynamic discharge threshold to generate a bounded integer pulse; based on the membrane potential of the current time step and the bounded integer pulse, and combined with the dynamic discharge threshold, it performs a soft reset update on the residual state of the current time step; it expands the bounded integer pulse into a binary pulse sequence under several virtual time steps according to the maximum discharge capacity, and performs a sparse accumulation operation on the synaptic weights of the next layer network and the binary pulse sequence, and uses the result of the sparse accumulation operation as the synaptic input of the next layer network; A spatial attention map for the current time step is generated based on a spatiotemporal attention mechanism. The bounded integer pulse is modulated using the spatial attention map to obtain the modulation output features for the current time step. The classification result of the image to be classified is obtained based on the modulation output features for each time step.

[0016] This invention offers the following advantages: By configuring adaptive integer leakage current integral firing neurons in a spiking neural network and utilizing a dynamic firing threshold to convert the membrane potential into bounded integer pulses, the quantization loss caused by traditional binary pulse hard truncation is effectively overcome, significantly improving the fine granularity of the network's feature representation. Simultaneously, by combining the dynamic firing threshold and bounded integer pulses to perform soft reset updates on the residual state, the state transition after neuron firing is smoothed, maintaining the temporal dynamic stability of the membrane potential. Furthermore, by utilizing a spatiotemporal attention mechanism to precisely weight and modulate bounded integer pulses with high information density, the network's ability to capture key spatiotemporal dependencies is effectively enhanced, fundamentally alleviating the problems of pulse jumps and error amplification under complex perturbation environments. Finally, end-to-end optimization significantly improves the model's robustness against interference and image classification accuracy in complex environments.

[0017] Furthermore, by introducing a dual regularization mechanism to construct the overall loss function based on the task loss, the spatial regularization term strictly constrains the Jacobian gradient of the attention map in the spatial dimension, effectively suppressing the amplification effect of small spatial perturbation errors on the network. At the same time, the temporal regularization term constrains the difference between attention maps in adjacent time steps, effectively preventing the continuous accumulation and temporal drift of perturbation errors in the network's recurrent dynamic mechanism. The synergistic effect of the two measures fundamentally stabilizes the dynamic modulation process of the attention mechanism from both spatial and temporal dimensions, significantly enhancing the structural stability and robustness of the spiking neural network in the face of complex noise and environmental distortions.

[0018] Furthermore, a piecewise surrogate gradient function was introduced during backpropagation, successfully overcoming the technical bottleneck of gradient non-differentiability caused by the non-differentiability of the floor operation during bounded integer pulse generation, thus establishing a link for end-to-end optimization of the network. Simultaneously, by limiting the effective propagation interval of the membrane potential and truncating the gradient to zero outside the interval, the risk of invalid updates and gradient explosion caused by extreme membrane potentials was effectively avoided, ensuring numerical stability during training. In addition, the gradient values ​​within the effective interval were adaptively replaced with the reciprocal of the dynamic firing threshold, allowing the gradient amplitude during backpropagation to be dynamically scaled according to the current state of the neurons, thereby significantly improving the smoothness of network weight optimization and the final convergence accuracy of the model.

[0019] Furthermore, during the inference phase, bounded integer pulses are expanded into binary pulse sequences over several virtual time steps. This allows the network to retain high-precision integer-level information representations without loss, while successfully replacing the expensive multiply-accumulate operations originally required by the underlying hardware with lightweight sparse accumulation operations. This mechanism perfectly matches the event-driven nature of spiking neural networks, helps reduce the additional computational overhead that may be caused by the introduction of integer pulses, and maintains low inference energy consumption while improving classification robustness. Attached Figure Description

[0020] To clearly illustrate the technical features of this solution, the embodiments of the present invention include the following figures: Figure 1 This is a flowchart illustrating the steps of a spiking neural network training method based on adaptive integer neurons according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the overall structure of the REST-SNN framework according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating the steps of a spiking neural network data processing method based on adaptive integer neurons, according to an embodiment of the present invention. Detailed Implementation

[0021] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings.

[0022] The following will describe in detail, with reference to the accompanying drawings, a method for training and processing a spiking neural network based on adaptive integer neurons provided by an embodiment of the present invention.

[0023] Training method example: Before introducing the training method of the embodiments of the present invention, the following is a brief introduction to the leakage current type integral-fire (LIF) neuron and the spatiotemporal attention mechanism in SNN: LIF neurons incorporate passive membrane potential decay into the integral-and-fire (IF) model, and their continuous-time form is as follows: (1) in, Indicates film capacitance; Indicates time step synaptic input current; Indicates time step Leakage conductivity; Indicates resting potential; Indicates time step The membrane potential of a neuron. After time discretization, the membrane potential is updated as follows: (2) in, and They represent the first The first in the layer At time step, one neuron and The membrane potential; Indicates the first The presynaptic neuron to the first The connection weights of each neuron; Indicates the bias parameter; Indicates the first Layer At time step, one neuron Pulse input; Indicates the first The first layer (i.e., the layer above) At time step, one neuron Pulse input; Indicates the discharge threshold; Leakage factor is defined as follows: , Represents discrete time intervals; This represents the membrane time constant.

[0024] Pulse delivery is determined by the following formula: (3) in, This represents the Heaviside step function, which outputs 1 when the input is non-negative and 0 otherwise.

[0025] The spatiotemporal attention mechanism in SNNs is a widely used mechanism that can simultaneously enhance spatial saliency and capture temporal dependencies in SNNs. In spatiotemporal attention-based structures, synaptic currents are generated by continuous attention signals. Modulation: (4) in, Indicates at time step Attention-modulated synaptic current (i.e., modulated current). Indicates at time step Baseline synaptic current without attention modulation.

[0026] At each time step The network generates pulse tensors and attention tensor signal The final output calculation result is: (5) in, Indicates at time step The weighted output feature tensor (i.e., the final output response) is modulated by spatiotemporal attention; ⊙ represents the Hadamard product, i.e., element-wise multiplication. This method can extract continuous-value sensitivity from discrete impulse events.

[0027] This embodiment provides a training method for spiking neural networks based on adaptive integer neurons. This method proposes a robust and efficient spatio-Temporal Attention (REST-SNN) training framework. The REST-SNN framework integrates adaptive integer leakage current integral release (AI-LIF) neurons and attention-level dual regularization mechanism, aiming to improve the robustness and computational efficiency of the spatio-temporal attention mechanism under perturbation conditions.

[0028] like Figure 1 As shown, the spiking neural network training method based on adaptive integer neurons provided in this embodiment specifically includes the following steps: Step S1: Obtain sample images and their true labels and input them into the spiking neural network to be trained. The spiking neural network is configured with adaptive integer leakage current integral firing neurons.

[0029] Obtaining sample images and their true labels provides a clear supervised learning data source and optimization benchmark for model training.

[0030] Sample images and their ground truth labels refer to the fundamental data sources used for supervised training of the spiking neural network model. Sample images provide specific visual feature input to the network, while ground truth labels represent the objective, actual category of the image (usually expressed as a category index, one-hot vector, or a soft label proportionally mixed with CutMix data augmentation), providing a clear optimization benchmark for the model's loss function.

[0031] While existing methods have improved the performance of SNNs to some extent, standard LIF neurons are still limited by their hard-threshold firing mechanism, whose outputs are strictly limited to discrete binary (0 / 1) pulses. This inherently restricts the ability to represent fine-grained information and weakens information propagation in deep networks. Furthermore, mapping continuous membrane potentials to binary pulses inevitably introduces quantization loss, leading to gradient flow instability during training. To balance representational power and energy efficiency, this embodiment introduces Adaptive Integer Leakage Current Integral Fire Neuron (AI-LIF) neurons into the spiking neural network. AI-LIF allows multi-bit integer activation during training to preserve high-density information, and utilizes virtual time-step expansion during inference to maintain the sparsity and efficiency of pulse-driven computation.

[0032] Step S2: At each time step, the adaptive integer leakage current integral firing neuron updates the membrane potential of the current time step based on the received synaptic input of the current time step and the residual state of the previous time step; it obtains the residual state of the previous time step and calculates the dynamic discharge threshold of the current time step, and converts the membrane potential to a bounded range based on the dynamic discharge threshold to generate a bounded integer pulse; based on the membrane potential of the current time step and the bounded integer pulse, and in combination with the dynamic discharge threshold, it performs a soft reset update on the residual state of the current time step.

[0033] The internal dynamic mechanisms of AI-LIF neurons, such as Figure 2 As shown. This AI-LIF neuron can be represented as a discrete-time dynamic system. At time step... synaptic input Compared with the previous time step residual state Integrating, we obtain the membrane potential. : (6) in, Indicates from the previous level The first neuron is connected to the current layer. Synaptic weights of individual neurons; Indicates the current layer number Bias of each neuron; Indicates the current time step No. The pulse input to a presynaptic neuron.

[0034] Obtain the residual state from the previous time step and calculate the dynamic discharge threshold for the current time step: = in, Indicates the current time step The dynamic discharge threshold; This indicates the preset initial reference discharge threshold; This represents the adaptive adjustment coefficient; Indicates the previous time step The residual state.

[0035] When a neuron outputs an integer, it is activated. At that time, the membrane potential is determined by the leakage factor. Controlled soft reset: (7) in, Indicates the current time step The residual state; Regulating the time decay process of historical information is equivalent to adjusting the membrane time constant.

[0036] This form echoes the burst firing mechanism of biological neurons: neurons do not merely express stimulus intensity through isolated binary events, but can encode richer information through high-frequency, multi-pulse bursts. By allowing integer outputs S[t], AI-LIF can retain high-density representational information that would otherwise be discarded in hard-threshold mechanisms. Simultaneously, due to the leakage factor... The regulated soft reset mechanism can simulate the natural recovery timescale after ion channel inactivation, thereby effectively preventing violent oscillations in membrane potential after a burst and suppressing pulse jitter caused by input noise. This stability achieved at the microscopic neuronal level provides the necessary biophysical basis for the macroscopic robustness boundary and drift inhibition strategy established in subsequent steps.

[0037] To mitigate the inherent accuracy loss of binary SNNs, AI-LIF employs quantized integer activation during the training phase. Current neuron output. By normalizing the membrane potential threshold and trimming it to a bounded interval: (8) in, Indicates the maximum discharge capacity; This represents the floor function; This represents a truncation function used to restrict the input value to a certain range. Within the range.

[0038] Step S3: Generate a spatial attention map for the current time step based on the spatiotemporal attention mechanism, and use the spatial attention map to modulate the bounded integer pulse to obtain the modulation output features of the current time step.

[0039] The spatial attention map for the current time step is generated based on the spatiotemporal attention mechanism. This spatial attention map corresponds to the attention tensor signal in equation (5), which can be obtained by combining conventional convolution and pooling operations with a normalized activation function. At the same time, the bounded integer pulses output by each neuron constitute the pulse tensor for the current time step. According to equation (5), the pulse tensor composed of bounded integer pulses is weighted and modulated using the spatial attention map through element-wise multiplication (Hadamard product) to obtain the modulated output features (i.e., the weighted output feature tensor) for the current time step.

[0040] Step S4: Map the modulation output features of each time step to the corresponding logical values, construct the task loss based on the logical values ​​of each time step and the true labels, and determine the spatial regularization term used to constrain the Jacobian gradient of the spatial attention map in the spatial dimension, and the temporal regularization term used to constrain the difference between the spatial attention maps of adjacent time steps; weight the spatial regularization term and the temporal regularization term based on preset weights and fuse them with the task loss to construct the overall loss function; perform end-to-end optimization and update of the spiking neural network based on the overall loss function.

[0041] The degradation of the robustness of the spatiotemporal attention mechanism in the face of input perturbations is mainly attributed to the sensitivity of the attention map to small changes in the input space (spatial error amplification) and the temporal drift of the error as the network dynamically evolves (temporal error accumulation). Based on this error decomposition theory, this embodiment does not adopt the conventional radical approach of directly constraining discontinuous impulses, but follows the design idea of ​​"from upper bound to target", transforming the theoretical boundary into a differentiable alternative optimization objective: that is, to suppress spatial sensitivity by constraining the Frobenius norm of the Jacobian gradient of the attention map through a spatial regularization term, and to block the accumulation of temporal drift of the error by limiting the difference of the attention map in adjacent time steps through a temporal regularization term; finally, to construct the overall loss function by fusing the double regularization term representing the stability of the network structure with the task loss representing the classification accuracy through a weighted mechanism, the task loss is obtained by comparing the logical value obtained by mapping the modulated output features of each time step with the true label, thereby forcibly guiding the network weights to converge toward the parameter space that balances high discriminability and strong environmental robustness from the underlying dynamic mechanism during backpropagation.

[0042] Specifically, during the end-to-end optimization and update of the spiking neural network, since the rounding and pruning operations within AI-LIF neurons are non-differentiable, an alternative gradient is used during training. Achieve end-to-end optimization. Gradient propagation will only occur when the membrane potential is within the effective range. K (9) During the inference phase, to maintain pulse-driven computation, this embodiment decomposes each integer pulse into a series of binary {0,1} pulses through virtual time steps. For example... Figure 2 As shown, the original The time step was expanded to A virtual time step, in which For the maximum discharge capacity, the first Layer at the original time step Bounded integer pulses From binary sequence It indicates that, and has: , (10) in, Indicates the first Layer at the original time step Bounded integer pulses; Indicates the first Layer at the original time step The unfolded first A binary pulse for each virtual time step, with a value of either 0 or 1. At the original time step... The number of binary pulses is determined by the maximum discharge capacity. Decide.

[0043] Since synaptic weights are linear operators, the input to the next layer can be strictly rewritten as: (11) in, Indicates the next level, i.e., the first Layer in time step Synaptic input; Indicates the first Synaptic weights from one layer to the next.

[0044] because For strictly binary values, the dense multiply-accumulate (MAC) operations required for integer activation can be replaced by sparse accumulation (AC). From a complexity perspective, quantized ANNs with equal precision typically require... The cost of dense MAC. Among them, This represents the total number of neurons in the network. This represents the computational cost of a single multiply-accumulate operation. In contrast, AI-LIF inference complexity only increases with the actual number of issued units, i.e. .in, This represents the number of neurons that are active (i.e., have actually generated and fired pulses) at the current time step. Due to the sparsity of event-driven processes... Furthermore, in hardware, AC operations are typically significantly less expensive than MAC operations. Therefore, the theoretical overhead of this virtual expansion is still within a controllable range and is significantly lower than that of dense integer networks.

[0045] It is important to emphasize that this extension merely reindexes the time axis to achieve finer-grained temporal integration over the same static input, while the spatial weights remain unchanged. Therefore, it retains both the original model capacity and parameter count, and differs from traditional ANN quantization. Furthermore, unlike the static, dense activation paradigm of quantized ANNs, AI-LIF makes fuller use of the time dimension through state-dependent integration and leakage mechanisms, improving accuracy while maintaining the sparsity of impulse computation events.

[0046] In actual operation, the virtual time step This does not lead to a significant increase in inference latency. This is because the extended extra micro-timesteps mainly perform lightweight sparse accumulation and thresholding operations, and each micro-timestep shares the same set of network weights. Combined with the event-driven impulse sparsity, these operations can be efficiently executed by the vectorized kernel of the underlying hardware. For rigorous algorithm performance comparison, this embodiment measured the average inference latency under a fixed hardware configuration (NVIDIA TeslaA 100 GPU, batch size of 32, standard input dimension). As shown in Table 2, compared with the standard LIF baseline under the same backbone and operating environment, the latency of AI-LIF only shows marginal changes. It should be noted that the highly parallel dense tensor computation paradigm of standard GPUs cannot fully reflect the asynchronous event-driven energy consumption and latency characteristics of dedicated neuromorphic hardware (such as Intel Loihi). Therefore, the GPU latency results in this embodiment are mainly used to verify that there is no significant increase in algorithm-level complexity, rather than to predict the absolute latency on the neuromorphic chip.

[0047] To theoretically characterize the robustness of LIF and AI-LIF frameworks under input perturbation, this embodiment derives a unified error bound that simultaneously quantifies the spatial sensitivity induced by the attention mechanism and the temporal instability of neuronal dynamic accumulation. Specifically, it considers a clean input x and its perturbated corresponding value. ,in, , ϵ represents the upper bound of the input perturbation. The output bias is mainly affected by two interrelated factors: (i) the spatiotemporal attention map. (ii) Input sensitivity; (ii) Propagation of perturbations through the recursive dynamics of LIF and AI-LIF neurons.

[0048] To analyze robustness under input perturbations, this embodiment introduces notation for attention-modulated AI-LIF dynamics. Let the neuron state at time step t be represented as... ,in, Represents membrane potential. Indicates the discharge threshold. This indicates pulse output. The following analysis focuses on the attention map. Coupled evolution with recursive membrane potential dynamics. When D=1, the neuron model proposed in this embodiment degenerates into the standard LIF model.

[0049] To derive an explanatory upper bound for the perturbation, this embodiment employs the following conventional regularity conditions: (i) Lipschitz continuity of attention. For each time step Attention map With respect to the local Lipschitz continuity of the input, and its Jacobian matrix being bounded under the Frobenius norm: in, This represents the Jacobian matrix of the attention map with respect to the input x; This represents the Frobenius norm.

[0050] This condition is used to control the amplification of spatial perturbations by attention modulation.

[0051] (ii) Bounded substitution gradient. Pulse generation process. Using the replacement gradient operator Conduct training and meet the requirements. This is to avoid excessively large gradients when activating backpropagation with discontinuous pulses. Among these, This represents the pulse generation / quantization activation function. This represents the membrane potential variable after threshold normalization. This indicates the upper bound of the alternative gradient magnitude.

[0052] (iii) The input current is bounded. The synaptic current before attention modulation is uniformly bounded: .in, Indicates pre-attention synaptic current. Represents the infinite norm, This represents a unified upper bound for the input current.

[0053] This condition is used to ensure that the membrane potential dynamics do not undergo unbounded fluctuations under perturbation input.

[0054] (iv) Regularity of time drift. The change in attention between adjacent time steps is bounded. One-step attention drift is defined as: and in, This represents the L2 norm difference between attention maps at adjacent time steps; Indicates single-step attention drift The upper bound is used to ensure that the temporal evolution of attention is controlled and does not induce the accumulation of unbounded errors.

[0055] Based on the spatiotemporal attention form described above and the output definition in equation (5), the network output is determined by the element-wise interaction between spatial attention and temporal impulse response. Therefore, the output perturbation can naturally be decomposed into two sources: spatial sensitivity caused by attention input dependence, and temporal drift caused by recursive impulse dynamics.

[0056] Theorem 1 (Unified Robustness Upper Bound Based on Double Regularity). Under input perturbation... and At that time, the instantaneous output deviation at time step t satisfies: (12) in, Indicates time step ; output deviation; This represents the Jacobian matrix of the attention map with respect to input x under the perturbation upper bound; Indicates membrane leakage factor and ; Indicates the summation index; Indicates the upper bound of the perturbation; Indicates a second-order minor quantity; The constant formed by the combination of the substitution gradient, input current, threshold, and adaptive coefficients is: (13) in, This represents the uniform upper bound of the pre-attention synaptic current (i.e., the synaptic current before attention modulation); Indicates the threshold adaptive coefficient; This is used to replace the upper bound of the gradient magnitude.

[0057] Furthermore, regarding equation (12) in Summing above, and using the upper bound of the geometric series: The upper bound of the cumulative deviation can be obtained: (14) in, and Let the scale constants representing the spatial sensitivity term and the temporal drift term be respectively: (15) in, Indicates the upper bound of the normalized membrane potential; This represents the initial threshold.

[0058] This embodiment demonstrates the above results through three steps: output perturbation decomposition, recursive membrane perturbation analysis, and time accumulation analysis.

[0059] Step 1: Output perturbation decomposition. This is defined by the output. Disturbance term: (16) It can be expanded as follows: (17) in, Indicates the input after introducing the disturbance. At time step The corresponding network output feature tensor; Represents the unperturbed raw input At time step The corresponding network output feature tensor; Represents the unperturbed raw input At time step The impulse response tensor; This represents element-wise multiplication; This indicates the amount of perturbation change in the attention map; This represents the amount of disturbance change in the impulse response.

[0060] Ignoring second-order interaction terms, we obtain a first-order upper bound: (18) Subsequently, the pulse perturbation is jointly controlled by the bounded substitution gradient and the threshold membrane potential dynamics. We can obtain: (19) in, This represents the amount of perturbation change in the membrane potential variable after threshold normalization; Indicating time step in theoretical analysis The membrane potential; This represents the amount of change in membrane potential perturbation; Indicates the initial threshold; Indicates time step The amount of disturbance change in the dynamic discharge threshold.

[0061] Further substitution After that, it can be written as: + = in, Indicates the threshold adaptive coefficient; Indicates the upper bound of the normalized membrane potential; Indicates all time steps Take the supremum, which represents the global maximum amplitude.

[0062] Step 2: Recursive analysis of membrane potential perturbations. Considering membrane potential dynamics. in, Indicates time step The membrane potential of a neuron; Indicates the previous time step The membrane potential of a neuron; Indicates pre-attention synaptic current; This represents the membrane potential leakage factor.

[0063] Under perturbation input, the membrane potential deviation satisfies: (20) in, This represents the cumulative difference caused by attention drift at time steps.

[0064] From the local Lipschitz continuity of attention, we can obtain: in, Indicates the upper bound of the perturbation The upper bound of the Frobenius norm within a limited local neighborhood.

[0065] Simultaneously, by performing telescopic accumulation on the drift terms of adjacent time steps, the cumulative drift at the k-th step can be obtained as follows: in, This represents the perturbation term accumulated by attention time drift up to step k; This indicates a one-step attention drift between adjacent time steps.

[0066] Expanding equation (20) yields: (twenty one) Step 3: Instantaneous upper bound and cumulative upper bound. Substituting equation (21) into equation (19), we can obtain the instantaneous upper bound in equation (12), where: For equation (12) in Summing. For the spatial term, estimated by geometric series: get: For the time term, it grows linearly: It can be seen that it has a double-cumulative form, therefore: The two contributions are combined to complete the proof.

[0067] Theorem 1 shows that the robustness degradation of attention-modulated SNNs is mainly determined by two structural factors: a spatial sensitivity term related to the Jacobian matrix of the attention map input, and a temporal drift term related to the attention changes at each time step. These two terms not only indicate where perturbations are amplified, but also indicate which quantities should be prioritized for control during training.

[0068] The reason this embodiment applies regularization to the attention map, rather than directly constraining the pulse output or membrane potential state, is that attention is the primary continuous modulation channel through which input perturbations interact with discrete pulse dynamics. In contrast, pulse sequences and threshold-triggered membrane potential transitions are highly non-smooth, unsuitable as direct optimization targets, and prone to causing training instability. Therefore, constraining the spatial and temporal variations of attention can suppress perturbation amplification and temporal accumulation in a more optimizable manner and closer to the structural mechanism.

[0069] However, the exact operator-level quantities appearing in Direct Optimization Theorem 1 are often impractical. In particular, accurately characterizing the local Lipschitz constant requires controlling the operator norm of the Jacobian matrix, which is computationally expensive and unstable in high-dimensional training. Therefore, the embodiment introduces an alternative objective that maintains a monotonic relationship with the dominant term of the theoretical bound while facilitating standard gradient optimization. It should be noted that this design does not claim to precisely minimize the theoretical bound, but rather provides a stable and computationally feasible approximation while maintaining consistency with the theoretical decomposition.

[0070] In this embodiment, the spatial regularization term is defined as: (twenty two) Define the time regularization term as: (twenty three) in, Represents the space regularization term; This represents the time regularization term, used to penalize differences in attention maps between adjacent time steps; This represents a function that takes the minimum value. This indicates that the attention map is applied to the input. The Jacobian matrix; Denotes the Frobenius norm; Represents the L2 norm; This represents the shearing threshold used to prevent gradient explosion, which is used to avoid excessively large Jacobian responses that could lead to optimization instability.

[0071] The spatial term in Theorem 1 is formally related to the local Lipschitz quantity, which is typically characterized by the operator norm. However, in practice, directly optimizing the spectral norm is costly and unstable. Therefore, this embodiment uses the Frobenius norm as a substitute for the smoothing upper bound because: (twenty four) Therefore, minimize It can serve as a principled and optimizable proxy target for reducing attention-induced spatial perturbation amplification and is fully compatible with end-to-end automatic differentiation.

[0072] The final training objective consists of the task loss and the proposed dual regularization term: Combining the task loss with the aforementioned double regularization term, we construct the overall loss function: (25) in, Represents the overall loss function; Each represents and Corresponding weights; task loss Based on time average logits calculate, Indicates time step The logical value vector output by the network is the class confidence score obtained by linearly mapping the modulated output features of the current time step to the fully connected classification layer. This represents the average vector of the network's output logical values ​​across all time steps. This indicates the total number of time steps.

[0073] When introducing a mixing ratio of After CutMix data augmentation, the task loss is defined as: (26) in, This represents the standard cross-entropy loss; , This represents the two true labels corresponding to the CutMix mixed sample; This represents the cross-entropy loss between the average logical value and the first true label; This represents the cross-entropy loss between the average logical value and the second true label.

[0074] To allocate weights in a reasonable manner ( This embodiment relates it to the two dominant components in Theorem 1. Schematably, the cumulative deviation can be written as: (27) in, and The theorem constants related to spatial sensitivity and temporal drift are collected separately. This indicates that... and This should reflect the relative importance of these two structural factors.

[0075] In actual training, the precise scales of these two terms depend on data statistics, perturbation strength, and optimization dynamics. Therefore, the implementation does not mandate the use of a rigid analytical scale. Instead, it normalizes by time range. and Subsequently, the equilibrium initialization provides a theoretically sound default choice: it treats spatial scaling and time accumulation as equally important first-order robustness factors, consistent with the symmetric decomposition of Theorem 1.

[0076] Empirically, this design choice was further supported by the subsequent sensitivity analysis: excessively weak or strong regularization degrades performance, while a balanced setting achieves a good trade-off between robustness and discriminability. Therefore, It should not be understood as a universally optimal value, but rather as a stable and theoretically consistent default configuration under the current training protocol.

[0077] The following section evaluates REST-SNN on a structured robustness benchmark and compares it with representative state-of-the-art SNN methods to verify its performance in static image classification tasks facing multi-class environmental perturbations. Furthermore, this embodiment separates the contributions of different modules through ablation experiments and efficiency analysis, and characterizes the trade-off between accuracy, robustness, and efficiency at different simulation time steps. To ensure reproducibility, this embodiment also provides implementation details and key hyperparameter settings. All methods use a consistent backbone configuration.

[0078] To rigorously evaluate the robustness of the REST-SNN framework, this embodiment employs a structured evaluation scheme incorporating various environmental perturbations, covering three typical types of visual distortion in the real world: (1) Geometric transformation: used to evaluate spatial invariance under translation and rotation on the coordinate plane; (2) Photometric variation: Simulates changes in illumination and noise caused by the sensor by random color dithering; (3) Structural occlusion: Simulates partial information loss caused by environmental obstacles through random erasure and fixed color block occlusion operations.

[0079] During the evaluation process, static image data is transformed into a non-stationary sequence by applying these perturbations at multiple intensity levels and combining a simulation process with T discrete time steps and random noise injection. This processing method effectively simulates the time-varying volatility and high-frequency noise characteristics unique to event-based data streams, fully verifying the robustness of the method provided in this embodiment in dealing with complex temporal perturbations and spatial distortions. To verify the robustness of REST-SNN, this embodiment conducts systematic experiments on three benchmark datasets: CIFAR-10, CIFAR-100, and ImageNet, and compares it with a group of representative state-of-the-art SNN models. Specifically, these include: the SpikingTransformer framework STATten SNN for capturing long-range dependencies, FSTA-SNN for spatiotemporal attention modeling using frequency domain information, STAA-SNN for optimizing feature fusion through temporal aggregation, and GAC-SNN for refining impulse representation using multiplicative gating. In addition, for a more comprehensive comparison with non-native training paradigms, this embodiment also incorporates the representative ANN to SNN conversion method ANN2SNN. Table 1 presents the quantitative comparison results on CIFAR-10, CIFAR-100, and ImageNet. Bold text in Table 1 indicates the best results, and underlined text indicates the second-best results. Overall, REST-SNN maintains competitive clean accuracy in most evaluation scenarios and demonstrates stable robustness advantages.

[0080] On CIFAR-10, REST-SNN achieves a clean accuracy of 96.87%, 0.13% higher than the second-best STAA-SNN. More importantly, under perturbation conditions, REST-SNN achieves an average accuracy of 83.46%, 1.88% higher than the competitive GAC-SNN. This advantage is even more pronounced on the more complex CIFAR-100. Although baseline methods struggle to balance complexity and stability simultaneously, REST-SNN still achieves a clean accuracy of 84.14%, a 2.59% improvement over the state-of-the-art STAA-SNN. In terms of robustness, REST-SNN achieves the highest average accuracy of 58.29%, demonstrating strong resistance to rotation (67.56%) and color perturbations (25.36%).

[0081] ImageNet results further demonstrate the robustness-oriented nature of REST-SNN. As shown in Table 1, methods such as FSTA-SNN and STATten-SNN achieve higher accuracy of approximately 79.9% on clean data, indicating stronger fitting ability under standard conditions. In contrast, REST-SNN achieves a clean accuracy of 68.62%, but its perturbation-average accuracy reaches the highest at 46.43%, which is 1.45% and 3.19% higher than FSTA-SNN and STATten-SNN, respectively. This result shows a clear trade-off between peak performance on clean data and robustness to perturbation inputs. REST-SNN does not only pursue maximizing clean accuracy, but focuses on improving structural stability in the presence of spatial and temporal perturbations. Notably, REST-SNN achieves best results on several challenging perturbation types, including rotation (52.83%), color jitter (20.81%), and erase occlusion (57.75%). These results demonstrate that AI-LIF neurons and a dual regularization strategy help the model learn more stable attention-modulated representations under distribution shifts, even though this robustness gain is accompanied by a modest decrease in clean accuracy on large-scale datasets.

[0082] Table 1 Comparison of Top-1 accuracy (%) under clean and perturbed conditions To systematically evaluate the contributions of AI-LIF neurons and the dual regularization strategy, this embodiment conducts two sets of complementary analyses. First, controlled ablation is performed under the same training and evaluation settings to quantify the accuracy improvement of AI-LIF relative to the standard LIF baseline. Second, the sensitivity of SSR and TCR regularization weights is analyzed to understand how these hyperparameters affect the trade-off between robustness and accuracy. Unless otherwise specified, all ablation settings share the same backbone structure and training schedule, changing only the neuron type or regularization term to ensure fair comparison.

[0083] While virtual timestep unrolling increases the nominal number of time indices, it does not introduce additional spatial computations, such as parameters or attention heads, and is implemented with low-cost pulse update sequences sharing weights. On GPUs, these updates consist primarily of memory-friendly accumulation and thresholding, benefiting from vectorization and kernel-level optimizations; in neuromorphic execution, they maintain event-driven sparsity, with inactive units incurring almost no additional cost. Therefore, the actual wall clock latency is mainly determined by the backbone network and attention modules, with AI-LIF unrolling introducing only marginal overhead, consistent with the almost constant latency in Table 2.

[0084] To quantify the contribution of the proposed neuron design, controlled ablation was performed under clean input conditions on the CIFAR-100. As shown in Table 2, the AI-LIF neuron achieved a Top-1 accuracy of 82.48%, a 2.03% improvement over the standard LIF baseline (80.45%). Notably, this performance improvement did not introduce additional space overhead. The number of model parameters remained at 12.68M, and the peak GPU memory remained at 273.76MB, indicating that the virtual timestep unfolding utilizes temporal dynamics rather than increasing spatial model complexity. Simultaneously, inference efficiency was maintained, with a latency of 13.8ms and a throughput of 72FPS. These results demonstrate that AI-LIF can enhance representational capabilities while preserving the inherent computational efficiency of SNNs.

[0085] Table 2 compares the controlled ablation results of AI-LIF neurons with the standard LIF baseline on CIFAR-100. To further evaluate the metabolic efficiency of the proposed REST-SNN framework in this embodiment, the spatiotemporal firing rate dynamics in the hierarchical network stages were monitored, and the results were compared with those of standard LIF and AI-LIF neurons. Monitoring results show that the normalized firing rate of AI-LIF neurons is highly consistent with the standard LIF baseline across all monitored sub-modules. Quantitative statistics further indicate that the average firing rate difference between AI-LIF and standard LIF at each layer is controlled within 1%. This result is significant because the energy consumption of neuromorphic hardware is essentially proportional to the number of generated impulse events. The almost identical firing statistics suggest that the performance improvement of REST-SNN does not stem from higher neuronal activity levels, but rather from the higher representation density resulting from integer quantization during the training phase. Therefore, REST-SNN alleviates the bottleneck of binary impulse signal information while maintaining the low-power characteristics of the underlying SNN layers.

[0086] To quantify the independent effects of spatial and temporal regularization, this embodiment adjusts the weighting coefficients. (SSR) and (TCR) is used to study hyperparameters, and these two coefficients together control the trade-off between robustness and accuracy.

[0087] With spatial smoothness term ( For example, the impact of ) As the value increased from 0.1 to 1.0, the accuracy steadily improved from 69.40% to 70.74%, indicating that SSR effectively mitigated perturbation amplification by constraining the Jacobian norm. Continuing... When the value was increased to 5.0, the accuracy dropped to 70.15%, indicating that excessively strong spatial regularization limits the model's ability to capture fine-grained discriminative features.

[0088] With time consistency items ( For example, the impact of ) : Similarly, when Increasing the value from 0.1 to 1.0 improved the accuracy from 69.16% to 70.65%, validating that TCR can reduce noise-induced pulse drift and enhance temporal stability. However, when... When the accuracy was further increased to 2.0, it dropped to 69.64%, indicating that excessive smoothing in the time domain may have suppressed the impulse dynamics necessary for effective information encoding.

[0089] Overall, The goal is to achieve the optimal balance between robustness and discriminability, that is, to maximize stable dynamics without compromising the fidelity of pulse characterization.

[0090] To characterize the trade-off between robustness and computational efficiency, this embodiment compares REST-SNN with a competitive GAC-SNN baseline under translational noise conditions, using multidimensional efficiency metrics and different simulation time windows. Unless otherwise specified, both methods use the same backbone and training protocol to ensure fairness in cost and performance comparisons.

[0091] In terms of computational efficiency and robustness, experimental results show that REST-SNN improves robustness while maintaining almost the same efficiency profile. Specifically, REST-SNN increases accuracy from 66.88% to 72.14% while maintaining the spatial parameterization scale (12.68M parameters), and peak GPU memory only slightly increases from 273.76MB to 274.21MB. Inference efficiency remains comparable: REST-SNN achieves a throughput of 72.77 FPS and a latency of 13.98ms, both comparable to GAC-SNN. These results indicate that the performance gains mainly come from the more stable spatiotemporal representation induced by the AI-LIF mechanism and theoretically guided dual regularization, rather than relying on increased model capacity or hardware costs.

[0092] In terms of temporal scalability and low-latency inference, the comparison of accuracy as a function of simulation time step T shows that REST-SNN reaches performance saturation earlier than GAC-SNN, indicating stronger temporal scalability under extremely short time windows. Especially with a very low time budget of T=2, REST-SNN achieves an accuracy of 70.92%, surpassing GAC-SNN's peak performance (66.97%) at T=8. This advantage remains consistent across different time windows (e.g., an improvement of +5.52% at T=2 and +5.26% at T=6). This demonstrates that REST-SNN can rapidly capture mission-critical spatiotemporal cues in the early stages of impulse integration, thereby achieving reliable inference with fewer simulation steps, which is particularly important for neuromorphic applications requiring low-latency deployment.

[0093] about The generality of this is also demonstrated by the experimental results, which show that the baseline setting... This allows for optimal overall performance. This choice is not arbitrary: robustness analysis shows that the upper bound decomposes into two structurally symmetrical components, namely... Scaling spatial sensitivity terms and by The scaled time drift term. After these two terms are normalized to a comparable scale, as shown... and As defined, a balanced weighting (i.e., a 1:1 ratio) becomes a theoretically sound default choice, simultaneously suppressing spatial magnification and temporal accumulation. In practice, ( , The absolute value of ) may still need slight adjustment due to differences in dataset noise intensity and time statistics, but the balance ratio provides a stable and transferable starting point and is consistent with the theoretical decomposition.

[0094] The above experiments were performed on an NVIDIA A100 GPU computing platform equipped with 80GB of video memory. The SGD optimizer was used during training with a momentum coefficient of 0.9. The learning rate used a cosine decay strategy, decreasing from 0.1 to 0 over 250 epochs. Weight decay was set according to the dataset: CIFAR-10 / 100 was [specific value missing]. ImageNet is .

[0095] To ensure a rigorous and fair evaluation and to distinguish the structural contributions of REST-SNNs from training techniques, all contrastive SNN baselines in this embodiment were reimplemented and evaluated under the same training protocol. Specifically, all contrastive models employed a completely consistent data augmentation strategy (i.e., CutMix) and used consistent alternative gradient approximations where applicable. By controlling these confounding variables, this embodiment ensures that the improvement in robust accuracy can be explicitly attributed to the structural stability provided by AI-LIF neurons and theoretically guided dual regularization, rather than differences in the training process.

[0096] Although this embodiment did not conduct explicit ablation experiments completely isolating SSR and TCR, the effectiveness of the dual regularization strategy can be inferred from the performance trends under different perturbation types. Specifically, the performance improvement in tests with geometric and occlusion-related distortions (such as translation and erasure perturbations) demonstrates that constraint space sensitivity can effectively mitigate the amplification effect of local perturbations, which is entirely consistent with the design purpose of SSR. Meanwhile, the performance gains in tests with random and appearance-related perturbations (such as color jitter and noise interference) indicate that the model's stability is significantly enhanced during the time accumulation process, which is consistent with the design goals of TCR. These experimental observations provide solid empirical evidence, fully demonstrating that the two regularization terms function in a complementary manner and are highly consistent with the aforementioned theoretical decomposition.

[0097] The theoretical analysis in this embodiment decomposes robustness degradation into two dominant factors: spatial sensitivity and temporal drift. Although the precise quantities in the upper bound are not easily measured directly during training, experimental results exhibit patterns consistent with this decomposition. In particular, the improved stability under multiple perturbations demonstrates that limiting attentional variations effectively reduces error amplification and accumulation. This consistency between theoretical insights and empirical performance supports the "from upper bound to target" design philosophy of this embodiment, which approximates the dominant terms in the theoretical upper bound with computable substitution regularization terms.

[0098] A key foundation of the training method provided in this embodiment is a unified robustness upper bound (Theorem 1), which depends on several regularity conditions, such as the local Lipschitz continuity of the attention map. The assumptions of continuous substitution gradients (K(·)) and bounded substitution gradients are important. While SSR and TCR actively optimize these conditions within the range of normal noise, these assumptions may still fail under extreme or adversarial perturbations. In the high-dimensional, non-convex optimization landscape of SNNs, the discontinuity of discrete pulses is a fundamental challenge. Severe noise injection can cause tiny continuous fluctuations in membrane potential to trigger abrupt changes in the macroscopic pulse sequence, i.e., catastrophic binary flips. In such boundary cases, continuous substitution gradients K(·) fail to accurately reflect real discrete jumps, potentially leading to an instantaneous increase in the empirical Lipschitz constant. Although AI-LIF partially smooths the landscape through integer activations, this phenomenon highlights the inherent tension between continuous optimization theory and discrete neuromorphic dynamics.

[0099] Traditional experience in deep learning often suggests that improving robustness requires larger model capacity, redundant parameterization, or computationally expensive defense mechanisms. A key insight from the experiments in this example is that REST-SNN can, to some extent, decouple robustness from metabolic costs. Discharge rate analysis clearly shows that AI-LIF maintains nearly identical impulse statistics to the standard LIF baseline.

[0100] This indicates that the performance gains of the method in this embodiment do not come from higher levels of neuronal activity, because in neuromorphic hardware, neuronal activity directly translates into energy consumption; instead, the gains primarily come from more efficient representations and structurally stable dynamics. This supports an important viewpoint: the robustness of SNNs should be achieved through representation refinement and principled regularization, rather than solely relying on a brute-force increase in computational cost.

[0101] ImageNet results show that while REST-SNN may have lower clean accuracy than some attention-based SNN baselines, it achieves stronger robustness under multiple perturbations. This reflects the inherent trade-off in robustness-oriented model design: imposing stability constraints on attention dynamics can improve the model's resistance to distorted inputs, but may also reduce the model's flexibility in fitting clean data distributions.

[0102] For REST-SNN, this trade-off aligns with its objective of improving the structural stability of attention-modulated impulse dynamics, rather than simply maximizing clean data performance. For robust perception tasks in noisy or uncertain environments, such as real-time neuromorphic perception, this property is often more practically significant than the highest clean accuracy under ideal conditions.

[0103] This embodiment uses a frame image dataset instead of a native neuromorphic event stream (such as DVS) because a frame benchmark with synthetic perturbations provides a better "controlled environment" for preliminary verification of the theoretical robustness upper bound. Real-world event data often contains both spatial contamination and asynchronous temporal jitter, making it difficult for researchers to distinguish whether performance improvements come from spatial smoothness or temporal stability. By integrating the perturbated static frames over T time steps, this embodiment can isolate and precisely quantify the suppression of attention-induced temporal drift, which is the core term of the upper bound derived in this embodiment.

[0104] Nevertheless, the robustness upper bound of this embodiment characterizes the propagation of perturbations in recursive neural dynamics, a problem prevalent in various modalities processing spatiotemporal signals. Therefore, the core ideas of REST-SNN are expected to be naturally extended to a wider range of neuromorphic tasks. Further extending this framework to asynchronous event datasets is a crucial next step in evaluating its modal generalization robustness in original neuromorphic perception scenarios.

[0105] It is important to emphasize that the regularization term proposed in this embodiment does not directly minimize the theoretical robustness upper bound, but rather serves as an optimizable alternative objective, maintaining a monotonic relationship with the dominant factor in the upper bound. This design strikes a trade-off between precise optimization and stable, efficient training, which is particularly crucial for end-to-end training of high-dimensional SNNs.

[0106] REST-SNN is designed with size, weight, and power (SWaP) constraints in mind. Since spatial parameters remain constant and no regularization overhead is introduced during inference, this approach is largely compatible with standard SNN deployment procedures. Nevertheless, several open challenges remain. While virtual timestep unrolling preserves sparse accumulation-driven operations, scheduling and synchronization overhead may occur on some highly parallel neuromorphic chips within extremely long time windows (T×D). Finally, the attention mechanism itself still incurs fundamental computational overhead; future research could explore dynamically activating attention only during highly uncertain integration phases to further reduce the overall cost.

[0107] The REST-SNN provided in this embodiment enhances the robustness of spiking neural networks by combining integer neuron modeling with theoretically guided attention regularization. By establishing a connection between robustness analysis and optimization objectives, it provides a path for stabilizing attention-modulated pulse dynamics that is both theoretically sound and practically feasible.

[0108] Experiments on CIFAR and ImageNet perturbation benchmarks demonstrate that REST-SNN consistently improves robustness under multi-class distortion compared to strong SNN baselines, while maintaining similar parameter count, memory usage, and inference latency. On large-scale datasets, these robustness gains may be accompanied by a moderate tradeoff in clean accuracy, reflecting the robustness-oriented design of REST-SNN. Overall, the results suggest that SNN robustness can be enhanced through structured regularization and representation refinement without relying on increased model complexity.

[0109] Data processing method example: Based on the same inventive concept, embodiments of the present invention also provide a data processing method for spiking neural networks based on adaptive integer neurons, such as... Figure 3 As shown, the method includes: Step S10: Acquire the image to be classified and input it into a pre-trained spiking neural network; the spiking neural network is configured with adaptive integer leakage current integral firing neurons.

[0110] Step S20: At each time step, the adaptive integer leakage current integral firing neuron updates the membrane potential of the current time step based on the received synaptic input of the current time step and the residual state of the previous time step; obtains the residual state of the previous time step and calculates the dynamic discharge threshold of the current time step; based on the dynamic discharge threshold, the membrane potential is converted to a bounded range to generate a bounded integer pulse; based on the membrane potential of the current time step and the bounded integer pulse, and combined with the dynamic discharge threshold, the residual state of the current time step is soft-reset updated; the bounded integer pulse is expanded into a binary pulse sequence under several virtual time steps according to the maximum discharge capacity, and the synaptic weights of the next layer network are sparsely accumulated with the binary pulse sequence, and the result of the sparse accumulation operation is used as the synaptic input of the next layer network.

[0111] Step S30: Generate a spatial attention map for the current time step based on the spatiotemporal attention mechanism, and use the spatial attention map to modulate the bounded integer pulse to obtain the modulation output features for the current time step; obtain the classification result of the image to be classified based on the modulation output features of each time step.

[0112] Since the implementation process of each step in this processing method has been described in detail in the above training method embodiments, it will not be repeated here.

[0113] It should be noted that the above embodiments are only used to illustrate the technical concept and technical solution of the present invention, and are not intended to limit the scope of protection of the present invention. Any modifications or partial improvements made to the technical solution of the present invention by those skilled in the art without departing from the inventive concept of the present invention, or equivalent substitutions or transformations of some technical features, shall fall within the scope of protection defined by the present invention.

Claims

1. A training method for a spiking neural network based on adaptive integer neurons, characterized in that, The method includes: Sample images and their true labels are acquired and input into a spiking neural network to be trained. The spiking neural network is configured with adaptive integer leakage current integral firing neurons. At each time step, the adaptive integer leakage current integral firing neuron updates the membrane potential of the current time step based on the received synaptic input of the current time step and the residual state of the previous time step; it obtains the residual state of the previous time step and calculates the dynamic discharge threshold of the current time step, and converts the membrane potential to a bounded range based on the dynamic discharge threshold to generate a bounded integer pulse; based on the membrane potential of the current time step and the bounded integer pulse, and in combination with the dynamic discharge threshold, it performs a soft reset update on the residual state of the current time step. A spatial attention map for the current time step is generated based on a spatiotemporal attention mechanism. The bounded integer pulse is modulated using the spatial attention map to obtain the modulation output characteristics of the current time step. The modulation output features at each time step are mapped to corresponding logical values. The task loss is constructed based on the logical values ​​at each time step and the true labels. The spiking neural network is then optimized and updated end-to-end based on the task loss.

2. The method for training a spiking neural network based on adaptive integer neurons according to claim 1, characterized in that, The membrane potential at the current time step is updated using the following formula: in, Indicates the current time step The membrane potential; Indicates the previous time step The residual state; Indicates the current time step received. Synaptic input; Indicates from the previous level The first neuron is connected to the current layer. Synaptic weights of individual neurons; Indicates the current layer number Bias of each neuron; Indicates the current time step No. The pulse input to a presynaptic neuron.

3. The method for training a spiking neural network based on adaptive integer neurons according to claim 1, characterized in that, Obtain the residual state from the previous time step and calculate the dynamic discharge threshold for the current time step. The corresponding calculation formula is as follows: = in, Indicates the current time step The dynamic discharge threshold; This indicates the preset initial reference discharge threshold; This represents the adaptive adjustment coefficient; Indicates the previous time step The residual state.

4. The method for training a spiking neural network based on adaptive integer neurons according to claim 1, characterized in that, The formula for generating bounded integer impulses is: in, Indicates the current time step Bounded integer pulses; Indicates the current time step The membrane potential; Indicates the current time step The dynamic discharge threshold; This indicates the preset maximum discharge capacity; This represents the floor function; This represents a truncation function used to restrict the input value to a certain range. Within the range.

5. The method for training a spiking neural network based on adaptive integer neurons according to claim 1, characterized in that, The residual state at the current time step is updated using a soft reset, and the corresponding calculation formula is as follows: in, Indicates the current time step The residual state; Indicates the current time step The membrane potential; Indicates the leakage factor. ; Indicates the current time step Bounded integer pulses; Indicates the current time step The dynamic discharge threshold.

6. The method for training a spiking neural network based on adaptive integer neurons according to claim 1, characterized in that, Based on the task loss, the spiking neural network is optimized and updated end-to-end, including: Determine the spatial regularization term used to constrain the spatial attention map in the spatial dimension Jacobian gradient, and the temporal regularization term used to constrain the difference between spatial attention maps in adjacent time steps, respectively. The spatial regularization term and the temporal regularization term are weighted based on preset weights and fused with the task loss to construct an overall loss function. Based on the overall loss function, the spiking neural network is optimized and updated end-to-end.

7. The method for training a spiking neural network based on adaptive integer neurons according to claim 6, characterized in that, The spatial regularization term used to constrain the Jacobian gradient of the spatial attention map is determined, and the corresponding calculation formula is as follows: in, Represents the space regularization term; Indicates input At the current time step Spatial attention map; This indicates that the attention map is applied to the input. Jacobian matrix; Denotes the Frobenius norm; This represents the shear threshold used to prevent gradient explosion; This represents a function that takes the minimum value. This represents the total number of time steps.

8. The method for training a spiking neural network based on adaptive integer neurons according to claim 6, characterized in that, The time regularization term used to constrain the differences in spatial attention maps between adjacent time steps is determined by the following formula: in, Represents the time regularization term; and These represent the input samples respectively. At the current time step and the next time step Spatial attention map; Represents the L2 norm; This represents the total number of time steps.

9. The method for training a spiking neural network based on adaptive integer neurons according to claim 1, characterized in that, The end-to-end optimization and update of the spiking neural network includes: When calculating the gradient of the network loss with respect to the membrane potential based on the task loss during backpropagation, a preset surrogate gradient function is used to replace the gradient of the membrane potential rounding operation. When the membrane potential at the current time step is greater than or equal to zero and less than or equal to the product of the maximum discharge capacity and the dynamic discharge threshold at the current time step, the gradient of the bounded integer pulse with respect to the membrane potential at the current time step is replaced with the reciprocal of the dynamic discharge threshold at the current time step; otherwise, the gradient of the bounded integer pulse with respect to the membrane potential at the current time step is replaced with zero.

10. A data processing method for a spiking neural network based on adaptive integer neurons, characterized in that, The method includes: The image to be classified is acquired and input into a pre-trained spiking neural network; the spiking neural network is configured with adaptive integer leakage current integral firing neurons; At each time step, the adaptive integer leakage current integral firing neuron updates the membrane potential of the current time step based on the received synaptic input of the current time step and the residual state of the previous time step; it obtains the residual state of the previous time step and calculates the dynamic discharge threshold of the current time step, and converts the membrane potential to a bounded range based on the dynamic discharge threshold to generate a bounded integer pulse; based on the membrane potential of the current time step and the bounded integer pulse, and combined with the dynamic discharge threshold, it performs a soft reset update on the residual state of the current time step; it expands the bounded integer pulse into a binary pulse sequence under several virtual time steps according to the maximum discharge capacity, and performs a sparse accumulation operation on the synaptic weights of the next layer network and the binary pulse sequence, and uses the result of the sparse accumulation operation as the synaptic input of the next layer network; A spatial attention map for the current time step is generated based on a spatiotemporal attention mechanism. The bounded integer pulse is modulated using the spatial attention map to obtain the modulation output features for the current time step. The classification result of the image to be classified is obtained based on the modulation output features for each time step.