Method and system for improving image classification performance by using spiking neural network
By embedding the Inception module in the pulsed neural network and using triple STDP rules for training, combined with simulated annealing optimization parameters, the problems of SNN's poor performance and slow convergence speed in image classification tasks are solved, achieving higher accuracy and robustness.
Patent Information
- Application Number
- CN202510426678.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-17
AI Technical Summary
The existing pulsed neural networks have poor performance in image classification tasks, slow convergence speed, and poor design of network structures for unsupervised learning, resulting in poor robustness of the results.
By improving the SNN network structure, the Inception module is embedded to improve network performance, so that the SNN can synchronously process input information at different abstract levels. Then, the network is trained using triple STDP unsupervised learning rules, and the parameter space is optimized through STDP fusion simulation annealing to achieve parameter fine-tuning.
It significantly improves the accuracy and convergence speed of image classification, improves the robustness and performance of the network, and is specifically manifested in the increase of the accuracy rate from 77.79% to 89.98%.
Smart Images

Figure CN120164045A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of spiking neural networks, and particularly to a method and system for improving image classification performance using spiking neural networks. Background Art
[0002] Artificial Neural Network (ANN) has achieved good performance results in many cognitive tasks (such as recognition, analysis, and reasoning). However, ANN is computationally intensive and requires a large amount of computing power during operation. Different from ANN that uses static and continuously activated neurons, Spiking Neural Network (SNN) calculates and transmits information by using discrete spikes, and its computing mechanism is similar to the working mechanism of the brain in biology. Since spike data can use event-driven computing methods, the power consumption and computing latency of SNN can be greatly reduced, significantly reducing the computing cost.
[0003] Existing SNN training methods can be divided into three categories: supervised, unsupervised, and ANN conversion. Supervised / unsupervised means training the SNN with / without label information. Specifically, supervised SNN requires labeled data and uses a loss function to guide its training process, aiming to reduce the difference between the output and the label; while unsupervised SNN adjusts its synaptic weights based on the theory of synaptic plasticity in biology, such as Spike-timing-dependent Plasticity (STDP) and its variant learning rules, and adjusts the synaptic weights according to the spike firing patterns of local neurons, so as to learn the laws of spike firing patterns. ANN conversion refers to an algorithm that converts a pre-trained ANN into an SNN to avoid directly training the SNN. However, the current supervised training methods have the following problems: 1. The non-differentiable nature of neuron spikes makes it difficult for the loss function to backpropagate during supervised learning; 2. Training SNN through supervised or ANN conversion requires labeled data, but the acquisition cost of such data is very high.
[0004] For the above reasons, the training method of unsupervised learning gradually shows advantages. However, most SNNs using unsupervised learning mostly adopt a fully connected method, and the network structure is not well designed, which often leads to convergence to a suboptimal result in the end, and the convergence speed of the network training process is extremely slow, and the final network training result has poor robustness. Summary of the Invention
[0005] This application provides a method and system for improving image classification performance using spiking neural networks, which can be used to solve the technical problems of poor network performance and slow convergence speed.
[0006] This application provides a method for improving image classification performance using spiking neural networks. The method includes:
[0007] Step 1: Improve the common SNN network structure by embedding the Inception module to enhance network performance, enabling the SNN to synchronously process input information at different abstraction levels;
[0008] Step 2: Train the network using the triplet STDP unsupervised learning rule;
[0009] Step 3: Co-optimize the parameter space through the fusion of STDP and simulated annealing to achieve parameter fine-tuning: Search the parameter space using the SA algorithm, estimate the performance of the search results using the evaluation algorithm, and combine with the STDP training algorithm to achieve fine-tuning of the network parameters;
[0010] Step 4: Substitute the above parameter optimization results into the network for retraining and complete the analysis of the classification results.
[0011] Furthermore, in Step 1: This step improves the common SNN network structure by embedding the Inception module to enhance network performance, enabling the SNN to synchronously process input information at different abstraction levels. Most common SNN network structures use global fully connected and local fully connected structures, which require high computing power in most cases and have many redundant neuron synaptic connections, resulting in limited performance with the same number of parameters. The Inception module designed in this application can well solve this problem, and the specific steps are as follows:
[0012] Step 1.1: Determine that the neuron model used in the network is the Leaky Integrate and Fire (LIF) model. The dynamic change of the neuron cell membrane potential is as follows:
[0013]
[0014] where τ m is the neuron cell membrane time constant, I syn (t) is the neuron synaptic current at time t, V is the neuron cell membrane potential, and V rest is the neuron resting potential. In actual neuron simulation, the neuron uses a subtractive soft reset mechanism, subtracting only the firing threshold instead of hard resetting to the reset potential after pulse firing, and uses the Euler method to simulate the above differential equation:
[0015]
[0016] where Δt is the simulation time step, which needs to be balanced between simulation accuracy and simulation step size, usually at the millisecond level. In the implementation of this application, 1 millisecond is used as the simulation step size, τm is the time constant of the neuron cell membrane, I syn (t) is the synaptic current of the neuron at time t, V is the potential of the neuron cell membrane, V rest is the resting potential of the neuron, V threshold is the threshold potential for neuron spike firing, S(t) is the neuron spike firing, and the corresponding spike firing conditions are as follows:
[0017]
[0018] where V(t) is the potential of the neuron cell membrane, V threshold is the threshold potential for neuron spike firing. Using subtractive soft reset can often retain more temporal information of the neuron membrane potential.
[0019] Step 1.2: The core component used is the spiking Inception module integrated with multi-scale convolution:
[0020] M inception = Concat(F 1×1 F 3×3 F 5×5 F pool )
[0021] F 1×1 = σ spike (W 1×1 * X)
[0022] F 3×3 = σ spike (W 3×3 * Reduce(X))
[0023] F 5×5 = σ spike (W 5×5 * Reduce(x))
[0024] F pool = σ spike (W pool * MaxPool(X))
[0025] where M inception is the Inception module, F n×n represents the convolutional layer with a convolutional kernel size of n×n, F pool represents the pooled feature, X is the input feature map, W n×n represents the n×n convolutional kernel, * represents the convolution operation, Reduce() is the dimensionality reduction function to ensure controllable overall computational complexity, MaxPool() is the max pooling function, and the operator σ spike is the spiking activation function: where Vthreshold is the neuron firing threshold; the Inception module can enhance the network's ability to capture features at different scales simultaneously.
[0026] Further, step 2: Train the network using the triplet STDP unsupervised learning rule, and the specific steps are as follows:
[0027] Step 2.1: Randomly initialize the network weights through a probability distribution: where represents the initial connection weight from neuron i to neuron j, represents a normal distribution with mean μ and variance σ 2 ; Random initialization ensures the diversity of the starting points in the weight space and avoids symmetry problems during the training process;
[0028] Step 2.2: Adjust the synaptic weights according to the temporal relationship between pre- and post-synaptic spikes through the STDP rule:
[0029]
[0030] where, represents the change in synaptic weight caused by the STDP rule, Δt = t j - t i represents the pre-synaptic spike time t i and the time stamp difference between the post-synaptic spike time t j , A + , A - are the synaptic potentiation and synaptic depression amplitudes respectively, τ + , τ - are two decay constants that control the time window of the learning rule; The STDP rule follows the Hebbian learning rule, that is, neurons that fire synchronously establish strong connections, and neurons that are asynchronously active weaken connections;
[0031] Step 2.3: Expand the triplet STDP to obtain:
[0032]
[0033] where, is the change in synaptic weight caused by the triplet STDP rule, represents the change in synaptic weight caused by the STDP rule, η is the learning rate, r i , r j represent the pre- / post-synaptic neuron firing frequencies respectively, r k represents the firing frequency of the regulatory neuron, that is, the interneuron or the cross-layer neuron; The triplet term can capture the high-order temporal correlations between spikes, enabling the network to learn complex spike patterns that cannot be captured by standard STDP.
[0034] Step 2.4: Track the changes in the traces of pre- and post-synaptic pulses:
[0035]
[0036]
[0037] where P pre (t), P post (t) are the traces of pre- and post-synaptic pulses at time t, S i (t), S j (t) are binary pulse events at time t, which is 1 when a pulse occurs and 0 otherwise; η pre , η post are the scaling coefficients of the traces, τ pre , τ post are the event constants of the exponential decay of the traces; the traces are used to indicate the recent pulse records of neurons and serve as eligibility traces in neuronal plasticity. At the same time, it is necessary to analyze the population activity of neurons to obtain the firing pattern and receptive field of neurons, which helps to enhance the selectivity of neurons:
[0038]
[0039] where A pattern is the population activity pattern of neurons, S(t) is the neuronal pulse firing, as described in Step 1.1, V(t) is the neuronal output, and T is the total simulation time.
[0040] Furthermore, Step 3: Since the firing behavior of neurons is jointly determined by multiple hyperparameters, different hyperparameter combinations will result in completely different dynamic responses of neurons, thus affecting the performance of the entire network. Therefore, Step 3 jointly optimizes the parameter space through STDP and simulated annealing to achieve parameter fine-tuning; searches the parameter space through the SA algorithm, and estimates the performance of the search results using a small sample through the STDP training algorithm, so as to achieve fine-tuning of the network parameters. The specific steps are as follows:
[0041] Step 3.1: Establish a complete parameter space Θ, covering all trainable elements of the network:
[0042] Θ = {τ + , τ - , A + , A - , V threshold , V reset , τ mem}
[0043] where τ + , τ -are the synaptic potentiation and synaptic depression time constants of STDP, respectively, A + , A - are the magnitudes of synaptic potentiation and synaptic depression of STDP, respectively, A threshold , A reset is the neuron firing threshold and reset potential, τ mem is the membrane potential time constant; the parameter space is optimized by SA:
[0044]
[0045] where, E new is the error of the new parameter combination, E current is the error of the current parameter combination, T is the temperature parameter; the error function E measures the performance of the network and will be described in detail in step 3.2.
[0046] Balance the parameter search range and search complexity by planning T: T(k) = T0·a k , where T0 is the initial temperature, α is the cooling coefficient, and k is the number of iterations; when the temperature is high, the algorithm freely explores the parameter space and accepts sub-optimal solutions to jump out of the local optimum. When the temperature decreases, the algorithm converges to the optimal solution; the parameters to be evaluated are generated by The higher the temperature, the larger the search space, where Θ k , Θ k+1 are the parameter combinations at the k-th and (k + 1)-th iterations, and ΔΘ k is the parameter perturbation at the k-th iteration;
[0047] Step 3.2: Bring the hyperparameters into the network and use a small sample for training and testing to obtain an estimated value of the network accuracy as an estimation result of the performance of the hyperparameter combination. Since the sample size needs to be balanced between running efficiency and estimation accuracy, this application uses 5000 samples as the estimation training set and 500 samples as the estimation test set. Use the new hyperparameters to perform STDP training on the network using the estimation training set, and then use the estimation test set to evaluate the results:
[0048]
[0049] where, E is the network performance evaluation function, is the network classification result, and y is the true data label.
[0050] Further, Step 4: Substitute the above parameter optimization results into the network for retraining and complete the classification result parsing. Since the network performance evaluation in the SA process uses the small-sample estimation method, the network is not fully trained. Therefore, substitute the hyperparameters finally optimized by SA into the network for retraining to obtain the final network result. At the same time, since the network needs to convert the pulse output result into a classification result, a corresponding conversion algorithm is required. The specific steps are as follows:
[0051] Step 4.1: Use STDP to train the network, and the process is as described in Step 2.
[0052] In traditional SNNs, a single neuron often corresponds to a fixed category. When converting pulses into category results, the pulses emitted by each neuron only vote for the category to which it is bound, and the category with the highest pulse count is statistically obtained as the classification result. The results of this type of method depend on the response intensity of neurons to different categories and are less effective when the category distribution of training samples is unbalanced. In this application, a voting layer is added in the inference stage to allow each neuron in the output layer to vote for all categories, thus avoiding the binding of neurons in the output layer to a single category. The neurons in the voting layer use the passive cell membrane model:
[0053]
[0054] Among them, V represents the neuron membrane potential, and w i is the synaptic connection strength of the output layer neuron, and S i represents the pulse output of the output layer neuron. For the i-th output layer neuron and the j-th voting layer neuron, the synaptic weight w ij is:
[0055]
[0056] Among them, S ij is the average number of pulses emitted by the i-th output layer neuron to the j-th voting layer neuron, and C is the total number of categories. Finally, select the category represented by the voting layer neuron with the highest membrane potential as the classification result.
[0057] This application also provides a system for improving image classification performance using a spiking neural network. This system is used to implement the method provided by this application. The system includes:
[0058] A network construction module for constructing network neurons and network structures, generating corresponding neuron cell models and instantiating corresponding neuron cell populations, and generating corresponding neuron synaptic connections;
[0059] A training module that writes the triple STDP learning rule into the synaptic dynamic model to achieve unsupervised learning of the network;
[0060] The parameter adjustment module is used to jointly optimize the parameter space through STDP fusion and simulated annealing to achieve parameter fine-tuning: search the parameter space through the SA algorithm, use the evaluation algorithm to estimate the performance of the search results, and combine the STDP training algorithm to achieve fine-tuning of the network parameters;
[0061] The fusion module brings the above parameter optimization results into the network for retraining and completes the analysis of the classification results.
[0062] In this application, by embedding the Inception module into the traditional SNN model and training through triple STDP, an accuracy rate of 95.88% is achieved; by further optimizing the network hyperparameters using the SA algorithm, the network convergence speed is greatly improved, from the accuracy rate of 77.79% of the original network to 89.98% in 10,000 iterations. Brief Description of the Drawings
[0063] Figure 1 It is a schematic diagram of the overall processing flow;
[0064] Figure 2 It is a schematic diagram of the network training process;
[0065] Figure 3 It is a diagram of the neuron connection strength matrix;
[0066] Figure 4 It is a diagram of the change in neuron activity intensity during the training process;
[0067] Figure 5 It is a diagram of the network classification accuracy rate during the training process;
[0068] Figure 6 It is a diagram of the network classification accuracy rate (moving average) during the training process;
[0069] Figure 7 It is a diagram of the average firing activity of each neuron;
[0070] Figure 8 It is a confusion matrix diagram of the training results of the present invention;
[0071] Figure 9 It is an operation example diagram of the present invention on the MNIST dataset;
[0072] Figure 10 It is an operation example diagram of the present invention on the EMNIST dataset;
[0073] Figure 11 It is an operation example diagram of the present invention in large-scale prediction;
[0074] Figure 12 It is an operation example diagram of the present invention on the USPS dataset;
[0075] Figure 13 This is a running example diagram of the present invention on the Fashion-MNIST dataset. Detailed implementation manners
[0076] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the implementation manners of the present application in detail with reference to the accompanying drawings.
[0077] The unsupervised SNN based on synaptic plasticity is similar to the learning process in biology and is more biologically reasonable; therefore, the present invention uses unsupervised classification for handwritten image classification.
[0078] In the handwritten digit classification task, five datasets are used: MNIST, USPS, Fashion MNIST, EMNIST, and SVHN. These standard datasets facilitate comparison with existing benchmarks. The present invention aims to improve the classification accuracy while minimizing the computational overhead of training spiking neural networks.
[0079] The preprocessing pipeline is applied to all datasets to ensure consistency and optimize the SNN performance. It includes removing artifacts from the images. Converting the preprocessed images with pulse coding into a pulse-based representation suitable for SNN processing. Training and testing are split at an 80 / 20 ratio respectively. The dataset division ensures a consistent class distribution between the two sets, which is crucial for accurate performance evaluation.
[0080] The SNN has good performance and efficiency in pattern recognition, especially for the handwritten digit recognition task. The SNN has advantages such as biological rationality (simulating the information processing of the human brain) and high energy efficiency. These characteristics make the SNN very suitable for our classification task, especially in resource-constrained environments. Initially, the model training was carried out on a local desktop (i7-11700K, 32GB RAM, RTX 3070). However, due to the computational requirements of larger datasets and the need for efficient storage, the project transitioned to the AutoDL platform (RTX 3090, 14 CPU cores), which provides superior resources and a dedicated framework. The AutoDL environment uses PyTorch 1.9.0, CUDA 11.3, and BindsNET 0.2.9. The reliability of the present invention is addressed through comprehensive documentation, data anonymization, secure storage, and fairness evaluation.
[0081] The present application provides a method for improving the image classification performance using a spiking neural network. The method includes:
[0082] Step 1: Improve the common SNN network structure to enhance network performance by embedding the Inception module, enabling the SNN to synchronously process input information at different abstraction levels;
[0083] Step 2: Train the network using the triplet STDP unsupervised learning rule;
[0084] Step 3: Co-optimize the parameter space through the combination of STDP and simulated annealing to achieve parameter fine-tuning: Search the parameter space using the SA algorithm, estimate the performance of the search results using the evaluation algorithm, and combine with the STDP training algorithm to achieve fine-tuning of the network parameters;
[0085] Step 4: Substitute the above parameter optimization results into the network for retraining and complete the parsing of the classification results.
[0086] Furthermore, in Step 1: This step improves the common SNN network structure to enhance network performance by embedding the Inception module, enabling the SNN to synchronously process input information at different abstraction levels. Most common SNN network structures use global fully connected and local fully connected structures, which generally require high computing power in most cases and have many redundant neuron synapse connections, resulting in limited performance with the same number of parameters. The Inception module designed in this application can well solve this problem, and the specific steps are as follows:
[0087] Step 1.1: Determine that the neuron model used in the network is the Leaky Integrate and Fire (LIF) model. The dynamic change of the neuron cell membrane potential is as follows:
[0088]
[0089] Among them, τ m is the neuron cell membrane time constant, I syn (t) is the neuron synapse current at time t, V is the neuron cell membrane potential, and V rest is the neuron resting potential. When actually simulating neurons, the neuron uses a subtractive soft reset mechanism, subtracting only the firing threshold instead of hard resetting to the reset potential after the pulse is emitted, and uses the Euler method to simulate the above differential equation:
[0090]
[0091] Among them, Δt is the simulation time step, which needs to be balanced between simulation accuracy and simulation step size, usually at the millisecond level. For example, in the implementation of this application, 1 millisecond is used as the simulation step size, τ m is the neuron cell membrane time constant, I syn(t) is the synaptic current of the neuron at time t, V is the membrane potential of the neuron, and V rest is the resting potential of the neuron, and V threshold is the threshold potential for neuron spike firing. S(t) represents neuron spike firing, and the corresponding spike firing conditions are as follows:
[0092]
[0093] where V(t) is the membrane potential of the neuron, and V threshold is the threshold potential for neuron spike firing. Using subtractive soft reset can often retain more temporal information of the neuron membrane potential.
[0094] Step 1.2: The core component used is the spiking Inception module integrated with multi-scale convolution:
[0095] M inception = Concat(F 1×1 F 3×3 F 5×5 F pool )
[0096] F 1×1 = σ spike (W 1×1 * X)
[0097] F 3×3 = σ spike (W 3×3 * Reduce(X))
[0098] F 5×5 = σ spike (W 5×5 * Reduce(x))
[0099] F pool = σ spike * W pool * MaxPool(X))
[0100] where M inception is the Inception module, F n×n represents the convolutional layer with a convolutional kernel size of n × n, F pool represents the pooled features, X is the input feature map, W n×n represents the n × n convolutional kernel, * represents the convolution operation, Reduce() is a dimensionality reduction function to ensure controllable overall computational complexity, MaxPool() is the max pooling function, and the operator σ spike is the spiking activation function: where V threshold is the neuron firing threshold; this module can enhance the network's ability to capture features at different scales simultaneously.
[0101] Further, Step 2: Train the network using the triple STDP unsupervised learning rule. The specific steps are as follows:
[0102] Step 2.1: Randomly initialize the network weights according to a probability distribution: where represents the initial connection weight from neuron i to neuron j, represents a normal distribution with mean μ and variance σ 2 ; Random initialization ensures the diversity of the starting points in the weight space and avoids symmetry problems during the training process;
[0103] Step 2.2: Adjust the synaptic weights according to the temporal relationship between pre- and post-synaptic spikes using the STDP rule:
[0104]
[0105] where represents the change in synaptic weight caused by the STDP rule, Δt = t j - t i represents the time stamp difference between the pre-synaptic spike time t i and the post-synaptic spike time t j , A + , A - are the synaptic potentiation and synaptic depression amplitudes respectively, τ + , τ - are two decay constants that control the time window of the learning rule; The STDP rule follows Hebb's learning law, that is, neurons that fire synchronously establish strong connections, and neurons with asynchronous activities weaken connections;
[0106] Step 2.3: Expand the triple STDP to obtain:
[0107]
[0108] where is the change in synaptic weight caused by the triple STDP rule, represents the change in synaptic weight caused by the STDP rule, η is the learning rate, r i , r j represent the pre- and post-synaptic neuron firing frequencies respectively, r k represents the firing frequency of the regulatory neuron, i.e., the interneuron or the cross-layer neuron; The triple term can capture the high-order temporal correlations between spikes, enabling the network to learn complex spike patterns that cannot be captured by standard STDP.
[0109] Step 2.4: Track the changes in the traces of pre- and post-synaptic spikes:
[0110]
[0111] Among them, P pre *t(,P post (t) is the trace of the pre- and post-synaptic pulses at time t, S i (t), S j (t) is the binary pulse event at time t, which is 1 when a pulse occurs and 0 otherwise; η pre , η post is the scaling coefficient of the trace, τ pre , τ post is the event constant of the exponential decay of the trace; the trace is used to indicate the most recent pulse record of the neuron and serves as the eligibility trace in neuronal plasticity. At the same time, it is necessary to analyze the population activity of neurons to obtain the firing pattern and receptive field of neurons, which helps to enhance the selectivity of neurons:
[0112]
[0113] Among them, A pattern is the population activity pattern of neurons, S(t) is the neuronal spike firing, as described in Step 1.1, V(t) is the neuronal output, and T is the total simulation time.
[0114] Furthermore, Step 3: Since the firing behavior of neurons is jointly determined by multiple hyperparameters, different combinations of hyperparameters will result in completely different dynamic responses of neurons, thus affecting the performance of the entire network. Therefore, this step jointly optimizes the parameter space through STDP and simulated annealing to achieve parameter fine-tuning; searches the parameter space through the SA algorithm, and estimates the performance of the search results using a small sample through the STDP training algorithm, so as to achieve fine-tuning of the network parameters. The specific steps are as follows:
[0115] Step 3.1: Establish a complete parameter space Θ, covering all trainable elements of the network:
[0116] Θ = {τ + , τ - , A + , A - , V threshold , V reset , τ mem}
[0117] Among them, τ + , τ - are the synaptic potentiation and synaptic depression time constants of STDP respectively, A + , A - are the amplitudes of synaptic potentiation and synaptic depression of STDP respectively, A threshold , A resetis the neuron firing threshold and reset potential, τ mem is the membrane potential time constant; optimize this parameter space through SA:
[0118]
[0119] where E new is the error of the new parameter combination, E current is the error of the current parameter combination, T is the temperature parameter; the error function E measures the performance of the network and is described in detail in step 3.2.
[0120] Balance the parameter search range and search complexity by planning T: T*k) = T0·a k , where T0 is the initial temperature, α is the cooling coefficient, and k is the number of iterations; when the temperature is high, the algorithm freely explores the parameter space and accepts suboptimal solutions to jump out of local optima. When the temperature decreases, the algorithm converges to the optimal solution; the parameters to be evaluated are generated through The higher the temperature, the larger the search space, where Θ k , Θ k+1 are the parameter combinations at the kth and k + 1th iterations, and ΔΘ k is the parameter perturbation at the kth iteration;
[0121] Step 3.2: Bring the hyperparameters into the network and use small samples for training and testing to obtain an estimated value of the network accuracy as an estimate of the performance of the hyperparameter combination. Since the sample size needs to be balanced between running efficiency and estimation accuracy, this application uses 5000 samples as the estimation training set and 500 samples as the estimation test set. Use the new hyperparameters to perform STDP training on the network using the estimation training set, and then use the estimation test set to evaluate the results:
[0122]
[0123] where E is the network performance evaluation function, is the network classification result, and y is the true data label.
[0124] Furthermore, step 4: Bring the above parameter optimization results into the network for retraining and complete the classification result analysis. Since the network performance evaluation in the SA process uses a small sample estimation method, the network has not been fully trained. Therefore, bring the hyperparameters finally optimized by SA into the network for retraining to obtain the final network result. At the same time, since the network needs to convert the pulse output result into a classification result, a corresponding conversion algorithm is required. The specific steps are as follows:
[0125] Step 4.1: Use STDP to train the network, and the process is as described in Step 2.
[0126] Step 4.2: In traditional SNNs, a single neuron often corresponds to a fixed category. When converting spikes into category results, the spikes emitted by each neuron only vote for the category it is bound to, and the category with the highest spike count is statistically obtained as the classification result. The results of such methods depend on the response intensity of neurons to different categories and are less effective when the category distribution of training samples is unbalanced. In this application, a voting layer is added during the inference stage to allow each neuron in the output layer to vote for all categories, thus avoiding the binding of neurons in the output layer to a single category. The neurons in the voting layer use a passive cell membrane model:
[0127]
[0128] where V represents the neuron membrane potential, and w i is the synaptic connection strength of the neurons in the output layer, and S i represents the spike output of the neurons in the output layer. For the i-th neuron in the output layer and the j-th neuron in the voting layer, the synaptic weight w ij is:
[0129]
[0130] where S ij is the average number of spikes emitted by the i-th neuron in the output layer to the j-th neuron in the voting layer, and C is the total number of categories. Finally, the category represented by the neuron in the voting layer with the highest membrane potential is selected as the classification result.
[0131] This application also provides a system for improving image classification performance using a spiking neural network. The technical details of the system provided in this application are consistent with those of the method provided in this application, and will not be elaborated here one by one. Please refer specifically to the method provided in this application. This system is used to implement the method provided in this application, and the system includes:
[0132] A network construction module for improving the SNN network structure and embedding the Inception module to enable the SNN to synchronously process input information at different abstraction levels;
[0133] A training module that writes the triplet STDP learning rule into the synaptic dynamic model to achieve unsupervised learning of the network;
[0134] A parameter adjustment module for jointly optimizing the parameter space through STDP and simulated annealing to achieve parameter fine-tuning: searching the parameter space through the SA algorithm, using an evaluation algorithm to estimate the performance of the search results, and combining the STDP training algorithm to achieve fine-tuning of the network parameters;
[0135] The fusion module takes the parameter optimization results and retrains the network, and completes the parsing of the classification results.
[0136] The technical details of the system part in this application correspond one by one to the technical details in the method, which will not be elaborated here. For details, please refer to the technical details in the method part.
[0137] The present invention reveals the synergistic effect of SA and STDP, highlights the key role of fine parameter initialization in SNN training, contributes to promoting the hybrid learning strategy of neuromorphic computing, and reveals the key mechanism for optimizing the learning parameters and network architecture of spiking neural networks to improve image classification accuracy. In addition, by precisely exploiting the spike timing characteristics, the SNN proposed in the present invention outperforms traditional artificial neural networks trained based on backpropagation in low-bit precision scenarios, which is consistent with the research conclusion emphasizing the importance of SNN temporal coding for high-precision classification. The present invention demonstrates the great potential of pulse-based neuromorphic systems and opens up a new path for developing bio-inspired efficient computing paradigms.
[0138] The digit recognition process starts with preprocessing the input image through normalization and noise reduction techniques. Then, the preprocessed image is encoded into a spike train using temporal coding, where the pixel intensity is converted into the spike firing rate. These three methods work in parallel to extract complementary features, while triple STDP learning captures the temporal correlations in the spike patterns to learn the discriminative features of digits through a biologically plausible mechanism. Pulse Inception module: Processes spike training simultaneously at multiple scales, enabling the detection of digit features at different levels of abstraction. STDP and simulated annealing: Uses the STDP training results on small samples as an estimate of network performance, and optimizes the network hyperparameters using simulated annealing. Feature integration and classification integrate the features extracted from each method through a weighted combination layer. Then, the integrated features are passed to the classification layer, where the temporal integration of spikes occurs for the final digit prediction. Performance evaluation. The recognized digits are evaluated according to the ground truth using standard metrics such as accuracy, precision, recall, and F1 score to assess the performance of the combined method.
[0139] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0140] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0141] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for improving image classification performance using a pulse neural network, characterized in that: The method comprises: Step 1: Improve the SNN network structure and embed the Inception module so that the SNN can synchronously process input information at different abstraction levels; Step 2: Train the network using the triplet STDP unsupervised learning rule; Step 3: Optimize the parameter space by fusing STDP with simulated annealing to achieve parameter fine-tuning: Search the parameter space by using the SA algorithm, use the evaluation algorithm to estimate the performance of the search results, and combine the STDP training algorithm to achieve fine-tuning of the network parameters; Step 4: Use the parameter optimization results to retrain the network and complete the classification result analysis.
2. The method according to claim 1, characterized in that Step 1: Improve the SNN network structure and embed the Inception module to enable the SNN to simultaneously process input information at different abstraction levels; including: Step 1.1: Determine that the neuron model used by the network is the leaky integral-fire model; the dynamics of its neuron cell membrane potential are as follows: Among them, τ m is the neuronal cell membrane time constant, I syn (t) is the neuron synaptic current at time t, V is the neuron cell membrane potential, V rest is the resting potential of the neuron; in the actual neuron simulation, the neuron uses a subtraction soft reset mechanism, which only subtracts the firing threshold after the pulse is fired instead of hard resetting to the reset potential, and uses the Euler method to simulate the differential equation: Among them, Δt is the simulation time step, which is in milliseconds. 1 millisecond is used as the simulation step. τ m is the neuronal cell membrane time constant, I syn (t) is the neuron synaptic current at time t, V is the neuron cell membrane potential, V rest is the neuron resting potential, V threshold is the threshold potential of neuronal pulse emission, S(t) is the neuronal pulse emission, and the corresponding pulse emission conditions are as follows: Where V(t) is the neuron cell membrane potential, V threshold Threshold potential for neuronal pulse firing; Step 1.2: The core component used is the pulse Inception module integrated with multi-scale convolution: M inception =Concat(F 1×1 F 3×3 F 5×5 F pool ) F 1×1 =s spike (W 1×1 *X) F 3×3 =s spike (W 3×3 *Reduce(X)) F 5×5 =s spike (W 5×5 *Reduce(x)) F pool =s spike (W pool *MaxPool(X)) Among them, M inception is the Inception module, F n×n represents a convolutional layer with a kernel size of n×n, F pool represents the pooling feature, X is the input feature map, W n×n represents an n×n convolution kernel, * represents a convolution operation, Reduce() is a dimensionality reduction function, MaxPool() is a maximum pooling function, and the operator σ spike is the impulse activation function: Where V threshold is the neuron firing threshold.
3. The method according to claim 1, characterized in that Step 2: Train the network using the triplet STDP unsupervised learning rule: Step 2.1: Initialize the network weights randomly using probability distribution: in represents the initial connection weight from neuron i to neuron j, Indicates that the mean is μ and the variance is σ 2 Normal distribution of Step 2.2: Adjust the synaptic weight according to the timing relationship between pre- and post-synaptic pulses using the STDP rule: in, represents the change in synaptic weight caused by the STDP rule, Δt = t j -t i represents the presynaptic spike time t i and the postsynaptic spike time t j The timestamp difference, A + ,A - are the amplitudes of synaptic potentiation and synaptic depression, respectively, and τ + ,τ - The distribution is two decay constants that control the time window of the learning rule; the STDP rule follows the Hebbian learning rule, that is, neurons that fire synchronously establish strong connections, and neurons that fire asynchronously weaken their connections; Step 2.3: Expand the triplet STDP to obtain: in, is the change of synaptic weight caused by the triple STDP rule, represents the change of synaptic weight caused by the STDP rule, η is the learning rate, r i ,r j Respectively represent the firing frequency of presynaptic / postsynaptic neurons, r k It is expressed as the firing rate of the regulating neuron, i.e., the interneuron or the cross-layer neuron; the triple term can capture the high-order temporal correlations between spikes, allowing the network to learn complex pulse patterns that cannot be captured by standard STDP; Step 2.4: Track the changes in the traces of pre- and post-synaptic spikes: Among them, P pre (t),P post (t) is the trace of the presynaptic and postsynaptic pulses at time t, S i (t),S j (t) is the binary pulse event at time t, which is 1 when the pulse occurs and 0 otherwise; η pre ,η post is the scaling factor of the trace, τ pre ,τ post is the event constant of the exponential decay of the trace; the trace is used to indicate that the neuron is the most recent spike record, which serves as a qualifying trace in neuronal plasticity; at the same time, the activity of the neuronal population is analyzed to obtain the spike firing pattern and the receptive field of the neuron: Among them, A pattern is the activity pattern of the neuron group, S(t) is the neuron pulse emission, V(t) is the neuron output, and T is the total simulation time.
4. The method according to claim 1, characterized in that: Step 3: Optimize the parameter space by fusing STDP with simulated annealing to achieve parameter fine-tuning, including: Step 3.1: Establish a complete parameter space Θ, covering all trainable elements of the network: Θ={τ + ,t - ,A + ,A - ,V threshold ,V reset ,t mem } Among them, τ + ,τ - are the synaptic potentiation and synaptic depression time constants of STDP, A + ,A - are the prominent enhancement and synaptic depression amplitudes of STDP, A threshold ,A reset is the neuron firing threshold and reset potential, τ mem is the cell membrane potential time constant; the parameter space is optimized by SA: Among them, E new is the error of the new parameter combination, E current is the error of the current parameter combination, T is the temperature parameter; the error function E measures the performance of the network; Balance the parameter search range and search complexity by planning T: T(k) = T0·a k , where T0 is the initial temperature, α is the cooling coefficient, and k is the number of iterations; when the temperature is high, the algorithm freely explores the parameter space and accepts suboptimal solutions to jump out of the local optimal value. When the temperature drops, the algorithm converges to the optimal solution; the parameters to be evaluated are Generate, the higher the temperature, the larger the search space, where Θ k ,Θ k+1 is the parameter combination for k and k+1 iterations, ΔΘ k is the parameter perturbation at k iterations; Step 3.2: By bringing the hyperparameters into the network and using small samples for training and testing, we get an estimate of the network accuracy as the estimated result of the hyperparameter combination performance; we use 5000 samples as the estimated training set and 500 samples as the estimated test set; we use the new hyperparameters to perform STDP training on the estimated training set, and then use the estimated test set to evaluate the results: Among them, E is the network performance evaluation function, is the network classification result, and y is the real data label.
5. The method according to claim 1, characterized in that Step 4: Use the above parameter optimization results to retrain the network and complete the classification result analysis, including: Step 4.1: Train the network using STDP. Step 4.2: Add a voting layer in the inference phase to let each neuron in the output layer vote for all categories, thus avoiding the neurons in the output layer being bound to a single category; the voting layer neurons use a passive cell membrane model: Where V represents the neuron membrane potential, w i is the synaptic connection strength of neurons in the output layer, S i Represents the pulse output of the output layer neuron. For the output layer neuron number i and the voting layer neuron number j, the synaptic weights w ij for: Among them, S ij is the average number of pulses emitted by the output layer neuron No. i to the voting layer neuron No. j, and C is the total number of categories. Finally, the category represented by the voting layer neuron with the highest membrane potential is selected as the classification result.
6. A system for improving image classification performance using a spiking neural network, characterized in that: The system comprises: The network construction module is used to construct network neurons and network structures, generate corresponding neuron cell models and instantiate corresponding neuron cell populations, and generate corresponding neuron synaptic connections; The training module writes the triplet STDP learning rules into the synaptic dynamic model, thereby achieving unsupervised learning of the network; The parameter adjustment module is used to optimize the parameter space by fusing STDP with simulated annealing to achieve parameter fine-tuning: the parameter space is searched by the SA algorithm, the performance of the search results is estimated by the evaluation algorithm, and the network parameters are fine-tuned by combining the STDP training algorithm; The fusion module brings the above parameter optimization results into the network for retraining and completes the classification result analysis.
Citation Information
Cited By
Multi-modal image resolution fusion enhancement method based on brain inspiration
CN121724848A