Pulse Neural Network Training Method, Device and Storage Medium Based on Knowledge Transfer
Through a knowledge migration-based method, the knowledge of the teacher ANN network is transferred to the student SNN network, and the performance degradation and gradient instability problems in SNN training are solved by using LIF neurons and approximate gradient functions, and the low-power consumption performance with fast convergence and high precision are achieved.
Patent Information
- Application Number
- CN202310712420.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-06-15
AI Technical Summary
The existing pulsed neural network training methods have problems with the ANN-SNN conversion method leading to degradation of model performance and gradient instability of SNN direct training methods, and lack of general and efficient training solutions.
Using a knowledge transfer-based method, by constructing a teacher ANN network and a student SNN network, the error backpropagation algorithm is used to transfer the knowledge of the teacher network to the student network, and combining LIF neurons and approximate gradient functions, the intermediate layer and output layer loss functions are defined to realize gradient descent training.
There is no need to constrain the original ANN model, maintain its performance, and no large simulation time step is required, and rapid convergence is achieved to achieve low power consumption, high precision, and low latency performance of SNN on resource-constrained devices.
Smart Images

Figure CN116702865B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a training method for a spiking neural network based on knowledge transfer. Background Art
[0002] A spiking neural network (SNN) is a new generation of artificial neural network inspired by the biological brain and performing information processing based on event-driven sparse computing, and is called the third-generation neural network. Different from the traditional artificial neural network (ANN) that uses continuous real values as the carrier for information propagation, the spiking neural network simulates the way of transmitting electrical pulse signals between neurons in the human brain and uses discrete pulse signals as the carrier for information, so it has an information processing ability closer to the real brain. The sparse and powerful computing ability of SNN is expected to break through the energy limitation and computing power bottleneck faced by traditional ANN and achieve the task goals of low power consumption, high precision, and low latency.
[0003] Compared with the mature and perfect state of the ANN field, the research in the SNN field is still in the stage of rapid development, especially the research on its training algorithm. In the field of supervised learning of spiking neural networks, since the discrete pulse activation function cannot be differentiated, the gradient-based error backpropagation optimization algorithm cannot be directly applied to SNN. For this reason, researchers have found some solutions and achieved good results. These methods can generally be divided into two categories: one is the ANN-SNN conversion method. This method first trains a well-performing ANN, and then directly converts the trained ANN into an SNN by mapping the real activation values of the activation function to the pulse emission frequency of the SNN neurons. The advantage of this method is that there is no need to train an SNN from scratch, and the mature and complete training system of SNN can be directly used to directly obtain an SNN that can perform inference, and the converted SNN has a high inference accuracy; the other is the direct training method of SNN. This method usually uses an approximate gradient function to replace the non-differentiable pulse activation function in SNN so that the backpropagation algorithm can be directly applied to SNN. The advantage of this method is that there is no need for the assistance of ANN, and SNN can complete the training task independently, thus giving full play to the biological similarity and performance advantages of SNN.
[0004] Although the above methods have been used to obtain SNNs with good performance, the above two methods still have significant drawbacks: The ANN-SNN conversion method requires certain constraints on the original ANN, such as restricting the bias to zero, being unable to use batch normalization methods, and having to use average pooling instead of max pooling, etc. This will cause a partial decline in the model performance, and in order to maintain the accuracy of the mapping, this method requires a large simulation time step, resulting in a high inference latency; The disadvantage of the SNN direct training method is that the introduction of the approximate gradient function makes the gradient backpropagation unstable, and the gradient vanishing problem will occur when the pulse neural network is deeper, and the error caused by the approximate gradient function will accumulate layer by layer, resulting in a significant performance decline problem. To sum up, there is currently no general and efficient SNN training method. Summary of the Invention
[0005] Objective of the Invention: The objective of the present invention is to provide an SNN training method based on knowledge transfer, which transfers the learned representation ability and label knowledge of a well-trained ANN with good learning, representation, and generalization abilities to the SNN, thereby accelerating the convergence of the SNN, improving the inference accuracy of the SNN, and enabling the SNN to achieve low power consumption, high accuracy, and low latency performance when deployed on resource-constrained devices.
[0006] Technical Solution: The pulse neural network training method based on knowledge transfer of the present invention includes the following steps:
[0007] Step 1: Construct a data set D, and divide the data set D into a training set D Train and a test set D Test ;
[0008] Step 2: Construct a teacher ANN network, and use the error backpropagation algorithm to train the teacher ANN network on D Train to obtain the trained teacher network Net T , and save the weight parameters in the teacher network Net T ;
[0009] Step 3: Construct a pulse neural network SNN, denoted as the student network Net S , and the student network Net S adopts LIF neurons with pulse accumulation-emission and leakage characteristics, and introduces an approximate gradient function for the LIF neurons and define the loss function for training the student network Net S ;
[0010] Step 4: Define the intermediate layer loss function S and the output layer loss function for training the student network Net
[0011] Step 5: Transfer the learned knowledge of the teacher network Net T to the student network Net S The specific method is as follows: Perform forward propagation on the teacher network Net T and the student network Net S respectively to obtain the intermediate layer outputs and the output layer outputs of both; Calculate the intermediate layer loss function S and the output layer loss function of the student network Net Execute the gradient descent algorithm using the neural network optimizer to train the student network Net S . If the model of the student network Net S converges or reaches the maximum number of training epochs, then execute Step 6; otherwise, return to Step 5;
[0012] Step 6: Test the performance of the student network Net Test on the test set D S to obtain the trained spiking neural network SNN.
[0013] Furthermore, Step 2 specifically includes:
[0014] S210: Randomly initialize the weight parameters T of Net
[0015] w t ~RandomInitialization
[0016] S220: For each training sample (x i , y i ) ∈ D Train , calculate the prediction result of the sample:
[0017]
[0018] where x i represents the input sample, w t represents the model parameters, is the network prediction result;
[0019] S230: Calculate the loss between the prediction results of all training samples and the actual labels:
[0020]
[0021] where l is defined as the cross-entropy between two inputs;
[0022] S240: Calculate the gradient of the loss function L with respect to the model parameters w t :
[0023]
[0024] S250: Update the model parameters using the gradient descent optimization algorithm:
[0025]
[0026] where α is the learning rate, which controls the step size of each parameter update;
[0027] Repeat steps S220 - S250 until the model converges or reaches the predetermined number of training epochs, and save the model parameters θ of Net T for subsequent use in training Net T S S when used.
[0028] Furthermore, the LIF neuron adopted in step 3 in Net S is a common biological neuron model. LIF stands for "Leaky Integrate-and-Fire". It is a discrete-time model used to simulate how a neuron accumulates and processes information in response to external input signals. The LIF neuron model consists of the following parts:
[0029] Input current I: The neuron receives input current from other neurons or the environment.
[0030] Membrane potential U: The membrane potential of the neuron is an inherent property representing the internal electric potential of the LIF neuron. The input current affects the membrane potential of the neuron, causing it to gradually increase.
[0031] Threshold When the membrane potential of the neuron exceeds a specific threshold, the neuron emits an electrical pulse (action potential).
[0032] Resting potential U reset : After the neuron emits an electrical pulse, the membrane potential quickly resets to a lower potential, and the neuron will not respond to the input current for a period of time.
[0033] The inherent property of the membrane potential of the LIF neuron depends on the membrane potential at the previous moment and the current input. In the spatial domain, the LIF neuron receives and integrates the stimulation of the input current, which acts on the membrane potential U, causing an increase in U; in the time domain, based on the membrane potential at the previous moment and the decay constant β, the membrane potential U of the LIF neuron has a decay. After the combined effects in the spatial and time domains, if the membrane potential exceeds the pulse emission threshold the LIF neuron will generate a pulse, and the membrane potential will fall back to the resting potential U reset; Otherwise, no pulse is generated and U remains its original value. The neural dynamics characteristics of the above LIF neurons can be described by the following formula:
[0034]
[0035]
[0036] Where and respectively represent the subthreshold membrane potential and input current of the i-th neuron in the n-th layer at time t. β is the membrane potential decay constant, is the pulse excitation threshold, is the constant input current of the i-th neuron in the n-th layer at time t, represents the pulse generated by the i-th neuron in the n-th layer at time t, which is 1 or 0, and the specific value is determined by the following Heaviside step function:
[0037]
[0038] The Heaviside step function is a discrete function. Its derivative is always 0 when x≠0, and the derivative is infinite when x = 0, which causes the backpropagation algorithm to be unable to be directly applied to SNN.
[0039] In the forward propagation stage, the SNN receives input samples at each time step (where i = 1…T, T is the total number of time steps). The information propagates through layers of LIF neurons to the C neurons in the output layer (where C is the number of categories). The logits of the SNN are defined as the pulse firing frequency of the neurons in the output layer, that is, the number of pulses emitted by the output layer per unit time:
[0040]
[0041] Where represents the total number of pulses emitted by the neurons in the output layer in T time steps.
[0042] We hope that the neurons in the output layer corresponding to the true labels can emit pulses at each time step, that is, we hope that the logits of the neurons in the output layer corresponding to the true labels are as close to 1 as possible.
[0043] Furthermore, in order to calculate the gradient of Let
[0044]
[0045] The shape of this function is close to the Heaviside function and its derivative can be calculated During forward propagation, behaves the same as the Heaviside step function. During backpropagation,
[0046]
[0047] Furthermore, step 4 is specifically as follows:
[0048] In the middle layer, the loss function of Net S is defined as the mean squared error between the firing rate of the neurons in this layer and the activation value of Net S in this layer: T the activation value of the neurons in this layer:
[0049]
[0050] where represents the output of the activation value of the neurons in the nth layer of Net T , represents the maximum output of the activation value of the neurons in the nth layer of Net T , represents the total number of spikes fired by the neurons in the nth layer of Net S over T time steps
[0051] Next, we derive the derivative of the weight parameter w in Net S during the backpropagation process. First, let s
[0052]
[0053] For time step t = T, we directly obtain from the expression of
[0054]
[0055] From the above equation, we can get
[0056]
[0057] For time step t < T, according to the iterative process of the LIF neuron in the time domain, we have
[0058]
[0059] Furthermore, we have
[0060]
[0061] Furthermore, we obtain the gradient of the model parameters in Net S
[0062]
[0063]
[0064] In Net S 's output layer, according to the knowledge distillation method, the loss function is defined as the linear combination of the cross - entropy between the softmax output of Net S and the true labels and the KL divergence between the softmax outputs of Net S and Net T Specifically, in each forward propagation process, Net T and Net S obtain their softmax outputs
[0065]
[0066]
[0067] where \(z_i\) represents the logits of each class output;
[0068] The loss between Net S and the true labels is defined as the cross - entropy between
[0069]
[0070] where \(C\) is the number of classification categories.
[0071] Furthermore, a distillation temperature \(\tau\) is introduced to obtain the distilled softmax outputs of Net T and Net S The KL divergence between the two is used as the second part of the output layer loss
[0072]
[0073]
[0074] The output layer loss
[0075]
[0076] is defined as the linear combination of and and where \(\lambda_1\) and \(\lambda_2\) are weight hyperparameters for balancing the two losses.
[0077]
[0078]
[0079] Thus, the loss function of the middle layer is obtained. The loss function of the output layer and the approximate gradient function can be used to train Net with the error backpropagation algorithm S for the weight parameters of each layer.
[0080] Furthermore, the specific algorithm steps of step 5 are as follows:
[0081] S510: Net T loads the saved model parameters θ T ;
[0082] S520: Input the batch samples into Net T and Net S respectively for forward propagation to obtain the middle layer outputs of both and the softmax output of the output layer, that is, and
[0083] S530: Introduce the distillation temperature to calculate the distillation softmax outputs of both, that is, and
[0084] S540: Calculate the loss of each layer of Net and according to the formulas of S ;
[0085] S550: Use the neural network optimizer to execute the error backpropagation algorithm to update the model parameters of Net T During the training process of reducing and transfer the knowledge already learned by Net T to Net S ;
[0086] Repeat steps S520 - S550 until the model converges or reaches the predetermined number of training epochs.
[0087] An electronic device includes: a processor and a memory for storing instructions executable by the processor;
[0088] wherein, the processor is configured to execute the instructions to implement the above - mentioned training method of the spiking neural network based on knowledge transfer.
[0089] A computer - readable storage medium, when the instructions in the computer - readable storage medium are executed by the processor of an electronic device, enable the electronic device to execute the above - mentioned training method of the spiking neural network based on knowledge transfer.
[0090] Advantages: Compared with the prior art, the present invention has the following remarkable advantages: (1) The ANN-SNN conversion method requires a series of constraints on the original ANN model, such as not being able to use batch normalization method, not being able to use max pooling, etc., which seriously affects the performance of the original ANN model and the converted SNN model. However, the present invention does not require any constraints on the original ANN model and completely retains the performance of the original ANN; (2) The present invention does not need to set a large simulation time step to accurately represent the real-valued activation of the original ANN; (3) In the present invention, the SNN is not trained autonomously from scratch, but is guided by the teacher ANN during the training process, with a fast convergence speed. Brief Description of the Drawings
[0091] Figure 1 is a flowchart of the present invention;
[0092] Figure 2 is an explanatory diagram of a commonly used approximate gradient function in the present invention;
[0093] Figure 3 is a schematic diagram of the neural network training algorithm of the present invention. Detailed Embodiments
[0094] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0095] As Figure 1 shown, a method for training a spiking neural network based on knowledge transfer of the present invention includes the following steps:
[0096] Step 1: Prepare a dataset for image classification tasks, such as CIFAR-10, and divide the dataset into a training dataset D Train and a test dataset D Test , and preprocess the two datasets. In particular, perform data augmentation on the training dataset D Train to prevent model overfitting.
[0097] The preprocessing operations on the dataset include normalization Normalize to make the model converge faster; tensorization ToTensor to enable data to be calculated on a general GPU for acceleration. The data augmentation means for the training set D Train include random cropping RandomCrop, random rotation RandomRotation, horizontal flipping RandomHorizontalFlip, and vertical flipping RandomVerticalFlip, etc. These means can enhance the complexity of the data and prevent the neural network model from overfitting during the convergence process;
[0098] Step 2: Train a teacher network with good representation, learning, and generalization abilities, such as the vgg16 network, using the classical error backpropagation algorithm. The specific training steps are as follows:
[0099] S210: Randomly initialize the weights of Net T parameters
[0100] w t ~Random Initialization
[0101] S220: For each training sample (x i , y i ) ∈ D Train , calculate the predicted result of the sample:
[0102]
[0103] S230: Calculate the loss between the predicted results of all training samples and the actual labels:
[0104]
[0105] S240: Calculate the gradient of the loss function L with respect to the model parameter w t parameters:
[0106]
[0107] S250: Update the model parameters using the gradient descent optimization algorithm:
[0108]
[0109] Repeat steps S220 - S250 until the model converges or reaches the predetermined number of training epochs, and save the model parameters θ T of Net T for use in subsequent training of Net S .
[0110] Step 3: Construct the student spiking neural network Net S . Among them, the spiking neuron uses the LIF neuron, and its membrane potential U depends on the membrane potential at the previous moment and the input current. When its membrane potential U exceeds the spike firing threshold , it will emit a spike, and then the membrane potential will fall back to the resting level U reset . Its neural dynamics characteristics are described by the following formula
[0111]
[0112]
[0113]
[0114] As mentioned above, the pulse activation function is the non-differentiable Heaviside step function. In order to enable normal differentiation during the backpropagation process, an approximate gradient function is introduced for the LIF neuron. During forward propagation, behaves the same as the Heaviside step function. During backpropagation, there is
[0115]
[0116] Common There are several types, such as the Sigmoid function, the Atan function, etc. The shapes of these two approximate gradient functions and their derivative functions are shown in the appendix Figure 2 as shown.
[0117] Net S The loss function for each layer is as follows:
[0118] In the middle layer of the network, Net S 's loss function is defined as the mean square error between the pulse firing frequency of the neurons in Net S and the activation value of the neurons in Net T of this layer:
[0119]
[0120] In the output layer, the loss function is based on a linear combination of two parts of losses. The first part is defined as the cross-entropy between the softmax output of Net S and the true label. The second part is defined as the KL divergence between the softmax output of Net S and Net T after introducing the distillation temperature:
[0121]
[0122] where
[0123]
[0124]
[0125] Step 4: Train the SNN using the knowledge transfer method. Figure 3 is a schematic diagram of the algorithm training of the vgg16 neural network model. The specific steps of the training algorithm are as follows:
[0126] S410: As Figure 3As shown, the trained Teacher ANN, i.e., Net T Load the saved model parameters θ T ;
[0127] S420: As Figure 3 shown, for the same dataset dataset, input the batch samples into Net T and Net S respectively for forward propagation. According to the obtained intermediate layers of the two, i.e., the outputs of layer1, layer2, …, layer15 in the figure, and the output layer, i.e., the softmax output of layer16 in the figure, and
[0128]
[0129]
[0130] S430: Introduce the distillation temperature to calculate the distillation softmax outputs of the two, i.e., and
[0131]
[0132]
[0133] S440: Calculate the intermediate layer loss of each layer from layer1 to layer15 in Net according to the formula of S , and calculate the output layer distillation loss of layer16 in Net according to the calculation formula of S .
[0134] S450: Use neural network optimizers such as Adam and SGD to execute the error backpropagation algorithm to update the model parameters of each layer in Net T , and transfer the learned knowledge of Net and to Net T during the training process of reducing S .
[0135] Repeat steps S420 to S450 until the model converges or reaches the predetermined number of training epochs, and save the model parameters θ T of Net T , for use in subsequent training of Net S .
[0136] Step 5: Test Net on D Test S Model performance
[0137] Figure 2 The function graphs of two commonly used approximate gradient functions, Sigmoid and Atan, are shown. Their functional expressions and the functional expressions of their derivative functions are as follows:
[0138]
[0139] Sigmoid'(x) = a(1 - Sigmoid(x))Sigmoid(x)
[0140]
[0141]
[0142] where α and β are hyperparameters that control the function shape. In Figure 2 , α = 4 and β = 2.
Claims
1. A training method for a spiking neural network based on knowledge transfer, characterized in that, It includes the following steps: Step 1: Construct the dataset D for the image classification task of CIFAR-10, and divide the dataset D into a training set D Train and a test set D Test ; Step 2: Construct a teacher ANN network and train the teacher ANN network on D using the error backpropagation algorithm Train to obtain the trained teacher network Net T , and save the weight parameters in the teacher network Net T ; Step 3: Construct a spiking neural network SNN, denoted as the student network Net S , the student network Net S adopts LIF neurons with spike accumulation-emission and leakage characteristics, and introduces an approximate gradient function for the LIF neurons Step 4: Define the loss functions for the intermediate layer and the output layer of the student network Net S for the intermediate layer and the output layer Step 5: Transfer the knowledge learned by the teacher network Net T to the student network Net S as follows: Perform forward propagation on the teacher network Net T and the student network Net S respectively to obtain the intermediate layer outputs and the output layer outputs of both; Calculate the intermediate layer loss function S and the output layer loss function of the student network Net Execute the gradient descent algorithm using a neural network optimizer to train the student network Net S ; If the model of the student network Net S converges or reaches the maximum number of training epochs, then execute Step 6, otherwise return to Step 5; Step 6: Test the performance of the student network Net Test on the test set D S to obtain the trained spiking neural network.
2. The pulse neural network training method based on knowledge transfer according to claim 1, characterized in that Step 2 specifically includes: S210: Randomly initialize the weight parameters of Net T w t ~Random Initialization S220: For each training sample (x i , y i ) ∈ D Train , calculate the predicted result of the sample: where x i represents the input sample, w t represents the model parameter, is the network prediction result; S230: Calculate the loss between the prediction results of all training samples and the actual labels: wherein is defined as the cross entropy between two inputs; S240: Calculate the gradient of the loss function L with respect to the model parameter w t : S250: Update the model parameters using the gradient descent optimization algorithm: where α is the learning rate, which controls the step size of each parameter update; Repeat steps S220 - S250 until the model converges or reaches a predetermined number of training epochs, and save the model parameters θ of Net T for subsequent use in training Net T when training Net S again.
3. A training method for a spiking neural network based on knowledge transfer according to claim 1, characterized in that, The LIF neuron model described in Step 3 includes the following parts: Input current I: The input current received by the neuron from other neurons or the environment; Membrane potential U: The membrane potential of the neuron is an inherent property representing the internal electric potential of the LIF neuron. The input current will affect the membrane potential of the neuron, causing it to gradually increase; Threshold When the membrane potential of a neuron exceeds a specific threshold, the neuron fires an electrical impulse; Resting potential U reset : After a neuron fires an electrical impulse, the membrane potential will quickly reset to a lower potential, and the neuron will not respond to input current for a period of time; The membrane potential U of the inherent property of the LIF neuron depends on the membrane potential at the previous moment and the current input. In the spatial domain, the LIF neuron receives and integrates the stimulation of the input current, which acts on the membrane potential U, causing an increase in U. In the time domain, based on the membrane potential at the previous moment and the decay constant β, there will be a decay in the membrane potential U of the LIF neuron. After the combined effects of the spatial domain and the time domain, if the membrane potential exceeds the pulse firing threshold the LIF neuron will generate a pulse, and the membrane potential will fall back to the resting potential U reset ; otherwise, no pulse is generated and U remains at its original value. The above-mentioned neural dynamics characteristics of the LIF neuron are described by the following formula: Among them, and respectively represent the subthreshold membrane potential and input current of the i-th neuron in the n-th layer at time t. β is the membrane potential decay constant, is the pulse firing threshold, is the constant input current of the i-th neuron in the n-th layer at time t, represents the pulse generated by the i-th neuron in the n-th layer at time t, which is 1 or 0, and the specific value is determined by the following Heaviside step function: The Heaviside step function is a discrete function whose derivative is constantly 0 when x≠0, and the derivative is infinite when x = 0; Since the pulse activation function is the non-differentiable Heaviside step function, in order to enable normal differentiation during the backpropagation process, an approximate gradient function is introduced for the LIF neuron During forward propagation, behaves the same as the Heaviside step function. During backpropagation, we have 4. A training method for a spiking neural network based on knowledge transfer according to claim 1, characterized in that, Step 4 specifically is: In the middle layer of the network, Net S 's loss function is defined as the mean square error between the firing frequency of the neurons in this layer in Net S and the activation values of the neurons in this layer in Net T : Among them represents Net T the output of the activation value of the neurons in the nth layer, represents Net T the output of the maximum activation value of the neurons in the nth layer, represents Net S the total number of spikes fired by the neurons in the nth layer over T time steps; The following derives the process of backpropagation for Net S the derivative of the weight parameter w s in it. First, let For the time step t = T, it is directly obtained from the expression of Obtained from the above formula For time step t < T, according to the iterative process of the LIF neuron in the time domain, there is Furthermore, there is Thus, for Net S the gradients of the model parameters in In Net S 's output layer, the loss function is a linear combination of two parts of losses. The first part is the cross-entropy between the softmax output of Net S and the true labels. The second part is the KL divergence between the softmax outputs of Net S and Net T after introducing the distillation temperature. Specifically, in each forward propagation process, Net T and Net S obtain their softmax outputs where z i represents the logits output for each class; The first part of the loss is defined as the cross-entropy with the true label, i.e., where C is the number of classification categories; Introduce the distillation temperature τ to obtain Net T and Net S The distillation softmax output of The first part of the loss is defined as the KL divergence between the two Output layer loss is defined as and a linear combination of where λ1 and λ2 are weight hyperparameters that balance the two losses.
5. A training method for a spiking neural network based on knowledge transfer according to claim 1, characterized in that Step 5 specifically includes: S510: Net T Load the saved model parameters θ T ; S520: Input the batch samples into Net T and Net S Perform forward propagation to obtain the intermediate layer outputs of both, as well as the softmax output of the output layer, i.e., and S530: Introduce the distillation temperature to calculate the distillation softmax outputs of the two, that is and S540: Calculate the loss of Net for each layer according to and using the formula S ; S550: Use a neural network optimizer to execute the error backpropagation algorithm to update the model parameters of Net, and transfer the knowledge already learned by Net T to Net S during the training process of reducing and . T during the process of reducing and transfer the knowledge already learned by Net T to Net S ; Repeat steps S520 - S550 until the model converges or reaches the predetermined number of training epochs.
6. An electronic device, characterized in that, It includes: A processor and a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the knowledge transfer - based spiking neural network training method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, When the instructions in the computer - readable storage medium are executed by the processor of the electronic device, the electronic device can execute the knowledge transfer - based spiking neural network training method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Facial expression image recognition method and device, electronic equipment and storage medium
CN115482578A
Pulse neural network lightweight method and system suitable for embedded device
CN116151335A