A Hardware Accelerator for a Convolutional Spiking Neural Network Based on STDP Online Learning
By designing a convolutional pulse neural network hardware accelerator based on STDP and RSTDP on the FPGA platform, the existing DCNNs are solved for the problem of large and time-consuming calculations in online learning tasks, and efficient online learning capabilities and low-power hardware implementation are achieved.
Patent Information
- Application Number
- CN202210220091.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-08
AI Technical Summary
Existing deep convolutional neural networks (DCNNs) are not suitable for online learning tasks, and their calculations are large and time-consuming, making it difficult to achieve efficient online learning on hardware platforms.
Using STDP and RSTDP as training algorithms, a convolutional pulse neural network hardware accelerator based on FPGA is designed to realize online learning capabilities. The hardware accelerator includes a PS terminal, a PL terminal and a PC terminal, and data transmission and processing are carried out through the TCP protocol and the AXI bus.
It realizes an efficient convolutional pulse neural network hardware accelerator on the FPGA platform, supports online learning, significantly reduces computing time and power consumption, and improves training speed and accuracy.
Smart Images

Figure CN114611684B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of hardware acceleration of neural networks, and relates to a hardware accelerator for a convolutional spiking neural network based on STDP online learning. Background Art
[0002] Neural networks are a core in the current field of artificial intelligence, and among them, deep convolutional neural networks (DCNNs) have powerful functions. However, DCNNs with large computational amounts and long time consumption are not suitable for online learning tasks that can adapt to environmental changes. Spiking neural networks (SNNs), on the other hand, simulate the operation mechanism of the biological brain, transmit information through spikes, require lower computational precision, are faster, and are more suitable for the field of online learning. In addition, the computational power consumption of SNNs is low, and implementing SNNs on a hardware platform can significantly reduce power consumption.
[0003] STDP is a local unsupervised SNN training rule, and its synaptic weight update does not depend on data labels, but on the order of spike firing of two local neurons. This training method makes memory (training) and computing (inference) more closely combined, the architecture is more similar to the human brain, and it is also suitable for hardware implementation. Training during the inference process is the way of online learning.
[0004] The hardware implementation of SNNs can be divided into two types from the circuit technology: mixed-signal circuit implementation and fully digital circuit implementation. The simulation of complex neuron models is relatively easy for analog circuits, but the stability is not high and the programmability is not high. Therefore, fully digital circuit implementation is more favored by researchers:
[0005] 1) TrueNorth: The most representative fully digital circuit neuromorphic chip of IBM. The computing cores are connected through a 2D network routing and can be extended with multiple chips. It supports multiple neuron models and has low computational power consumption.
[0006] 2) Loihi: An online learning neuromorphic chip released by Intel, which supports variable synaptic formats, supports on-chip online learning, changes the synaptic state according to historical spike activities, and updates the synaptic weights in real time according to the STDP rule.
[0007] 3) Tianji chip: A heterogeneous fusion neuromorphic chip released by Tsinghua University in 2019, which can support artificial neural networks and spiking neural networks at the same time.
[0008] 4) Darwin chip: Developed by Zhejiang University and Hangzhou Dianzi University in 2015, it supports 2048 neurons and 4,194,304 synaptic connections, and supports 15 synaptic delays.
[0009] 5) FPGA: Implementing spiking neural networks through FPGA programmable devices has a short development cycle, rich hardware computing resources, fast computing speed and low power consumption. Summary of the Invention
[0010] The purpose of the present invention is to overcome the deficiencies of the prior art. The present invention uses STDP and RSTDP as training algorithms to provide an online learning hardware accelerator architecture for implementing a convolutional spiking neural network on an FPGA. The specific content is as follows:
[0011] The present invention proposes a hardware accelerator for a convolutional spiking neural network based on STDP online learning, which includes an FPGA composed of a PS side and a PL side, and a PC side;
[0012] The PC side preprocesses the input sample image data into pulse data, transmits it to the PS side through the TCP protocol, and then the PL side reads it through the AXI bus;
[0013] For the data returned by the PL side to the PS side, when sending specific data with a large quantity such as pulses, the PS side sends a data write address to the PL side during initialization. The network output data of the PL side will be stored in a FIFO. When the data storage quantity reaches a certain value, the PL side will give an interrupt to the PS side to notify it to read the data. At the same time, the PL side will write the data in the FIFO into the DDR at the corresponding address, and the PS side will read the data from it and then return it to the PC side; if the sent is a single control signal such as the number of training layers, the PS side directly sends it to the PL side through the AXI_Lite protocol.
[0014] As a preferred solution of the present invention, the PL side includes a training calculation mode and an inference calculation mode; in the training calculation mode, the input is a labeled sample for learning, and the synaptic weights in the network are modified according to the specified STDP / RSTDP training rules;
[0015] In the inference calculation mode, an unlabeled sample is input, and the predicted label for the sample is output.
[0016] As a preferred solution of the present invention, the PL side includes a pulse data transmission module, a control signal transmission module, an input pulse FIFO, a predicted label FIFO, and a three-layer network training layer; the pulse data transmission module and the control signal transmission module are respectively used to transmit specific data and control signals to and from the PS side; the input pulse FIFO and the predicted label FIFO are respectively used to cache the input pulses of the network and the output predicted labels of the network; the first two training layers include a convolutional layer, a pooling layer, and a padding layer, and the third training layer only includes a convolutional layer; the convolutional layers of the first two layers adopt the STDP learning rule, and the convolutional layer of the third layer adopts the R-STDP learning rule, and the training is carried out layer by layer.
[0017] As a preferred embodiment of the present invention, the first two convolutional layers, namely convolutional layer 1 and convolutional layer 2, are convolutional modules adopting the STDP learning rule; the third convolutional layer, namely convolutional layer 3, is a convolutional module adopting the R-STDP learning rule; pooling layer 1 and pooling layer 2 are pulse-based pooling modules, and padding layer 1 and padding layer 2 are for zero-padding the image edges; each convolutional module corresponds to a weight cache module and a weight sum cache module.
[0018] As a preferred embodiment of the present invention, the convolutional layer includes a convolutional overall control module (CONV_CTRL), a shift module (SH_CTRL), a computing unit control module (CU_CTRL), a convolutional kernel computing unit array (CU), a membrane potential computing module (POT), a pointwise inhibition module (POI_CTRL), a winner selection module (WIN_CTRL), and an STDP / RSDTP learning module (LEARN_STDP, LEARN_RSTDP).
[0019] As a preferred embodiment of the present invention, the training calculation mode adopts a layer-by-layer training method, and the three convolutional layers are trained in sequence. During training, only the weights of the training layer are changed, and the weights of other layers are not changed. Among them, the membrane potential computing module stores the output membrane potential in the membrane potential FIFO. The pointwise inhibition module reads the membrane potentials of several channels, selects the most excellent neurons according to the membrane potentials, and outputs the inhibited neurons to the pointwise inhibition RAM. After all input neurons have undergone pointwise inhibition calculation, the winner selection module reads the inhibited neurons from the pointwise inhibition RAM, selects the winner from them, and stores it in the winner neuron FIFO. The STDP / RSTDP learning module reads the winner from the winner neuron FIFO and updates the weights of its connected synapses.
[0020] As a preferred embodiment of the present invention, the membrane potential computing module stores the membrane potential in the membrane potential FIFO in the order of the number of channels where the neuron is located, the x and y coordinates in the feature map, and the pulse time. The pointwise inhibition module selects the most excellent neuron from the neurons of all channels at the same coordinate in the feature map and reads the membrane potential in the same order. The two modules cache data through the FIFO, so they can perform calculations independently and in parallel; similarly, every time the winner selection module outputs a winner neuron, the STDP / RSTDP learning module can update the weights of the synapses connected to the winner neuron, and the two modules cache data through the FIFO, so they can also perform calculations independently and in parallel.
[0021] As a preferred embodiment of the present invention, the inference calculation mode is that the convolution overall control module reads pulse data from the input pulse FIFO and then sends it to the shift control module. When the shift control module stores a certain number of pulses, it outputs one pulse to the convolution kernel calculation unit. The convolution kernel calculation unit calculates the weight sum, and then aggregates it to the membrane potential calculation module to calculate the membrane potential, and outputs the pulse and the membrane potential; among them, every time the shift control module sends one pulse to the convolution kernel calculation unit, the latter will cumulatively calculate the weight sums of the output channels, so it will also output the pulses and membrane potentials of multiple channels. The convolution overall control module will wait until they are completely calculated and output before continuing to input the next pulse.
[0022] As a preferred embodiment of the present invention, the calculation implementation process of convolution in the convolution layer is as follows: One fixed convolution kernel slides on the input feature map, and the overlapping part of the convolution kernel and the feature map performs multiplication and addition convolution calculation to obtain an output value. Every time the convolution kernel slides once, one output value is obtained. When the convolution kernel slides to cover the entire input feature map, the calculation ends, and an output feature map is obtained; when the input feature map has multiple channels, the convolution kernel is in a three-dimensional form, and a three-dimensional overlapping part needs to perform multiplication and addition calculation; if there are multiple convolution kernels, multiple channels of output feature maps will be generated; in the spiking neural network, convolution is reflected in the synapses connecting two layers of neurons; Neurons within one convolution kernel region in the upper layer are connected to one neuron in the lower layer, and the weights on the convolution kernel correspond to the corresponding synapses; this neuron accumulates the sum of the synaptic weights with pulse inputs, that is, the convolution calculation is realized; Since the convolution kernel is shared in one layer and there are many synapses with the same weights; Therefore, each position of the convolution kernel is set as one calculation unit, and the input pulses are transmitted to the calculation unit by a shift register, and then the calculation results of all calculation units are summed to realize the convolution calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is the physical architecture of the hardware accelerator of the convolutional spiking neural network based on STDP online learning of the present invention;
[0024] Figure 2 It is a schematic diagram of the hardware module division on the PL side of the present invention;
[0025] Figure 3 Schematic diagram of the STDP training part inside the convolution kernel;
[0026] Figure 4 Schematic diagram of the RSTDP training part inside the convolution kernel;
[0027] Figure 5 Schematic diagram of the winner neuron selection structure;
[0028] Figure 6 Schematic diagram of the STDP learning module structure;
[0029] Figure 7 Schematic diagram of the RSTDP learning module structure;
[0030] Figure 8 Schematic diagram of the in-kernel inference calculation part;
[0031] Figure 9 Schematic diagram of the connection relationship between the shift register and the calculation unit;
[0032] Figure 10 Structure diagram of the convolution kernel calculation unit. Specific implementation manner
[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0034] As Figure 1 shown, the present invention uses Xilinx's ZYNQ system. The hardware accelerator of the convolutional spiking neural network based on STDP online learning is divided into three parts: a computer (PC side), a software processing system (PS side, Processing System) of the FPGA, and a hardware programmable logic (PL side, Programmable logic) of the FPGA.
[0035] The input sample image data is preprocessed into pulse data by the PC side, transmitted to the PS side through the TCP protocol, and then read by the PL side through the AXI bus. For the data returned by the PL side to the PS side, at initialization, the PS side will send a data write address to the PL side. The network output data of the PL side will be stored in a FIFO. When the number of stored data reaches a certain value, the PL side will give an interrupt to the PS side to notify it to read the data. At the same time, the PL side will write the data in the FIFO into the DDR at the corresponding address, and the PS side will read the data from it and then return it to the PC side. The above-mentioned data sent is specific data such as pulses with a large quantity. If a single control signal such as the number of training layers is sent, the PS side only needs to directly send it to the PL side through the AXI_Lite protocol.
[0036] The PL side includes two working modes: a training calculation mode and an inference calculation mode; in the training calculation mode, the input is a labeled sample for learning, and the synaptic weights in the network are modified according to the specified STDP / RSTDP training rules; in the inference calculation mode, the input is an unlabeled sample, and the predicted label for the sample is output.
[0037] The PL side performs actual network calculations. The module division of the PL side of the present invention is as Figure 2As shown, the pulse data transmission module and the control signal transmission module are respectively used to transmit specific data and control signals to the PS side. The input pulse FIFO and the prediction label FIFO are respectively used to cache the input pulses of the network and the output prediction labels of the network. According to the network topology and its algorithm, the PL side is divided into three layers. The first two layers include convolutional layers, pooling layers, and padding layers, and the third layer only contains a convolutional layer. The convolutions in the first two layers adopt the STDP learning rule, and the convolution in the third layer adopts the R-STDP learning rule, and the training is carried out layer by layer. Convolution layer 1 and convolution layer 2 are convolutional modules that adopt the STDP learning rule, and their training architectures are as Figure 3 shown. Convolution layer 3 is a convolutional module that adopts the R-STDP learning rule, and its training architecture is as Figure 4 shown. Pooling layer 1 and pooling layer 2 are pulse-based pooling modules, and padding layer 1 and padding layer 2 are used to pad zeros at the image edges. Each convolutional module corresponds to a weight cache module and a weight sum cache module.
[0038] As Figure 3 and 4 shown, the calculations of the present invention mainly occur inside the convolutional layer, and are further divided into a convolutional overall control module (CONV_CTRL), a shift module (SH_CTRL), a calculation unit control module (CU_CTRL), a convolutional kernel calculation unit array (CU), a membrane potential calculation module (POT), a pointwise inhibition module (POI_CTRL), a winner selection module (WIN_CTRL), and an STDP / RSDTP learning module (LEARN_STDP, LEARN_RSTDP). Among them, the pointwise inhibition module, the winner selection module, and the STDP / RSDTP learning module belong to the training calculation part. When calculating, the input is a labeled sample for learning, and the synaptic weights in the network are modified according to the specified STDP / RSTDP training rules; while the remaining modules inside the convolutional layer and the pooling layer and the padding layer all belong to the inference calculation part. The input is an unlabeled sample, and the predicted label for this sample is output.
[0039] The described training calculation mode adopts a layer-by-layer training method, training the three convolutional layers in sequence. During training, only the weights of the training layer are changed, and the weights of other layers are not changed. Among them, the membrane potential calculation module stores the output membrane potential in the membrane potential FIFO. The pointwise inhibition module reads the membrane potentials of several channels, selects the most excellent neurons according to the membrane potentials, and outputs the inhibited neurons to the pointwise inhibition RAM. After all input neurons have undergone pointwise inhibition calculation, the winner selection module reads the inhibited neurons from the pointwise inhibition RAM, selects the winner from them, and stores it in the winner neuron FIFO. The STDP / RSTDP learning module reads the winner from the winner neuron FIFO and updates the weights of its connected synapses. Among them, the membrane potential calculation module stores the membrane potential in the membrane potential FIFO in the order of the number of channels where the neurons are located, the x and y coordinates in the feature map, and the pulse time. The pointwise inhibition module selects the most excellent neuron from the neurons of all channels at the same coordinate in the feature map and reads the membrane potential in the same order. The two modules cache data through the FIFO, so they can perform calculations independently and in parallel. Similarly, every time the winner selection module outputs a winner neuron, the STDP / RSTDP learning module can update the weights of the synapses connected to the winner neuron, and the two modules cache data through the FIFO, so they can also perform calculations independently and in parallel.
[0040] In a specific embodiment, to increase the feature extraction ability, a pointwise inhibition mechanism is adopted during training. After pointwise inhibition, the most matching feature at each position can be selected, and at the same time, the number of pulse firings of the neurons in this layer is reduced. At a position in the feature map at a certain moment, as long as there are neurons firing pulses on the channels therein, the membrane potentials of all the neurons excited by the pulses are compared, and the most excellent neuron in one channel is selected. Figure 5 Taking the data of 6 channels as an example, the membrane potentials of each channel at the same position in the feature map are sequentially read from the membrane potential storage FIFO. If it exceeds the threshold voltage and can generate a pulse, it is compared with the current maximum membrane potential stored in the register. If the newly input membrane potential is larger, it is stored in the most excellent neuron register. After inputting and comparing the 6 channels, the most excellent neuron at a position in the feature map is obtained.
[0041] In the training calculation, the STDP / RSTDP learning module is used to update the weights of the synapses connected to a small number of winner neurons selected by the previous module. As Figure 6 shown, the STDP learning module first judges the pulse sequence of the pre- and post-synaptic neurons according to the STDP rule of the algorithm, and then increases or decreases the synaptic weight, and the amount of increase or decrease is the learning rate. For the RSTDP module, as Figure 7As shown, it is also necessary to compare whether the predicted label is the same as the true label to determine the increase or decrease of the weight. The algorithm also stipulates the weight clamping rule and the learning rate update rule, which limit the range of the weight and update the learning rate according to a fixed rule after a certain number of input images have been reached.
[0042] The described inference calculation mode is that the convolution overall control module reads pulse data from the input pulse FIFO and then sends it to the shift control module. When the shift control module stores a certain number of pulses, it outputs 1 pulse to the convolution kernel calculation unit. The convolution kernel calculation unit calculates the weight sum, then aggregates it to the membrane potential calculation module to calculate the membrane potential, and outputs the pulse and the membrane potential. Among them, every time the shift control module sends 1 pulse to the convolution kernel calculation unit, the latter will cumulatively calculate the weight sums of the output channel numbers, so it will also output the pulses and membrane potentials of multiple channels. The convolution overall control module will wait for them to be completely calculated and output before continuing to input the next pulse.
[0043] The module division within the convolution kernel of the present invention is as Figure 8 shown, where the convolution overall control module is used to overall control the calculation progress of the convolution layer. The shift register module transfers the input pulse to the calculation unit, as Figure 9 Taking the first convolution layer of the network as an example, when the size of the input feature map is 32×32 and the size of the convolution kernel is 5×5, a total of 25 CUs are set in the convolution module CONV1 of the first layer to calculate the weight sum. At the same time, 4 shift registers are set to cache the input pulses, and the length of each shift register is set to 32. The first and last 5 ports of each shift register are led out and connected to the 5 rightmost edges of the 25 CUs ((0,4),(1,4),(2,4),(3,4),(4,4)). At the same time, two adjacent CUs on the left and right are connected pairwise, and the pulses are transmitted from the neuron on the right to the neuron on the left. Whenever a pulse is input, it is input from the input port of the 1st shift register, and the data in the register is shifted once. Until the pulse data fills the register, when another pulse is input, the output of the 1st register is input to the 2nd register, and so on. While the pulse is being stored in the register, the pulse is also transmitted from the connection port between the register and the CU to the CU. The pulse is transmitted from right to left in a row of CUs and gradually fills all 25 CUs from bottom to top. When all CUs are filled with pulses, the CUs start to calculate the weight sum. The calculation unit control module and the convolution kernel calculation unit array calculate the weight sum according to the input pulse, and aggregate the calculation results to the membrane potential calculation module to obtain the neuron membrane potential. As Figure 10It is the structure of the convolution kernel calculation unit. The input pulse data will be stored in a register inside the CU. A pulse data contains the pulse information of neurons in the number of channels. When the external control signal indicates the start of calculation, the CU will traverse all channels, and based on whether the pulse of the corresponding channel in the register is 1 or 0, it will choose whether to accumulate the input weight to the weight sum. After waiting for all channels to be traversed, the CU will output the weight sum and send the pulse data in the register to other CUs.
[0044] The MINIST dataset is used to test the present invention. First, the parameters of the original algorithm are quantized, and the floating-point numbers are converted into fixed-point numbers suitable for hardware processing. Then, the input pulse coding is performed. The test results show that the accuracy of the present invention can reach 95%, the training speed is 16 times faster than that of the cpu platform, the power consumption can be reduced by two orders of magnitude, and compared with hardware projects using the same model, the hardware resource consumption is equivalent.
[0045] The above is only an example of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A hardware accelerator for a convolutional spiking neural network based on STDP online learning, characterized in that it includes: an FPGA composed of a PS side and a PL side, and a PC side; The PC side preprocesses the input sample image data into pulse data, transmits it to the PS side via the TCP protocol, and then the PL side reads it through the AXI bus; For the data returned from the PL side to the PS side, when the specific data with a large number of pulses is sent, the PS side sends a data write address to the PL side during initialization. The network output data of the PL side will be stored in a FIFO. When the data storage quantity reaches a certain value, the PL side will give an interrupt to the PS side to notify it to read the data. At the same time, the PL side will write the data in the FIFO into the DDR at the corresponding address, and the PS side will read the data from it and then return it to the PC side; if the control signal with a single training layer is sent, the PS side will directly send it to the PL side through the AXI_Lite protocol; The PL side includes two working modes: a training calculation mode and an inference calculation mode; in the training calculation mode, the input is a labeled sample for learning, and the synaptic weights in the network are modified according to the specified STDP / RSTDP training rules; In the inference calculation mode, an unlabeled sample is input, and the predicted label for the sample is output; The training calculation mode adopts a layer-by-layer training method, and trains the 3 convolutional layers in sequence. During training, only the weights of the training layer are changed, and the weights of other layers are not changed. Among them, the membrane potential calculation module stores the output membrane potential in the membrane potential FIFO. The point-by-point inhibition module reads the membrane potentials of several channels, selects the most excellent neurons according to the membrane potential, and outputs the inhibited neurons to the point-by-point inhibition RAM. After all input neurons have undergone point-by-point inhibition calculation, the winner selection module reads the inhibited neurons from the point-by-point inhibition RAM, selects the winner from them, and stores it in the winner neuron FIFO; the STDP / RSTDP learning module reads the winner from the winner neuron FIFO and updates the weights of its connected synapses; In the inference calculation mode, the convolutional overall control module reads the pulse data from the input pulse FIFO, and then sends it to the shift control module. When the shift control module stores a certain number of pulses, it outputs 1 pulse to the convolutional kernel calculation unit. The convolutional kernel calculation unit calculates the weight sum, and then aggregates it to the membrane potential calculation module to calculate the membrane potential, and outputs the pulse and the membrane potential; among them, every time the shift control module sends 1 pulse to the convolutional kernel calculation unit, the latter will cumulatively calculate and output the weight sums of the number of output channels, so it will also output the pulses and membrane potentials of multiple channels. The convolutional overall control module will wait until they are completely calculated and output before continuing to input the next pulse.
2. The hardware accelerator for a convolutional spiking neural network based on STDP online learning according to claim 1, characterized in that, The PL side includes a pulse data transmission module, a control signal transmission module, an input pulse FIFO, a prediction label FIFO, and a three-layer network training layer; the pulse data transmission module and the control signal transmission module are respectively used to transmit specific data and control signals to and from the PS side; the input pulse FIFO and the prediction label FIFO are respectively used to cache the input pulses and the output prediction labels of the network; the first two training layers include a convolutional layer, a pooling layer, and a padding layer, and the third training layer only contains a convolutional layer; the convolutional layers of the first two layers adopt the STDP learning rule, and the convolutional layer of the third layer adopts the R-STDP learning rule, and the training is carried out layer by layer.
3. The hardware accelerator of the convolutional spiking neural network based on STDP online learning according to claim 2, characterized in that, The first two convolutional layers, namely convolutional layer 1 and convolutional layer 2, are convolutional modules adopting the STDP learning rule; the convolutional layer of the third layer, namely convolutional layer 3, is a convolutional module adopting the R-STDP learning rule; pooling layer 1 and pooling layer 2 are pulse-based pooling modules, and padding layer 1 and padding layer 2 are for zero-padding the image edges; each convolutional module corresponds to a weight cache module and a weight sum cache module.
4. The hardware accelerator of the convolutional spiking neural network based on STDP online learning according to claim 2 or 3, characterized in that, The convolutional layer includes a convolutional overall control module (CONV_CTRL), a shift module (SH_CTRL), a computing unit control module (CU_CTRL), a convolutional kernel computing unit array (CU), a membrane potential computing module (POT), a point-by-point inhibition module (POI_CTRL), a winner selection module (WIN_CTRL), and an STDP / RSDTP learning module (LEARN_STDP, LEARN_RSTDP).
5. The hardware accelerator of the convolutional spiking neural network based on STDP online learning according to claim 1, characterized in that, The membrane potential computing module stores the membrane potential in the membrane potential FIFO in the order of the number of channels where the neuron is located, the x and y coordinates in the feature map, and the pulse time. The point-by-point inhibition module selects the most outstanding neuron from all the neurons in the same coordinate of all channels in the feature map and reads the membrane potential in the same order. The data between the two modules is cached by the FIFO, so they can be calculated separately and in parallel; similarly, every time the winner selection module outputs a winner neuron, the STDP / RSTDP learning module can update the weights of the synapses connected to the winner neuron, and the data between the two modules is cached by the FIFO, so they can also be calculated separately and in parallel.
6. The hardware accelerator of the convolutional spiking neural network based on STDP online learning according to claim 2 or 3, characterized in that, The calculation implementation process of convolution in the convolutional layer is as follows: One fixed convolution kernel slides on the input feature map, and the overlapping part between the convolution kernel and the feature map performs multiplication and addition convolution calculation to obtain an output value. Each time the convolution kernel slides once, one output value is obtained. When the convolution kernel slides to cover the entire input feature map, the calculation ends, and an output feature map is obtained. When the input feature map has multiple channels, the convolution kernel is in three-dimensional form, and multiplication and addition calculation need to be performed on a three-dimensional overlapping part. If there are multiple convolution kernels, output feature maps with multiple channels will be generated. In a spiking neural network, convolution is reflected in the synapses connecting two layers of neurons. Neurons within one convolution kernel region in the upper layer are connected to one neuron in the lower layer, and the weights on the convolution kernel correspond to the corresponding synapses. This neuron accumulates the sum of the synaptic weights with pulsed inputs, thus implementing convolution calculation. Since the convolution kernel is shared within one layer and there are many synapses with the same weights, each position of the convolution kernel is set as one calculation unit. A shift register is used to transfer the input pulses to the calculation units, and then the calculation results of all calculation units are summed to implement convolution calculation.
Citation Information
Patent Citations
Method for learning and recognizing image pulse data space-time information based on Spike cube SNN
CN110210563A
Target recognition method, device and system and computer readable storage medium
CN111275742A