A Pulse Convolutional Neural Network Accelerator
By designing the pulse convolutional neural network accelerator and using the first-order Euler method to optimize neuronal computing, the problem of single pulse neural network model and large resource occupation is solved, and efficient and low-power pulse neural network calculation is realized.
Patent Information
- Application Number
- CN202210987300.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-08-17
AI Technical Summary
In the prior art, the neuron model of the pulsed neural network is single and cannot support the construction of multiple different pulsed neurons. The computing unit occupies hardware resources, computing volume and power consumption.
A pulse convolutional neural network accelerator is designed, including a controller, an image sorting unit, a PE computing array, a fully connected computing unit and an internal image buffering unit. By optimizing the differential calculations of LIF neurons and Izhikevich neurons with a first-order Euler method, the computational complexity is reduced and the network of different pulsed neurons is supported.
It realizes efficient pulse neural network computing, reduces hardware resource usage and calculation amount, improves computing speed, supports a variety of different pulse neurons, and improves resource and power consumption problems in the existing technology.
Smart Images

Figure CN115329934B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of spiking neural network acceleration, and particularly to a spiking convolutional neural network accelerator. Background Art
[0002] A spiking neural network (SNN) is the third-generation neural network with spiking neurons as computing units, which can simulate the encoding and processing process of human brain information. Different from the traditional artificial neural network (ANN), the SNN uses discrete spike signals to transmit information, and its neurons have memory functions, so it has a high degree of biological realism and real-time processing ability. At present, SNNs have been widely studied and applied to various high-performance computing tasks. Generally speaking, the scale of the SNN network is large, and the requirement for operation parallelism is high. Therefore, software simulation on traditional PC terminals is inefficient and consumes a large amount of computing power. How to design and implement a hardware system suitable for SNN computing and acceleration has become a research hotspot in the field of neural network computing.
[0003] The neuron models in the prior art are single, unable to support SNNs built with multiple different spiking neurons, and the hardware resources, computing volume, and power consumption occupied by the computing units are very large. Summary of the Invention
[0004] This application provides a spiking convolutional neural network accelerator, which is used to improve the technical problems in the prior art that the neuron models are single, unable to support SNNs built with multiple different spiking neurons, and the hardware resources, computing volume, and power consumption occupied by the computing units are very large.
[0005] In view of this, in the first aspect of this application, a spiking convolutional neural network accelerator is provided, including: a controller, an image sorting unit, a PE computing array, a fully connected computing unit, and an internal image buffer unit;
[0006] The controller is configured to, after receiving an input image, start the image sorting unit to sort the image, and after the sorting is completed, start the PE computing array or the fully connected computing unit to perform calculations, and cache the results generated by the calculations into the internal image buffer unit. If the current network layer is the output layer, then count the spike firing frequency to complete image recognition. If the current network layer is a non-output layer, then start the image sorting unit;
[0007] The PE computing array is used to complete the calculations of the convolutional layer and the pooling layer in the spiking neural network;
[0008] The fully connected computing unit is used to complete the calculations of the fully connected layer in the spiking neural network;
[0009] Among them, both the PE computing array and the fully connected computing unit include optimized LIF neurons and optimized Izhikevich neurons obtained by using the first-order Euler method to eliminate the differential calculations solved by LIF neurons and Izhikevich neurons.
[0010] Optionally, the image sorting unit is specifically configured to:
[0011] When receiving a start signal, determine whether it is the input layer to select a data source. If so, sort the input image; if not, sort the intermediate buffered image in the internal image buffer unit;
[0012] After the sorting is completed, generate a pulse sequence sorting completion signal.
[0013] Optionally, the optimization process of the LIF neuron is as follows:
[0014] Use the first-order Euler method to eliminate the differential calculation solved by the LIF neuron, and the calculation formula for eliminating the solved differential calculation is:
[0015]
[0016] For the parameters in the calculation formula of the eliminated solved differential calculation Perform solidification, use arithmetic phase shift to replace the multiplication operation in the calculation formula of the eliminated solved differential calculation, and use the second-order predictor-corrector method to optimize the calculation formula of the eliminated solved differential calculation to obtain an optimized calculation formula. The optimized calculation formula is:
[0017]
[0018]
[0019]
[0020]
[0021] In the formula, V[n] is the membrane state at the current moment, V[n + 1] is the magnitude of the membrane potential to be obtained, is the membrane potential voltage obtained by using the first-order Euler method, t n is the nth discrete time step during the calculation, v n is the magnitude of the input membrane potential, α is the R of the neuron m I is the influence magnitude of the neuron's current membrane potential, R mis the neuron membrane resistance constant, I is the input current value, f1(t, v) is the first target equation, f2(t, v) is the second target equation, β1 and β2 are the magnitude biases of V in different stages, and V reset is the neuron resting potential, h is the time step, and τ m is the time constant.
[0022] Optionally, the calculation formula of the optimized Izhikevich neuron is:
[0023] V[n + 1] = (0.04V 2 + 5V + 140 - U + I)·h;
[0024] U[n + 1] = [a(bV - U)]·h;
[0025] In the formula, V is the membrane voltage, U is the membrane potential recovery variable, I is the input current of the neuron, a, b, c, and d are the model parameters of the neuron, and V threhold is the membrane voltage threshold.
[0026] Optionally, it further includes: a weight storage unit for storing and reading the weights of convolution calculations.
[0027] Optionally, it further includes: a membrane potential storage unit for storing and reading the neuron membrane potential.
[0028] Optionally, the PE calculation array is specifically used for:
[0029] Reading and writing the weights and membrane potential from the weight storage unit and the membrane potential storage unit respectively to update the neuron membrane potential.
[0030] Optionally, the PE calculation array includes a number of PE units;
[0031] The PE unit is used to, when performing a convolution operation, accumulate the weights into a register when detecting an input pulse, and after completing the convolution operation, add the membrane potential of the input neuron to the register to obtain the final membrane potential; when performing a pooling operation, count the number of pulses in the input image through a finite state machine and save the currently required accumulated value determined according to the number of pulses into the register to obtain the final neuron membrane potential and internal pulses.
[0032] It can be seen from the above technical solutions that the present application has the following advantages:
[0033] The present application provides a pulsed convolutional neural network accelerator, comprising: a controller, an image sorting unit, a PE computing array, a fully-connected computing unit, and an internal image buffer unit; the controller is configured to, after receiving an input image, start the image sorting unit to sort the image, and after the sorting is completed, start the PE computing array or the fully-connected computing unit to perform calculations, and cache the results generated by the calculations into the internal image buffer unit. If the current network layer is an output layer, the pulse firing frequency is statistically calculated to complete image recognition. If the current network layer is a non-output layer, the image sorting unit is started; the PE computing array is used to complete the calculations of the convolutional layer and the pooling layer in the pulsed neural network; the fully-connected computing unit is used to complete the calculations of the fully-connected layer in the pulsed neural network; wherein, both the PE computing array and the fully-connected computing unit include optimized LIF neurons and optimized Izhikevich neurons obtained by using the first-order Euler method to eliminate the differential calculations of the LIF neurons and the Izhikevich neurons.
[0034] In the present application, the neurons in the neuron computing unit are optimized neurons obtained by using the first-order Euler method to eliminate the differential calculations of the LIF neurons and the Izhikevich neurons, which optimize the calculation processes of the LIF neurons and the Izhikevich neurons, reduce the calculation complexity, enable the neuron computing unit to update the neuron membrane potential faster, occupy less hardware resources, and can support pulsed neural networks of different pulsed neurons, improving the technical problems in the prior art that the neuron model is single, unable to support SNNs built by multiple different pulsed neurons, and the hardware resources, calculation amount, and power consumption occupied by the computing unit are very large. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic structural diagram of a pulsed convolutional neural network accelerator provided by an embodiment of the present application;
[0037] Figure 2 It is a schematic diagram of the controller state update change provided by an embodiment of the present application. Detailed Embodiments
[0038] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0039] For ease of understanding, please refer to Figure 1 , the embodiments of this application provide a pulse convolutional neural network accelerator, including:
[0040] A controller, an image sorting unit, a PE computing array, a fully connected computing unit, and an internal image buffer unit;
[0041] The controller is used to start the image sorting unit to sort the image after receiving the input image, and after the sorting is completed, start the PE computing array or the fully connected computing unit for calculation, and cache the result generated by the calculation into the internal image buffer unit. If the current network layer is the output layer, the pulse firing frequency is statistically calculated to complete image recognition. If the current network layer is a non-output layer, the image sorting unit is started;
[0042] The PE computing array is used to complete the calculations of the convolutional layer and the pooling layer in the pulse neural network;
[0043] The fully connected computing unit is used to complete the calculations of the fully connected layer in the pulse neural network;
[0044] Among them, both the PE computing array and the fully connected computing unit include optimized LIF neurons and optimized Izhikevich neurons obtained by using the first-order Euler method to eliminate the differential calculations of solving LIF neurons and Izhikevich neurons.
[0045] In the embodiments of this application, in order to improve the calculation speed and reduce the occupation of computing resources, the neurons in the neuron computing unit are optimized to provide a pulse convolutional neural network accelerator that supports optimized LIF neurons and Izhikevich neurons. Optimizing the exponential operation in the process of solving LIF neurons includes using the first-order Euler method to eliminate the differential calculations of solving, and fixing the parameters, using arithmetic shift to replace multiplication calculation. Using the first-order Euler method to eliminate the differential calculations of solving LIF neurons, the calculation formula for eliminating the differential calculations of solving is obtained as:
[0046]
[0047] For the parameters in the calculation formula for eliminating the differential calculations of solving Perform solidification, use arithmetic phase shift to replace the multiplication operation in the calculation formula for eliminating the differential calculation of the solution, use arithmetic shift to replace the multiplication calculation, including achieving the effect of multiplying by 0.125 by arithmetic right shift three bits on hardware, and adopt the second-order predictor-corrector method to optimize the calculation formula for eliminating the differential calculation of the solution, obtaining the optimized calculation formula. The optimized calculation formula is:
[0048]
[0049]
[0050]
[0051]
[0052] In the formula, V[n] is the membrane state at the current moment, and V[n + 1] is the magnitude of the membrane potential to be obtained. is the membrane potential voltage obtained by using the first-order Euler method, and t n is the nth discrete time step during the calculation, and v n is the magnitude of the input membrane potential, α is the R of the neuron m The influence of I on the membrane potential of the neuron at this time, R m is the neuron membrane resistance constant, I is the input current value, f1(t, v) is the first target equation, f2(t, v) is the second target equation, β1 and β2 are the magnitude biases of V in different stages, and V reset is the neuron resting potential, h is the time step, and τ m is the time constant. In the embodiments of the present application, some of the above parameters are fixed. Among them, V reset is set to 0, α is set to 0.5, is set to 0.125, and the simplified calculation formula is:
[0053] V[n + 1] = V[n] - (y1 + y2);
[0054] y1 = (V[n] - β1) >> 4;
[0055]
[0056]
[0057] In the formula, >>4 is a logical left shift by 4 bits, >>3 is a logical left shift by 3 bits, and y1 and y2 are intermediate parameters.
[0058] For the Izhikevich neuron (IZH neuron), its optimization process is similar to the process of optimizing the LIF neuron. The first-order Euler method is used to implement dV / dt and dU / dt in the formula, that is:
[0059]
[0060]
[0061] if V>V threhold , Spike and V=c, U=U + d;
[0062] Wherein, V is the membrane voltage, U is the membrane potential recovery variable, I is the input current of the neuron, a, b, c, d are the model parameters of the neuron, and V threhold is the membrane voltage threshold.
[0063] Although there is still multiplication calculation in the optimized IZH neuron, the exponential calculation has been greatly reduced. The calculation formula of the optimized Izhikevich neuron is:
[0064] V[n + 1] = (0.04V 2 + 5V + 140 - U + I)·h;
[0065] U[n + 1] = [a(bV - U)]·h.
[0066] The embodiment of the present application uses the first-order Euler method to eliminate the differential calculation in the neuron solving process, reduces the calculation complexity, helps to improve the calculation speed of the neuron computing unit, can support the spiking neural network of different spiking neurons, and can support the spiking neural network of different topologies including multi-layer perceptrons, traditional, depthwise separable, and residual convolutional networks, etc.
[0067] The embodiment of the present application constructs an accelerator based on the optimized neuron, including a controller, an image sorting unit, a neuron computing unit, and an internal image buffer unit. The neuron computing unit includes a PE computing array and a fully connected computing unit. After receiving the input image, the controller sends a start signal to the image sorting unit to start the image sorting unit to sort the image. After the sorting is completed, the controller starts the PE computing array or the fully connected computing unit for calculation according to the current network layer. If the current network layer is a convolutional layer or a pooling layer, the PE computing array is started for calculation. If the current network layer is a fully connected layer, the fully connected computing unit is started for calculation, and the result generated by the calculation is cached to the internal image buffer unit. If the current network layer is an output layer, the spike firing frequency is counted to complete image recognition. If the current network layer is a non-output layer, the image sorting unit is started.
[0068] When receiving a start signal, the image sorting unit determines whether it is the input layer to select a data source. If so, it sorts the input images. If not, it sorts the intermediate buffered images in the internal image buffer unit. After the sorting is completed, a pulse sequence sorting completion signal is generated to indicate the completion of sorting. The image sorting unit mainly pre-fetches data for convolution calculation and pooling calculation. There is a Finite State Machine (FSM) in the image sorting unit to control the working state of the image sorting unit. When the received start signal is at a high level, it starts the pre-fetching of pulse data. For the input layer, it is necessary to complete the Poisson coding design of the image data. For the convolutional layer, it is necessary to complete the pre-fetching of pulse data and internal sorting. For the pooling operation, only pre-fetching is required. When performing a convolution operation, due to the limited size of the PE computing array, the current designed size of the PE computing array is 16×16. Therefore, a pulse image cannot be fetched in one go, so there is a buffered state in the FSM. When the calculation of the pre-fetched pulses is completed, the data pre-fetching operation continues until the convolution calculation is completed. At the same time, the pulse pre-fetching for the convolution operation is stored in a multi-bit wide register, and the output of the systolic data is completed according to the array scheduling signal of the controller. For the pooling operation, the image sorting unit is only responsible for generating the pre-fetch signal and address, and transmitting them to the PE computing array. The PE computing array will cache the corresponding data of the pooling block into the array according to the pre-fetch signal. The buffer is in the PE computing array for convenient direct pooling operation. When the input data is ready, the image sorting unit generates a pulse sorting completion signal to indicate completion.
[0069] The internal image buffer unit is used to perform two-dimensional / three-dimensional expansion on the one-dimensional / two-dimensional data output by the PE computing array, including data processing for convolution, pooling, and fully connected layer calculations.
[0070] The fully connected computing unit is used to implement the fully connected calculation in the pulse neural network. The fully connected layer computing unit first sequentially obtains pulses from the internal image buffer unit. When there is a high-level pulse, it buffers the address of the pulse into the corresponding FIFO. After the traversal is completed, it updates the corresponding fully connected neurons according to the empty / full state of the FIFO to complete the fully connected calculation. Among them, the fully connected computing unit updates the fully connected neurons based on the optimized neurons.
[0071] The accelerator in this application further includes a weight storage unit and a membrane potential storage unit. The weight storage unit is used to store and read weights, mainly for reading the weights of convolution operations. There is also an FSM in the weight storage unit. When the start signal sent by the controller is at a high level, it starts to read the weights. Since the images of one channel may need to be multiplexed several times to complete a convolution calculation, the weights read for the first time can be multiplexed to improve the calculation efficiency. When the weights are ready, an input weight ready signal will be generated to indicate completion. The membrane potential storage unit is used to store and read and write the neuron membrane potential.
[0072] The PE calculation array is used to complete the calculations of the convolution layer and the pooling layer in the spiking neural network, and reads and writes weights and membrane potentials from the weight storage unit and the membrane potential storage unit respectively to realize the update of the neuron membrane potential. The PE calculation array includes the design of a single PE unit and the combination of PE array data. For the PE unit, when performing a convolution operation, it includes a data selector and a register to save the intermediate calculation value. When an input pulse is detected, the weight is accumulated into the above register. When a single convolution operation is completed, the final membrane potential is obtained by adding the membrane potential of the input neuron and the above register. When the potential exceeds the threshold, a pulse is emitted and the potential is set to 0. When performing a pooling operation, the current embodiment of this application can implement 2×2 pooling calculation. There is an FSM for the pooling operation in the PE unit to count the number of pulses in the input image, and then the value to be accumulated this time obtained according to the number of pulses is saved to the above register to obtain the final neuron membrane potential and internal pulses. The PE calculation array, which contains 16×16 PE units, is mainly responsible for integrating and outputting the results of the systolic array, including the membrane potential of the convolution neuron and the internal output pulses. For an image with a single-channel input, a neuron is only related to the pixels of one channel, that is, the calculation is performed once. However, for an image with a multi-channel input, a neuron is related to all the pixels of the input channels. In the embodiment of this application, the OR method is adopted. Specifically, when the input image has 16 channels, for the first channel, the basic calculation method is used to obtain the membrane potential and output pulses. However, when calculating the following 15 channels, the membrane potential is still calculated normally. However, for the output pulses, an OR operation needs to be performed with the previous calculation, so that multi-channel convolution calculation can be realized.
[0073] Numerical constraints of the convolution intermediate result, Weight represents the weight, internal_result represents the accumulated intermediate result, and result represents the final accumulated result:
[0074]
[0075] Numerical constraints of the PE computing array, where MP represents the membrane potential of neurons and Spike represents the spike firing situation:
[0076] if MP > 1:
[0077] MP = 0;
[0078] Spike = 1;
[0079] else if MP < -10:
[0080] MP = -10;
[0081] Spike = 0;
[0082] else:
[0083] MP = MP;
[0084] Spike = 0.
[0085] The controller completes the calculations of each layer of the spiking neural network (including convolution, pooling, fully connected, etc.) through the coordinated control of the above units. Before the calculation, the controller needs to read and configure the calculation parameters. There is a ROM in the controller, which stores the detailed parameters of each layer of the network, as shown in Table 1. The FSM state transition in the controller is as Figure 2 shown.
[0086] Table 1
[0087] Signal Name Corresponding Bit Number Signal Meaning Input_Channel 79-64 Number of Channels of the Input Map PE_Array_Mult_Times 63-48 Number of Times of Reusing the PE Array Kernel_Size 47-32 Size of the Convolution Kernel Window_Size 31-16 Image Size of the Output Map Stride 15-0 Convolution Stride
[0088] In the embodiments of the present application, the neurons in the neuron computing unit are optimized neurons obtained by using the first-order Euler method to eliminate the differential calculations of LIF neurons and Izhikevich neurons. The calculation processes of LIF neurons and Izhikevich neurons are optimized, the calculation complexity is reduced, the speed of updating the membrane potential of neurons in the neuron computing unit is accelerated, less hardware resources are occupied, and it can support spiking neural networks of different spiking neurons, improving the technical problems in the prior art that there is a single neuron model, it is impossible to support SNNs built by multiple different spiking neurons, and the hardware resources, calculation amount, and power consumption occupied by the computing unit are very large.
[0089] In the description of the present application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0090] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0091] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0092] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0093] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0094] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks, or optical discs.
[0095] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A pulse convolutional neural network accelerator, characterized in that, Including: A controller, an image sorting unit, a PE computing array, a fully-connected computing unit, and an internal image buffer unit; The controller is configured to, after receiving an input image, start the image sorting unit to sort the image, and after the sorting is completed, start the PE computing array or the fully-connected computing unit to perform calculations, and cache the results generated by the calculations to the internal image buffer unit. If the current network layer is the output layer, the spike firing frequency is statistically calculated to complete image recognition. If the current network layer is a non-output layer, the image sorting unit is started; The PE computing array is used to complete the calculations of the convolutional layer and the pooling layer in the spiking neural network; The fully-connected computing unit is used to complete the calculations of the fully-connected layer in the spiking neural network; Wherein, both the PE computing array and the fully-connected computing unit include optimized LIF neurons and optimized Izhikevich neurons obtained by using the first-order Euler method to eliminate the differential calculations of the LIF neurons and the Izhikevich neurons; The optimization process of the LIF neurons is as follows: Using the first-order Euler method to eliminate the differential calculations of the LIF neurons, the calculation formula for eliminating the differential calculations is obtained as: ; For the parameters in the calculation formula of the differential calculation for elimination solution Solidify, use arithmetic phase shift to replace the multiplication operation in the calculation formula of the differential calculation for elimination solution, and optimize the calculation formula of the differential calculation for elimination solution by using the second-order predictor-corrector method to obtain the optimized calculation formula, and the optimized calculation formula is: ; ; , ; , ; Where, V[n] is the membrane state at the current moment, and V[n+1] is the magnitude of the membrane potential to be obtained. is the membrane potential voltage obtained by using the first-order Euler method. is the nth discrete time step during calculation. is the magnitude of the input membrane potential. is the R of the neuron m I is the influence magnitude of the membrane potential of the neuron at this time, and R m is the neuron membrane resistance constant, I is the input current of the neuron, f1(t, v) is the first target equation, and f2(t, v) is the second target equation. 、 are the magnitude biases of V in different stages. is the neuron resting potential, h is the time step. is the time constant; The calculation formula for the optimized Izhikevich neurons is: ; ; In the formula, V is the membrane voltage, U is the membrane potential recovery variable, I is the input current of the neuron, and a and b are the model parameters of the neuron.
2. The pulse convolutional neural network accelerator according to claim 1, wherein The image sorting unit is specifically configured to: When receiving a start signal, determine whether it is the input layer to select the data source. If so, sort the input image. If not, sort the intermediate buffered image in the internal image buffer unit; After the sorting is completed, generate a spike sequence sorting completion signal.
3. The pulse convolutional neural network accelerator according to claim 1, wherein It further includes: A weight storage unit for storing and reading the weights of the convolutional calculations.
4. The pulse convolutional neural network accelerator according to claim 3, wherein It further includes: A membrane potential storage unit for storing and reading the neuron membrane potential.
5. The pulse convolutional neural network accelerator according to claim 4, characterized in that, The PE computing array is specifically configured to: Read and write the weights and the membrane potential from the weight storage unit and the membrane potential storage unit respectively to implement the update of the neuron membrane potential.
6. The pulse convolutional neural network accelerator according to claim 1, characterized in that The PE computing array includes a plurality of PE units; The PE unit is configured to, when performing a convolutional operation, accumulate the weights to a register when detecting an input spike, and after the convolutional operation is completed, add the membrane potential of the input neuron to the register to obtain the final membrane potential; when performing a pooling operation, count the number of spikes in the input image through a finite state machine, and save the currently required accumulated value determined according to the number of spikes to the register to obtain the final neuron membrane potential and internal spikes.
Citation Information
Patent Citations
Impulse neural network reward optimization method and device, electronic equipment and storage medium
CN113822416A
Spiking neural network-based short-range tracking method and system
WO2021012752A1