Storage optimization method for pulse convolutional neural network accelerator
By building a network with less pulse training in the SCNN accelerator and using the time and space mask modules, the energy consumption problem of SCNN while maintaining high accuracy is solved, and significant energy consumption reduction and accuracy loss control are achieved.
Patent Information
- Application Number
- CN202510340922.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-21
AI Technical Summary
How to reduce its energy consumption while maintaining the high accuracy of pulsed convolutional neural networks (SCNNs), especially in inference accelerators. How to make SCNN accelerators play their energy-efficient advantages.
A storage optimization method for SCNN accelerator is proposed. By building a SCNN trained with less pulses and adding a time and space mask module, the number of pulses and memory accesses are reduced, thereby reducing energy consumption.
The energy consumption of SCNN is significantly reduced, the energy consumption of memory access is reduced by 20.04%-54.93%, the total energy consumption is reduced by 50.48%, and the accuracy loss is less than 1%.
Smart Images

Figure CN120146124A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural network processors, and particularly relates to a storage optimization method for a pulse convolutional neural network accelerator. Background Art
[0002] Whether it is an artificial neural network or a spiking neural network, processing complex tasks often comes at the cost of consuming more computing resources, and methods such as increasing the number of neurons and deepening the network structure are used to improve the network accuracy. In order to meet the demand for efficiently processing complex tasks, in recent years, people's research on spiking convolutional neural networks (SCNNs) has gradually deepened. In addition, accelerators for SCNNs are generally designed as dedicated computing chips in the field, which have higher energy consumption and computing efficiency compared to general computing chips such as CPUs and GPUs. Currently, the research goals for SCNNs mainly focus on improving network accuracy, while the advantages of low energy consumption and high energy efficiency ratio of spiking neural networks are often ignored in the process of repeatedly improving accuracy. Especially for the accelerator of SCNN inference, how to make the SCNN accelerator still play the advantage of high energy efficiency while maintaining a high accuracy rate is a challenge faced by the research on SCNNs.
[0003] Similar to the shallow spiking neural network (VSNN) with time encoding, the SCNN commonly uses the LIF neuron model. The difference is that the SCNN has a network structure similar to that of a convolutional neural network, its weight scale is more massive, and in the convolutional operation of its accelerator, as long as there are spikes in the input feature map of a certain channel, it is inevitable to access the weights corresponding to this channel. Summary of the Invention
[0004] The problem to be solved by the present invention is to reduce the access to heavy weights, and a storage optimization method for a pulse convolutional neural network accelerator is proposed.
[0005] To achieve the above object, the present invention is realized through the following technical solutions:
[0006] A storage optimization method for a pulse convolutional neural network accelerator includes the following steps:
[0007] S1. Construct less-spike training of a pulse convolutional neural network including a spike rate, train the pulse convolutional neural network to obtain a pulse convolutional neural network accelerator with fewer spikes;
[0008] S2. For the low-pulse pulse convolutional neural network accelerator obtained in step S1, equip each neuron with a time masking module. The time masking module consists of a comparator and an AND gate. Design a time masking method to perform time masking processing on the low-pulse pulse convolutional neural network;
[0009] S3. For the low-pulse pulse convolutional neural network accelerator after time masking processing in step S2, equip each neuron with a spatial masking module. The spatial masking module consists of a counter and a comparator. Design a spatial masking method to perform spatial masking processing on the low-pulse pulse convolutional neural network accelerator after time masking processing, and complete a storage optimization for the pulse convolutional neural network.
[0010] Further, the specific implementation method of step S1 includes the following steps:
[0011] S1.1. On the basis of the original training objective function L, add a pulse rate multiplied by a preset coefficient d to obtain the objective function L with the pulse rate. s The expression is as follows:
[0012]
[0013] Among them, L is the original training objective function, which is the cross-entropy between the output of the neural network and the training set labels. is the average pulse rate, which is calculated using the number of pulses and the number of time windows counted by the previous layers. d is a preset coefficient used to adjust the proportion of the pulse rate in L. s in;
[0014] S1.2. Use the objective function L with the pulse rate. s , and use the gradient descent method for training. Calculate the partial derivatives of the objective function with respect to the weight w and the neuron threshold voltage v. th of, and update the parameters according to the magnitude of the derivative values. The calculation formulas are as follows;
[0015] The calculation formula for training the weight W is as follows:
[0016]
[0017] Among them, w ij is the element in the i-th row and j-th column of the weight matrix;
[0018] has no functional relationship with w ij ,
[0019] The calculation formula for training the neuron threshold voltage v th is as follows:
[0020]
[0021] Among them, is the membrane potential of the \(i\)-th neuron in the \(l\)-th layer, \(t\) represents the time step, indicating the time progress calculated by the neuron within a time window \(T\) w where \(0\leq t\leq T\) w ; \(l\) represents the layer number of the neuron, is the output spike of the \(i\)-th neuron in the \(l\)-th layer, \(o\) represents the output spike. For the adopted neuron model, \(a\) is the coefficient for fitting the spike emission process, set to 0.5, and \(\text{sign}(\cdot)\) represents the sign function, outputting the positive or negative sign of the input value of the sign function. For example, if the input is a positive number, it outputs 1, otherwise it outputs -1.
[0022] Furthermore, the design at the hardware level in step S2 is to equip each neuron in the SCNN with a time mask module. The time mask module consists of a comparator and an AND gate. The comparator has two input ports. The first input port is the output of the time step counter, and the second input port inputs the value \(T\) ap , the output of the comparator and the output of the neuron are connected to the AND gate, and the output of the AND gate is used as the final output of the neuron.
[0023] Furthermore, the specific implementation method of the time mask method designed in step S2 is that for the neurons in the SCNN, a comparator is added at the output end of the neuron. According to the time step \(t\) performed by the neuron, if the set threshold \(T\) ap is not reached, the neuron is allowed to emit a spike when the membrane potential \(u\) reaches \(v\) th , if the time step \(t\) has reached \(T\) ap , the neuron is prohibited from emitting a spike; \(T\) ap In the inference process, the calculation is performed according to the following formula:
[0024]
[0025] Among them, \(T\) w represents the set time window and is an input parameter of the SCNN; can be calculated based on the output spikes of the neurons before the \(l\)-th layer during the inference calculation process of the SCNN.
[0026] Furthermore, the design at the hardware level in step S3 is to add a counter to each channel in the convolutional layer, and use the accumulative method to count the input spikes of each channel. Then, the output \(C\) of the counter and the threshold of each convolutional layer are jointly input into a comparator. If The comparator outputs 1, indicating that the convolution calculation for this channel continues; otherwise, the comparator outputs 0, indicating that the calculation for this channel is skipped. The output of the comparator is connected to the memory access controller and the convolution layer calculation component of the memory, and the start / stop of these two components is controlled through the output signal of the comparator, thereby realizing the control of whether the convolution layer is skipped.
[0027] Further, the specific implementation method of the spatial mask method in step S3 is that SCNN has a convolution structure. The input data of the convolution layer is a feature map, and the feature map is the output pulse of the previous layer. The convolution layer performs a convolution operation on the input feature map and the weights in units of channels, and a threshold for each convolution layer is set for each convolution layer. Count the number of input pulses for each channel. When the number of input pulses is lower than the threshold of each convolution layer, this channel is skipped, and the corresponding weight access and convolution calculation are ignored. The calculation of the threshold for each convolution layer is as follows:
[0028]
[0029] where f h is the height of the feature map input to the convolution layer.
[0030] Advantages of the present invention:
[0031] A storage optimization method for a pulse convolutional neural network accelerator according to the present invention completes a memory access optimization method for reducing the memory access times of SCNN by reducing the number of pulses of SCNN and then further reducing the memory access times of SCNN. Since the memory access overhead of the SCNN accelerator accounts for a high proportion in the total energy consumption, the present invention can significantly reduce the energy consumption of SCNN by reducing the memory access overhead of SCNN; on the other hand, the pulse reduction method included in the present invention can also reduce the number of pulses, thereby reducing the amount of calculation. The present invention optimizes the energy consumption of the SCNN accelerator from both aspects of memory access and calculation, and can achieve more energy consumption optimization effects with lower precision loss compared with the same type of methods (such as pruning, quantization, etc.). Description of the drawings
[0032] Figure 1 is a flowchart of a storage optimization method for a pulse convolutional neural network accelerator according to the present invention;
[0033] Figure 2 is a schematic diagram of the memory access optimization mechanism according to the present invention;
[0034] Figure 3 is a schematic diagram of an SCNN accelerator with a memory access optimization mechanism according to the present invention. Detailed implementation manners
[0035] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the specific embodiments described are only a part of the embodiments of the present invention, rather than all of the specific embodiments. The components of the specific embodiments of the present invention usually described and shown in the accompanying drawings here can be arranged and designed in various different configurations, and the present invention can also have other embodiments.
[0036] Therefore, the following detailed description of the specific embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents the selected specific embodiments of the present invention. All other specific embodiments obtained by those skilled in the art based on the specific embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0037] To further understand the content, features and effects of the present invention, the following specific embodiments are exemplified and accompanied by Figure 1 - Appendix Figure 3 The details are as follows:
[0038] Example 1:
[0039] A storage optimization method for a pulse convolutional neural network accelerator includes the following steps:
[0040] S1. Construct few-pulse training of a pulse convolutional neural network including a pulse rate, train the pulse convolutional neural network to obtain a few-pulse pulse convolutional neural network accelerator;
[0041] Furthermore, training the pulse convolutional neural network is a training method that linearly relates the pulse rate and incorporates it into the training objective function;
[0042] Furthermore, the specific implementation method of step S1 includes the following steps:
[0043] S1.1. On the basis of the original training objective function L, add a pulse rate multiplied by a preset coefficient d to obtain the objective function L s The expression of is as follows:
[0044]
[0045] where L is the original training objective function, which is the cross-entropy between the output of the neural network and the training set labels, is the average pulse rate, calculated using the number of pulses and the number of time windows counted by the previous layers, and d is a preset coefficient used to adjust the proportion of the pulse rate in L s ;
[0046] S1.2. Using the objective function L containing the pulse rate s , training will be carried out using the gradient descent method, calculating the partial derivatives of the objective function with respect to the weight w and the neuron threshold voltage v th , and updating the parameters according to the magnitudes of the derivative values. The calculation formulas are as follows;
[0047] The calculation formula for training the weight W is as follows:
[0048]
[0049] where, w ij is the element in the i-th row and j-th column of the weight matrix;
[0050] has no functional relationship with w ij ,
[0051] The calculation formula for training the neuron threshold voltage v th is as follows:
[0052]
[0053] where, is the membrane potential of the i-th neuron in the l-th layer, t represents the time step, indicating the time progress calculated by the neuron within a time window T w , 0 ≤ t ≤ T w ; l represents the layer number of the neuron, is the output pulse of the i-th neuron in the l-th layer, o represents the output pulse. For the adopted neuron model, a is the coefficient for fitting the pulse emission process, set to 0.5, sign(·) represents the sign function, outputting the positive or negative sign of the input value of the sign function. For example, if the input is a positive number, it outputs 1, otherwise it outputs -1.
[0054] Furthermore, the SCNN contains a large number of spiking neurons. The neurons have variables such as the membrane potential u and the threshold voltage v th . The membrane potential u is a variable that gradually accumulates and increases as the neuron receives input pulses. When it reaches the set threshold voltage v th , the neuron emits a pulse to the neurons in the next layer and resets the membrane potential u to 0. Here, the threshold voltage v th affects the accuracy of the SCNN in the same way as the weight W. Therefore, the optimal value can be obtained through training instead of manually inputting a fixed value. Under such a neuron model, to calculate the neuron threshold voltage v thThe training calculation formula requires calculating the gradient of the objective function with respect to the membrane potential first, and then substituting it to calculate the gradient of the objective function with respect to the threshold voltage v th of.
[0055] S2. For the low-pulse pulse convolutional neural network accelerator obtained in step S1, equip each neuron with a time masking module. The time masking module consists of a comparator and an AND gate. Design a time masking method to perform time masking processing on the low-pulse pulse convolutional neural network;
[0056] Furthermore, the hardware-level design of step S2 is to equip each neuron in the SCNN with a time masking module. The time masking module consists of a comparator and an AND gate. The comparator has two input ports. The first input port is the output of the time step counter, and the second input port inputs the value T ap , the output of the comparator and the output of the neuron are connected to the AND gate, and the output of the AND gate is used as the final output of the neuron;
[0057] Furthermore, the specific implementation method of the time masking method in step S2 is that for the neurons in the SCNN, a comparator is added at the output end of the neuron. According to the time step t performed by the neuron, if the set threshold T is not reached ap then allow the neuron to emit a pulse when the membrane potential u reaches v th , if the time step t has reached T ap , then prohibit the neuron from emitting a pulse; T ap In the inference process, the calculation is performed according to the following formula:
[0058]
[0059] where, T w represents the set time window and is an input parameter of the SCNN; can be calculated based on the output pulses of the neurons before the l-th layer during the inference calculation of the SCNN; Therefore, an appropriate T can be set for each layer of neurons in real time ap .
[0060] Furthermore, time masking is different from directly shortening the length of the time window of the SCNN. Directly shortening the length of the time window will reduce the size of the tensors transmitted internally, which will have a greater impact on the inference accuracy; while time masking only reduces pulses and computational load by pausing the operation of neurons, and does not change the size of the output tensors of neurons, having a smaller impact on accuracy.
[0061] S3. For the few-pulse pulse convolutional neural network accelerator after time mask processing in step S2, equip each neuron with a spatial mask module. The spatial mask module consists of a counter and a comparator. Design a spatial mask method to perform spatial mask processing on the few-pulse pulse convolutional neural network accelerator after time mask processing, and complete a storage optimization for the pulse convolutional neural network.
[0062] Further, the hardware-level design of step S3 is to add a counter to each channel in the convolutional layer, accumulate to count the input pulses of each channel, and then input the output C of the counter and the threshold of each convolutional layer into a comparator together. If the comparator outputs 1, it means that this channel continues to perform convolutional calculation; otherwise, the comparator outputs 0, indicating to skip the calculation of this channel. The output of the comparator is connected to the memory access controller and the convolutional layer calculation component of the memory, and the start / stop of these two components is controlled by the output signal of the comparator, realizing the control of whether the convolutional layer is skipped.
[0063] Further, the specific implementation method of designing the spatial mask method in step S3 is that the SCNN has a convolutional structure. The input data of the convolutional layer is a feature map, and the feature map is the output pulse of the previous layer. The convolutional layer performs a convolutional operation on the input feature map and the weight value in units of channels, and sets the threshold of each convolutional layer for each convolutional layer to count the number of input pulses of each channel. When the number of input pulses is lower than the threshold of each convolutional layer, this channel is skipped, and the corresponding weight access and convolutional calculation are ignored. The calculation of the threshold of each convolutional layer is as follows:
[0064]
[0065] where f h is the height of the feature map input to the convolutional layer.
[0066] For the storage optimization method for the pulse convolutional neural network accelerator described in this embodiment, design an SCNN accelerator in the 28nm process library, and use Cadence Genus 15.0 software to evaluate energy consumption and delay. In addition, accuracy tests are performed on three models, namely ResNet18, ResNet34, and 7B-wideNet, on a software platform based on the open-source PyTorch framework. The method proposed in this embodiment can reduce the pulses of the SCNN by 49.86%, and the energy consumption of memory access is reduced by 20.04% - 54.93%, and the total energy consumption drops by up to 50.48%, while the accuracy loss of all tests is less than 1%.
[0067] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.
[0068] Although the present application has been described above with reference to specific embodiments, various improvements can be made thereto and components thereof can be replaced with equivalents without departing from the scope of the present application. In particular, as long as there is no structural conflict, the various features in the specific embodiments disclosed in the present application can be combined with each other in any way, and the exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A storage optimization method for a pulse convolutional neural network accelerator, characterized in that: The steps include: S1. Constructing a pulse convolutional neural network including a few-pulse training of a pulse rate, training the pulse convolutional neural network, and obtaining a few-pulse pulse convolutional neural network accelerator; S2. For the pulse convolutional neural network accelerator with few pulses obtained in step S1, a time mask module is equipped for each neuron, the time mask module is composed of a comparator and an AND gate, a time mask method is designed, and a time mask processing is performed on the pulse convolutional neural network with few pulses; S3. For the pulse convolutional neural network accelerator with few pulses after the time mask processing in step S2, a spatial mask module is equipped for each neuron. The spatial mask module is composed of a counter and a comparator. A spatial mask method is designed to perform spatial mask processing on the pulse convolutional neural network accelerator with few pulses after the time mask processing, and a storage optimization for the pulse convolutional neural network is completed.
2. The storage optimization method for a pulse convolutional neural network accelerator according to claim 1, characterized in that: The specific implementation method of step S1 includes the following steps: S1.
1. On the basis of the original training objective function L, add a pulse rate multiplied by the preset coefficient d to obtain the objective function L containing the pulse rate s The expression is as follows: Among them, L is the original training objective function, which is the cross entropy between the output of the neural network and the training set label. is the average pulse rate, which is calculated using the number of pulses and time windows counted in the previous layer, and d is the preset coefficient used to adjust the pulse rate in L s The proportion of S1.
2. Using the objective function L containing the pulse rate s , the gradient descent method will be used for training to calculate the objective function relative to the weight w and the neuron threshold voltage v th The partial derivative of , and the parameters are updated according to the value of the derivative. The calculation formula is as follows; The calculation formula for weight W training is as follows: Among them, w ij is the element in the i-th row and j-th column of the weight matrix; With w ij There is no functional relationship. Neuron threshold voltage v th The training calculation formula is as follows: in, is the membrane potential of the i-th neuron in the l-th layer, t represents the time step, which means that the neuron is in a time window T w The time progress calculated within, 0≤t≤T w ; l represents the number of neuron layers, is the output pulse of the ith neuron in the lth layer, o represents the output pulse, and for the neuron model used, a is the coefficient of the pulse emission process, which is set to 0.
5. sign(·) represents the sign function, which outputs the sign of the input value of the sign function. If a positive number is input, 1 is output, otherwise -1 is output.
3. The storage optimization method for a pulse convolutional neural network accelerator according to claim 2, characterized in that: The hardware level design of step S2 is to equip each neuron in SCNN with a time mask module. The time mask module consists of a comparator and an AND gate. The comparator has two input ports. The first input port is the output of the time step counter, and the second input port is the input value T. ap The output of the comparator and the output of the neuron are connected to the AND gate, and the output of the AND gate is used as the final output of the neuron.
4. The storage optimization method for a pulse convolutional neural network accelerator according to claim 3, characterized in that: Step S2 designs a time mask method. The specific implementation method is to add a comparator to the output end of the neuron in the SCNN, and make a judgment based on the time step t of the neuron. If it does not reach the set threshold T ap Then the neuron is allowed to reach v when the membrane potential u th When a pulse is emitted, if the time step t has reached T ap , the neuron is prohibited from emitting pulses; T ap The calculation is performed according to the following formula during the inference process: Among them, T w Represents the set time window, which is the input parameter of SCNN; It can be calculated during the inference calculation process of SCNN based on the output pulses of neurons before layer l.
5. The storage optimization method for a pulse convolutional neural network accelerator according to claim 4, characterized in that: The hardware level design of step S3 is to add a counter to each channel of the convolution layer, count the input pulses of each channel in a cumulative manner, and then compare the output C of the counter with the threshold of each convolution layer. Common input to a comparator, if The comparator outputs 1, indicating that the channel continues to perform convolution calculations; otherwise, the comparator outputs 0, indicating that the calculation of the channel is skipped; the output of the comparator is connected to the memory access controller and the convolution layer calculation component of the memory, and the start / stop of these two components is controlled by the output signal of the comparator to control whether the convolution layer is skipped.
6. The storage optimization method for a pulse convolutional neural network accelerator according to claim 5, characterized in that: The specific implementation method of the spatial masking method designed in step S3 is that SCNN has a convolutional structure, the input data of the convolutional layer is the feature map, the feature map is the output pulse of the previous layer, the convolutional layer convolves the input feature map with the weight in units of channels, and sets the threshold of each convolutional layer for each convolutional layer. The number of input pulses of each channel is counted. When the number of input pulses is lower than the threshold of each convolution layer, the channel is skipped and the corresponding weight access and convolution calculation are ignored. The threshold of each convolution layer is calculated as follows: Among them, f h It is the height of the feature map input to the convolutional layer.
Citation Information
Patent Citations
Spatial correlation feature extraction in audio processing based on neural network
CN115497495A
Neuromorphic convolution calculation accelerator based on scalable channel
CN118468963A
Hardware efficient weight structure for sparse deep neural networks
US20230244746A1
Cited By
Liquid state machine dynamics optimization system and method based on self-feedback mask
CN121436060A