Multi-model fusion neural network accelerator supporting out-of-order sparse output
By designing a multi-model fusion neural network accelerator that supports out-of-order sparse output and decoupling the calculation process of the computing unit, the problems of hardware resource waste and low utilization in the existing technology are solved, and efficient calculation of multiple neural network models is achieved.
Patent Information
- Application Number
- CN202510478711.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing neural network accelerators are unable to unify different models, resulting in large area overhead and waste of hardware resources. In particular, they are powerless for models with completely different computing mechanisms, resulting in low hardware utilization.
A multi-model fusion neural network accelerator that supports out-of-order sparse output is designed, including a data cache module, a computing unit, a prediction module, a serial-in-serial-out register, a post-processing module, and an output module. Through a dynamic allocator and an asynchronous handshake protocol, the calculation process of the computing unit is decoupled to support multiple neural network models.
It improves hardware utilization, reduces computing unit waiting time, and achieves efficient computing and resource optimization for different models.
Smart Images

Figure CN120633742A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-model fusion neural network accelerator that supports out-of-order sparse output. Background Art
[0002] The computation process of neural network models requires hardware support. Because neural network training and inference require extensive computation and storage, traditional central processing units (CPUs) often cannot meet the high-performance and low-latency requirements, while large GPUs cannot meet the power consumption and latency constraints of the edge. Therefore, to overcome this shortcoming, deploying neural network models on mobile devices or edge devices requires hardware accelerators specifically designed for neural network computation. At the data flow level, due to factors such as data sparsity and data size, the payloads of various computational units in a neural network accelerator often vary significantly, resulting in inconsistent computation progress. This inevitably results in units that complete computation earlier waiting for later units to complete, reducing hardware utilization. In related technologies, accelerators designed for different model computation methods and input data models are typically specialized. For example, spiking neural networks (SNNs) incorporate an additional time dimension compared to convolutional neural networks (CNNs), requiring continuous computation across multiple time steps to achieve reliable results. Furthermore, the data format is single-bit pulses rather than multi-bit values. This results in significant differences in the data storage and computational data flow of the two accelerators. Therefore, accelerator fusion is necessary for different models. Traditional fusion methods include heterogeneous and homogeneous modes. The heterogeneous mode designs multiple dedicated accelerators that can communicate data with each other, reducing the cost of data exchange, but it still cannot unify different accelerators, resulting in a large area overhead. While the homogeneous mode can provide support for multiple models at the computing unit level or the functional module and data flow level, these methods are currently limited to convolutional models such as CNN and SCNN. They are powerless for models with completely different computing mechanisms, and their versatility is still low.
[0003] On the other hand, to increase the accelerator's computing power, a larger number of computing units are typically configured within the accelerator. During the computational process, these units may not receive the same tasks. For example, some units may be configured to compute the first layer of a neural network, while others may compute the second layer. This results in different data dimensions and quantities, leading to different computational speeds for each unit. Because the accelerator typically reads data from the cache uniformly and sends data to all computing units simultaneously, some faster units may complete their calculations first, waiting for later units to complete. During this waiting period, the earlier units remain idle, wasting both time and energy. Data sparsity can also contribute to this phenomenon: Since data 0 multiplied by any number always yields 0, it has no effect on the result and is therefore skipped during the computation. This can cause some units with more data 0 to compute faster than others, leading to lower utilization. Summary of the Invention
[0004] The present invention provides a multi-model fusion neural network accelerator that supports out-of-order sparse output, which is used to solve the defects of traditional multi-model fusion neural network accelerators, such as the inability to unify different accelerators, the large area overhead, the inability to handle models with completely different computing mechanisms, and the waste of hardware resources.
[0005] The present invention provides a multi-model fusion neural network accelerator supporting out-of-order sparse output, comprising: A data cache module is used to store weight data, input data and calculation parameters. The data cache module determines to send the stored data to the corresponding idle computing unit through a dynamic allocator; a plurality of computing units, configured to obtain weight data, input data, and computing parameters from the data cache module, each computing unit being configured to perform a multiplication-accumulation operation based on the weight data, input data, and computing parameters; Multiple prediction modules, each prediction module is connected to a corresponding computing unit, and predicts whether to terminate the calculation of subsequent time steps based on the multiplication and accumulation operation results; under the action of the prediction module, the multiplication and accumulation operation results of each computing unit are output at different times; Serial-in-serial-out registers, used to receive and cache the multiplication and accumulation results of each computing unit; a post-processing module, configured to extract the multiplication-accumulation operation results of each computing unit from the serial-input-serial-output register according to the calculated neural network type, and process the multiplication-accumulation operation results of each computing unit; The output module is used to output the processing result of the post-processing module based on the asynchronous handshake protocol.
[0006] The multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention also includes a preprocessing module for detecting the data in the data cache module and removing zero values or invalid time step data in the input data.
[0007] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the dynamic allocator uses a polling algorithm to distribute preprocessed data to idle computing units to balance the computing load.
[0008] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the data cache module is further used for: For convolutional neural network data width is 8 bits; For the pulse neural network data width of 1 bit, it supports neural network calculations of no more than 8 time steps. The data of these 8 time steps are stored in 8-bit data in sequence from the first time step to the last time step, and are multiplexed with the convolutional neural network data space.
[0009] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the data cache module supports a multi-mode storage structure and switches between convolution sliding window block storage and matrix multiplication row storage through control signals.
[0010] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the data cache module is further used for: For convolutional neural networks, the convolutional sliding window block storage is used to split the input feature map into blocks according to the convolutional sliding window. The data points in these blocks from the upper left to the lower right are stored together according to their corresponding positions and stored in different storage areas. One row in the storage area can store data from multiple channels at the same position. During calculation, the position after sliding is compared with the position before sliding to update the changed area. For network models that use matrix multiplication operations, matrix multiplication rows are stored and data is read row by row, so that the input data and the stored data correspond one to one.
[0011] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the multiple computing units are further used to: For convolutional neural networks or Transformer models, multi-bit data will directly enter the corresponding computing unit; For spiking neural network or spiking-transformer models, calculations are performed according to time serial rules or time parallel rules; The time serial rule includes taking out the pulse data of only one time step and not processing other time steps temporarily; the time parallel rule includes loading the pulse data of multiple time steps, taking out the pulse data of one time step in sequence according to the time step order and sending it to the calculation unit until the required time steps are taken.
[0012] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the post-processing module includes an activation module; The activation module uses a lookup table to perform nonlinear mapping for convolutional neural networks and Transformer models to process data output by the computing unit.
[0013] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the post-processing module includes an early exit module; The early exit module uses an early exit algorithm for the spiking neural network and the Spiking-Transformer model to determine whether to terminate the network model calculation early.
[0014] According to the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the present invention, the output module is further used for: For pulse neural networks, the membrane level data of the time step is updated, and the address of the output of the calculation result of each set of data is calculated according to the number of times the input data is read.
[0015] The multi-model fusion neural network accelerator supporting disordered sparse output provided by the present invention includes a data cache module for storing weight data, input data and calculation parameters, and the data cache module determines to send the stored data to the corresponding idle computing unit through a dynamic allocator; multiple computing units are used to obtain weight data, input data and calculation parameters from the data cache module, and each computing unit is configured to perform multiplication and accumulation operations based on the weight data, input data and calculation parameters; multiple prediction modules, each prediction module is connected to the corresponding computing unit, and predicts whether to terminate the calculation of subsequent time steps according to the result of the multiplication and accumulation operation; the multiplication and accumulation operation results of each computing unit under the action of the prediction module are The output time is different; a serial input-serial output register is used to receive and cache the multiplication and accumulation operation results of each computing unit; a post-processing module is used to extract the multiplication and accumulation operation results of each computing unit from the serial input-serial output register according to the calculated neural network type, and process the multiplication and accumulation operation results of each computing unit; an output module is used to output the processing results of the post-processing module based on an asynchronous handshake protocol. The present invention combines the characteristics of various neural network models, decouples the calculation processes of each computing unit, and makes them independent, and outputs sparse predictions through a prediction module to provide better dynamics, which can compress the waiting time of the computing unit to a minimum, thereby achieving maximum hardware utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 This is a functional structure diagram of a multi-model fusion neural network accelerator supporting out-of-order sparse output provided by an embodiment of the present invention; Figure 2 1. It is a comparison diagram between the embodiment of the present invention and the prior art; Figure 3 This is an example diagram of a data organization method that can reduce the number of data reads when the convolution window slides, provided by an embodiment of the present invention; Figure 4 This is a neural network acceleration data flow graph supporting out-of-order sparse output provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0019] Figure 1 The functional module diagram of the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the embodiment of the present invention is as follows: Figure 1 As shown, the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the embodiment of the present invention includes: A data cache module is used to store weight data, input data and calculation parameters. The data cache module determines to send the stored data to the corresponding idle computing unit through a dynamic allocator; a plurality of computing units, configured to obtain weight data, input data, and computing parameters from the data cache module, each computing unit being configured to perform a multiplication-accumulation operation based on the weight data, input data, and computing parameters; In this embodiment of the present invention, there are 16 computational units, each of which can receive up to 144 pairs of 8-bit input and 8-bit weight data (a 16-channel 3×3 convolution kernel). Each pair is a synapse. The computational unit primarily multiplies these 144 pairs of data and accumulates the products.
[0020] Multiple prediction modules, each prediction module is connected to a corresponding computing unit, and predicts whether to terminate the calculation of subsequent time steps based on the multiplication and accumulation operation results; under the action of the prediction module, the multiplication and accumulation operation results of each computing unit are output at different times; The prediction module predicts whether the received data is positive, negative, or zero. It uses algorithms such as bit-by-bit product calculation and new coded product calculation to achieve spatial output sparsity. For example, the calculation of subsequent bits can be terminated based on the result. Each computation unit is equipped with a prediction module, supporting the prediction of the multiplication and accumulation results of up to 16 channels. To provide output prediction capabilities for a larger number of channels, a shared prediction module is also included for every 2, 4, 8, or 16 computation units, supporting output prediction for up to 256 channels.
[0021] Serial-in-serial-out registers, used to receive and cache the multiplication and accumulation results of each computing unit; In the embodiment of the present invention, the output results of each calculation unit will not be output simultaneously under the action of the prediction module. Therefore, a serial-in-serial-out register (FIFO) is designed to save all the calculation results to prevent loss and pass them through subsequent processing in sequence.
[0022] a post-processing module, configured to extract the multiplication-accumulation operation results of each computing unit from the serial-input-serial-output register according to the calculated neural network type, and process the multiplication-accumulation operation results of each computing unit; The output module is used to output the processing result of the post-processing module based on the asynchronous handshake protocol.
[0023] In the embodiment of the present invention, the output module sends the result data and its corresponding coordinates out of the accelerator in sequence.
[0024] Neural network models include CNN, SNN, Transformer, etc., which can play a role in different applications or environments. However, common neural network accelerators are usually designed for a certain application or a certain network model, and usually perform specific optimization and preprocessing based on the special conditions of the model. Since the calculation methods and input data of different models vary greatly, the design of accelerators is usually dedicated, which limits the development of accelerators. Figure 2 As shown in (1), these methods are basically limited to convolutional models such as CNN and SCNN, and are powerless for models with completely different computing mechanisms. Their versatility is still not high. Related technologies are usually implemented through software or hardware methods, such as Figure 2 As shown in (2), some backlogged tasks are deployed from the originally deployed computing units to other units, and the task load between the units is balanced as much as possible, so that the computing units maintain similar computing speeds in each round of calculation, thereby reducing the waiting time between them. However, this method can only reduce, but not completely eliminate, the waiting time, and there is still room for improvement in hardware utilization.
[0025] To reduce the idle time of computing resources, after each computing unit completes a round of tasks and outputs the multiplication and accumulation results, it determines whether all data written to the accelerator's cache SRAM has been calculated. If so, no more data is read out. Otherwise, the control logic reads another set of data from the cache, preprocesses it, and sends it to an idle computing unit. In other words, the reading process is sequential, while the calculation process is out of order. Therefore, it is not necessary for all computing units to be idle. As long as there are idle hardware resources, data can be read and calculated. This decouples the calculation processes between different computing units and compresses the idle time of computing units to the cache read time, which cannot be reduced.
[0026] The multi-model fusion neural network accelerator supporting disordered sparse output provided by the embodiment of the present invention includes a data cache module for storing weight data, input data and calculation parameters, and the data cache module determines to send the stored data to the corresponding idle computing unit through a dynamic allocator; multiple computing units are used to obtain weight data, input data and calculation parameters from the data cache module, and each computing unit is configured to perform multiplication and accumulation operations based on the weight data, input data and calculation parameters; multiple prediction modules, each prediction module is connected to the corresponding computing unit, and predicts whether to terminate the calculation of subsequent time steps according to the result of the multiplication and accumulation operation; the multiplication and accumulation operations of each computing unit under the action of the prediction module are performed. The result output time is different; a serial input-serial output register is used to receive and cache the multiplication and accumulation operation results of each computing unit; a post-processing module is used to extract the multiplication and accumulation operation results of each computing unit from the serial input-serial output register according to the calculated neural network type, and process the multiplication and accumulation operation results of each computing unit; an output module is used to output the processing results of the post-processing module based on an asynchronous handshake protocol. The present invention combines the characteristics of various neural network models, decouples the calculation processes of each computing unit, and makes them independent, and outputs sparse predictions through a prediction module to provide better dynamics, which can compress the waiting time of the computing unit to a minimum, thereby achieving maximum hardware utilization.
[0027] Based on any of the above embodiments, the multi-model fusion neural network accelerator that supports out-of-order sparse output also includes a preprocessing module for detecting the data in the data cache module and removing zero values or invalid time step data in the input data.
[0028] In the embodiment of the present invention, the preprocessing module is further used to generate data that may be required in the output sparse prediction process.
[0029] In an embodiment of the present invention, the dynamic allocator uses a round-robin algorithm to distribute the pre-processed data to idle computing units to balance the computing load.
[0030] In an embodiment of the present invention, after the prediction module processes, some units may still be calculating the value of a certain output point, while some units have already moved to other output locations. By identifying the output direction of the result of each calculation unit through the output address, the order will not be disordered, thereby reducing the idle time of the computing resources. In order to make the computing load distributed between the various computing units more balanced, when sending data after preprocessing, the distributor adopts a polling algorithm to distribute data, that is, the computing unit that received the data most recently has the lowest priority when it has data to send next time, and the priorities of the remaining computing units increase in sequence. Each time, data is sent to the idle unit with the highest priority. In this way, it is possible to avoid the situation where some units are always calculating while some units are idle for a long time.
[0031] Based on any of the above embodiments, the data cache module is further configured to: For convolutional neural network data width is 8 bits; For the pulse neural network data width of 1 bit, it supports neural network calculations of no more than 8 time steps. The data of these 8 time steps are stored in 8-bit data in sequence from the first time step to the last time step, and are multiplexed with the convolutional neural network data space.
[0032] In the embodiment of the present invention, calculation parameters such as the membrane potential of neurons, electrical parameter values, thresholds used in some calculations, etc., input data, weights and calculation parameters are all stored in SRAM.
[0033] In an embodiment of the present invention, the multi-model fusion neural network accelerator supporting out-of-order sparse output also includes a global register, and some important indicative parameters such as the model category calculated this time, the current time step, the current number of network layers, data coordinates, etc. are stored in the global register.
[0034] The parameters or weights of neurons can be reconfigured to fine-tune the network, supporting a large number of neuron models, such as the accumulation-release (IF) model, the leakage-accumulation-release (LIF) model, the degenerate Izhikevich model, etc. The number of synapses supported by each neuron is configurable, up to 2304. In order to support different data storage types at the same time, the CNN data width is specified to be 8 bits, and the SNN pulse width is specified to be 1 bit. Neural network calculations of up to 8 time steps are supported. The data of these 8 time steps are stored in 8-bit data in order from the first time step to the last time step. This achieves reuse of the CNN data space and improves storage efficiency.
[0035] In an embodiment of the present invention, the data cache module supports a multi-mode storage structure, and switches between convolution sliding window block storage and matrix multiplication row storage through a control signal.
[0036] In an embodiment of the present invention, the data cache module is further configured to: (1) For convolutional neural networks, the input feature map is split into blocks according to the convolution sliding window through convolution sliding window block storage, and the data points from the upper left to the lower right in these blocks are stored together according to the corresponding positions and stored in different storage areas respectively; One row in the storage area can store data from multiple channels at the same position. During calculation, the position after sliding is compared with the position before sliding to update the changed area. In this embodiment of the present invention, the input feature map is split into blocks using a convolutional sliding window. Data points within these blocks are stored together at corresponding positions from the top left to the bottom right, with different colors stored in separate storage areas. To improve data parallelism, a single row in a storage area can store data from multiple channels at the same position. During calculations, the post-sliding position is simply compared with the pre-sliding position and the changed areas are updated.
[0037] like Figure 3As shown in the figure, a 3×3 convolution kernel is used as an example. The input feature map is split into 3×3 blocks. The nine data points within each block, from the top left to the bottom right, are stored together at corresponding positions. Points of the same color represent data points at the same position within each block, while points of different colors are stored in nine separate storage areas. To improve data parallelism, a row in a storage area can store data for 16 channels at the same position, totaling 144 bits. During calculation, since the convolution window needs to slide, only the position after the slide is compared with the position before the slide, and the changed area is updated. For example, if the sliding step size is 1, all nine storage areas need to be opened simultaneously at the beginning, and the first row of data (including the channel corresponding to the point) is read out, corresponding to the first convolution window: A0, A1, A2, A3, A4, A5, A6, A7, and A8. The second calculation of the convolution window slides horizontally to A1, A2, A9, A4, A5, A12, A7, A8, and A15. Compared with the previous data, only three points need to be updated: A0→A9, A3→A12, and A6→A15. Therefore, only storage areas 0, 3, and 6 need to be opened. If all data needs to be retrieved again, all 9 storage areas need to be opened, which greatly reduces the overhead of reading the cache area. Similarly, the next time storage areas 1, 4, and 7 are opened, and the next time storage areas 2, 5, and 8 are opened, and then back to storage areas 0, 3, and 6. If vertical sliding is required, such as A0, A1, A2, A3, A4, A5, A6, A7, and A8 sliding to A3, A4, A5, A6, A7, A8, A21, A22, and A23, then storage areas 0, 1, and 2 need to be opened. Similarly, only 1 / 3 of the data needs to be updated, and so on. To reduce the frequency of window data updates, when sliding horizontally to the end of the input feature map, the window will not restart at the next row, but will slide down directly to keep 2 / 3 of the data unchanged. If the calculation is a spike network with time steps, the calculation is performed by taking 1 bit of data at each time step as described above until all data are taken.
[0038] (2) For network models that use matrix multiplication operations, matrix multiplication rows are stored and data is read row by row, so that the input data and the stored data correspond one to one.
[0039] Based on any of the foregoing embodiments, the multiple computing units are further configured to: For convolutional neural networks or Transformer models, multi-bit data will directly enter the corresponding computing unit; For spiking neural network or spiking-transformer models, calculations are performed according to time serial rules or time parallel rules; The time serial rule includes taking out the pulse data of only one time step and not processing other time steps temporarily; the time parallel rule includes loading the pulse data of multiple time steps, taking out the pulse data of one time step in sequence according to the time step order and sending it to the calculation unit until the required time steps are taken, without waiting until the current time step is calculated before taking the next time step.
[0040] Based on any of the above embodiments, the post-processing module includes an activation module; The activation module uses a lookup table to perform nonlinear mapping for convolutional neural networks and Transformer models to process data output by the computing unit.
[0041] In the embodiment of the present invention, the lookup table supports most nonlinear functions such as ReLu, Sigmoid, Softmax, etc., to process the results of the calculation unit. The distributor can select which unit and its corresponding serial input and output register to receive data for processing.
[0042] In the embodiment of the present invention, the post-processing module further includes an early exit module; The early exit module uses an early exit algorithm for the spiking neural network and the Spiking-Transformer model to determine whether to terminate the network model calculation early.
[0043] In an embodiment of the present invention, the early-retirement algorithm determines whether to terminate the calculation of subsequent time steps by comparing the pulse emission rate of the current time step with a preset threshold, thereby achieving output sparseness in the time dimension.
[0044] Based on any of the above embodiments, the output module is further configured to: For pulse neural networks, the membrane level data of the time step is updated, and the address of the output of the calculation result of each set of data is calculated according to the number of times the input data is read.
[0045] Based on any of the above embodiments, the multi-model fusion neural network accelerator that supports out-of-order sparse output controls data access, issuance, output address calculation, and other processes, as well as generates other required control signals. In the output address calculation, CNN and SNN can directly calculate according to the order of input data, while Transformer and Spiking-Transformer need to determine whether transposition is required based on the properties of the matrix being calculated. If transposition is required, the horizontal and vertical coordinates in the space can be directly exchanged, and the other dimensions remain unchanged.
[0046] The working process of this multi-model fusion neural network accelerator that supports out-of-order sparse output is divided into five stages: data acquisition, distribution, calculation, post-processing, and output. Before calculation, the data to be processed is written into each data buffer through the input frame, and the key parameters required for the accelerator operation are configured. After the calculation starts: ① First, the required data is retrieved from the cache. If it is a CNN, Transformer or other model, the multi-bit data will go directly to the next step. If it is an SNN, Spiking-Transformer or other model, the accelerator can calculate according to the time serial rule (only one time step, i.e. 1 bit of pulse data, is retrieved, and other time steps are not processed for the time being, which is determined by the time step value in the global register. This is more in line with the biological properties of the pulse neural network). It can also be calculated according to the time parallel rule (the 8-bit data is retrieved, and one time step, i.e. 1 bit of pulse is retrieved in turn and sent to the subsequent module until the required time step is taken. There is no need to wait until the current time step is calculated before taking the next time step, which reduces the repeated reading of weights).
[0047] ② In the preprocessing module, preprocessing is performed according to different model requirements, and the distributor decides to send it to the corresponding idle computing unit through the arbitration algorithm.
[0048] ③ After receiving the data, the computing unit combines the prediction algorithm to complete the multiplication and accumulation operation, and sends the result to the serial input and serial output register.
[0049] ④ Depending on the type of model processed by the accelerator, the activation module and the early exit module are started as appropriate, and the calculation results are read from the 16 serial-input and serial-output registers and processed accordingly.
[0050] ⑤ Output to the accelerator. If it is a spike-type neural network, the membrane level data of the time step will be updated. At the same time, based on the number of times the input data is read, the control logic calculates the address of the output of the calculation result of each set of data and connects the output port to the corresponding data frame, which is then handed off to other accelerators or modules for processing.
[0051] After processing the above data stream, the multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the embodiment of the present invention unifies the computing paradigms of different models, and can support the current mainstream CNN, SNN, Transformer, Spiking-Transformer and other models. It can support both spatial and temporal calculations, and also supports sparse output of data in different dimensions.
[0052] like Figure 4As shown, after each calculation is completed, the calculation result will be sent to the subsequent processing module, and the computing unit is in an idle state. In order to reduce the idle time, after each computing unit completes a round of tasks and outputs the multiplication and accumulation results, the control logic will determine whether all the data written to the accelerator cache SRAM has been calculated. If so, the data will no longer be read out. Otherwise, the control logic will read a set of data from the cache again, pre-process it, and then send it to the idle computing unit. The reading process is sequential, while the calculation process is disordered. It is not necessary for all computing units to be idle. As long as there are idle hardware resources, data can be read and calculated, thereby decoupling the calculation processes between different computing units and compressing the idle time of the computing unit to the cache read time that cannot be reduced.
[0053] Based on any of the above embodiments, the specific structure of the multi-model fusion neural network accelerator that supports out-of-order sparse output includes: (1) Data cache: including weight SRAM, input data SRAM, calculation parameter SRAM, and important parameter global registers. The number of synapses and models supported by each neuron are configurable. SRAM also supports different data storage types, realizing reuse with CNN data space and improving storage efficiency. (2) Preprocessing module: Detecting input data or SNN time steps that are 0 and removing them, and generating data that may be needed in the output sparse prediction process. (3) Computing unit: The most important computing module, which completes the multiplication and accumulation operation together with the prediction unit. (4) Prediction unit: Using the prediction algorithm to predict whether the result of the received data is positive, negative, or 0, supporting output prediction functions of up to 256 channels, and realizing sparse spatial dimensions. (5) Serial input and output register: Save the calculation results of all computing units to prevent loss, and pass them through subsequent processing in sequence. (6) Activation module: For models such as CNN and Transformer, use a lookup table to support nonlinear function operators. (7) Early Exit Module: For models such as SNN and Spiking-Transformer, the early exit algorithm can be used to determine whether to terminate the model calculation early to achieve output sparseness in the time dimension. (8) Output Module: Using the asynchronous handshake protocol, the data and its output coordinates are sent out of the accelerator in sequence. (9) Control Logic: Controls the data access, release, output address calculation and other processes, as well as generates other required control signals. The working process is divided into five stages: data acquisition, distribution, calculation, post-processing, and output. The data acquisition stage supports the reading of multi-bit data, serial time step pulses, and parallel time step pulses. The distributor decides to send it to the corresponding idle computing unit through the arbitration algorithm. After receiving the data, the computing unit combines the prediction algorithm to complete the multiplication and accumulation operation, and the result is sent to the serial input and output register. In the post-processing stage, the activation module and the early exit module are started according to the type of model processed by the accelerator this time, and the calculation results are read from the 16 serial input and output registers and processed accordingly. Then the output is output to the outside of the accelerator. If it is a pulse-type neural network, the membrane level data of the time step will also be updated. This data flow and architectural design unifies the computing paradigms of different models, can provide support for all types of currently mainstream models, and can achieve higher computing power and energy efficiency benefits.
[0054] Table 1 lists the test results comparison of similar neural network accelerator hardware solutions. It can be seen from Table 1 that the technology proposed in this invention achieves higher computing power and energy efficiency while maintaining a higher model accuracy.
[0055] Table 1 Comparison of different design schemes
[0056] The embodiments of this invention are applicable to both ASICs and FPGAs. ASIC design and implementation have been performed using tools such as Synopsys' DesignCompiler and PrimeTime, and have passed comprehensive testing. The implementation can reach a neuron scale of 35,000 and a synapse scale of 4.5 million. At a frequency of 500 MHz, the peak energy efficiency (90% output sparsity) is 61.19 TOPs / W. This ensures the feasibility and high performance of the algorithm and hardware design.
[0057] Relatively few existing neural network accelerators support multiple computing modes simultaneously, including CNN, SNN, Transformer, and Spiking-Transformer. The dynamic nature of workloads within accelerators results in varying computational speeds across computing units, which often limits hardware resource utilization. Using output sparsity technology makes data processing within computing units even more uncontrollable. If synchronous computing is still used, this results in longer waiting times for computing units, leading to wasted hardware resources. Existing accelerators generally use software or hardware methods to sort computing task sizes and then schedule them to achieve a more balanced computing process. While these approaches reduce potential waiting times, they are still insufficient for optimizing the utilization of accelerator hardware resources and may also introduce other hardware overhead.
[0058] The multi-model fusion neural network accelerator supporting out-of-order sparse output provided by the embodiment of the present invention models the data processing and calculation processes of different models and extracts common points to design an accelerator that can support the calculation operations of multiple models, unifies the data access methods and calculation paradigms of different models, uses the same computing circuit, and completes the calculation processes required by different models through different data scheduling commands. It can support the hybrid intelligence of multiple models such as CNN / SNN / Transformer / Spiking-Transformer, and can obtain higher computing power and energy efficiency benefits; the process of issuing data to the computing unit is studied, and a neural network acceleration data flow supporting out-of-order sparse output is proposed. The inference algorithm is used to predict the sparsity of the output data, and new workload data is issued to it when idle resources appear, successfully compressing the useless waiting time in the calculation process and maximizing the utilization of hardware computing resources as much as possible. In data flow design, the present invention also proposes a data organization method that can reduce the number of data readings when the convolution window slides. The input data is divided into 9 categories according to the relative position. Each time the input data allocated to the computing unit only needs to be updated by 1 / 3, and the weight data originally stored in the unit only needs to be simply shifted, which greatly reduces the high-consumption interaction process between the computing module and the storage module.
[0059] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A multi-model fusion neural network accelerator supporting out-of-order sparse output, characterized in that: include: A data cache module is used to store weight data, input data and calculation parameters. The data cache module determines to send the stored data to the corresponding idle computing unit through a dynamic allocator; a plurality of computing units, configured to obtain weight data, input data, and computing parameters from the data cache module, each computing unit being configured to perform a multiplication-accumulation operation based on the weight data, input data, and computing parameters; Multiple prediction modules, each prediction module is connected to a corresponding computing unit, and predicts whether to terminate the calculation of subsequent time steps based on the multiplication and accumulation operation results; under the action of the prediction module, the multiplication and accumulation operation results of each computing unit are output at different times; Serial-in-serial-out registers, used to receive and cache the multiplication and accumulation results of each computing unit; a post-processing module, configured to extract the multiplication-accumulation operation results of each computing unit from the serial-input-serial-output register according to the calculated neural network type, and process the multiplication-accumulation operation results of each computing unit; The output module is used to output the processing result of the post-processing module based on the asynchronous handshake protocol.
2. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 1, characterized in that: It also includes a pre-processing module for detecting the data in the data cache module and removing zero values or invalid time step data in the input data.
3. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 2, characterized in that: The dynamic allocator uses a round-robin algorithm to distribute the pre-processed data to the idle computing units to balance the computing load.
4. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 1, characterized in that: The data cache module is also used for: For convolutional neural network data width is 8 bits; For the pulse neural network data width of 1 bit, it supports neural network calculations of no more than 8 time steps. The data of these 8 time steps are stored in 8-bit data in sequence from the first time step to the last time step, and are multiplexed with the convolutional neural network data space.
5. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 1, characterized in that: The data cache module supports a multi-mode storage structure and switches between convolution sliding window block storage and matrix multiplication row storage through control signals.
6. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 5, characterized in that: The data cache module is also used for: For convolutional neural networks, the convolutional sliding window block storage is used to split the input feature map into blocks according to the convolutional sliding window. The data points from the upper left to the lower right in these blocks are stored together according to their corresponding positions and stored in different storage areas. One row in the storage area can store data from multiple channels at the same position. During calculation, the position after sliding is compared with the position before sliding to update the changed area. For network models that use matrix multiplication operations, matrix multiplication rows are stored and data is read row by row, so that the input data and the stored data correspond one to one.
7. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 1, characterized in that: The multiple computing units are further configured to: For convolutional neural networks or Transformer models, multi-bit data will directly enter the corresponding computing unit; For spiking neural network or spiking-transformer models, calculations are performed according to time serial rules or time parallel rules; The time serial rule includes taking out the pulse data of only one time step and not processing other time steps temporarily; the time parallel rule includes loading the pulse data of multiple time steps, taking out the pulse data of one time step in sequence according to the time step order and sending it to the calculation unit until the required time steps are taken.
8. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 1, characterized in that: The post-processing module includes an activation module; The activation module uses a lookup table to perform nonlinear mapping for convolutional neural networks and Transformer models to process data output by the computing unit.
9. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 1, characterized in that: The post-processing module includes an early exit module; The early exit module uses an early exit algorithm for the spiking neural network and the Spiking-Transformer model to determine whether to terminate the network model calculation early.
10. The multi-model fusion neural network accelerator supporting out-of-order sparse output according to claim 1, characterized in that: The output module is also used for: For pulse neural networks, the membrane level data of the time step is updated, and the address of the output of the calculation result of each set of data is calculated according to the number of times the input data is read.
Citation Information
Patent Citations
PYNQ-based YOLOv4-tiny neural network accelerator and acceleration method
CN116306851A
Random calculation CIM circuit and MAC operation circuit suitable for machine learning training
CN119356640A
Convolution Processing Method and Apparatus for Convolutional Neural Network, and Storage Medium
US20210350205A1