A Convolutional Neural Network Acceleration Method with Low Off-chip Transmission Bandwidth Requirements
By performing reusability analysis and scheduling strategy optimization on the data flow of the convolutional neural network model, the data transmission bandwidth limitation problem caused by limited on-chip storage resources of FPGA is solved, and the inference speed and adaptability of the accelerator are improved.
Patent Information
- Application Number
- CN202210430624.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-22
AI Technical Summary
FPGAs have limited on-chip storage resources and cannot store all parameters of the convolutional neural network model at one time, resulting in the data transmission bandwidth limiting the accelerator's inference speed when reading and writing directly from off-chip memory devices.
By performing reusability analysis on data flows based on "slice" scheduling strategy, a separate scheduling strategy is designed for each convolutional layer, using the reusability of input data to reduce off-chip memory access overhead and reduce off-chip transmission bandwidth requirements.
It effectively reduces the pressure on transmission caused by high throughput, avoids limited bandwidth to become a bottleneck in the overall performance of the system, solves the problem of memory access congestion, and improves the inference speed and adaptability of the convolutional neural network model on the FPGA platform.
Smart Images

Figure CN114638347B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of transmission, and relates to a method for accelerating a convolutional neural network with low off-chip transmission bandwidth requirements. Background Art
[0002] As a branch field of deep learning, convolutional neural networks have been widely applied in various scenarios, such as image processing, speech processing, etc., due to their excellent accuracy. Along with the improvement of the accuracy of convolutional neural networks, today's top convolutional neural network models need to construct a quite deep convolutional layer structure to convert the input image data into a highly abstract representation form, as Figure 1 shown. The increase in the depth of the neural network poses new challenges to computing. Using a very deep neural network is accompanied by a large number of multiply-accumulate (MAC) operations, and the architecture of a general CPU is not suitable for this kind of computing. A simple way to solve this problem is to use a GPU suitable for parallel computing, and its existing operation framework can better meet the computing requirements of convolutional neural networks. With the advent of the Internet of Things era, the application requirements for edge computing have also increased. People have begun to notice that GPUs are not suitable for edge computing applications due to their low energy efficiency. To solve this problem, current research has started to develop hardware accelerators in FPGAs and ASICs, which have a high energy efficiency ratio. In addition, due to the flexibility of the reconfigurable logic structure of FPGAs, they have become a solution widely used in the field of hardware accelerators. However, when designing an FPGA-based convolutional neural network accelerator, the following problems are usually faced.
[0003] The on-chip storage resources of FPGAs are limited, usually only a few megabits, and it is impossible to store all the parameters of a convolutional neural network model, which are often hundreds of megabits. If directly read and written from off-chip storage, tens of millions of calculations make the fragmented data reading and writing very time-consuming. Figure 2 shows a schematic diagram of the execution of the convolution behavior in a single-layer convolutional neural network. The left side is the input feature map, where in_h represents the height of the input feature map, in_w represents the width of the input feature map, and N represents the number of channels of the input feature map. The middle part is the weights. K represents the height and width of the weights (which are usually the same). Every N weights (each weight corresponds to an input channel) form a group of filters, as shown in the part enclosed by the ellipse in the figure. Each layer of the convolutional neural network model has M groups of filters (each group of filters corresponds to an output channel). The right side is the output feature map, where R represents the height of the output feature map, C represents the width of the output feature map, and M represents the number of channels of the output feature map. The process of a single convolution execution is shown by Figure 3 shown, and it can be specifically divided into the following three steps:
[0004] 1. A filter multiplies all elements of the input feature map at a fixed stride (usually smaller than the kernel size), and a total of N*K*K products can be obtained.
[0005] 2. Within each weight, the results of multiplying the K*K input feature maps by the weight are accumulated.
[0006] 3. In the direction of the input channels, the results of the multiply-accumulate operations within the N weights are accumulated.
[0007] In this way, the convolution result at a single position of the output feature map is obtained. In fact, there is a one-to-one correspondence between the width and height of the input feature map, the width and height of the output feature map, the width and height of the weights, and the stride, which can be represented by the following equations. Here, S represents the stride of the convolution, i represents the width of the convolution kernel, and j represents the height of the convolution kernel.
[0008] in_h = R×S + j
[0009] in_w = C×S + i
[0010] During the inference process, a large number of parameter transfers are required to support each calculation process. The design principle of a convolutional neural network accelerator is to meet the characteristics of "high parallel computing". To meet this characteristic, it usually requires extremely high data throughput, which will cause the phenomenon of "frequent data interaction". The on-chip storage resources of an FPGA cannot store all the parameters of the convolutional neural network model at one time. A common approach is to place them in off-chip storage devices, such as DDR (Double DataRate SDRAM). If directly read and written from the DDR, the fragmented data reading and writing is very time-consuming and there are data access conflicts. Some scholars have proposed the "slicing" strategy, that is, dividing the input and output feature maps and weights of the convolutional neural network into several small parts and loading them into the on-chip storage in multiple times, so that the FPGA can process these data in turn. An example of its implementation is Figure 4As shown in the figure, the input channel N is divided into several small parts Tn, and the output channel M is divided into several small parts Tm. Each time, only a corresponding part of the input feature map and weight data is loaded onto the chip and stored in the on-chip cache. The computing unit can directly obtain the data through the on-chip interconnection network without directly accessing off-chip storage. After this part of the calculation is completed, the remaining data is loaded and the same steps are repeated. Although this approach avoids fragmented data reading and writing and reduces the overall time overhead of memory access, the bandwidth that can be provided by off-chip storage devices is limited under the constraints of fixed data frequency and bus width. During the period when the data loading process is initiated, if the peak bandwidth requirement for data transmission exceeds the maximum bandwidth that can be provided by off-chip storage devices, it will cause bus contention, making the data loading speed unable to fully match the processing speed of the computing unit. In the architecture design of the system pipeline, the computing unit can only idle and wait until the data loading is completed, and the inference speed of the entire accelerator will also be slowed down.
[0011] Therefore, in the application of a convolutional neural network accelerator system, in order to avoid the limitation of data transmission rate by off-chip bandwidth, starting from the perspective of utilizing data reusability, the present invention designs a convolutional neural network accelerator architecture with low off-chip transmission bandwidth requirements, further optimizing the pipeline design in the convolutional neural network accelerator, thereby avoiding the overall performance bottleneck caused by transmission bandwidth and improving the inference speed of the convolutional neural network accelerator. Summary of the Invention
[0012] In view of this, the purpose of the present invention is to provide a convolutional neural network acceleration method with low off-chip transmission bandwidth requirements. By analyzing the reusability of the data stream based on the "slicing" scheduling strategy and designing a separate scheduling strategy for each convolutional layer of the convolutional neural network model under the constraints of resources such as on-chip computing, storage, and logic in the FPGA, the pressure on transmission caused by high throughput is reduced, and the limited bandwidth is prevented from becoming the bottleneck of the overall system performance, thereby solving the problem of memory access congestion in practical applications and improving the adaptability of the convolutional neural network model deployed on the FPGA platform and expanding its application scenarios.
[0013] To achieve the above object, the present invention provides the following technical solutions:
[0014] A convolutional neural network acceleration method with low off-chip transmission bandwidth requirements, the method comprising:
[0015] Data loading (load), convolutional calculation (conv), and data write-back (store);
[0016] According to the calculation formula of the theoretical bandwidth of DDR3 transmission:
[0017] Bandwidth DDR= Core_frq × Bus_bitwidth × Mult_factor / 8 bits
[0018] Among them, Core_frq represents the memory core frequency, which is equal to the DDR3 data frequency divided by 8 bits, Bus_bitwidth represents the memory bus bit width, the maximum DDR3 data frequency supported by the 5CSEBA6 chip is 800Mhz, and the supported bus bit width is 32 bits; Mult_factor represents the memory multiplication factor;
[0019] DDR uses both the rising and falling edges of the clock pulse to transmit data once each. One clock signal transmits twice the data of SDRAM, which is called double data rate SDRAM; its multiplication factor is 2,
[0020] DDR2 uses the technology of transmitting data once each on both the rising and falling branches of the clock pulse, and prefetches 4 bits of data each time, which is twice that of DDR. Its multiplication factor is 4,
[0021] As an improvement of DDR2, DDR3 prefetches 8 bits of data each time, which is twice that of DDR2 and four times that of DDR. Its multiplication factor is 8. Through the following formula:
[0022] Bandwidth DDR = 800Mb / 8 bits × 32 bits × 8 bits / 8 bits
[0023] The maximum DDR side transmission bandwidth supported by the selected 5CSEBA6 chip is obtained as 3200Mhz; according to the bandwidth calculation formula of the convolutional neural network accelerator at the application side:
[0024] Bandwidth APP = Data_frq × Data_bitwidth / 8 bits
[0025] Among them, Data_frq represents the application side clock frequency, Data_bitwidth represents the application side data bit width. Substitute the current design data bus frequency of 150Mhz and bit width of 128 bits:
[0026] Bandwidth APP = 150Mhz × 128 bits / 8 bits
[0027] The required bandwidth for current data loading is obtained as 2400Mhz; the required bandwidth for data write-back is 2400Mhz;
[0028] To reduce the bandwidth requirement for data transmission, memory access is divided into two levels: off-chip memory access and on-chip memory access; a loop-nested pseudocode of the current data scheduling strategy is constructed to analyze which data streams are for off-chip DDR memory access and which are for on-chip cache memory access;
[0029] Compared with on-chip memory access, each off-chip memory access requires handshake signals stipulated by the protocol, resulting in greater time and power consumption overheads and occupying the limited off-chip transmission bandwidth; in the outer loop to of each corresponding output channel slice Tm, by utilizing the reusability of input data, the scheduling strategy of one input feature map data corresponding to one weight data is converted into one input feature map data corresponding to several weight data, reducing the off-chip memory access overhead caused by transmitting input feature map data and lowering the requirement for off-chip transmission bandwidth; after calculating multiple weights, output feature map data corresponding to multiple output channels will be obtained. Under the constraint of on-chip storage resources, an appropriate input feature map data reuse rate is selected to control the number of weight copies processed by each convolution calculation process, avoiding exceeding the on-chip cache space due to excessive output data;
[0030] After the data scheduling strategy is converted and implemented on the board, for the layer with a smaller output feature map size, the on-chip cache space it occupies is smaller. Each time, one input feature map data and two weight data are loaded from the off-chip DDR3 memory, and a certain degree of input feature map data reusability is utilized.
[0031] Optionally, the convolutional neural network acceleration method specifically includes:
[0032] Engine calculation, used to process the convolutional operations of the accelerator core;
[0033] On-chip cache, used to buffer the input feature map and weight data loaded from off-chip storage;
[0034] Process control, used to control the interaction of each module, initiate and stop the operation process; the operation process includes calculation, loading, and storage.
[0035] Optionally, in the engine calculation, there are several processing units PE, and each processing unit includes Tn multipliers. The multiplication operation of weights and input feature map data is executed in parallel in the input channel slice dimension;
[0036] The products of every two multipliers are added by an adder, and the sums of every two adders are added by an adder. Finally, the Tn multiplications are accumulated into one value, achieving the effect of stacking on the input channels;
[0037] All the input feature map data of the PEs are the same. By reusing the input data, there is no need to separately load data from the DDR for each PE's calculation, reducing the off-chip memory access overhead.
[0038] Optionally, in the on-chip cache, the storage module is divided into two parts, namely the input cache and the output cache. The output feature map data obtained by the engine calculation and processing is mapped to the on-chip RAM in the following manner:
[0039] Set the data bit width of the RAM to the output channel slice * data precision. Concatenate the data of all channels at the same width and height positions of the output feature map into an output result and store it in the RAM;
[0040] Set the address depth of the RAM to the width C × height R of the output feature map. Store the data at the i-th row and j-th column positions of the output feature map in the RAM according to the address mapping method of the following formula:
[0041] data_address = j * C + i
[0042] Allocate storage space for the input data. Store the input feature map and weight data into the input cache in sequence according to the width and height of the input and output channels; Each output feature map data is calculated from the corresponding weight and input feature map data. Through the on-chip data interconnection network, address and fetch values from the input cache in the order of output feature map calculation;
[0043] All on-chip caches adopt a double-buffer strategy to ensure that during the execution of the current calculation process, data required for the next calculation is loaded from off-chip to on-chip simultaneously, realizing pipelined processing
[0044] Optionally, the convolutional neural network acceleration method is specifically divided into the following steps:
[0045] (1) At the initial moment, the system is in the idle state ST_IDLE; after receiving the start command issued by the host, it enters the ST_CONFIG state; The accelerator first starts to configure the model parameters of this layer of the convolutional neural network, including the stride, filter size, and input and output channel slice sizes and feature map width and height sizes, which is called the parameter_config operation; After the configuration is completed, it enters the ST_FIRST_LOAD_0 state;
[0046] (2) The ST_FIRST_LOAD_0 state represents the first data loading operation of each convolutional layer. Load the input feature map and weight data of the first channel slice in the DDR into the on-chip input cache IRAM_0, which is called Load0. Load1 is the same; Only loading operations are performed in this state, and no convolution or write-back to the DDR operations are performed; After the data loading is completed, jump to the ST_FIRST_LOAD_1 state;
[0047] (3) The ST_FIRST_LOAD_1 state represents the first initiation of the Load1 operation and simultaneously initiates the convolution task. It calculates the sliced data that has been loaded into the input cache IRAM_0 and writes it into the output cache ORAM_0, which is called the Conv0 operation. The Conv1 operation is the same, and there is no write-back to the DDR operation. After completing the data loading and calculation, it enters the ST_LOAD_1 state.
[0048] (4) After entering the ST_LOAD_0 state, it initiates the Conv1, Load0 operations and writes the data in the output cache ORAM_0 back to the DDR, which is called the Store0 operation. The Store1 operation is the same, and it increments the data counter by the amount of data transferred in this Load operation. After completion, it jumps to the ST_LOAD_1 state.
[0049] (5) After entering the ST_LOAD_1 state, it initiates the Conv0, Load1, and Store1 operations, increments the data counter by the amount of data transferred in this Load operation, and determines whether all the data has been loaded. If so, it jumps to the ST_LAST_CONV state; otherwise, it jumps back to the ST_LOAD_0 state.
[0050] (6) After entering the ST_LAST_CONV state, it initiates the Conv1 and Store0 operations. Since all the input feature maps and weight data have been loaded, there is no need to initiate the Load operation anymore. After completion, it enters the ST_LAST_STORE state.
[0051] (7) After entering the ST_LAST_STORE state, it initiates the Store1 operation, which represents that all the operations and related operations in a convolutional layer have been completed. It determines whether this convolutional layer is the last layer in the convolutional neural network model. If so, it jumps back to the ST_IDLE state; otherwise, it jumps to the ST_CONFIG state to prepare for the calculation of the next layer.
[0052] The beneficial effects of the present invention are as follows: It has the advantages of low cost, high integration, low consumption of hardware resources, simple structure, high reliability, and easy implementation. It can effectively reduce the pressure on transmission caused by high throughput and avoid the limited bandwidth from becoming the bottleneck of the overall system performance, thus solving the problem of memory access congestion in practical applications and improving the adaptability of the convolutional neural network model deployed on the FPGA platform and broadening its application scenarios.
[0053] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings
[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail and preferably below with reference to the drawings, where:
[0055] Figure 1 is the convolutional neural network model structure;
[0056] Figure 2 is the schematic diagram of a single-layer convolutional behavior;
[0057] Figure 3 is the schematic diagram of a single convolutional behavior;
[0058] Figure 4 is the schematic diagram of the scheduling strategy for data "slicing";
[0059] Figure 5 is the overall architecture of the accelerator of the present invention;
[0060] Figure 6 is the parallel engine computing module;
[0061] Figure 7 is the on-chip cache structure;
[0062] Figure 8 is the system process control state diagram. Detailed Embodiments
[0063] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0064] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, rather than physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0065] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and cannot be understood as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0066] A convolutional neural network acceleration method with low off-chip transmission bandwidth requirements uses an Altera CycloneV SoC FPGA (5CSEBA6) as the deployment platform for the convolutional neural network accelerator. It includes the designed engine computing, interconnection network, input / output cache, and control unit. In addition, off-chip Micron DDR3 SDRAM is used to store the convolutional neural network model parameters.
[0067] For the existing "slice" data stream, according to its data loop scheduling method, an accelerator architecture is designed. Separate time counters are designed for each process including data loading (load), convolutional calculation (conv), and data write-back (store) to analyze the processing time overhead of each module and find the improvement direction for optimizing performance. In the traditional "slice" data stream accelerator, there is a situation where the data loading time (load_time) is greater than the convolutional calculation time (conv_time). Since the accelerator adopts a dual-buffer technology and a multi-stage pipeline architecture, after the input cache finishes loading data once, the process of the next data loading is immediately initiated, and at the same time, the convolutional calculation unit starts to process the data that was loaded last time. Therefore, the execution of the data loading and convolutional calculation processes overlaps in time. However, if the data loading time is greater than the convolutional calculation time, this will cause the convolutional calculation unit to wait for the completion of the previous data loading process before it can initiate the next convolutional calculation process after completing one process. This will cause a pause (bubble) in the pipeline, making the calculation unit idle during the time waiting for the data to be loaded completely, reducing the calculation efficiency and slowing down the overall inference speed. To verify the problem, the accelerator design is simulated. The processing process of the accelerator system includes a total of three processes, namely data loading (load), convolutional calculation (conv), and data write-back (store). Focusing on analyzing the load and conv processes, it is found that the overly long data loading time causes a delay before the next calculation can be initiated after each convolutional calculation process is completed. The pause in the pipeline of the system results in a slowdown of the overall inference speed.
[0068] To address the impact of long data loading time and locate the cause of this problem, it is necessary to analyze the theoretical maximum bandwidth supported by DDR3 and the bandwidth required by the accelerator application side. According to the calculation formula for the transmission theoretical bandwidth of DDR3:
[0069] Bandwidth DDR = Core_frq × Bus_bitwidth × Mult_factor / 8bits
[0070] Among them, Core_frq represents the memory core frequency, which is equal to the DDR3 data frequency divided by 8 bits. Bus_bitwidth represents the memory bus bit width. The maximum DDR3 data frequency supported by the 5CSEBA6 chip is 800Mhz, and the supported bus bit width is 32 bits. Mult_factor represents the memory multiplication factor. DDR transmits data once on both the rising and falling edges of the clock pulse. One clock signal can transmit twice as much data as SDRAM, so it is also called Double Data Rate SDRAM. Its multiplication factor is 2. DDR2 still uses the technology of transmitting data once on both the rising and falling branches of the clock pulse. The difference is that it pre-reads 4 bits of data each time, which is twice that of DDR. Therefore, its multiplication factor is 4. As an improvement of DDR2, DDR3 is characterized by pre-reading 8 bits of data each time, which is twice that of DDR2 and four times that of DDR. Therefore, its multiplication factor is 8. Through the following formula:
[0071] Bandwidth DDR = 800Mb / 8bits × 32bits × 8bits / 8bits
[0072] The maximum DDR side transmission bandwidth supported by the selected 5CSEBA6 chip is obtained as 3200Mhz. However, according to the bandwidth calculation formula of the application side (convolutional neural network accelerator):
[0073] Bandwidth APP = Data_frq × Data_bitwidth / 8bits
[0074] Among them, Data_fiq represents the application side clock frequency, and Data_bitwidth represents the application side data bit width. Substitute the current design data bus frequency of 150Mhz and bit width of 128bits:
[0075] Bandwidth APP = 150Mhz × 128bits / 8bits
[0076] The currently required bandwidth for data loading is 2400Mhz. Similarly, the required bandwidth for data write-back can be obtained as 2400Mhz. Since the accelerator system adopts a pipeline architecture, the data write-back operation of the previously completed calculation overlaps with the data loading operation of the next calculation waiting to be processed. Coupled with the memory overhead brought by the simultaneously running Linux operating system on the SoC processor, its peak bandwidth consumption has exceeded the maximum bandwidth on the DDR side. This has led to bus contention. Therefore, neither increasing the data reading frequency nor the transmission bandwidth can effectively increase the data transmission rate.
[0077] To reduce the data transmission bandwidth requirements, analyze from the data scheduling strategy. According to the traditional "slicing" scheduling strategy, I divide memory access into two levels: off-chip memory access and on-chip memory access. Construct the loop nest pseudocode of the current data scheduling strategy and analyze which data streams are for off-chip DDR memory access and which are for on-chip cache memory access.
[0078] Compared with on-chip memory access, each off-chip memory access requires handshake signals stipulated by the protocol, resulting in greater time and power consumption overhead and occupying the limited off-chip transmission bandwidth. By analyzing the data reusability, it is found that the traditional "slicing" data stream repeatedly loads the same input feature map data, causing unnecessary memory access overhead.
[0079] Therefore, by leveraging the reusability of the input data, the scheduling strategy of one input feature map data corresponding to one weight data can be converted into one input feature map data corresponding to several weight data, thereby reducing the off-chip memory access overhead caused by transmitting the input feature map data. Thus, the requirement for off-chip transmission bandwidth is reduced. After calculating multiple weights, output feature map data corresponding to multiple output channels will be obtained. Therefore, a larger output cache is needed to store this part of the results. So, under the condition of meeting the on-chip storage resource constraints, select an appropriate input feature map data reuse rate to control the number of weight copies processed by each convolution calculation process, and avoid exceeding the on-chip cache space due to excessive output data.
[0080] For the layers with smaller output feature map sizes, the on-chip cache space they occupy is smaller. Therefore, each time, one input feature map data and two weight data are loaded from the off-chip DDR3 memory, taking advantage of a certain degree of input feature map data reusability. With the convolution calculation time conv_time remaining unchanged, the data loading time load_time is reduced by half, overcoming the performance bottleneck caused by the off-chip transmission bandwidth, eliminating the waiting time of the convolution calculation unit, and improving the overall inference time by 11 milliseconds, accelerating the overall inference time of the convolutional neural network accelerator. According to the simulation analysis, in the process pipeline of the proposed accelerator architecture, considering the reusability of the input feature map data, the data scheduling strategy is changed, reducing the off-chip memory access overhead generated by transmitting the input feature map data, reducing the execution time of the data loading process, thus eliminating stalls in the pipeline. Two consecutive convolution calculation processes can be initiated, thereby avoiding the idling of the calculation unit and improving the utilization rate of computing resources.
[0081] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. As Figure 5 shown, a convolutional neural network acceleration method with low off-chip transmission bandwidth requirements according to the present invention mainly consists of three parts, namely:
[0082] 1. An engine computing unit for processing the convolution operation of the accelerator core;
[0083] 2. An on-chip cache unit for buffering the input feature map and weight data loaded from off-chip storage;
[0084] 3. A process control unit for controlling the interaction of each module and initiating and stopping the operation processes (such as calculation, loading, storage).
[0085] The architecture of the engine calculation is as Figure 6 shown. Each system within the dashed box is called a processing element (PE). It contains Tn multipliers, so the multiplication operation of the weight and input feature map data can be performed in parallel in the input channel slice dimension. The products of every two multipliers are added by an adder. Similarly, the sums of every two adders are also added by an adder. Finally, the Tn multiplications are accumulated into one value to achieve the effect of superimposing on the input channels. It should be noted that all PEs have the same input feature map data. By reusing the input data, there is no longer a need to separately load data from the DDR for each PE's calculation, reducing the off-chip memory access overhead, avoiding the slowdown of the data caching speed due to the limited transmission bandwidth, causing the waiting of the calculation process, and thus restricting the overall inference speed of the accelerator.
[0086] The schematic diagram of the on-chip cache architecture is asFigure 7 As shown in the figure, the storage module is mainly divided into two parts, namely the input cache and the output cache. The output feature map data obtained by the engine calculation and processing is mapped to the on-chip RAM in the following way:
[0087] 1. Set the data width of the RAM to the output channel slice * data precision. Concatenate the data of all channels at the same width and height positions of the output feature map into an output result and store it in the RAM.
[0088] 2. Set the address depth of the RAM to the width * height of the output feature map (C * R). Then, the data at the i-th row and j-th column positions of the output feature map can be stored in the RAM according to the address mapping method of the following formula.
[0089] data_address = j * C + i
[0090] The storage space for the input data can be allocated in a similar way. The input feature map and weight data are stored in the input cache in sequence according to dimensions such as input and output channels, width, and height. From a single convolution operation, it can be seen that each output feature map data is calculated from the corresponding weight and input feature map data and can be calculated through a formula. In this way, through the on-chip data interconnection network, the values can be addressed and fetched from the input cache in the order of output feature map calculation.
[0091] It should be noted that all on-chip caches adopt a double-buffer strategy to ensure that during the execution of the current calculation process, the data required for the next calculation can be loaded from off-chip to on-chip simultaneously, thereby realizing pipelined processing and improving the calculation efficiency of the system.
[0092] The flow control of the accelerator is as Figure 8 shown, and it can be specifically divided into the following steps:
[0093] 1. At the initial moment, the system is in the idle state ST_IDLE. After receiving the start command issued by the host, it enters the ST_CONFIG state. The accelerator first starts to configure the model parameters of this layer of the convolutional neural network, including the stride, filter size, input and output channel slice sizes, and feature map width and height sizes, etc. (referred to as the parameter_config operation). After the configuration is completed, it enters the ST_FIRST_LOAD_0 state.
[0094] 2. The ST_FIRST_LOAD_0 state represents the first data loading operation for each convolutional layer. At this time, the input feature map of the first channel slice in the DDR and the weight data are loaded into the on-chip input cache IRAM_0 (referred to as the Load0 operation, and Load1 is the same). Only the loading operation is performed in this state, without performing convolution or writing back to the DDR operation. After the data loading is completed, it jumps to the ST_FIRST_LOAD_1 state.
[0095] 3. The ST_FIRST_LOAD_1 state represents the first initiation of the Load1 operation and simultaneously initiates the convolution task. The sliced data that has been loaded into the input cache IRAM_0 is calculated and written into the output cache ORAM_0 (referred to as the Conv0 operation, and Conv1 is the same). No writing back to the DDR operation is performed in this state. After the data loading and calculation are completed, it enters the ST_LOAD_1 state.
[0096] 4. After entering the ST_LOAD_0 state, initiate the Conv1, Load0, and write-back to the DDR of the data in the output cache ORAM_0 (referred to as the Store0 operation, and Store1 is the same), and increase the data counter by the amount of data transferred in this Load operation. After completion, it jumps to the ST_LOAD_1 state.
[0097] 5. After entering the ST_LOAD_1 state, initiate the Conv0, Load1, and Store1 operations. After all these operations are completed, increase the data counter by the amount of data transferred in this Load operation, and determine whether all the data has been loaded. If so, it jumps to the ST_LAST_CONV state; otherwise, it jumps back to the ST_LOAD_0 state.
[0098] 6. After entering the ST_LAST_CONV state, initiate the Conv1 and Store0 operations. In this state, all the input feature maps and weight data have been loaded, and there is no need to initiate the Load operation anymore. After completion, it enters the ST_LAST_STORE state.
[0099] 7. After entering the ST_LAST_STORE state, initiate the Store1 operation. After this operation is completed, it represents that all the operations and related operations in a convolutional layer have been completed. Determine whether this convolutional layer is the last layer in the convolutional neural network model. If so, it jumps back to the ST_IDLE state; otherwise, it jumps to the ST_CONFIG state to prepare for the calculation of the next layer.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A convolutional neural network acceleration method with low off-chip transmission bandwidth requirements, characterized in that: This method includes: Data loading (load), convolutional calculation (conv), and data write-back (store); According to the calculation formula of the theoretical bandwidth of DDR3 transmission: Bandwidth DDR = Core_frq × Bus_bitwidth × Mult_factor / 8 bits Among them, Core_frq represents the memory core frequency, which is equal to the DDR3 data frequency divided by 8 bits, Bus_bitwidth represents the memory bus bit width, the maximum DDR3 data frequency supported by the 5CSEBA6 chip is 800Mhz, and the supported bus bit width is 32 bits; Mult_factor represents the memory multiplication factor; DDR uses the rising and falling edges of the clock pulse to transmit data once each, and 1 clock signal transmits twice the data of SDRAM, which is called double data rate SDRAM; its multiplication factor is 2, DDR2 uses the technology of transmitting data once each on the rising and falling branches of the clock pulse, and pre-reads 4 bits of data each time, which is twice that of DDR. Its multiplication factor is 4, As an improvement of DDR2, DDR3 pre-reads 8 bits of data each time, which is twice that of DDR2 and 4 times that of DDR. Its multiplication factor is 8. Through the following formula: Bandwidth DDR = 800 Mb / 8 bits × 32 bits × 8 bits / 8 bits The maximum DDR-side transmission bandwidth supported by the selected 5CSEBA6 chip is obtained as 3200Mhz; according to the bandwidth calculation formula of the convolutional neural network accelerator at the application end: Bandwidth APP = Bata_frq × Data_bitwidth / 8 bits Among them, Data_frq represents the application-side clock frequency, Data_bitwidth represents the application-side data bit width, and substitute the current design data bus frequency of 150Mhz and bit width of 128 bits: Bandwidth APP = 150Mhz × 128bits / 8bits The required bandwidth for current data loading is obtained as 2400Mhz; the required bandwidth for data write-back is 2400Mhz; To reduce the data transmission bandwidth requirements, the memory access is divided into two levels: off-chip memory access and on-chip memory access; construct the loop-nested pseudocode of the current data scheduling strategy, and analyze which data streams are off-chip DDR memory access and which data streams are on-chip cache memory access; Compared with on-chip memory access, off-chip memory access requires handshake signals specified by the protocol each time a transmission is initiated, with greater time and power consumption overhead and occupying limited off-chip transmission bandwidth; in the outer loop to of each corresponding output channel slice Tm, using the reusability of the input data, convert the scheduling strategy of one input feature map data corresponding to one weight data into one input feature map data corresponding to several weight data, reducing the off-chip memory access overhead caused by transmitting the input feature map data and reducing the demand for off-chip transmission bandwidth; after calculating multiple weights, output feature map data corresponding to multiple output channels will be obtained. Under the condition of meeting the on-chip storage resource constraints, select an appropriate input feature map data reuse rate to control the number of weight copies processed by each convolutional calculation process, avoiding exceeding the on-chip cache space due to excessive output data; After the data scheduling strategy is converted and put on the board, each time one input feature map data and two weight data are loaded from the off-chip DDR3 memory, using a certain degree of input feature map data reusability.
2. A convolutional neural network acceleration method with low off-chip transmission bandwidth requirements according to claim 1, characterized in that: The convolutional neural network acceleration method specifically includes: Engine computing, which is used to process the convolutional operations of the accelerator core; On-chip cache, which is used to buffer the input feature maps and weight numbers loaded from off-chip storage; Process control, where the control unit is used to control the interaction of each module, the initiation and stop of the operation process; the operation process includes computing, loading, and storing.
3. A convolutional neural network acceleration method with low off-chip transmission bandwidth requirements according to claim 2, characterized in that: In the engine computing, it includes a number of processing units PE, each processing unit includes Tn multipliers, and the multiplication operation of the weights and the input feature map data is executed in parallel on the input channel slice dimension; The product of every two multipliers is added by an adder, and the sum of every two adders is added by an adder. Finally, the Tn multiplications are accumulated into one value, achieving the effect of superimposing on the input channels; All the input feature map data of the PEs are the same. By reusing the input data, there is no need to separately load data from the DDR for the calculation of each PE, reducing the overhead of off-chip memory access.
4. A convolutional neural network acceleration method with low off-chip transmission bandwidth requirements according to claim 2, characterized in that: In the on-chip cache, the storage module is divided into two parts, namely the input cache and the output cache. The output feature map data obtained by the engine computing is mapped to the on-chip RAM in the following way: Set the data bit width of the RAM to the output channel slice * data precision, and splice the data of all channels at the same width and height position of the output feature map into an output result and store it in the RAM; Set the address depth of the RAM to the width C × height R of the output feature map, and store the data at the i-th row and j-th column position of the output feature map in the RAM according to the address mapping method of the following formula: data_address = j * C + i Allocate storage space for the input data, and store the input feature map and weight data into the input cache in sequence according to the width and height of the input and output channels; each output feature map data is calculated from the corresponding weight and input feature map data, and through the on-chip data interconnection network, address and retrieve values from the input cache in the order of output feature map calculation; All on-chip caches adopt a double-buffering strategy to ensure that during the execution of the current calculation process, the data required for the next calculation is simultaneously loaded from off-chip to on-chip, realizing pipelined processing.
5. A convolutional neural network acceleration method with low off-chip transmission bandwidth requirements according to claim 2, characterized in that: The convolutional neural network acceleration method is specifically divided into the following steps: (1) At the initial moment, the system is in the idle state ST_IDLE; after receiving the start command issued by the host, it enters the ST_CONFIG state; the accelerator first starts to configure the model parameters of this layer of the convolutional neural network, including the stride, filter size, and the slice sizes of the input and output channels and the width and height sizes of the feature map, which is called the parameter_config operation; after the configuration is completed, it enters the ST_FIRST_LOAD_0 state; (2) The ST_FIRST_LOAD_0 state represents the first data loading operation for each convolutional layer. The input feature map of the first channel slice and the weight data in the DDR are loaded into the on-chip input cache IRAM_0, which is called Load0, and Load1 is the same; Only the loading operation is performed in this state, and no convolution or write-back to the DDR operation is performed; when the data loading is completed, it jumps to the ST_FIRST_LOAD_1 state; (3) The ST_FIRST_LOAD_1 state represents the first initiation of the Load1 operation, and at the same time initiates the convolution task. The sliced data that has been loaded into the input cache IRAM_0 is calculated and written into the output cache ORAM_0, which is called the Conv0 operation, and Conv1 is the same, and no write-back to the DDR operation is performed; after the data loading and calculation are completed, it enters the ST_LOAD_1 state; (4) After entering the ST_LOAD_0 state, initiate Conv1, Load0, and write the data in the output cache ORAM_0 back to the DDR, which is called the Store0 operation, and Store1 is the same, and increase the data counter by the amount of data transferred in this Load operation; after completion, jump to the ST_LOAD_1 state; (5) After entering the ST_LOAD_1 state, initiate Conv0, Load1, and Store1 operations; increase the data counter by the amount of data transferred in this Load operation, and determine whether all the data has been loaded. If so, jump to the ST_LAST_CONV state, otherwise jump back to the ST_LOAD_0 state; (6) After entering the ST_LAST_CONV state, initiate Conv1 and Store0 operations. All the input feature maps and weight data have been loaded, and there is no need to initiate the Load operation anymore; after completion, enter the ST_LAST_STORE state; (7) After entering the ST_LAST_STORE state, initiate the Store1 operation, which means that all the operations and related operations in a convolutional layer have been completed; determine whether this convolutional layer is the last layer in the convolutional neural network model. If so, jump back to the ST_IDLE state, otherwise jump to the ST_CONFIG state to prepare for the calculation of the next layer.
Citation Information
Patent Citations
Lightweight convolutional neural network reconfigurable deployment method based on FPGA
CN111931909A
Hardware accelerator of full-frequency-domain convolutional neural network, acceleration method and image classification method
CN112712174A