An Optimization Method for Recurrent Neural Network Algorithms Oriented to FPGA Platforms
By porting deep learning algorithms to the FPGA platform and performing storage optimization and algorithm optimization, the problem of high power consumption and low real-time performance when training deep learning models is solved, efficient computing and low power consumption are achieved, and it is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202210415163.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-15
AI Technical Summary
When training deep learning models, the prior art has problems such as excessive power consumption, poor real-time performance and insufficient mobile performance, and it is difficult to use in application scenarios that require low power consumption, strong real-time performance and pursue mobile performance.
The deep learning algorithm is ported to the FPGA platform, and through storage optimization and algorithm optimization methods, distributed RAM and BRAM are built using LUTs in FPGA for storage, expand loop operations and pipeline optimization.
It improves the computing efficiency and resource utilization of recurrent neural networks, reduces computing delay and power consumption, and is suitable for application scenarios with low power consumption, strong real-time performance and good mobile performance.
Smart Images

Figure CN114841332B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to an optimization method for recurrent neural network algorithms for FPGA platforms. Background Art
[0002] Machine learning and deep learning are applied to different scenarios, such as in text classification, object recognition, speech recognition, and autonomous driving. Even in some specific application scenarios, the performance of artificial intelligence is better than human judgment. For example, the alpha go robot trained by Google with reinforcement learning defeated the world champion in the Go game. Data-driven is an important feature of deep learning. A large amount of high-quality data is required to train a model with higher accuracy and stronger capabilities. However, this also poses a challenge to the computing power of computers. Currently, most deep learning models are trained using CPU and GPU solutions. Due to the structural characteristics of the CPU, the efficiency and computing speed are not as good as those of the dedicated graphics processor GPU when using the CPU to train deep learning models. However, when accelerating the training model based on the GPU, there are weaknesses such as excessive power consumption and insufficient mobility, and it is difficult to use in application scenarios that require low power consumption, strong real-time performance, and pursuit of mobility. Summary of the Invention
[0003] To solve the above problems, this application provides an optimization method and system for recurrent neural network algorithms for FPGA platforms, which can transplant deep learning algorithms to application scenarios with low power consumption, strong real-time performance, and pursuit of mobility.
[0004] The optimization method for recurrent neural network algorithms for FPGA platforms provided by this application mainly includes:
[0005] Storage optimization: For intermediate variables in neural network algorithms, use the distributed RAM constructed by LUT in the FPGA for storage. For parameters in neural network algorithms that require more storage resources than the threshold, use the random block memory BRAM in the FPGA for storage;
[0006] Algorithm optimization: For the vector multiplication and addition operations in the forward calculation process of neural network algorithms, expand the outer multi-loop operations and optimize the calculation process with a pipeline.
[0007] Preferably, in the storage optimization, if there is multi-dimensional array storage calculation, perform memory alignment operations on the FPGA.
[0008] Preferably, the FPGA platform uses a ZYNQ-7z100 series development board.
[0009] Preferably, the parameters of the PS part of the FPGA platform include an application processor based on a dual-core ARM Cortex-A9, and the ARM-v7 architecture is selected.
[0010] Preferably, the parameters of the PL part of the FPGA platform include Logic Cells: 444K; Look-Up Tables (LUTs): 277,400; flip-flops: 554,800; 18x25 MACCs multipliers: 2,020; Block RAM: 26.5 Mb.
[0011] This application improves the computing efficiency and resource utilization rate of the recurrent neural network, and reduces the computing delay and power consumption. Description of the Drawings
[0012] Figure 1 It is a flowchart of a preferred embodiment of the recurrent neural network algorithm optimization method for the FPGA platform in this application. Detailed Embodiments
[0013] To make the purpose, technical solutions, and advantages of the implementation of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the drawings in the embodiments of this application. In the drawings, the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The described embodiments are some, but not all, of the embodiments of this application. The embodiments described below by referring to the drawings are exemplary and are intended to explain this application and should not be construed as limiting this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the drawings.
[0014] This application provides a recurrent neural network algorithm optimization method for an FPGA platform, as Figure 1 shown, mainly including:
[0015] Storage optimization: For intermediate variables in the neural network algorithm, a distributed RAM constructed by LUTs in the FPGA is used for storage. For parameters in the neural network algorithm that require more storage resources than the threshold, a random block memory (BRAM) in the FPGA is used for storage;
[0016] Algorithm optimization: For the vector multiplication and addition operations in the forward calculation process of the neural network algorithm, expand the outer multi-loop operations and optimize the calculation process through pipelining.
[0017] An important part of the optimization design of neural network algorithms for FPGA platforms is the optimization of storage. FPGAs contain different types of memory to meet different design requirements. For the large number of intermediate variables in neural network models, although they do not occupy much memory, they are frequently read and written. Therefore, LUT RAM is used for storage. For parameters that require relatively large storage resources in the network, BRAM storage is optimized. At the same time, for multi-dimensional array storage calculations, memory alignment operations are adopted. During the forward calculation process of recurrent neural networks, there are a large number of vector multiply-accumulate operations. Since the outer layer of recurrent neural networks includes many loop operations, the loop calculation of data can be unfolded, and the calculation process can be optimized by pipelining.
[0018] In the storage optimization process of this application, LUTRAM is used for storage, which can ensure high read and write efficiency. BRAM is used for storage, and at the same time, memory is used for its operation. While ensuring high read efficiency, large memory can be separately stored in small fragmented storage areas, reducing BRAM resource consumption within a reasonable range. In the calculation optimization process, pipelining optimization and loop unrolling are carried out, greatly increasing the throughput of FPGA calculations and improving calculation efficiency.
[0019] In this application, the random block memory is the on-chip memory of the FPGA and is the core data cache of the FPGA, used to cache the hot data of the hash table and as the cache of the packet processing unit in the KVS communication link. Compared with off-chip caches, the on-chip cache BRAM has the characteristics of low latency and high bandwidth.
[0020] In some alternative embodiments, in the storage optimization, if there are multi-dimensional array storage calculations, memory alignment operations are performed on the FPGA.
[0021] In some alternative embodiments, the FPGA platform uses a ZYNQ-7z100 series development board.
[0022] In some alternative embodiments, the parameters of the PS part of the FPGA platform include an application processor based on an ARM dual-core CortexA9, and the ARM-v7 architecture is selected, with a frequency of up to 800 MHz. In addition, each CPU includes 32KB of level-1 instruction and data caches and 512KB of level-2 caches, and the two CPUs of the FPGA share the above storage space. On-chip boot ROM, and 256KB of on-chip RAM.
[0023] In some alternative embodiments, the parameters of the PL part of the FPGA platform include Logic Cells: 444K; Look-Up Tables (LUTs): 277,400; flip-flops: 554,800; 18x25 MACCs multipliers: 2,020; Block RAM: 26.5 Mb.
[0024] Compared with CPUs and GPUs, when using the FPGA of the present application and the accompanying optimization method, the comparison of resource occupancy ratios before and after optimization is shown in Table 1, and the comparison of computing latencies with other processors is shown in Table 2. The present application designs an optimization method for recurrent neural network algorithms for the FPGA platform. Through operations such as storage optimization, computing parallelization optimization, and loop unrolling, parallel matrix calculations are designed, enabling the network model to perform calculations in a pipelined manner and allowing internal parallel computing. This greatly improves the computing efficiency and resource utilization efficiency. By comparing the average computing latency and power consumption of CPUs and GPUs, it can be seen that there is approximately a 7-fold improvement in computing latency compared to CPUs and approximately a 2-fold improvement compared to GPUs. At the same time, in terms of power consumption, it is only 37.8% and 16% of that of CPUs and GPUs respectively.
[0025] Table 1 Comparison of resource occupancy ratios before and after optimization
[0026] On-chip resources BRAM_18K DSP48E FF LUT Before optimization 48% 14% 16% 47% After optimization 29% 12% 12% 30%
[0027] Table 2 Comparison of computing latencies
[0028]
[0029] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An optimization method for recurrent neural network algorithms oriented to the FPGA platform, characterized in that, Including: Storage optimization: For intermediate variables in neural network algorithms, a distributed RAM constructed by LUTs in the FPGA is used for storage. For parameters in neural network algorithms that exceed the threshold in terms of storage resource requirements, a random block memory BRAM in the FPGA is used for storage; Algorithm optimization: For the vector multiply-accumulate operations in the forward calculation process of neural network algorithms, the outer multi-loop operations are expanded, and pipeline optimization is performed on the calculation process; Among them, in the storage optimization, if there are multi-dimensional array storage calculations, memory alignment operations are performed on the FPGA; the FPGA platform uses a ZYNQ-7z100 series development board; the parameters of the PS part of the FPGA platform include an application processor based on an ARM dual-core CortexA9 and the selection of the ARM-v7 architecture; the parameters of the PL part of the FPGA platform include Logic Cells: 444K; Look-Up Tables (LUTs): 277,400; flip-flops: 554,800; 18x25 MACCs multipliers: 2020; Block RAM: 26.5Mb.
Citation Information
Patent Citations
Artificial neural network implementation in field-programmable gate arrays
US20200257986A1