A software-hardware collaborative acceleration method for convolutional neural networks based on FPGA
By implementing software and hardware collaborative acceleration methods on the FPGA platform, including training and quantization of convolutional neural networks, generating hardware circuits and optimizing network layer IP, the problem of low hardware resource utilization efficiency when deploying and accelerating convolutional neural networks at the edge is solved, efficient computing and model deployment effects are achieved, and universality is maintained on different FPGA platforms.
Patent Information
- Application Number
- CN202311624785.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-11-30
AI Technical Summary
When deploying and accelerating convolutional neural networks at the edge end, it is difficult for the existing technology to effectively utilize limited hardware resources, and acceleration solutions are usually not universal on different FPGA platforms, resulting in waste of resources and poor performance.
The software and hardware collaborative acceleration method based on FPGA is adopted, and the convolutional neural network and data set are selected, the hardware circuit is generated and the network layer IP is optimized. The pipeline structure is used to increase parallelism, the execution efficiency is improved through matrix division and loop expansion, and the memory flow is used to optimize data access.
It realizes a software and hardware collaborative optimization strategy suitable for a variety of convolutional neural networks, reduces hardware resource consumption, improves computing efficiency and model deployment effect, and realizes universality on different FPGA platforms, meeting the real-time processing needs in edge scenarios.
Smart Images

Figure CN117610626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to a software and hardware collaborative acceleration method for implementing a convolutional neural network based on FPGA. Background Art
[0002] With the continuous development of machine learning, convolutional neural networks (CNNs) have been widely used in computer vision, natural language processing, and data processing, producing huge social and economic benefits. However, while convolutional neural networks have excellent performance, they also consume a lot of computing and storage resources for training and reasoning. Take the convolutional neural network LeNet-5 proposed by Yann Lecun of New York University in 1998 as an example. This simple 5-layer neural network (excluding the pooling layer) has 8087 neurons and 61706 parameters; while the deeper AlexNet has 8 layers (the pooling layer is included in the convolution layer), the number of neurons is 650,000, and the number of parameters has reached an astonishing 60 million. Therefore, the demand for resources for deeper and more efficient convolutional neural networks is exponentially increasing, and various methods are needed to accelerate the neural network in practical applications.
[0003] At present, the acceleration of neural networks is mainly carried out from two aspects: software and hardware. In terms of software, based on deep learning frameworks such as Pytorch and Tensorflow, the algorithm is improved on the CPU and GPU platforms, without sacrificing the accuracy of the neural network, and the calculation speed is accelerated by optimizing the network structure. For example, classic neural networks such as GoogleNet and ResNet proposed by optimizing the algorithm based on simple CNN run faster than general neural networks with the same number of layers.
[0004] In terms of hardware, the basic hardware architecture for running neural networks is optimized to improve computing efficiency. For example, the CUDA parallel computing architecture launched by NVIDA allows developers to efficiently accelerate neural network computing tasks on the GPU. ASIC (application-specific integrated circuit) can customize dedicated hardware circuits to adapt to specific neural network structures, thereby achieving efficient computing. FPGA (field programmable gate array) uses Verilong language to write programs. Developers can connect the logic modules inside the FPGA according to their needs, thereby realizing a specific hardware structure to accelerate the neural network.
[0005] Calling local GPU resources or remote server resources for acceleration on a computer platform can generally meet the needs, but it is difficult to do so in edge scenarios such as industrial production and real-time processing. In edge scenarios, the computing resources of the device are extremely limited, and at the same time, there are high requirements for latency, requiring real-time processing of tasks, and not waiting for remote processing to complete before returning data. On the other hand, many studies on algorithm acceleration are for specific platforms or architectures, and cannot be directly applied on edge devices, or cannot achieve the expected performance. Therefore, for the problem of deploying and accelerating neural networks at the edge, a complete set of feasible solutions should be designed to adapt the designed hardware circuits to the adopted algorithms and achieve the highest efficiency. Instead of just running the advanced network structure directly on the device, there will be an error due to insufficient hardware resources; or blindly using matrix blocking, loop expansion and other technologies without considering the conditions of each layer of the network, resulting in resource waste or conflicts between computing and memory access.
[0006] GPUs have high performance but also very high cost and power consumption, making them unsuitable for accelerating edge devices in edge scenarios. ASICs lack flexibility, and their hardware structure cannot be reconstructed, which means they cannot be freely iterated or updated. In addition, their tape-out costs are high, making them only suitable for specific situations. Summary of the invention
[0007] The purpose of the present invention is to address the deficiencies in the prior art and propose a software-hardware collaborative acceleration method for implementing a convolutional neural network based on FPGA.
[0008] The object of the present invention is to achieve the following technical solution: a method for realizing software and hardware collaborative acceleration of convolutional neural network based on FPGA, the method comprising:
[0009] S1. Select a convolutional neural network and a data set, and train and quantize the convolutional neural network;
[0010] S2. Model deployment: Save the quantized model as a binary machine language file according to the weights and parameters;
[0011] S3, software and hardware collaboration: Generate hardware circuits based on saved data, and generate network layer IP based on functions; use pipeline structure to increase parallelism; specify the storage method of matrix through matrix partitioning; improve execution efficiency by changing the proportion of conditional judgment instructions in the loop body; pipeline the operation process of FPGA reading and writing data from the cache;
[0012] S4. Hardware implementation: Regenerate hardware IP for data that achieves software-hardware co-optimization, and complete hardware settings, constraints, and routing based on the hardware IP.
[0013] Furthermore, in the selection of the neural network and the data set, the size of the data set is selected to be able to take advantage of parallelism according to the hardware resource performance of the target FPGA.
[0014] Furthermore, after the convolutional neural network is trained on the personal device, the neural network is quantized using a quantization compression tool under the Pytorch framework.
[0015] Furthermore, during the quantization process, if the prediction accuracy of the model after quantization drops by more than 5%, the quantization precision is adjusted or the quantization strategy is changed. If the accuracy loss after quantization is less than 1%, a lower quantization precision is selected to improve the compression rate.
[0016] Furthermore, the method of saving the quantization model as a binary machine language file according to the weights and parameters is specifically as follows: saving the weights and bias parameters as a binary machine language file .bin according to the network layer to which they belong, and then transferring them to the FPGA cache through the AXI bus for reading according to needs.
[0017] Furthermore, the use of a pipeline structure to increase parallelism includes: using a first-level pipeline structure to increase parallelism and using a second-level pipeline to increase parallelism, wherein the second-level pipeline is two internal nested loops, and when used, the order between the loop nests needs to be adjusted to resolve conflicts in access and processing, ensuring that no data dependency is generated.
[0018] Furthermore, the storage method of the specified matrix by matrix partitioning includes: storing all elements in the matrix into different BRAMs, using multiple BRAMs to store the matrix, or storing several adjacent elements into different BRAMs.
[0019] Furthermore, the process of pipelining the operation of reading and writing data from the cache of the FPGA is specifically as follows: the data is moved from the off-chip storage to the faster on-chip storage, and after the operation, the result is transmitted back to the off-chip storage. When multiple tasks are performed in parallel, the logic of data access needs to be planned to avoid memory access bottlenecks.
[0020] Furthermore, in the hardware implementation, an interactive interface for observing the operation of the hardware needs to be designed to observe the operating results of the FPGA and change input parameters in real time to meet the reasoning requirements in actual scenarios.
[0021] Beneficial effects of the present invention:
[0022] 1) A software-hardware collaborative optimization strategy applicable to a variety of convolutional neural networks is proposed. Faced with the problem that hardware circuits cannot adapt well to different neural network structures or algorithms, the software-hardware collaborative optimization strategy adopted by the present invention realizes operator-level optimization, that is, it is not designed for each layer, but a general module is designed for different operators. For example, the core convolution operator, that is, the part that implements the convolution kernel sliding window operation in the convolutional neural network, can be instantiated as a circuit module, and this module can be called each time a convolution operation is performed. In this way, no matter how many convolution layers there are in the network structure, the designed convolution module can be used efficiently; and even if different layers of convolution require different calculation methods, they can be implemented by calling different operators.
[0023] 2) Adopt a complete set of neural network quantization and model compression solutions to ensure the actual deployment effect of the model. Neural network quantization compresses the model size by quantizing the network weights from 32-bit floating point to low-bit fixed-point numbers (such as 8-bit integers), thereby reducing the hardware resource consumption of running the network and reducing the bandwidth delay of accessing data from the memory during calculation. However, general neural network quantization research only verifies the effect on the host's CPU or GPU, and does not test its deployment effect on the FPGA hardware platform. This patent adopts a complete set of solutions from quantization strategy, quantization verification, quantization result export to quantization result deployment.
[0024] 3) It has the versatility of multiple FPGA platforms, and the project has been implemented on both Zynq 3EG and Ultra96 model FPGAs. Different models of FPGAs are not completely universal due to differences in the total amount of hardware resources, cache and interface configuration. Many related studies are based on a specific model of FPGA and have failed to verify the performance on other models of FPGAs. The projects implemented by this patent on both FPGAs can run smoothly and achieve the expected results, and are not limited to a specific model. FPGAs are highly versatile, combining the low-cost flexibility of programmable hardware circuits with the advantages of reconfigurable and low-power consumption of semi-custom circuits, and can efficiently accelerate convolutional neural networks in edge scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is an overview flow chart of a method for co-accelerating convolutional neural networks using FPGA according to the present invention;
[0026] Figure 2 It is a basic network model structure diagram adopted by the present invention;
[0027] Figure 3 It is a schematic diagram of multi-channel two-bit convolution calculation used to establish the mathematical model of the convolution operator;
[0028] Figure 4 Schematic diagram of pipeline parallel optimization;
[0029] Figure 5 It is a schematic diagram of realizing a two-stage pipeline to increase parallelism through data flow optimization;
[0030] Figure 6 It is a schematic diagram of dividing the matrix into different BRAMs;
[0031] Figure 7 This is a comparison diagram of the principle of loop unrolling;
[0032] Figure 8 This is a code sample for processing Ping-Pong Buffer data streams in memory access pipeline.
[0033] Fig. 9 This is the result diagram of the model quantitative verification. DETAILED DESCRIPTION
[0034] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0035] like Figure 1 As shown, the present invention provides a method for hardware-software coordinated acceleration of a convolutional neural network using FPGA, and the specific steps are as follows:
[0036] S1. Select basic neural network model and data set
[0037] When the task type is limited in a specific industrial application scenario, the data set will not be particularly large, and the hardware conditions of the edge device also limit the size of the selected data set, so the selected data set does not have to be too large. For example, in LeNet5, the first layer of convolution input is 28*28, and the number of channels is 6. If this layer is to be stored in parallel on the FPGA (stored as an integer eight-bit INT8), then at least 4.7KB of storage space is required; however, compared with the VGG16 network, the first layer of convolution input is 224*224, the number of channels is 64, and it is also stored in INT8 type, requiring at least 3.2MB of storage space. The current Zynq series FPGA has a maximum on-chip cache BRAM of 2180KB, which cannot take advantage of parallelism when running VGG16.
[0038] like Figure 2 As shown, the focus of this embodiment is to demonstrate the acceleration method and the whole deployment plan of software and hardware collaboration, so the general and representative public data set Cifar10 is selected. Since the neural network needs to be implemented on hardware, it is impossible to use various techniques such as residual and batch normalization to optimize the network structure, so a common 9-layer convolutional neural network is selected to merge the RELU activation layer into the convolution layer or the fully connected layer (the activation layer is not calculated separately).
[0039] This method can be applied to most convolutional neural networks. The selected embodiment is not special and is a typical "4-layer convolution + 3-layer full connection" structure without adding residual, batch normalization and other parts. The basic neural network structure is composed of convolution layer, activation layer, pooling layer and full connection layer, as shown in the following table.
[0040]
[0041]
[0042] Although this network has 9 layers, its actual structure is very simple. Convolution calculation accounts for a large proportion of the entire network, and it has four convolution layers of different sizes, which is of great significance to the design and application of subsequent convolution operators.
[0043] S2. Neural network training and quantization scheme
[0044] Because the embodiment of the present invention is aimed at inference acceleration in edge application scenarios, it is necessary to train the neural network first, and then export the parameter files such as weights and biases for subsequent use in hardware inference. First, train the network with a GPU on a personal computer, and the training framework is Pytorch. Since the BaseCNN used is not a specially designed high-accuracy network, the accuracy of the trained model will not be very high. First, train the network with a GPU on a personal computer, the training framework is Pytorch, and the learning rate is set to 0.01. Generally, 50 epochs (training cycles) can reach convergence. After training, a model with an accuracy rate that meets the requirements is obtained, and its structure and parameters are saved as a pt format file.
[0045] As mentioned earlier, the size of some layers may exceed the cache limit during the operation, affecting parallelism. The largest convolution layer in BaseCNN is Conv1. If stored in FT32 (floating point 32-bit), the required space is 28×28×32×32bit, that is, 100KB; if stored in INT8 type, the required space is 25KB, which reduces the hardware resource requirements by 75%. This is the intuitive effect of neural network quantization, storing neural network parameters as low-bit data, which facilitates its deployment on hardware or improves efficiency under the same circumstances. Generally speaking, quantizing neural networks will lose a certain amount of accuracy, but choosing appropriate quantization methods and strategies can reduce the accuracy loss, making the quantization benefits far greater than the losses.
[0046] For subsequent deployment, the quantization compression tool under the Pytorch framework is used to quantize the neural network. The specific quantization method is static quantization after training, the quantization algorithm is linear quantization, the quantization accuracy is integer INT8, and the quantization strategy selects the qnnpack strategy of channel quantization in the Pytorch library.
[0047] In fact, the above quantization scheme is not fixed and needs to be adjusted according to the verification results of the quantization model. If the prediction accuracy of the model after quantization drops too much (more than 5%), it means that the quantization loss is too large, and you should consider adjusting the quantization precision or changing the quantization strategy, that is, selecting FP16 higher precision quantization or selecting a layer-by-layer quantization strategy. In addition, if the precision loss after quantization is not large (less than 1%), you can choose a lower quantization precision, such as INT4, for a higher compression rate.
[0048] During the verification process, the parameter type should be considered. The quantized parameters cannot be used directly in the original Pytorch model because their type is QINT8, not the default FP32 of Pytorch. For this reason, the running device should be specified as CPU in the verification code, otherwise the CUDA parallel framework cannot be used, which will cause errors.
[0049] S3. Quantization model deployment method
[0050] The Pytorch framework itself does not support the export of int8 type parameters. The directly output parameters are still of FP32 type, and the configuration needs to be changed manually. Here, you can call the int_repr() function to get the integer INT8 parameters. The principle is to calculate the quantized parameters based on the quantization configuration (scaling factor scale, quantization zero point zero_point).
[0051] The model deployment scheme adopted in this embodiment is to save the weights and bias parameters as binary machine language files.bin according to the network layer to which they belong, which can be subsequently transferred to the FPGA cache for reading through the AXI bus as needed. In addition, when generating the hardware circuit, it is not necessary to directly produce the IP of the entire network, but to generate convolution IP, pooling IP, etc. according to the function, so that the code is modified to call different network layer IPs to generate hardware-level neural networks in real time during the hardware experiment. However, this requires additional AXI bus transmission modules and interface designs, and additional signal transmission may bring more delays. For this patent, the biggest advantage of this method is that it can achieve operator-level optimization and achieve regulation of each refined operator. For example, the first layer of convolution has a large amount of calculation, which mainly relies on increasing parallelism and matrix division to accelerate; while the fourth layer of convolution has a much smaller amount of calculation, and the additional operation overhead of matrix division is greater than the benefits, so it is not necessary to adopt it. Through operator-level optimization, targeted acceleration of different algorithms, different data volumes, and different logical network layers can be achieved, which is the basis for the next step of software and hardware collaborative optimization.
[0052] S4. Software and hardware collaborative design
[0053] The first step of hardware-software co-design is to generate the hardware circuit of the previous BaseCNN. Currently, there is a relatively mature set of development tools in this regard. The Vitis series of development tools form a complete development chain, including HLS, which generates hardware description language IP cores through C / C++ language, Vivado platform, which uses hardware description language to generate hardware circuits and simulations, and Vitis SDK platform, which provides software development kits such as user interfaces.
[0054] After successfully generating the hardware circuit of BaseCNN, it is necessary to optimize the hardware implementation of the neural network on the HLS platform to improve the efficiency and speed of the hardware circuit, which is also a key step in the co-design of software and hardware. This patent mainly uses the following four technologies to perform co-optimization of software and hardware.
[0055] 1) Convolution parallel optimization
[0056] Attached Figure 3 The figure shows the convolution calculation in BaseCNN of this embodiment, where the number of channels of the input feature map is CHin, the convolution kernel size is K×K, the output feature map size is R×C, and the number of output channels is CHout. By determining these six parameters, the computational size of the entire convolution can be obtained. Other convolution settings are set to the default stride of 1 and padding of 0.
[0057] For this convolution, the total number of convolution kernels is CHin×CHout. A convolution operation of four layers of loop nesting is actually performed on each convolution kernel. The convolution calculation of the entire convolution layer is 6 layers of loop nesting, as shown in the following formula.
[0058]
[0059] Based on the above description, the following mathematical expression can be used to calculate the value of the output element. The purpose of convolution parallel optimization is to accelerate this calculation process and generate a general convolution operator with an accelerated structure.
[0060]
[0061] The first step of parallel optimization is to use pipeline structure to increase parallelism and improve hardware efficiency. Figure 4 As shown, a one-level pipeline is used; the initial pipeline is that when executing multiple tasks, there is no need to wait for the previous task to be completed, but the next task can be started after a certain detailed operation (such as reading data) is completed, thus realizing the pipeline work of task processing.
[0062] Furthermore, a two-level pipeline can be used to increase the degree of parallelism, which means that the two inner nested loops can be executed in parallel. However, the two-level nested loops cannot be pipelined in all cases; this requires adjusting the order between the loop nests to resolve the conflict between data access and storage. Figure 5 As an example, we describe the method of changing nested loops to achieve pipelining. The left side of the figure shows two different nested loop methods. Data processing is still performed in each loop to calculate formula (2). Here, the part shown in the red box in the figure is regarded as a computing unit. The optimization goal is to prevent conflicts between different computing units that execute this task in parallel. The right side uses the data access to the output Out to illustrate the conflict situation. If the above nested method is used, when the red box computing unit is processed in parallel, after one loop ends, the kc of the 4th layer increases by one, but the access position of Out remains unchanged. Therefore, the two Iterations in the pipeline processing start reading and calculating from Out[0][0][0]. The write operation of the previous one and the read operation of the next one generate data dependency, resulting in a conflict (according to the logical relationship, the previous iteration must write back to the next one before it can be read, otherwise a program error will occur). After adjustment, the fourth layer is c. After one loop, c is increased by one and the access position of Out is changed. The previous Iteration starts reading from Out[0][0][0], and the next Iteration starts reading from Out[0][0][1]. No data dependency is generated as before, so secondary pipelining can be performed to further improve the degree of parallelism.
[0063] 2) Matrix partitioning
[0064] When declaring the matrix A[m][n] in C language, there is generally no need to consider the storage order and position of the elements, because the processor will solve these problems in subsequent use and complete the target tasks of the written program, and the developer generally does not need to consider how to implement it. In FPGA, how these m×n elements are stored in the FPGA needs to be set by the developer, because it is related to how the subsequent hardware level is executed. By default, without using optimization instructions, the elements in the matrix A are stored continuously in a BRAM, so only one data can be read in one cycle, and the on-chip bandwidth often needs to be very large, which directly causes the problem of reduced efficiency. To solve this problem, you can specify the storage method of the matrix through matrix partitioning (Array Partition). For example, store all elements in the matrix in different BRAMs, or use multiple BRAMs to store the matrix, and store adjacent elements in different BRAMs to avoid conflicts. As shown in the attached figure Figure 6As shown, by dividing the matrix into n BRAM memories, the parallelism of the task can be greatly improved; at the same time, according to the relationship between the actual amount of calculation and the cache, using n close to the target parallelism (8 is used in this embodiment) can achieve a balance between memory access and calculation to solve the conflict problem.
[0065] 3) Loop unrolling
[0066] The principle of loop unrolling is that, in terms of the most basic hardware instruction execution, loop condition judgment has overhead, but it does not actually perform task operations. The execution efficiency can be improved by changing the proportion of condition judgment instructions in the loop body. Especially in large-scale loop body structures, the overhead of condition judgment statements cannot be ignored, and loop unrolling has a good effect. Combining loop unrolling with the above-mentioned matrix partitioning technology can greatly improve the efficiency of task parallelism. Figure 7 As shown in the figure, for example, a normal loop is used, and each execution needs to be judged once; while in loop unrolling, multiple executions only need to be judged once. In this embodiment, k=5 is set, and the execution efficiency is improved by nearly 7%.
[0067] 4) Memory access pipeline
[0068] Memory access pipelining refers to the pipelining of the operation process of FPGA reading and writing data from the cache. An important corresponding metric is the memory access cycle, which is the shortest time required to complete such a complete read and write operation. For this embodiment, this process can be understood as first moving the data from the off-chip storage (DRAM) to the faster on-chip storage (BRAM), and after the operation (such as convolution operation), the result is transferred back to DRAM. Due to possible memory conflicts and other reasons, it is necessary to plan the logic of data access when multiple tasks are in parallel, otherwise the calculation will be completed but the data is not in place, and the effect of high-performance computing cannot be exerted, that is, the "memory access bottleneck".
[0069] Based on the Dtaflow structure of HLS, the embodiment of the present invention uses Ping-Pong Buffer to perform memory access pipeline, and designs access functions for reading and writing input data and parameters. Figure 8 The following is a sample code of Ping-PongBuffer data flow processing in memory access pipeline. Two different internal cache buffers are designed. For tasks at the same time, only one buffer is used to read data from DRAM, and the data in the other buffer is on standby or for task operations. In this way, the two buffers take turns to load data from DRAM like a ping-pong ball, and the concurrent execution of data loading and computing operations is achieved while ensuring that the read and write data do not conflict.
[0070] Since this project designs a general convolution operator, even if the neural network structure is different, it can be simply called without having to modify it from scratch, which greatly reduces the development threshold.
[0071] Step S5. Hardware implementation:
[0072] After achieving the above-mentioned hardware and software co-optimization, regenerate the hardware IP and import it into the user IP library of the Vivado platform. Complete the setting, constraint and routing of the obtained hardware IP on the Vivado platform to obtain the top-level design file Block Design. After the Block Design is synthesized and correct, generate the BIT stream file readable by the FPGA, and the preparatory work of hardware implementation is completed.
[0073] However, there is another issue in hardware implementation that cannot be ignored, which is that it is necessary to design an interactive interface for observing the operation of the hardware, to observe the operating results of the FPGA and to change the input parameters in real time (i.e., the requirements for reasoning in actual scenarios). Otherwise, the FPGA will not be able to automatically print out the output results. This patent adopts a deployment solution based on the VITIS development platform: load the BIT file to generate a hardware platform project, create an APP project on the VITIS platform to obtain an interactive interface, and use AXI bus communication to transmit network structure parameters and inputs to the FPGA, thereby calling the obtained hardware IP to run the FPGA, and at the same time obtain the real-time results of the FPGA operation through the UART serial port to implement the entire process. According to the above detailed implementation method, the embodiment of the present invention adopts the attached Figure 2 The convolutional neural network BaseCNN shown in the figure is used for inference acceleration. The implementation focuses on two aspects: the first is the effect of quantization compression deployment and whether the accuracy loss of model quantization is acceptable, which determines the practical value of the invention; the second is the effect of FPGA on inference acceleration and whether the use of this solution for inference on FPGA reduces latency and memory access cycles.
[0074] After sufficient experiments and comparisons, the quantization scheme with the best quantization effect was retained. It was verified that its accuracy only dropped by 0.71%, and the overall storage size of the model was reduced from 834KB to 218KB, reducing the amount of calculation by 73.9%. The results are attached. Fig. 9 .
[0075] The FPGA deployment effect is analyzed. After the software and hardware coordinated optimization is adopted, the effect comparison is shown in the table below. The overall inference delay is reduced from 462ms to 115ms, and the speed is increased by 4 times (the FPGA clock is set to 10ns). The effect of memory access optimization is analyzed. In this embodiment, the memory access cycle (the shortest cycle between two reads and writes) after pipelining is reduced from 66 cycles to 17 cycles, and the entire data reading process time is reduced from 3146 cycles to 1547 cycles, and the total time utilization is doubled.
[0076]
[0077] The above embodiments are used to illustrate the present invention rather than to limit the present invention. Any modification and change made to the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.
Claims
1. A software and hardware collaborative acceleration method for implementing a convolutional neural network based on FPGA, characterized in that: The method includes: S1. Select a convolutional neural network and a data set, and train and quantize the convolutional neural network; S2, model deployment: save the weights and bias parameters as binary machine language files .bin according to the network layer to which they belong, and then transfer them to the FPGA cache through the AXI bus for reading as needed; S3, software and hardware collaboration: Generate hardware circuits based on the saved data, where the network layer IP is generated based on the function; use pipeline structure to increase parallelism; specify the storage method of the matrix through matrix partitioning; improve execution efficiency by changing the proportion of conditional judgment instructions in the loop body; pipeline the operation process of FPGA reading and writing data from the cache; the use of pipeline structure to increase parallelism is specifically: use a first-level pipeline structure to increase parallelism and use a second-level pipeline to increase parallelism, where the second-level pipeline is two internal nested loops, and when used, it is necessary to adjust the order between the loop nests to solve the conflict problem of access and processing to ensure that no data dependency is generated; S4. Hardware implementation: Regenerate hardware IP for data that achieves software-hardware co-optimization, and complete hardware settings, constraints, and routing based on the hardware IP.
2. According to claim 1, a method for realizing software and hardware coordinated acceleration of convolutional neural network based on FPGA, characterized in that: The neural network and the data set are selected, and a data set size that can take advantage of parallelism is selected according to the storage space of the target FPGA.
3. The method for realizing software and hardware coordinated acceleration of convolutional neural network based on FPGA according to claim 1, characterized in that: After the convolutional neural network is trained on a personal device, the neural network is quantized using a quantization compression tool under the Pytorch framework.
4. The method for realizing software and hardware coordinated acceleration of convolutional neural network based on FPGA according to claim 1, characterized in that: During the quantization process, if the prediction accuracy of the model after quantization drops by more than 5%, the quantization precision is adjusted or the quantization strategy is changed. If the accuracy loss after quantization is less than 1%, a lower quantization precision is selected to improve the compression rate.
5. The method for realizing software and hardware coordinated acceleration of convolutional neural network based on FPGA according to claim 1, characterized in that: The storage method of the specified matrix by matrix partitioning includes: storing all elements in the matrix into different BRAMs, using multiple BRAMs to store the matrix, or storing several adjacent elements into different BRAMs.
6. The method for realizing software and hardware coordinated acceleration of convolutional neural network based on FPGA according to claim 1, characterized in that: The process of pipelining the FPGA's reading and writing data from the cache is specifically as follows: data is moved from the off-chip storage to the faster on-chip storage, and after the operation, the result is transmitted back to the off-chip storage. When multiple tasks are performed in parallel, the logic of data access needs to be planned to avoid memory access bottlenecks.
7. The method for realizing software and hardware coordinated acceleration of convolutional neural network based on FPGA according to claim 1, characterized in that: In the hardware implementation, an interactive interface for observing the operation of the hardware needs to be designed to observe the operating results of the FPGA and change input parameters in real time to meet the reasoning requirements in actual scenarios.
Citation Information
Patent Citations
Method and system for deep learning algorithm acceleration on field-programmable gate array platform
CN106228238A
Depth separable convolution acceleration method, storage medium and application
CN111079904A
Software and hardware cooperation acceleration method based on FPGA
CN111178518A