Special ASIC implementation method for deep learning algorithm based on ZYNQ architecture

By designing general IP modules and state machines on the ZYNQ architecture, the coordinated work between the PS and PL ends is achieved, and the problem of high development costs of customized ASIC circuits and insufficient flexibility utilization of FPGA circuits is solved, and a dedicated ASIC for deep learning algorithms is realized with rapid deployment and efficient computing.

CN119990195APending Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510085482.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, custom ASIC circuits have high cost development, poor flexibility, insufficient flexibility and insufficient utilization of FPGA circuits, and too high requirements for staff, resulting in a longer deployment time for neural network acceleration circuits.

Method used

The dedicated ASIC implementation method for deep learning algorithms based on ZYNQ architecture is adopted. By designing a general IP module on the FPGA side, configuring a state machine on the ARM side, scheduling the IP module on the FPGA side, realizing the cooperation between the PS and PL side, configuring neural network parameters and calling multiplexable IP modules, realizing the flexible deployment of accelerated computing circuits.

Benefits of technology

It greatly shortens the time for staff to deploy neural network acceleration circuits when facing new neural networks, improves resource utilization and computing efficiency, and reduces hardware resource consumption and development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990195A_ABST
    Figure CN119990195A_ABST
Patent Text Reader

Abstract

The invention provides an ASIC implementation method special for a deep learning algorithm based on a ZYNQ architecture, belongs to the technical field of deep learning and relates to the field of deep learning, and the method comprises the steps: designing a universal IP module at an FPGA end according to a universal neural network module; an ARM end state machine is configured and used for processing and scheduling an IP module special for FPGA end hardware; the universal processor analyzes the neural network configuration information and the weight data, and transmits the neural network configuration information to the ARM end cache module; the FPGA end reads instructions of the PS end through the AXILite to execute specified operation, reads related weights through the AXIDMA, and stores the related weights in the corresponding cache modules; and the FPGA end performs operations such as convolution and pooling according to an instruction of the ARM end, and the ARM end reads an interrupt signal value of the FPGA end to judge a convolution state, sends a subsequent instruction and sends a final generation result back to the ARM end for storage so as to facilitate subsequent processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and specifically relates to a dedicated ASIC implementation method for deep learning algorithms based on a ZYNQ architecture. Background Art

[0002] In the 1950s and 1960s, people imitated the brain's synaptic connections to create the concept of neural networks and built a mathematical model for processing information. From the initial perceptron to the multi-layer perceptron and the back-propagation algorithm, neural networks have played an important role in solving linear inseparable problems. It was not until 2006 that deep learning was proposed, and the great potential of neural networks began to be reflected. Deep neural networks consist of input layers, hidden layers, and output layers. They learn features layer by layer, effectively approximate the mapping relationship between original data and features, and have the function of rapid learning for various things. However, in the training and prediction process, because the requirements for the accuracy of input results are gradually increasing, a large amount of data needs to be processed at the same time. However, when processing large-scale data, people need to face problems such as large computing resource requirements and long computing time. Therefore, people began to think about how to accelerate the computing process in big data networks. Therefore, using the characteristics of hardware circuits to help achieve repeated computing has become a popular choice for people. Among them, the mainstream hardware platforms are ASIC and FPGA. ASIC focuses on improving hardware architecture to achieve algorithm acceleration, with high computing efficiency but high development cost and poor flexibility. Currently, there are two main methods for using FPGA to accelerate neural networks:

[0003] One is to deploy the complete neural network model directly on the FPGA to achieve acceleration: This architecture can greatly improve the utilization of computing resources and make full use of the parallel computing capabilities of hardware circuits for acceleration. However, resources need to be allocated reasonably to prevent performance bottlenecks caused by insufficient resources. The hardware resource requirements are very high, and it is difficult to have enough resources to support different calculations for each layer of the complete neural network.

[0004] The second is to deploy some operations of the specific network to the FPGA, and deploy some related network structures to the FPGA chip according to the specific type of the network: Although this method can reduce the resource occupancy rate, it greatly reduces the flexibility of the FPGA and reduces the reusability. In addition, when transplanting different networks, a lot of modifications need to be made to the underlying neural network modules, which is undoubtedly a huge test for the workload and experience of the staff. Summary of the invention

[0005] The purpose of the present invention is to provide a dedicated ASIC implementation method for deep learning algorithms based on the ZYNQ architecture, which effectively solves the problems of high development cost, poor flexibility, insufficient flexibility utilization of FPGA circuits, and excessive requirements on staff of customized ASIC circuits. The present invention realizes a method of flexibly deploying accelerated computing circuits by using ZYNQ as the architecture, cooperating between the PS and PL ends, configuring neural network parameters on the PS end, and calling reusable IP modules on the PL end, which greatly shortens the time for staff to deploy neural network acceleration circuits when facing new neural networks.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture includes the following steps:

[0008] Step 1: Design a general IP module on the FPGA side according to the general neural network module, wherein the general neural network module includes data acquisition, convolution, pooling, activation, and upsampling modules;

[0009] Step 2: Decompose different neural network models into multi-layer networks, determine the neural module to be called, and configure the state machine on the ARM side to schedule the IP module on the FPGA side.

[0010] As a preferred solution of the present invention, for IP modules with different functions, a hardware acceleration solution is designed according to data flow characteristics to construct an integrated circuit, wherein the hardware acceleration solution includes multi-channel convolution, DSP multiplication calculation multiplexing, and lookup table to implement activation function.

[0011] As a preferred solution of the present invention, a reusable IP acceleration module is designed according to network characteristics, and relevant state machines are designed according to parameters such as whether to fill, convolution kernel size, input size, and sending data type to ensure that the data flow direction is consistent with the propagation direction of the neural network.

[0012] As a preferred solution of the present invention, a state machine is designed on the PS side of the ZYNQ architecture to indicate the forward propagation direction of the neural network, and an instruction parsing module is designed on the PL side to parse the instruction information transmitted by the PS side via AXI_Lite and direct the data to the correct storage or usage module.

[0013] As a preferred solution of the present invention, the data path between the PS end and the PL end is configured, including configuring AXI_VDMA, AXI_DMA, AXI_Lite and interrupt settings, wherein AXI_VDMA is used to expand the interface to support subsequent application expansion of real-time display solutions, AXI_DMA is used to read and fill in image data in the DDR of the PS end, AXI_Lite is used to set the register module and transmit instructions and status information, and the interrupt setting is used to transmit the interrupt signal after the PL end processing is completed.

[0014] As a preferred solution of the present invention, a top-level module is designed at the PL end, and the state machine in the top-level module is used to convert data reading and writing, convolution, padding, upsampling, calculation completion and other operations into semaphores, which are transmitted between modules, and a dedicated acceleration IP module is designed according to the data flow direction, bit width, data type and acceleration scheme.

[0015] As a preferred solution of the present invention, a software optimization solution is adopted, including the fusion of BN layer and CONV layer and Int8 quantization, wherein the fusion of BN layer and CONV layer improves the forward inference speed of the model by merging the BN layer parameters into the convolution layer, and Int8 quantization converts the input data and weights from floating point numbers to integers through a specific formula, and dequantizes the output data, and performs a 14-bit left shift operation when the FPGA processes the convolution result; in the fusion of BN layer and convolution layer:

[0016] The convolutional layer formula is: conv =w×x+b;

[0017] The BN layer formula is:

[0018] And x i is the result of the previous convolution, so after the two formulas are combined:

[0019] make:

[0020]

[0021] This completes the fusion of the convolutional layer and the BN layer;

[0022] In Int8 quantization, the following steps are included:

[0023] S1. Determine the scaling factor (s) and zero point (z) quantization processing of each layer according to the distribution and range of the input data and weights, and convert the input data and weights of each layer from floating point numbers to integers. The calculation formula is as follows:

[0024]

[0025] S2, then dequantize the output data of each layer and convert the integer back to floating point number. The calculation formula is as follows:

[0026] X=S(Q(X)-z)

[0027] The purpose is to re-perform the quantization operation when the next data is input to avoid data overflow or loss.

[0028] As a preferred solution of the present invention, the multi-channel convolution uses DMA to transmit multi-channel weights and image data parameters at one time, and adopts multi-RAM and multi-Fifo storage modes to perform fast convolution operations; the DSP multiplication calculation reuse uses the principle of (A+D)B=AB+DB in the DSP48 multiplication resource, fills the low-order Int8 data and expands the high-order sign bit to calculate, so as to realize the DSP48 multiplier resource reuse; the lookup table realizes the activation function using the LeakyReLu function, caches it in the form of a lookup table inside the FPGA, reads the PS end data through the PL end for caching operation, and uses the input data as the read address to obtain the activation processing output value, wherein:

[0029] The Leaky ReLu function is defined as:

[0030] Where: The Leaky ReLu function is used to cache the activation function in the form of a lookup table inside the FPGA, and the PS end data is read through the PL end for cache operation. When the activation operation is to be performed, the input data is used as the cache read address, and its output is the output value of the activation process.

[0031] As a preferred solution of the present invention, a storage medium stores a PS-side basic state machine code and a configuration file, which can configure a specific network structure and send a data read instruction.

[0032] An embedded device that contains a dedicated ASIC driver module, which brings together most of the basic neural network operation IPs and top-level module IPs, and can be used with the PS end to implement a neural network hardware acceleration solution.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] 1. The present invention adopts a batch normalization layer and a convolutional layer fusion solution to reduce hardware resource consumption. Each network layer reduces one layer of batch normalization network calculation, and the data calculation rate is significantly increased;

[0035] 2. The present invention adds an Int8 quantization module, and uses Int8 quantization operations for all operation data types. Under the premise of losing very little precision, the large bit width calculation of the 32-bit data type is abandoned and replaced by 8-bit bandwidth calculation, which can further reduce the time consumed in the neural network calculation; compared with the original network, data transmission and data calculation data are greatly improved;

[0036] 3. The present invention adopts a multi-channel convolution method in the network layer operation module, transmits multi-channel convolution data at one time through AXI_DMA, saves output weights in a multi-port multi-RAM manner, and realizes the acceleration effect of multi-channel simultaneous operation through three-row and same-column output modules and DSP multiplexing modules; the speed effect is geometrically multiple compared with single-channel calculation speed;

[0037] 4. When using DSP48 multiplier resources, the present invention utilizes the multiplier rule of DSP48 and uses a method of multiplying high and low bits by the same factor to realize the multiplexing operation of DSP48 multiplier, thereby realizing the maximum utilization of hardware resources, further reducing its operation time, and utilizing hardware resources to the maximum extent;

[0038] 5. When considering the convolution layer calculation, the size of the convolution kernel is generally 3X3 or 1X1. Therefore, in the design, the same product principle is used to determine whether the weight data in the calculation is spliced ​​according to the size of the convolution kernel. If it is a 3x3 convolution, 9 8-bit weight data need to be spliced, and a write enable signal is generated after each splicing. If it is a 1x1 convolution, no splicing is required, and the 8-bit weight data is directly written to the RAM, and a write enable signal is generated for each data. Therefore, compared with other invention schemes, the multiplexing operation of the 3X3 convolution module is realized, and there is no need to add an additional 1X1 convolution module circuit, thereby improving resource utilization;

[0039] 6. In the activation function module, the present invention uses a lookup table method to pre-calculate the function value and store it in an array, taking into account the quantization operation and the resource occupancy rate caused by the function calculation; when the function value is needed, the system quickly obtains the approximate value by looking up this table; thereby achieving the activation function acceleration effect, which improves the acceleration effect and reduces the resource occupancy rate compared to the direct calculation scheme;

[0040] 7. The design method of the present invention is universal. The PS side has designed the driver file for the IP of the scheduling configuration PL side. The user only needs to fill the parameters of the convolution layer into the corresponding list and modify the corresponding weight file. The network layer of the neural network can be accelerated according to the correspondence, and the state machine can be modified. The configuration on the PS side conforms to the selected network algorithm structure to build a complete deep learning algorithm hardware acceleration solution; it provides application examples and solutions for the engineering application of embedded accelerated neural networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0042] Figure 1 The resource configuration of the ZYNQ architecture used in the embodiment of the present invention;

[0043] Figure 2 is a practical framework diagram in an embodiment of the present invention;

[0044] Figure 3 is the PS-side network register configuration in the embodiment of the present invention;

[0045] Figure 4 is a schematic diagram of an instruction parsing module in an embodiment of the present invention;

[0046] Figure 5 This is a schematic diagram of the first weight data buffer module in an embodiment of the present invention.

[0047] Figure 6 is a schematic diagram of a second weight data buffer module in an embodiment of the present invention;

[0048] Figure 7 is a filling module scheme diagram in an embodiment of the present invention;

[0049] Figure 8 It is a flowchart of three rows and columns of output modules during network layer calculation in an embodiment of the present invention;

[0050] Fig. 9 is a flowchart of DSP resource reuse during network layer calculation in an embodiment of the present invention;

[0051] Fig.10 is a flow chart of batch convolution during network layer calculation in an embodiment of the present invention;

[0052] Fig.11 This is a flow chart of the comparison of the pooling layer beat delay in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] Example 1

[0055] See also Figure 1-Figure 11 , the present invention provides the following technical solutions:

[0056] A method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture includes the following steps:

[0057] Step 1: Design a general IP module on the FPGA side according to the general neural network module, wherein the general neural network module includes data acquisition, convolution, pooling, activation, and upsampling modules;

[0058] Step 2: Decompose different neural network models into multi-layer networks, determine the neural module to be called, and configure the state machine on the ARM side to schedule the IP module on the FPGA side.

[0059] In a specific embodiment of the present invention, a general IP module is designed on the FPGA side according to a general neural network module. Common neural network modules include data acquisition, convolution, pooling, activation, upsampling, etc.

[0060] The PS side is mainly responsible for data preprocessing, scheduling, and analysis:

[0061] First, according to the configuration of the deep learning algorithm that you want to achieve hardware acceleration, obtain its standard network structure and configure the PS-side state machine. The main contents of the PS-side state machine include determining the number of network layers of the algorithm, the network type of each layer, the input and output data flow, etc. According to the data flow, the jump condition of the state machine is set to the interrupt signal set when the calculation of the previous layer of network is completed, so that after the calculation of the PL end is completed, it will jump to the next-end state machine of the PS end, thereby receiving the new instruction size of the next layer;

[0062] Extract the characteristics of the network layer, such as the input image dimension size, the output image dimension size, whether to fill, whether to pool, the data type, the convolution kernel size, the step size, etc. In order to save the characteristics of the convolution layer on the PS side, the present invention saves the conventional characteristics of each convolution layer in sequence on the PS side in the form of a structure array. During the data calculation process, after each new network layer is run, the PS side will transmit the network layer configuration information to the PL side through AXI_Lite. At this time, the instruction parsing module on the PL side will divert data according to the configuration information of the network layer, and transmit the currently transmitted data to the corresponding acceleration calculation module;

[0063] Use the training set data suitable for the current network, use Python to retain and output the weight parameters, bias parameters, etc. of the neural network after training, and save them in the SD card. Use the SD card peripheral on the PL side of the ZYNQ architecture to save the data on the PS side.

[0064] The PS side obtains image data input from the outside, such as camera input or SD card reading, and saves the data to the PS side DDR3; reads the network parameter file (bin file) obtained by Pytorch training from the SD card, including weights, biases, etc.; the PS side sends a control signal to the PL side through the AXI_Lite interface to guide the calculation level and the required data transmission;

[0065] The PL side (FPGA) is responsible for performing computationally intensive operations such as convolution, activation function processing, pooling, and upsampling; these operations are optimized for neural network algorithms, such as reducing bit precision, fusing normalization layers and convolution layers, etc. to reduce memory space and bandwidth requirements

[0066] The PL side will first build a top-level module, which is mainly used to control the orderly transmission between the modules on the PL side. When receiving data from the PS side, it will determine the write data and read data status. After reading or writing, it will first call the corresponding module according to the instruction obtained by the instruction parsing module;

[0067] The PS-side instructions are first transmitted to the instruction parsing module on the PL side via the top-level module. The instruction parsing module mainly parses the corresponding bit width of the registers set by AXI_Lite. After parsing, it is transmitted back to the top-level module to help the top-level module determine the subsequent calling module.

[0068] After receiving the network layer acceleration operation instruction,

[0069] First, the PS side reads the weight parameters of the current layer from the SD card, including the convolution kernel and bias, and sends them to the weight cache module on the PL side through AXI_DMA;

[0070] Further, the PS side informs the PL side through AXI_Lite that the current transmission is the weight parameter;

[0071] Then, the PS side reads the output result of the previous layer network from DDR3 and sends it to the filling module on the PL side through the AXI_DMA module;

[0072] At the same time, the PS side informs the PL side through AXI_Lite that the current data being transmitted is image data;

[0073] Furthermore, the PL side calls the padding module to perform padding operations on the image data according to the instruction judgment, expands it to the desired size, and sends the padded data to the convolution module; if there is no padding instruction, this step will be skipped;

[0074] After that, the convolution module on the PL side performs convolution calculation on the padded data according to the weight parameters stored in the weight cache module, and sends the result to the activation module;

[0075] Furthermore, the activation module on the PL side calculates the activation function on the convolution result according to the transmission parameters on the PS side, and sends the result to the pooling module;

[0076] Finally, the pooling module performs the maximum pooling operation on the activation result and writes the result into DDR3 as the input of the next layer of network;

[0077] After a layer of network layer acceleration, the PL side will trigger an interrupt instruction to be transmitted back to the PS side, and the PS side will call the next instruction according to the network structure;

[0078] After multiple network layers are cyclically run, the output result of the last layer is finally obtained, which conforms to the selected neural network result structure. The result of the last output layer is output to the PS end, and the PS end will perform corresponding result decoding operations on it through the protocol of each network output result.

[0079] Specifically, the present invention also provides a storage medium, which includes a PS-side basic state machine code and a configuration file, and can configure a specific network structure and send a data read instruction.

[0080] Specifically, the present invention also provides an embedded device, which includes an exclusive ASIC driver module, which integrates most of the basic operation IPs of the neural network and the top-level module IP, and can be used in conjunction with the PS end to implement a neural network hardware acceleration solution.

[0081] Specifically, for IP modules with different functions, hardware acceleration solutions are designed according to data flow characteristics to construct integrated circuits. The hardware acceleration solutions include multi-channel convolution, DSP multiplication calculation multiplexing, and lookup table implementation of activation functions.

[0082] In this embodiment: for different functional IP modules, different hardware acceleration schemes are designed according to the flow characteristics of data, so as to construct different integrated circuits; according to different neural network models, it is necessary to disassemble them into multi-layer networks to determine the called neural module, so as to configure the corresponding state machine on the ARM side to perform orderly scheduling of the IP module of FPGA; the neural module includes: convolution, pooling, activation, etc.

[0083] Specifically, a reusable IP acceleration module is designed based on network characteristics, and relevant state machines are designed based on parameters such as whether to fill, convolution kernel size, input size, and sent data type to ensure that the data flow is consistent with the propagation direction of the neural network.

[0084] In this embodiment: Based on the characteristics of the network, a reusable IP acceleration module is designed, and based on relevant parameters, such as whether to fill, the convolution kernel size, the input size, and the type of data being sent, a related state machine is designed so that the data flow direction conforms to the propagation direction of the neural network.

[0085] Specifically, a state machine is designed on the PS side of the ZYNQ architecture to indicate the forward propagation direction of the neural network, and an instruction parsing module is designed on the PL side to parse the instruction information transmitted by the PS side via AXI_Lite and direct the data to the correct storage or usage module.

[0086] In this embodiment: a relevant state machine is designed on the PS side (ARM side) in the ZYNQ architecture to indicate the forward propagation direction of the neural network, and a relevant instruction parsing module is designed on the PL side (FPGA side) to parse the instruction information transmitted by the PS side through AXI_Lite, thereby directing the data transmitted from the PS side to the correct storage module or usage module.

[0087] Specifically, configure the data path between the PS and PL ends, including configuring AXI_VDMA, AXI_DMA, AXI_Lite and interrupt settings, where AXI_VDMA is used to expand the interface to support subsequent application expansion of real-time display solutions, AXI_DMA is used to read and fill in the image data in the DDR of the PS end, AXI_Lite is used to set the register module and transmit instructions and status information, and the interrupt setting is used to transmit the interrupt signal after the PL end processing is completed.

[0088] In this embodiment: configure the relevant data paths between the PS end and the PL end, configure AXI_VDMA to design an expansion interface so that subsequent applications such as cameras can expand the real-time display solution, configure AXI_DMA to read and fill in the PS end DDR3, mainly the data for saving image data, configure AXI_Lite to set the relevant register module, design the bit interpretation of the corresponding bit width, and implement the corresponding protocol with the PL end, the main transmission data is instruction information and status information; configure the relevant interrupt settings, the main data is the interrupt signal generated after the PL end code processing is completed, which is used to transmit the PL end status.

[0089] Specifically, a top-level module is designed on the PL side, and the state machine in the top-level module is used to convert data reading and writing, convolution, padding, upsampling, calculation completion and other operations into semaphores, which are transmitted between modules. An exclusive acceleration IP module is designed based on the data flow direction, bit width, data type and acceleration solution.

[0090] In this embodiment: In the overall architecture, the PL-side top-level module is first designed, and the state machine in the PL-side top-level module is used to convert operations such as data reading and writing, data convolution, data filling, data upsampling, and data calculation completion into semaphores, which are transmitted to each module to help each module perform each part of the operation in an orderly manner, and design a dedicated acceleration IP module based on the data flow direction, bit width, data type, acceleration scheme, etc.

[0091] Specifically, the exclusive ASIC implementation based on the ZYNQ architecture mentioned in this method is implemented by using a code optimization solution combined with a hardware circuit optimization solution, and a software optimization solution is adopted, including the fusion of the BN layer and the CONV layer and Int8 quantization, wherein the fusion of the BN layer and the CONV layer improves the forward inference speed of the model by merging the BN layer parameters into the convolution layer, and the Int8 quantization converts the input data and weights from floating point numbers to integers through a specific formula, and dequantizes the output data, and performs a 14-bit left shift operation when the FPGA processes the convolution result;

[0092] At present, when training deep network models, most neural networks will perform batch normalization (BN) layers after the convolutional layers. After the BN layer normalizes the data, it can effectively solve the gradient vanishing and gradient exploding problems, and can accelerate network convergence and control overfitting. However, although the BN layer plays a positive role in training, it adds some layers of calculations during network forward inference, which affects the performance of the model and occupies more memory or video memory space. By merging the parameters of the BN layer into the convolutional layer, the speed of the model forward inference can be greatly improved. In the fusion of the BN layer and the convolutional layer:

[0093] The convolutional layer formula is: conv =w×x+b;

[0094] The BN layer formula is:

[0095] And x i is the result of the previous convolution, so after the two formulas are combined:

[0096] make:

[0097]

[0098] This completes the fusion of the convolutional layer and the BN layer;

[0099] In Int8 quantization, in order to avoid the low efficiency of floating-point calculation, fixed-point operations are often used inside FPGAs. When processing the convolution result, the 14th power of 2 is enlarged by shifting left by 14 bits, which helps to convert the floating-point number into a fixed-point number after preparing it in integer form. The specific steps include the following:

[0100] S1. Determine the scaling factor (s) and zero point (z) quantization processing of each layer according to the distribution and range of the input data and weights, and convert the input data and weights of each layer from floating point numbers to integers. The calculation formula is as follows:

[0101]

[0102] S2, then dequantize the output data of each layer and convert the integer back to floating point number. The calculation formula is as follows:

[0103] X=S(Q(X)-z)

[0104] The purpose is to re-perform the quantization operation when the next data is input to avoid data overflow or loss.

[0105] The process is as follows:

[0106] A. When processing the convolution result, the FPGA will perform a 14-bit left shift operation to expand the result by 2 to the 14th power. This operation prepares for the subsequent floating-point to fixed-point conversion;

[0107] B. The quantization parameters of the Int8 type obtained in Python will be transmitted to the instruction parsing module through the AXI-Lite interface;

[0108] C. The parsing module will parse the received data and send the parsed data to the quantization module for quantization operations. These operations may include quantizing the data according to parameters such as thresholds, weights, and biases.

[0109] Hardware strategy:

[0110] Multi-channel convolution:

[0111] By taking advantage of the fact that DMA can transmit multi-bit width information at one time, multi-channel weights and image data parameters can be transmitted at one time. By utilizing the storage mode of multiple RAMs and multiple Fifos, data can be quickly convolved in a multi-channel manner, achieving maximum resource utilization and accelerated computing on the ZYNQ architecture.

[0112] DSP multiplication calculation reuse:

[0113] DSP48 is a hardware resource dedicated to digital signal processing (DSP) in Xilinx FPGA. It can realize high-speed and high-precision arithmetic operations, such as addition, subtraction, multiplication, accumulation, etc.; using the principle of (A+D)B=AB+DB in DSP48 multiplication resource, when DSP performs 25*18 bit calculation, we fill the low-order Int8 data, extend the high-order sign bit, fill the low-order with other Int8 data for calculation, and finally obtain the high and low bits of the data respectively. , You can realize the maximum reuse of DSP48 multiplier resources.

[0114] Specifically, the multi-channel convolution uses DMA to transfer multi-channel weights and image data parameters at one time, and adopts multi-RAM and multi-Fifo storage modes to perform fast convolution operations; the DSP multiplication calculation reuse uses the principle of (A+D)B=AB+DB in the DSP48 multiplication resource, fills the low-order Int8 data and expands the high-order sign bit, and then calculates to realize DSP48 multiplier resource reuse; the lookup table implements the activation function using the LeakyReLu function, caches it in the form of a lookup table inside the FPGA, reads the PS end data through the PL end for caching operations, and uses the input data as the read address to obtain the activation processing output value;

[0115] Lookup table implements activation function:

[0116] The activation function in the present invention uses Leaky ReLu as an example, which is a variant of the correction line unit. When the input is a negative value, it will not be completely suppressed to 0, but a small part of the negative value is allowed to pass. The Leaky ReLu function is defined as:

[0117] The Leaky ReLu function is defined as:

[0118] Where: The Leaky ReLu function is used to cache the activation function in the form of a lookup table inside the FPGA, and the PS end data is read through the PL end for cache operation. When the activation operation is to be performed, the input data is used as the cache read address, and its output is the output value of the activation process.

[0119] Example 2

[0120] Specifically, the present invention also provides a specific implementation method of a deep learning algorithm dedicated ASIC implementation method based on the ZYNQ architecture, which is used to further explain and illustrate the working principle of the method provided in Example 1, as follows:

[0121] The ZYNQ architecture used in the present invention is shown in FIG. Figure 1 As shown, it includes SD peripherals, DDR3, HP interface, GP interface, etc. The main purpose is to save network data parameters and AXI bus protocol;

[0122] Figure 2 This is a practical framework diagram of an embodiment of the present invention, and the overall data flow is in this diagram:

[0123] The PS side first analyzes the specific network framework to be arranged and writes the network model. It also saves the relevant weight parameters to the SD card on the PL side.

[0124] The PL side obtains data streams through peripherals such as cameras and transmits them to the PS side DDR3 through AXI_VDMA for subsequent processing;

[0125] After the PS end has built the network model, Figure 3 As shown, fill in the relevant network layer parameters, start sending command data to the PL end, and prepare to call the relevant modules;

[0126] First, the instruction will be handed over to the instruction parsing module for processing via the top-level module of PL;

[0127] like Figure 4 Shown

[0128] By setting the registers that need to correspond to the requirements in the AXI_Lite protocol IP, the PS and PL sides can change and read the same address of the same register to determine the operations that both parties need to perform at this time. And according to the corresponding signals, the BIt bits in the register are divided accordingly, so as to achieve coordinated acceleration of each module and the generation of completion interrupt when the operation is completed;

[0129] The main operation is to parse the calculator information of Axi4-lite, including parameters such as data type, convolution type, convolution kernel size, number of input and output channels, input and output feature map size, etc.; this information is stored in four registers, and is extracted and assigned to the corresponding signals according to the bit allocation given in the table;

[0130] Based on the parsed calculator information, feedback is generated to the PS end to coordinate and control the various modules inside the accelerator, such as the arithmetic logic unit (ALU), data storage unit (DSU), data transmission unit (DTU), etc. The main control module needs to send corresponding control signals to make each module work as expected;

[0131] The modules inside the accelerator include Padding, convolution calculation, transposition, upsampling, etc. The main control module is responsible for coordinating and controlling the working timing of these modules, transmitting relevant parameters, and receiving completion signals of each module.

[0132] After a certain stage of work is completed, the main control module will send an interrupt request signal to the PS end to notify the PL end to prepare to receive the next command or process the next set of data. The interrupt request is a high-level signal that lasts for 200 clock cycles;

[0133] The interrupt request is implemented through a state machine, which controls the state change of the interrupt signal according to different operation types and completion status; this function is used to send a notification to the PS end after the accelerator completes the current convolution or upsampling operation;

[0134] Then, through the instruction command, the PL side knows what type of data the PS side is transmitting through AXI_DMA at this time;

[0135] The PL side receives the image input data from the DMA, with each transmission of 32 bits. In this example, a 4-channel convolution scheme is set. To ensure that 4-channel convolution calculations are performed simultaneously, a 4Fifo storage design is performed.

[0136] According to the timing relationship of the stream interface, the state signal is used to control the generation and pull-down of the tready signal; the state signal is transmitted through the instruction parsing module to indicate the current state of the module; when the state is in the write state, it indicates that the PL end is ready to receive data;

[0137] Output different types of data: According to the data type signal, generate valid flag signals of different types of data, and output the signal to the subsequent modules; the data type signal is transmitted from the PS end through the Axi4-lite register, indicating the type of data currently sent. According to the value of data type, pull up the valid flag signal of the corresponding type of data, and pull down the valid flag signal of other types; output tdata signal to the subsequent modules to distinguish and process different types of data.

[0138] At this time, in this instance, the weight data will be transmitted first. If the command is parsed as weight data, the data will be input into the weight cache module.

[0139] The bias data is stored in a bin file as a 32-bit data type. Each bin file corresponds to the bias parameters of all channels of a layer of convolution. DMA sends 32 bits of data each time, which is 1 bias parameter.

[0140] Therefore, each 64-bit data contains the bias parameters of one channel. The current accelerator solution will calculate four transfer channels at the same time, so it is necessary to ensure that the bias cache module can output the bias parameters of four channels at the same time. Therefore, a 32-bit wide RAM cannot be used directly, but four 16-bit wide RAMs need to be used, and the 32-bit data must be allocated to different RAMs according to certain rules. Figure 5 shown.

[0141] Specifically, if Figure 6 As shown in the figure, each time a 32-bit data is received, the upper 16 bits are written to RAM0 and the lower 16 bits are written to RAM1; when the second 32-bit data is received, the upper 16 bits are written to RAM2 and the lower 16 bits are written to RAM3; and so on. The purpose of this is to allow each RAM to store the bias parameters of two adjacent channels for subsequent reading.

[0142] In order to achieve multi-channel output, such as the four-channel output in this example, the outputs of the four RAMs can be bound together at the same time, and then the four output ports are combined to transmit the data of four channels at a time.

[0143] For each output channel's weight cache module, the weight parameters of the four input channels need to be output simultaneously.

[0144] Use RAM for storage: the bit width is: 3*3*8=72 bits; both reading and writing are 72 bits (reusing 1*1 convolution and 3*3 convolution), and the depth is 256.

[0145] Then, according to the state machine scheduling of the PS end, the image data will be passed in according to the same data caching principle. In this example, it will first determine whether the data is filled, such as Figure 7 shown.

[0146] The input image data will be divided into three scenes for filling. This data parameter is also saved in the instruction information:

[0147] 1. Contains the first row of data;

[0148] 2. Contains the last row of data;

[0149] 3. It does not include either the first row or the last row of data;

[0150] The PS side tells the instruction parsing module on the PL side the current input data requirements through AXI_Lite.

[0151] If the first row of data is included, add another row of 0s before the first row, and add 0s on both sides;

[0152] If you want to include data in the middle, just add 0 at the beginning and end of each line;

[0153] If the last row of data is included, add another row of 0 after the last row;

[0154] There will be no situation where both the first and last rows are included

[0155] Determine the position type according to set_type in the instruction parsing module (0 means including the first line, 1 means the middle position, and 2 means including the last line);

[0156] Get the number of columns and rows of the current cache according to the col_cnt and row_cnt in the register;

[0157] like Figure 8 As shown, the image data will then enter the three rows and columns for corresponding output.

[0158] The main purpose is to convert the input feature map data into the data of the same column of 3 adjacent rows when performing 3x3 convolution calculations, so as to construct a 3x3 matrix for convolution operations; this module helps to improve parallelism and processing efficiency;

[0159] But for the reuse module, the fixed length of Fifo will limit the size of the neural network. Therefore, we make improvements on this solution:

[0160] The register size output is defined as the maximum image size when the network is input. In the network with a size smaller than the maximum, the three rows of the network are output in advance during transmission, and the variable-length three-row same-column output module can be realized:

[0161] Use two variable length shift registers named Shift Reg A and Shift Reg B;

[0162] When the first row of data is input, store it in the first position of Shift Reg A (shift_reg_a[0]).

[0163] When the second row of data comes in, the second row of data is stored in the first position of Shift Reg A (shift_reg_a[0]), and the first row of data stored in Shift Reg A is shifted to the second position (shift_reg_a[1]);

[0164] When the third row of input data comes in, the third row of data is stored in the first position of Shift Reg A (shift_reg_a[0]), and the second row of data stored in Shift Reg A is shifted to the second position (shift_reg_a[1]), and the first row of data stored in Shift Reg A is shifted to the third position (shift_reg_a[2]);

[0165] Then a buffer is used to cache the first two columns of data, and when the third column of data comes out, it is combined with the first two columns of data in the buffer to form a 3x3 image matrix;

[0166] After that, the three lines of output data will be convolved with the previously cached weight data. At this time, the DSP48 multiplier will be used. Using the principle of the DSP48 multiplier, such as Fig. 9 As shown, DSP48 multiplexing calculation operation is performed;

[0167] The input sequence and convolution kernel are respectively mapped to the input ports A and B of DSP48, each element is an 18-bit signal, and the output port P of DSP48 is connected to a 48-bit register R for storing the accumulated result;

[0168] In each clock cycle, DSP48 calculates AB and adds the result to R, that is, R = R + (A + D) B. When all elements of the input sequence and the convolution kernel are calculated, the value in R is the output value of the convolution operation. The principle of using DSP48 to complete the multiplication operation in convolution is:

[0169] The DSP48 contains an 18x18-bit multiplier, a 48-bit accumulator, and some control logic and registers. The input ports A and B of the DSP48 can choose whether to perform sign extension, that is, to complement the high bit with a sign bit or 0, to adapt to different data types.

[0170] Through the architecture, we can know that a normal DSP can perform 25*18 bit calculations. In the process of simple convolution, we can only use DSP resources once for multiplication. Therefore, in order to maximize the utilization of resources, we adopt a DSP multiplexing scheme (the same factor is used when multiplying);

[0171] At this time, A and D are used as different weights, and B is used as the input image data for multiplexing multiplication; finally, when performing convolution calculation on a single convolution kernel, multiple channels of the convolution kernel can be calculated at one time;

[0172] After we calculate the value of a single convolution kernel, because the overall architecture of this example is 4-channel parallel calculation, the data of a single channel will be accumulated first, and then the total number of channels will be accumulated. Therefore, the data will be input into Fig.10 The batch convolution module shown.

[0173] By processing in batches, the consumption of computing resources can be reduced, especially in large-scale convolution operations; calculating the convolution of only a part of the channels each time can effectively reduce the amount of calculation, save memory usage and increase the calculation speed.

[0174] Thereby improving the computational efficiency. The calculation results of each batch are finally accumulated to obtain the overall convolution output, while ensuring the correctness of the calculation results.

[0175] It can improve memory access patterns, reduce conflicts between memory accesses, and increase data reuse rates, thereby further improving computing efficiency.

[0176] First, implement the single-channel accumulation module

[0177] Design ideas (accumulation logic):

[0178] 1. Use a FIFO to cache the convolution results of the previous batch and set the batch_type signal to decide whether to perform accumulation operation;

[0179] 2. When batch_type == 0, it means that the input data is the first element of the first batch, that is, the input data is directly written into the FIFO;

[0180] 3. When batch_type == 1, it means that the input data is not the first element of the first batch, that is, the input data is added to the data in FIFO and then written into FIFO;

[0181] 4. When batch_type == 2, it means that the input data is the first element of the second batch, that is, the input data is added to the data in FIFO and then output;

[0182] 5. The read enable of FIFO is determined by batch_type and data_in_vld (data validity), ensuring that FIFO reads data only when it needs to be accumulated or output.

[0183] 1. The accumulation logic of each channel in the multi-channel accumulation module is the same as that of the single-channel accumulation module. The input data and output data of each channel are passed to the corresponding single-channel accumulation module instance;

[0184] 2. Four single-channel accumulation modules are instantiated internally in the example, corresponding to the accumulation operation of each channel respectively; thus a four-channel batch volume accumulation module is realized;

[0185] The final convolution completed data will enter the activation module according to the instruction information, and the activation function will be arranged and quantized using the lookup table method;

[0186] First, the activation module will split the input 32-bit data into four 8-bit data and store them in RAM; the RAM IP core is configured as a dual-port, with a write bit width of 32 bits and a depth of 64; the read bit width is 8 bits and a depth of 256;

[0187] For each channel, the corresponding convolution output value is used as the read address of the RAM, for example, channel 0 corresponds to RAM0; this allows the corresponding activation function value to be obtained at the same time, speeding up the data processing process;

[0188] This design can effectively implement cache storage and read operations of activation function data, and quickly obtain activation function values ​​after convolution calculation, thereby accelerating the overall data processing process;

[0189] The activated data will be put into the pooling module, which performs pooling operations on the data.

[0190] Reduce the amount of calculation and prevent overfitting;

[0191] First, if Fig.11 As shown, the first row of data passed in is processed by a register. This means that the data in adjacent positions are compared and processed to get a valid result;

[0192] Data comparison and obtaining the maximum value:

[0193] Compare the packaged data with the original data, find the maximum value data at each adjacent position, and write it into the FIFO;

[0194] In this way, the depth of the FIFO is reduced while retaining important data, because only the maximum value of each line needs to be recorded.

[0195] Line-by-line operation:

[0196] 1. A similar method is used to process the second row of data to obtain the maximum value of the row;

[0197] 2. Compare the existing data in FIFO with the new maximum value to obtain the maximum value of the two rows of data;

[0198] Final result: By comparing with the data in FIFO, the maximum value from the two rows of data is taken as the final pooling output result; this method cleverly reduces the depth of FIFO and successfully implements data pooling.

[0199] Functional implementation: Through the above method, the data is effectively pooled and the maximum value in each area is extracted, thereby achieving the purpose of optimizing data processing and reducing storage requirements.

[0200] The use of register beat processing, data comparison and maximum value acquisition helps improve module performance and simplify the data processing process. Through this method, effective data pooling processing is achieved, providing a more compact data structure for subsequent tasks;

[0201] In this example, the final output data will enter the instruction again to determine whether upsampling is required;

[0202] The main purpose of the upsampling module is to expand the input data so that the input size of the image increases while keeping the number of channels unchanged;

[0203] The core principle is to achieve data expansion by replicating one piece of data into four pieces of data.

[0204] The feature map data is stored in FIFO as a temporary queue;

[0205] Use the buffer_rd_en enable signal to control the up-sampling function of the data. In odd rows, the buffer_rd_en enable signal is pulled high every other clock cycle;

[0206] During the high level period, data is read from the FIFO, and the two read data are packaged and cached in another FIFO;

[0207] In odd-numbered rows, buffer_rd_en is activated clock cycle by clock cycle to read, process and package data;

[0208] The processed data is cached and read out again when it enters the even-numbered row. This alternating operation completes the upsampling of the data;

[0209] Finally, repeating the above network layer calculations, we will get the final result of the deep learning algorithm, which will be transmitted to the PS end for storage, and the PS end will perform decoding operations according to the network characteristics;

[0210] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiment method, the C language code on the PS side can be modified to instruct the relevant hardware to complete the process, and the program can be stored in a computer-readable storage medium.

[0211] Technicians can choose different models of ZYNQ according to their own resource conditions. The choice of storage media is not limited to SD cards, DDR3 and other storage peripherals. As for image input acquisition, they can choose real-time input from the camera or local storage of image data, etc., and can utilize the relevant technical solutions in the present invention.

[0212] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for implementing a deep learning algorithm dedicated to ASIC based on ZYNQ architecture, characterized in that: The steps include: Step 1: Design a general IP module on the FPGA side according to the general neural network module, wherein the general neural network module includes data acquisition, convolution, pooling, activation, and upsampling modules; Step 2: Decompose different neural network models into multi-layer networks, determine the neural module to be called, and configure the state machine on the ARM side to schedule the IP module on the FPGA side.

2. The method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture according to claim 1, characterized in that: For IP modules with different functions, hardware acceleration solutions are designed according to data flow characteristics to construct integrated circuits. The hardware acceleration solutions include multi-channel convolution, DSP multiplication calculation multiplexing, and lookup table implementation of activation functions.

3. The method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture according to claim 2, characterized in that: A reusable IP acceleration module is designed based on network characteristics, and related state machines are designed based on parameters such as whether to fill, convolution kernel size, input size, and sent data type to ensure that the data flow is consistent with the propagation direction of the neural network.

4. The method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture according to claim 3, characterized in that: A state machine is designed on the PS side of the ZYNQ architecture to indicate the forward propagation direction of the neural network. An instruction parsing module is designed on the PL side to parse the instruction information transmitted by the PS side via AXI_Lite and direct the data to the correct storage or usage module.

5. The method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture according to claim 4, characterized in that: Configure the data path between the PS and PL ends, including configuring AXI_VDMA, AXI_DMA, AXI_Lite and interrupt settings. AXI_VDMA is used to expand the interface to support subsequent application expansion of real-time display solutions, AXI_DMA is used to read and fill in the image data in the DDR of the PS end, AXI_Lite is used to set the register module and transfer instructions and status information, and the interrupt setting is used to transmit the interrupt signal after the PL end processing is completed.

6. The method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture according to claim 5, characterized in that: Design the top-level module on the PL side, use the state machine in the top-level module to convert data reading and writing, convolution, padding, upsampling, calculation completion and other operations into semaphores, transmit them between modules, and design exclusive acceleration IP modules based on data flow, bit width, data type and acceleration solution.

7. The method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture according to claim 6, characterized in that: A software optimization solution is adopted, including the fusion of BN layer and CONV layer and Int8 quantization. The fusion of BN layer and CONV layer improves the forward inference speed of the model by merging BN layer parameters into the convolution layer. Int8 quantization converts input data and weights from floating point numbers to integers through a specific formula, and dequantizes the output data. At the same time, a 14-bit left shift operation is performed when the FPGA processes the convolution result. In the fusion of BN layer and convolution layer: The convolutional layer formula is: conv =w×x+b; The BN layer formula is: And x i is the result of the previous convolution, so after the two formulas are combined: make: This completes the fusion of the convolutional layer and the BN layer, where: conv represents the output result of the convolution layer, w represents the weight parameter of the convolution kernel, x is the input data of the convolution layer, and b is the bias parameter. represents the i-th output value after BN layer processing, γ is the scaling factor, x i is the i-th element in the convolution result of the previous layer, is the input data of the BN layer, and μ is the input data x i The mean of σ represents the input data x i The standard deviation of , ∈ is a small positive number, which is used to prevent the denominator from being 0 and ensure the stability and effectiveness of the formula calculation. β is the offset parameter; In Int8 quantization, the following steps are included: S1. Determine the scaling factor (s) and zero point (z) quantization processing of each layer according to the distribution and range of the input data and weights, and convert the input data and weights of each layer from floating point numbers to integers. The calculation formula is as follows: Where: Q(X) represents the quantized result, X represents the input data to be quantized, S is the scaling factor, and z is the zero point; S2, then dequantize the output data of each layer and convert the integer back to floating point number. The calculation formula is as follows: X = S (Q (X) - z); Where: X here represents the result after inverse quantization, and the others are the same as the above quantization formula; The purpose is to re-perform the quantization operation when the next data is input to avoid data overflow or loss.

8. The method for implementing a deep learning algorithm dedicated ASIC based on the ZYNQ architecture according to claim 7, characterized in that: The multi-channel convolution uses DMA to transfer multi-channel weights and image data parameters at one time, and adopts multi-RAM and multi-Fifo storage modes for fast convolution operations; the DSP multiplication calculation reuse uses the principle of (A+D)B=AB+DB in the DSP48 multiplication resource, fills the low-order Int8 data and expands the high-order sign bit to calculate, so as to realize the DSP48 multiplier resource reuse; the lookup table realizes the activation function using the LeakyReLu function, caches it in the form of a lookup table inside the FPGA, reads the PS end data through the PL end for cache operation, and uses the input data as the read address to obtain the activation processing output value, wherein: The Leaky ReLu function is defined as: Where: f(x) is the output value of the Leaky ReLu function, x is the input value of the function, the Leaky ReLu function caches the activation function inside the FPGA in the form of a lookup table, and the PS end data is read through the PL end for cache operation. When the activation operation is to be performed, the input data is used as the cache read address, and its output is the output value of the activation process.

9. A storage medium, characterized in that: The basic state machine code and configuration file of the PS side in the method described in claim 1 are stored, and a specific network structure can be configured and a data reading instruction can be sent.

10. An embedded device, characterized in that: It contains a dedicated ASIC driver module for the method described in claim 1, which brings together most of the basic neural network operation IPs and top-level module IPs, and can be used with the PS side to implement a neural network hardware acceleration solution.

Citation Information

Cited By

  • Data control channel reconstruction method, data decoupling transmission method, equipment and medium

    CN121029655A