AI accelerator card based on distributed FPGA
By adopting a multi-chip collaborative design scheme on FPGA, the problem of insufficient resources of monolithic FPGA is solved, efficient computing of large-scale convolutional neural networks is achieved, and the performance of edge AI devices is improved.
Patent Information
- Application Number
- CN202411379121.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-13
AI Technical Summary
Insufficient resources are needed when deploying convolutional neural networks in a monolithic FPGA, resulting in limited computing performance and inability to effectively handle complex deep learning models.
Using the design scheme of collaborative work of multiple FPGAs, a large-scale convolutional neural network model is deployed on multiple FPGAs. Each FPGA performs different operations and transmits and communicates through high-speed serial channels and concurrent data buses.
It effectively improves the operation efficiency of large convolutional neural networks in FPGAs, solves the problem of insufficient resources, and realizes efficient computing on edge AI devices.
Smart Images

Figure CN119990205A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an AI acceleration card based on a distributed FPGA, belonging to the technical field of FPGA hardware accelerator devices. Background Art
[0002] With the rapid development of artificial intelligence (AI) and the Internet of Things (IoT) technologies, edge computing has gradually attracted widespread attention. However, this also brings about the trade-off between energy consumption and computing power. Convolutional Neural Networks (CNN) improve the reasoning accuracy of deep learning tasks by mapping between specific weights and input features, and are therefore widely used in computer vision, voiceprint recognition and other fields. However, since CNN involves a large number of convolution calculations, these calculations require high computing power, making it difficult to deploy CNN in power-constrained edge computing devices. In order to meet this challenge, low-power, high-computing-power CNN accelerators have received increasing attention in recent years.
[0003] FPGA (Field Programmable Gate Array) has attracted great attention in the field of neural network hardware acceleration due to its high programmability and parallel processing capabilities. First, the high parallel processing capability of FPGA enables it to perform a large number of convolution operations at the same time, thereby significantly improving computing efficiency. Compared with traditional general-purpose processors (CPUs), FPGAs can greatly shorten computing time by processing multiple convolution kernel operations in parallel. Secondly, the programmability of FPGAs provides a high degree of flexibility. Designers can dynamically adjust hardware configurations and optimize the allocation of computing resources according to specific CNN model requirements. This flexibility enables FPGAs to adapt to changing deep learning algorithms and maintain efficient computing performance. FPGAs also have the advantage of low power consumption. Compared with GPUs (graphics processing units), FPGAs generally consume less power when performing the same computing tasks. This makes FPGAs particularly suitable for edge computing scenarios, providing efficient computing capabilities in power-constrained environments such as mobile devices and Internet of Things (IoT) devices.
[0004] Although FPGAs have shown many advantages in CNN acceleration, they also face the problem of insufficient resources in practical applications. First, CNNs usually require a lot of computing resources to process convolution operations, pooling, and fully connected layers, which require a large number of DSP units and LUTs. When designing high-performance CNN accelerators, frequent convolution operations will quickly consume the DSP resources on the FPGA, and the number and depth of convolution kernels will also make LUT resources tight. Secondly, the on-chip storage (BRAM) capacity on FPGAs is limited. For large CNN models, especially those containing millions or even hundreds of millions of parameters, BRAM may not be able to accommodate all model weights and intermediate calculation results. At this time, frequent data exchange needs to be carried out through off-chip storage, resulting in data transmission bottlenecks, which seriously affects the overall computing performance. In addition, as the complexity of CNN models increases, deep models such as residual networks (ResNet) and densely connected convolutional networks (DenseNet) require far more resources than the load of general FPGAs. In this case, FPGAs cannot fully load and process the entire model while maintaining high performance, and can only partially alleviate the problem of insufficient resources through methods such as model pruning, quantization, and block computing.
[0005] In order to cope with the lack of resources, researchers usually adopt several strategies. One is to optimize the hardware architecture and improve resource utilization through resource sharing, operator reuse and other technologies. The other is model compression and optimization, such as using sparse matrices and low-precision calculations to reduce computing resource requirements. And so on. Summary of the invention
[0006] In view of the problem of insufficient resources in deploying convolutional neural networks on a single FPGA at this stage, the present invention proposes a convolutional neural network acceleration card based on the collaborative work of multiple FPGAs, which can deploy large-scale convolutional neural network models on multiple FPGAs and perform different computing operations on different FPGAs to achieve the purpose of efficient acceleration of large-scale convolutional neural networks.
[0007] The technical solution of the present invention is an AI acceleration card based on distributed FPGA, which contains a distributed FPGA acceleration card and a convolutional neural network deployment based on the acceleration card. The distributed FPGA acceleration card adopts a distributed FPGA design scheme, the main control unit uses a ZYNQ chip with rich LUT resources, and the computing unit uses an FPGA chip with rich DSP resources; the main control unit and the computing unit are connected through a GTH bidirectional high-speed serial channel and a 16 / 32bit bidirectional concurrent data bus to ensure efficient data transmission and communication; each FPGA chip is equipped with 2GB DDR3 memory for temporary storage, and NVME hard disk and SD card interfaces are reserved for persistent data storage; the convolutional neural network deployment scheme based on the acceleration card is mainly composed of a main control unit module, a computing unit module and a data cache module; the main control unit module is deployed in the ZYNQ chip, responsible for configuring and managing the computing unit, data preprocessing and post-processing, sending and synchronizing control instructions, and distributing and collecting data; the computing unit module is deployed in the distributed computing unit FPGA chip to perform convolution calculations, quantization, activation, pooling and upsampling operations.
[0008] The AI acceleration card can be a main control unit that controls four computing units.
[0009] The main control unit model includes but is not limited to XCZU19EG, and the computing unit model includes but is not limited to XCZU15EG.
[0010] The main control unit module can be divided into a PS for reconfigurable software parameter design, an off-chip storage DDR, an input and output data cache module, a data management module, etc. Among them, DDR and on-chip data communication adopts the form of DMA. The PS for reconfigurable software parameter design completes the parameter instruction design according to the currently executed convolutional neural network model, including convolution type, pooling selection, activation and quantization parameters, data address and data volume, etc. The distributed computing unit module can be divided into a data transceiver control layer, a computing control layer, a data input and output cache, a convolution operation module, a quantization module, an activation module, a pooling module, an upsampling module, etc. Among them, the convolution operation module is a general convolution module including 3×3 and 1×1 convolutions; the convolution module includes a convolution control layer, a padding layer, a convolution matrix construction layer and a convolution kernel; each convolution kernel is composed of 9 DSP units, and the data of two channels are processed by feature multiplexing; 32-channel convolution arrays are deployed in the four computing units, and 20 3×3 convolution kernels are configured for each channel. The AI accelerator card contains any one or more of a UART interface, an HDMI interface, a DP interface, a Gigabit Ethernet interface, a QSFP optical port, and a PCIe interface. The AI accelerator card can be powered by a 12V 6PIN×2 interface.
[0011] This distributed FPGA design can effectively improve the operating efficiency of large convolutional neural networks in FPGAs, and the modularity and flexibility of the design provide a good foundation for subsequent optimization and expansion. Through further optimization and improvement, higher performance and efficiency can be achieved in more complex neural network architectures and a wider range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A schematic diagram of an architecture of the present invention.
[0013] Figure 2 This is a schematic diagram of the main control unit module structure when deploying a convolutional neural network in the present invention.
[0014] Figure 3 This is a schematic diagram of the structure of a computing unit module when deploying a convolutional neural network in the present invention.
[0015] Figure 4 This is a timing diagram of the first data transmission when the present invention deploys YOLOv3-Tiny. DETAILED DESCRIPTION
[0016] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The specific embodiments are used to illustrate the present invention rather than to limit the present invention.
[0017] The present invention includes two parts: one is an AI acceleration card design based on distributed FPGA, and the other is a convolutional neural network deployment solution based on the AI acceleration card. An AI acceleration card based on distributed FPGA, the AI acceleration card adopts a distributed FPGA design solution, integrating multiple FPGAs into a PCB board in a star-ray topology structure, wherein the main control unit uses a ZYNQ chip with more LUT resources, while the computing unit uses an FPGA chip with more DSP resources. This architectural design fully utilizes the flexibility and control capabilities of the main control ZYNQ chip, while using the powerful DSP resources of DPGA to perform complex computing tasks. The overall architecture is as follows: Figure 1 As shown, the connection relationship between each component is clearly demonstrated, making the system well maintainable and scalable.
[0018] In the distributed computing card solution, the communication between the main control unit and a computing unit requires 68 IOs, so the specific distributed computing solution can be expanded to multiple chips according to the IO of the main control unit, and the computing unit can also select FPGA chips with different DSP resources according to different needs.
[0019] Figure 1Here, the invention is described by taking the main control unit as XCZU19EG and four external XCZU15EG as distributed computing units as an example.
[0020] Communication and data transmission: The computing unit and the main control unit are connected by a GTH bidirectional high-speed serial channel, and the data channel is a 16 / 32bit bidirectional concurrent data bus. This design ensures efficient data transmission and communication, which is beneficial to the performance of the accelerator card when processing large-scale data, and also provides sufficient bandwidth to support complex neural network calculations.
[0021] PCB design: In the PCB design of the distributed FPGA neural network accelerator card, the star ray structure and the JTAG daisy chain structure are two key parts. The star ray structure is a common network topology used to connect multiple FPGA chips so that each FPGA chip can communicate directly with the central control unit. In the distributed FPGA neural network accelerator card, the application of the star ray structure can ensure that each FPGA computing unit can efficiently receive and send data. The central node of the structure is usually the master FPGA or the master processor, which is connected to each computing unit FPGA through a high-speed serial communication channel (such as GTH or GTY). The advantage of this topology is its low latency and high bandwidth, which is suitable for the needs of large-scale data transmission; JTAG daisy chain structure: The JTAG (Joint Test Action Group) interface is used for hardware debugging and programming. The daisy chain structure is a serial connection method that connects multiple FPGA chips in series so that debugging and programming signals can be passed to each chip in turn. In the distributed FPGA neural network accelerator card, the use of the JTAG daisy chain structure simplifies the hardware debugging process. In specific implementation, the TDI (Test Data In) port of the JTAG interface is connected to the input of the first FPGA, the TDO (Test Data Out) port is connected to the input of the next FPGA, and so on, until the TDO port of the last FPGA is connected to the debugging device. The advantage of the daisy chain structure is that its hardware wiring is simple and resource-saving, but it is necessary to ensure that the signal transmission between each FPGA chip is reliable.
[0022] Storage design: Each FPGA chip is equipped with 2GB of DDR4 memory for temporary storage of data required for calculations, such as model parameters and intermediate results. At the same time, each FPGA chip is equipped with QSPI-FLASH storage for circuit solidification. The main control unit is also equipped with EMMC memory for configuration file storage. In addition, NVME hard disk and SD card interfaces are reserved. These storage devices can be used for persistent data storage, such as training data and model files. Such a storage design not only meets computing needs, but also provides a flexible data management method.
[0023] Interface design: In addition to the data storage interface, the accelerator card is also designed with a variety of interfaces, including UART interface for control and communication, HDMI, DP, Gigabit Ethernet, QSFP optical port for data transmission, and PCIE interface for high-speed communication with PC. These interfaces can meet the needs of different scenarios and enhance the flexibility and scalability of the accelerator card.
[0024] Power design: The power supply interface of the accelerator card uses a 12V 6PIN×2 interface. To ensure stable operation of the system, the accelerator card can also be equipped with a power management module to monitor and manage the power supply of each component. This ensures that each component can work at the appropriate voltage and current, thereby improving the reliability and stability of the system.
[0025] The deployment scheme of the convolutional neural network based on this accelerator card is as follows:
[0026] According to the structure of the accelerator card, the computing unit module, the main control unit module, and the data cache module are designed separately when deploying the convolutional neural network. The main control unit module is mainly used to configure and manage the computing unit module, data preprocessing and post-processing, control the sending and synchronization of instructions, and data distribution and collection; the computing unit module performs specific convolutional neural network operation functions, mainly including convolution calculation, quantization, activation, pooling calculation, upsampling, etc.; the data cache module is a data buffer designed in each FPGA, which is used to temporarily store parameter data used for calculation and result data after calculation.
[0027] Figure 2 FIG. 1 is a schematic diagram of a main control unit module structure when deploying a convolutional neural network in the present invention. The design framework of the main control unit module is as follows: Figure 2As shown in the figure, it is mainly composed of PS for reconfigurable software parameter design, off-chip storage DDR, input and output data cache module, and data management module. Among them, the parameter instruction design is completed on the PS side according to the currently executed convolutional neural network model, including convolution type, pooling selection, activation and quantization parameters, data address and data volume, etc. After the system is powered on, the PS will first cache the weight, bias, activation and other parameter files stored in the SD card or NVME hard disk in the DDR, wait for the feature data input from the peripheral interface and cache a frame of data in the DDR, and then send the data to the on-chip buffer in batches, and the data management module will send the data to four parallel convolution calculation units in batches. Finally, the calculation unit module sends the data after the convolution to the main control unit cache, and redistributes the data before the next operation. When sending data to the calculation unit module, the main control unit module will first send a preamble containing the convolution control instruction, which is generated by the instruction pool in the control layer.
[0028] Figure 3 FIG. 1 is a schematic diagram of a computing unit module structure when deploying a convolutional neural network in the present invention. The computing unit module is as follows: Figure 3 As shown, it includes a data transceiver control layer, a computing control layer, a data input and output cache, a convolution operation module, a quantization module, an activation module, a pooling module, and an upsampling module. First, the data transceiver control layer receives the feature and parameter data sent by the main control unit module through the GTH bus and caches it in the corresponding on-chip buffer. After the data is cached, the computing unit module control layer sends a command signal to start the convolution operation. After the convolution operation is completed, quantization and activation processing are performed, and then the next step is performed according to whether the current batch of data needs pooling or upsampling. Finally, all the processed batch data is cached in the output cache and sent to the main control unit module through the data transceiver control layer.
[0029] The design of the convolution operation module is the core of the entire computing unit. The present invention designs a universal convolution operation array for the common 3×3 and 1×1 convolutions in the network. The convolution operation module is composed of a convolution control layer, a filling layer, a convolution matrix construction layer and a convolution kernel. When the convolution instruction sent by the computing unit control layer is received, the control layer will first determine whether the current feature data needs to be sent to the filling module for filling according to the instruction, and then send the filled data to the matrix construction module for convolution matrix construction, and also store the weight value in the weight FIFO in a flattened manner. After the matrix construction is completed, the weight and feature data are sent to the convolution kernel for convolution operation. Finally, the operation result is temporarily stored in the output data cache module.
[0030] The convolution kernel design is a 3×3 convolution. When performing a 1×1 convolution, it is only necessary to control the size of the convolution matrix in the matrix construction layer to achieve a 1×1 convolution. Since a single DSP can complete P=(A+D)×B, two sets of different weights are added and multiplied by the features in the form of feature data reuse to obtain data for two output channels. Taking XCZU15EG as an example, the chip has 3528 DSP resources. According to the solution of the present invention, a 32-channel convolution array is designed and deployed, and each channel is configured with 20 3×3 convolution kernels. Each convolution kernel is composed of 9 DSP units, and each convolution kernel can process data from two channels through feature reuse. In this way, a total of 2880 DSP resources are used, and the DSP resource utilization rate is 81.6%, which fully utilizes the computing power on the FPGA to achieve efficient convolution operations.
[0031] Exemplarily, the present invention is used to complete the deployment of YOLOv3-Tiny.
[0032] The method is as follows:
[0033] Step 1: According to the network structure of YOLOv3-Tiny, design the control instructions of the PS side of the main control unit module, including convolution, pooling, quantization, activation, upsampling, data transmission address and transmission amount, etc.
[0034] Step 2: Allocate data according to the convolution design of each computing unit module. Taking the first layer of YOLOv3-Tiny as an example, the total data volume is 416×416×3, and the data volume of each computing unit module with 32 channels requires 34 rows of data, so the data is divided into 13 batches, of which the first and last batches contain one row of padding data, and the data is sent to the four computing units in parallel for calculation in 4 times. Figure 4 It is a schematic diagram of the data transmission timing when the first layer of YOLOv3-Tiny is deployed using the present invention.
[0035] Step 3: The four computing units cache the data sent by the main control unit to DDR3 respectively, and read the weight data saved in the SD card to DDR3, and then perform filling, convolution, quantization, and activation operations according to the instructions of the main control unit, and then cache the output data to DDR3 again, waiting for the master control to send instructions.
[0036] Step 4: After receiving the data sending instruction from the master control, each computing unit sends the calculation results cached in the DDR to the master control unit for caching, and performs pooling or the next layer of convolution according to the corresponding instructions of the network parameters, and redistributes the data.
[0037] Step 5: Similarly, pooling and upsampling are also deployed in four computing units. The computing units can choose whether to pool or upsample according to the instructions after activation, and then directly send the pooled or upsampled data back to the main control unit, thereby reducing the number of data transmission and reception times to improve the network's operating efficiency.
[0038] Step 6: After all calculations are completed, the operation results of the entire network are output. When designing the accelerator card, multiple interfaces such as Gigabit Ethernet, fiber optic interface, HDMI, DP, serial port, etc. are reserved for data input and output.
[0039] In summary, the present invention proposes an AI acceleration card based on multiple distributed FPGAs. When deploying a convolutional neural network model, data can be sent in batches to different FPGAs for calculation and a high-speed bus can be used for communication. This solves the problem of insufficient resources of a single FPGA when deploying a convolutional neural network model, thereby improving the operating efficiency of the network model on the edge AI device.
Claims
1. An AI accelerator card based on distributed FPGA, characterized in that: The AI accelerator card contains: a distributed FPGA accelerator card and its convolutional neural network.
2. The AI acceleration card based on distributed FPGA according to claim 1, characterized in that: The distributed FPGA acceleration card has a ZYNQ chip as the main control unit and multiple FPGA chips containing DSP as the computing unit. The main control unit and the computing unit are connected through a GTH bidirectional high-speed serial channel and a 16 / 32-bit bidirectional concurrent data bus. Each FPGA chip is equipped with 2GB DDR3 memory for temporary storage, and NVME hard disk and SD card interfaces are reserved for persistent data storage.
3. The AI acceleration card based on distributed FPGA according to claim 1, characterized in that: The convolutional neural network is composed of a main control unit module, a computing unit module and a data cache module; the main control unit is deployed in the ZYNQ chip and is responsible for configuring and managing the computing unit, data pre-processing and post-processing, sending and synchronizing control instructions, and distributing and collecting data; The computing unit modules are deployed into distributed FPGA chips to perform convolution calculations, quantization, activation, pooling, and upsampling operations.
4. The AI acceleration card based on distributed FPGA according to claim 1 or 2, characterized in that: The AI accelerator card is a main control unit that controls four distributed computing units.
5. The distributed FPGA-based AI acceleration card according to claim 3, characterized in that: The main control unit module model is XCZU19EG, and the computing unit module model is XCZU15EG.
6. The AI acceleration card based on distributed FPGA according to claim 3, characterized in that: The data cache module is a data buffer of the FPGA, which is used to temporarily store parameter data used for calculation and result data after calculation.
7. The AI acceleration card based on distributed FPGA according to claim 5, characterized in that: The computing unit module is divided into a data transceiver control layer, a computing control layer, a data input and output cache, a convolution operation module, a quantization module, an activation module, a pooling module, and an upsampling module.
8. The AI acceleration card based on distributed FPGA according to claim 7, characterized in that: The convolution operation module is a general convolution module including 3×3 and 1×1 convolutions; the convolution module includes a convolution control layer, a padding layer, a convolution matrix construction layer and a convolution kernel; each convolution kernel is composed of 9 DSP units and processes data of two channels through feature multiplexing; 32-channel convolution arrays are deployed in four computing units respectively, and 20 3×3 convolution kernels are configured for each channel.
9. The AI acceleration card based on distributed FPGA according to claim 1, characterized in that: The AI accelerator card also contains any one or more of a UART interface, an HDMI interface, a DP interface, a Gigabit Ethernet interface, a QSFP optical port, and a PCIe interface.
10. The AI acceleration card based on distributed FPGA according to claim 1, characterized in that: The AI accelerator card is powered by a 12V 6PIN×2 interface.