Heterogeneous processing system based on FPGA and GPU

By configuring the FPGA chip as an RDMA network card and integrating it with the GPU chip, the communication bottleneck problem of traditional GPU clusters is solved, achieving high-efficiency GPU computing performance and efficiency improvement, and reducing network latency and resource waste.

CN223897879UActive Publication Date: 2026-02-10EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202520351067.3
Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2026-02-10
Estimated Expiration
2035-03-03

AI Technical Summary

Technical Problem

When training large models using traditional GPU clusters, communication accounts for a large proportion of the workload, and insufficient network performance leads to a waste of computing resources. Furthermore, traditional network protocols are prone to congestion and packet loss, which affects computing power.

Method used

The FPGA chip is configured as an RDMA network card and integrated with the GPU chip to achieve a unified system. The FPGA chip can directly access remote nodes, reducing the dependence on the kernel processing system and improving the computing performance and efficiency of the GPU chip.

Benefits of technology

It improves the computing performance and efficiency of GPU chips, reduces communication latency, enhances the overall computing power of GPU clusters, and reduces resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN223897879U_ABST
    Figure CN223897879U_ABST
Patent Text Reader

Abstract

The utility model discloses a heterogeneous processing system based on an FPGA (Field Programmable Gate Array) and a GPU (Graphics Processing Unit), which comprises a plurality of cascaded CPU (Central Processing Unit) servers, the CPU server comprises a CPU chip and a plurality of heterogeneous computing modules, each heterogeneous computing module comprises an FPGA chip and a GPU chip, and the FPGA chips and the GPU chips are interconnected through PCIE; the FPGA chip and the GPU chip respectively realize interaction with the CPU chip through PCIE (Peripheral Component Interface Express) Wherein the FPGA chip is configured to realize an RDMA (Remote Direct Memory Access) network card function. A heterogeneous scheme formed by the FPGA and the GPU is adopted, the FPGA chip is configured to achieve the function of the RDMA network card, the GPU chip is used for data processing and calculation, integration of the RDMA network card and the GPU computing power card is achieved through integration of the FPGA chip and the GPU chip, and the computing performance of the GPU chip can be effectively improved; the GPU chip can directly access the far-end node through the FPGA chip, a kernel processing system is not needed any more, and the operation efficiency of the GPU chip can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This utility model belongs to the field of cloud computing, and in particular relates to a heterogeneous processing system based on FPGA and GPU. Background Technology

[0002] As the number of parameters in large AI models skyrockets from hundreds of millions to trillions, the massive computing power required to support their training has become increasingly important. The computational power for these large training tasks necessitates a computing cluster comprised of servers with numerous Graphics Processing Units (GPUs). These servers connect via a network, exchanging massive amounts of data. However, data shows that the synchronous communication between servers behind each computation can reach hundreds of gigabytes. Therefore, even with the most powerful individual GPUs, if network performance is insufficient, the overall computing power of the cluster will be significantly reduced. Thus, a large cluster does not equate to high computing power; on the contrary, the larger the GPU cluster, the greater the additional communication overhead.

[0003] For large-scale models with hundreds of billions or trillions of parameters, communication can account for up to 50% of the training process, which is far beyond the bandwidth of traditional low-speed networks. Simultaneously, traditional network protocols are prone to network congestion, high latency, and packet loss; even a mere 0.1% packet loss can lead to a 50% loss of computing power, ultimately resulting in a severe waste of computing resources. Therefore, there is an urgent need for a system that can achieve maximum network performance, enabling GPU clusters to reach maximum throughput, while simultaneously supporting GPUs to achieve high utilization and computing power. Utility Model Content

[0004] The purpose of this invention is to address the aforementioned problems by providing a heterogeneous processing system based on FPGA and GPU. It employs a heterogeneous scheme of FPGA chip + GPU chip, where the FPGA chip is configured to implement RDMA network interface card (NIC) functionality, and the GPU chip is used for data processing and computation. The integration of the FPGA chip and GPU chip achieves the integration of the RDMA NIC and the GPU computing card, effectively improving the computing performance of the GPU chip. Furthermore, the GPU chip can directly access remote nodes through the FPGA chip without going through the kernel processing system, thus significantly improving the GPU chip's computing efficiency.

[0005] To achieve the above objectives, the technical solution adopted by this utility model is as follows:

[0006] A heterogeneous processing system based on FPGA and GPU includes: multiple cascaded CPU servers, which interact with each other through a switch; each CPU server includes a CPU chip and multiple heterogeneous computing modules, which include FPGA chips and GPU chips, and the FPGA chips and GPU chips are interconnected through PCIe.

[0007] The FPGA chip and the GPU chip interact with the CPU chip via PCIe respectively;

[0008] The FPGA chip is configured to implement RDMA network card functionality.

[0009] The FPGA chip and the GPU chip are interconnected via a direct PCIe_0 connection.

[0010] The FPGA chip supports multiple 100Gbps network interfaces and is connected to Ethernet.

[0011] The GPU chip accesses remote nodes through the FPGA chip and the Ethernet.

[0012] The FPGA chip transmits the access request of the remote node to the GPU chip via PCIe, thereby enabling the interaction between the GPU chip and the remote node.

[0013] The CPU server further includes a storage module for storing data exchanged between the GPU chip and the remote node.

[0014] The FPGA chip and the CPU chip interact with each other via a PCIe_1 interface.

[0015] The CPU chip interacts with the remote node through the FPGA chip.

[0016] The GPU chip is used for data processing and computation, and interacts with the CPU chip through PCIe.

[0017] The heterogeneous processing system further includes a power supply unit, which provides power to the CPU server.

[0018] The beneficial effects of this utility model are:

[0019] 1. In this application, the FPGA chip is configured to implement the RDMA network card function, and the GPU chip is used for data processing and calculation. By integrating the FPGA chip and the GPU chip, the RDMA network card and the GPU computing card can be integrated, which can effectively improve the computing performance of the GPU chip.

[0020] 2. The GPU chip in this application can directly access remote nodes through the FPGA chip without going through the kernel processing system, which can effectively improve the computing efficiency of the GPU chip.

[0021] To make the above and other objects, features and advantages of this utility model more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the specific embodiments of this utility model, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this utility model. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of a heterogeneous processing system based on FPGA and GPU, provided for an embodiment of the present invention.

[0024] Figure 2 A schematic diagram of the CPU server structure provided for an embodiment of this utility model.

[0025] Figure 3 This is an architecture diagram of an FPGA chip used to implement RDMA network card functionality, provided in an embodiment of this utility model. Detailed Implementation

[0026] The foregoing and other technical contents, features, and effects of this utility model will be clearly presented in the following detailed description of a preferred embodiment with reference to the accompanying drawings. The directional terms mentioned in the following embodiments, such as up, down, left, right, front, or back, are only for reference to the accompanying drawings. Therefore, the directional terms used are for illustrative purposes and not for limiting the scope of this utility model.

[0027] This application adopts a heterogeneous solution of FPGA chip + GPU chip. The FPGA chip is configured to implement RDMA network card function, and the GPU chip is used for data processing and calculation. By integrating the FPGA chip and the GPU chip, the RDMA network card and the GPU computing card are integrated, which can effectively improve the computing performance of the GPU chip. The GPU chip can directly access the remote node through the FPGA chip without going through the kernel processing system, which can effectively improve the computing efficiency of the GPU chip.

[0028] Example:

[0029] like Figure 1As shown, a heterogeneous processing system based on FPGA and GPU includes: multiple cascaded CPU servers, which interact with each other through a switch; each CPU server includes a CPU chip and multiple heterogeneous computing modules, which interact with the CPU chip through PCIe.

[0030] Specifically, such as Figure 2 As shown, the heterogeneous computing module includes an FPGA (Field-Programmable Gate Array) chip and a GPU (Graphics Processing Unit) chip. The FPGA and GPU in the heterogeneous computing module are interconnected at high speed via a PCIe (Peripheral Component Interconnect Express) bus. The GPU chip can directly implement RDMA drivers and application calls to the FPGA to access remote nodes. For access requests from remote nodes, the FPGA chip notifies the GPU chip via PCIe, enabling interaction between the GPU chip and the remote node, and synchronously storing the data in the storage module. The FPGA chip also acts as an intermediary, receiving access requests from remote nodes and then notifying the GPU chip via PCIe to perform corresponding operations, ensuring that data is synchronously stored in the storage module.

[0031] The FPGA chip is configured to implement RDMA network card functionality, supporting multiple 100Gbps network interfaces for Ethernet connection. The FPGA chip and GPU chip are interconnected via PCIe_0 direct connection (e.g., implementing Gen4.0 x8 to match the 100G bandwidth rate). The GPU chip can directly access the GPU computing card of the remote node through the FPGA chip without going through the CPU, which can effectively improve the computing efficiency of the GPU chip.

[0032] The CPU server also includes a storage module, which is used to store the interaction data between the GPU chip and the remote node.

[0033] like Figure 2 As shown, the FPGA chip and GPU chip are interconnected with the CPU via PCIe, forming a high-efficiency heterogeneous processing system. Specifically, the FPGA chip and CPU chip are interconnected via the PCIe_1 interface to achieve data interaction (Gen4.0 x8, matching the 100G bandwidth rate, ensuring high-speed data exchange between the FPGA and the CPU). This connection method not only supports efficient data interaction but also enables general-purpose network card functions.

[0034] Furthermore, by using the latest PCIe standard, high-bandwidth and low-latency communication is guaranteed between the FPGA and the CPU, as well as between the GPU and the CPU. The FPGA is configured to support multiple 100Gbps network interfaces, making it an ideal high-performance network solution. It can not only handle large amounts of network traffic, but also implement various customized network protocols and acceleration functions through its programmability. The programmability of the FPGA allows for the rapid implementation of new network protocols or the optimization of existing protocols to adapt to the ever-changing network environment requirements.

[0035] GPU chips are used for data processing and computation. They are interconnected with CPU chips via PCIe to enable interaction between the CPU and GPU. Specifically, CPU chips can also implement RDMA drivers and access the CPU or GPU of remote nodes through RDMA applications to achieve interaction with remote nodes.

[0036] like Figure 3 As shown, the FPGA chip is based on the ROCEv2 and iWARP architectures, implementing the IB transport layer protocol, UDP / IP protocol, TCP / IP protocol, and data link layer protocol. RDMA is an ultra-high-speed network memory access technology that allows programs to access the memory of remote computing nodes at extremely high speeds. The FPGA chip implements RDMA network card functionality based on the ROCEv2 / iWARP architecture and Internet protocols, without relying on dedicated IB switches, saving cluster deployment costs while leveraging the inherent reconfigurability and customizability of the FPGA itself.

[0037] The heterogeneous processing system also includes a power supply unit, which provides power to the CPU server.

[0038] This application adopts a heterogeneous solution of FPGA chip + GPU chip. The FPGA chip is configured to implement RDMA network card function, and the GPU chip is used for data processing and calculation. By integrating the FPGA chip and the GPU chip, the RDMA network card and the GPU computing card are integrated, which can effectively improve the computing performance of the GPU chip. The GPU chip can directly access the remote node through the FPGA chip without going through the kernel processing system, which can effectively improve the computing efficiency of the GPU chip.

[0039] Although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0040] The detailed descriptions listed above are merely specific descriptions of feasible implementations of this utility model, and are not intended to limit the scope of protection of this utility model. All equivalent implementations or modifications made without departing from the spirit of this utility model should be included within the scope of protection of this utility model.

Claims

1. A heterogeneous processing system based on FPGA and GPU, characterized in that, include: Multiple cascaded CPU servers interact with each other via a switch; each CPU server includes a CPU chip and multiple heterogeneous computing modules, the heterogeneous computing modules include FPGA chips and GPU chips, the FPGA chips and the GPU chips are interconnected via PCIe; The FPGA chip and the GPU chip interact with the CPU chip via PCIe respectively; The FPGA chip is configured to implement RDMA network card functionality.

2. The heterogeneous processing system based on FPGA and GPU according to claim 1, characterized in that, The FPGA chip and the GPU chip are interconnected via a direct PCIe_0 connection.

3. The heterogeneous processing system based on FPGA and GPU according to claim 2, characterized in that, The FPGA chip supports multiple 100Gbps network interfaces and is connected to Ethernet.

4. The heterogeneous processing system based on FPGA and GPU according to claim 3, characterized in that, The GPU chip accesses remote nodes through the FPGA chip and the Ethernet.

5. A heterogeneous processing system based on FPGA and GPU according to claim 4, characterized in that, The FPGA chip transmits the access request of the remote node to the GPU chip via PCIe, thereby enabling the interaction between the GPU chip and the remote node.

6. The heterogeneous processing system based on FPGA and GPU according to claim 5, characterized in that, The CPU server further includes a storage module for storing data exchanged between the GPU chip and the remote node.

7. A heterogeneous processing system based on FPGA and GPU according to claim 3, characterized in that, The FPGA chip and the CPU chip interact with each other via a PCIe_1 interface.

8. A heterogeneous processing system based on FPGA and GPU according to claim 4, characterized in that, The CPU chip interacts with the remote node through the FPGA chip.

9. A heterogeneous processing system based on FPGA and GPU according to claim 1, characterized in that, The GPU chip is used for data processing and computation, and interacts with the CPU chip through PCIe.

10. A heterogeneous processing system based on FPGA and GPU according to claim 1, characterized in that, The heterogeneous processing system further includes a power supply unit, which provides power to the CPU server.