Method for realizing large-scale systolic array by adopting FPGA (Field Programmable Gate Array) cluster

By building an FPGA cluster, using the collaborative work of multiple CPU servers and FPGA boards, the problem of resource limitation and data transmission bottlenecks of a single FPGA is solved, and efficient computing and flexible expansion of large-scale pulsating arrays are achieved, reducing the difficulty of system development and usage costs.

CN120086181APending Publication Date: 2025-06-03SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510229420.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

When implementing large-scale pulsating arrays, due to the resource limitations and data transmission bottlenecks of a single FPGA, it is difficult to expand the array scale and improve performance. In addition, multi-FPGA systems have flexibility and complexity problems in configuration and operation, which increases the development difficulty and usage cost.

Method used

The virtual two-dimensional structure is built using FPGA clusters, and the parallel computing of large-scale pulsating arrays is realized through the collaborative work of multiple CPU servers and FPGA boards. The specific steps include building an FPGA cluster, decomposing a large-scale pulsating array into multiple blocks and assigning it to an FPGA board for calculation, and obtaining the calculation results through the CPU server for further processing.

Benefits of technology

It improves the computing speed of large-scale pulsating arrays, can complete complex computing tasks in a shorter time, adapt to different computing needs, improves the system's resource adaptability and scalability, ensures the accuracy and reliability of data transmission, and reduces the system's usage cost and development difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086181A_ABST
    Figure CN120086181A_ABST
Patent Text Reader

Abstract

The invention discloses a method for realizing a large-scale systolic array by adopting an FPGA (Field Programmable Gate Array) cluster, which relates to the technical field of systolic arrays and comprises the following steps: S1, constructing a virtual two-dimensional FPGA cluster, communicating different CPU (Central Processing Unit) servers of the FPGA cluster through a server interaction machine network, and interconnecting and communicating different FPGA board cards through a switch; s2, a large-scale systolic array to be tested is decomposed into a plurality of blocks, the multiple systolic array blocks are correspondingly distributed to the multiple FPGA board cards of the FPGA cluster, and the FPGA board cards calculate the distributed systolic array blocks; and S3, after all the FPGA board cards complete the calculation task of the systolic array block, obtaining a calculation result through the CPU server, and performing further processing. According to the method, the calculation speed of the large-scale systolic array can be increased, complex calculation tasks can be completed in a shorter time, different large-scale systolic array calculation requirements are met, and the resource adaptability and expandability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of systolic arrays, in particular to a method for realizing a large-scale systolic array by using an FPGA cluster. Background Art

[0002] Systolic arrays are an extremely important large-scale parallel computing structure, which is composed of many processing elements (PEs) in an orderly manner. These processing elements work together, just like the pulsation of a heart, to process and calculate data in a highly regular and coordinated manner. Each processing element has the ability to independently perform simple operations, such as efficiently completing basic arithmetic operations such as multiplication and addition. They do not operate in isolation, but are closely connected through a carefully designed communication network, thus building a powerful parallel computing system.

[0003] In the field of high-performance computing, systolic arrays have shown outstanding advantages. They can decompose complex computing tasks into multiple parallel subtasks and assign them to various processing units for synchronous execution, greatly improving computing efficiency and significantly shortening the processing time of large-scale scientific computing, password cracking and other tasks that require extremely high computing speed. In digital signal processing, systolic arrays can be used to process real-time audio and video signals, perform fast filtering and transformation operations on signals, and ensure the efficiency and accuracy of signal processing. Especially in recent years, with the rapid development of artificial intelligence technology, systolic arrays play a key role in artificial intelligence computing. In the training and reasoning process of neural networks, it can accelerate a large number of matrix operations in parallel, greatly improving the speed of model training and the real-time performance of reasoning, and providing powerful computing support for applications such as deep learning and machine learning.

[0004] FPGA, or Field-Programmable Gate Array, is a highly flexible integrated chip. Its core feature is reconfigurability, which means that users can customize the logic circuit inside the chip through programming according to specific application requirements. In modern digital circuit design, FPGA is widely used. From the construction of simple digital logic circuits to the realization of complex digital signal processing systems and communication systems, FPGA can quickly respond to different design requirements with its reconfigurable characteristics, greatly shortening the product development cycle and reducing development costs. For example, in communication equipment, FPGA can be configured according to different communication protocols to realize data modulation, demodulation, encoding and decoding and other functions; in the field of image processing, FPGA can be flexibly configured to complete a series of operations such as image acquisition, preprocessing, and feature extraction.

[0005] Currently, common methods for implementing systolic arrays mostly rely on a single FPGA. In this implementation, researchers mainly improve system performance by optimizing the data transmission path and increasing the data reuse rate. For example, by carefully designing the data caching mechanism to reduce the number of data transmissions between different processing units, enabling data to flow efficiently between processing units, thereby improving the computing efficiency to a certain extent. However, a single FPGA itself has significant limitations. The internal resources, such as the number of logic units and the capacity of storage units, are relatively limited, and there are also bottlenecks in the data transmission bandwidth, which severely restricts the scale expansion and performance improvement of the constructed systolic array. Moreover, this implementation scheme based on a single FPGA performs poorly in terms of generality and is often only applicable to specific types of computing tasks, making it difficult to flexibly adapt to diverse application scenarios; there are also deficiencies in scalability, and it is difficult to expand in a simple way when further improving computing power is required.

[0006] Although multi-FPGA systems can, to a certain extent, break through the limitations of a single FPGA, existing multi-FPGA systems also have many problems. Some systems lack flexibility during the configuration process, and it is difficult for users to conveniently adjust system parameters and functions according to actual needs; there are also some systems with complex operations that require professionals to spend a lot of time and effort on debugging and maintenance, which undoubtedly increases the usage cost and development difficulty of the system and limits the popularization and application of multi-FPGA systems in practice. Summary of the Invention

[0007] In view of the current needs and deficiencies in the development of technology, the present invention provides a method for implementing a large-scale systolic array using an FPGA cluster.

[0008] The method for implementing a large-scale systolic array using an FPGA cluster according to the present invention adopts the following technical solutions to solve the above technical problems:

[0009] A method for implementing a large-scale systolic array using an FPGA cluster, which includes the following steps:

[0010] S1. Construct a virtual two-dimensional FPGA cluster. Different CPU servers in the FPGA cluster communicate through a server interaction machine network, and different FPGA boards are interconnected and communicate through a switch;

[0011] S2. Decompose the large-scale systolic array to be measured into multiple blocks, and correspondingly allocate the multiple systolic array blocks to multiple FPGA boards in the FPGA cluster. The FPGA boards calculate the allocated systolic array blocks;

[0012] S3. After all FPGA boards complete the computing tasks of the systolic array blocks, the CPU server is used to obtain the computing results and perform further processing.

[0013] Optionally, the FPGA cluster constructed in step S1 includes multiple CPU servers, and two FPGA boards are installed on each CPU server. The two FPGA boards communicate with the CPU server on which they are installed through the PCIE communication module.

[0014] Each FPGA board is configured with two optical communication interfaces. The optical communication interfaces of the FPGA boards are connected to the corresponding ports of the 100G switch through optical fibers. The switch is responsible for forwarding and exchanging the data sent by each FPGA board, enabling efficient data interaction between different FPGA boards.

[0015] Further optionally, the internal logic implementation of the involved FPGA board involves the following structure:

[0016] Physical ports for implementing multiple virtual ports.

[0017] The PCIE communication module is integrated into the FPGA board and is used to achieve high-speed, two-way data communication between the FPGA board and the CPU in the CPU server where it is located, providing a communication bridge for the collaborative work of the FPGA board and the CPU.

[0018] The MIG module is used to control external DDR reading and writing.

[0019] The high-speed reading and writing DMA module.

[0020] The crossbar switch module, as a communication and routing control component, is used to route different virtual ports and systolic array modules.

[0021] The systolic array module is used to achieve algorithm acceleration. Among them, the MUX module and the DMUX module are the interfaces between the systolic array module and the virtual network. The MUX module is used to achieve time-division multiplexing of multiple data streams, receive data from multiple virtual ports, and multiplex the received data onto a transmission line according to a preset rule for transmission to the systolic array module for processing. The DMUX module is used to demultiplex the multiplexed data output after being processed by the systolic array module, and distribute the data to different virtual ports according to the original source or preset rule for subsequent processing or output of the data.

[0022] The encoding module connected to the MUX module is used to encapsulate the Ethernet header information and valid data into an Ethernet frame for transmission according to a set format. Among them, the valid data includes the source ID and destination ID of the virtual port, the source address and destination address, and the channel number.

[0023] The decoding module connected to the DMUX module is used to receive the Ethernet frame data from the DMUX module, separate the Ethernet header information and the valid data according to the set format, and extract the virtual port source ID and destination ID, source address and destination address, and channel number from the valid data according to the corresponding rules when the encoding module encapsulates, and perform further processing or pass it to the subsequent system modules to achieve the correct use and processing of the data;

[0024] The eth_ctrl module contains a MAC table and is used to write to its MAC table through the CPU server to configure the virtual network working topology, where the working topology describes the connection relationship and data transmission path between various components in the virtual network.

[0025] Preferably, the involved MUX module realizes the time-division multiplexing of multiple data streams through different data processing modes:

[0026] a) In the round-robin mode, the MUX module sequentially selects data from each virtual port in a cyclic order for multiplexing;

[0027] b) In the preset sequential mode, the MUX module processes the data of the virtual ports according to the preset priority order or specified order;

[0028] c) When a certain virtual port has a large data volume sending requirement, the MUX module preferentially multiplexes the data of the high-priority port according to the priorities of each virtual port.

[0029] Preferably, the involved DMUX module demultiplexes the multiplexed data output after being processed by the systolic array module. In this process, the DMUX module parses and distributes the input multiplexed data according to the identification information carried by the data during the multiplexing process or the preset demultiplexing rules, ensuring that the data can accurately return to the corresponding virtual port, realizing the complete docking with the function of the MUX module, and ensuring the accuracy and integrity of the entire data transmission and processing process.

[0030] Preferably, the MAC table contained in the involved eth_ctrl module stores the following two items of information:

[0031] One is the ID of the message virtual port, which is used to identify which virtual port the data message comes from or is sent to;

[0032] The other is the destination address, that is, the address information of the target location where the data is to be sent;

[0033] Through the recording of these two pieces of information, the MAC table can manage and control the flow direction of the data message.

[0034] Preferably, FIFOs exceeding the requirements are set in the input and output channels of each FPGA board. The input data is queued through the FIFOs, enabling FPGA boards with different processing speeds to maintain consistency in the data input rhythm, thereby synchronizing the array calculations of all FPGA boards.

[0035] After all FPGA boards complete the calculation tasks of the systolic array blocks they are responsible for, the FPGA boards establish a connection with the CPU server through the PCIE communication module.

[0036] Each FPGA board transmits the calculation results to the CPU server through the FIFO in the output channel according to the set data format and communication protocol. During the transmission process, the FIFO in the output channel ensures that the data can be stably and orderly sent to the CPU server.

[0037] After the CPU server obtains the calculation results, it performs further processing with its powerful general computing capabilities and rich software resources.

[0038] Optionally, in the interaction with the FPGA board, the CPU server uses its own processing capabilities and communication interfaces to send configuration instructions to the eth_ctrl module and crossbar module of the FPGA board, so as to set up a virtual network topology according to the pre-planned network structure and data transmission requirements, enabling the data to be accurately transmitted from the source node to the target node according to the planning of the virtual network topology.

[0039] Optionally, in step S2, the large-scale systolic array to be tested is decomposed into M×N independent and parallel-running blocks according to the set rules, and the decomposed M×N systolic array blocks are centrally placed in the memory of a main CPU server.

[0040] Based on the network connection and load balancing strategy between CPU servers, the main CPU server distributes the M×N systolic array blocks in its own memory to the memories of the CPU servers connected to each FPGA board in the FPGA cluster.

[0041] Each CPU server connected to the FPGA board directly transfers the systolic array block in its own memory to the memory of the connected FPGA board through the PCIE XDMA technology.

[0042] The beneficial effects of a method for implementing a large-scale systolic array using an FPGA cluster according to the present invention compared with the prior art are:

[0043] 1. The present invention utilizes parallel processing of multiple FPGA boards in the FPGA cluster, greatly improving the computing speed of the large-scale systolic array and enabling complex computing tasks to be completed in a shorter time. It can flexibly adjust the configuration of the FPGA cluster according to different computing tasks and data scales, such as increasing or decreasing the number of CPU servers, FPGA boards, and adjusting the virtual network topology, etc., to adapt to different large-scale systolic array computing requirements and improve the resource adaptability and scalability of the system.

[0044] 2. The present invention distributes large-scale systolic array blocks to the memories of different CPU servers based on load balancing and then to each FPGA board, which can make full use of the computing resources of each server and FPGA board in the FPGA cluster, avoiding the situation where some devices are overloaded while some are idle, and enabling the efficient utilization of the resources of the entire system.

[0045] 3. In the FPGA cluster of the present invention, different CPU servers communicate through the server interaction machine network, and the FPGA boards of different servers are interconnected through a switch. This architecture provides a high-speed and stable communication channel, ensuring the rapid transmission of data between devices, reducing data transmission latency, and facilitating the improvement of the collaborative work efficiency of the entire system. By configuring the eth_ctrl module and cross-switch module to establish a virtual network topology and data transmission path, data can be accurately transmitted from the source node to the target node, ensuring the accuracy and reliability of data transmission and reducing the error rate and packet loss rate during data transmission.

[0046] 4. The present invention realizes pipelined data processing by setting FIFOs in the FPGA input and output channels, enabling data to continuously flow and be processed between FPGA boards, further enhancing the overall computing efficiency, avoiding data processing interruptions and waiting, and improving resource utilization. The use of FIFOs not only realizes the synchronization of computing but also enhances the stability of the system. Even if the data processing times of each FPGA board are different, the FIFOs can queue and buffer the data to ensure the consistency and coherence of data processing and avoid system errors and crashes caused by data asynchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] FIG. Figure 1 is a schematic diagram of the FPGA cluster according to the embodiment of the present invention;

[0048] FIG. Figure 2 is a structural diagram of the internal logic implementation of the FPGA board according to the embodiment of the present invention;

[0049] FIG. Figure 3 is an example diagram of the FPGA cluster of the present invention implementing a large-scale systolic array. Detailed implementation manners

[0050] To make the technical solutions, the technical problems to be solved, and the technical effects of the present invention clearer and more understandable, the following describes the technical solutions of the present invention clearly and completely in conjunction with specific embodiments.

[0051] Embodiment 1:

[0052] This embodiment proposes a method for implementing a large-scale systolic array using an FPGA cluster, which includes the following steps:

[0053] S1. Construct a virtual two-dimensional FPGA cluster. Different CPU servers in the FPGA cluster communicate through a server interaction machine network, and different FPGA boards are interconnected and communicated through a switch.

[0054] Combined with the attached Figure 1 , the FPGA cluster constructed in this step includes multiple CPU servers. Two FPGA boards are installed on each CPU server. The two FPGA boards communicate with the CPU server on which they are installed through a PCIE communication module. Each FPGA board is configured with two optical communication interfaces. The optical communication interfaces of the FPGA board are connected to the corresponding ports of a 100G switch through optical fibers. The switch is responsible for forwarding and exchanging the data sent by each FPGA board, enabling efficient data interaction between different FPGA boards.

[0055] Combined with the attached Figure 2 , the internal logic implementation of the FPGA board mentioned in this step involves the following structures:

[0056] Physical ports, used to implement multiple virtual ports;

[0057] PCIE communication module, integrated on the FPGA board, used to implement high-speed, two-way data communication between the FPGA board and the CPU in the CPU server where it is located, providing a communication bridge for the collaborative work of the FPGA board and the CPU;

[0058] MIG module, used to control external DDR reading and writing;

[0059] DMA module for high-speed reading and writing;

[0060] Crossbar switch module, as a communication and routing control component, used to route different virtual ports and systolic array modules;

[0061] The systolic array module is used to accelerate algorithms. Among them, the MUX module and the DMUX module are the interfaces between the systolic array module and the virtual network. The MUX module is used to implement time-division multiplexing of multiple data streams, receive data from multiple virtual ports, and multiplex the received data onto a transmission line according to a preset rule for transmission to the systolic array module for processing. The DMUX module is used to demultiplex the multiplexed data output after being processed by the systolic array module, and distribute the data to different virtual ports according to the original source or preset rule for subsequent processing or output of the data;

[0062] The encoding module connected to the MUX module is used to encapsulate Ethernet header information and valid data into Ethernet frames for transmission according to a set format. Among them, the valid data includes the virtual port source ID (used to identify the virtual port from which the data is sent), the destination ID (identifying the virtual port to which the data is to arrive), the source address (the network address of the data sender), the destination address (the network address of the data receiver), and the channel number (used to identify the channel used for data transmission);

[0063] The decoding module connected to the DMUX module is used to receive Ethernet frame data from the DMUX module, separate the Ethernet header information and valid data according to a set format, and extract the virtual port source ID and destination ID, source address and destination address, and channel number from the valid data according to the corresponding rule when the encoding module encapsulates, for further processing or passing to subsequent system modules to achieve the correct use and processing of the data;

[0064] The eth_ctrl module contains a MAC table and is used to perform a write operation on its MAC table through the CPU server to configure the working topology of the virtual network. Among them, the working topology describes the connection relationship and data transmission path between various components in the virtual network. It should be added that the eth_ctrl module usually refers to the Ethernet Control module, which is a functional module used to control and manage Ethernet data transmission and related operations in a system involving Ethernet communication.

[0065] In the interaction with the FPGA board, the CPU server uses its own processing power and communication interface to send configuration instructions to the eth_ctrl module and the crossbar module of the FPGA board to set the virtual network topology according to the pre-planned network structure and data transmission requirements, so that the data can be accurately transmitted from the source node to the target node according to the planning of the virtual network topology.

[0066] It should be added that the involved MUX module realizes time-division multiplexing of multiple data streams through different data processing modes:

[0067] a) In the round-robin mode, the MUX module sequentially selects data from each virtual port in a cyclic order for multiplexing;

[0068] b) When in the preset sequence mode, the MUX module processes the data of the virtual ports according to the preset priority order or specified order;

[0069] c) When a certain virtual port has a large data volume sending requirement, the MUX module preferentially multiplexes the data of the high-priority port according to the priorities of each virtual port.

[0070] The involved DMUX module demultiplexes the multiplexed data output after being processed by the systolic array module. In this process, the DMUX module parses and distributes the input multiplexed data according to the identification information carried by the data during the multiplexing process or the preset demultiplexing rules, ensuring that the data can accurately return to the corresponding virtual port, realizing the complete docking with the function of the MUX module, and ensuring the accuracy and integrity of the entire data transmission and processing process.

[0071] The involved eth_ctrl module contains a MAC table (Media Access Control Table) that stores the following two items of information:

[0072] One is the ID of the message virtual port, which is used to identify which virtual port the data message comes from or is sent to;

[0073] The other is the destination address, that is, the address information of the target location where the data is to be sent;

[0074] By recording these two pieces of information, the MAC table can manage and control the flow direction of the data message.

[0075] S2. Decompose the to-be-tested large-scale systolic array into M×N independent and parallel-running blocks according to the set rules, and centrally place the decomposed M×N systolic array blocks in the memory of a main CPU server; based on the network connection and load balancing strategy between CPU servers, the main CPU server distributes the M×N systolic array blocks in its own memory to the memories of the CPU servers connected to each FPGA board in the FPGA cluster; each CPU server connected to the FPGA board directly transfers the systolic array block in its own memory to the memory of the connected FPGA board through the PCIE XDMA technology, and the FPGA board calculates the allocated systolic array block. Refer to the appendix Figure 3 。

[0076] It should be added that FIFOs exceeding the requirements are set in the input and output channels of each FPGA board. The input data is queued through the FIFOs, enabling FPGA boards with different processing speeds to maintain consistency in the data input rhythm, thereby synchronizing the array calculations of all FPGA boards;

[0077] After all FPGA boards complete the calculation tasks of the systolic array blocks they are responsible for, the FPGA boards establish a connection with the CPU server through the PCIE communication module;

[0078] Each FPGA board transmits the calculation results to the CPU server through the FIFO in the output channel according to the set data format and communication protocol; during the transmission process, the FIFO in the output channel ensures that the data can be stably and orderly sent to the CPU server;

[0079] After the CPU server obtains the calculation results, it performs further processing with its powerful general computing capabilities and rich software resources.

[0080] S3. After all FPGA boards complete the calculation tasks of the systolic array blocks, the CPU server is used to obtain the calculation results and perform further processing.

[0081] In summary, by using the method for implementing a large-scale systolic array with an FPGA cluster of the present invention, the calculation speed of the large-scale systolic array can be increased, complex calculation tasks can be completed in a shorter time, different large-scale systolic array calculation requirements can be adapted, and the resource adaptability and scalability of the system can be improved.

[0082] The above application of specific examples has elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, those skilled in the art of the present technology, without departing from the principle of the present invention, any improvements and modifications made to the present invention shall fall within the scope of patent protection of the present invention.

Claims

1. A method for implementing a large-scale systolic array using an FPGA cluster, characterized in that: The steps include: S1. Build a virtual two-dimensional FPGA cluster. Different CPU servers in the FPGA cluster communicate with each other through a server interactive network, and different FPGA boards communicate with each other through switches. S2, decomposing the large-scale systolic array to be tested into multiple blocks, and assigning the multiple systolic array blocks to multiple FPGA boards of the FPGA cluster, and the FPGA boards perform calculations on the assigned systolic array blocks; S3. After all FPGA boards complete the calculation tasks of the systolic array blocks, the calculation results are obtained through the CPU server and further processed.

2. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 1, characterized in that: The FPGA cluster constructed in step S1 includes multiple CPU servers, each of which is equipped with two FPGA boards, and the two FPGA boards communicate with the CPU servers on which they are installed through a PCIE communication module; Each FPGA board is equipped with two optical communication interfaces. The optical communication interfaces of the FPGA board are connected to the corresponding ports of the 100G switch through optical fibers. The switch is responsible for forwarding and exchanging the data sent by each FPGA board, so that efficient data interaction can be carried out between different FPGA boards.

3. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 2, characterized in that: The internal logic implementation of the FPGA board involves the following structures: Physical ports, used to implement multiple virtual ports; The PCIE communication module is integrated in the FPGA board and is used to realize high-speed, bidirectional data communication between the FPGA board and the CPU in the CPU server where it is located, providing a communication bridge for the collaborative work of the FPGA board and the CPU; MIG module, used to control external DDR reading and writing; DMA module for high-speed reading and writing; A crossbar switch module, as a communication and routing control component, is used to route different virtual ports and systolic array modules; The systolic array module is used to realize algorithm acceleration, wherein the MUX module and the DMUX module are interfaces between the systolic array module and the virtual network, the MUX module is used to realize time division multiplexing of multiple data streams, receive data from multiple virtual ports, and multiplex the received data to a transmission line according to preset rules, so as to transmit the received data to the systolic array module for processing, and the DMUX module is used to demultiplex the multiplexed data output after processing from the systolic array module, and distribute the data to different virtual ports according to the original source or preset rules, so as to perform subsequent processing or output of the data; An encoding module connected to the MUX module is used to encapsulate the Ethernet header information and valid data into an Ethernet frame for transmission according to a set format, wherein the valid data includes a virtual port source ID and a destination ID, a source address and a destination address, and a channel number; The decoding module connected to the DMUX module is used to receive the Ethernet frame data from the DMUX module, separate the Ethernet header information and valid data according to the set format, and extract the virtual port source ID and destination ID, source address and destination address, and channel number of the valid data according to the rules corresponding to the encapsulation of the encoding module, and further process or pass it to the subsequent system module to realize the correct use and processing of the data; The eth_ctrl module includes a MAC table, which is used to write to the MAC table through the CPU server to configure the virtual network working topology, where the working topology describes the connection relationship and data transmission path between various components in the virtual network.

4. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 3, characterized in that: The MUX module implements time division multiplexing of multiple data streams through different data processing modes: a) In round-robin mode, the MUX module selects data from each virtual port in turn for multiplexing in a cyclic order; b) In the preset sequential mode, the MUX module processes the data of the virtual port according to the pre-set priority order or specified order; c) When a virtual port has a large data volume transmission requirement, the MUX module prioritizes the data of the high-priority port based on the priority of each virtual port.

5. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 3, characterized in that: The DMUX module demultiplexes the multiplexed data output from the systolic array module after processing. During this process, the DMUX module parses and distributes the input multiplexed data according to the identification information carried by the data during the multiplexing process or the pre-set demultiplexing rules, ensuring that the data can be accurately returned to the corresponding virtual port, achieving complete docking with the MUX module function, and ensuring the accuracy and integrity of the entire data transmission and processing process.

6. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 3, characterized in that: The MAC table contained in the eth_ctrl module stores the following two pieces of information: The first is the ID of the message virtual port, which is used to identify the virtual port to which the data message comes from or is sent; The second is the destination address, which is the address information of the target location where the data is to be sent; By recording these two pieces of information, the MAC table can manage and control the flow of data packets.

7. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 3, characterized in that: Set up FIFOs that exceed the demand in the input and output channels of each FPGA board. Use FIFOs to queue the input data so that FPGA boards with different processing speeds can maintain consistent data input rhythms, thereby synchronizing the array calculations of all FPGA boards. After all FPGA boards complete the computing tasks of the systolic array blocks they are responsible for, the FPGA boards establish a connection with the CPU server through the PCIE communication module; Each FPGA board transmits the calculation results to the CPU server through the FIFO of the output channel according to the set data format and communication protocol; During the transmission process, the FIFO of the output channel ensures that the data can be sent to the CPU server stably and orderly; After the CPU server obtains the calculation results, it performs further processing with its powerful general computing capabilities and rich software resources.

8. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 3, characterized in that: In the interaction with the FPGA board, the CPU server uses its own processing power and communication interface to send configuration instructions to the eth_ctrl module and cross switch module of the FPGA board to set the virtual network topology according to the pre-planned network structure and data transmission requirements, so that data can be accurately transmitted from the source node to the target node according to the planning of the virtual network topology.

9. The method of implementing a large-scale systolic array using an FPGA cluster according to claim 1, characterized in that: Execute step S2, decompose the large-scale systolic array to be tested into M×N blocks that are independent and run in parallel according to a set rule, and centrally place the decomposed M×N systolic array blocks in the memory of a main CPU server; Based on the network connection and load balancing strategy between CPU servers, the main CPU server distributes the M×N systolic array blocks in its own memory to the memory of the CPU servers connected to each FPGA board in the FPGA cluster; Each CPU server connected to the FPGA board directly transfers the systolic array blocks in its own memory to the connected FPGA board memory through PCIE XDMA technology.