Universally distributed cross-connection architecture design for improving throughput and computing power of big data artificial intelligence to maximum extent

By designing connected FPGA architectures in the field of big data artificial intelligence and using cross switches to achieve fast communication between FPGAs, the problem of modern big data artificial intelligence's demand for high computing power is solved, and high throughput and low latency computing effects are achieved.

CN120068761APending Publication Date: 2025-05-30SHANGHAI ELEPHANT TENSOR NANOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311619393.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The huge demand for computing power in the field of modern big data artificial intelligence is difficult to effectively meet the computing needs of high throughput and low latency.

Method used

By designing a scattered connection architecture, a large number of FPGAs are used for computing, and the fast and flexible communication between FPGAs is achieved through cross-switches to meet the needs of high computing power.

Benefits of technology

It realizes high-speed, low-power consumption, high throughput, and low-latency computing effects, and can effectively handle application tasks such as artificial intelligence, big data, FFT, and image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068761A_ABST
    Figure CN120068761A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a pervasive cross-connection architecture design capable of improving the throughput and computing power of big data artificial intelligence to the greatest extent. The AI processor architecture is a highly integrated AI processor architecture realized in the FPGA, is an architecture which is highly pipelined and can generate high-speed operation due to parallel operation of a plurality of threads, and can process big data more quickly compared with a common solution. The AI processor architecture comprises a system with a plurality of RISC-V cores, is used for monitoring loading, selection and operation of FPGA configuration and other tasks, and supports a plurality of independent concurrent operation memory access channels, a plurality of independent concurrent operation I / O channels and a plurality of independent concurrent operation 100GigE. Better performance is achieved by connecting local processing nodes (ACUs) with crossbar switches, and better inter-board I / O performance is achieved by using crossbar switches on multiple boards in a chassis. Cross-links exist in multiple high-speed serial input non-blocking 16 x 16 (or some other size) crossbars implemented on FPGA-specific hardware, which may be connected together as larger crossbars. The crossbar switch provides a large number of data paths, can efficiently process application tasks such as artificial intelligence, big data, FFT, image recognition and the like, and uses the crossbar switch to realize data distribution, load sharing, demand supply, load balancing and the like as a part of the whole task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data artificial intelligence computing power, and particularly to a cross-connect architecture design that maximally improves the throughput and computing power of big data artificial intelligence. Background Art

[0002] Since the technological breakthrough in 2012, this important technological need of big data artificial intelligence has suddenly "broken through the circle" and immediately become "extremely popular". It has caused the demand for AI computing power by humans to increase explosively by 300,000 times in just six years. It can be said that it doubles on average every 100 days. Therefore, application programs in the modern big data artificial intelligence field all require a large amount of computing power. Summary of the Invention

[0003] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides an architecture that meets the computing power requirements by means of widespread connection and use of a large number of FPGAs. Application tasks (such as image detection, FFT, big data problems, etc.) are processed by implementing multiple parallel pipeline execution cores in the FPGA, which have the advantages of high speed, low power consumption, high throughput, low latency, etc. The present invention combines the implementation of RISC-V cores in the FPGA and accelerator FPGAs for core computing processing.

[0004] Huge computing power requires more FPGA support. To effectively use a large number of FPGAs, a fast and flexible communication method and data transmission path between FPGAs have become the core part of the architecture design. In the present invention, the flexible communication method between FPGAs is realized by using a crossbar switch. The crossbar switch provides a simultaneously available way for all connections between FPGAs required by application tasks (without the need for time-division multiplexing like a bus structure). Moreover, this connection structure can be changed at any time to meet the connection requirements between FPGA resources on a single board or the connection requirements between 16 boards in a chassis.

[0005] The large number of data paths provided by the crossbar switch can efficiently process application tasks such as artificial intelligence, big data, FFT, image recognition, etc., including using the crossbar switch to achieve data distribution, load sharing, demand supply, load balancing, etc., as part of the overall task. The connection method will vary according to specific application tasks. For example, the connection of FPGAs for implementing the YOLO image detection task is very different from the connection for implementing FFT or big data problems. Brief Description of the Drawings

[0006] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments that conform to the present invention, and are used together with the specification to explain the principles of the present invention.

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0008] Figure 1 It is the architecture design layout and size diagram of the single board;

[0009] Figure 2 It is different perspectives of the architecture design of the single board;

[0010] Figure 3 It is the size of each content of the architecture design of the single board;

[0011] Figure 4 Sequencer schematic diagram;

[0012] Figure 5 It is the connection schematic diagram of cross switches 1 - 5;

[0013] Figure 6 It is the connection schematic diagram of cross switch 10;

[0014] Figure 7 It is the connection schematic diagram of cross switches 6 - 9;

[0015] Figure 8 It is the connection schematic diagram of the single board in the chassis and the other 15 boards;

[0016] Figure 9 It is the maximum connection number between any two FPGAs;

[0017] Figure 10 It is the non-blocking cross switch design diagram;

[0018] Figure 11 It is the cross switch logic schematic diagram;

[0019] Figure 12 It is the design layout diagram of the single-bit cross switch;

[0020] Figure 13 It is the SerDes parameter diagram;

[0021] Figure 14 It is the input / output schematic diagram of SerDes;

[0022] Figure 15 It is the connection schematic diagram of 16 pairs of FPGAs to GRISC;

[0023] Figure 16 It is the supply-demand balance algorithm schematic diagram;

[0024] Figure 17Schematic diagram of transmitting multiple logical data streams through one port;

[0025] Figure 18 is Figure 14 the example data stream in

[0026] Figure 19 is the timing diagram of multiplexing the data stream in Figure 14 by using time-division multiplexing;

[0027] Figure 20 Schematic diagram of the FPGA code module related to the present invention. Detailed implementation manners

[0028] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0029] It should be understood that the steps recited in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0030] There are 16 powerful Starline_FFT_SoCs on the single board. Each SoC has a global control RISC-V core, which is implemented by 2 FPGAs (GRISC) and 16 accelerator units (ACU). Each unit is composed of 2 FPGAs, one is a local RISC-V core (LRISC), and the other implements the accelerator function.

[0031] For applications that do not require LRISC, the LRISC FPGA can also be configured as an accelerator FPGA, providing the computing power of 32 accelerator FPGAs. After such configuration, a small state machine needs to be created in the configurable FPGA resources to perform a certain degree of local control over the I / O and reorder the steps in the accelerator. This state machine can be a state machine based on very long instruction word (VLIW) microcode created in the FPGA logic and memory resources. It is smaller than the RISC-V core, so more resources can be left for the accelerator logic. Since this part of the FPGA configuration is loaded when powered on and may be replaced by another configuration of GRISC later, the state machine can be customized according to the needs of the application rather than being used as a standard RISC-V core that uses a large amount of FPGA resources. According to the logical requirements, there may be several state machines based on Figure 4 the sequencer shown implemented on a single FPGA.

[0032] All these FPGAs on a single board are connected together through a 16x16 high-speed non-blocking crossbar also implemented in the FPGA. In addition, there are 2 additional crossbars that can connect to the remaining boards in a chassis (there are 16 boards in a chassis). The communication with each board can be board-to-board communication or broadcast to multiple boards simultaneously.

[0033] Each board also has 16 high-speed I / O ports, 8 of which are suitable for high-speed Ethernet (RJ45) and 8 are suitable for GBIC (fiber optic). Each chassis can provide a total of 256 I / O ports for data and connections to other chassis.

[0034] The single-board design layout and dimensions as described above are as Figure 1 , where SPF and ETH are the I / Os on the board. The Starline_FFT_SoC includes a global control RISC-V core (GRISC), a non-blocking crossbar, a local RISC-V core (LRISC), accelerator FPGAs. A pair of LRISC and accelerator FPGAs form an acceleration processing unit (ACU), and there are 16 ACUs in total on each SoC. Only the top-level global control RISC-V (GRISC) needs to be configured from the EPROM, and the crossbar and ACU units are configured by the GRISC. Each board is capable of starting the operating system when powered on and has code for configuring the remaining FPGAs (LRISK, accelerator, crossbar). The GRISC, LRISC, and accelerator FPGAs are connected through a large number of 16x16 serial input non-blocking crossbars implemented in the FPGA.

[0035] The connections in crossbars 1 to 5 are as Figure 5) Responsible for inputting data into 16 working ACUs and outputting the results. These connections can also be used to move data between the LRISC FPGA and the accelerator FPGA that do not belong to the same ACU. The connections in the crossbar 10 (such as Figure 6 ) are also responsible for inputting data into 16 working ACUs and outputting the results. These connections can also be used to move data between the LRISC FPGA and the accelerator FPGA that do not belong to the same ACU. Combining Figure 5 with Figure 6 , it can be seen that there are at least 2 channels from GRISC to each FPGA. The connections in the crossbars 6 - 9 (such as Figure 7 ) are responsible for moving data between the ACUs. Two of the switches are used to move data between the LRISCs, and two switches are used to move data between the accelerator FPGAs.

[0036] The connection of a single board in the chassis to the remaining 15 boards is as Figure 8 . In fact, this figure needs to be repeated 16 times in total, and each board in the chassis is connected in this way. There is a bidirectional lane on the backplane for connecting any two boards, with a total of 16 x 15 (240 in total) board-to-board connections. Each lane occupies 2 pairs of traces, so there are a total of 960 traces. Each FPGA on the single board has a path to other boards and can perform up to 240 bidirectional board-to-board tasks simultaneously.

[0037] The maximum number of connections from FPGA to FPGA is as Figure 9 . Any FPGA can be connected to other FPGAs through one (the numerical part in Figure 9 ) or 2 crossbars (the blank part in Figure 9 ). Most connections only require one crossbar. In most cases, if the ports on the FPGA are available (not yet used for another connection), there are multiple connection paths for a specific FPGA-to-FPGA connection. For two FPGAs that only require one crossbar to connect, Figure 9 lists the maximum number of connections between these two FPGAs. An empty connection number indicates that a connection between these two FPGAs requires 2 crossbars.

[0038] Implementing SerDes for achieving the maximum throughput of a non-blocking 16x16 switching matrix in an FPGA is a complex task that requires a large amount of Verilog code. The specific steps are as follows (please note that this is a general description, and the actual implementation details will depend on the FPGA development platform and development requirements):

[0039] Step 1: Define the SerDes interface. Define the interface of the SerDes module, including the number of data channels (e.g., number of channels = 16) and the data rate.

[0040] Step 2: Implement the Serializer. In the present invention, the Serializer is implemented in the hardware of the FPGA. Create a module for the Serializer that receives parallel data from 16 input channels and serializes it into a single high-speed serial data stream. Use appropriate clocking and data serialization techniques according to the FPGA development platform and data rate requirements.

[0041] Step 3: Implement the Deserializer. In the present invention, the Deserializer is implemented in the hardware of the FPGA. Create a module for the Deserializer that receives the high-speed serial data stream and converts it back into parallel data for 16 output channels.

[0042] Step 4: Design the switch matrix. Implement a 16x16 switch matrix using multiplexers or other switch logic that allows any input channel to be routed to any output channel. Ensure that the switch matrix is non-blocking, i.e., any input can be routed to any output without contention.

[0043] Step 5: Synchronize clock domains. If there are multiple clock domains in the system (e.g., between the SerDes and the switch matrix), clock domain crossing techniques need to be implemented to ensure correct data synchronization.

[0044] Step 6: Testbench. Develop a comprehensive testbench to verify the functionality of the SerDes and the switch matrix, test various input patterns and routing scenarios, and ensure that the system works properly.

[0045] Step 7: Simulation and synthesis. Use a Verilog simulator to simulate the design and verify its correctness. Perform synthesis for a specific FPGA platform to generate a programming bitstream.

[0046] Step 8: Constraints and timing analysis. Define timing constraints to ensure that all operations are correct and meet the timing requirements. Use the FPGA synthesis tool for timing analysis and optimize the design to achieve maximum throughput.

[0047] Step 9: Implementation and deployment. Program the FPGA with the generated bitstream. Test the system on the actual hardware to ensure that the performance and throughput requirements are met.

[0048] Step 10: Debugging and optimization. Debug any issues that may arise during the testing process and optimize the design.

[0049] The design of the non-blocking crossbar switch is as Figure 10, the configuration of the crossbar switch is controlled by the GRISC processor. The purpose of designing 16 end-to-end registers is to broadcast the same data to multiple ports. In addition, 2 ports connected to registers (n, n + 1) are designed so that when a new configuration is established, all contents can be loaded at once. There are 8 16x16 crossbar switches in the non-blocking crossbar switch, totaling 2048 single-bit crossbar switches. The reason for designing the crossbar switch as 8 single-bit crossbar switches instead of 1 eight-bit 16x16 crossbar switch is to better adapt to the FPGA routing resources.

[0050] The crossbar switch is logically a set of single-bit switches, as Figure 11 shown. For ease of use, the crossbar switch is implemented through a set of 16 input multiplexers (MUX), as Figure 12 shown, where the CLB is a configurable logic block. Enabling two paths through the crossbar switch can achieve bidirectional functions, such as from input 2 (In 2) to output 10 (Out 10) and from input 10 (In 10) to output 2 (Out 2).

[0051] The SerDes parameters are as Figure 13 shown. The FPGA has 8 pairs of SerDes. Each pair of SerDes shares a time generator. Each memory bank has a pair of SerDes, and they share the SerDes common clock block, totaling 16 SerDes channels (ports). Each SerDes includes an input serial-to-parallel part and an output parallel-to-serial part, as Figure 14 shown.

[0052] From the user's perspective, the control of the crossbar switch is completed by loading the configuration from the master RISC-V. We can assume that the master RISC-V retains a copy of the crossbar switch configuration and does not provide a read-back configuration.

[0053] The crossbar switch configuration is a bit string serially sent from the master RISC-V using a simple self-clocking protocol bidirectional interface, and any useful configuration among 256 crossbar switches can be set.

[0054] When all ports (connected to 16 devices) are in normal use, 16 crossbar switches will be turned off.

[0055] The command bit stream consists of 4 parts: start bit, 8-bit mode command, 128-bit configuration string, and parity bit:

[0056] Mode x01 is the normal mode for setting individual switches using a 128-bit string. 1 indicates that the switch is closed, and 0 indicates that the switch is open.

[0057] Each group has 8 bits, forming a set of input-to-output ports. The first group is for output 0, the second group is for output 1, and so on. In each group, the bit pattern selects the input encoded in 4 bits, and the other bits are used for the enable signal.

[0058] Typically, we want to establish a bidirectional connection between each pair of input and output ports, so a corresponding reverse switch needs to be set up for acknowledgment and flow control.

[0059] If for some reason only a unidirectional connection is needed, then half of the unused port SerDes connections can be used to connect to any other available port.

[0060] The crossbar switch can also be set to broadcast data to multiple ports simultaneously. For example, if more than one group of 8-bit outputs is set to the same bit pattern, the same data will be sent to each output.

[0061] Generally, the entire configuration is stored in the main control RISC-V. The crossbar switch is changed in the copy, and the entire bit string is sent to the crossbar switch. After several FPGA clock cycles, the entire configuration will be updated.

[0062] Other modes can be used to change the SerDes parameters without reloading the initial FPGA switch configuration and other pending functions.

[0063] Each of the 32 FPGAs (16 working node pairs) has a 6-wire interface to the main control RISC-V (host single), as Figure 15 shown. Some of these wires are used for command and status communication between the main control RISC-V and the 16 working pairs (both the RISC-V FPGA and the accelerator FPGA can communicate with the main control RISC-V).

[0064] This design can dynamically allocate tasks through a supply-demand balancing algorithm, as Figure 16 shown. When a node pair has completed all the assigned tasks, it can signal that it is idle. If another node pair is doing the same type of task at the same time but still has many steps unfinished, the master node can set the accelerator crossbar switch to connect the two accelerator FPGAs and send a command for data transfer to the two node pairs, so as to make full use of the idle node pair. The command steps in this case are:

[0065] In the first step, a node pair signals that it is idle or has completed the specified task.

[0066] In the second step, the master node checks (such as by querying) other busy node pairs to see if there is work suitable for the idle node pair.

[0067] Step 3: If available, the master node sets the crossbar switch as needed to connect the accelerator FPGAs of the idle node pair and the accelerator FPGAs of the busy node pair.

[0068] Step 4: The master node sends a command to the idle node pair, waits for incoming data from the crossbar switch, and processes it.

[0069] Step 5: The idle node pair acknowledges the command.

[0070] Step 6: The master node sends a command to the busy node pair to send data to the crossbar switch.

[0071] Step 7: The busy node pair acknowledges the command.

[0072] Step 8: The busy node pair sends data to the idle node pair.

[0073] Step 9: The idle node pair receives the data and starts processing.

[0074] Step 10: After processing is complete, the idle node pair signals the master node that it is idle again or has completed the task.

[0075] In some applications, we encounter situations where a large number of logical data streams need to be transferred from one FPGA to another through a single physical connection. A straightforward way to achieve this is to share the physical connection in the time domain, such as Figure 17 , which can be effectively achieved by using FPGA logic to combine the logical data streams into one data stream and split it into separate logical streams in the destination FPGA. Figure 17 The example data stream in Figure 18 has a header character in front of each data block indicating the logical data stream it belongs to. Figure 17 The demultiplexing (Demux) in Figure 17 and Figure 18 is a logical demultiplexing. In fact, the data stream can be applied to all FIFOs and transmitted to the correct FIFO by the control logic. For demonstration purposes, Figure 18 in Figure 19 we use three logical data streams A, B, and C. In actual applications, more can be set, and the data streams do not need to be in the strict order as shown in Figure 17

[0076] The FPGA code modules related to the present invention are as shown in Figure 20 , including a control module, a decoding module, and eight single-bit 16x16 crossbar switch modules. SerDes has hardware support as shown in Figure 14 , but the control of each SerDes requires a section of FPGA code.​Figure 20 The numbers beside the slashes are the number of signals for each module connected to a certain function. Most of the signals are data, excluding the signals from the decoding module. Figure 20 SerDes control lines are not shown.

Claims

1. The present invention designs a highly integrated AI processor architecture implemented in an FPGA.

2. The present invention designs a device that processes big data faster than ordinary solutions.

3. The present invention provides a design that enables high-speed operation due to being highly pipelined and capable of running multiple threads in parallel.

4. The present invention designs an efficient non-blocking crossbar implemented in an FPGA.

5. The crossbar according to claim 4 provides a large number of data paths and can efficiently process application tasks such as artificial intelligence, big data, FFT, image recognition, etc., including using the crossbar to achieve data distribution, load sharing, demand supply, load balancing, etc., as part of the overall task.

6. The present invention designs a system including multiple RISC-V cores for supervising the loading, selection, and operation of FPGA configurations and other tasks.

7. The present invention provides the ability to support multiple independent concurrent operation memory access channels.

8. The present invention provides the ability to support multiple independent concurrent operation I / O channels to achieve fast I / O.

9. The present invention provides the ability to support multiple independent concurrent operation 100GigE to achieve fast I / O.

10. The present invention provides a method to achieve better performance by connecting local processing nodes (ACUs) to the crossbar.

11. The present invention provides a method to achieve better inter-board I / O performance by using crossbars on multiple boards in a chassis.

12. The present invention provides a method to implement multiple 16 x 16 (or some other size) crossbars using FPGA-specific hardware, and these crossbars can be connected together to be used as a larger crossbar.

13. The present invention designs a high-speed serial input non-blocking 16x16 crossbar implemented in an FPGA.