Collaborative acceleration method and device based on distributed FPGA

By utilizing a distributed FPGA architecture and the collaborative work of the master FPGA chip and multiple slave FPGA chips, the problem of insufficient computing resources and scheduling difficulties in complex scenarios under a single FPGA solution is solved, achieving efficient task processing and resource utilization.

CN120909805BActive Publication Date: 2025-12-16XIAN LINGKONG ELECTRONICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511445849.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-12-16
Estimated Expiration
2045-10-11

AI Technical Summary

Technical Problem

Existing single FPGA solutions suffer from insufficient computing resources, severe resource contention, and difficulty in coordinating and scheduling multiple models when dealing with complex scenarios, making it difficult to meet the needs of modern UAV systems for large-scale data processing and complex algorithms.

Method used

The system adopts a distributed FPGA architecture, with the first FPGA chip serving as the main control unit, responsible for task parsing and subtask allocation. Multiple second FPGA chips perform parallel computing, and data interaction and synchronization are achieved through a high-speed communication interface. Combined with a clock synchronization module, a JTAG download module, and a power management module, the system ensures stability.

Benefits of technology

It improves the resource utilization and computing efficiency of the computing platform, achieves balanced distribution of task load and reliability of the computing platform, solves the performance bottleneck of existing solutions, and meets the needs of complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909805B_ABST
    Figure CN120909805B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on distributed FPGA's cooperative acceleration method and device, it is related to model acceleration technical field, the method includes: one first FPGA chip is connected with multiple second FPGA chips and constructs computing platform;Clock synchronization is carried out to first FPGA chip and multiple second FPGA chips, based on first FPGA chip receives and analyzes original task package from external system, original task package is decomposed into multiple subtasks;The load state of each second FPGA chip is monitored, and subtask is distributed based on load state;Subtask is processed in parallel by each second FPGA chip, and processing result is fed back to first FPGA chip;Based on first FPGA chip, each processing result is checked, integrated, and task execution result is obtained, and task execution result is fed back to external system.The technical problem that the acceleration effect and performance of existing model acceleration scheme exist bottleneck, it is difficult to meet the demand of complex scene is solved.Can improve the reliability, flexibility and computing efficiency of computing platform.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of model acceleration, and particularly relates to a cooperative acceleration method and device based on distributed FPGA. BACKGROUND

[0002] With the rapid development of aerial models, especially modern unmanned aerial vehicle systems, field programmable gate array (FPGA) chips gradually become the mainstream computing platform in the field of modern unmanned aerial vehicle system acceleration due to their hardware reconfigurability, low latency, high energy efficiency ratio and other characteristics. Modern unmanned aerial vehicle systems integrate complex sensor networks, high-precision navigation algorithms, real-time image / video processing, multi-channel communication systems and diversified control mechanisms. And with the continuous improvement of the technical complexity and scene diversification of modern unmanned aerial vehicle systems, more stringent requirements are put forward for the computing platform in terms of computing power, real-time response and system reliability.

[0003] The existing computing platform uses a single FPGA scheme to realize model acceleration. The single FPGA scheme can develop computing units customized for specific algorithms (such as neural network inference, signal processing, etc.) through hardware logic reconfiguration characteristics, fully utilize the advantages of parallel computing to improve data processing efficiency.

[0004] However, the single FPGA scheme has significant limitations when dealing with complex scenarios: first, for large models with high computational complexity (such as deep neural networks) or multiple independent small models running in parallel, the logic units, memory bandwidth and I / O (input / output) interfaces of a single FPGA are difficult to meet the demand of multiple models / large models for computing resources and data throughput, resulting in resource competition or processing bottlenecks; second, the architecture of a single FPGA is closed, making it difficult to coordinate multiple models, and it is difficult to achieve efficient pipeline or distributed processing of cross-model tasks, which seriously limits the overall computing efficiency; third, in the face of the comprehensive demand of large-scale data processing (such as high-resolution image streams, multi-sensor fusion data) and complex algorithms (such as multi-modal perception, dynamic path planning) in the field of aerial models, the resource upper limit and scalability of a single FPGA are difficult to meet the demand.

[0005] To solve the above technical problems, the embodiment of the present application provides a kind of based on distributed FPGA's collaborative acceleration method and device, specifically, the method is by first FPGA chip as master unit, responsible for task analysis, subtask division and result integration, while real-time monitoring the load state of each second FPGA chip, dynamically adjusts task allocation strategy;Each second FPGA chip is as a slave computing unit, focuses on the parallel execution of subtask, and realizes data interaction and synchronization through high-speed communication interface.In addition, the present application also introduces clock synchronization module, JTAG download module and power management module and other auxiliary units, ensure the clock consistency of multi-FPGA system, program updateability and power supply stability. SUMMARY

[0006] The embodiment of the present application provides a kind of based on distributed FPGA's collaborative acceleration method and device, the technical problem that the acceleration effect and performance of existing model acceleration scheme exist bottleneck, it is difficult to meet the demand of complex scene is solved.

[0007] Firstly, the embodiment of the present application provides a kind of based on distributed FPGA's collaborative acceleration method, comprising: a first FPGA chip and multiple second FPGA chips are connected by high-speed communication link to build computing platform;Clock synchronization is carried out to the first FPGA chip and multiple second FPGA chips, and task processing step is executed based on the computing platform to realize task acceleration;Wherein, the task processing step includes: based on the first FPGA chip receives and analyzes the original task package from external system, the original task package is decomposed into multiple subtasks;The load state of each second FPGA chip is monitored by the first FPGA chip, and multiple second FPGA chips are dynamically allocated the subtasks based on the load state;Subtask is processed in parallel by each second FPGA chip, and processing result is fed back to first FPGA chip;Based on the first FPGA chip, the processing result of each second FPGA chip is checked, integrated, and task execution result is obtained, and task execution result is fed back to external system.

[0008] With reference to the first aspect, in a possible implementation manner, the clock synchronization of the first FPGA chip and the plurality of second FPGA chips comprises: metal wiring inside the first FPGA chip and the second FPGA chip to construct a clock tree; receiving, by the first FPGA chip, an original clock signal of a clock source, outputting a first clock signal through a first DCM of the first FPGA chip; transmitting the first clock signal to the second FPGA chip and a feedback input end of the first FPGA chip respectively; performing phase detection on the original clock signal and a second clock signal received by the feedback input end, and performing phase adjustment to align the original clock signal and the second clock signal; determining a first working clock based on the aligned second clock signal; transmitting, by the clock tree of the first FPGA chip, the first working clock to each register inside the first FPGA chip; the second FPGA chip generates a second working clock based on the first clock signal received from the first FPGA chip; and transmitting, by the clock tree of the second FPGA chip, the second working clock to each register inside the second FPGA chip.

[0009] With reference to the first aspect, in a possible implementation manner, the metal wiring inside the first FPGA chip and the second FPGA chip to construct a clock tree comprises: taking a clock origin as a center, taking a first direction passing through the center as a clock main stem, extending a plurality of clock branch stems on the clock main stem in a second direction perpendicular to the first direction, extending a plurality of first-direction local clock lines on each clock branch stem to be connected to each register respectively, and setting configurable switches at the connection positions of the clock main stem and the clock branch stem, the clock branch stem and the local clock line respectively, to construct the clock tree in the metal wiring layer inside the first FPGA chip and the second FPGA chip.

[0010] With reference to the first aspect, in a possible implementation manner, the receiving and parsing, by the first FPGA chip, of the original task package from the external system comprises: receiving, by the first FPGA chip, the original task package from the external system and transmitting the original task package to a task buffer area of the first FPGA chip; structurally disassembling the original task package to extract key information in the original task package; wherein the key information comprises a task type, a data size and a task dependency degree; and evaluating a task parallelism feature of the original task package based on the key information, comprising: if the task type is a data parallel task, evaluating data correlation; and if the task type is a task parallel task, drawing a directed acyclic graph according to the task dependency degree to identify a parallel node therein.

[0011] With reference to the first aspect, in a possible implementation manner, the decomposing the original task package into a plurality of sub-tasks comprises: determining a hardware configuration of the computing platform; determining a division number and a division strategy of the sub-tasks based on the key information and the hardware configuration; and decomposing the original task package into a plurality of sub-tasks according to the division number and the division strategy.

[0012] With reference to the first aspect, in a possible implementation manner, after the original task package is decomposed into a plurality of sub-tasks, the method further comprises: marking a data boundary of each of the sub-tasks based on the first FPGA chip; determining a data dependency relationship between the sub-tasks based on the data boundary; merging the sub-tasks with the data dependency relationship to avoid data cross-sub-task dependency; generating scheduling information of each of the sub-tasks, and transmitting the scheduling information to the first FPGA chip.

[0013] With reference to the first aspect, in a possible implementation manner, the monitoring, by the first FPGA chip, of a load state of each of the second FPGA chips and the dynamically assigning, by the first FPGA chip, of the sub-tasks to the plurality of second FPGA chips based on the load state comprises: periodically sending, by the first FPGA chip, a load request frame to each of the second FPGA chips at a first frequency to obtain the load state of each of the second FPGA chips; traversing the second FPGA chips to perform a task allocation step according to the load state; wherein the task allocation step comprises: assigning a sub-task to each of the second FPGA chips according to the load state of each of the second FPGA chips; performing a feasibility check on the second FPGA chip to which the sub-task is assigned; if the feasibility check passes, issuing, by the first FPGA chip, a task instruction frame to the corresponding second FPGA chip; if the feasibility check does not pass, traversing other second FPGA chips to perform the task allocation step until all second FPGA chips are traversed; and based on the currently updated load state, traversing an idle second FPGA chip to perform the task allocation step.

[0014] In a second aspect, the embodiments of the present application provide a collaborative acceleration device based on distributed FPGA, comprising a plurality of FPGA chips and a clock synchronization module; the plurality of FPGA chips comprise a first FPGA chip and a plurality of second FPGA chips, and the first FPGA chip is connected to the plurality of second FPGA chips through high-speed communication links; the first FPGA chip is configured to: be communicatively connected to an external system; perform calculation on a raw task package received from the external system; and / or, analyze the raw task package and decompose it into a plurality of sub-tasks; monitor the load status of each second FPGA chip and dynamically distribute the sub-tasks to the second FPGA chips based on the load status; collect the processing results of the sub-tasks by each second FPGA chip, and send them to the external system after verification and integration; the second FPGA chip is configured to: receive and execute the sub-tasks from the first FPGA chip, and feed back the processing results of the sub-tasks to the first FPGA chip; send the load status of the second FPGA chip to the first FPGA chip; and the clock synchronization module is configured to: realize clock synchronization of the first FPGA chip and the plurality of second FPGA chips.

[0015] In combination with the second aspect, in a possible implementation, a Flash module is further included; the Flash module is configured to be communicatively connected to the first FPGA chip and the plurality of second FPGA chips respectively; and a communication path is switched by a multiplexer to realize communication with a single FPGA chip.

[0016] In combination with the second aspect, in a possible implementation, a JTAG download module is further included; the JTAG download module is configured to connect the first FPGA chip and the plurality of second FPGA chips in a daisy chain; and programs or firmware are transmitted to the first FPGA chip and / or the second FPGA chip.

[0017] The one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0018] By constructing a multi-FPGA collaborative computing platform, the embodiments of the present application can dynamically allocate computing tasks to a plurality of FPGA chips for parallel processing, thereby improving resource utilization and computing efficiency; by allocating sub-tasks according to the load status, efficient use of computing resources and balanced distribution of task loads can be achieved; and by clock synchronization, data consistency can be ensured. The technical problem of existing model acceleration scheme that the acceleration effect and performance exist bottlenecks and are difficult to meet the needs of complex scenarios is effectively solved. The reliability, flexibility and computing efficiency of the computing platform are further improved. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments of the present application or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort under the premise of not paying creative effort.

[0020] Figure 1 A flowchart of a collaborative acceleration method based on distributed FPGA provided by the embodiments of the present application;

[0021] Figure 2 A structural example diagram of a clock tree provided by the embodiments of the present application;

[0022] Figure 3 A cascade structural example diagram of clock synchronization provided by the embodiments of the present application;

[0023] Figure 4 A structural schematic diagram of a JTAG download module provided by the embodiments of the present application;

[0024] Figure 5 A structural schematic diagram of a collaborative acceleration device based on distributed FPGA provided by the embodiments of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort belong to the scope of protection of the present application.

[0026] The following describes some technologies related to the embodiments of the present application to help understanding, which should be considered only as exemplary. Therefore, those skilled in the art should realize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, in order to be clear and concise, the description of some well-known functions and structures is omitted in the following description.

[0027] Figure 1 A flowchart of a collaborative acceleration method based on distributed FPGA provided by the embodiments of the present application, including steps 101 to 106. Among them, Figure 1 Only one execution order is shown for the embodiments of the present application, which does not represent the only execution order of a collaborative acceleration method based on distributed FPGA, and Figure 1 The steps shown can be executed in parallel or in reverse.

[0028] Step 101: connect a first FPGA chip with a plurality of second FPGA chips through high-speed communication links to build a computing platform. In the embodiment of the present application, a star topology is adopted, taking the first FPGA chip as the center, and connecting with a plurality of second FPGA chips through high-speed communication links respectively to build a computing platform.

[0029] Exemplarily, the present application takes three second FPGA chips as an example for illustration. A large amount of data exchange is required between the first FPGA chip and the plurality of second FPGA chips, so the GTX (high-speed serial transceiver) interface is adopted to ensure the efficiency and reliability of data exchange. That is, the high-speed communication link adopts the GTX link. The computing platform of the present application is a collaborative acceleration device based on distributed FPGA.

[0030] Step 102: clock synchronization of the first FPGA chip and the plurality of second FPGA chips. In the embodiment of the present application, metal wiring is performed inside the first FPGA chip and the second FPGA chip to build a clock tree. The first FPGA chip receives a raw clock signal of a clock source, and outputs a first clock signal through a first DCM of the first FPGA chip. The first clock signal is transmitted to the second FPGA chip and the feedback input end of itself respectively. The raw clock signal and the second clock signal received by the feedback input end are phase detected and phase adjusted to align the raw clock signal and the second clock signal. The first working clock is determined based on the aligned second clock signal. The first working clock is transmitted to each register inside the first FPGA chip through the clock tree of the first FPGA chip. The second FPGA chip generates a second working clock based on the first clock signal received from the first FPGA chip. The second working clock is transmitted to each register inside the second FPGA chip through the clock tree of the second FPGA chip.

[0031] In the embodiment of the present application, taking the clock origin as the center, taking the first direction passing through the center as the clock trunk, extending a plurality of clock branches in the second direction perpendicular to the first direction on the clock trunk, and extending a plurality of local clock lines in the first direction on each clock branch to be connected to each register respectively, and setting configurable switches at the connection between the clock trunk and the clock branch, the clock branch and the local clock line, to build a clock tree in the metal wiring layer inside the first FPGA chip and the second FPGA chip.

[0032] Exemplarily, as shown in Figure 2 The structure example diagram of the clock tree of the present application is shown in the figure, where the yellow dot represents the clock origin, the red line represents the clock trunk, the green line represents the clock branch, the black line represents the local clock line, and the black square represents the register.

[0033] Specifically, the application arranges metal wires dedicated to clock synchronization in the independent metal wiring layers inside the first FPGA chip and the second FPGA chip, avoiding interference from other signals. The clock origin is taken as the center, and a clock main stem is arranged in the vertical direction passing through the center as the first direction. A clock branch stem is arranged on the clock main stem in the second direction (i.e., the horizontal direction) perpendicular to the first direction, as shown in Figure 2 The application exemplarily arranges three uniform clock branch stems, each covering a different FPGA region. A plurality of local clock lines are arranged on each clock branch stem in the first direction (i.e., the vertical direction), and the ends of each local clock line extend to a register (the physical positions of the registers in the metal wiring layer are as uniformly distributed as possible). Configurable switches are arranged at the connections between each clock main stem and clock branch stem and between each clock branch stem and local clock line, obtaining a clock tree.

[0034] In the same clock tree, the clock signals reaching each register from the clock origin all pass through approximately the same path and the same number of configurable switches. Through dynamic adjustment of the configurable switches (such as optimization of switch state or path selection), the actual delay of the clock tree is calibrated, ensuring that the clock signals of the registers on the same clock tree have minimal deviation in the initial state, thereby eliminating the deviation introduced by the internal clock network itself in design.

[0035] It should be noted that, Figure 2 This is only one embodiment of the application, and those skilled in the art can arrange the clock branch stem into other numbers (such as four or five) according to the inventive concept of the application, as long as the registers are uniformly distributed in the metal wiring layer, the clock tree is a symmetrical structure, and the number of configurable switches on each path (such as the number of configurable switches of the clock main stem and the clock branch stem) and the path length (such as the length matching of the clock branch stem and the local clock line) are strictly controlled to ensure that the path delay from the clock origin to any register is highly consistent. The balanced configuration of the configurable switches and the wiring directly determines the clock deviation performance of the clock tree and is the prerequisite for realizing clock synchronization.

[0036] The application tests 16 clock trees simultaneously constructed by different FPGA chips, wherein the shortest delay of the clock trees is 2.5 ns (nanoseconds), the longest delay is 3.6 ns, and the maximum clock deviation is 0.9 ns.

[0037] Further, after the clock tree is constructed, the delay generated by the clock signal transmission (such as the I / O interface transmission between different FPGA chips, the printed circuit board trace on the FPGA chip) needs to be measured and compensated, so as to align the clock trees inside the multiple FPGA chips in phase. The present application exemplarily adopts a DLL (Delay Locked Loop), a DCM (Digital Clock Manager) and a board-level feedback loop to compensate the clock transmission delay between different FPGA chips, so as to realize the clock synchronization between the multiple FPGA chips.

[0038] Specifically, the original clock signal generated by the clock source is directly transmitted to the first DCM of the first FPGA chip as the phase reference of the DLL. After receiving the original clock signal, the first DCM outputs a first clock signal, and transmits the first clock signal to the feedback input end (i.e. the feedback input end of the DLL) of the first FPGA chip itself through a special clock I / O interface (such as LVDS, i.e. Low Voltage Differential Signal interface or HSTL, i.e. High Speed Transceiver Logic interface), forming a closed loop of transmitting the first clock signal output by the first DCM to the feedback input end of the DLL through the board-level transmission. After the first clock signal is transmitted to the feedback input end through the board-level loop, the feedback input end receives a second clock signal.

[0039] It should be noted that in this step, the parameters (frequency division, frequency multiplication ratio, phase offset mode, etc.) of the first DCM need to be controlled so that the first clock signal output by the first DCM is synchronized in frequency and phase with the original clock signal. The configurable switch of the clock tree also needs to be controlled so that the wiring paths of the original clock signal and the first clock signal transmitted in the first FPGA chip are symmetrical, ensuring consistent distribution delay, i.e. only the time delay difference of the board-level transmission, providing a reliable input signal for the phase detection of the DLL.

[0040] The PFD (Phase Detector) inside the DLL performs phase detection on the original clock signal and the second clock signal. If the second clock signal lags behind the original clock signal, it means that the board-level transmission delay is too large, and the delay chain of the DLL needs to be shortened. If the second clock signal leads the original clock signal, the length of the delay chain of the DLL needs to be increased. Based on the result of the phase detection, the delay chain of the DLL is adjusted to reduce the phase difference between the second clock signal and the original clock signal and align them.

[0041] The second DCM of the first FPGA chip uses the aligned second clock signal as input to generate a first working clock, and transmits the first working clock to each register inside the first FPGA chip through the clock tree inside the first FPGA chip.

[0042] The first clock signal is transmitted to the second FPGA chip. The third DCM in the second FPGA chip receives the first clock signal and generates a second operating clock. Using the clock tree inside the second FPGA chip, the second operating clock is transmitted to its internal registers, thereby achieving clock synchronization between the first and second FPGA chips.

[0043] Similarly, such as Figure 3 As shown in the diagram, I represents input, O represents output, D represents diode, Q represents transistor, FB represents feedback element, and FF represents flip-flop. FPGA chips can be constructed in a cascaded structure, where a first FPGA chip is connected in series with multiple second FPGA chips. After the first FPGA chip and an adjacent second FPGA chip are clocked synchronously using a first clock signal (clk1) and a second clock signal (clk2) in the diagram, the third clock signal (clk3) output from the third DCM of this adjacent second FPGA chip can be transmitted to the fourth DCM of the next adjacent second FPGA chip, outputting a fourth clock signal (clk4). The fourth clock signal can then be transmitted to the fifth DCM of the next adjacent second FPGA chip, outputting a fifth clock signal (clk5), thus achieving clock synchronization between the first FPGA chip and multiple second FPGA chips.

[0044] Those skilled in the art should realize that, in addition to the first FPGA chip, each subsequent second FPGA chip transmits the clock signal output by its (DCM) to the next adjacent second FPGA chip, and so on, which can achieve clock synchronization between multiple FPGA chips.

[0045] By synchronizing the clocks of the first FPGA chip with multiple second FPGA chips, the consistency of parallel processing of the FPGA chips can be guaranteed.

[0046] Step 103: The first FPGA chip receives and parses the original task package from the external system, decomposing the original task package into multiple subtasks. In this embodiment, the first FPGA chip receives the original task package from the external system and transmits it to its own task buffer. The original task package is structurally decomposed to extract key information. This key information includes task type, data size, and task dependency. Based on the key information, the task parallelism characteristics of the original task package are evaluated, including: if the task type is a data-parallel task, the data correlation is evaluated; if the task type is a task-parallel task, a directed acyclic graph is drawn based on the task dependency to identify parallelizable nodes.

[0047] In the embodiment of the present application, the hardware configuration of the computing platform is determined. The hardware configuration includes the number of FPGA chips and the available computing units, memory capacity and data processing bandwidth of each FPGA chip. The number of subtask division and the division strategy are determined based on the key information and the hardware configuration. The original task package is divided into multiple subtasks according to the number of division and the division strategy.

[0048] Specifically, the first FPGA chip is connected with the external system through a PCIe (Peripheral Component Interconnect Express) interface. After receiving the original task package from the external system, the first FPGA chip transmits the original task package to its own task cache area first, and then performs structured decomposition on the original task package to extract the key information of the original task package, including task type, data size (such as image pixel size, data frame number, matrix dimension, etc.) and task dependency.

[0049] If the task type is a data parallel task, such as image batch processing and large matrix block operation, the data correlation is evaluated, that is, whether the data can be input to the calculation independently after being split, without relying on other data or the processing results of other data.

[0050] If the task type is a task parallel task, such as image recognition, posture control and communication decoding modules that need to be run simultaneously, a directed acyclic graph is drawn according to the task dependency to identify the parallel nodes. The task dependency is to determine whether each module (task) has independent input and output logic and whether it needs to rely on the output results of other modules. According to the task dependency, different tasks are taken as nodes, and the dependency relationship between tasks is taken as directed edges to draw a directed acyclic graph. The level of different nodes in the directed acyclic graph is determined by topological sorting. Nodes with the same level are parallel nodes, and the tasks represented by the parallel nodes can be executed in parallel.

[0051] The hardware configuration of the computing platform is obtained, including the number of FPGA chips and the available computing units (such as the number of logic gates and the number of DSP slices, where DSP represents digital signal processor) of each FPGA chip, memory capacity and data processing bandwidth (GTX interface transmission rate). Then, the number of subtask division (according to the number of FPGA chips and the maximum task size they can bear) and the division strategy are determined according to the key information of the original task package and the hardware configuration of the computing platform. Specifically, if the data parallel degree of the original task package is high, the input data is divided into multiple data blocks according to the number of division. If the task parallel degree of the original task package is high, the tasks in the original task package are divided into multiple task segments according to the number of division. The original task package is divided into multiple subtasks according to the number of division and the division strategy.

[0052] In the embodiment of the present application, after the original task package is decomposed into multiple sub-tasks, the data boundaries of each sub-task can be marked based on the first FPGA chip. The data dependency relationship between the sub-tasks is determined based on the data boundaries. The sub-tasks with data dependency relationship are merged to avoid data cross-sub-task dependency. The scheduling information of each sub-task is generated and transmitted to the first FPGA chip. The scheduling information includes sub-task ID, data boundary, task configuration (in-frame block position, resource requirement, etc.), and output information (output data type, output address, etc.).

[0053] Specifically, after the original task package is split into multiple sub-tasks, each sub-task can be verified. A task ID is generated for each sub-task, and the data boundary is determined according to the data size of the original task package and the data processing amount of each sub-task. For example, if the original data package is divided into 4 sub-tasks, and the data size of the original task package is 100 frames, then each sub-task processes 25 frames, the data boundary of the first sub-task is 1-25 frames, the data boundary of the second sub-task is 26-50 frames, and so on. The data dependency relationship between different sub-tasks is obtained by judging whether the data boundaries of the sub-tasks cross or overlap. If there is a dependency relationship between the sub-tasks (the data boundaries cross or overlap), the corresponding sub-tasks are merged to avoid data cross-sub-task dependency, and then the situation of task overload, data dirty read / write, deadlock, data transmission cost, etc. may occur. The scheduling information of the last remaining sub-task is determined and transmitted to the first FPGA chip, and each sub-task is allocated and scheduled by the first FPGA chip.

[0054] Step 104: Monitor the load state of each second FPGA chip through the first FPGA chip, and dynamically allocate sub-tasks to multiple second FPGA chips based on the load state. In the embodiment of the present application, the first FPGA chip periodically sends a load request frame to each second FPGA chip at a first frequency to obtain the load state of each second FPGA chip. The task allocation step is executed according to the load state by traversing the second FPGA chip. The task allocation step is as follows: according to the load state of each second FPGA chip, the sub-tasks are allocated to it. The second FPGA chip to which the sub-tasks are allocated performs a feasibility check. If the feasibility check passes, the first FPGA chip issues a task instruction frame to the corresponding second FPGA chip. If the feasibility check does not pass, the task allocation step is executed by traversing other second FPGA chips until all second FPGA chips are traversed. Based on the current updated load state, the task allocation step is executed by traversing the idle second FPGA chip.

[0055] The load state includes a computing unit utilization rate, a memory occupancy rate, a bandwidth utilization rate and a task remaining time length. Specifically, the first FPGA chip periodically sends a load request frame to each second FPGA chip at a first frequency (exemplarily set as 10 milliseconds / time), and after receiving the load request frame sent by the first FPGA chip, the second FPGA chip verifies whether the information (target second FPGA chip ID) in the load request frame matches itself, then reads the load state of itself, and packs the load state into a load state frame and sends the load state frame to the first FPGA chip. The frame structure of the load request frame is: [request identification (1 byte)+target second FPGA chip ID (1 byte)], and the frame structure of the load state frame is: [second FPGA chip ID (1 byte)+computing unit utilization rate (1 byte)+memory occupancy rate (1 byte)+bandwidth utilization rate (1 byte)+task remaining time length (2 bytes)+CRC check bit (2 bytes)], and the size of a single load state frame is 8 bytes, which is transmitted every 10 milliseconds, and compared with the data amount of 1 billion bits transmitted per second by the GTX interface, the bandwidth occupied by the load state frame can be ignored.

[0056] After the first FPGA chip receives the load state of each second FPGA chip, the first FPGA chip first performs CRC (cyclic redundancy check) verification, re-sends the load request frame if the verification fails, and stores the corresponding load data into a load data table if the verification passes, and then allocates sub-tasks to the second FPGA chips according to the load state.

[0057] After the sub-tasks are allocated, the second FPGA chips to which the sub-tasks are allocated are subjected to a feasibility check, that is, whether the remaining resources of the second FPGA chips after the sub-tasks are allocated are sufficient to run the system basic logic or cache intermediate data. If the feasibility check of the second FPGA chip passes, the first FPGA chip generates a task instruction frame and sends the task instruction frame to the corresponding second FPGA chip through the GTX interface. If the feasibility check of the second FPGA chip does not pass, the second FPGA chip is temporarily shelved, and the task allocation step is performed on other idle second FPGA chips until all sub-tasks are divided.

[0058] Exemplarily, the frame structure of the task instruction frame is: [allocation identification (1 byte)+target second FPGA chip ID (1 byte)+sub-task ID (1 byte)+sub-task data address (4 bytes, corresponding to the Flash address of the data boundary)+output address of the sub-task (4 bytes, corresponding to the DDR address of the second FPGA chip)], wherein the Flash is a non-volatile memory, and the DDR is a double data rate synchronous dynamic random access memory.

[0059] Step 105: Each second FPGA chip processes the sub-tasks in parallel and feeds the processing result back to the first FPGA chip. In the embodiment of the present application, after receiving the task instruction frame, each second FPGA chip parses the instruction, verifies whether the sub-task data address is accessible, and then replies to the first FPGA chip with a task confirmation frame (containing the second FPGA chip ID, the sub-task ID and the confirmation mark), and then starts processing the corresponding sub-task, sends the processing result to the first FPGA chip after completing the sub-task, and finally releases the occupied resources to become an idle second FPGA chip, waiting to be assigned a task again.

[0060] Step 106: The first FPGA chip checks and integrates the processing results of each second FPGA chip to obtain the task execution result, and feeds the task execution result back to the external system. In the embodiment of the present application, after receiving the processing result sent by each second FPGA chip, the first FPGA chip first performs CRC check to ensure the integrity of the data. Then, the processing results of each second FPGA chip are merged to obtain the task execution result of the original task package, which is sent to the external system through the PCIe interface.

[0061] The present application can effectively improve the resource utilization and the computing efficiency by connecting multiple FPGA chips into a distributed computing cluster, and each FPGA chip can be used as an independent computing unit to independently process tasks and work cooperatively with other modules (or FPGA chips).

[0062] Although the present application provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps can be included based on conventional or non-creative labor. The order of steps listed in the embodiments is only one of the many execution orders, and does not represent the only execution order. In actual device or client product execution, the method order shown in the embodiments or the drawings can be executed in sequence or in parallel (for example, in a parallel processor or multi-thread processing environment).

[0063] As shown in FIG. 1, the present application also provides a collaborative acceleration device based on distributed FPGA, which includes a plurality of FPGA chips, a clock synchronization module, a Flash module and a JTAG download module, and specifically as follows. Figure 5 The plurality of FPGA chips includes a first FPGA chip and a plurality of second FPGA chips, and the first FPGA chip is connected to the plurality of second FPGA chips through a high-speed communication link.

[0064]

[0065] ​The first FPGA chip is configured to be in communication connection with an external system, to perform calculation on a raw task package received from the external system, and / or to parse the raw task package and decompose it into a plurality of sub-tasks, to monitor load states of the second FPGA chips and dynamically distribute the sub-tasks to the second FPGA chips based on the load states, and to collect processing results of the sub-tasks by the second FPGA chips and send them to the external system after verification and integration.

[0066] Specifically, the first FPGA chip and the second FPGA chips are connected by high-speed GTX links, each FPGA chip is equipped with a plurality of GTX transceivers for high-speed data transmission, to ensure stable data transmission performance under high-speed data flow, and the GTX transceivers use differential signal transmission and have strong anti-interference ability, to ensure the reliability of data transmission. The first FPGA chip and the external system are connected by a PCIe link.

[0067] The second FPGA chip is configured to receive and execute the sub-tasks from the first FPGA chip and feed back processing results of the sub-tasks to the first FPGA chip, and to send its load state to the first FPGA chip.

[0068] The clock synchronization module is configured to realize clock synchronization of the first FPGA chip and the plurality of second FPGA chips.

[0069] In the embodiments of the present application, the Flash module is configured to be in communication connection with the first FPGA chip and the plurality of second FPGA chips respectively, and to switch communication paths by a multiplexer to realize communication with a single FPGA chip.

[0070] Specifically, in order to realize that the four FPGA chips share one Flash module for configuration and data storage, a parallel configuration circuit is adopted. The circuit connects the Flash module and the four FPGA chips through a BPI bus to realize parallel transmission and control of data. Specifically, the control command lines of the Flash module are connected to the corresponding control pins of the four FPGA chips, and the data lines and address lines of the Flash module are also connected to the data and address ports of the four FPGA chips through a multiplexer or a buffer circuit element. In this way, during the configuration process of the FPGA chip, the state of the multiplexer or the buffer can be controlled to selectively transmit data or commands in the Flash module to the designated FPGA chip.

[0071] The four FPGA chips in the present application share one Flash module for configuration and data storage, so the capacity of the Flash module needs to meet the storage requirements of the four FPGA chips, and the data transmission rate also meets the requirements of the FPGA chips.

[0072] In this embodiment, the JTAG download module is configured to: connect a first FPGA chip and multiple second FPGA chips in a daisy chain; and transfer programs or firmware to the first FPGA chip and / or the second FPGA chip.

[0073] Specifically, such as Figure 4 As shown in the diagram (in the figure, the JTAG Header is the physical interface used to connect the JTAG debugging tool and the target device; TDI represents the serial data input port; TDO represents the serial data output port; TMS represents the mode selection switch; and TCK represents the clock pulse signal), this daisy-chain structure can save PCB space occupied by multiple JTAG ports. If the collaborative acceleration device based on distributed FPGA is in a closed environment, and sometimes it is necessary to upgrade the program of the FPGA chip in the device online or remotely, it is necessary to bring the JTAG port outside the chassis. This single JTAG port daisy-chain structure is the most convenient.

[0074] In addition, this application also includes a power management module and a reset module. The power management module is configured to power multiple FPGA chips, ensuring the stability of the power supply. The reset module is configured to attempt to restore normal operation by restarting in the event of a fault or abnormality.

[0075] Specifically, each FPGA chip has an independent reset signal input pin, which can be used to perform reset operations separately. The reset signal comes from the physical reset button of the external reset source.

[0076] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0077] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0078] The methods, apparatuses or modules described in the present application can be implemented in a computer readable program code in any appropriate manner, for example, the controller can take the form of, for example, a microprocessor or processor and a computer readable medium storing computer readable program code (for example, software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASIC), programmable logic controllers and embedded microcontrollers, examples of the controller include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in pure computer readable program code, the same function can be achieved by logically programming the method steps in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers. Therefore, such a controller can be considered as a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both a software module for implementing the method and a structure within the hardware component.

[0079] The embodiments of the present application also provide a device, which comprises: a processor; a memory for storing processor executable instructions; and the processor implements the method as described in the embodiments of the present application when executing the executable instructions.

[0080] The embodiments of the present application also provide a non-volatile computer readable storage medium, which stores a computer program or instructions, and when the computer program or instructions are executed, the method as described in the embodiments of the present application is implemented.

[0081] In addition, the functional modules in the various embodiments of the present application can be integrated in one processing module, or each module can exist independently, or two or more modules can be integrated in one module.

[0082] The storage medium described above includes but is not limited to random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD) or memory card. The memory can be used to store computer program instructions.

[0083] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary hardware. Based on such an understanding, the technical solutions of the present application can be embodied in the form of a software product or can be embodied in the form of data migration in the implementation process. The computer software product can be stored in a storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the method described in each embodiment or some parts of the embodiments of the present application.

[0084] The various embodiments in the specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment mainly describes the difference from other embodiments. The whole or part of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, etc.

[0085] The above embodiments are only used to illustrate the technical solutions of the present application, and not to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present application.

Claims

1. A method for distributed FPGA-based collaborative acceleration, the method comprising: The application relates to a method for constructing a computing platform by connecting a first FPGA chip and a plurality of second FPGA chips through a high-speed communication link. The first FPGA chip and the plurality of second FPGA chips are clock-synchronized, and a task processing step is executed based on the computing platform to realize task acceleration. The clock synchronization of the first FPGA chip and the plurality of second FPGA chips comprises the following steps: metal wiring is performed in the first FPGA chip and the second FPGA chip to construct a clock tree; a first FPGA chip receives an original clock signal of a clock source, and outputs a first clock signal through a first DCM of the first FPGA chip; the first clock signal is transmitted to the second FPGA chip and a feedback input end of the first FPGA chip; phase detection is performed on the original clock signal and a second clock signal received by the feedback input end, and phase adjustment is performed to align the original clock signal and the second clock signal; a first working clock is determined based on the aligned second clock signal; the first working clock is transmitted to each register in the first FPGA chip through the clock tree of the first FPGA chip; the second FPGA chip generates a second working clock based on the first clock signal received from the first FPGA chip; and the second working clock is transmitted to each register in the second FPGA chip through the clock tree of the second FPGA chip. The metal wiring in the first FPGA chip and the second FPGA chip to construct the clock tree comprises the following steps: a clock origin is taken as a center, a first direction passing through the center is taken as a clock main stem, a plurality of clock branch stems are extended on the clock main stem along a second direction perpendicular to the first direction, a plurality of first-direction local clock lines are extended on each clock branch stem and connected to each register, and configurable switches are arranged at the connection positions of the clock main stem and the clock branch stem and the clock branch stem and the local clock line to construct the clock tree in the metal wiring layer in the first FPGA chip and the second FPGA chip. The task processing step comprises the following steps: The first FPGA chip receives and analyzes an original task package from an external system, and decomposes the original task package into a plurality of subtasks; The first FPGA chip monitors the load states of the second FPGA chips, and dynamically allocates the subtasks to the second FPGA chips based on the load states; The second FPGA chips perform parallel processing on the subtasks, and feed back processing results to the first FPGA chip; The first FPGA chip checks and integrates the processing results of the second FPGA chips, obtains a task execution result, and feeds back the task execution result to the external system. The first FPGA chip receives an original task package from an external system, and transmits the original task package to a task buffer area of the first FPGA chip.

2. The method of claim 1, wherein, ​ ​ Structurally disassembling the original task package to extract key information in the original task package, wherein the key information comprises a task type, a data scale, and a task dependency degree; Evaluating a task parallelism feature of the original task package based on the key information, comprising: If the task type is a data parallel task, evaluating data correlation; If the task type is a task parallel task, drawing a directed acyclic graph according to the task dependency degree to identify parallel nodes therein.

3. The method of claim 2, wherein, The task package is divided into a plurality of sub-tasks, comprising: Determining a hardware configuration of a computing platform; Based on the key information and the hardware configuration, determining a division number and a division strategy of a sub-task; According to the division number and the division strategy, the original task package is divided into a plurality of sub-tasks.

4. The method of claim 1, wherein, After the original task package is divided into a plurality of sub-tasks, it further comprises: Based on the first FPGA chip, marking the data boundary of each sub-task; Based on the data boundary, determining the data dependency relationship between sub-tasks; Merging sub-tasks with data dependency relationship to avoid data cross-sub-task dependency; Generating scheduling information for each sub-task and transmitting it to the first FPGA chip.

5. The method of claim 1, wherein, The first FPGA chip monitors the load state of each second FPGA chip, and dynamically allocates the sub-tasks to the plurality of second FPGA chips based on the load state, comprising: The first FPGA chip periodically sends a load request frame to each second FPGA chip at a first frequency to obtain the load state of each second FPGA chip; Iterate through the second FPGA chip and execute the task allocation step according to the load state; The task allocation step is as follows: According to the load state of each second FPGA chip, allocate sub-tasks to it; The second FPGA chip to which the sub-tasks are allocated performs feasibility verification; If the feasibility verification is passed, the first FPGA chip issues a task instruction frame to the corresponding second FPGA chip; If the feasibility verification is not passed, iterate through other second FPGA chips to execute the task allocation step until all second FPGA chips are iterated through; Based on the current updated load state, iterate through the idle second FPGA chip to execute the task allocation step.

6. A distributed FPGA based co-accelerator for implementing the method of any of claims 1-5, characterized by, Comprise a plurality of FPGA chips and a clock synchronization module; The plurality of FPGA chips comprise a first FPGA chip and a plurality of second FPGA chips, and the first FPGA chip is connected to the plurality of second FPGA chips through a high-speed communication link; The first FPGA chip is configured to: communicate with an external system; perform calculation on the original task package received from the external system; and / or, parse the original task package and divide it into a plurality of sub-tasks; monitor the load state of each second FPGA chip and dynamically distribute the sub-tasks to the second FPGA chips based on the load state; Collect the processing results of each second FPGA chip on the sub-tasks, and send them to the external system after verification and integration; The second FPGA chip is configured to receive and execute the sub-tasks from the first FPGA chip, and feed back the processing results of the sub-tasks to the first FPGA chip; and send its load state to the first FPGA chip. The clock synchronization module is configured to realize clock synchronization of the first FPGA chip and the plurality of second FPGA chips, including: performing metal wiring inside the first FPGA chip and the second FPGA chip to construct a clock tree; receiving an original clock signal of a clock source through the first FPGA chip, outputting a first clock signal through a first DCM of the first FPGA chip; transmitting the first clock signal to the second FPGA chip and a feedback input end of the first FPGA chip respectively; performing phase detection on the original clock signal and a second clock signal received by the feedback input end, and performing phase adjustment to align the original clock signal and the second clock signal; determining a first working clock based on the aligned second clock signal; transmitting the first working clock to each register inside the first FPGA chip through the clock tree of the first FPGA chip; the second FPGA chip generates a second working clock based on the first clock signal received from the first FPGA chip; and transmitting the second working clock to each register inside the second FPGA chip through a clock tree of the second FPGA chip. The metal wiring inside the first FPGA chip and the second FPGA chip to construct the clock tree includes: taking a clock origin as a center, taking a first direction passing through the center as a clock main stem, extending a plurality of clock branch stems on the clock main stem along a second direction perpendicular to the first direction, extending a plurality of local clock lines of the first direction on each clock branch stem to be connected to each register respectively, and setting configurable switches at the connection positions of the clock main stem and the clock branch stem, and the clock branch stem and the local clock line respectively, to construct the clock tree in the metal wiring layer inside the first FPGA chip and the second FPGA chip.

7. The apparatus of claim 6, wherein, Further comprising a Flash module; The Flash module is configured to be communicatively connected to the first FPGA chip and the plurality of second FPGA chips respectively; and switch a communication path through a multiplexer to realize communication with a single FPGA chip.

8. The apparatus of claim 6, wherein, Further comprising a JTAG download module; The JTAG download module is configured to connect the first FPGA chip and the plurality of second FPGA chips in series in a daisy chain form; and transmit a program or firmware to the first FPGA chip and / or the second FPGA chip.

Citation Information

Patent Citations

  • CPU+FPGA-based heterogeneous computing system and acceleration method thereof

    CN108776649A

  • Multi-FPGA (Field Programmable Gate Array) cooperative multi-deep neural network pipeline acceleration method

    CN119808859A