Soft and hard collaborative computing framework based on FPGA (Field Programmable Gate Array)
By adopting a software-hard and hard-core collaborative computing framework based on FPGA in the graph computing engine, combining hardware acceleration and software computing advantages, the problems of inefficient computing efficiency and data transmission efficiency in large-scale graph data processing are solved, and higher computing efficiency and data transmission speed are achieved.
Patent Information
- Application Number
- CN202510106164.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing graph computing engines rely on traditional CPU computing and memory transmission bandwidth, resulting in low computing efficiency and low data transmission efficiency during large-scale graph data processing, and fail to make full use of the parallel computing capabilities of modern computing devices.
Using a software and hard-core collaborative computing framework based on FPGA, the coordinated work of the PS side and the PL side is combined with the advantages of FPGA hardware acceleration and software computing to achieve efficient processing of graph computing tasks. The specific implementation includes generating graph calculation instructions on the PS side and sending them to the PL side through the DMA channel. The command processor on the PL side performs graph calculation operations and returns the results to the PS side to form a closed-loop data stream.
Through the combination of FPGA hardware acceleration and parallel computing capabilities, the computing efficiency and data transmission speed of graph computing are significantly improved, breaking through the performance bottleneck of traditional CPU computing, and improving the overall performance and real-time performance of the system.
Smart Images

Figure CN120029947A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of graph data processing technology, and more specifically, to a FPGA-based soft-hard collaborative computing framework. Background Art
[0002] In the existing technology, graph computing, as a typical parallel computing mode, has been widely used in computer science, network analysis, artificial intelligence and other fields. Its core task is to efficiently process graph data, including the storage of graph structures, the operation of nodes and edges, and complex pathfinding calculations. With the expansion of data scale and the increase of computing needs, it is difficult to meet the requirements of efficient and real-time computing by simply relying on traditional CPU processing capabilities.
[0003] Currently, most graph computing engines rely on traditional computer architectures, with CPUs as the main computing unit. However, this approach has several significant problems. For example, due to the limited processing power of a single CPU core, it is unable to provide efficient processing efficiency when faced with large-scale parallel computing tasks, especially when performing large-scale path-finding calculations such as depth-first search (DFS) and breadth-first search (BFS), which easily leads to waste of computing resources and degradation of system performance. In addition, traditional graph computing engines mainly rely on memory storage and cache mechanisms to complete data exchange, which makes memory bandwidth one of the main obstacles to performance improvement, especially when processing large-scale graph data, where low data transmission efficiency seriously affects overall performance. Existing graph computing engines generally lack effective hardware acceleration module support and fail to fully utilize the powerful parallel computing capabilities provided by modern computing devices. Therefore, they are inefficient when processing complex graph computing tasks and fail to fully utilize the potential of hardware.
[0004] Therefore, how to develop a new soft-hard collaborative computing framework to solve performance bottlenecks by combining the advantages of hardware acceleration and software computing, and improve computing efficiency and data transmission speed has become one of the research directions for technical personnel in this field. Summary of the invention
[0005] The purpose of this application is to provide an FPGA-based soft-hard collaborative computing framework in order to overcome the existing technical defects, which completes graph computing tasks through the collaborative work of the PS side and the PL side, combining the advantages of FPGA hardware acceleration and software computing.
[0006] The purpose of this application is achieved through the following technical solutions:
[0007] In the first aspect, the present application proposes a FPGA-based soft-hard collaborative computing framework, the framework is based on a Z19 FPGA board, including a PS side and a PL side, and the interface between the PS side and the PL side adopts a DMA-FIFO data transmission method;
[0008] The PS side sends the graph calculation instructions generated by the command management software to the command request FIFO through the DMA channel for the command processor on the PL side to read;
[0009] The command processor on the PL side receives instructions from the command request FIFO, performs corresponding operations according to the instructions, and generates results. It puts the generated results into the command response FIFO and returns them to the PS side, forming a closed-loop data flow.
[0010] In a possible implementation, the PS side adopts the Petalinux operating system, uses NVMe SSD as persistent storage of graph data, and uses the storage space as a graph resource management module.
[0011] In a possible implementation, the PL side includes a graph data reading and writing module and a 480-way pathfinder, and the pathfinder supports depth-first and breadth-first strategies.
[0012] In a possible implementation, the PL side also includes a command processor, which is responsible for receiving commands from the PS side and starting corresponding parallel computing units to accelerate graph computing tasks, and is also used to execute creation and deletion of nodes and edges, and graph pathfinding operations.
[0013] In one possible implementation, the framework includes a computing unit having a unique resource ID and attribute information, the input of the computing unit being a resource description derived from software definition, and being used to generate an access address according to the resource ID and to perform read and write operations on a graph data entity.
[0014] In a possible implementation, the PL side executes space mapping logic to convert the ID into an access address in the addressable address space. The space mapping logic is: DDR4 =(ID×Size 节点、边 )+Base 节点、边 , where Addr DDR4 is the final address mapped to the DDR4 memory, ID is the unique number used to identify the data entity, Size 节点、边 Base is the number of bytes occupied by each node or edge. 节点、边 The base address of the starting position of the node or edge data block in the DDR4 memory.
[0015] In one possible implementation, the computation unit includes a DSP macro for implementing a multiplier that performs spatial mapping logic.
[0016] The above-mentioned main scheme of the present application and its further options can be freely combined to form multiple schemes, all of which are schemes that can be adopted and claimed for protection in the present application; and in the present application, (non-conflicting options) options and other options can also be freely combined. After understanding the scheme of the present application, those skilled in the art can understand that there are multiple combinations based on the prior art and common knowledge, all of which are technical schemes to be protected by the present application, and they are not exhaustively listed here.
[0017] The present application discloses a FPGA-based soft-hard collaborative computing framework, which is based on the Z19 FPGA board and includes a PS side and a PL side. The interface between the PS side and the PL side adopts a DMA-FIFO data transmission method. The PS side sends the graph computing instructions generated by the command management software to the command request FIFO through the DMA channel for the command processor on the PL side to read. The command processor on the PL side receives instructions from the command request FIFO, and performs corresponding operations according to the instructions to obtain the generated results, and puts the generated results into the command response FIFO and returns them to the PS side, forming a closed-loop data flow. Through the collaborative work of the PS side and the PL side, combined with the advantages of FPGA hardware acceleration and software computing, the graph computing tasks are completed together, the performance bottleneck problem is solved, and the computing efficiency and data transmission speed can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 A schematic diagram of a FPGA-based soft-hard collaborative computing framework proposed in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0020] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0021] Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of this application.
[0022] In the prior art, traditional graph computing engines are difficult to achieve efficient processing of large-scale graph computing due to their reliance on CPU computing and memory transmission bandwidth. When processing parallel computing in graph computing, the CPU often cannot fully utilize the parallel computing capabilities of the hardware, resulting in low computing efficiency. The data transmission speed between the CPU and the memory is slow, especially when processing large-scale graph data, data transmission becomes a performance bottleneck, affecting the overall computing efficiency. Existing graph computing engines lack hardware acceleration modules and cannot improve computing efficiency through hardware timing logic and parallel computing.
[0023] Therefore, in order to solve the above-mentioned technical problems, the embodiment of the present application proposes a FPGA-based soft-hard collaborative computing framework, which combines the advantages of FPGA hardware acceleration and software computing, fully utilizes the parallel computing capabilities and efficient data transmission mechanism of FPGA, and improves the overall performance and real-time performance of the graph computing engine.
[0024] Please refer to Figure 1 , Figure 1 A schematic diagram of a FPGA-based soft-hard collaborative computing framework proposed in an embodiment of the present application is shown, the framework is based on a Z19 FPGA board, includes a PS side and a PL side, and the interface between the PS side and the PL side adopts a DMA-FIFO data transmission method;
[0025] The PS side sends the graph calculation instructions generated by the command management software to the command request FIFO through the DMA channel for the command processor on the PL side to read;
[0026] The command processor on the PL side receives instructions from the command request FIFO, performs corresponding operations according to the instructions, and generates results. It puts the generated results into the command response FIFO and returns them to the PS side, forming a closed-loop data flow.
[0027] The data exchange between the PS side and the PL side adopts a combination of DMA (direct memory access) and FIFO (first-in, first-out queue). This method can achieve efficient transmission of batch data, reduce the burden on the CPU, and ensure the speed and reliability of data transmission.
[0028] On the PS side, the command management software generates graph calculation instructions and sends these instructions to the command request FIFO on the PL side through the DMA channel. This step ensures that the instructions can be quickly and reliably delivered to the PL side. The command processor on the PL side will continuously monitor the command request FIFO. Once a new instruction is detected, it will read the instruction and complete the corresponding graph calculation operation according to its requirements. When the PL side completes the operation specified by the instruction, it will put the processing result into the command response FIFO. The PS side reads these results from the command response FIFO to form a closed-loop data flow.
[0029] The whole process forms an efficient closed-loop system, in which the PS side is responsible for generating and managing instructions, and the PL side is responsible for executing specific graph computing tasks and returning the results to the PS side. This not only improves the overall efficiency of the system, but also ensures the real-time and accuracy of data processing.
[0030] The PS side uses the Petalinux operating system, uses NVMe SSD as persistent storage for graph data, and uses the storage space as a graph resource management module.
[0031] The PS side uses the Petalinux operating system, which is an operating system customized for embedded Linux applications. It is suitable for hardware platforms such as FPGA and can provide a hardware abstraction layer to make software development more convenient. The graph data persistent storage uses a 2TB NVMe SSD as a persistent storage device for graph data, which can significantly improve the efficiency of loading and saving graph data. About 1GB of storage space is dedicated to the graph resource management module, which is responsible for managing and scheduling the resources required for graph computing tasks, such as nodes and edges.
[0032] The PL side includes a graph data reading and writing module and a 480-way pathfinder, and the pathfinder supports depth-first and breadth-first strategies.
[0033] The graph data reading and writing module on the PL side is responsible for data exchange with the PS side. Specifically, it efficiently transmits batch graph data through the DMA-FIFO mechanism, ensuring the fast transfer of graph data between the PS side and the PL side. Each of the 480 pathfinders can independently perform depth-first search (DFS) or breadth-first search (BFS). This parallel processing capability greatly improves the speed and efficiency of graph traversal. Each pathfinder has a certain amount of local cache and control logic to optimize its workflow.
[0034] The PL side also includes a command processor, which is responsible for receiving commands from the PS side and starting the corresponding parallel computing units to accelerate graph computing tasks. It is also used to perform node and edge creation and deletion, and graph pathfinding operations.
[0035] The command processor is located on the PL side and is one of the core components of the entire system. It is responsible for receiving graph computing instructions sent by the PS side from the command request FIFO. After receiving the instruction, the command processor will parse the instruction content and start the corresponding parallel computing unit to accelerate the graph computing task. These parallel computing units include but are not limited to 480-way pathfinders, which can perform operations such as depth-first search (DFS) and breadth-first search (BFS) in parallel.
[0036] The command processor is also responsible for executing operations related to the graph structure, such as creating and deleting nodes and edges, which is crucial for dynamic graph computing and ensures the flexibility and variability of the graph data structure. When path finding or traversal is required, the command processor schedules the pathfinder to perform specific pathfinding tasks. Due to the use of parallel processing, the system can significantly improve the efficiency of pathfinding operations at the hardware level.
[0037] Through parallel computing units, the system can process multiple graph computing tasks simultaneously, thereby greatly improving computing efficiency. Especially in scenarios that require a large number of concurrent operations (such as depth-first and breadth-first traversal of large-scale graph data), efficient parallel processing capabilities enable the system to complete complex graph computing tasks in a shorter time, improving the system's response speed and real-time performance.
[0038] The framework includes a computing unit with a unique resource ID and attribute information. The input of the computing unit is a resource description derived from software definition, which is used to generate an access address according to the resource ID and perform read and write operations on graph data entities.
[0039] Each computing unit has a unique resource ID and related attribute information. These attributes include the type of computing unit, configuration parameters, etc., which are used to identify and describe the function and behavior of the computing unit. The input of the computing unit comes from the software-defined resource description, which is generated by the command management software on the PS side and passed to the computing unit on the PL side through the DMA-FIFO channel. According to the received resource ID, the computing unit can generate the corresponding access address @(ID), which is used to locate and access the specific graph data entity to ensure that the computing unit can correctly read or write the required data. After generating the access address, the computing unit can perform read and write operations on the graph data entity. The operation may involve reading or updating the data of nodes and edges, thereby realizing various tasks in graph computing, such as creating and deleting nodes or edges, and path finding.
[0040] The PL side executes the space mapping logic to convert the ID into an access address in the addressable address space. The space mapping logic is: DDR4 =(ID×Size 节点、边 )+Base 节点、边 , where Addr DDR4 is the final address mapped to the DDR4 memory, ID is the unique number used to identify a specific data entity, Size 节点、边 Base is the number of bytes occupied by each node or edge. 节点、边 The base address of the starting position of the node or edge data block in the DDR4 memory.
[0041] In the computing logic on the PL side, all graph data entities are addressed using IDs, and the space mapping logic can convert IDs into access addresses within the addressable address space.
[0042] The computational unit includes a DSP macro that implements the multipliers that perform the spatial mapping logic.
[0043] In order to improve the efficiency of parallel execution of computing logic, a DSP (digital signal processing) macro is assigned to each computing unit. A 480-way pathfinder and a data entity access unit require a total of 481 DSP macros, which are mainly used to implement the multipliers in the execution space mapping logic.
[0044] It is worth noting that the management of graph data and graph resources is persisted through 2TB SSD storage. An efficient DMA (direct memory access) channel is used to exchange data with the PL side to ensure fast reading and writing of data during graph calculation. The computing unit uses DSP macros to accelerate multiplication operations, combined with high-speed DDR4 memory, to achieve efficient graph data access and processing.
[0045] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0046] First, by implementing a 480-way parallel pathfinder on the PL side and using the DSP macro in the FPGA to perform efficient multiplication operations, the system can efficiently perform graph computing tasks, greatly improving the computing speed and efficiency, and breaking through the performance bottleneck of traditional CPU computing.
[0047] Second, the use of efficient DMA channels and FIFO queue mechanisms reduces CPU intervention, allowing large-scale graph data to be exchanged quickly and stably between the PS and PL, greatly improving data transmission efficiency and avoiding the impact of traditional memory bandwidth bottlenecks.
[0048] Third, the timing logic characteristics of FPGA are used to accelerate graph computing tasks, especially graph pathfinding operations, to achieve higher parallelism and lower latency. At the same time, a dedicated DSP macro is assigned to each computing task, which further enhances the processing power and response speed of the computing unit and improves the overall computing power of the system.
[0049] Fourth, through modular hardware design and software management architecture, the system can flexibly configure computing resources according to actual needs to ensure efficient management and scheduling, which not only improves the stability and reliability of the system, but also facilitates future expansion and upgrades to adapt to ever-changing computing needs.
[0050] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A soft-hard collaborative computing framework based on FPGA, characterized in that: The software-hardware collaborative computing framework is based on the Z19FPGA board, including the PS side and the PL side, and the interface between the PS side and the PL side adopts the DMA-FIFO data transmission mode; The PS side sends the graph calculation instructions generated by the command management software to the command request FIFO through the DMA channel for the command processor on the PL side to read; The command processor on the PL side receives instructions from the command request FIFO, performs corresponding operations according to the instructions, and generates results. It puts the generated results into the command response FIFO and returns them to the PS side, forming a closed-loop data flow.
2. The software-hardware collaborative computing framework according to claim 1, characterized in that: The PS side adopts the Petalinux operating system, uses NVMe SSD as the persistent storage of graph data, and uses the storage space as a graph resource management module.
3. The software-hard collaborative computing framework according to claim 1, characterized in that: The PL side includes a graph data reading and writing module and a 480-way pathfinder, and the pathfinder supports depth-first and breadth-first strategies.
4. The software-hard collaborative computing framework according to claim 3, characterized in that: The PL side also includes a command processor, which is responsible for receiving commands from the PS side and starting the corresponding parallel computing unit to accelerate the graph computing task, and is also used to execute the creation and deletion of nodes and edges, and graph pathfinding operations.
5. The software-hardware collaborative computing framework according to claim 1, characterized in that: The framework includes a computing unit with a resource ID and attribute information. The input of the computing unit is a resource description derived from software definition, and is used to generate an access address according to the resource ID and perform read and write operations on the graph data entity.
6. The software-hardware collaborative computing framework according to claim 5, characterized in that: The PL side executes the space mapping logic to convert the ID into an access address in the addressable address space. The space mapping logic is: DDR4 =(ID×Size 节点、边 )+Base 节点、边 , where Addr DDR4 is the final address mapped to the DDR4 memory, ID is the number used to identify the data entity, Size 节点、边 Base is the number of bytes occupied by each node or edge. 节点、边 The base address of the starting position of the node or edge data block in the DDR4 memory.
7. The soft-hard collaborative computing framework according to claim 6, characterized in that: The computational unit includes DSP macros that implement multipliers that perform spatial mapping logic.
Citation Information
Patent Citations
Video processing system and method realized by combining software with hardware and device thereof
CN102065288A
Artificial neural network processing device
CN107679621A
Design method for heterogeneous reconfigurable diagram calculation accelerator system on the basis of FPGA (Field Programmable Gate Array)
CN108563808A
Method and system for realizing DMA (Direct Memory Access) read operation and FPGA (Field Programmable Gate Array) equipment
CN117033273A
PCIE DMA data transmission method and system based on FPGA
CN118427135A