Unified memory and data transmission system and management method of FPGA heterogeneous platform

By integrating the memory resources of the CPU and FPGA into a unified virtual address space in the FPGA heterogeneous platform and using the PCIe bus and IP engine for data transmission, the complexity of memory management and data transmission in traditional FPGA acceleration systems is solved, efficient memory management and data transmission are achieved, and system performance and memory utilization are improved.

CN120631827APending Publication Date: 2025-09-12YANGZHOU WANFANG ELECTRONICS TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510726697.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In traditional FPGA-accelerated heterogeneous computing systems, the memory management of the CPU and FPGA cannot be unified, resulting in high programming complexity, large data transmission overhead, low memory usage efficiency, performance bottlenecks and increased latency. The FPGA on-chip memory utilization is low, and the computing power cannot be fully utilized.

Method used

By integrating CPU components and FPGA components into a unified virtual address space, using the PCIe bus for data transmission, using the IP engine to achieve high-speed zero-copy, dual-level TLB for address translation, the FPGA memory management unit ensures data consistency, and the FPGA on-chip memory stack striping to increase bandwidth.

Benefits of technology

It achieves seamless memory address space sharing between CPU and FPGA, simplifies programming model, improves development efficiency, reduces data transmission complexity, improves memory access speed and system performance, and ensures data consistency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631827A_ABST
    Figure CN120631827A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer system structures, in particular to a unified memory and data transmission system and management method for an FPGA heterogeneous platform, and the system comprises a CPU component which is used for detecting and initializing an FPGA component, constructing a globally unified virtual address space view, managing a global page table, maintaining a mapping table, managing PCIe communication and processing exceptions; the FPGA component is used for PCIe DMA data transmission, conversion from a virtual address to a physical address, striped on-chip storage and storage of user acceleration function logic implementation; and the PCIe bus is used for data transmission of the CPU component and the FPGA component. Physical memory resources of the CPU component and the FPGA component are integrated and dynamically mapped to a unified virtual address space in a software and hardware cooperation mode, a simple, safe and flexible unified memory system and management method are provided, and efficient memory management and data transmission between the CPU and the FPGA are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer architecture, and in particular to a unified memory and data transmission system and management method for an FPGA heterogeneous platform. Background Art

[0002] With the continuous growth of computing needs, heterogeneous computing systems, such as CPU + FPGA (Central Processing Unit; Field-Programmable Gate Array), have been widely used in data centers, artificial intelligence, and other fields, and have become an important means to improve system performance. However, in traditional FPGA-accelerated heterogeneous computing systems, the CPU and FPGA typically use a partitioned memory model (Partitioned Memory Model), that is, the CPU and FPGA memories are mutually independent, making it impossible to achieve a unified memory abstraction similar to the CPU for memory management between the FPGA and CPU. In addition, there are some other problems.

[0003] Programming is complex, requiring developers to manually manage data transmission and synchronization between the host and FPGA, which not only increases development difficulty but also reduces program maintainability. Data movement overhead is high, requiring explicit data copying for data exchange between the host and FPGA, resulting in high output transmission overhead and impacting overall computing efficiency. Memory utilization is inefficient, with memory walls and slow memory access speeds limiting performance when processing large amounts of data. Frequent data transmission creates performance bottlenecks, increasing latency and reducing overall system performance. FPGA on-chip memory utilization is low, with on-chip memories such as HBM (High Bandwidth Memory) and DDR (Double Data Rate) not being fully utilized to cache frequently accessed data. Limited bandwidth prevents the FPGA from fully utilizing its computing power. Summary of the Invention

[0004] In light of the above technical issues and shortcomings, the present invention aims to break the barrier between CPU and FPGA memory in heterogeneous computing environments, enabling CPUs and FPGAs to securely and efficiently share memory resources, achieving efficient memory management and data transmission. This invention proposes a unified memory and data transmission system for heterogeneous FPGA platforms. The technical solutions employed are as follows:

[0005] CPU components, FPGA components, and PCIe bus;

[0006] The CPU component is responsible for detecting and initializing the FPGA component, building a global unified virtual address space view, managing the global page table and maintaining the mapping table, managing PCIe communication, and handling exceptions;

[0007] FPGA components for PCIe DMA data transfer, virtual address to physical address translation, striping on-chip storage, and storing user acceleration function logic implementation;

[0008] PCIe bus is used for data transmission between CPU components and FPGA components.

[0009] Preferably, the CPU component includes a kernel driver and a user application and host memory respectively connected to the kernel driver; the FPGA component includes a bus topology and an IP engine, an FPGA memory management unit, an FPGA on-chip memory stack and an FPGA user logic respectively connected to the bus topology, and an FPGA on-chip memory is connected to the FPGA on-chip memory stack. The bus topology transmits data in parallel to the host memory, FPGA memory management unit, FPGA on-chip memory and FPGA user logic through the IP engine.

[0010] Preferably, the IP engine is any one of an XDMA IP engine and a QDMA IP engine.

[0011] Preferably, the FPGA on-chip memory stack includes a DMA controller, a strip distributor and a memory controller arranged in sequence, the memory controller is connected to the FPGA on-chip memory, and the strip distributor evenly divides the large page physical address according to the number of channels, and stripes the FPGA on-chip memory to the available memory controller.

[0012] Preferably, the memory controller is any one of a DDR controller, an HBM controller, or a combination of a DDR controller and an HBM controller, and the DDR controller and / or the HBM controller are respectively connected to their corresponding FPGA on-chip memories.

[0013] Preferably, the FPGA memory management unit includes a dual-level TLB and a read-write engine. The dual-level TLB is used for large pages and small pages respectively, and supports dynamic memory allocation and recycling, dynamic page table loading, permission security verification and page table missing exception handling; the read-write engine is used for memory access by different users.

[0014] Preferably, the FPGA user logic is used to store user-specified and / or user-defined general and / or special hardware acceleration function logic implementations.

[0015] To solve the above problems, the present application further provides: a method for managing a unified memory and data transmission system of an FPGA heterogeneous platform, for managing a unified memory and data transmission system of an FPGA heterogeneous platform as described in any of the above items, the method comprising:

[0016] The kernel driver detects and initializes the FPGA component, integrates the memory resources of the CPU and FPGA components into a unified virtual address space, constructs a unified virtual address space mapping to form a mapping table, loads the initial address mapping configuration, and sets the control registers of the FPGA memory management unit.

[0017] Call memory allocation based on user applications, determine resource usage, allocate physical space in the CPU component and FPGA component, and update the mapping table at one time;

[0018] The user application and FPGA user logic access data to the memory of the CPU component and / or FPGA component according to the updated mapping table.

[0019] The kernel driver migrates data that is not in the corresponding location through the IP engine and PCIe bus based on the memory access pattern;

[0020] The kernel driver handles exceptions generated during data access and data migration;

[0021] After the data processing is completed, the user application reclaims the memory and updates the mapping table again.

[0022] Preferably, the user application and the FPGA user logic access data from the memory of the CPU component and / or the FPGA component according to the updated mapping table, including:

[0023] When accessing data at corresponding locations, the host memory and FPGA on-chip memory are directly accessed through the user application and FPGA user logic respectively;

[0024] When accessing data that is not in the corresponding location, the CPU component accesses the FPGA on-chip memory or the FPGA user logic accesses the host memory. The user application performs the host read or write operation, and the FPGA user logic performs the local read or write operation to achieve data access.

[0025] A cache consistency protocol is determined based on the CPU component and the FPGA component. When a user application accesses the memory of the CPU component and the FPGA component according to an updated mapping table, a synchronization operation is performed according to the cache consistency protocol.

[0026] Preferably, the cache consistency protocol is that when the CPU component or the FPGA component modifies shared data, a consistency message is sent through the PCIe bus to notify the FPGA component or the CPU component to update the cache.

[0027] Preferably, the kernel driver performs data migration on data not in a corresponding location through the IP engine and the PCIe bus according to the memory access mode, including:

[0028] The kernel driver determines the location of the data based on the memory access pattern. If the data is not in the location corresponding to the access request, it triggers data migration through the IP engine and PCIe bus.

[0029] When data is in host memory and the user application performs a local offload operation, the kernel driver migrates the host memory data to the FPGA on-chip memory. Alternatively, when the user application performs a local transfer operation, the kernel driver triggers a bidirectional data migration, transferring the data from the host memory to the FPGA user logic via the IP engine and the PCIe bus. The FPGA user logic then accelerates the data and returns the processed data to the host memory.

[0030] When data is in the FPGA on-chip memory and needs to be returned to the CPU component, the FPGA component sends an interrupt message through the PCIe bus to notify the CPU component that the FPGA on-chip memory data is ready. The kernel driver triggers a local synchronization operation to perform reverse data migration, migrating the data from the FPGA on-chip memory back to the host memory.

[0031] Preferably, the kernel driver handles exceptions generated during data access and data migration, including:

[0032] The exceptions include page miss faults generated in the dual-level TLB, access permission violations, and transmission errors generated in data transmission;

[0033] When a page fault occurs, the FPGA memory management unit sends an interrupt signal to the CPU component through the PCIe bus, notifying the kernel driver to allocate a new physical page, update the mapping table and the two-level TLB, and re-execute the memory access that caused the page fault;

[0034] When access rights are violated, the FPGA memory management unit denies the access request and sends an interrupt signal to the CPU component to notify the kernel driver to handle the issue.

[0035] When a transmission error occurs, the kernel driver performs either data retransmission or error recovery operations.

[0036] Preferably, after data processing is completed, the user application reclaims the memory and updates the mapping table a second time, including:

[0037] After data processing is completed, when the user application no longer uses the allocated memory, the user application calls the memory release function and sends a memory release request to the kernel driver. The kernel driver reclaims the memory and updates the mapping table a second time to mark the released memory as available memory. If the released memory is located in the FPGA component, the kernel driver drives the FPGA memory management unit to reclaim it.

[0038] The present invention has the following beneficial effects:

[0039] 1. Through the CPU component and FPGA component, the physical memory resources of the CPU and FPGA are integrated and dynamically mapped to a unified virtual address space, so that the memory address space between the two components is seamlessly shared, avoiding data copying between kernel space and user space, simplifying the programming model, that is, reducing the complexity of data transmission and improving development efficiency; and through the IP engine in the FPGA component, no explicit data copy is required, achieving high-speed zero-copy and low-latency access to data between the CPU component and the FPGA component; and through the two-level TLB (Translation Lookaside Buffer) in the FPGA component, software and hardware collaboration, and dynamic TLB management, flexible and fast conversion of virtual addresses to physical addresses is achieved, reducing address conversion delays and improving memory access speed; through the stripe allocator of the FPGA on-chip memory stack, the FPGA on-chip memory is striped, maximizing the transmission bandwidth of the FPGA on-chip memory, accelerating sequential access speed and reducing pipeline pauses, and improving the utilization efficiency of the FPGA on-chip memory. The FPGA memory management unit, namely FMMU (FPGAMemoryManagement The CPU and FPGA components do not interfere with each other when accessing shared memory, ensuring data consistency and improving stability and security. The entire system, consisting of the CPU, FPGA, and PCIe (peripheral component interconnect) bus, provides a simple, secure, and flexible unified memory system and management technology for FPGA heterogeneous acceleration environments, enabling efficient memory management and data transfer between the CPU and FPGA.

[0040] 2. The present invention also provides a management method for a unified memory and data transmission system of an FPGA heterogeneous platform, which is used to prepare the unified memory and data transmission system of the FPGA heterogeneous platform provided above. This method has the same beneficial effects as the unified memory and data transmission system of the FPGA heterogeneous platform provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 A structural block diagram of a unified memory and data transmission system for an FPGA heterogeneous platform provided by one embodiment of the present invention;

[0043] Figure 2 A flowchart of a method for managing a unified memory and data transmission system for an FPGA heterogeneous platform provided by one embodiment of the present invention;

[0044] Figure 3 A schematic block diagram of the overall data transmission abstract model structure of the CPU component and the FPGA component of a management method for a unified memory and data transmission system of an FPGA heterogeneous platform provided by one embodiment of the present invention;

[0045] Figure 4 A schematic block diagram of a data access abstract model structure of a method for managing a unified memory and data transmission system for an FPGA heterogeneous platform provided by one embodiment of the present invention;

[0046] Figure 5 A schematic block diagram of the data migration abstract model structure of a method for managing a unified memory and data transmission system of an FPGA heterogeneous platform provided by one embodiment of the present invention. DETAILED DESCRIPTION

[0047] To further illustrate the technical means and effectiveness of the present invention in achieving its intended objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effectiveness of a unified memory and data transmission system and management method for a heterogeneous FPGA platform proposed in accordance with the present invention. In the following description, references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0048] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0049] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, any technician familiar with the art can make many possible changes and modifications without departing from the technical solution of the present invention, or modify them into equivalent embodiments with equivalent changes, which does not affect the essential content of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of protection of the technical solution of the present invention.

[0050] like Figure 1 FIG2 is a block diagram of a unified memory and data transmission system for an FPGA heterogeneous platform provided by an embodiment of the present invention, wherein the system includes: a CPU component, an FPGA component, and a PCIe bus;

[0051] The CPU component is responsible for detecting and initializing the FPGA component, building a global unified virtual address space view, managing the global page table and maintaining the mapping table, managing PCIe communication, and handling exceptions;

[0052] FPGA components for PCIe DMA data transfer, virtual address to physical address translation, striping on-chip storage, and storing user acceleration function logic implementation;

[0053] PCIe bus is used for data transmission between CPU components and FPGA components.

[0054] To better illustrate, in this embodiment, the system constructs a unified memory management architecture for FPGA and CPU through software and hardware collaboration, enabling seamless collaboration between FPGA and CPU components to transparently access and operate memory resources, share memory space, and avoid single memory management issues. It can fully utilize the parallel processing capabilities of the FPGA and the general computing capabilities of the CPU, and the two will not interfere with each other when sharing memory, ensuring the stability and reliability of the entire system while improving overall performance and efficiency. Among them, DMA stands for Direct Memory Access Controller.

[0055] Furthermore, the CPU component includes a kernel driver and a user application and a host memory respectively connected to the kernel driver; the FPGA component includes a bus topology and an IP engine, an FPGA memory management unit, an FPGA on-chip memory stack and an FPGA user logic respectively connected to the bus topology, and an FPGA on-chip memory is connected to the FPGA on-chip memory stack. The bus topology transmits data in parallel to the host memory, the FPGA memory management unit, the FPGA on-chip memory and the FPGA user logic through the IP engine.

[0056] Preferably, the bus topology refers to an AXI4-Stream (Advanced eXtensible Interface) bus topology, which supports high-speed data stream transmission through a flexible interface protocol.

[0057] Furthermore, the FPGA on-chip memory stack includes a DMA controller, a stripe distributor, and a memory controller, which are arranged in sequence. The memory controller is connected to the FPGA on-chip memory. The stripe distributor evenly divides the large page physical addresses according to the number of channels and stripes the FPGA on-chip memory to the available memory controllers. Among them, the FPGA on-chip memory can be used to cache data, store intermediate calculation results, save configuration information, and implement various buffers and queues, etc., which can reduce dependence on external memory and reduce system power consumption and cost.

[0058] Furthermore, the memory controller is any one of a DDR controller, an HBM controller, or a combination of a DDR controller and an HBM controller, and the DDR controller and / or HBM controller are respectively connected to their corresponding FPGA on-chip memories; that is, the DDR controller is connected to the DDR memory, and the HBM controller is connected to the HBM memory. Optionally, the selection and quantity of the DDR controller and / or HBM controller and the corresponding FPGA on-chip memories are determined according to actual needs and applications. The DMA controller directly transfers data between the memory and the peripherals without the intervention of the CPU component to improve data transmission efficiency. DDR memory has a high data transmission rate and low cost, suitable for most general computing scenarios, while HBM memory has extremely high bandwidth and low latency, and is particularly suitable for high-performance computing scenarios that require rapid transmission of large amounts of data. For example, in scenarios requiring rapid transmission of large amounts of data, HBM memory can achieve higher bandwidth and lower latency. In cost-sensitive application scenarios, DDR memory is used to reduce costs. The two types of memory can be configured according to the needs of the specific application scenario to provide high-speed, low-latency storage for intermediate result caching, further improving data transmission efficiency, reducing delays and bottlenecks in the data transmission process, and thereby improving overall system performance.

[0059] Furthermore, the FPGA memory management unit includes a dual-level TLB and a read-write engine. The dual-level TLB is used for large pages and small pages respectively, and supports dynamic memory allocation and recycling, dynamic page table loading, permission security verification and page table missing exception handling; the read-write engine is used for memory access by different users.

[0060] As an optional implementation, in this embodiment, large pages refer to 2M large pages; small pages refer to 4K small pages; it is explained that the read-write engine is used for memory access by different users, that is, to realize the access of user applications or FPGA user logic to memory, specifically including: user applications access FPGA on-chip memory in the form of virtual memory through the IP engine and the read-write engine; FPGA user logic accesses FPGA on-chip memory through the FPGA on-chip memory stack in the form of virtual memory through the read-write engine, and accesses host memory through the IP engine.

[0061] Specifically, the kernel driver is responsible for detecting and initializing the FPGA component, including reading device information, configuring the operating mode of the IP engine, setting bus topology parameters, setting the initial page table of the FPGA memory management unit, and configuring the control registers of the FPGA memory management unit. It is also responsible for building a unified virtual address space view, mapping the physical memory of the CPU component and the FPGA component to a unified virtual address space, providing a transparent memory access interface for user applications and FPGA user logic, and copying data through the kernel driver, so that the two components can operate the CPU and FPGA memory in a unified manner or directly interact with each other. In addition, the kernel driver manages the global page table, maintains a mapping table from virtual addresses to physical addresses, allocates physical space in the host memory or FPGA on-chip memory according to memory requests from user applications and FPGA user logic, and updates the mapping table. It can be understood that the kernel driver of the CPU component is mainly responsible for PCIe bus communication, triggering DMA controller transfers, data migration, and handling TLB page misses or exceptions. The bus topology connects the internal modules of the FPGA component, supporting parallel data transmission to the host memory, FPGA memory management unit, FPGA on-chip memory, and FPGA user logic through the IP engine.

[0062] The FPGA memory management unit converts virtual addresses to physical addresses based on mapping information written by the kernel driver. A dual-level TLB structure with configurable large and small pages allows dynamic resource allocation based on page size, improving conversion efficiency and flexibility. FPGA user logic accesses the FPGA on-chip memory through the FPGA on-chip memory stack in the form of virtual memory via the read / write engine, and accesses host memory through the IP engine. The dual-level TLB in the FPGA memory management unit supports dynamic memory allocation and deallocation, dynamic page table loading, permission security verification, and page table miss exception handling. Optionally, the dual-level TLB also supports multi-level page size management. The stripe allocator in the FPGA on-chip memory stack evenly distributes the physical addresses of large pages based on the number of channels, striping FPGA on-chip memory accesses across all available controllers. This ensures that the data buffer, i.e., the memory area used to temporarily store data in the FPGA component, is equally partitioned across all HBM or DDR memory areas. Channel resources are dynamically allocated through the kernel driver to ensure that each channel fully utilizes bandwidth, thereby maximizing bandwidth, enabling faster sequential access, and reducing pipeline stalls.

[0063] Furthermore, the IP engine is either an XDMA IP engine or a QDMA IP engine; it can be explained that the XDMA (Xilinx Direct Memory Access) IP engine is a direct memory access technology developed by Xilinx, which is used to achieve efficient data transmission without the intervention of the CPU, reducing the burden on CPU components, avoiding the data transmission bottleneck caused by the intervention of traditional CPUs, and making data transmission more efficient and faster; the QDMA (Queue Direct Memory Access Subsystem) IP engine mainly manages data transmission through a queue mechanism, improves the flexibility and efficiency of data transmission, and can process multiple data streams simultaneously by creating multiple data transmission queues, thereby optimizing the performance of data transmission and improving the concurrent processing capability of the system.

[0064] It is explained that the IP engine adopts a DMA transmission mechanism based on AXI Stream flow control, combined with the AXI Stream interface, without the need for explicit data copying, and realizes high-speed, zero-copy, and low-latency access to data between the CPU component and the FPGA component.

[0065] Furthermore, the FPGA user logic is used to store user-specified and / or user-defined general and / or special hardware acceleration function logic implementations.

[0066] As an optional implementation, FPGA user logic refers to the accelerator FPGAAccelerator, which is used to store user-specified general hardware acceleration function logic implementation, user-defined hardware acceleration function logic implementation or dedicated hardware acceleration function logic implementation. It is the acceleration core of the FPGA component and accesses the unified memory through the AXI4-Stream interface to perform computing tasks or acceleration functions.

[0067] It can be understood that the physical memory resources of the CPU and FPGA are integrated and dynamically mapped to a unified virtual address space through the CPU component and FPGA component, so that the memory address space between the two components is seamlessly shared, avoiding data copying between the kernel space and the user space, simplifying the programming model, that is, reducing the complexity of data transmission and improving development efficiency; and through the IP engine in the FPGA component, no explicit data copy is required, achieving high-speed zero-copy and low-latency access to data between the CPU component and the FPGA component; and through the two-level TLB (Translation Lookaside Buffer) in the FPGA component, software and hardware collaboration, and dynamic TLB management, flexible and fast conversion of virtual addresses to physical addresses is achieved, reducing the delay of address conversion and improving memory access speed; the FPGA on-chip memory is striped by the strip allocator of the FPGA on-chip memory stack, maximizing the transmission bandwidth of the FPGA on-chip memory, accelerating the sequential access speed and reducing pipeline pauses, and improving the utilization efficiency of the FPGA on-chip memory, and the FPGA memory management unit, namely FMMU (FPGA Memory Management Unit The CPU and FPGA components do not interfere with each other when accessing shared memory, ensuring data consistency and improving stability and security. The entire system, consisting of the CPU, FPGA, and PCIe (peripheral component interconnect) bus, provides a simple, secure, and flexible unified memory system and management technology for FPGA heterogeneous acceleration environments, enabling efficient memory management and data transfer between the CPU and FPGA.

[0068] like Figure 2 FIG2 is a flowchart illustrating an implementation method for managing a unified memory and data transmission system for an FPGA heterogeneous platform provided in a second embodiment of the present invention. The method is used to manage the unified memory and data transmission system for an FPGA heterogeneous platform provided in the first embodiment of the present invention. The method includes:

[0069] Step S1: Detect and initialize the FPGA component through the kernel driver, integrate the memory resources of the CPU component and the FPGA component into a unified virtual address space, build a unified virtual address space mapping to form a mapping table, load the initial address mapping configuration and set the control register of the FPGA memory management unit;

[0070] Step S2: Memory allocation is called according to the user application, resource usage is determined, physical space is allocated in the CPU component and the FPGA component, and a mapping table is updated once;

[0071] Step S3: The user application and the FPGA user logic access data from the memory of the CPU component and / or the FPGA component according to the updated mapping table;

[0072] Step S4: The kernel driver performs data migration on the data that is not in the corresponding location through the IP engine and the PCIe bus according to the memory access mode;

[0073] Step S5: The kernel driver processes the exceptions generated during data access and data migration;

[0074] Step S6: After the data processing is completed, the user application reclaims the memory and updates the mapping table for the second time.

[0075] It can be explained that step S1 is based on the initialization of the entire system. Specifically, the system is started, the kernel driver detects and identifies the FPGA component, and reads the memory size, address range and other information of the FPGA component through the PCIe bus; then the FPGA component is initialized through the kernel driver, that is, any initialization operation such as configuring the working mode of the IP engine, setting interrupt processing, setting the channel configuration of the DMA controller, setting the parameters of the bus topology, etc. is performed; then a unified virtual address space view is constructed, and the memory resources of the CPU and FPGA are registered to the memory management system of the unified memory and data transmission system of the FPGA heterogeneous platform to construct a unified virtual address space mapping, provide a consistent virtual address view for user applications and FPGA user logic, initialize and establish a mapping table from virtual address to physical address, and prepare for subsequent memory allocation and mapping; and the FPGA memory management unit is processed, the initial address mapping configuration is loaded, and the control register of the FPGA memory management unit is set.

[0076] Step S2 is to allocate memory for the system. Specifically, the user application calls the memory allocation function and specifies the required memory size. The user application only needs to specify the required memory size and does not need to care about the specific location of the CPU memory or FPGA memory, and transmits the call to the memory allocation function to the kernel driver. After receiving the memory allocation request, the kernel driver chooses to allocate physical space in the CPU memory, that is, the host memory, or the FPGA memory, that is, the FPGA on-chip memory, based on resource usage. If there is enough free space in the FPGA on-chip memory and the access mode of the user application is suitable for execution on the FPGA component, the FPGA on-chip memory is allocated first. Then, the kernel driver updates the virtual address to physical address mapping table once and records the virtual address and corresponding physical address of the newly allocated memory. If the allocated memory is located in the CPU component, it is directly allocated. If the allocated memory is located in the FPGA component, the kernel driver writes the relevant mapping information into the FPGA memory management unit through the PCIe bus.

[0077] like Figure 3 and Figure 4 Shown are respectively a schematic block diagram of the overall data transmission abstract model structure of the CPU component and the FPGA component of a management method for a unified memory and data transmission system of an FPGA heterogeneous platform provided by the second embodiment of the present invention, and a schematic block diagram of the data access abstract model structure; wherein, Figures B1 and B2 respectively represent the host read and write modes; Figures B3 and B4 respectively represent the local read and write modes; that is, they show different ways of describing data access in the entire system.

[0078] Optionally, in this embodiment, the user application accesses the memory through a unified virtual address without having to care about the actual storage location of the data.

[0079] Furthermore, step S3 includes:

[0080] Step S31: When accessing data at a corresponding location, the host memory and the FPGA on-chip memory are directly accessed through the user application and the FPGA user logic respectively.

[0081] It is explained that data is accessed at corresponding locations, that is, when the data is located in the host memory and the FPGA on-chip memory respectively, when the address accessed by the user application is mapped to the host memory, the CPU component directly accesses the local memory; when the address accessed by the FPGA user logic is mapped to the FPGA on-chip memory, the FPGA user logic accesses the FPGA on-chip memory through the memory controller.

[0082] Step S32: When accessing data that is not in the corresponding location, the CPU component accesses the FPGA on-chip memory or the FPGA user logic accesses the host memory, the user application performs a host read or write operation, and the FPGA user logic performs a local read or write operation to achieve data access.

[0083] To clarify, data that is not in the corresponding position means that the location of the data is not on the side that needs to be accessed. When the CPU component or the FPGA component is used, the CPU component needs to access the FPGA on-chip memory or the FPGA user logic needs to access the host memory. Specifically, when the CPU component accesses data to the FPGA on-chip memory, the user application performs a host read or write operation, and the CPU component sends the access request to the FPGA component through the PCIe bus. Then the FPGA memory management unit converts the virtual address into a physical address according to an updated mapping table, and then accesses the FPGA on-chip memory through the FPGA on-chip memory stack through the XDMA or QDMA IP engine and the PCIe bus without page copying; when the FPGA user logic accesses the host memory, the kernel driver identifies the data location. When the data is in the host memory, the FPGA user logic performs a local read or write operation, and the kernel driver triggers data transmission to realize the FPGA user logic's access to the host memory.

[0084] Step S33: a cache coherence protocol is determined based on the CPU component and the FPGA component. When the user application accesses the memory of the CPU component and the FPGA component according to the updated mapping table, a synchronization operation is performed according to the cache coherence protocol.

[0085] Furthermore, the cache coherence protocol is that when the CPU component or the FPGA component modifies shared data, a coherence message is sent via the PCIe bus to notify the FPGA component or the CPU component to update the cache.

[0086] It can be explained that shared data refers to data that can be accessed by the CPU component and the FPGA component together. Specifically, if shared data is accessed, the entire system performs synchronization operations according to the cache consistency protocol to ensure data consistency. In addition, if an exception occurs during the data access process, the exception handling process is entered.

[0087] Specifically, when either the CPU component or the FPGA component modifies shared data, a consistency message is sent via the PCIe bus. For example, when the CPU component modifies a data block located in the FPGA on-chip memory, an invalidation message is sent to the FPGA. After receiving this invalidation message, the FPGA component marks the corresponding cache line as invalid to ensure that the data obtained the next time it is accessed is the latest data. Conversely, when the FPGA component modifies shared data, the FPGA component will also notify the CPU component to update the cache, thereby ensuring that the entire system guarantees consistent access to shared data by the CPU and FPGA components based on cache consistency.

[0088] like Figure 5 The figure shows a schematic block diagram of the data migration abstract model structure of a method for managing a unified memory and data transmission system of an FPGA heterogeneous platform provided by the second embodiment of the invention, wherein Figure C1 represents the local offloading mode, i.e., migrating data from the host memory to the FPGA on-chip memory; Figure C2 represents the local synchronization mode, i.e., migrating data from the FPGA on-chip memory to the host memory; and Figure C3 represents the local transmission mode, i.e., bidirectional data migration between the host memory and the FPGA user logic.

[0089] Furthermore, step S4 includes:

[0090] Step S41: The kernel driver determines the location of the data according to the memory access mode. If the data is not at the location corresponding to the access request, the kernel driver triggers data migration through the IP engine and the PCIe bus.

[0091] Step S42: When the data is in the host memory, the user application performs a local unload operation, and the kernel driver migrates the host memory data to the FPGA on-chip memory; or the user application performs a local transfer operation, and the kernel driver triggers bidirectional data migration, transferring the data in the host memory to the FPGA user logic through the IP engine and the PCIe bus. The FPGA user logic accelerates the data processing and returns the processed data to the host memory.

[0092] Specifically, the user application initiates a local offload access, and the kernel driver sends a data migration instruction to the XDMA or QDMA IP engine through the PCIe bus, specifying the data address and target address to be migrated, that is, the FPGA on-chip memory. After receiving the data migration instruction, the XDMA or QDMA IP engine adopts a flow control-based DMA transmission mechanism combined with the AXIStream interface to transfer the data from the host memory to the FPGA on-chip memory. During the transmission process, the bus topology is responsible for coordinating the parallel transmission of data to improve transmission efficiency; or, the user application performs a local transmission operation, and the kernel driver triggers bidirectional migration of data, that is, the user application initiates a local transmission access, the kernel driver triggers data migration, and the data in the host memory is passed to the FPGA user logic through the unified virtual memory for streaming processing. After the FPGA user logic completes the accelerated processing of the data, the processing results are returned to the host memory. The bidirectional migration process enables the FPGA to fully exert its parallel processing capabilities while maintaining close cooperation with the CPU components.

[0093] Step S43: When the data is in the FPGA on-chip memory and needs to be returned to the CPU component, the FPGA component sends an interrupt message to the CPU component through the PCIe bus to notify the CPU component that the FPGA on-chip memory data is ready. The kernel driver triggers a local synchronization operation to perform reverse data migration, migrating the data in the FPGA on-chip memory back to the host memory.

[0094] Specifically, when the data is in the FPGA on-chip memory and needs to be returned to the CPU component, an interrupt message is sent through the PCIe bus to notify the CPU component that the FPGA on-chip memory data is ready. After receiving the interrupt message, the kernel driver initiates a reverse data migration request, that is, local synchronous access, to transfer the data from the FPGA on-chip memory to the host memory.

[0095] It can be explained that different migration processes are performed on the data. After the data migration is completed, the kernel driver updates the memory mapping table for the second time to ensure that subsequent memory accesses can correctly point to the new data location.

[0096] Furthermore, step S5 includes:

[0097] Exceptions include page miss faults generated in the dual-level TLB, access permission violations, and transfer errors generated during data transfer;

[0098] When a page fault occurs, the FPGA memory management unit sends an interrupt signal to the CPU component through the PCIe bus, notifying the kernel driver to allocate a new physical page, update the mapping table and the two-level TLB, and re-execute the memory access that caused the page fault;

[0099] When access rights are violated, the FPGA memory management unit denies the access request and sends an interrupt signal to the CPU component to notify the kernel driver to handle the issue.

[0100] When a transmission error occurs, the kernel driver performs either data retransmission or error recovery operations.

[0101] Specifically, when the FPGA memory management unit is performing address translation and the corresponding mapping is not found in the page table, a page fault exception is triggered. The FPGA memory management unit then sends an interrupt signal to the CPU component via the PCIe bus, notifying the kernel driver to allocate a new physical page, update the mapping table and the two-level TLB, and then re-execute the memory access operation that caused the page fault exception. When the FPGA memory management unit detects that the access request permissions, namely read, write, and execute permissions, are illegal, a permission violation exception is triggered. The FPGA memory management unit then rejects the access request and sends an interrupt signal to the CPU component. The kernel driver then handles the situation according to the specific situation, such as terminating the offending application or logging operations, to prevent further illegal operations, facilitate subsequent analysis and debugging, and effectively protect itself from unauthorized access and potential security threats. During data migration or data transfer, if errors such as DMA controller transmission failure and checksum mismatch occur, a data transfer error exception is triggered. The kernel driver then attempts to retransmit the data or perform error recovery operations. If the problem cannot be solved after multiple attempts, the error information is recorded and the user is notified.

[0102] Furthermore, step S6 includes:

[0103] After data processing is completed, when the user application no longer uses the allocated memory, the user application calls the memory release function and sends a memory release request to the kernel driver. The kernel driver reclaims the memory and updates the mapping table a second time to mark the released memory as available memory. If the released memory is located in the FPGA component, the kernel driver drives the FPGA memory management unit to reclaim it.

[0104] Specifically, after data processing is completed and the user application no longer needs to use the allocated memory, the user application calls the memory release function to request the release of memory. After receiving the memory release request, the kernel driver updates the mapping table a second time and marks the released memory as "available" so that the system can reallocate these memory resources to other processes or applications in need; if the released memory is located in the FPGA component, the kernel driver will send the corresponding instructions and data to the FPGA memory management unit through the PCIe bus, and recycle the memory to the free pool, which is conducive to the reallocation and utilization of memory, thereby improving overall resource utilization and operating efficiency.

[0105] It can be understood that this management method simplifies the development process and the programming burden of developers, improves memory utilization, system performance and resource utilization, reduces data transmission delays, ensures that CPU components and FPGA components can share memory resources safely and efficiently, and realizes efficient memory management and data transmission. In addition, the present application can be further expanded. For example, in a heterogeneous computing system containing a CPU, a GPU (Graphics Processing Unit, i.e., a graphics processor), and an FPGA, data access and transmission between FPGA memory resources and GPU memory resources can be achieved by bypassing the CPU; or data access and transmission between local memory resources and remote node memory resources can be achieved through a network interface, such as TCP / IP (Transmission Control Protocol / Internet Protocol, i.e., Transmission Control Protocol / Internet Protocol) and RDMA (Remote Direct Memory Access, i.e., remote direct data access).

[0106] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and accompanying drawings, this does not require or imply that the operations must be performed in this specific order, or that all illustrated operations must be performed to achieve the desired results. Therefore, the order of the steps in the embodiments should not be considered a limitation of the present invention. Preferably or optionally, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

Claims

1. A unified memory and data transmission system for FPGA heterogeneous platforms, characterized in that: The system includes: a CPU component, an FPGA component and a PCIe bus; The CPU component is responsible for detecting and initializing the FPGA component, building a global unified virtual address space view, managing the global page table and maintaining the mapping table, managing PCIe communication, and handling exceptions; FPGA components for PCIe DMA data transfer, virtual address to physical address translation, striping on-chip storage, and storing user acceleration function logic implementation; PCIe bus is used for data transmission between CPU components and FPGA components.

2. The unified memory and data transmission system for FPGA heterogeneous platforms according to claim 1, characterized in that: The CPU component includes a kernel driver and a user application and host memory respectively connected to the kernel driver; the FPGA component includes a bus topology and an IP engine, an FPGA memory management unit, an FPGA on-chip memory stack and an FPGA user logic respectively connected to the bus topology, and an FPGA on-chip memory is connected to the FPGA on-chip memory stack. The bus topology transmits data in parallel to the host memory, FPGA memory management unit, FPGA on-chip memory and FPGA user logic through the IP engine.

3. The unified memory and data transmission system for FPGA heterogeneous platforms according to claim 2, characterized in that: The IP engine is either an XDMAIP engine or a QDMAIP engine.

4. The unified memory and data transmission system for FPGA heterogeneous platforms according to claim 2, characterized in that: The FPGA on-chip memory stack includes a DMA controller, a stripe distributor and a memory controller arranged in sequence. The memory controller is connected to the FPGA on-chip memory. The stripe distributor evenly divides the large page physical address according to the number of channels and stripes the FPGA on-chip memory to the available memory controllers.

5. The unified memory and data transmission system for FPGA heterogeneous platforms according to claim 4, characterized in that: The memory controller is any one of a DDR controller, an HBM controller, or a combination of a DDR controller and an HBM controller, and the DDR controller and / or the HBM controller are respectively connected to their corresponding FPGA on-chip memories.

6. The unified memory and data transmission system for FPGA heterogeneous platforms according to claim 4, characterized in that: The FPGA memory management unit includes a dual-level TLB and a read-write engine. The dual-level TLB is used for large pages and small pages respectively, and supports dynamic memory allocation and recycling, dynamic page table loading, permission security verification and page table missing exception handling; the read-write engine is used for memory access by different users.

7. The unified memory and data transmission system for FPGA heterogeneous platforms according to claim 2, characterized in that: The FPGA user logic is used to store user-specified and / or user-defined general and / or special hardware acceleration function logic implementations.

8. A method for managing a unified memory and data transmission system for an FPGA heterogeneous platform, characterized in that: A unified memory and data transmission system for managing an FPGA heterogeneous platform according to any one of claims 1 to 7, the method comprising: The kernel driver detects and initializes the FPGA component, integrates the memory resources of the CPU and FPGA components into a unified virtual address space, constructs a unified virtual address space mapping to form a mapping table, loads the initial address mapping configuration, and sets the control registers of the FPGA memory management unit. Call memory allocation based on user applications, determine resource usage, allocate physical space in the CPU component and FPGA component, and update the mapping table at one time; The user application and FPGA user logic access data to the memory of the CPU component and / or FPGA component according to the updated mapping table. The kernel driver migrates data that is not in the corresponding location through the IP engine and PCIe bus based on the memory access pattern; The kernel driver handles exceptions generated during data access and data migration; After the data processing is completed, the user application reclaims the memory and updates the mapping table again.

9. The method for managing a unified memory and data transmission system of an FPGA heterogeneous platform according to claim 8, characterized in that: The user application and FPGA user logic access data from the memory of the CPU component and / or FPGA component based on the updated mapping table, including: When accessing data at corresponding locations, the host memory and FPGA on-chip memory are directly accessed through the user application and FPGA user logic respectively; When accessing data that is not in the corresponding location, the CPU component accesses the FPGA on-chip memory or the FPGA user logic accesses the host memory. The user application performs the host read or write operation, and the FPGA user logic performs the local read or write operation to achieve data access. A cache consistency protocol is determined based on the CPU component and the FPGA component. When a user application accesses the memory of the CPU component and the FPGA component according to an updated mapping table, a synchronization operation is performed according to the cache consistency protocol.

10. The method for managing a unified memory and data transmission system of an FPGA heterogeneous platform according to claim 9, characterized in that: The cache consistency protocol is that when the CPU component or the FPGA component modifies shared data, a consistency message is sent through the PCIe bus to notify the FPGA component or the CPU component to update the cache.

11. The method for managing a unified memory and data transmission system of an FPGA heterogeneous platform according to claim 8, characterized in that: The kernel driver migrates data that is not in the corresponding location through the IP engine and PCIe bus based on the memory access pattern, including: The kernel driver determines the location of the data based on the memory access pattern. If the data is not in the location corresponding to the access request, it triggers data migration through the IP engine and PCIe bus. When data is in host memory and the user application performs a local offload operation, the kernel driver migrates the host memory data to the FPGA on-chip memory. Alternatively, when the user application performs a local transfer operation, the kernel driver triggers a bidirectional data migration, transferring the data from the host memory to the FPGA user logic via the IP engine and the PCIe bus. The FPGA user logic then accelerates the data and returns the processed data to the host memory. When data is in the FPGA on-chip memory and needs to be returned to the CPU component, the FPGA component sends an interrupt message through the PCIe bus to notify the CPU component that the FPGA on-chip memory data is ready. The kernel driver triggers a local synchronization operation to perform reverse data migration, migrating the data from the FPGA on-chip memory back to the host memory.

12. The method for managing a unified memory and data transmission system of an FPGA heterogeneous platform according to claim 8, wherein: The kernel driver handles exceptions generated during data access and data migration, including: The exceptions include page miss faults generated in the dual-level TLB, access permission violations, and transmission errors generated in data transmission; When a page fault occurs, the FPGA memory management unit sends an interrupt signal to the CPU component through the PCIe bus, notifying the kernel driver to allocate a new physical page, update the mapping table and the two-level TLB, and re-execute the memory access that caused the page fault; When access rights are violated, the FPGA memory management unit denies the access request and sends an interrupt signal to the CPU component to notify the kernel driver to handle the issue. When a transmission error occurs, the kernel driver performs either data retransmission or error recovery operations.

13. The method for managing a unified memory and data transmission system of an FPGA heterogeneous platform according to claim 8, characterized in that: After data processing is completed, the user application reclaims the memory and updates the mapping table again, including: After data processing is completed, when the user application no longer uses the allocated memory, the user application calls the memory release function and sends a memory release request to the kernel driver. The kernel driver reclaims the memory and updates the mapping table a second time to mark the released memory as available memory. If the released memory is located in the FPGA component, the kernel driver drives the FPGA memory management unit to reclaim it.

Citation Information

Cited By

  • Heterogeneous computing system and heterogeneous computing method of computer

    CN122111696A

  • Heterogeneous computing systems and methods for computers

    CN122111696B