Method, device and equipment for verifying HBM storage function on V19P and storage medium
By instantiating the self-encapsulated Aurora IP core on V19P and using XDMA drivers, the optical port communication and HBM functions between V19P and the VCU128 development board are realized, solving the problem that traditional verification platforms cannot support hyper-large HBM channels, and achieving high bandwidth and low latency storage system performance.
Patent Information
- Application Number
- CN202510510026.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
AI Technical Summary
On high-end FPGA platforms such as Xilinx V19P, how to effectively integrate and verify high-bandwidth memory (HBM) functions to solve the problem that traditional verification platforms cannot support hyper-large-scale HBM channels.
By instantiating the self-encapsulated Aurora IP core on V19P, optical port communication between V19P and the VCU128 development board is realized, and combined with the XDMA driver and the AXI4-Stream Interconnect IP core, the transmission and verification of HBM write requests and read requests are realized.
It realizes seamless connection between high-end FPGA V19P logic resources and HBM, supports ultra-large-scale HBM channels, with a single channel bandwidth greater than 32GB/s, and a total bandwidth up to 460GB/s, which significantly improves the performance of the storage system and meets the real-time processing needs of massive data in scenarios such as AI inference and genomics.
Smart Images

Figure CN120045400A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of high - bandwidth memory, and particularly to a method, device, equipment, and storage medium for verifying the HBM storage function on V19P. Background Art
[0002] With the rapid development of information technology, especially the growing demands in fields such as big data analysis, artificial intelligence (AI), cloud computing, and high - performance computing (HPC), the requirements for storage systems are also constantly increasing. Although traditional DDR memory has certain advantages in terms of cost and compatibility, it is gradually unable to meet the needs of modern applications in terms of bandwidth, latency, and power consumption.
[0003] High Bandwidth Memory (HBM) has become an ideal choice for application fields such as data centers and artificial intelligence accelerators due to its excellent performance. By stacking multiple DRAM chips and directly connecting them to a processor or FPGA, HBM greatly improves the data transfer rate and reduces power consumption. Compared with traditional memory architectures, HBM can provide data transfer speeds several times higher than traditional DDR memory, which is crucial for applications that need to process large amounts of data.
[0004] Due to its compact design and short - distance data transfer path, HBM can complete data exchange in a shorter time, reducing waiting time. In addition, when designing HBM, energy efficiency issues are fully considered, and it consumes less power under the same performance conditions compared with traditional solutions.
[0005] However, HBM has a complex structure and technical requirements. For high - end FPGA platforms such as Xilinx V19P, V19P is a high - end FPGA chip launched by Xilinx in 2019. Although V19P provides rich logic resources and I / O interfaces, how to effectively integrate and verify the function of HBM on V19P still remains a difficult problem. Summary of the Invention
[0006] In order to effectively integrate and verify the function of HBM on V19P, this application provides a method for verifying the HBM storage function on V19P.
[0007] In a first aspect, this application provides a method for verifying the HBM storage function on V19P. The method is based on VeriTiger - V19P, VCU128 development board, Aurora IP core, XDMA, XDMA driver, Aurora Lane, AXI Interconnect, and a host computer, and adopts the following technical solutions: Instantiate the self - encapsulated Aurora IP core based on the VeriTiger - V19P to obtain the V19P Aurora Core, and instantiate the self - encapsulated Aurora IP core based on the VCU128 development board to obtain the VCU128 Aurora Core; Receive the HBM write request sent by the host computer, where the HBM write request includes the first file data and the storage information of the first file data; Transmit the first file data to the HBM through the V19P Aurora Core and the VCU128 Aurora Core according to the HBM write request; Initiate an HBM read request according to the storage information; Read the HBM according to the HBM read request to obtain the second file data, and transmit the second file data to the host computer through the VCU128 Aurora Core and the V19P Aurora Core; Compare the first file data and the second file data to verify the HBM storage function.
[0008] Through the above technical solution, the V19P and the VCU128 are interconnected and communicate based on the optical port and the self - encapsulated Aurora IP, achieving seamless docking between the logic resources of the high - end FPGA V19P and the HBM, solving the problem that the traditional verification platform cannot support ultra - large - scale HBM channels, and through the HBM multi - channel parallel access architecture, achieving a continuous bandwidth of more than 32GB / s per channel and a total bandwidth of up to 460GB / s, which is more than 18 times higher than the traditional DDR4 memory, significantly improving the storage system performance of the V19P and meeting the real - time processing requirements of massive data in scenarios such as AI inference and genomics.
[0009] In a specific feasible implementation, the instantiating the self - encapsulated Aurora IP core based on the VeriTiger - V19P to obtain the V19P Aurora Core and instantiating the self - encapsulated Aurora IP core based on the VCU128 development board to obtain the VCU128 Aurora Core includes: Instantiate the self - encapsulated Aurora IP core based on the VeriTiger - V19P to obtain the V19P Aurora Core, and instantiate the self - encapsulated Aurora IP core based on the VCU128 development board to obtain the VCU128 Aurora Core; Set the user interfaces of the V19P Aurora Core and the VCU128 Aurora Core as AXI4-Stream interfaces, and adopt an AXI4-Stream Interconnect IP core according to the AXI4-Stream interfaces.
[0010] Through the above technical solution, according to the method of self-packaging IP, self-package the IP cores V19P Aurora Core and VCU128 Aurora Core based on the V19P and VCU128 platforms respectively, and use these IP cores to realize data communication based on optical ports and the Aurora protocol between the V19P and VCU128 platforms, laying a foundation for the subsequent V19P to read and write HBM. The self-packaged Aurora IP can realize a user-adaptive usage mode, improving the simplicity and usability of the Aurora IP.
[0011] In a specific feasible implementation, the process of transmitting the first file data to the HBM through the V19P Aurora Core and the VCU128 Aurora Core according to the HBM write request includes: Use the XDMA driver to drive the first XDMA to transmit the first file data to the V19P Aurora Core through the AXI4-Stream Interconnect IP core; Transmit the first file data in the V19P Aurora Core to the VCU128 Aurora Core through the Aurora Lane; Transmit the first file data in the VCU128 Aurora Core to the host computer through the AXI4-Stream Interconnect IP core and the second XDMA; Use the XDMA driver to drive the second XDMA to write the first file data into the HBM.
[0012] Through the above technical solution, for writing file data to the HBM on the V19P, the V19P and the VCU128 are interconnected and communicate based on optical ports and self-packaged Aurora IP, realizing seamless docking between the high-end FPGA V19P logic resources and the HBM, and solving the problem that the traditional verification platform cannot support ultra-large-scale HBM channels.
[0013] In a specific feasible implementation, the process of using the XDMA driver to drive the second XDMA to write the first file data into the HBM includes: Set the second XDMA to Bypass mode, where the Bypass mode includes a Bypass path and a Bypass interface; Use the AXI Interconnect to convert the Bypass interface into an AXI interface; Use the XDMA driver to drive the second XDMA to write the first file data into the HBM through the Bypass path and the AXI interface.
[0014] Through the above technical solution, a Bypass interface is added. The Bypass interface can bypass DMA, that is, it can directly read and write the Bar space like a normal PCIE. An AXI Interconnect is added to convert the Bypass path of the XDMA into an AXI4 interface, enabling the smooth transmission of file data.
[0015] In a specific feasible implementation, the step of reading the HBM according to the HBM read request to obtain the second file data and transmitting the second file data to the host computer through the VCU128 Aurora Core and the V19P Aurora Core includes: Use the XDMA driver to drive the second XDMA to read the HBM to obtain the second file data; Transmit the second file data to the VCU128 Aurora Core through the AXI4-Stream Interconnect IP core; Transmit the second file data in the VCU128 Aurora Core to the V19P Aurora Core through the Aurora Lane; Output the second file data in the V19P Aurora Core to the host computer through the AXI4-Stream Interconnect IP core and the first XDMA.
[0016] Through the above technical solution, for reading file data from the HBM on the V19P, the V19P and the VCU128 are interconnected and communicate based on the optical port and the self-packaged Aurora IP, achieving seamless docking between the high-end FPGA V19P logic resources and the HBM, and solving the problem that the traditional verification platform cannot support ultra-large-scale HBM channels.
[0017] In a specific feasible implementation, the step of using the XDMA driver to drive the second XDMA to read the HBM to obtain the second file data includes: Set the second XDMA to Bypass mode, where the Bypass mode includes a Bypass path and a Bypass interface; Use the AXI Interconnect to convert the Bypass interface to an AXI interface; Use the XDMA driver to drive the second XDMA to read the HBM through the Bypass path and the AXI interface to obtain second file data.
[0018] Through the above technical solution, a Bypass interface is added. The Bypass interface can bypass DMA, that is, it can directly perform read and write operations on the Bar space like a common PCIE. An AXI Interconnect is added to convert the Bypass path of the XDMA into an AXI4 interface, enabling the smooth transmission of file data.
[0019] In a specific feasible implementation, the verifying the HBM storage function by comparing the first file data and the second file data includes: Compare whether the first file data and the second file data are consistent; If the first file data and the second file data are consistent, it is determined that the HBM storage function verification passes; Otherwise, it is determined that the HBM storage function verification fails.
[0020] Through the above technical solution, the same file data is written and read in the HBM on V19P, and by comparing the original file data and the file data after reading and writing, the read and write verification of the HBM on the VCU128 is realized.
[0021] In a second aspect, the present application provides a device for verifying the HBM storage function on V19P, which is used to implement the above method for verifying the HBM storage function on V19P, and adopts the following technical solution: The device includes: An Aurora IP core instantiation module, which is used to instantiate the self - encapsulated Aurora IP core based on the VeriTiger - V19P to obtain a V19P Aurora Core, and instantiate the self - encapsulated Aurora IP core based on the VCU128 development board to obtain a VCU128 Aurora Core; An HBM write request receiving module, which is used to receive an HBM write request, and the HBM write request includes first file data and storage information of the first file data; A file data writing module, configured to transmit the first file data to the HBM through the V19P Aurora Core and the VCU128 Aurora Core according to the HBM write request; An HBM read request initiating module, configured to initiate an HBM read request according to the storage information; A file data reading module, configured to read the HBM according to the HBM read request to obtain second file data, and transmit the second file data to the host computer through the VCU128 Aurora Core and the V19P Aurora Core; An HBM storage function verification module, configured to verify the HBM storage function by comparing the first file data and the second file data.
[0022] In a third aspect, the present application provides a computer device, adopting the following technical solution: including a memory and a processor, and a computer program capable of being loaded and executed by the processor, such as the method for verifying the HBM storage function on the V19P as described above, is stored on the memory.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: storing a computer program capable of being loaded and executed by the processor, such as the method for verifying the HBM storage function on the V19P as described above.
[0024] In summary, the present application has the following beneficial technical effects: (1) By implementing the interconnection and communication between the V19P and the VCU128 based on the optical port and the self-packaged Aurora IP, the present design provides a high-performance, low-latency, and high-reliability heterogeneous platform communication solution, which can achieve ultra-high bandwidth and long-distance transmission. The self-packaged Aurora IP can realize a user-adaptive usage mode, improving the simplicity and usability of the Aurora IP.
[0025] (2) The interconnection and communication between the V19P and the VCU128 based on the optical port and the self-packaged Aurora IP realizes the seamless docking of the high-end FPGA V19P logic resources and the HBM, solves the problem that the traditional verification platform cannot support ultra-large-scale HBM channels, and through the HBM multi-channel parallel access architecture, realizes a continuous bandwidth of more than 32GB / s per channel, and the total bandwidth can reach 460GB / s, which is more than 18 times higher than that of the traditional DDR4 memory, significantly improving the storage system performance of the V19P and meeting the real-time processing requirements of massive data in scenarios such as AI inference and genomics. Description of the Drawings
[0026] Figure 1 It is the design framework for verifying the HBM storage function.
[0027] Figure 2 It is a flowchart of a method for verifying the HBM storage function on V19P in an embodiment of the present application.
[0028] Figure 3 It is a schematic diagram of a self - encapsulated IP.
[0029] Figure 4 It is a structural block diagram of a method for verifying the HBM storage function on V19P in an embodiment of the present application.
[0030] Reference numerals: 401, Aurora IP core instantiation module; 402, HBM write request receiving module; 403, file data writing module; 404, HBM read request initiating module; 405, file data reading module; 406, HBM storage function verification module. Detailed implementation manners
[0031] The following further elaborates on the present application Figures 1 - 4 with reference to the accompanying drawings.
[0032] An embodiment of the present application discloses a method for verifying the HBM storage function on V19P, which is used to effectively integrate and verify the function of HBM on V19P.
[0033] With the rapid development of information technology, especially the growing demands in fields such as big data analysis, artificial intelligence (AI), cloud computing, and high - performance computing (HPC), the requirements for storage systems are also constantly increasing. Although traditional DDR memory has certain advantages in terms of cost and compatibility, it gradually fails to meet the needs of modern applications in terms of bandwidth, latency, and power consumption.
[0034] High Bandwidth Memory (HBM) has become an ideal choice for application fields such as data centers and artificial intelligence accelerators due to its excellent performance. By stacking multiple DRAM chips and directly connecting them to a processor or FPGA, HBM greatly improves the data transfer rate and reduces power consumption. Compared with traditional memory architectures, HBM can provide data transfer speeds several times higher than traditional DDR memory, which is crucial for applications that need to process large amounts of data.
[0035] Due to its compact design and short - distance data transfer path, HBM can complete data exchange in a shorter time, reducing waiting time. In addition, when designing HBM, energy efficiency issues are fully considered, and it consumes less power under the same performance conditions compared with traditional solutions.
[0036] However, HBM has a complex structure and technical requirements. For high-end FPGA platforms such as Xilinx V19P, which is a high-end FPGA chip launched by Xilinx in 2019. Although V19P provides rich logic resources and I / O interfaces, how to effectively integrate and verify the functions of HBM on V19P still remains a difficult problem.
[0037] Therefore, this application proposes a method for verifying the HBM storage function on V19P, and uses this method to effectively integrate and verify the functions of HBM on V19P.
[0038] Before specifically elaborating on the method, it is necessary to first elaborate on the design framework for verifying the HBM storage function, such as Figure 1 shown below: Host Program: The host service program, which sends control commands and data processing tasks. It mainly communicates with the XDMA controller through the XDMADriver to implement the read and write tasks of a large amount of data with files as parameter inputs initiated by the user side.
[0039] XDMA Driver: The XDMA driver provided by Xilinx official manages the DMA (Direct Memory Access) operations between the PC and the FPGA, bypasses the CPU to directly read and write memory, and provides software interfaces (such as Linux kernel modules or Windows drivers) for user programs to call. It mainly includes several commonly used device files such as xdma0_events_x, xdma0_h2c_x / xdma0_c2h_x, xdma0_user, and xdma0_bypass. The corresponding application programming is to operate these device files to implement the corresponding business logic.
[0040] XDMA0 / XDMA1: The high-performance, configurable SG-mode DMA provided by Xilinx for PCIE2.0 and PCIE3.0. It provides user-selectable AXI4-full, AXI4-Lite, and AXI4-Stream interfaces, and can be set to Bypass mode to directly bypass the DMA, that is, it can directly read and write the Bar space like a normal PCIE. The read and write operations on the Bar space will then be converted into AXI read and write operations for subsequent modules.
[0041] AXIS0 / AXIS1 / AXIS2 / AXIS3: All are AXI4-Stream Interconnect IP cores. Since the user interfaces of both the V19P Aurora Core and the VCU128 Aurora Core are AXI4-Stream interfaces, AXI4-Stream Interconnect IP cores are used to connect XDMA0 / XDMA1 with the V19P Aurora Core / VCU128 Aurora Core.
[0042] V19P Aurora Core: Instantiate the Aurora 64B / 66B IP core based on the V19P platform, and self-package this IP core into the V19P Aurora Core IP core. Through this IP core, optical port communication between V19P and VCU128 is achieved.
[0043] VCU128 Aurora Core: Instantiate the Aurora 64B / 66B IP core based on the VCU128 platform, and self-package this IP core into the VCU128 Aurora Core IP core. Through this IP core, optical port communication between VCU128 and V19P is achieved.
[0044] AXI4: AXI Interconnect is the core IP used to manage the AXI (Advanced eXtensible Interface) bus interconnection in Xilinx FPGA designs. It is responsible for connecting multiple AXI masters and slaves, and implementing functions such as address routing, protocol conversion, and clock domain isolation.
[0045] HBM: HBM is a 3D stacked high-bandwidth memory integrated in high-end Xilinx FPGAs (such as the Virtex UltraScale+ HBM series), providing throughput capabilities far exceeding those of traditional DDR (such as 460GB / s). The HBM IP core is used to manage the read / write access, channel allocation, and interface protocol conversion of HBM, and is a core component in scenarios such as high-performance computing and AI acceleration.
[0046] As Figure 1 shown, the method includes: S10, instantiate and self-package the Aurora IP core based on VeriTiger-V19P to obtain the V19P Aurora Core, and instantiate and self-package the Aurora IP core based on the VCU128 development board to obtain the VCU128 Aurora Core.
[0047] Specifically, instantiate the Aurora 64B / 66B IP core based on the V19P platform, self-encapsulate this IP core into the V19P Aurora Core IP core, and implement optical port communication between V19P and VCU128 through this IP core; instantiate the Aurora 64B / 66B IP core based on the VCU128 platform, self-encapsulate this IP core into the VCU128 Aurora Core IP core, and implement optical port communication between VCU128 and V19P through this IP core.
[0048] S20. Receive the HBM write request sent by the host computer. The HBM write request includes the first file data and the storage information of the first file data.
[0049] Specifically, for the read and write verification of the HBM on VCU128, it is necessary to write to and read from the HBM. Use VeriTiger-V19P and VCU128 for interconnection communication to implement the writing and reading of file data to and from the HBM. Write the host computer program in the Host PC (i.e., the client). In the host computer program, use the xdma0_h2c_x device file provided by the XDMA driver to initiate the HBM write request for the file data. The HBM write request includes the first file data and the storage information of the first file data. The storage information includes information such as the storage location of the first file data in the HBM.
[0050] S30. Transmit the first file data to the HBM through the V19P Aurora Core and the VCU128 Aurora Core according to the HBM write request.
[0051] Specifically, in the software protocol of the method of this application, use the self-encapsulated Aurora IP core to implement the interconnection communication between VeriTiger-V19P and VCU128, so as to transmit the first file data from VeriTiger-V19P to the HBM located on VCU128 through the V19P Aurora Core and the VCU128 Aurora Core, thereby completing the writing of the file data to the HBM.
[0052] S40. Initiate an HBM read request according to the storage information.
[0053] Specifically, for the read and write verification of the HBM on VCU128, it is necessary to write to and read from the HBM. Use VeriTiger-V19P and VCU128 for interconnection communication to implement the writing and reading of file data to and from the HBM. After completing the writing of the file data to the HBM, it is also necessary to read the first file data.
[0054] First, initiate an HBM read request according to the stored information, and read the file in the HBM according to the stored information of the first file data. The upper computer program in the Host PC can directly operate the PCIE Bar space using the Bypass path of XDMA1 to perform a read operation on the VCU128 HBM, and save the read data to a file.
[0055] S50, read the HBM according to the HBM read request to obtain the second file data, and transmit the second file data to the upper computer through the VCU128 Aurora Core and the V19P Aurora Core.
[0056] Specifically, the upper computer program in the Host PC can directly operate the PCIE Bar space using the Bypass path of XDMA1 to perform a read operation on the VCU128 HBM and obtain the second file data.
[0057] S60, compare the first file data and the second file data to verify the HBM storage function.
[0058] Specifically, store the first file data in the specified location of the HBM, then read the HBM according to the stored information of the first file data to obtain the second file data, and the HBM storage function can be verified by comparing whether the first file data and the second file data are consistent.
[0059] In this application, the V19P and the VCU128 are interconnected and communicate based on the optical port and the self - encapsulated Aurora IP, realizing the seamless docking of the high - end FPGA V19P logic resources and the HBM, solving the problem that the traditional verification platform cannot support ultra - large - scale HBM channels, and through the HBM multi - channel parallel access architecture, achieving a continuous bandwidth of more than 32GB / s per channel and a total bandwidth of up to 460GB / s, which is more than 18 times higher than the traditional DDR4 memory, significantly improving the storage system performance of the V19P and meeting the real - time processing requirements of massive data in scenarios such as AI inference and genomics.
[0060] If the Aurora 64B / 66B provided by Xilinx is directly used by users, gt_reset and global reset, as well as the processing of gt_clk and initialization clock, are still required, which is not user-friendly. In the Example project of the Aurora 64B / 66B IP, gt_reset and global reset, as well as the processing of gt_clk and initialization clock, are carried out, and data generation and reception verification are performed through the aurora_64b66b_0_FRAME_GEN and aurora_64b66b_0_FRAME_CHECK modules. Therefore, the user interface is not led out externally, and the port signals that users can operate are only the clock reset pins, transceiver ports, and some error status pins, which is also not suitable for users to directly use.
[0061] In one embodiment, in order to effectively integrate and verify the function of HBM on V19P, the step of instantiating the self-packaged Aurora IP core based on VeriTiger-V19P to obtain V19P Aurora Core and instantiating the self-packaged Aurora IP core based on the VCU128 development board to obtain VCU128 Aurora Core can be specifically executed as follows: First, instantiate the self-packaged Aurora IP core based on VeriTiger-V19P to obtain V19P Aurora Core, and instantiate the self-packaged Aurora IP core based on the VCU128 development board to obtain VCU128 Aurora Core. Specifically, instantiate the Aurora 64B / 66B IP core based on the V19P platform, self-package this IP core into the V19P Aurora Core IP core, and realize the optical port communication between V19P and VCU128 through this IP core; instantiate the Aurora 64B / 66B IP core based on the VCU128 platform, self-package this IP core into the VCU128 Aurora Core IP core, and realize the optical port communication between VCU128 and V19P through this IP core.
[0062] Then, set the user interfaces of V19P Aurora Core and VCU128 Aurora Core to AXI4-Stream interfaces, and use the AXI4-Stream Interconnect IP core according to the AXI4-Stream interface. Specifically, the self-packaging IP process is as follows: Create a new top-level file, comment the aurora_64b66b_0_FRAME_GEN and aurora_64b66b_0_FRAME_CHECK modules, bring out the AXI4-Stream input and output ports for user reading and writing, change the initialization clock from differential to single-ended input, bring out the loopback control port to be controlled by external input, and output the user clock user_clk_out. When encapsulating Ports and Interfaces, for the convenience of users, set user_clk_out as the clock and associate it with the m_axi_rx and s_axi_tx interfaces, associate RESET reset with init_clk, encapsulate GTYQ0_p and GTYQ0_n as differential clock ports gt_clk, encapsulate RXP, RXN, TXP and TXN as gt_seria_io, and finally encapsulate the IP core interface as follows Figure 3 As shown, Figure 3 The instructions are as follows: s_axi_tx: AXI4-Stream Slave interface, used to receive user requests; gt_clk: gt clock input; init_clk: Initialize the clock; loopback_i: Set the Aurora data loopback mode; RESET: global reset signal; PMA_INIT: gt reset signal; m_axi_rx: AXI4-Stream Master interface, used to send user requests; gt_serial_io: GTY data sending and receiving port; user_clk_out: gt outputs user clock; HARD_ERR: Indicator signal used to indicate that a serious error that cannot be recovered has occurred in the Aurora link; SOFT_ERR: Indicator signal used to indicate that a recoverable error has occurred in the Aurora link; LANE_UP: Indicator signal used to indicate that the channel in the Aurora link has been successfully initialized and is ready to transmit data; CHANNEL_UP: Indicator signal used to indicate that the entire Aurora link has been successfully initialized and is ready to transmit data.
[0063] According to the method of self - encapsulating IP, self - encapsulate the IP cores V19P AuroraCore and VCU128 Aurora Core based on the V19P and VCU128 platforms respectively. Use these IP cores to implement data communication based on optical ports and Aurora protocol between the V19P and VCU128 platforms, laying a foundation for subsequent V19P to read and write HBM. The self - encapsulated Aurora IP can achieve a user - adaptive usage mode, improving the simplicity and usability of the Aurora IP.
[0064] In one embodiment, in order to effectively integrate and verify the function of HBM on V19P, combined with Figure 1 , the step of transmitting the first file data to HBM through V19P Aurora Core and VCU128 Aurora Core according to the HBM write request can be specifically executed as follows: First, use the XDMA driver to drive the first XDMA to transmit the first file data to V19P Aurora Core through the AXI4 - Stream Interconnect IP core. Specifically, write a host computer program in the Host PC (i.e., the client). In the host computer program, use the xdma0_h2c_x device file provided by the XDMA driver (i.e., the XDMA Driver shown in the figure) to initiate an HBM write request for the file data. The user interface of V19P Aurora Core is the AXI4 - Stream interface. So set the AXI4 - Stream Interconnect IP core (i.e., AXIS0 shown in the figure) so that the file data is transmitted to V19P Aurora Core. Therefore, the flow of the first file data is: XDMA0 sends the write request of this file data to the user interface of V19P Aurora Core through AXIS0 and then transmits it to V19P Aurora Core.
[0065] Then, transmit the first file data in V19P Aurora Core to VCU128 Aurora Core through the Aurora Lane. Specifically, V19P and VCU128 are interconnected and communicate through optical ports. At the same time, V19P Aurora Core sends the first file data to VCU128 Aurora Core through Aurora Lane 1~Aurora Lane n.
[0066] Next, the first file data in the VCU128 Aurora Core is transferred to the host computer through the AXI4-Stream Interconnect IP core and the second XDMA. Specifically, the user interface of the VCU128 Aurora Core is the AXI4-Stream interface. Therefore, the AXI4-Stream Interconnect IP core (i.e., AXIS1 shown in the figure) is set so that the file data is transferred to the VCU128 Aurora Core. The XDMA only acts as a transmission medium, and the XDMA driver drives the second XDMA (i.e., XDMA1 shown in the figure) to transfer the first file data. Therefore, the flow of the first file data is that the VCU128 Aurora Core transfers the first file data through the AXI4-Stream user interface and AXIS1, and transfers the first file data to the host computer through XDMA1.
[0067] Finally, the XDMA driver is used to drive the second XDMA to write the first file data into the HBM. Specifically, after the first file data is transferred to the host computer, it is also the host computer program that calls the XDMA driver to drive the XDMA to write the first file data into the HBM. The host computer program in the Host PC uses XDMA1 to realize the read and write of the VCU128 HBM and writes the first file data to the HBM.
[0068] For writing file data to the HBM on the V19P, the V19P and the VCU128 are interconnected and communicate based on the optical port and the self-packaged Aurora IP, realizing the seamless connection between the high-end FPGA V19P logic resources and the HBM, solving the problem that the traditional verification platform cannot support ultra-large-scale HBM channels, and through the HBM multi-channel parallel access architecture, achieving a continuous bandwidth of more than 32 GB / s per channel and a total bandwidth of up to 460 GB / s, which is more than 18 times higher than the traditional DDR4 memory, significantly improving the storage system performance of the V19P and meeting the real-time processing requirements of massive data in scenarios such as AI inference and genomics.
[0069] In one embodiment, in order to effectively integrate and verify the function of the HBM on the V19P, the step of using the XDMA driver to drive the second XDMA to write the first file data into the HBM can be specifically executed as follows: First, set the second XDMA to the Bypass mode. The Bypass mode includes a Bypass path and a Bypass interface. Specifically, Figure 1The AXI4 shown in the figure is the AXI Interconnect, which is the core IP for managing the AXI (Advanced eXtensible Interface) bus interconnect in Xilinx FPGA designs. It is responsible for connecting multiple AXI masters and slaves, and implementing functions such as address routing, protocol conversion, and clock domain isolation. In the VCU128 platform project, when adding AXI4, since the AXI user read / write interface is the AXI interface, while the self - encapsulated V19P Aurora Core and VCU128 AuroraCore are AXI4 - Stream interfaces, the DMA interface in the XDMA0 / XDMA1 IP core pre - selects the AXI Stream interface. Therefore, a Bypass interface is added. The Bypass interface can bypass the DMA, that is, it can directly perform read and write operations on the Bar space like a normal PCIE. For data transmission, the second XDMA is set to the Bypass mode, and the Bypass mode includes a Bypass path and a Bypass interface.
[0070] Then, the AXI Interconnect is used to convert the Bypass interface into an AXI interface. Specifically, AXI4 is used to convert the Bypass path into an AXI4 interface.
[0071] Finally, the XDMA driver is used to drive the second XDMA to write the first file data into the HBM through the Bypass path and the AXI interface. The upper - computer program in the Host PC directly operates on the PCIE Bar space using the Bypass path of XDMA1, so as to realize the read and write of the VCU128 HBM and write the first file data sent by V19P to the HBM.
[0072] Adding a Bypass interface, the Bypass interface can bypass the DMA, that is, it can directly perform read and write operations on the Bar space like a normal PCIE. Adding the AXI Interconnect is used to convert the Bypass path of XDMA into an AXI4 interface, enabling the file data to be transmitted smoothly.
[0073] In one embodiment, in order to effectively integrate and verify the function of the HBM on V19P, the step of the VCU128 AuroraCore and V19P Aurora Core transmitting the second file data to the upper - computer can be specifically executed as: First, use the XDMA driver to drive the second XDMA to read the HBM to obtain the second file data. Specifically, the upper computer program in the HostPC can use XDMA1 to perform a read operation on the HBM on the VCU128 to obtain the second file data.
[0074] Then, transfer the second file data to the VCU128 Aurora Core through the AXI4-Stream Interconnect IP core. Specifically, the user interface of the VCU128 Aurora Core is the AXI4-Stream interface. Therefore, when executing the read request of the HBM, set the AXI4-Stream Interconnect IP core (i.e., AXIS2 shown in the figure) so that the file data is transferred to the VCU128 Aurora Core. Therefore, the direction of the second file data is that the VCU128 Aurora Core sends the file data to the VCU128 Aurora Core through the AXI4-Stream user interface and AXIS2.
[0075] Next, transfer the second file data in the VCU128 Aurora Core to the V19P Aurora Core through the Aurora Lane. Specifically, the V19P and the VCU128 are interconnected and communicate through the optical port. At the same time, the VCU128 Aurora Core sends the second file data to the V19P Aurora Core through Aurora Lane 1 to Aurora Lane n.
[0076] Secondly, output the second file data in the V19P Aurora Core to the upper computer through the AXI4-Stream Interconnect IP core and the first XDMA. Specifically, the user interface of the V19P Aurora Core is the AXI4-Stream interface. Therefore, when executing the read request of the HBM, set the AXI4-Stream Interconnect IP core (i.e., AXIS3 shown in the figure) so that the file data is transferred to the V19P Aurora Core. Therefore, the direction of the second file data is that the V19P Aurora Core sends the second file data, and the upper computer program in the Host PC uses the xdma0_c2h_x device file provided by the XDMA driver to receive the second file data from the XDMA0. In one embodiment, in order to effectively integrate and verify the function of the HBM on the V19P, using the XDMA driver to drive the second XDMA to read the HBM to obtain the second file data can be specifically executed as: First, set the second XDMA to Bypass mode. The Bypass mode includes a Bypass path and a Bypass interface. Use the AXI Interconnect to convert the Bypass interface into an AXI interface. Specifically, AXI4 is added to the VCU128 platform project. Since the AXI4 user read / write interface is an AXI interface, and the self-packaged V19P Aurora Core and VCU128 AuroraCore are AXI4-Stream interfaces, the DMA interface in the XDMA0 / XDMA1 IP core pre-selects the AXI Stream interface. Therefore, a Bypass interface is added. The Bypass interface can bypass the DMA, that is, like a normal PCIE, it can directly perform read and write operations on the Bar space. For data transmission, set the second XDMA to Bypass mode. The Bypass mode includes a Bypass path and a Bypass interface. AXI4 is used to convert the Bypass path into an AXI4 interface.
[0077] Then, use the XDMA driver to drive the second XDMA to read the HBM through the Bypass path and the AXI interface to obtain the second file data. Specifically, the upper computer program in the Host PC uses the Bypass path of XDMA1 to directly operate the PCIE Bar space, thereby realizing the read and write of the VCU128 HBM, and reads the HBM according to the storage information to obtain the second file data.
[0078] Add a Bypass interface. The Bypass interface can bypass the DMA, that is, like a normal PCIE, it can directly perform read and write operations on the Bar space. Add AXI Interconnect to convert the Bypass path of XDMA into an AXI4 interface, so that the file data can be transmitted smoothly.
[0079] In one embodiment, in order to effectively integrate and verify the function of the HBM on the V19P, the step of comparing the first file data and the second file data to verify the HBM storage function can be specifically performed as follows: First, compare whether the first file data and the second file data are the same. Specifically, store the first file data in a specified location of the HBM, and then read the HBM according to the storage information of the first file data to obtain the second file data. By comparing whether the first file data and the second file data are the same, the HBM storage function can be verified.
[0080] Then, if the first file data is consistent with the second file data, it is determined that the HBM storage function verification passes; otherwise, it is determined that the HBM storage function verification fails. Specifically, if the first file data written and the second file data read are consistent, it indicates that the HBM storage function is normal and the verification passes; if the first file data written and the second file data read are inconsistent, it indicates that there is an abnormality in the HBM storage function and the verification fails.
[0081] On V19P, the same file data is written to and read from the HBM, and by comparing the original file data with the file data after reading and writing, the read and write verification of the HBM on VCU128 is realized.
[0082] Based on the above method, an embodiment of the present application also discloses a device for verifying the HBM storage function on V19P. As Figure 4 , the device includes the following modules: The Aurora IP core instantiation module 401 is used to instantiate the self - encapsulated Aurora IP core based on VeriTiger - V19P to obtain the V19P Aurora Core, and instantiate the self - encapsulated Aurora IP core based on the VCU128 development board to obtain the VCU128 AuroraCore; The HBM write request receiving module 402 is used to receive the HBM write request, and the HBM write request includes the first file data and the storage information of the first file data; The file data writing module 403 is used to transmit the first file data to the HBM through the V19P AuroraCore and the VCU128 Aurora Core according to the HBM write request; The HBM read request initiating module 404 is used to initiate the HBM read request according to the storage information; The file data reading module 405 is used to read the second file data from the HBM according to the HBM read request, and transmit the second file data to the host computer through the VCU128 Aurora Core and the V19P Aurora Core; The HBM storage function verification module 406 is used to verify the HBM storage function by comparing the first file data and the second file data.
[0083] In one embodiment, the Aurora IP core instantiation module 401 is specifically configured to instantiate the self - encapsulated Aurora IP core based on VeriTiger - V19P to obtain the V19P Aurora Core, and instantiate the self - encapsulated Aurora IP core based on the VCU128 development board to obtain the VCU128 Aurora Core; set the user interfaces of the V19P Aurora Core and the VCU128 Aurora Core as AXI4 - Stream interfaces, and adopt the AXI4 - Stream Interconnect IP core according to the AXI4 - Stream interfaces.
[0084] In one embodiment, the file data writing module 403 is specifically configured to drive the first XDMA using the XDMA driver to transmit the first file data to the V19P Aurora Core through the AXI4 - Stream Interconnect IP core; transmit the first file data in the V19P Aurora Core to the VCU128 Aurora Core through the Aurora Lane; transmit the first file data in the VCU128 Aurora Core to the host computer through the AXI4 - Stream Interconnect IP core and the second XDMA; drive the second XDMA using the XDMA driver to write the first file data into the HBM.
[0085] In one embodiment, the file data writing module 403 is specifically configured to set the second XDMA to the Bypass mode, and the Bypass mode includes a Bypass path and a Bypass interface; convert the Bypass interface to an AXI interface using the AXI Interconnect; drive the second XDMA using the XDMA driver to write the first file data into the HBM through the Bypass path and the AXI interface.
[0086] In one embodiment, the file data reading module 405 is specifically configured to drive the second XDMA using the XDMA driver to read the HBM to obtain the second file data; transmit the second file data to the VCU128 Aurora Core through the AXI4 - Stream Interconnect IP core; transmit the second file data in the VCU128 Aurora Core to the V19P Aurora Core through the Aurora Lane; output the second file data in the V19P Aurora Core to the host computer through the AXI4 - Stream Interconnect IP core and the first XDMA.
[0087] In one embodiment, the file data reading module 405 is specifically configured to set the second XDMA to the Bypass mode. The Bypass mode includes a Bypass path and a Bypass interface. Convert the Bypass interface to an AXI interface by using the AXI Interconnect. Use the XDMA driver to drive the second XDMA to read the first file data through the Bypass path and the AXI interface to obtain the second file data.
[0088] In one embodiment, the HBM storage function verification module 406 is specifically configured to compare whether the first file data and the second file data are consistent. If the first file data and the second file data are consistent, it is determined that the HBM storage function verification passes; otherwise, it is determined that the HBM storage function verification fails.
[0089] The embodiment of the present application also discloses a computer device.
[0090] Specifically, the computer device includes a memory and a processor. A computer program capable of being loaded and executed by the processor for the above method of verifying the HBM storage function on V19P is stored on the memory.
[0091] The embodiment of the present application also discloses a computer-readable storage medium.
[0092] Specifically, the computer-readable storage medium stores a computer program capable of being loaded and executed by the processor for the method of verifying the HBM storage function on V19P as described above. The computer-readable storage medium includes, for example, various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0093] This specific embodiment is only an explanation of the present invention and does not limit the present invention. After reading this specification, those skilled in the art can make modifications without creative contributions to this embodiment as needed, but as long as it is within the scope of the claims of the present invention, it is protected by the patent law.
Claims
1. A method for verifying HBM storage function on V19P, characterized in that: The method is based on VeriTiger-V19P, VCU128 development board, self-packaged Aurora IP core and host computer, and the method includes: Based on the VeriTiger-V19P, the self-packaged Aurora IP core is instantiated to obtain the V19P Aurora Core, and based on the VCU128 development board, the self-packaged Aurora IP core is instantiated to obtain the VCU128 Aurora Core; Receiving an HBM write request issued by the host computer, wherein the HBM write request includes first file data and storage information of the first file data; Transmitting the first file data to the HBM through the V19P Aurora Core and the VCU128Aurora Core according to the HBM write request; Initiate an HBM read request according to the storage information; Read the HBM according to the HBM read request to obtain second file data, and transmit the second file data to the host computer through the VCU128 Aurora Core and the V19P Aurora Core; The first file data and the second file data are compared to verify the HBM storage function.
2. The method according to claim 1, characterized in that: The instantiation of the self-packaged Aurora IP core based on the VeriTiger-V19P to obtain the V19P Aurora Core, and the instantiation of the self-packaged Aurora IP core based on the VCU128 development board to obtain the VCU128 Aurora Core include: Based on the VeriTiger-V19P, the self-packaged Aurora IP core is instantiated to obtain the V19P Aurora Core, and based on the VCU128 development board, the self-packaged Aurora IP core is instantiated to obtain the VCU128 Aurora Core; The user interface of the V19P Aurora Core and the VCU128 Aurora Core is set to an AXI4-Stream interface, and an AXI4-Stream Interconnect IP core is used according to the AXI4-Stream interface.
3. The method according to claim 2, characterized in that: The method is also based on XDMA, XDMA driver and AuroraLane, wherein the first file data is transmitted to the HBM through the V19P Aurora Core and the VCU128 Aurora Core according to the HBM write request, including: Using the XDMA driver to drive a first XDMA to transfer the first file data to the V19P Aurora Core through the AXI4-StreamInterconnect IP core; Transmitting the first file data in the V19P Aurora Core to the VCU128 Aurora Core through the Aurora Lane; The first file data in the VCU128 Aurora Core is transferred to the host computer through the AXI4-Stream Interconnect IP core and the second XDMA; The XDMA driver is used to drive the second XDMA to write the first file data into the HBM.
4. The method according to claim 3, characterized in that: The method is also based on AXI Interconnect, and using the XDMA driver to drive the second XDMA to write the first file data into the HBM includes: Setting the second XDMA to a Bypass mode, wherein the Bypass mode includes a Bypass path and a Bypass interface; Converting the Bypass interface into an AXI interface using the AXI Interconnect; The XDMA driver is used to drive the second XDMA to write the first file data into the HBM through the Bypass path and the AXI interface.
5. The method according to claim 4, characterized in that: The step of reading the HBM according to the HBM read request to obtain the second file data, and transmitting the second file data to the host computer through the VCU128 Aurora Core and the V19P Aurora Core includes: Using the XDMA driver to drive the second XDMA to read the HBM to obtain second file data; Transmitting the second file data to the VCU128Aurora Core through the AXI4-Stream Interconnect IP core; Transmitting the second file data in the VCU128 Aurora Core to the V19P Aurora Core through the Aurora Lane; The second file data in the V19P Aurora Core is output to the host computer through the AXI4-Stream Interconnect IP core and the first XDMA.
6. The method according to claim 5, characterized in that: The using the XDMA driver to drive the second XDMA to read the HBM to obtain the second file data includes: Setting the second XDMA to a Bypass mode, wherein the Bypass mode includes a Bypass path and a Bypass interface; Converting the Bypass interface into an AXI interface using the AXI Interconnect; The XDMA driver is used to drive the second XDMA to read the HBM through the Bypass path and the AXI interface to obtain the second file data.
7. The method according to claim 1, characterized in that: The comparing the first file data and the second file data to verify the HBM storage function includes: comparing whether the first file data and the second file data are consistent; If the first file data and the second file data are consistent, it is determined that the HBM storage function verification is passed; Otherwise, it is determined that the HBM storage function verification has failed.
8. A device for verifying HBM storage function on V19P, used to implement the method for verifying HBM storage function on V19P as claimed in claim 1, characterized in that: The device comprises: Aurora IP core instantiation module (401), used for instantiating the self-packaged Aurora IP core based on the VeriTiger-V19P to obtain the V19P Aurora Core, and instantiating the self-packaged Aurora IP core based on the VCU128 development board to obtain the VCU128 Aurora Core; An HBM write request receiving module (402) is used to receive an HBM write request sent by the host computer, wherein the HBM write request includes first file data and storage information of the first file data; A file data writing module (403), configured to transmit the first file data to the HBM through the V19Aurora Core and the VCU128 Aurora Core according to the HBM write request; An HBM read request initiating module (404), configured to initiate an HBM read request according to the storage information; A file data reading module (405), configured to read the HBM according to the HBM read request to obtain second file data, and transmit the second file data to the host computer through the VCU128 Aurora Core and the V19P Aurora Core; The HBM storage function verification module (406) is used to compare the first file data and the second file data to verify the HBM storage function.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program which can be loaded by the processor and executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Device and method based on FPGA for achieving automatic read-write testing of DDR interface
CN107239374A
Access processing method and device for high-bandwidth memory on FPGA (Field Programmable Gate Array)
CN118133733A
Receiving and transmitting fusion processing method based on AURORA simplex IP core
CN119483873A