Low-latency transmission system for RDMA network cards

By designing a low-latency transmission system on the RDMA network card, using the combination of the driver software module and the FPGA hardware module to dynamically select the SWQE dispatch strategy, the problem of high transmission delay of the RDMA network card is solved, and lower network transmission delay and higher transmission performance are achieved.

CN118869477BActive Publication Date: 2025-06-10ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410838401.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-26
Publication Date
2025-06-10
Estimated Expiration
2044-06-26

AI Technical Summary

Technical Problem

The existing RDMA network cards have high transmission delays, which affects network transmission performance.

Method used

A low-latency transmission system is designed, including a driver software module and an FPGA hardware module. By dynamically selecting the SWQE dispatch strategy, the packet transmission process is optimized and the number of times the FPGA hardware reads the host memory is reduced.

Benefits of technology

It effectively reduces the time from sending SWQE from software to actually sending data packets from hardware, reduces network transmission delay, and improves the high-speed transmission performance of network cards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118869477B_ABST
    Figure CN118869477B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-latency transmission system for an RDMA network card, comprising: a driver software module and an FPGA hardware module; the driver software module is set in the host RDMA network card driver, and is used to receive the Send WR sent by the upper-layer interface, parse it into the SWQE required by the hardware, and is also responsible for dynamically selecting a strategy to send the parsed SWQE to the FPGA hardware module; the FPGA hardware module is set in the RDMA network card, and is used to obtain the information of the transmission data packet through the SWQE sent by the driver software module, and then read the data to be transmitted from the memory or the SWQE, so as to complete the actual transmission of the RDMA network card. The low-latency transmission system for an RDMA network card of the present invention can reduce the number of times the FPGA hardware reads the host memory when using the RDMA network, thereby reducing the time from the software sending the SWQE to the hardware actually sending the data packet, which helps to exert the high-speed transmission performance of the network card and greatly reduces the network transmission latency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network technology, and particularly relates to a low-latency transmission system for an RDMA network card. Background Art

[0002] With the rapid development of communication and electronic technologies, data has exploded in growth, and large-scale real-time data acquisition systems have increasingly high requirements for high-speed network transmission. RDMA (Remote Direct Memory Access) networks, with their excellent characteristics of low latency, low CPU overhead, and high bandwidth, have become the mainstream of data center networks in recent years. Their types mainly include Infiniband networks and RoCE networks. The RDMA technology eliminates the overhead of packet copying and context switching between the user space and the kernel space, thus liberating the memory bandwidth and CPU cycles for improving the performance of application systems.

[0003] RDMA provides a set of software transmission interfaces. Application programs can call the driver through the interfaces to issue Work Queue Elements (WQE for short). The WQE describes the addresses, lengths, keys, etc. of the data to be sent / received. Taking the Send operation request as an example, the sending-end driver writes the information such as the address, length, and key of the data to be sent into the SWQE (Send WQE). The sending-end network card parses the SWQE and completes the sending of the packet.

[0004] The process of most drivers issuing SWQE is as follows: The driver fills the SWQE into the memory and notifies the hardware through the register. After receiving the register instruction, the hardware reads the SWQE from the memory and parses it. Subsequently, the hardware network card reads the data to be sent from the memory according to the SWQE and completes the packet sending. In order to further reduce the sending latency of the RDMA network card, it is necessary to further study and optimize the above process. Summary of the Invention

[0005] The present invention provides a low-latency transmission system for an RDMA network card to solve the above-mentioned technical problems, and specifically adopts the following technical solutions:

[0006] A low-latency transmission system for an RDMA network card, comprising: a driver software module and an FPGA hardware module;

[0007] The driver software module is set in the host RDMA network card driver, and is used to receive the Send WR issued by the upper-layer interface, parse it into the SWQE required by the hardware, and is also responsible for dynamically selecting a strategy to issue the parsed SWQE to the FPGA hardware module;

[0008] The FPGA hardware module is set in the RDMA network card and is used to obtain information of the sending data packet through the SWQE sent by the driver software module, and then read the data to be sent from the memory or the SWQE, so as to complete the actual sending of the RDMA network card.

[0009] Further, the driver software module includes:

[0010] A task parsing unit, which is used to receive the Send WR sent by the upper-layer interface, read the address, length, and key of the sending data from it, and fill them into the SWQE. At the same time, if the length of the data to be sent is less than or equal to 16 bytes, the sending data is directly filled into the SWQE, so as to obtain the SWQE required by the hardware.

[0011] A task decision-making unit. If the length of the data to be sent is greater than 16 bytes, the task decision-making unit is used to dynamically select the SWQE sending strategy, and then send the SWQE to the subsequent unit.

[0012] A DMA sending unit. When the task decision-making unit selects the DMA sending mode, the task decision-making unit sends the SWQE to the DMA sending unit. The DMA sending unit sends the SWQE to the hardware in the DMA manner. First, the SWQE is filled into the DMA circular buffer allocated by the driver software module, and then the hardware is notified to read the SWQE from the DMA circular buffer by writing the Doorbell register.

[0013] An MMIO sending unit. When the task decision-making unit selects the MMIO sending mode, the task decision-making unit sends the SWQE to the MMIO sending unit, which is used to send the SWQE in the MMIO manner and write 64-byte SWQE into the MMIO_512 register at one time using the avx512 instruction and transfer it to the hardware.

[0014] Further, the FPGA hardware module includes:

[0015] A hardware sending unit, which is used to obtain information of the sending data packet according to its content after receiving the SWQE sent by the driver software module, and then read the data to be sent from the memory or the SWQE, so as to complete the corresponding packet sending task.

[0016] Further, the rules for the task decision-making unit to select the SWQE sending strategy are as follows:

[0017] Initially, the MMIO sending mode is preferentially selected. The queue depth of the MMIO sending mode is limited. After the MMIO queue is full, the mode is switched to the DMA sending mode.

[0018] When in the DMA transmission mode, and both the MMIO transmission queue and the DMA transmission queue are empty, and all transmission tasks are completed, switch back to the MMIO mode again.

[0019] Furthermore, the DMA transmission unit manages a DMA circular buffer, which is a continuous memory space with a length that is an integer multiple of the length of a single SWQE. Each access will offset by the length of one SWQE based on the previous access. If the end is reached, the next access starts from the beginning. The DMA transmission unit writes the SWQE, and the hardware transmission unit reads the SWQE.

[0020] Furthermore, the hardware transmission unit includes a Doorbell register, an MMIO_512 register, a DMA transmission queue, and an MMIO transmission queue.

[0021] Furthermore, the DoorBell register is an FPGA hardware register. After address mapping, it can be directly accessed by the DMA transmission unit and is used for the DMA transmission unit to notify the FPGA hardware module to read the SWQE after writing the DMA buffer.

[0022] The MMIO_512 register is an FPGA hardware register with a length of 512 bits. After address mapping, it is directly accessed by the MMIO transmission unit and is used for the MMIO transmission unit to write the SWQE.

[0023] Furthermore, the DMA transmission queue is a queue managed by the FPGA hardware. When the DMA transmission unit writes the Doorbell register, the DMA transmission queue enqueues. When the packet in the DMA transmission mode is completed, the DMA transmission queue dequeues.

[0024] The MMIO transmission queue is a queue managed by the FPGA hardware. When the MMIO transmission unit writes the MMIO_512 register, the MMIO transmission queue enqueues. When the packet in the MMIO transmission mode is completed, the MMIO transmission queue dequeues.

[0025] Furthermore, the low-latency transmission system for the RDMA network card is applicable to RDMA network card devices based on the Infiniband network or the RoCE network.

[0026] The advantage of the present invention lies in the provided low-latency transmission system for the RDMA network card, which can reduce the number of times the FPGA hardware reads the host memory when using the RDMA network, thereby reducing the time from the software to issue the SWQE to the actual packet sending by the hardware, helping to exert the high-speed transmission performance of the network card and greatly reducing the network transmission latency. Description of the Drawings

[0027] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a general schematic diagram of the low-latency transmission system for the RDMA network card of the present invention;

[0029] Figure 2 It is an internal structure schematic diagram of the low-latency transmission system for the RDMA network card of the present invention;

[0030] Figure 3 It is a flowchart of the SWQE issuing strategy of the task decision unit of the present invention;

[0031] Figure 4 It is a general flowchart of the driver software module issuing SWQE of the present invention;

[0032] Figure 5 It is a general flowchart of the FPGA hardware module completing packet transmission through SWQE of the present invention. Specific embodiments

[0033] The following will describe in detail the embodiments of the present application. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.

[0034] As Figure 1 shown, the low-latency transmission system 100 for the RDMA network card of the present application includes a driver software module 10 and an FPGA hardware module 20.

[0035] The driver software module 10 is set in the host RDMA network card driver, and is used to receive the Send WR (Work Request) sent by the upper-layer interface, parse it into the SWQE (Send Work Queue Element, abbreviated as SWQE) required by the hardware, and is also responsible for dynamically selecting the SWQE issuing strategy to issue the parsed SWQE to the FPGA hardware module 20.

[0036] The FPGA hardware module 20 is set in the RDMA network card and is used to obtain the information of the sending data packet through the SWQE sent by the driver software module 10, and then read the data to be sent from the memory or SWQE to complete the actual sending of the RDMA network card.

[0037] As Figure 2 shown is the internal structure schematic diagram of the low-latency sending system for the RDMA network card of the present application.

[0038] In the embodiment of the present application, the driver software module 10 includes a task parsing unit 11, a task decision-making unit 12, a DMA sending unit 13, and an MMIO sending unit 14.

[0039] The task parsing unit 11 is used to receive the Send WR sent by the upper-layer interface, read the information such as the address, length, and key of the sending data from it, and fill it into the SWQE. At the same time, if the length of the data to be sent is less than or equal to 16 bytes, the sending data is directly filled into the SWQE, thereby obtaining the SWQE required by the hardware.

[0040] The task decision-making unit 12 is used to dynamically select the SWQE sending strategy. If the length of the data to be sent is greater than 16 bytes, the task decision-making unit is used to dynamically select the SWQE sending strategy, and then send the SWQE to the subsequent unit, that is, the DMA sending unit 13 or the MMIO sending unit 14.

[0041] When the task decision-making unit 12 selects the DMA sending mode, the task decision-making unit 12 sends the SWQE to the DMA sending unit 13, and the DMA sending unit 13 sends the SWQE to the hardware through DMA, that is, first fills the SWQE into the DMA circular buffer allocated by the driver software module 10, and then notifies the hardware to read the SWQE from the DMA circular buffer by writing the Doorbell register.

[0042] When the task decision-making unit 12 selects the MMIO sending mode, the task decision-making unit 12 sends the SWQE to the MMIO sending unit 14, and the MMIO sending unit 14 sends the SWQE through MMIO and uses the avx512 instruction to write the 64-byte SWQE into the MMIO_512 register at one time and transfer it to the hardware.

[0043] In the embodiment of the present application, the FPGA hardware module 20 includes a hardware sending unit 21.

[0044] After receiving the SWQE sent by the driver software module 10, the hardware sending unit 21 obtains the information of the sending data packet according to its content, and then reads the data to be sent from the memory or SWQE to complete the corresponding packet sending task.

[0045] In an embodiment of the present application, the DMA transmission unit 13 manages a DMA circular buffer. The DMA circular buffer is a continuous memory space, the length of which is an integer multiple of the length of a single SWQE. Each access will offset by the length of one SWQE based on the previous access. If the end is reached, the next access starts from the beginning. The DMA transmission unit 13 writes the SWQE, and the hardware transmission unit 21 reads the SWQE.

[0046] Furthermore, the hardware transmission unit 21 includes a Doorbell register, an MMIO_512 register, a DMA transmission queue, and an MMIO transmission queue.

[0047] The DoorBell register is essentially an FPGA hardware register. After address mapping, it can be directly accessed by the DMA transmission unit 13, and is used for the DMA transmission unit 13 to notify the FPGA hardware module 20 to read the SWQE after writing the DMA buffer.

[0048] The MMIO_512 register is an FPGA hardware register with a length of 512 bits. After address mapping, it can be directly accessed by the MMIO transmission unit 14, and is used for the MMIO transmission unit 14 to write the SWQE.

[0049] The DMA transmission queue is a queue managed by the FPGA hardware. When the DMA transmission unit 13 writes the Doorbell register, the DMA transmission queue is enqueued. When the packet in the DMA transmission mode is completed, the DMA transmission queue is dequeued.

[0050] The MMIO transmission queue is a queue managed by the FPGA hardware. When the MMIO transmission unit 14 writes the MMIO_512 register, the MMIO transmission queue is enqueued. When the packet in the MMIO transmission mode is completed, the MMIO transmission queue is dequeued.

[0051] As Figure 3 shown, in an embodiment of the present application, the rules for the task decision unit 12 to dynamically select the SWQE distribution strategy are as follows:

[0052] Initially, the MMIO transmission mode is preferentially selected. Since the queue depth of the MMIO transmission mode is limited, after the MMIO queue is full, it will switch to the DMA transmission mode.

[0053] When it is in the DMA transmission mode at this time, and both the MMIO transmission queue and the DMA transmission queue are empty, that is, when all transmission tasks are completed, it will switch back to the MMIO mode.

[0054] In an embodiment of the present application, elements in the MMIO transmission queue and the DMA transmission queue are both queued by the driver software to issue SWQEs, and packet transmission and dequeueing are completed by the FPGA hardware.

[0055] As Figure 4 shown in the general flowchart of the driver software module 10 of the present invention for issuing SWQEs, the process is as follows:

[0056] S11: The IB_Verbs interface calls Post_send to send the Send WR to the driver software module 10.

[0057] S12: After the driver software module 10 receives the Send WR, the task parsing unit 11 parses it and fills information such as the length, address, and key of the transmission data into the SWQE.

[0058] S13: The task parsing unit 11 judges the length of the transmission data. If it is greater than 16 bytes, step S15 is implemented; if it is less than or equal to 16 bytes, step S14 is implemented.

[0059] S14: The task parsing unit 11 fills the transmission data into the SWQE.

[0060] S15: The task decision unit 12 dynamically selects the SWQE issuing strategy. If the MMIO transmission mode is adopted, the SWQE is issued to the MMIO transmission unit 14 and step S18 is implemented; if the DMA transmission mode is adopted, the SWQE is issued to the DMA transmission unit 13 and step S16 is implemented.

[0061] S16: The DMA transmission unit 13 fills the SWQE into the DMA buffer.

[0062] S17: The DMA transmission unit 13 notifies the hardware to read the SWQE from the DMA buffer by writing the Doorbell register.

[0063] S18: The MMIO transmission unit 14 uses the avx512 instruction to write 64 bytes of SWQE into the MMIO_512 register at one time and transfers it to the hardware.

[0064] S19: The hardware transmission unit 21 of the FPGA hardware module 20 receives the SWQE issued by the driver software module 10 and completes the packet transmission task.

[0065] As Figure 5 shown in the general flowchart of the FPGA hardware module 20 of the present invention for completing packet transmission through the SWQE, the process is as follows:

[0066] S21: The FPGA hardware module 20 receives the SWQE issued by the driver software module 10.

[0067] S22: Judge the MMIO transmission queue. If the queue is empty, execute step S23; if the queue is not empty, execute step S24.

[0068] S23: The hardware transmission unit 21 reads the SWQE from the corresponding position of the DMA buffer.

[0069] S24: The hardware transmission unit 21 directly obtains the register from the MMIO transmission queue.

[0070] S25: The hardware transmission unit 21 judges the length of the transmission data. If it is greater than 16 bytes, execute step S26; if it is less than or equal to 16 bytes, execute step S27.

[0071] S26: The hardware transmission unit 21 reads the transmission data from the memory according to the information such as the address, length, and key of the transmission data in the SWQE.

[0072] S27: The hardware transmission unit 21 directly obtains the transmission data from the SWQE.

[0073] S28: The hardware transmission unit 21 packetizes the transmission data and sends it out.

[0074] The present invention belongs to the field of network technologies, and particularly relates to a low-latency transmission system for an RDMA network card.

[0075] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form. Any technical solutions obtained by using equivalent replacements or equivalent transformations fall within the protection scope of the present invention.

Claims

1. A low-latency transmission system for an RDMA network card, characterized in that: Contains: driver software module and FPGA hardware module; The driver software module is set in the host RDMA network card driver, and is used to receive the SendWR sent by the upper layer interface, parse it into the SWQE required by the hardware, and is also responsible for dynamically selecting a strategy to send the parsed SWQE to the FPGA hardware module; The FPGA hardware module is arranged in the RDMA network card, and is used to obtain information of sending data packets through the SWQE issued by the driver software module, and then read the sent data from the memory or SWQE to complete the actual sending of the RDMA network card; The driver software module includes: The task parsing unit is used to receive the Send WR sent by the upper layer interface, read the address, length and key of the sent data, and fill them into the SWQE. At the same time, if the length of the sent data is less than or equal to 16 bytes, the sent data is directly filled into the SWQE, thereby obtaining the SWQE required by the hardware; A task decision unit, if the length of the sent data is greater than 16 bytes, the task decision unit is used to dynamically select a SWQE delivery strategy, and then deliver the SWQE to a subsequent unit; A DMA sending unit, when the task decision unit selects the DMA sending mode, the task decision unit sends the SWQE to the DMA sending unit, and the DMA sending unit sends the SWQE to the hardware through DMA, first fills the SWQE into the DMA ring buffer allocated by the driver software module, and then notifies the hardware to read the SWQE from the DMA ring buffer by writing the Doorbell register; The MMIO sending unit, when the task decision unit selects the MMIO sending mode, sends the SWQE to the MMIO sending unit, which is used to write the 64-byte SWQE into the MMIO_512 register once and pass it to the hardware through the MMIO mode by using the avx512 instruction.

2. The low-latency transmission system for RDMA network card according to claim 1, characterized in that: The FPGA hardware module includes: The hardware sending unit is used to obtain the information of sending data packet according to the content of the SWQE after receiving the SWQE sent by the driver software module, and then read the sent data from the memory or SWQE to complete the corresponding packet sending task.

3. The low-latency transmission system for RDMA network card according to claim 2, characterized in that: The rules for the task decision unit to select the SWQE delivery strategy are as follows: Initially, the MMIO send mode is preferred. The queue depth of the MMIO send mode is limited. After the MMIO queue is full, it switches to the DMA send mode. When the DMA transmission mode is in DMA transmission mode, and both the MMIO transmission queue and the DMA transmission queue are empty and all transmission tasks are completed, switch back to MMIO mode.

4. The low-latency transmission system for RDMA network card according to claim 3, characterized in that: The DMA sending unit manages a DMA ring buffer, which is a continuous memory space whose length is an integer multiple of the length of a single SWQE. Each access will offset the length of a SWQE based on the previous access. If the access reaches the end, the next access will start from the beginning. The DMA sending unit writes the SWQE, and the hardware sending unit reads the SWQE.

5. The low-latency transmission system for RDMA network card according to claim 4, characterized in that: The hardware sending unit includes a Doorbell register, an MMIO_512 register, a DMA sending queue and an MMIO sending queue.

6. The low-latency transmission system for RDMA network card according to claim 5, characterized in that: The DoorBell register is an FPGA hardware register, which can be directly accessed by the DMA sending unit after address mapping, and is used by the DMA sending unit to notify the FPGA hardware module to read SWQE after writing the DMA buffer; The MMIO_512 register is an FPGA hardware register, which has a length of 512 bits and is directly accessed by the MMIO sending unit after address mapping, and is used for the MMIO sending unit to write SWQE.

7. The low-latency transmission system for RDMA network card according to claim 5, characterized in that: The DMA sending queue is a queue managed by the FPGA hardware. When the DMA sending unit writes the Doorbell register, the DMA sending queue is queued. When the data packet through the DMA sending mode is completed, the DMA sending queue is dequeued. The MMIO sending queue is a queue managed by FPGA hardware. When the MMIO sending unit writes the MMIO_512 register, the MMIO sending queue is queued. When the data packet in the MMIO sending mode is completed, the MMIO sending queue is dequeued.

8. The low-latency transmission system for RDMA network card according to claim 1, characterized in that: The low-latency sending system for an RDMA network card is applicable to an RDMA network card device based on an Infiniband network or a RoCE network.

Citation Information

Patent Citations

  • Control layer data kernel bypass system for RDMA network card

    CN118158088A