Loading command processing method and system based on remote direct memory access protocol
By combining the load commands of the graphics processor at the source side as read commands and decomposing them at the target side, the problem of low bandwidth utilization in RoCE v2 technology is solved, and efficient graphics processor loading operations are achieved.
Patent Information
- Application Number
- CN202510846096.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-23
AI Technical Summary
When the existing RoCE v2 technology solution uses remote direct memory access to load across chips, the bandwidth utilization rate is low and occupies the send queue resources, and it is impossible to effectively synthesize multiple read commands to transmit data.
The source-side graphics processor converts multiple load commands into multiple read commands, merges and generates work queue items, and transmits them to the target end through the Ethernet link. The target end is then decomposed into multiple read commands, improving bandwidth utilization and saving queue resources.
In the graphics processor loading operation, the bandwidth utilization is significantly improved, the transmission queue resource occupation is reduced, and the system operation is maintained in an orderly manner.
Smart Images

Figure CN120378389A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer hardware and network communication, and particularly to a method and system for processing load commands based on a Remote Direct Memory Access (RDMA) protocol. Background Art
[0002] With the rapid development of high-performance computing and artificial intelligence technologies, the converged Ethernet protocol based on Remote Direct Memory Access (RDMA) technology has become a key technology for achieving low-latency and high-bandwidth communication in data centers and high-performance computing clusters. Among them, Remote Direct Memory Access over Converged Ethernet v2 (RoCE v2) can reduce the latency and improve the throughput rate during Ethernet transmission by eliminating the overhead of the traditional network protocol stack, significantly enhancing the data transmission efficiency in the Ethernet environment. However, there are still some technical bottlenecks in the existing RoCE v2 technology solutions. Especially when a Graphics Processing Unit (GPU) performs a load operation across chips through RoCE v2, there are problems of low bandwidth utilization and waste of the Send Queue (SQ).
[0003] Specifically, in the traditional RoCE v2 technology solution, the send queue needs to store Work Queue Entries (WQEs). The RoCE Network Interface Card (RNIC) reads each work queue entry in sequence and performs data transmission according to the transmission attributes and data length defined therein. If it is desired that the graphics processing unit directly completes the load operation through remote direct memory access, each 64B / 128B GPU load command needs to be converted into an Advanced eXtensible Interface (AXI) read command through the Advanced eXtensible Interface bus, and then a work queue entry with a read semantics is generated and stored in the corresponding send queue. Then, the RoCE network interface card reads and processes the work queue entry for a read operation.
[0004] However, due to the limited resources of the send queue for storing work queue items, multiple queue pairs (QP) in RDMA will quickly exhaust the number of entries in the send queue. Moreover, when transmitting over an Ethernet link based on remote direct memory access, the protocol header part will occupy bandwidth. At this time, if the amount of data carried increases, the bandwidth utilization rate will increase accordingly. This expects to combine multiple graphics processor load commands into one remote direct memory access network interface card read work queue item to save send queue resources and increase bandwidth utilization. However, multiple read commands will carry multiple read addresses, and in the InfiniBand (IB) protocol, read transmissions do not transmit data (payload) packets, so multiple addresses cannot be transmitted to the target chip.
[0005] Therefore, there is an urgent need for a technical solution that can effectively solve the problem of low bandwidth utilization and a large amount of occupation of send queue resources when the graphics processor performs a load operation across chips through RoCE v2. Summary of the Invention
[0006] In view of this, the present invention provides a method and system for processing load commands based on the remote direct memory access protocol to solve the above technical problems in the prior art.
[0007] According to one aspect of the present invention, a method for processing load commands based on the remote direct memory access protocol is provided, wherein the processing method includes: Initiate multiple load commands through the source graphics processor, and convert the multiple load commands into multiple source read commands through the source Advanced Extensible Interface bus. The multiple source read commands include multiple queue pairs; Merge the source read commands located in the same queue pair to generate a source work queue item; Transmit the source write data included in the source work queue item to the target end through the Ethernet link; Store the source write data in the target end receive buffer, and each of the source write data is stored in an entry of the target end receive buffer; Decompose the write data in each entry of the target end receive buffer to obtain multiple read commands; Send the multiple read commands to the target end Advanced Extensible Interface bus, and send the multiple read commands to the target end memory through the target end Advanced Extensible Interface bus and read back the data; Send the read-back data to the target end send buffer through the target end Advanced Extensible bus; Merge the read-back data to generate a target end work queue item; Read the target end work queue item, and transmit the target end write data included in the target end work queue item to the source end through the Ethernet link; Receive write data at the destination end; Decompose the write data at the destination end to obtain multiple read data, and transmit the multiple read data to the source graphics processor.
[0008] According to another aspect of the present invention, a load command processing system based on the Remote Direct Memory Access (RDMA) protocol is provided. The system includes: A source chip, which includes a source graphics processor, a source Advanced Extensible Interface (AXI) bus, a source Remote Direct Memory Access (RDMA) network interface card, a source merge generation module, and a source decomposition module; A destination chip, which includes a destination RDMA network interface card, a destination decomposition module, a destination memory, a destination merge generation module, and a destination AXI bus; Among them, the source graphics processor initiates multiple load commands, and converts the multiple load commands into multiple source read commands through the source AXI bus. The multiple source read commands include multiple queue pairs; The source merge generation module merges the source read commands in the same queue pair to generate a source work queue entry; The source RDMA network interface card reads the source work queue entry, and transmits the source write data included in the source work queue entry to the destination chip through an Ethernet link; The destination RDMA network interface card receives the source write data through the Ethernet link, and stores the source write data in the destination receive buffer. Each piece of the source write data is stored in an entry of the destination receive buffer; The destination decomposition module decomposes the write data in each entry of the destination receive buffer into multiple read commands, and sends the multiple read commands to the destination AXI bus; The destination AXI bus sends the multiple read commands to the destination memory and reads back the data, and then sends the read-back data to the destination transmit buffer; The destination merge generation module merges the read-back data to generate a destination work queue entry; The destination RDMA network interface card reads the destination work queue entry, and transmits the destination write data included in the destination work queue entry to the source chip through the Ethernet link; The source RDMA network interface card receives the destination write data; The source decomposition module decomposes the destination write data into multiple read data, and transmits the multiple read data to the source graphics processor.
[0009] According to another aspect of the present invention, there is provided an electronic device, including: one or more processors and a memory, wherein the memory is used for storing data; the one or more processors are configured to implement the above-mentioned method by reading and processing the data in the memory.
[0010] According to still another aspect of the present invention, there is provided a computer network communication device, which includes a storage medium and a data processing medium. When the data in the storage medium is executed by a processor, the processor is caused to execute the above-mentioned method.
[0011] As can be seen from the above technical solutions, the technical solutions provided by the present invention have at least the following advantages: By combining multiple (number of pens ≤ 8) graphics processor load commands transmitted to the Advanced eXtensible Interface bus into multiple read commands, and then merging them into a single work queue item with the RDMA semantics of write, and using multiple (number of pens ≤ 8) read addresses and related read information as write data, and sending the write data to the Ethernet link through a Remote Direct Memory Access network interface card, and then decomposing the aforementioned write data into multiple read commands on the Advanced eXtensible Interface bus at the target end, it is possible to improve the bandwidth utilization rate of RDMA and save the resources of the work queue (a memory dedicated to the graphics processor for loading / storing) in the case where the graphics processor sends a large amount of load data based on the Remote Direct Memory Access protocol. The technical solution provided by the present invention is applicable to scenarios where there is a large amount of graphics processor load command transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings are used to provide a further understanding of the technical solutions of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, but do not constitute a limitation to the technical solutions of the present invention.
[0013] Figure 1 It shows a flowchart of the processing method provided in an exemplary embodiment of the present invention; Figure 2 It shows a structural diagram of a source-end read command in an exemplary embodiment of the present invention; Figure 3 It shows a schematic diagram of the system composition in an exemplary embodiment of the present invention; Figure 4 It shows a schematic diagram of a target-end Remote Direct Memory Access network interface card in an exemplary embodiment of the present invention; Figure 5 It shows a schematic diagram of a source-end Remote Direct Memory Access network interface card in an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. The description of the exemplary embodiments is merely illustrative and is not intended to impose any limitation on the present invention and its application or use. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make the present invention thorough and complete and to fully convey the scope of the present invention to those skilled in the art.
[0015] Unless otherwise specified, if the number of elements is not specifically defined, the element can be one or more. The term "a plurality of" means two or more, the term "based on" should be interpreted as "at least partially based on", and the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations.
[0016] Please refer to Figure 1 , which shows a flowchart of the processing method provided by the exemplary embodiment of the present invention.
[0017] Specifically, the present invention provides a method for processing load commands based on the Remote Direct Memory Access (RDMA) protocol. The processing method includes: Initiate multiple load commands through the source graphics processor, and convert the multiple load commands into multiple source read commands through the source Advanced eXtensible Interface (AXI) bus. The multiple source read commands include multiple queue pairs; Merge the source read commands located in the same queue pair to generate source work queue entries; Transmit the source write data included in the source work queue entries to the target end through an Ethernet link; Store the source write data in the target end receive buffer. Each of the source write data is stored in an entry of the target end receive buffer; Decompose the write data in each entry of the target end receive buffer to obtain multiple read commands; Send the multiple read commands to the target end AXI bus, and send the multiple read commands to the target end memory through the target end AXI bus and read back the data; Send the read-back data to the target end transmit buffer through the target end Advanced eXtensible Bus (AXB); Merge the read-back data to generate target work queue entries; Read the target work queue entries, and transmit the target write data included in the target work queue entries to the source end through an Ethernet link; Receive the target write data; Decompose the target write data to obtain multiple read data, and transmit the multiple read data to the source graphics processor.
[0018] In an exemplary embodiment, the semantics of the source-side work queue entry and the destination-side work queue entry are write operations.
[0019] In an exemplary embodiment, the source-side read command includes an address, a read identifier (RID), and a beat. The source-side read command is stored in multiple entries of the source-side transmit buffer. Among them, when Beat is 0, it means the read data is 64B, and when Beat is 1, it means the read data is 128B, and the 64B data is placed at the low address bit. Since the read operation carries the read identifier during the transmission process on the Ethernet link, the memory and the Remote Direct Memory Access Controller (RDMA Controller) can jointly ensure the sequentiality of the read operation during the entire transmission process and no out-of-order situation will occur. Specifically, the Remote Direct Memory Access network interface card can ensure that the work queue entries sent and received are not out of order. The read command sent by the source side can resolve the sequence number of each queue pair according to the read address, and then send it to the corresponding queue pair of the RDMA. At the same time, each source-side read command will carry the RID as the read identifier, and these read commands will be sent to the destination side as write operations after being merged. The destination side parses the write data and then decomposes it to generate read commands. For example, for the destination-side read command arid = RID + QP, it means that the read address identifier is set to the read identifier plus the sequence number of the queue pair. At the memory side, a reorder buffer is set up to ensure the order of reads. The read-back data can be decomposed into the read identifier and the queue pair sequence number, and then transmitted to the corresponding queue pair of the destination side, and then merged into write data and sent back to the source side. The entire path identifies each read command through the read address identifier, the read data identifier, and the sequence number of the queue pair, and realizes the order preservation by correctly identifying these identifiers and sequence numbers through the Remote Direct Memory Access network interface card and the memory.
[0020] In an exemplary embodiment, when transmitting through the Ethernet link, a padding character (pad) is added at the end of each source-side read command to make each source-side read command 8B aligned. An exemplary structure is as Figure 2 shown.
[0021] In an exemplary embodiment, within a fixed time window, the source-side read commands located in the same queue pair are taken out from multiple entries of the source-side transmit buffer in an indexed manner for merging operations, and the number of source-side read commands n merged each time is less than or equal to 8.
[0022] Please refer to Figures 3 to 5 together, which respectively show schematic diagrams of the system, the destination-side Remote Direct Memory Access network interface card, and the source-side Remote Direct Memory Access network interface card provided by the exemplary embodiments of the present invention.
[0023] According to another aspect of the present invention, a loading command processing system based on the Remote Direct Memory Access (RDMA) protocol is provided. The system includes a source chip and a destination chip. The source chip includes a source graphics processing unit, a source Advanced Extensible Interface (AXI) bus, a source Remote Direct Memory Access network interface card, a source merge generation module, and a source decomposition module. The destination chip includes a destination Remote Direct Memory Access network interface card, a destination decomposition module, a destination memory, a destination merge generation module, and a destination Advanced Extensible Interface bus.
[0024] In an exemplary embodiment, the overall processing flow of the loading command by the system provided by the present invention is as follows: 1. The source graphics processing unit initiates multiple loading commands, and converts the multiple loading commands into multiple source read commands through the source Advanced Extensible Interface bus. The multiple source read commands include multiple queue pairs.
[0025] 2. The source merge generation module merges the source read commands in the same queue pair to generate a source work queue entry. In an exemplary embodiment, within a fixed time window, the source read commands in the same queue pair are retrieved from multiple entries in the source send buffer by indexing for the merging operation, and the number of source read commands merged each time, n, is less than or equal to 8.
[0026] 3. The source Remote Direct Memory Access network interface card reads the source work queue entry, and transmits the source write data included in the source work queue entry to the destination chip through an Ethernet link. In an exemplary embodiment, when transmitting through the Ethernet link, a padding character (pad) is added at the end of each source read command in the source work queue entry to make each source read command 8B-aligned.
[0027] 4. The destination Remote Direct Memory Access network interface card receives the source write data through the Ethernet link and stores the source write data in the destination receive buffer. Each source write data is stored in an entry in the destination receive buffer.
[0028] 5. The destination decomposition module decomposes the write data in each entry of the destination receive buffer into multiple read commands, and sends the multiple read commands to the destination Advanced Extensible Interface bus.
[0029] 6. The destination Advanced Extensible Interface bus sends the multiple read commands to the destination memory and reads back the data, and then sends the read-back data to the destination send buffer.
[0030] 7. The destination merge generation module merges the read-back data to generate a destination work queue entry.
[0031] 8. The destination - side Remote Direct Memory Access (RDMA) network interface card reads the destination - side work queue entry and transfers the destination - side write data included in the destination - side work queue entry to the source - side chip through an Ethernet link.
[0032] 9. The source - side Remote Direct Memory Access (RDMA) network interface card receives the destination - side write data.
[0033] 10. The source - side decomposition module decomposes the destination - side write data into multiple read data and transfers the multiple read data to the source - side graphics processing unit.
[0034] In an exemplary embodiment, the source - side write data is stored in the destination - side receive buffer.
[0035] In an exemplary embodiment, the source - side write data is stored in the destination - side receive First - In - First - Out (RX FIFO) buffer.
[0036] In an exemplary embodiment, the size of the write data (Writepayload) corresponding to the source - side work queue entry with the semantics of a write operation is 8B×n (n is the number of source - side read commands merged). In addition, when transmitting through the Ethernet link, the Remote Direct Memory Access network interface card reads the source - side work queue entry, packs the write data according to the InfiniBand protocol, and adds a header field. Some information that needs to be carried to the destination - side, such as the merge count information, can use the unused fields in the base transport header (BTH), such as the partition key (p_key) field, as part of the header field to be carried to the destination - side.
[0037] It should be understood that Figures 3 to 5 the system shown in, as well as the structures of the destination - side Remote Direct Memory Access network interface card and the source - side Remote Direct Memory Access network interface card, can correspond to the processing methods described in the previous part of this specification. Thus, the operations, features, and advantages described above for the processing methods also apply to the system provided by the present invention and its components and modules included. The operations, features, and advantages described above for the system, the structures of the destination - side Remote Direct Memory Access network interface card and the source - side Remote Direct Memory Access network interface card, and their components and modules also apply to the processing methods provided by the present invention. For the sake of brevity, some operations, features, and advantages will not be described again.
[0038] Although the specific functions have been discussed above with reference to specific modules, it should be noted that the functions of each module in the technical solution of the present invention can also be implemented by dividing them into multiple modules, and / or at least some functions of multiple modules can be combined into a single module for implementation. The manner in which a specific module in the technical solution of the present invention performs an action includes that the specific module itself performs the action, or the specific module calls or otherwise accesses another module that performs the action (or performs the action in combination with the specific module). Therefore, the specific module that performs the action may include the specific module itself that performs the action and / or another module that performs the action and is called or otherwise accessed by the specific module.
[0039] In addition to the above technical solution, the present invention also provides an electronic device, which includes one or more processors and a memory for storing data. Among them, the one or more processors are configured to implement the above method by reading data in the processing memory.
[0040] The present invention also provides a computer network communication device, which includes a storage medium and a data processing medium. When the data in the storage medium is executed by a processor, the processor is caused to execute the above method.
[0041] The technical solution described in the present invention is not limited to the specific examples of the described technical solution. The descriptions and explanations made on the present invention in the foregoing and the accompanying drawings are not restrictive. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can also be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, the scope of protection required by the present invention is defined by the claims rather than the above description, and all changes falling within the meaning and scope of the equivalent elements of the claims are covered by the protection scope of the present invention.
Claims
1. A method for processing a load command based on a Remote Direct Memory Access (RDMA) protocol, characterized in that, The processing method includes: Initiating multiple load commands through a source - side graphics processor, and converting the multiple load commands into multiple source - side read commands through a source - side Advanced Extensible Interface (AXI) bus, where the multiple source - side read commands include multiple queue pairs; Merging the source - side read commands in the same queue pair to generate source - side work queue entries; Transmitting the source - side write data included in the source - side work queue entries to the target side through an Ethernet link; Storing the source - side write data in a target - side receive buffer, where each of the source - side write data is stored in an entry of the target - side receive buffer; Decomposing the write data in each entry of the target - side receive buffer to obtain multiple read commands; Sending the multiple read commands to the target - side AXI bus, and sending the multiple read commands to the target - side memory through the target - side AXI bus and reading back the data; Sending the read - back data to the target - side transmit buffer through the target - side AXI bus; Merging the read - back data to generate target - side work queue entries; Reading the target - side work queue entries, and transmitting the target - side write data included in the target - side work queue entries to the source side through the Ethernet link; Receiving the target - side write data; Decomposing the target - side write data to obtain multiple read data, and transmitting the multiple read data to the source - side graphics processor.
2. The processing method according to claim 1, characterized in that The semantics of the source - side work queue entries and the target - side work queue entries are write operations.
3. The processing method according to claim 1, characterized in that The source - side read commands include an address, a read identifier, and a beat, and the source - side read commands are stored in entries of a source - side transmit buffer.
4. The processing method according to claim 1, characterized in that, When transmitting through the Ethernet link, a padding character is added at the end of each source - side read command to make each source - side read command 8B - aligned.
5. The processing method according to claim 3, characterized in that The source - side read commands in the same queue pair are taken out from the entries of the source - side transmit buffer through an index within a fixed time window for merging, and the number of source - side read commands merged each time is less than or equal to 8.
6. A loading command processing system based on the Remote Direct Memory Access protocol, characterized in that, The system includes: A source - side chip, which includes a source - side graphics processor, a source - side AXI bus, a source - side Remote Direct Memory Access (RDMA) network interface card, a source - side merge and generation module, and a source - side decomposition module; A target - side chip, which includes a target - side RDMA network interface card, a target - side decomposition module, a target - side memory, a target - side merge and generation module, and a target - side AXI bus; Wherein, the source - side graphics processor initiates multiple load commands, and converts the multiple load commands into multiple source - side read commands through the source - side AXI bus, and the multiple source - side read commands include multiple queue pairs; The source - side merge and generation module merges the source - side read commands in the same queue pair to generate source - side work queue entries; The source - side RDMA network interface card reads the source - side work queue entries, and transmits the source - side write data included in the source - side work queue entries to the target - side chip through an Ethernet link; The target - side Remote Direct Memory Access (RDMA) network interface card receives the source - side write data through the Ethernet link and stores the source - side write data in the target - side receive buffer. Each piece of the source - side write data is stored in an entry of the target - side receive buffer; The target - side decomposition module decomposes the write data in each entry of the target - side receive buffer into multiple read commands and sends the multiple read commands to the target - side Advanced Extensible Interface (AXI) bus; The target - side AXI bus sends the multiple read commands to the target - side memory, reads back the data, and then sends the read - back data to the target - side transmit buffer; The target - side merge and generation module merges the read - back data to generate a target - side work queue entry; The target - side RDMA network interface card reads the target - side work queue entry and transmits the target - side write data included in the target - side work queue entry to the source - side chip through the Ethernet link; The source - side RDMA network interface card receives the target - side write data; The source - side decomposition module decomposes the target - side write data into multiple read data and transmits the multiple read data to the source - side graphics processing unit; 7. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing data; The one or more processors are configured to implement the method according to any one of claims 1 to 5 by reading and processing the data in the memory.
8. A computer network communication device, characterized in that, It includes a storage medium and a data - processing medium. When the data in the storage medium is executed by a processor, the processor executes the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Cache access method, multi-level cache system and computer system
CN109753445A
Signal sending method, routing node, storage medium and electronic equipment
CN118573629A
Multi-path data access method, accelerator card, storage system and readable storage medium
CN118646699A
Intelligent network card and distributed object access method based on intelligent network card
CN120111107A
Remote memory management
US20200192857A1