Storage command merging and decomposition method and system

By combining and decomposing storage commands, the problem of low bandwidth utilization of small and medium-sized data packets in RoCE v2 technology is solved, which improves RDMA bandwidth utilization, optimizes system performance and reduces hardware costs.

CN120316058BActive Publication Date: 2025-08-22VASTAI TECH (SHANGHAI) INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510805753.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-22
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The existing RoCE v2 technology has low bandwidth utilization when processing small packet transmissions, especially in cross-chip storage operations of graphics processors, resulting in a significant reduction in RDMA bandwidth utilization, affecting the system performance of applications such as high-performance computing and artificial intelligence training.

Method used

By combining multiple storage commands on the source side as advanced extensible interface bus write operation commands and transmitting the combined work queue items on the Ethernet link, the target side decomposes these commands to achieve data transmission, optimize data transmission sequence and processing logic, and improve bandwidth utilization.

Benefits of technology

It significantly improves bandwidth utilization, saves hardware resources, reduces costs, and optimizes system performance to adapt to data transmission needs of different sizes, ensuring that data is restored to the original write data command at the target end.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316058B_ABST
    Figure CN120316058B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for merging and decomposing storage commands. Specifically, at the source end, the merging method converts multiple small data storage commands issued by a graphics processor into source-end advanced extensible interface bus write operation commands, merges the write data located in the same queue pair, and generates a work queue item; at the target end, the decomposition method decomposes the work queue item and restores it to the original command. The technical solution provided by the present invention can optimize remote direct memory access transmission efficiency and improve bandwidth utilization. In addition, the present invention also adopts a shared buffer design, which not only saves hardware resources, but also can improve overall processing efficiency by optimizing the transmission order.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer hardware and network communication, and in particular to a storage command merging and decomposition method and system based on a remote direct memory access protocol. Background Art

[0002] With the rapid development of data centers and high-performance computing applications, Remote Direct Memory Access (RDMA) technology has gained widespread adoption due to its low latency and high throughput. Specifically, RDMA over converged Ethernet v2 (RoCE v2) significantly improves data transmission efficiency in Ethernet environments by eliminating the overhead of traditional network protocol stacks. However, existing RoCE v2 solutions still have some technical bottlenecks, particularly low bandwidth utilization when handling the transmission of small data packets.

[0003] Specifically, in the traditional RoCE v2 technology solution, a send queue (SQ) stores work queue entries (WQEs). The RoCE Network Interface Card (RNIC) sequentially reads each WQE and performs data transfers based on the transmission attributes and data length defined within it. Because the RDMA reliable connection (RC) model requires strict transmission order, the WQEs in each queue pair (QP) must be processed and sent out in sequence. If most WQEs have small data lengths, then, given the same data bit width, the WQE processing time of the RNIC will be longer than the data fetching time, resulting in bandwidth loss. Assuming the data transmission path and the WQE processing module have the same clock frequency, the WQE processing time is t1, the data transmission time is t2, and the ideal transmission bandwidth is Bandwidth. The resulting bandwidth loss is (t1 - t2) × Bandwidth / t1.

[0004] For example, cross-chip memory operations within a graphics processing unit (GPU) typically involve data sizes of 64 or 128 bytes. In this scenario, bandwidth utilization for data transfers via a remote direct memory access (RDMA) network interface card (NIC) is not maximized. Assuming a 1GHz clock frequency and 512-bit data width for each transaction, a 128-byte Advanced eXtensible Interface (AXI) transaction requires only two clock cycles. However, if the RDMA NIC takes more than two clock cycles to process a work queue item, bandwidth is wasted.

[0005] To address this issue, the read time for a work queue item must be greater than or equal to the processing time. As the size of a work queue item increases, the read time will also increase for a given data bit width. If the read time increases to a value greater than or equal to the processing time, the data transmission bandwidth of the remote direct memory access network interface card can be fully utilized.

[0006] Existing technologies primarily increase bandwidth utilization by increasing the data size of individual work queue items. However, this approach is limited in scenarios where large numbers of small data packets are generated, such as in graphics processing units. In particular, in applications such as artificial intelligence training and high-performance computing, frequent communication of small data packets can significantly reduce RDMA bandwidth utilization, impacting overall system performance. Therefore, a technical solution that can effectively improve RDMA bandwidth utilization in small data packet scenarios is urgently needed. Summary of the Invention

[0007] In view of this, the present invention provides a method and system for merging and decomposing storage commands based on a remote direct memory access protocol, so as to solve the above-mentioned technical problems in the prior art.

[0008] According to one aspect of the present invention, a storage command merging method is provided, wherein the merging method includes:

[0009] Initiating multiple storage commands through a source-side graphics processor, and converting the multiple storage commands into source-side Advanced eXtensible Interface bus write operation commands and transmitting them to a source-side remote direct memory access network interface card, wherein the source-side Advanced eXtensible Interface bus write operation commands include multiple queue pairs, and the multiple queues include multiple write data;

[0010] Merge the write data in the same queue pair to generate merged write data;

[0011] A work queue item is generated based on the merged write data and is transmitted to the target end via an Ethernet link.

[0012] According to one aspect of the present invention, a storage command decomposition method is provided, wherein the decomposition method includes:

[0013] receiving and storing the work queue items generated according to the aforementioned merging method via the Ethernet link;

[0014] Decompose the work queue item and send the decomposed write operation command to the target storage structure through the target advanced extensible interface bus.

[0015] According to another aspect of the present invention, a storage command processing system is provided, wherein the system includes:

[0016] The source chip includes a source graphics processor, a source advanced extensible interface bus, a source remote direct memory access network interface card, and a source write-merging module;

[0017] The source-side graphics processor initiates multiple storage commands via the source-side Advanced eXtensible Interface bus, converts the multiple storage commands into source-side Advanced eXtensible Interface bus write operation commands, and transmits them to the source-side remote direct memory access network interface card, wherein the source-side Advanced eXtensible Interface bus write operation commands include multiple queue pairs, and the multiple queues include multiple write data;

[0018] The source-side write-merging module merges write data in the same queue pair to generate merged write data, and generates a work queue item based on the merged write data.

[0019] The remote direct memory access network interface card at the source side receives the work queue item and transmits the work queue item to the target side via the Ethernet link;

[0020] A target-side chip, the target-side chip including a target-side remote direct memory access network interface card, a target-side write decomposition module, and a target-side advanced extensible interface bus;

[0021] Wherein, the target-side remote direct memory access network interface card receives and stores the work queue item via the Ethernet link;

[0022] The target-side write decomposition module decomposes the work queue items and sends the decomposed write operation commands to the target-side storage structure through the target-side advanced extensible interface bus.

[0023] According to another aspect of the present invention, an electronic device is provided, comprising: one or more processors and a memory, wherein the memory is used to store data; and the one or more processors are configured to implement the above method by reading and processing the data in the memory.

[0024] According to yet another aspect of the present invention, a computer network communication device is provided, which includes a storage medium and a data processing medium. When the data in the storage medium is executed by the processor, the processor executes the above method.

[0025] It can be seen from the above technical solutions that the technical solution provided by the present invention has at least the following advantages:

[0026] 1. Significantly improve bandwidth utilization: By combining multiple small data items into larger work queue items, data transmission time is extended to match the processing time of the work queue items, thereby maximizing RDMA bandwidth utilization;

[0027] 2. Save hardware resources and reduce costs: The shared buffer design reduces the overall storage area, is applicable to application scenarios with multiple queue pairs, and reduces hardware costs;

[0028] 3. Optimize system performance: The Ethernet link uses a transmission method that transmits the data portion first and then the address portion and the selection signal portion. This simplifies the processing logic at the receiving end, facilitates alignment and data retrieval at the receiving end, and improves system efficiency.

[0029] 4. Flexibility and compatibility: Supports dynamic merging of up to 8 commands to adapt to data transmission requirements of different sizes. The write decomposition module ensures that data is restored to the original write data command on the target end without affecting the original function. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention, but do not constitute a limitation to the technical solution of the present invention.

[0031] Figure 1 A flow chart showing a merging method provided in an exemplary embodiment of the present invention is shown;

[0032] Figure 2 A schematic diagram showing a write data format in an exemplary embodiment of the present invention is shown;

[0033] Figure 3 A flow chart showing a decomposition method provided in an exemplary embodiment of the present invention is shown;

[0034] Figure 4 A schematic diagram showing the system composition in an exemplary embodiment of the present invention is shown;

[0035] Figure 5 A structural block diagram of an electronic device provided by an exemplary embodiment of the present invention is shown. DETAILED DESCRIPTION

[0036] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. The description of the exemplary embodiments is merely illustrative and is not intended to limit the invention, its application, or use. The present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make this disclosure thorough and complete and to fully convey the scope of the invention to those skilled in the art.

[0037] Unless otherwise indicated, if the number of an element is not specifically limited, the element may be one or more. The term "plurality" means two or more, the term "based on" should be interpreted as "based at least in part on", and the terms "and / or" and "at least one of..." cover any one and all possible combinations of the listed items.

[0038] Please refer to Figure 1 , which shows a flow chart of a merging method provided by an exemplary embodiment of the present invention.

[0039] Specifically, the present invention provides a storage command merging method based on a remote direct memory access protocol, wherein the merging method includes:

[0040] Initiating multiple storage commands through a source-side graphics processor, and converting the multiple storage commands into source-side Advanced eXtensible Interface bus write operation commands and transmitting them to a source-side remote direct memory access network interface card, wherein the source-side Advanced eXtensible Interface bus write operation commands include multiple queue pairs, and the multiple queues include multiple write data;

[0041] Merge the write data in the same queue pair to generate merged write data;

[0042] A work queue item is generated based on the merged write data and is transmitted to the target end via an Ethernet link.

[0043] In the technical solution provided by this invention, the graphics processor of the source chip initiates multiple data storage commands. These storage commands are transmitted as AEXI bus write commands on the source's AEXI bus to the source's remote direct memory access network interface card (RDMA). The source's RDMA network interface card receives the multiple AEXI bus write data transmissions and stores them in a send buffer. Within a queue pair, n (n <= 8) data writes stored in the send buffer trigger a merge operation within a fixed time window or when the number of data reaches 8. After merging, the merged data is transmitted to the target as a work queue item and then decomposed and restored on the target. This ensures that the data transmission time of a work queue item is greater than or equal to the processing time of the work queue item. Specifically, a maximum of eight data items can be merged at a time, within a fixed time window of 32 clock cycles. If eight data items are received before 32 clock cycles (for example, 16 clock cycles), the merge of these eight data items is triggered directly. If n (n <= 8) data items are received after 32 clock cycles, the merge operation is also triggered. In other words, the conditions for triggering a merge are: (1) when the fixed time window is not full and eight records are received, the merge operation is triggered directly; or (2) when the fixed time window expires, the merge operation is triggered regardless of the number of records received. This can avoid the problem of system efficiency being reduced due to long waits for data to be merged.

[0044] In an exemplary embodiment, a source-side remote direct memory access network interface card receives a source-side advanced extensible interface bus write operation command and stores multiple write data in a send buffer. The send buffer is a shared storage structure in which queues stored therein are distinguished and merged by indexes. Each write data in the send buffer includes an address portion (Address), a strobe signal portion (Strobe), and a data portion (Data). The size of the data portion is fixed at 128B, and the strobe signal can be used to select 64B or 128B for data transmission according to actual conditions. The specific storage format is as follows: Figure 2 shown.

[0045] In an exemplary embodiment, source-side AE ​​bus write operation commands are merged within a fixed time window.

[0046] In one exemplary embodiment, when a work queue item is transmitted over an Ethernet link, the data portion is transmitted first, followed by the address portion and the strobe portion. This transmission method achieves 64B alignment and facilitates the reading of the address, strobe, and data portions of the write data by the target end.

[0047] Please refer to Figure 3 , which shows a flow chart of a decomposition method provided by an exemplary embodiment of the present invention.

[0048] Specifically, the present invention further provides a storage command decomposition method based on the remote direct memory access protocol, wherein the decomposition method includes:

[0049] receiving and storing the work queue items generated according to the aforementioned merging method via the Ethernet link;

[0050] Decompose the work queue item and send the decomposed write operation command to the target storage structure via the target advanced extensible interface bus, wherein the work queue item is stored in the receive buffer.

[0051] Please refer to Figure 4 , which shows a schematic diagram of system composition in an exemplary embodiment of the present invention.

[0052] Specifically, the present invention also provides a storage command processing system, which includes a source chip and a target chip.

[0053] Specifically, the source chip includes a source graphics processor, a source advanced extensible interface bus, a source remote direct memory access network interface card, and a source write-merge module.

[0054] The source-side graphics processor initiates multiple storage commands via the source-side Advanced eXtensible Interface bus, converts the multiple storage commands into source-side Advanced eXtensible Interface bus write operation commands, and transmits them to the source-side remote direct memory access network interface card, wherein the source-side Advanced eXtensible Interface bus write operation commands include multiple queue pairs, and the multiple queues include multiple write data;

[0055] The source-side write-merging module merges write data in the same queue pair to generate merged write data, and generates a work queue item based on the merged write data.

[0056] The remote direct memory access network interface card on the source side receives the work queue item and transmits the work queue item to the target side through the Ethernet link.

[0057] The target chip includes a target-side remote direct memory access (RDMA) network interface card (NIC), a target-side write-decomposition module (WDM), and a target-side advanced extensible interface (AEI) bus. The RDMA NIC receives and stores work queue items via an Ethernet link. The WDM module decomposes work queue items and sends the resulting write operation commands to the target storage structure via the AEI bus.

[0058] In an exemplary embodiment, a remote direct memory access network interface card on the source side receives a source side Advanced eXtensible Interface bus write operation command and stores multiple write data in a sending buffer.

[0059] It is understood that to save overall storage space, the send buffer can be designed as a shared storage structure. That is, the write data information of all queue pairs is stored in different work items in the same send buffer. The write-merging module can use indexes to distinguish and merge these queue pairs stored in the same write data buffer.

[0060] The specific process of merging and decomposing write data by the storage command processing system provided by the present invention is as follows:

[0061] 1. In the source chip, the graphics processor issues multiple storage commands, which are sent to the remote direct memory access network interface card through the Advanced eXtensible Interface bus as Advanced eXtensible Interface bus write operations;

[0062] 2. The send buffer in the remote direct memory access network interface card stores the data / select signal / address from the Advanced Extensible Interface bus into each send buffer work item. The write data information of all queue pairs is stored in different work items of the same send buffer. The write merge module selects and merges the write data based on the index.

[0063] 3. The write-merge module merges n (n<=8) write data items of the same queue pair within a fixed time window and generates a work queue item and sends it to the remote direct memory access network interface card;

[0064] 4. The remote direct memory access network interface card reads and parses the work queue items and sends them to the Ethernet link. Each work queue item is transmitted on the Ethernet link first as the data portion, followed by the address portion and the selection signal portion.

[0065] 5. In the target chip, the receive buffer stores the received data part / select signal part / address part to each work item;

[0066] 6. The write decomposition module decomposes the data part / select signal part / address part in each work item and sends the decomposed write data command to the target end storage structure.

[0067] It should be understood that Figure 4 The system shown in the figure may correspond to the merging and decomposing method described earlier in this specification. Thus, the operations, features, and advantages described above for the merging and decomposing method are also applicable to the system provided by the present invention and its components and modules, and the operations, features, and advantages described above for the system and its components and modules are also applicable to the merging and decomposing method provided by the present invention. For the sake of brevity, certain operations, features, and advantages will not be described in detail.

[0068] While specific functions have been discussed above with reference to specific modules, it should be noted that the functions of each module in the technical solution of the present invention may also be implemented as multiple modules, and / or at least some functions of multiple modules may be combined into a single module for implementation. The manner in which a specific module in the technical solution of the present invention performs an action includes the specific module itself performing the action, or being called or otherwise accessed by the specific module to perform the action (or performing the action in conjunction with the specific module). Therefore, the specific module that performs the action may include the specific module itself that performs the action and / or another module that is called or otherwise accessed by the specific module to perform the action.

[0069] In addition to the above technical solutions, the present invention also provides an electronic device, which includes one or more processors and a memory for storing data. The one or more processors are configured to implement the above method by reading and processing data in the memory. The present invention also provides a computer network communication device, which includes a storage medium and a data processing medium. When the data in the storage medium is executed by the processor, the processor uses the above method. In the following part of this specification, Figure 5 To describe the aforementioned illustrative examples of electronic devices and computer network communication devices.

[0070] Figure 5 An example configuration of an electronic device 300 that can be used to implement the methods described herein is shown. The technical solutions of the present invention can also be implemented in whole or in part by electronic device 300 or similar devices / systems. Electronic device 300 can be a variety of different types of devices. Examples of electronic devices 300 include, but are not limited to, desktop computers, server computers, laptop or netbook computers, mobile devices, wearable devices, entertainment devices, televisions or other display devices, and automotive computers.

[0071] The electronic device 300 may include at least one processor 302, memory 304, communication interface(s) 309, a display device 301, other input / output (I / O) devices 310, and one or more mass storage devices 303, all capable of communicating with each other via a system bus 311 or other appropriate connections.

[0072] The processor 302 may be a single or multiple processing units, all of which may include a single or multiple computing units or multiple cores. The processor 302 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, graphics processors, state machines, logic circuits, and / or any device that manipulates signals based on operational instructions. Among other capabilities, the processor 302 may be configured to retrieve and execute computer-readable instructions stored in the memory 304, mass storage device 303, or other computer-readable media, such as program code of an operating system 305, application programs 306, or other programs 307.

[0073] Memory 304 and mass storage device 303 are examples of computer-readable storage media for storing instructions. When processor 302 issues instructions to be executed and the stored command data is transmitted via communication interface 309, the data merging and decomposition functions are performed by the hardware devices described above. For example, memory 304 may generally include both volatile memory and non-volatile memory. Furthermore, mass storage device 303 may generally include a hard disk drive, a solid-state drive, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks, storage arrays, network attached storage, storage area networks, and the like. In the present invention, a buffer is provided between communication interface 309 and bus 311 to accelerate data exchange by reading buffer data and performing merging and decomposition.

[0074] A plurality of programs may be stored on the mass storage device 303. These programs include an operating system 305, one or more application programs 306, other programs 307, and program data 308, and they may be loaded into the memory 304 for execution. Examples of such applications or program modules may include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functions: the methods provided by the present invention (including any suitable steps of the methods) and / or other embodiments described herein.

[0075] Although Figure 5304 of the electronic device 300, but the modular operating system 305, application programs 306, other programs 307, and program data 308, or portions thereof, may be implemented using any form of computer-readable media accessible by the electronic device 300. Here, a computer-readable medium may be any available computer-readable storage medium or communication medium accessible to a computer. Communication media include media such as communication signals for transmitting computer-readable instructions, data structures, program modules, or other data from one system to another. Communication media may include guided transmission media and wireless media capable of propagating energy waves. Computer-readable instructions, data structures, program modules, or other data may be embodied as, for example, modulated data signals in a wireless medium.

[0076] For example, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memory, such as random access memory (RAM, DRAM, SRAM); and non-volatile memory, such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other known media or later developed media capable of storing computer-readable information / data for use by a computer system.

[0077] One or more communication interfaces 309 are used to exchange data with other devices, such as via a network or direct connection. This communication interface can be one or more of the following: any type of network interface, wired or wireless (e.g., WLAN) interface, Wi-MAX interface, Ethernet interface, USB interface, cellular network interface, Bluetooth interface, NFC interface, etc. Communication interface 309 can facilitate communication within a variety of network and protocol types, including wired and wireless networks, the Internet, etc. Communication interface 309 can also provide communication with external storage devices (not shown) such as storage arrays, network-attached storage, and storage area networks. In the present invention, communication interface 309 cooperates with on-chip storage and the merging and decomposition unit to achieve inter-network communication.

[0078] In some examples, a display device 301 such as a monitor may be included for displaying information and images to a user. Other I / O devices 310 may be devices that receive user input and provide output to the user, and may include touch / gesture input devices, cameras, keyboards, remote controls, mice, audio input / output devices, etc.

[0079] In an exemplary embodiment, the present invention can be applied to artificial intelligence (AI) reasoning in a data center. The process for starting AI reasoning in a data center is as follows:

[0080] 1. Load the operating system and corresponding applications into the memory 304 and initialize the computer or server. During this process, each chip in the server will be initialized and interconnected through the communication interface 309;

[0081] 2. The processor reads the AI ​​model from the mass storage device 303 and performs inference based on the model. Because the inference computing power of a single large model machine is insufficient, it is necessary to interconnect with multiple computers or servers through the communication interface 309 to exchange data and jointly complete the inference work;

[0082] 3. The final result obtained by reasoning is stored in the large-capacity storage device 303.

[0083] In the present invention, the aforementioned exemplary embodiment optimizes the communication interface 309 by adding a storage medium and a corresponding storage data processing medium, thereby enabling data to be processed quickly and accurately when multiple machines are interconnected.

[0084] The technical solutions described in the present invention can be supported by these various configurations of the electronic device 300 and are not limited to the specific examples of the technical solutions described in the present invention. The illustrations and descriptions of the present invention in the foregoing text and the accompanying drawings are not restrictive. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can also be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, the scope of protection claimed by the present invention is defined by the claims rather than the above description, and all changes that fall within the meaning and scope of the equivalent elements of the claims are included in the scope of protection of the present invention.

Claims

1. A storage command processing system based on remote direct memory access protocol, characterized in that: The system comprises: A source chip, comprising a source graphics processor, a source advanced extensible interface bus, a source remote direct memory access network interface card, and a source write-merging module; The source-side graphics processor initiates multiple storage commands via the source-side Advanced eXtensible Interface (AEUI) bus, converts the multiple storage commands into source-side Advanced eXtensible Interface (AEUI) bus write operation commands, and transmits the commands to the source-side remote direct memory access (RDMA) network interface card, wherein the source-side Advanced eXtensible Interface (AEUI) bus write operation commands include multiple queue pairs, and the multiple queues include multiple write data. The source-side write-merging module merges write data in the same queue pair to generate merged write data, and generates a work queue item based on the merged write data; The remote direct memory access network interface card of the source end receives the work queue item and transmits the work queue item to the target end via an Ethernet link; A target-side chip, comprising a target-side remote direct memory access network interface card, a target-side write decomposition module, and a target-side advanced extensible interface bus; wherein the target-side remote direct memory access network interface card receives and stores the work queue item via the Ethernet link; The target-side write decomposition module decomposes the work queue item and sends the decomposed write operation command to the target-side storage structure through the target-side advanced extensible interface bus.

2. The system according to claim 1, wherein: Each of the multiple write data includes an address portion, a strobe signal portion, and a data portion.

3. The system according to claim 2, characterized in that The size of the data portion is 64B or 128B.

4. The system according to claim 1, wherein: When merging write data in the same queue pair, the number of write data records merged each time is less than or equal to 8.

5. The system according to claim 1, wherein: The source-end remote direct memory access network interface card receives the source-end advanced extensible interface bus write operation command and stores the multiple write data in a sending buffer.

6. The system according to claim 5, characterized in that The sending buffer is a shared storage structure, and the write-merge module distinguishes and merges the queue pairs stored in the sending buffer through indexes.

7. The system according to claim 1, wherein: The source-side remote direct memory access network interface card merges the source-side advanced extensible interface bus write operation commands within a fixed time window.

8. The system according to claim 1, wherein: The target-side remote direct memory access network interface card stores the work queue item in a receiving buffer.

9. The system according to claim 2, wherein: When transmitting the work queue item, the remote direct memory access network interface card first transmits the data portion, and then transmits the address portion and the selection signal portion.

10. The system according to claim 1, wherein: When multiple data are transmitted continuously, the data transmission time of the work queue item is greater than or equal to the processing time of the work queue item.

11. A storage command merging method based on remote direct memory access protocol, the merging method being executed by a system according to any one of claims 1 to 10, characterized in that: The merging method includes: Initiating multiple storage commands through a source-side graphics processor, and converting the multiple storage commands into source-side Advanced eXtensible Interface bus write operation commands and transmitting them to a source-side remote direct memory access network interface card, wherein the source-side Advanced eXtensible Interface bus write operation commands include multiple queue pairs, and the multiple queues include multiple write data; Merge the write data in the same queue pair to generate merged write data; A work queue item is generated based on the merged write data, and the work queue item is transmitted to a target end via an Ethernet link.

12. The merging method according to claim 11, characterized in that: Each of the multiple write data includes an address portion, a strobe signal portion, and a data portion.

13. The merging method according to claim 12, characterized in that: The size of the data portion is 64B or 128B.

14. The merging method according to claim 11, characterized in that: When merging write data in the same queue pair, the number of write data records merged each time is less than or equal to 8.

15. The merging method according to claim 11, characterized in that: The source-end remote direct memory access network interface card receives the source-end advanced extensible interface bus write operation command and stores the multiple write data in a sending buffer.

16. The merging method according to claim 15, characterized in that: The sending buffer is a shared storage structure, and the queue pairs stored in the sending buffer are distinguished and merged through indexes.

17. The merging method according to claim 11, characterized in that: The source-side AEIB bus write operation commands are merged within a fixed time window.

18. The merging method according to claim 11, characterized in that: When the work queue item is transmitted on the Ethernet link, the data portion is transmitted first, and the address portion and the selection signal portion are transmitted subsequently.

19. The merging method according to claim 11, characterized in that: When multiple data are transmitted continuously, the data transmission time of the work queue item is greater than or equal to the processing time of the work queue item.

20. A storage command decomposition method based on remote direct memory access protocol, the decomposition method being executed by the system according to any one of claims 1 to 10, characterized in that: The decomposition method comprises: receiving and storing the work queue item generated according to claim 1 via an Ethernet link; The work queue item is decomposed, and the write operation command obtained after the decomposition is sent to the target-side storage structure through the target-side advanced extensible interface bus.

21. The decomposition method according to claim 20, characterized in that: The work queue item is stored in a receive buffer.

22. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing data; The one or more processors are configured to implement the method of any one of claims 11 to 21 by reading and processing the data in the memory.

23. A computer network communication device, characterized in that: The device comprises a storage medium and a data processing medium, and when the data in the storage medium is executed by a processor, the processor is caused to execute the method according to any one of claims 11 to 21.

Citation Information

Patent Citations

  • Command read-write method and device and computer storage medium

    CN110910921A

  • FPGA (Field Programmable Gate Array) virtualization hardware system stack design for cloud deep learning reasoning

    CN113420517A