Remote node control using RDMA

Optimizing the data retrieval process in the network storage system through RDMA connection solves the high latency and bandwidth problems of the initiator nodes when searching and retrieving data, and realizes more efficient data retrieval.

CN120112898APending Publication Date: 2025-06-06HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280039894.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-03
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In network storage systems, the processing time, latency, and network bandwidth required by the initiator node to find and retrieve data is high, especially if the data is not stored locally.

Method used

Through the RDMA connection, the first node receives a read request. If the data is not available, the location of the data in the third node is determined, and the third node sends the data to the second node through the RDMA connection, avoiding the second node requesting data directly from the third node.

Benefits of technology

The processing time, data retrieval delay and network bandwidth requirements of the second node are significantly reduced, and the efficiency of data retrieval is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112898A_ABST
    Figure CN120112898A_ABST
Patent Text Reader

Abstract

The invention relates to controlling a remote node over an RDMA connection. A first node, a second node, and one or more third nodes are provided. The first node is connected to the second node and the third node through the RDMA connection, and the second node is connected to the third node through the RDMA connection. The first node receives a read request for data from the second node, determines at which third node the data is available, and causes the third node to send the data to the second node. The third node transmits the data by receiving an RDMA write request or transmission for a command from the first node. In response to the command that the first node directly writes to a transmit queue at the third node, the third node transmits the data by performing an RDMA write operation or a transmit operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to remote direct memory access (RDMA) used in the field of network computing and storage systems. In particular, the present invention relates to controlling remote RDMA nodes through RDMA connections (ie, by using RDMA). Background Art

[0002] Modern, efficient storage systems are designed to provide access to ever-growing amounts of data. A common design consists of one front-end node to which applications connect, and multiple back-end nodes that access and / or hold the data itself. Applications served by the storage system typically communicate only with the front-end node, so any input / output (I / O) operation requires another additional I / O operation between the front-end node and one or more back-end nodes.

[0003] The front-end node can maintain a "map" of the location of all data in the back-end nodes. Using multiple back-end nodes enables partitioning of data (called "sharding") to reduce access latency, allow concurrent parallel access to data, provide fault tolerance, and / or allow upgrades without interrupting service. Cloud environments are a typical use case for this type of partitioning. Another example is a database cache, where the front-end nodes cache frequently accessed data and the back-end nodes maintain all data.

[0004] Regardless of whether the storage system is centralized or network distributed, writing data to and reading data from the storage system requires the participation of the core processing unit (CPU) of each of the participating nodes running front-end software or back-end software or both.

[0005] Some traditional protocols enable applications to use and access remote non-local storage as if the storage were installed locally. To this end, a target (i.e., a node directly attached to one or more storage devices) can support multiple I / O queues, each of which can support multiple I / O commands. This enables a large degree of parallelization and control of I / O operations. An initiator can send one or more commands to a target, which can control and manipulate one or more remote I / O queues at the target. The target can execute commands on behalf of the initiator.

[0006] Control over remote I / O queues reduces processing and memory requirements on the initiator CPU and network interface card (NIC), respectively, because it enables the target to perform the "heavy lifting." For example, the target can set the priority of servicing I / O queues and I / O commands (e.g., based on factors unknown to the initiator, such as the state of its local storage devices), or can set relative priorities between different initiators without its knowledge (e.g., based on the number of connected initiators, fairness, and QoS).

[0007] RDMA is a technology that allows applications to perform memory access operations on remote memory installed in remote network nodes. The RDMA protocol stack is relatively easy to offload to the RDMA NIC (RDMA NIC, RNIC), thereby reducing the CPU requirements of the node to perform network functions. RDMA is now widely used in modern data centers and computer clusters because it provides low-latency remote memory access operations and high network bandwidth. Summary of the invention

[0008] In view of the above background technology, Figure 1 The following description illustrates some specific problems solved by the present invention.

[0009] Figure 1 An exemplary scenario with three nodes, denoted as "A", "B", and "C", is outlined. Node A runs the application, Node B maintains a map of locations and data, and Node C maintains the data itself. This scenario is based on the assumption that I / O operations in any network-based distributed and / or aggregated memory or storage can be further accelerated by using a delegation node that completely offloads (i.e., avoids) any use of the node's CPU and only distributes the I / O commands to another third-party node. Figure 1 In the scenario, node B is such a delegate node.

[0010] If node A wishes to read data, it sends a read request with a pair {key, A-addr} to node B, where key identifies the requested data and "A-addr" is the address in A's memory space to which A expects to read the data (using RDMA). Node B may have the data (for example, if the data is cached in node B's memory space), so node B can immediately respond to node A's request by sending a read response including the data. However, in some cases, node B does not have the data, but node B knows where the data is. Therefore, node B uses the key and its mapping to locate the data in node C.

[0011] In this case, Node B provides Node A with the location of the requested data in Node C, or responds that it does not hold the data and lets Node A request the data directly from Node C. Due to this role, Node B can be considered a delegation node. Node C can be considered a service node. It is worth noting that Figure 1 The scenario is a general scenario and can represent the case where node A is an application server, node B is a page cache server, and node C is a storage server. In this case, the key can be the page address and the data can be the page content. Node C may not even be a storage server, but another cache level where another server (node ​​D) can be the server that holds the page content.

[0012] Figure 1 The problem with the type of operation shown is the requirement for serial operation of node A. That is, node A first initiates a communication to receive the location of the data, then performs some processing, and then initiates another communication to receive the data itself. The total delay in retrieving the data also depends on the processing time of nodes B and C, and in addition on the roundtrip time (RTT) of the network.

[0013] This causes the CPU of node A to take considerable processing time to find the location of the data and eventually retrieve the data. In addition, in the case of cache server misses, a single read requires high network traffic and bandwidth. And the overall latency for the completion of the read request sent by the initiator node A is quite high.

[0014] It is therefore an object of the present invention to reduce the processing time, latency, and network bandwidth required for an initiator node in a network storage system to find and retrieve data.

[0015] These and other objects are achieved by the present invention as described in the independent claims. Advantageous implementations are further defined in the dependent claims.

[0016] A first aspect of the present invention provides a first node, which is connected to a second node and one or more third nodes respectively through RDMA connections, and the first node is used to: receive a read request for data from the second node; if the requested data is not available at the first node: determine at which node of the one or more third nodes the requested data is available; and enable the determined third node to send the requested data to the second node through the RDMA connection.

[0017] Therefore, the first node does not have to respond to the second node with a negative answer about the data. In addition, the second node does not have to request data from a third node. Therefore, the processing time of the second node, the latency of data retrieval, and the bandwidth required for the second node to find and retrieve data are significantly reduced.

[0018] In an implementation form of the first aspect, the read request includes information identifying and authorizing access to memory in the first node and memory in the second node, an address space of the first node, and an address space of the second node.

[0019] This allows the first node to communicate to the third node relevant information for writing data to the second node.

[0020] In an implementation form of the first aspect, in order to enable the determined third node to provide the requested data to the second node, the first node is used to trigger the third node to write the requested data into the address space of the second node.

[0021] In an implementation form of the first aspect, the first node is further configured to control the RDMA connection from the third node to the second node.

[0022] Therefore, the third node does not have to perform any processing, or only significantly reduced processing.

[0023] In an implementation form of the first aspect, the first node includes an address of a sending queue at the third node for sending from the third node to the second node, and an address of a doorbell register of the sending queue.

[0024] Therefore, the first node can directly trigger the sending queue of the third node to execute writing of the requested data to the second node.

[0025] For example, a send queue is a queue of one or more commands. A first node may be used to attach a command to a send queue of a third node, and may trigger the third node to execute the attached command. It is worth noting that a send queue does not execute the one or more commands, but is used to hold the one or more commands to be processed.

[0026] In one implementation form of the first aspect, in order to trigger the third node to write the requested data, the first node is used to send at least one RDMA write request to the address of the sending queue at the third node and / or to the address of the doorbell register of the sending queue.

[0027] In an implementation form of the first aspect, the at least one RDMA write request includes information identifying and authorizing access to memory in the third node and memory in the second node, an address space of the third node, and the address space of the second node.

[0028] This enables data to be written from the third node to the second node by the third node.

[0029] In an implementation form of the first aspect, the at least one RDMA write request includes one of the following: an RDMA immediate write request for the data and for a notification; an RDMA write request for the data and a send request for a notification; an RDMA write request for the data and an immediate send request for a notification.

[0030] In this way, the second node can obtain information that the requested data retrieval has been completed.

[0031] In one implementation form of the first aspect, in order to enable the determined third node to provide the requested data to the second node, the first node is used to send a read request to the third node, wherein the read request indicates that the requested data will be written by the third node into the address space of the second node.

[0032] In this case, the read request received by the third node from the first node is just like the read request received by the third node from the second node.

[0033] In an implementation form of the first aspect, the read request includes information identifying and authorizing access to the memory in the third node and the memory in the second node and the address space of the second node.

[0034] This enables data to be sent and written directly to the second node by the third node.

[0035] In an implementation form of the first aspect, the first node is further configured to: if the requested data is available at the first node, send the requested data to the second node through the RDMA connection with the second node.

[0036] Thus, if the data is, for example, cached at a first node, the second node can quickly obtain the data.

[0037] In an implementation form of the first aspect, the first node is further configured to: if the requested data is not available at the first node and if the first node cannot determine the third node at which the requested data is available, send a response to the second node indicating that the requested data is not found.

[0038] A second aspect of the present invention provides a second node, which is connected to a first node and one or more third nodes respectively through RDMA connections, and the second node is used to: send a read request for data to the first node; and receive the requested data from one of the third nodes.

[0039] Therefore, if the data is at the third node but not at the first node, the second node does not have to request the data from the third node. Therefore, the processing time of the second node, the latency of data retrieval, and the network bandwidth required for the second node to find and retrieve data are significantly reduced.

[0040] In an implementation form of the second aspect, the read request includes information identifying and authorizing access to memory in the first node and memory in the second node, an address space of the first node, and an address space of the second node.

[0041] In an implementation form of the second aspect, the second node is also used to obtain a completion notification indicating that the requested data has been written to the address space of the second node, wherein the completion notification is obtained by one of the following: after sending the read request, polling the completion queue of the second node to obtain the completion notification; receiving an event indicating that the completion notification can be polled from the completion queue.

[0042] In this way, the second node knows that the data has been received and can end the data retrieval.

[0043] The third aspect of the present invention provides a third node, which is connected to a first node and a second node respectively through an RDMA connection, and the third node is used to: receive an RDMA write request or send for a command from the first node; and in response to the command directly written by the first node to the send queue at the third node, provide the data to the second node by performing an RDMA write operation; or in response to the command directly written by the first node to the send queue at the third node, provide the data to the second node by performing a send operation.

[0044] Therefore, the third node can service the request of the second node even though it did not receive the request directly from the second node, but rather received the request from the first node. In this case, the second node does not have to request data from the third node. Therefore, the processing time of the second node, the latency of data retrieval, and the network bandwidth required for the second node to find and retrieve data are significantly reduced.

[0045] In an implementation form of the third aspect, the third node is used to: perform an immediate write, or a write and send operation, or a write and immediate send operation as the RDMA write operation.

[0046] In an implementation form of the first aspect, the third node is used to execute the command written into the send queue at the third node to provide the data into the address space of the second node.

[0047] In one implementation form of the first aspect, if the first node directly writes the command into the sending queue of the third node, the third node is used to provide the data to the second node without processing at the third node and / or without controlling the receiving queue at the third node for sending from the third node to the second node and the doorbell register of the sending queue.

[0048] This significantly reduces the processing load on the third node CPU.

[0049] In one implementation form of the first aspect, if the third node receives the command by sending from the first node, the third node is used to provide the data to the second node by initiating execution of the operation indicated by the RDMA command to the second node and / or by controlling a receiving queue at the third node to send from the third node to the second node.

[0050] A fourth aspect of the present invention provides a method for a first node, wherein the first node is respectively connected to a second node and one or more third nodes via RDMA connections, the method comprising: receiving a read request for data from the second node; if the requested data is not available at the first node: determining at which of the one or more third nodes the requested data is available; and causing the determined third node to send the requested data to the second node via the RDMA connection.

[0051] The method of the fourth aspect may be extended by an implementation form corresponding to the implementation form of the first node of the first aspect. The method of the fourth aspect and its implementation form may provide the same advantages as described above for the first node of the first aspect and its corresponding implementation form.

[0052] A fifth aspect of the present invention provides a method for a second node, wherein the second node is connected to a first node and one or more third nodes respectively via RDMA connections, the method comprising: sending a read request for data to the first node; and receiving the requested data from one of the third nodes.

[0053] The method of the fifth aspect may be extended by an implementation form corresponding to the implementation form of the second node of the second aspect. The method of the fifth aspect and its implementation form may provide the same advantages as described above for the second node of the second aspect and its corresponding implementation form.

[0054] A sixth aspect of the present invention provides a method for a third node, which is connected to a first node and a second node respectively through an RDMA connection, and the method includes: receiving an RDMA write request or transmission for a command from the first node; and in response to the command directly written by the first node to a sending queue at the third node, providing the data to the second node by performing an RDMA write operation; or in response to the command directly written by the first node to a sending queue at the third node, providing the data to the second node by performing a sending operation.

[0055] The method of the sixth aspect can be extended by an implementation form corresponding to the implementation form of the third node of the third aspect. The method of the sixth aspect and its implementation form can provide the same advantages as described above for the third node of the third aspect and its corresponding implementation form.

[0056] A seventh aspect of the present invention provides a computer program, which includes instructions. When the program is executed by a computer, the instructions cause the computer to execute the method according to any one of the fourth aspect, the fifth aspect or the sixth aspect.

[0057] An eighth aspect of the present invention provides a non-transitory storage medium storing executable program code, which, when executed by a processor, enables execution of the method according to the fourth aspect, the fifth aspect or the sixth aspect.

[0058] In the overview of the above aspects and implementation forms of the present invention, all nodes are RDMA connected to enable sending RDMA requests and receiving and servicing these RDMA requests. The first node (delegating node) can directly access the command queue of the third node (service node or storage node). The first node can directly forward (for example, without CPU involvement) the RDMA request received from the second node (requester node or initiator node) to the third node (storage node) that actually services these requests. The first node can use the metadata of the operation to create a command (for example, a work queue element (WQE)) in a special remote control queue of the third node, so that the third node can immediately service the request without additional information.

[0059] It should be noted that all devices, elements, units and devices described in this application can be implemented in software or hardware elements or any type of combination thereof. All steps performed by the various entities described in this application and the functions to be performed by the various entities described are intended to refer to the corresponding entities being suitable for or used to perform the corresponding steps and functions. Even in the following description of a specific embodiment, the specific functions or steps to be performed by an external entity are not reflected in the description of the specific detailed elements of the entity that performs the specific steps or functions, but it should be clear to the technician that these methods and functions can be implemented in the corresponding software or hardware elements or any type of combination thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The above aspects and implementation forms will be described in the following description of specific embodiments in conjunction with the attached drawings, in which:

[0061] Figure 1 An exemplary implementation of data access with three nodes is shown.

[0062] Figure 2 The first node, the second node and the third node provided by the embodiment of the present invention are shown.

[0063] Figure 3 The conceptual flow of a message for returning data to a second node provided by the first example of the present invention is shown.

[0064] Figure 4 The RDMA operations and fields for returning data to the second node for the first example of the present invention are shown.

[0065] Figure 5 The RDMA operation and fields for returning data to the second node provided by the second example of the present invention are shown.

[0066] Figure 6 The method provided by the present invention and used for the first node is shown.

[0067] Figure 7 The method for the second node provided by the present invention is shown.

[0068] Figure 8 The method provided by the present invention and used for the third node is shown. DETAILED DESCRIPTION

[0069] Figure 2The first node 110, the second node 120 and the third node 130 provided by the embodiment of the present invention are shown. The first node 110 is connected to the second node 120 and one or more third nodes 130 (only one third node 130 is shown as an example) through RDMA connections. The second node 120 is correspondingly connected to the first node 110 through an RDMA connection, and is also connected to the one or more third nodes 130 through an RDMA connection. The third node 130 is correspondingly connected to the first node 110 and the second node 120 through RDMA connections.

[0070] The second node 120 may be a requester node or an initiator node, the first node 110 may be a target node or a delegate node, and the third node 130 may be a service node or a storage node. Each node 110, 120, 130 may be referred to as an RDMA node and may include at least a processor or a processing circuit (e.g., a CPU) and / or an RNIC.

[0071] Typically, each node 110, 120, 130 may include a processor or processing circuit (not shown) for executing, implementing or initiating various operations of the corresponding node 110, 120, 130 as described below. The processing circuit may include hardware and / or the processing circuit may be controlled by software. The hardware may include analog circuits or digital circuits, or both analog circuits and digital circuits. The digital circuit may include components such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP) or a multi-purpose processor. Each node 110, 120, 130 may also include a memory circuit that stores one or more instructions that can be executed by a processor or processing circuit (specifically, executed under the control of software). For example, the memory circuit may include a non-transitory storage medium that stores executable software code, which, when executed by a processor or processing circuit, causes the various operations of the corresponding node 110, 120, 130 to be executed as described below. In one embodiment, the processing circuit includes one or more processors and a non-transitory memory connected to the one or more processors. The non-transitory memory may carry executable program code that, when executed by the one or more processors, causes the respective node 110, 120, 130 to perform, implement or initiate operations or methods as described below.

[0072] Specifically, the second node 120 is used to send a read request 121 for data 131 to the first node 110 . Accordingly, the first node 110 is used to receive a read request 121 for data 131 from the second node 120 .

[0073] If the requested data 131 is available at the first node 110, the first node 110 itself may be used to connect to the second node 120 via the RDMA connection between the first node 110 and the second node 120 ( Figure 1 The requested data 131 is sent back to the second node 120 via a reverse path (not explicitly shown in FIG. 1 ). However, the present invention is particularly concerned with the case where the requested data 131 is not available at the first node 110.

[0074] If the requested data 131 is not available at the first node 110, the first node 110 is used to determine at which of the one or more third nodes 130 the requested data 131 is available (in this case, the first node 110 determines that the data 131 is in the third node 130 shown), and then causes the determined third node 130 to send the requested data 131 to the second node 120 through the RDMA connection between the third node 130 and the second node 120. If the requested data 131 is not available at the first node 110 and if the first node 110 is also unable to determine the third node 130 at which the requested data 131 is available, the first node 110 may be used to send a response to the second node 120 through the RDMA connection between the first node 110 and the second node 120, wherein the response indicates that the requested data was not found.

[0075] For example, in order to enable the determined third node 130 to provide the requested data to the second node 120, the first node 110 may be used to trigger the third node 130 to write the requested data 131 to the address space of the second node 120. In order to trigger the third node 130 to write the requested data 131, the first node 110 may be used to send at least one RDMA write request 111 to at least one of the address of the send queue at the third node 130 and the address of the doorbell register of the send queue. Alternatively, in order to enable the determined third node 130 to provide the requested data to the second node 120, the first node 110 may also be used to send a read request 112 to the third node 130, wherein the read request 112 indicates that the requested data 131 will be written by the third node 130 to the address space of the second node 120.

[0076] Accordingly, the third node 130 is used to receive an RDMA write request 111 or send (which corresponds to a read request 112) for a command from the first node 110. These commands can be directly written by the first node 110 to a send queue at the third node 130. Then, the third node 130 is used to provide data 131 to the second node 120 by performing an RDMA write operation or a send operation in response to the command directly written by the first node 110 to the send queue at the third node 130.

[0077] Accordingly, the second node 120 is used to receive the requested data 131 from the third node 130 through the RDMA connection between the second node 120 and the third node 130. In addition, the second node 120 may obtain a completion notification indicating that the requested data 131 has been written into the address space of the second node 120.

[0078] According to the above, the three nodes 110, 120 and 130 are all connected to each other via RDMA connections. The RDMA connections on each node 110, 120, 130 can be under the same RDMA protection domain to allow, for example, active RDMA verification checks of at least one local key (local key, L_Key) and at least one remote key (remote key, R_Key) of one or more nodes 110, 120, 130.

[0079] In the following, the present invention proposes two examples of implementing the above-mentioned solution of the present invention. The first example uses direct doorbell triggering of queue pairs (QPs), while the second example does not. The first example achieves lower overall latency. For both examples, a management process can be created in the RNIC of the first node 110 to handle requests 111, 112 created by the first node 110 to be sent to the third node 130. For the second example, another management process can be created in the RNIC of the third node 130 to process and convert the request 112 received from the first node 110 with the destination being the second node 120. If the request from the first node 110 is a write request 111 command, the third node 130 does not need such a management process. If the request from the first node 110 is a send command, the send command can hold the converted read request 112, and the management process can be used by the third node 130 because the send can consume RQE from the RQ in the third node 130. This consumption can require a management process to (re)fill RQE in this RQ.

[0080] First example reference Figure 3 and Figure 4Explanation. In the first example, the first node 110 may know the QP used for communication between the third node 130 and the second node 120, in particular, the send queue of the third node 130 for sending to the second node 120. Specifically, the first node 110 may know the location of the send queue in the memory space of the third node 130, may know the size of the send queue, and may know the location of the control and doorbell registers of the send queue at the third node 130. For example, the first node 110 may include the address of the send queue at the third node 130, and may include the address of the doorbell register of the send queue. It is worth noting that the QP of the third node 130 may be stored in the host memory or in the RNIC or in the data processing unit (DPU) memory.

[0081] The first node 110 may be the “owner” of a QP between the second node 120 and the third node 130 at the third node 130. The first node 110 may be operable to directly and exclusively write requests (e.g., WQEs) to the send queue of the QP. For example, the first node 110 may send an RDMA write request 111 to the address of the send queue at the third node 130 and / or the address of the doorbell register of the send queue. The write request 111 may include information identifying and authorizing access to memory in the third node 130 and memory in the second node 120, may include the address space of the third node 130, and may include the address space of the second node 120. The “exclusive access” of the first node 110 may mean that no other process on the third node 130 may modify the state or contents of the send queue.

[0082] The first node 110 may create one or more requests 111, 112 (e.g., WQEs) based on the metadata of the RDMA operation (specifically, in this case, the read request 121) it received from the second node 120. The first node 110 may then be used to ring the doorbell mechanism of the QP at the third node 130 to notify the third node 130 that such new requests 111, 112 (e.g., WQEs) are available. The first node 110 may also obtain all required information about the QP between the second node 120 and the third node 130 at the third node 130 during the RDMA connection establishment phase.

[0083] In addition, the first node 110 may use auxiliary, system-related data structures to convert the request 121 from the second node 120 to the corresponding request 111, 112 (e.g., WQE) to the appropriate third node 130. The implementation of these data structures may affect the latency of the conversion and the amount of time required for the CPU of the first node 110 to complete the conversion.

[0084] The second node 120 may also send only the metadata of the operation as a request 121 to the first node 110, which may then be converted by the first node 110 into one or more requests 111, 112 (appropriate WQE) destined for the appropriate third node 130. For example, an RDMA read operation may have been defined by the InfiniBand (IB) specification (e.g., see InfiniBand Architecture Specification Volume 1, Release 1.4, April 2020) to include only metadata. For another example, a write operation (from the second node 120 to the first node 110) may include only metadata and may be converted by the first node 110 into a read operation in memory to the third node 130.

[0085] Figure 3 An example of a conceptual flow of messages returning data between nodes 110 , 120 , 130 provided by the first example of the present invention is shown, and some internal operations of the nodes 110 , 120 , 130 are shown.

[0086] exist Figure 3 In step 1 of , the second node 120 sends a request 121 for data 131 to the first node 110, which may include the identification key X. In step 2, the first node 110 may perform processing to locate the data 131. If the first node 110 has a cached copy of the data 131, it will immediately return the data 131 to the second node 120. However, in this example, the first node 110 does not hold the data 131, but rather holds information about the location of the data 131 in the third node 130.

[0087] In step 3 , the first node 110 writes a write request 111 (WQE) to a transmission queue at the third node 130 and rings the doorbell of the transmission queue at the third node 130 , wherein the write request instructs the third node 130 to directly write the requested data to the second node 120 .

[0088] In step 4, the third node 130 processes the request 111 (e.g., WQE), which is a write operation to the memory space of the second node 120. The operation may also generate a completion signal for the data request 121 of the second node 120, such as a completion queue entry (CQE). Different methods may be used here (e.g., write immediately, write first and then send, etc.). In step 5, the second node 120 processes the completion signal of the request, thereby indicating that the required data 131 is now stored in the memory space of the second node 120.

[0089] Figure 4 1 is an exemplary description of the RDMA operation and some important RDMA fields for returning the data 131 to the second node 120 provided by the first example of the present invention. Figure 4 and Figure 3 Similar, but shows more RDMA details. Specifically, Figure 4 The read request 121 of the second node 120 is shown to include information identifying and authorizing access to memory in the first node 110 and memory in the second node 120 (e.g., the R_Key of the second node 120 and the first node 110, respectively), the address space of the first node 110 (e.g., the virtual address (VA) and / or offset, and / or data length at the first node 110), and the address space of the second node 120 (e.g., the VA and / or offset at the second node 120).

[0090] Figure 4 It is also shown that the RDMA write request 111 may include information identifying and authorizing access to memory in the third node 130 and memory in the second node 120 (e.g., the L_Key of the third node 130 and the R_Key of the second node 120), the address space of the third node 130 (e.g., the VA and / or offset at the third node 130), and the address space of the second node 120 (e.g., the VA and / or offset at the second node 120).

[0091] Figure 4 It is also shown that sending data 131 from the third node 130 to the second node 120 is accomplished by, for example, an immediate write operation based on information identifying and authorizing access to memory in the second node 120 (e.g., the R_Key of the second node 120) and the address space of the second node 120 (e.g., the VA and / or offset at the second node 120).

[0092] Second example reference Figure 5 Explanation. In the second example, the third node 130 uses its general receiver queue (RQ) to process any incoming requests, or uses an additional RQ (QP) to process requests that arrive indirectly to it. The first node 110 can send all required information to the third node 130 so that it can be processed quickly at the third node 130. Special QPs are managed by a management process on the host or SmartNIC / DPU.

[0093] Figure 5Specifically, an exemplary description of the RDMA operation and some important RDMA fields for returning data to the second node 120 provided by the second example of the present invention is shown. In this second example, the first node 110 does not directly write to the send queue and doorbell of the QP as in the first example, but the first node 110 makes some modifications to the request 121 received from the second node 120 and then relays it to the third node 130. These modifications enable the third node 130 to service the request 121, just as if the request 121 was received directly from the second node 120.

[0094] exist Figure 5 In step 1, the second node 120 sends a sending operation (read request 121) with embedded read information to the first node 110, which has the destination information in the memory space of the second node 120 and the data source identifier in the memory space of the first node 110.

[0095] In step 2, the management process at the first node 110 modifies the request 121 of the second node 120 into a read request 112 to be processed by the third node 130. The first node 110 may replace its own information (e.g., VA and / or offset, R_Key) with the matching information and corresponding context in the memory space of the third node 130.

[0096] In step 3, the management process at the third node 130 converts the incoming transmission with the read request 112 into an immediate write operation from the third node 130 to the second node 120. In this way, the data 131 is provided by the third node 130 to the second node 120.

[0097] The first example of the present invention has the advantage that no CPU involvement is required at the third node 130 and it has the lowest latency. However, the first node 110 may need some information about the QP at the third node 130 (e.g., send queue address, CI pointer, L_Key, etc.). In some cases, the first node 110 may also need synchronization (e.g., read back CI value). The first node 110 may also need to handle errors of the QP of the third node 130.

[0098] The second example of the present invention has the advantage that it is an easy modification to an existing service (e.g., regarding the QP type), and the first node 110 does not require any synchronization because the third node 130 is controlling the send queue that holds the commands that are converted into outbound packets sent from the third node 130 over the network. However, the CPU of the third node 130 may need to process the request (e.g., post a receive queue element (RQE) into the RQ to service the incoming send operation from the first node 110). In addition, the RNIC of the third node 130 may require additional management procedures. The second example of the present invention also results in slightly higher latency than the first example of the present invention.

[0099] In an enhancement of the first example of the present invention, different nodes 110, 120, 130 may each have a different mapping of the WQE structure, for example, because the RNICs are from different manufacturing vendors. In this case, each node may construct the required WQE based on the type and version of the RNIC installed in the node it is intended to execute on.

[0100] Both examples of the present invention may use an immediate write RDMA opcode to send data 131 from the (storage) third node 130 to the (initiator) second node 120. Immediate write may be used to generate a notification to the application on the second node 120 that the data 131 is ready. Other methods of achieving this goal are also possible, for example:

[0101] Use two separate messages: one for writing the data, and then one for sending the notification.

[0102] Use two separate messages: a write for the data, followed by an immediate send for the notification.

[0103] Use two separate messages: a write for the data, followed by an immediate write for the notification.

[0104] Use any action that can trigger a notification after the first action.

[0105] • The second node 120 may poll for completion.

[0106] The third node 130 may also be another cache system with more data access hierarchy levels. For example, the third node 130 may be a storage system in which the data 131 is not in its memory but is stored in an HDD or SSD array. The third node 130 may load the required data into its memory before sending the data 131 directly to the second node 120. Alternatively, if the third node 130 is a storage system that supports the Non Volatile Memory express over Fabrics (NVMe-oF) protocol carried by Fabric, the solution of the present invention may be integrated with the NVMe-oF protocol so that the protocol is responsible for sending data 131 to the second node 120. Any cascade from the third node 130 to another service node may also be used, and such a cascade may occur if specific data 131 (e.g., a block) is temporarily unavailable at the third node 130.

[0107] The solution of the present invention can be extended to remote control of a graphics processing unit (GPU) so that the GPU's post-processing of data 131 is sent directly to the initiator (second node 120), for example, to accelerate video games over a network where rendered images are sent directly to the second node 120 rather than through any type of front-end server.

[0108] The solution of the present invention can also be extended for use with remote management applications, for example, to transfer data from multiple distributed application nodes. The solution of the present invention can also utilize DPUs, and all intended processing can be performed in one or more DPUs instead of on the CPU of the node.

[0109] Figure 6 The method 600 provided by the present invention for the first node 110 is shown. The first node 110 may perform the method 600. The first node 110 is connected to the second node 120 and one or more third nodes 130 respectively through RDMA connections, as described above.

[0110] The method comprises step 601: receiving a read request 112 for data 131 from a second node 120. If the requested data 131 is not available at the first node 110, the method comprises step 602: determining at which of the one or more third nodes 130 the requested data 131 is available, and step 603: causing the determined third node 130 to send the requested data 131 to the second node 120 via an RDMA connection.

[0111] Figure 7The method 700 provided by the present invention for the second node 120 is shown. The second node 120 may execute the method 700. The second node 120 is respectively connected to the first node 110 and one or more third nodes 130 through RDMA connections, as described above.

[0112] The method 700 comprises step 701 : sending a read request 121 for data 131 to the first node 110 , and step 702 : receiving the requested data 131 from one of the third nodes 130 .

[0113] Figure 8 The method 800 provided by the present invention for the third node 120 is shown. The third node 130 may perform the method 800. The third node 130 is connected to the first node 110 and the second node 120 respectively through RDMA connections, as described above.

[0114] The method 800 includes step 801: receiving an RDMA write request 111 or a send 112 for a command from the first node 110. Then, the method 800 includes step 802: in response to the command that the first node 110 directly writes to the send queue at the third node 130, providing the requested data 131 to the second node 120 by performing an RDMA write operation; or the method 800 includes step 803: in response to the command that the first node 110 directly writes to the send queue at the third node 130, providing the requested data 131 to the second node 120 by performing a send operation.

[0115] The solution of the present invention achieves faster I / O operations when more than two nodes participate in the exchange of data 131, thereby achieving faster operation of the application on the initiator node (the second node 120) because the remote data becomes available more quickly. The solution of the present invention reduces the requirements for CPU processing in the requesting node (the second node 120 in this article), and for the first example of the present invention, also reduces the requirements for the service node (the third node 130 in this article). This reduction in CPU usage means that a single delegation node (the first node 110 in this article) can handle and connect more service nodes, thereby improving the scale of the separated distributed storage system. The solution of the present invention reduces overall network traffic because fewer network packets are required for control. The solution of the present invention also achieves more efficient operation of the initiator node because it reduces the processing of negative responses of the type "the requested data is located in another node". Efficiency is also improved by reducing the number of required requests to be generated and the total time required to wait for a response.

[0116] The invention has been described in conjunction with various embodiments as examples and implementations. However, from a study of the drawings, the invention and the independent claims, other variations will be understood and implemented by those skilled in the art in implementing the claimed subject matter. In the claims and in the specification, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single element or other unit may fulfil the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

Claims

1. A first node (110), It is characterized in that The first node is connected to the second node (120) and one or more third nodes (130) respectively via remote direct memory access (RDMA) connections, and the first node (110) is used to: receiving a read request (121) for data (131) from the second node (120); If the requested data (131) is not available at the first node (110): - determining at which node of the one or more third nodes (130) the requested data (131) is available; - causing the determined third node (130) to send the requested data (131) to the second node (120) via an RDMA connection.

2. The first node (110) according to claim 1, It is characterized in that The read request (121) includes information identifying and authorizing access to memory in the first node (110) and memory in the second node (120), an address space of the first node (110), and an address space of the second node (120).

3. The first node (110) according to claim 1 or 2, It is characterized in that In order to enable the determined third node (130) to provide the requested data to the second node (120), the first node (110) is used to trigger the third node (130) to write the requested data (131) into the address space of the second node (120).

4. The first node (110) according to claim 3, It is characterized in that It is also used to control the RDMA connection from the third node (130) to the second node (120).

5. The first node (110) according to claim 3 or 4, It is characterized in that The first node (110) includes an address of a transmission queue at the third node (130) for transmission from the third node (130) to the second node (120), and an address of a doorbell register of the transmission queue.

6. The first node (110) according to any one of claims 3 to 5, It is characterized in that In order to trigger the third node (130) to write the requested data (131), the first node (110) is used to send at least one RDMA write request (111) to the address of the sending queue at the third node (130) and / or to the address of the doorbell register of the sending queue.

7. The first node (110) according to claim 6, It is characterized in that The at least one RDMA write request (111) includes information identifying and authorizing access to memory in the third node (130) and memory in the second node (120), an address space of the third node (130), and the address space of the second node (120).

8. The first node (110) according to claim 6 or 7, It is characterized in that The at least one RDMA write request (111) comprises one of the following: an RDMA immediate write request for the data and for the notification; an RDMA write request for the data and a send request for a notification; An RDMA write request for the data and an immediate send request for notification.

9. The first node (110) according to claim 1 or 2, It is characterized in that In order to enable the determined third node (130) to provide the requested data to the second node (120), the first node (110) is used to send a read request (112) to the third node (130), wherein the read request (112) indicates that the requested data (131) will be written by the third node (130) to the address space of the second node (120).

10. The first node (110) according to claim 9, It is characterized in that The read request (112) includes information identifying and authorizing access to memory in the third node (130) and memory in the second node (120) and the address space of the second node (120).

11. The first node (110) according to any one of claims 1 to 10, It is characterized in that The method is further configured to: if the requested data (131) is available at the first node (110), send the requested data (131) to the second node (120) via the RDMA connection with the second node (120).

12. The first node (110) according to any one of claims 1 to 11, It is characterized in that Also used for: if the requested data (131) is not available at the first node (110) and if the first node (110) cannot determine the third node (130) at which the requested data (131) is available, sending a response to the second node (120) indicating that the requested data was not found.

13. A second node (120), It is characterized in that The second node is connected to the first node (110) and one or more third nodes (130) respectively via remote direct memory access (RDMA) connections, and the second node (120) is used to: Sending a read request (121) for data (131) to the first node (110); The requested data (131) is received from one of the third nodes (130).

14. The second node (120) according to claim 13, It is characterized in that The read request (121) includes information identifying and authorizing access to memory in the first node (110) and memory in the second node (120), an address space of the first node (110), and an address space of the second node (120).

15. The second node (120) according to claim 13 or 14, It is characterized in that Also used for: Obtaining a completion notification indicating that the requested data (131) has been written to the address space of the second node (120), wherein the completion notification is obtained by one of: - after sending the read request (121), polling the completion queue of the second node (120) to obtain the completion notification; - receiving an event indicating that the completion notification can be polled from the completion queue.

16. A third node (130), It is characterized in that The third node is connected to the first node (110) and the second node (120) respectively via a remote direct memory access (RDMA) connection, and the third node (130) is used to: Receiving (111) or sending (112) an RDMA write request for a command from the first node (110); In response to the command by the first node (110) to directly write to the send queue at the third node (130), providing data (131) to the second node (120) by performing an RDMA write operation; or In response to the command by the first node (110) to directly write to the send queue at the third node (130), the data is provided to the second node (120) by performing a send operation.

17. The third node (130) according to claim 16, It is characterized in that Used for: As the RDMA write operation, an immediate write, or a write and send operation, or a write and immediate send operation is performed.

18. The third node (130) according to claim 16 or 17, It is characterized in that The command for executing write to a transmit queue at the third node (130) provides the data (131) into the address space of the second node (120).

19. The third node (130) according to any one of claims 16 to 18, It is characterized in that If the first node (110) directly writes the command into the sending queue of the third node (130), the third node (130) is configured to: The data (131) is provided to the second node (120) without processing at the third node (130) and / or without controlling a receive queue at the third node (130) for transmission from the third node (130) to the second node (120) and a doorbell register of the transmit queue.

20. The third node (130) according to any one of claims 16 to 18, It is characterized in that If the third node (130) receives the command through the transmission (112) from the first node (110), the third node (130) is configured to: The data (131) is provided to the second node (120) by initiating execution of the operation indicated by the RDMA command to the second node (120) and / or by controlling a receiving queue at the third node (130) for transmission from the third node (130) to the second node (120).

21. A method (600) for a first node (110), It is characterized in that The first node is connected to the second node (120) and one or more third nodes (130) respectively via remote direct memory access (RDMA) connections, and the method (600) includes: receiving (601) a read request (112) for data (131) from the second node (120); If the requested data (131) is not available at the first node (110): - determining (602) at which node of the one or more third nodes (130) the requested data (131) is available; - causing (603) the determined third node (130) to send the requested data (131) to the second node (120) via an RDMA connection.

22. A method (700) for a second node (120), It is characterized in that The second node is connected to the first node (110) and one or more third nodes (130) respectively via remote direct memory access (RDMA) connections, and the method (700) includes: Sending (701) a read request (121) for data (131) to the first node (110); The requested data (131) is received (702) from one of the third nodes (130).

23. A method (800) for a third node (130), It is characterized in that The third node is connected to the first node (110) and the second node (120) respectively via remote direct memory access (RDMA) connections, and the method (800) includes: Receiving (801) an RDMA write request (111) or sending (112) for a command from the first node (110); In response to the command by the first node (110) to write directly to the send queue at the third node (130), providing (802) the requested data (131) to the second node (120) by performing an RDMA write operation; or In response to the command by the first node (110) to directly write to the send queue at the third node (130), the requested data (131) is provided (803) to the second node (120) by performing a send operation.

24. A computer program, It is characterized in that The program comprises instructions which, when executed by a computer, cause the computer to perform the method (600, 700, 800) according to any one of claims 21 to 23.