Adaptive Optimization Method and System for Improving the Scalability of RDMA Unilateral Operations
The thread-aware RDMA resource allocation and credit point mechanism control the depth of the send queue, which solves the problem of message throughput decrease when the number of RDMA unilateral operation threads or coroutines increases, and achieves better scalability and performance stability.
Patent Information
- Application Number
- CN202211569801.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-12-07
AI Technical Summary
In the prior art, when the number of RDMA unilateral operation threads or coroutines increases, the message throughput rate decreases and cannot meet the scalability requirements of the application.
When establishing a connection between the local node and the remote node, RDMA resources are allocated in a thread-aware manner, and a credit point mechanism is used to limit the depth of the sending queue, thereby preventing cache jitter.
When the number of threads or coroutines increases, the RDMA message throughput remains or is basically unchanged, avoiding the negative impact of inter-thread lock synchronization and cache jitter on performance.
Smart Images

Figure CN116185603B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of system software, and particularly to an adaptive optimization method and system for improving the scalability of RDMA one-sided operations. Background Art
[0002] Remote Direct Memory Access (RDMA) is a commonly used network communication technology in data centers. RDMA supports one-sided operations such as read, write, atomic increment, and atomic exchange. This process bypasses the operating system of the remote host, thus saving a large amount of CPU resources of the remote host. At the same time, RDMA also reduces network communication latency, improves message transmission rate and bandwidth. Taking ConnectX-6 as an example, its minimum latency is 600ns, it can process 200M messages per second, and the maximum bandwidth can be 200Gbps.
[0003] The Memory disaggregation architecture has gradually been widely used in the industrial and academic fields due to its good memory resource utilization and scalability. The computing unit and memory unit of a traditional server are separated to form a computing pool and a storage pool respectively. The nodes in the computing pool have a large number of processor cores and a small amount of memory (used as a cache); the nodes in the storage pool have a large-capacity DRAM or NVM storage space, but do not have (or only have limited) computing power. The nodes in the computing pool mainly rely on various RDMA one-sided operations to actively access the memory data located in the storage pool, while the nodes in the storage pool hardly need to actively initiate any operations (except for establishing connections).
[0004] To implement RDMA one-sided operations, a connection-oriented reliable service (RC) communication mode needs to be used when establishing an RDMA connection. In the RC communication mode, one local queue pair (QP) can and can only establish a connection with one remote QP to access the memory of the corresponding remote node. Therefore, the RC communication mode is more likely to have scalability problems than the connectionless unreliable service (UD). Previous studies have shown that if the number of QPs is too large, there will be jitter in the RNIC on-board cache during the sending of RDMA requests, thus reducing the message processing throughput.
[0005] Since the nodes in the computing pool have a large number of processor cores, memory disaggregation applications can enable a large number of threads. These threads may all use RDMA one-sided operations. FaRM and LITE proposed a QP sharing method, that is, for a certain remote node, k threads fixedly share one QP. When the process has N threads and R remote nodes, the process needs to create RN / k QPs. Although this optimization reduces the number of QPs required by the process, it is at the cost of increasing the synchronization cost between threads. Therefore, when the number of threads increases, this method still cannot achieve good performance.
[0006] Memory decoupled applications such as FORD and Sherman avoid cross-thread sharing of QPs by allocating one QP for each remote node per thread in their open-source implementations. However, even with this strategy, the message throughput rate still decreases as the number of threads or coroutines increases, thus failing to meet the scalability requirements of applications. Through experiments, we found that this phenomenon is not simply due to an excessive number of QPs, but rather the combined result of the following two factors:
[0007] 1. Different QPs may share the same Doorbell Register, so there is still synchronization overhead between threads;
[0008] 2. If the size of the Work Requests (WRs) located in all Send Queues (SQ) exceeds the on-board cache, it will cause cache thrash, thereby affecting the message processing throughput rate.
[0009] In summary, the prior art fails to well solve the problem of the decreasing message throughput rate when the number of threads or coroutines increases. Distributed applications based on RDMA need to be optimized from two aspects: RDMA resource allocation and active work request rate limiting, to avoid the decrease in message processing rate when the number of threads or coroutines increases, so as to ensure that applications mainly relying on unilateral RDMA operations, such as memory decoupled architectures, can achieve better scalability without modifying the application logic. Summary of the Invention
[0010] In view of the above problems, the present invention provides an adaptive optimization method and system for improving the scalability of RDMA unilateral operations.
[0011] An adaptive optimization method for improving the scalability of RDMA unilateral operations, the method comprising:
[0012] During the establishment of a connection between a local node and a remote node, allocate RDMA resources in a thread-aware manner;
[0013] Set credit points at fixed intervals for determining whether to submit RDMA unilateral work requests;
[0014] Whenever a user submits a set of RDMA unilateral work requests, if the credit points of the current thread are greater than the number of the work requests, submit the work requests; otherwise, postpone the submission of the work requests until the credit points of the current thread are greater than the number of the work requests before submitting.
[0015] Further, the allocation of RDMA resources includes initializing credit points, allocating QP objects, and allocating queue CQ objects.
[0016] Furthermore, during the establishment of the connection between the local node and the remote node, RDMA resources are allocated in a thread-aware manner, which specifically includes:
[0017] The local node creates a completion queue CQ object;
[0018] The local node allocates a doorbell register that can be shared by multiple QP objects for each thread, and assigns the credit value Ct with an initial value of CI;
[0019] The local node creates a QP object for each thread respectively, and all QP objects of the same thread are bound to the doorbell register and the queue CQ object allocated by this thread;
[0020] The remote node creates the same number of QP objects, and these QP objects can be bound to any doorbell register and CQ object;
[0021] The local node and the remote node exchange necessary information and set the RC communication mode;
[0022] If there are still remote nodes not connected by the current thread, then select another remote node and repeat the above operations until all remote nodes are connected.
[0023] Furthermore, whenever the user submits a set of RDMA one-sided work requests, if the credit points of the current thread are greater than the number of the work requests, then submit the work requests; otherwise, postpone the submission of the work requests until the credit points of the current thread are greater than the number of the work requests before submitting, which specifically includes:
[0024] Check whether the credit points Ct of the current thread are greater than the number N of this set of work requests; if not, suspend the execution of the current coroutine until the credit points Ct change;
[0025] If it meets the condition, allocate credit points, that is, set the credit points Ct of the current thread to Ct - N;
[0026] Use the QP corresponding to the target remote node within the current thread to send a request to the RDMA device;
[0027] When this set of work requests has been processed, recycle the credit points, that is, set the credit points Ct of the current thread to Ct + N.
[0028] Furthermore, the credit points used to judge whether to submit RDMA one-sided work requests are set at regular intervals, which specifically includes:
[0029] Each fixed time period includes two stages: a normal execution mode and a sampling test mode;
[0030] The sampling test mode uses a microbenchmarking method to determine the maximum credit points used in the same-cycle normal execution mode.
[0031] Furthermore, each fixed time period includes two stages: the normal execution mode and the sampling test mode, specifically including:
[0032] Under the normal execution mode, the maximum send queue depth remains unchanged; after a certain time, it switches to the sampling test mode;
[0033] Under the sampling test mode, each element Di in the set of the maximum send queue depths needs to be tested separately to obtain an index f(Dk) reflecting its message processing rate.
[0034] Furthermore, the obtaining of the index f(Dk) reflecting the message processing rate specifically includes:
[0035] Select the maximum send queue depth Dk to be tested;
[0036] Obtain the current maximum send queue depth D, and adjust the current credit points of each thread according to the formula Ct = Ct + (Dk - D); at the same time, modify the current maximum send queue depth D to Dk;
[0037] Start measuring the number of RDMA messages processed by each thread, and stop measuring after a fixed time to obtain the total number of RDMA messages f(Dk) processed by all threads per unit time;
[0038] If there are still new maximum send queue depths available for testing, re-select the maximum send queue depth Dk to be tested and repeat the above steps; otherwise, select the Di that maximizes f(Di), and obtain the current maximum send queue depth D; adjust the current credit points of each thread according to the formula Ct = Ct + (Dk - D), and modify the current maximum send queue depth D to Dk;
[0039] After execution, switch to the normal execution mode.
[0040] Furthermore, the credit points for judging whether an RDMA unilateral work request is submitted are set at fixed intervals, which is completed by a separate thread.
[0041] An adaptive optimization system for improving the scalability of RDMA unilateral operations includes an initialization module, a credit point dynamic adjustment module, and a credit point audit module;
[0042] The initialization module is used to initialize the credit points and allocate RDMA resources in a thread-aware manner;
[0043] The credit point dynamic adjustment module is used to set the credit points for judging whether an RDMA unilateral work request is submitted at fixed intervals;
[0044] A credit point auditing module, which is used to, whenever a user submits a set of RDMA unilateral work requests, if the credit points of the current thread are greater than the number of the work requests, allow the work requests and allocate credit points; otherwise, postpone the submission of the work requests until the credit points of the current thread are greater than the number of the work requests before submitting; and recycle the credit points after the work requests are executed.
[0045] Furthermore, the allocation of RDMA resources includes QP objects and queue CQ objects.
[0046] Compared with the adaptive optimization method for the scalability of conventional RDMA unilateral operations, the technical solution proposed by the present invention has at least the following advantages:
[0047] 1. The adaptive optimization method for the scalability of RDMA unilateral operations proposed by the present invention ensures that there is no inter-thread lock synchronization in the data path for sending RDMA requests. On the one hand, by allocating one QP per thread, it is ensured that different threads do not need to use QP-level spin locks when sending RDMA requests to the same target node; on the other hand, since the doorbell register is allocated by thread and different threads' QPs use different doorbell registers, the implicit inter-thread synchronization overhead caused by the sharing of the doorbell register is also avoided. In addition, compared with allocating one device context per thread (such as Sherman), the working principle of the present invention is the same, but it reduces the number of RDMA resource allocations on the local node and shortens the connection establishment time.
[0048] 2. The adaptive optimization method for the scalability of RDMA unilateral operations proposed by the present invention restricts the send queue depth through the credit point mechanism, thereby preventing cache thrashing. Experiments have proved that as the number of threads or coroutines increases, the maximum credit points gradually decrease, while the RDMA message throughput rate remains rising or basically unchanged. This reflects that the mechanism of the present invention can effectively prevent the negative impact of cache thrashing on performance and scalability, and at the same time can maximize the message processing performance by improving the utilization rate of the network card link.
[0049] Other features and advantages of the present invention will be described in the subsequent specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures pointed out in the specification and the drawings. Description of the Drawings
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 This is the main flowchart of the adaptive optimization method for the scalability of RDMA unilateral operations proposed by the present invention.
[0052] Figure 2 This is the network topology diagram of the adaptive optimization system according to the embodiment of the present invention.
[0053] Figure 3 This is the storage layout of the doorbell register and the mapping relationship between the QP and the doorbell register in the embodiment of the present invention.
[0054] Figure 4 This is the allocation status of the RDMA object at the end of step S1 in the embodiment of the present invention.
[0055] Figure 5 This is the curve graph of the throughput rate of the 8B random read operation varying with the number of threads under 16 coroutines in the embodiment of the present invention.
[0056] Figure 6 This is the curve graph of the throughput rate of the 8B random read operation varying with the number of coroutines under 96 threads in the embodiment of the present invention.
[0057] Figure 7 This is the curve graph of the throughput rate of the read operation of the distributed hash table based on the decoupled memory architecture varying with the number of threads (fixed 16 coroutines) in the embodiment of the present invention. Detailed implementation manners
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] In the field of remote direct memory access, the prior art fails to well solve the problem that the message throughput rate decreases when the number of threads or coroutines increases.
[0060] Therefore, the present invention proposes an adaptive optimization method and system for improving the scalability problem of RDMA unilateral operations, including an adaptive optimization method for improving the scalability problem of RDMA unilateral operations and an adaptive optimization system for improving the scalability problem of RDMA unilateral operations.
[0061] The object of the present invention is to improve the scalability of distributed applications (such as memory decoupling architecture applications) based on RDMA one-sided operations and increase the utilization rate of network resources. More specifically, the invention avoids the decrease in the throughput rate of RDMA message processing when the number of node threads or coroutines increases.
[0062] In a first aspect, as Figure 1 shown, the present invention provides an adaptive optimization method for improving the scalability problem of RDMA one-sided operations, and the method includes:
[0063] Initializing credit points and allocating RDMA resources in a thread-aware manner during the establishment of a connection between the local node and the remote node;
[0064] Setting credit points at fixed intervals for judging whether an RDMA one-sided work request is submitted;
[0065] Whenever a user submits a set of RDMA one-sided work requests, if the credit points of the current thread are greater than the number of the work requests, then submit the work requests; otherwise, postpone the submission of the work requests until the credit points of the current thread are greater than the number of the work requests before submitting.
[0066] In specific implementation, based on credit point review, whenever a user submits a set of RDMA one-sided work requests, if the credit points of the current thread are greater than the number of the work requests, then allow the work requests and allocate credit points; otherwise, postpone the submission of the work requests until the credit points of the current thread are greater than the number of the work requests before submitting; recycle the credit points after the work requests are executed.
[0067] In this embodiment, the allocation of RDMA resources includes QP objects and queue CQ objects.
[0068] In this embodiment, during the establishment of a connection between the local node and the remote node, the allocation of RDMA resources in a thread-aware manner specifically includes:
[0069] The local node creates a completion queue CQ object;
[0070] The local node allocates a doorbell register shared by multiple QP objects for each thread, and assigns the credit point Credit value Ct to the initial value CI;
[0071] The local node creates a QP object for each thread respectively, and binds all QP objects of the same thread to the doorbell register and the queue CQ object allocated by this thread;
[0072] The remote node creates the same number of QP objects, and the QP objects can be bound to any doorbell register and CQ object;
[0073] The local node and the remote node exchange necessary information to set the RC communication mode;
[0074] If there are still remote nodes not connected by the current thread, reselect a remote node and repeat the above operations until all remote nodes are connected.
[0075] In specific implementation, the local node creates a QP object for each thread, and all QP objects of the same thread are bound to the doorbell register and the queue CQ object allocated to this thread; the remote node creates the same number of QP objects, and these QP objects can be bound to any doorbell register and CQ object; that is, the local node and the remote node use different allocation strategies.
[0076] In this embodiment, whenever the user submits a set of RDMA one-sided work requests, if the credit points of the current thread are greater than the number of the work requests, the work requests are allowed and credit points are allocated; otherwise, the submission of the work requests is postponed until the credit points of the current thread are greater than the number of the work requests before submission; when the work requests are completed, the credit points are recovered, which specifically includes:
[0077] Check whether the credit points Ct of the current thread are greater than the number N of this set of work requests (equivalent to the depth of the send queue being less than or equal to the threshold); if not, suspend the execution of the current coroutine until the credit points Ct change;
[0078] If it meets the condition, allocate credit points, that is, set the credit points Ct of the current thread to Ct - N;
[0079] Use the QP corresponding to the target remote node in the current thread to send a request to the RDMA device;
[0080] When this set of work requests has been processed, recover the credit points, that is, set the credit points Ct of the current thread to Ct + N.
[0081] In specific implementation, the formula for allocating credit points is Ct = Ct - N. In fact, the length of the send queue in the QP also increases by N; the formula for releasing credit points is Ct = Ct + N. In fact, since N work requests have been completed, the length of the send queue in the QP also decreases by N accordingly. Assume the initial value of Ct is CI. When Ct is equal to 0, the length of the send queue in the QP is CI. Therefore, it can be obtained that whether the credit points Ct are greater than the number N of this set of work requests is equivalent to the depth of the send queue being less than or equal to the threshold. Set the credit points for judging whether to submit the RDMA one-sided work requests at fixed intervals, that is, adjust the threshold.
[0082] In this embodiment, setting the credit points for judging whether to submit the RDMA one-sided work requests at fixed intervals specifically includes:
[0083] Each fixed time period includes two phases: a regular execution mode and a sampling test mode;
[0084] The sampling test mode uses a microbenchmarking method to determine the maximum credit points used in the regular execution mode of the same period.
[0085] In this embodiment, each fixed time period includes two phases: a regular execution mode and a sampling test mode, specifically including:
[0086] Under the regular execution mode, the maximum send queue depth remains unchanged; after a certain time, it switches to the sampling test mode;
[0087] Under the sampling test mode, it is necessary to test each element Di in the maximum send queue depth set respectively to obtain an index f(Dk) reflecting its message processing rate.
[0088] In this embodiment, obtaining the index f(Dk) reflecting the message processing rate specifically includes:
[0089] Select the maximum send queue depth Dk to be tested;
[0090] Obtain the current maximum send queue depth D, and adjust the current credit points of each thread according to the formula Ct = Ct + (Dk - D); at the same time, modify the current maximum send queue depth D to Dk;
[0091] Start measuring the number of RDMA messages processed by each thread, and stop measuring after a fixed time to obtain the total number f(Dk) of RDMA messages processed by all threads per unit time;
[0092] If there are still new maximum send queue depths available for testing, re - select the maximum send queue depth Dk to be tested and repeat the above steps; otherwise, select the Di that maximizes f(Di), and obtain the current maximum send queue depth D; adjust the current credit points of each thread according to the formula Ct = Ct + (Dk - D), and modify the current maximum send queue depth D to Dk;
[0093] After execution, switch to the regular execution mode.
[0094] In this embodiment, the credit points for judging whether an RDMA unilateral work request is submitted are set at fixed intervals and are completed by a separate thread.
[0095] In a second aspect, the present invention provides an adaptive optimization system for improving the scalability problem of RDMA unilateral operations, including an initialization module, a credit point dynamic adjustment module, and a credit point audit module;
[0096] An initialization module, configured to initialize credit points and allocate RDMA resources in a thread-aware manner;
[0097] A credit point dynamic adjustment module, configured to set credit points for determining whether an RDMA unilateral work request is submitted at fixed intervals;
[0098] A credit point audit module, configured to, whenever a user submits a set of RDMA unilateral work requests, if the credit points of the current thread are greater than the number of the work requests, allow the work requests and allocate credit points; otherwise, postpone the submission of the work requests until the credit points of the current thread are greater than the number of the work requests and then submit; and recycle the credit points after the work requests are executed.
[0099] In this embodiment, the allocation of RDMA resources includes QP objects and queue CQ objects.
[0100] Specifically, in implementation, the adaptive optimization system for improving the scalability of RDMA unilateral operations and the adaptive optimization method for improving the scalability of RDMA unilateral operations of the present invention have corresponding implementation processes, which will not be elaborated herein.
[0101] To enable those skilled in the art to better understand the present invention, the principle of the present invention is described below with reference to the accompanying drawings:
[0102] The main technical idea of the present invention is as follows: During thread initialization, allocate RDMA resources in a thread-aware manner to ensure that there is no thread-interlock synchronization in the data path for sending RDMA requests; limit the depth of the send queue through a credit point mechanism to prevent cache thrashing; adjust the maximum send queue depth in an adaptive manner without the need for prior training and / or identifying load characteristics. Compared with the existing system, the adaptive optimization method disclosed in the present invention can avoid performance degradation when the number of threads / coroutines increases, and can also maximize the message processing performance by improving the utilization rate of the network card link.
[0103] The present invention proposes an adaptive optimization method and system for the scalability of RDMA unilateral operations. The method is divided into three parts: (1) Thread-aware allocation of RDMA resources; (2) Send queue depth control based on a credit point mechanism; and (3) Periodic update of the maximum credit points.
[0104] Figure 2 This is the network topology diagram of the embodiment of the present invention. The computing pool includes a node CN1, and a memory decoupled application program is executed on CN1 with t threads. The storage pool includes two nodes MN1 and MN2.
[0105] The following details the design of the three parts of the present invention with reference to the embodiments.
[0106] S1. Thread-Aware Allocation of RDMA Resources
[0107] As mentioned in the background section, in order to reduce the synchronization overhead between threads, the application located on CN1 needs to allocate t QP pairs to MN1 and MN2 respectively, and they are in one-to-one correspondence with the threads of the application. This avoids the synchronization overhead on its spin lock caused by two threads using the same QP, which is beneficial to improving the request processing performance. However, when the number of threads exceeds a certain threshold, the request processing throughput still decreases.
[0108] The reason for this phenomenon is the many-to-one mapping between QP and the doorbell register. Figure 3 Shows the main causes of this phenomenon. In the default configuration, a device context (ibv_context) contains 16 doorbell registers, among which 4 doorbell registers are exclusively used by 1 QP respectively, and the remaining 12 doorbell registers can be shared by multiple QPs. When creating a QP, the doorbell register to be used needs to be determined according to the fixed round-robin rule, and the mapping relationship cannot be changed during the lifetime of the QP. Since the allocation of the doorbell register has nothing to do with the thread corresponding to the QP, even if different QPs are used between threads, synchronization overhead may still occur.
[0109] Therefore, the present invention adopts a thread-aware RDMA resource allocation method, aiming to avoid the synchronization overhead between threads and share resources as much as possible to avoid resource waste. Specifically, during the initialization of a certain thread T, the following steps need to be executed:
[0110] S1.1. Allocate a doorbell register that can be shared by multiple QP objects for thread T, create a completion queue (CQ) object, and assign the credit value Ct as the initial value CI;
[0111] S1.2. Select a remote node and create a QP object, which is bound to the doorbell register and CQ object obtained in step S1.1;
[0112] In implementation, the doorbell register is implicitly allocated during the creation of the QP object, and the mapping relationship cannot be intervened; therefore, several QP objects can be created first and placed in different free queues according to the doorbell register address. Each free queue corresponds to a thread, and thread T needs to obtain the QP object from the corresponding free queue (other threads cannot use this queue). In an actual system, it may also be necessary to use the switch option of the driver to expand the number of available doorbell registers in the current device context.
[0113] In addition, all QP objects of a thread should share a CQ object. Therefore, by polling only this CQ object, it can be determined whether there is a completed work request for this thread. This not only saves RDMA resources but also reduces the polling overhead. In the implementation, since the CQ object (as a parameter) needs to be determined before the QP object is created, before the QP object is created, it is necessary to determine which idle queue the QP should be in, then determine the thread corresponding to this idle queue, and the pointer to the CQ object to be used.
[0114] S1.3. Request the remote node described in step S1.2 to create a QP object, which can use any doorbell register and CQ object;
[0115] Since in the RC communication mode, a QP object must establish a one-to-one connection with another QP object at the remote end. Assume that CN1 includes 2t QPs, then MN1 and MN2 each include t QPs (even if the number of threads in MN1 and MN2 is not t). In the decoupled memory model, neither MN1 nor MN2 actively issues RDMA requests. Therefore, MN1 and MN2 can use any doorbell register and CQ object, which has no obvious impact on network performance.
[0116] S1.4. The local node and the remote node described in step S1.2 exchange necessary information and set the RC communication mode;
[0117] In the embodiment, t QPs in CN1 (1 per thread) exchange QP numbers with t QPs in MN1 respectively, and complete the setting of the RC communication mode according to the RDMA standard process; similarly, the remaining t QPs in CN1 exchange QP numbers with t QPs in MN2 respectively, and also set the RC communication mode.
[0118] S1.5. If there are still remote nodes that the current thread has not connected to, jump to step S1.2; otherwise, end.
[0119] After the above steps, the thread-aware allocation of RDMA resources is completed, ensuring that the data path for sending RDMA requests will not have inter-thread synchronization. Figure 4 The allocation status of RDMA objects at the end of step S1 in the embodiment is given.
[0120] S2. Transmission queue depth control based on the credit point mechanism
[0121] In addition to avoiding the synchronization overhead between threads, it is also necessary to prevent the on-board cache jitter of the RNIC to avoid the decrease in message processing throughput. Previous work believed that the increase in the number of QPs was the cause of this phenomenon. However, through experiments, it was found that this phenomenon would only be triggered when the send queue depth SQ exceeded a certain threshold. Among them, SQ = ∑SQi, and SQi represents the send queue depth of QPi. The determination and update of this threshold are described in step S3, and this step presents the send queue depth control algorithm. When the user initiates an RDMA unilateral operation (such as RDMA READ), the following operations are performed:
[0122] S2. Whenever the user submits a set of RDMA unilateral work requests (using the ibv_post_send API), if the current send queue depth exceeds the threshold, the submission of this operation is postponed until the send queue depth is less than the threshold before submitting. The present invention adopts a method based on credit points. Specifically, the following steps are performed:
[0123] S2.1. Check whether the credit point Ct of the current thread is greater than the number N of this set of work requests (equivalent to the send queue depth being less than or equal to the threshold). If not, jump to step S2.2; otherwise, jump to step S2.3.
[0124] Note: In step S1.1, each thread sets a credit point variable Ct, and its initial value is CI. Ct represents the maximum number of work requests (WR) that this thread allows to enter the send queue (SQ) in the current state. All QPs of the same thread are jointly restricted by Ct, but the send work request operations of other threads will not cause the credit point of this thread T to decrease. To send a work request, the credit point of this thread T needs to be deducted by 1.
[0125] S2.2. Pause the execution of the current coroutine until the credit point Ct changes, and then jump to step S2.1.
[0126] Note: If the credit point Ct of the current thread is less than 0, the execution of the current coroutine needs to be paused. At this time, the coroutine scheduler should schedule to other coroutines to avoid program suspension. Since our reference implementation needs to wait for the success of the RDMA polling operation to determine that the work request is completed, it is also necessary to actively initiate the RDMA polling operation during the suspension of the current coroutine and restore the credit point Ct of the current thread according to the execution result.
[0127] S2.3. Set the credit point Ct of the current thread to Ct - N. This step indicates that the send queue depth has increased, so the process of other threads submitting RDMA unilateral work requests can be delayed.
[0128] S2.4. Use the QP corresponding to the target remote node within the current thread to send a request to the RDMA device. In this step, according to the allocation scheme described in S1, there will be no lock competition among the QPs of different threads due to the sharing of locks or doorbell registers, so good scalability performance can be achieved.
[0129] S2.5. When this group of work requests has been processed (i.e., the corresponding RDMA completion record has been obtained), set the credit point Ct of the current thread to Ct + N. This step indicates that the send queue depth has decreased, so other threads can retry.
[0130] Note: The program needs to actively initiate an RDMA polling operation (e.g., using the ibv_poll_cq API) to determine that the RDMA request has been processed. When a new completion message is obtained in the CQ, it means that the number of work requests in the send queue has decreased accordingly, so the credit point of this thread T can be incremented by 1. For performance considerations, multiple CQ objects can be obtained in one polling operation, so the increment of the credit point may also be greater than 1.
[0131] S3. Periodic update of the maximum credit point
[0132] As mentioned above, only when the send queue depth SQ exceeds a certain threshold will the message processing throughput decrease. This threshold is related to multiple factors, including the number of enabled threads, the number of remote nodes (i.e., the total number of QPs), the load read-write ratio, etc. Therefore, it is difficult to determine the maximum credit point of each thread using an empirical formula. After completing step S2, the application adjusts the maximum send queue depth at regular intervals and dynamically adjusts the credit point accordingly. This operation can also be completed by a separate thread. Each cycle includes two phases: a regular execution mode and a sampling test mode. The sampling test mode uses a microbenchmarking method to determine the maximum credit point used in the regular execution mode of the same cycle. The application is not aware of these two phases and executes the same logic. The specific steps are as follows:
[0133] S3.1. In the regular execution mode, the maximum send queue depth remains unchanged. After a period of time (longer than the sampling phase), switch to the sampling test mode (step S3.2).
[0134] S3.2. In the sampling test mode, it is necessary to test each element Di in the set of the maximum send queue depths respectively to obtain an index f(Dk) reflecting its message processing rate. For the short-time test phase, the embodiments sequentially use the following maximum credit point values: 4, 6, 8, 10, 12, 14, 16, 256. Specifically, execute the following steps:
[0135] S3.2.1. Select the maximum transmission queue depth Dk to be tested: Assume that the maximum credit value used in the previous cycle execution phase is 12, and 4 is used as the new maximum credit value in this cycle first.
[0136] S3.2.2. Obtain the current maximum transmission queue depth D, and adjust the current credit of each thread according to the formula Ct = Ct + (Dk - D). At the same time, modify the current maximum transmission queue depth D to Dk. According to the formula, the credit value Ct of each thread in the embodiment is Ct + (4 - 12) = Ct - 8, that is, it is adjusted down by 8 units.
[0137] S3.2.3. Start measuring the number of RDMA messages processed by each thread, and stop the measurement after a fixed time to obtain the total number of RDMA messages f(Dk) processed by all threads per unit time. In the embodiment, this time can be 5 ms, because too short a time will lead to too much measurement randomness, and too long a time will increase the time ratio of the test phase and have an adverse impact on performance.
[0138] S3.2.4. If there is still a new maximum transmission queue depth available for testing, jump to step S3.2.1. For example, in the embodiment, 6 is also required to be used as the new maximum credit value. According to the formula, the credit value Ct of each thread is Ct + (8 - 4) = Ct + 4, and measure the number of RDMA messages processed in a short period of time,... and so on until all maximum credit values are tested.
[0139] Finally, when all maximum transmission queue depths have been tested, we obtain a mapping table of the maximum credit value and the message processing rate. At this time, it is necessary to select Di that maximizes f(Di), and obtain the current maximum transmission queue depth D. Adjust the current credit of each thread according to the formula Ct = Ct + (Dk - D), and modify the current maximum transmission queue depth D to Dk. Switch to the normal execution mode (step S3.1).
[0140] To verify the actual use effect of the method proposed by the present invention, several servers with a total of 96 cores (hyper-threading enabled) of two processors were used as the experimental environment. Among them, 1 is a computing node and 2 are memory nodes. All servers are installed with ConnectX-6 network cards and are interconnected through an InfiniBand / RDMA network.
[0141] First, we measured the throughput rate of 8B random read operations with the change of the number of threads under 16 coroutines respectively. The test results are as Figure 5As shown. When the thread-aware QP allocation mechanism is turned off, even though each thread has its own private QP, the performance cannot be improved when the number of threads exceeds 16 because of the contention of the doorbell register. By enabling the thread-aware allocation of RDMA resources, especially ensuring that different threads' QPs use different doorbell registers, the inter-thread lock synchronization caused by sending work requests can be effectively reduced, thus significantly improving scalability. When the number of threads is 96, the message throughput rate that can be achieved with the optimization enabled is 108.414 MOP / s, basically reaching full load; it is up to 8.6 times higher than that with the optimization disabled.
[0142] On the other hand, we fixed the number of threads at 96 and adjusted the number of coroutines, and measured the change in the throughput rate of 8B random read operations with the number of threads. The test results are shown in Figure 6 . By restricting the send queue depth through the credit point mechanism, the cache jitter phenomenon can be avoided. As the number of threads increases, the maximum credit point gradually decreases, while the RDMA message throughput rate remains rising or basically unchanged. The throughput rate of this operation is up to 1.85 times higher than that without the optimization enabled.
[0143] For other RDMA one-sided operations, their performance characteristics are basically the same. The method proposed by the present invention shows good scalability in the micro-benchmark experiment and can adapt to different numbers of threads and coroutines.
[0144] To evaluate whether this method can improve its scalability in actual applications, we also measured the change in the throughput rate of the pure read operation of the distributed hash table based on the decoupled memory architecture with the number of threads. The test results are shown in Figure 7 . It can be seen from the test results that the method proposed by the present invention is beneficial to improving the performance of applications with an actual decoupled memory architecture, and both optimizations contribute to the performance improvement. After the optimization is enabled, the throughput rate of this operation is up to 3 times higher than that without the optimization enabled at most.
[0145] Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive optimization method for improving the scalability of RDMA unilateral operations, characterized in that, the method includes: During the establishment of a connection between the local node and the remote node, allocate RDMA resources in a thread-aware manner; Set a credit point at fixed intervals for determining whether an RDMA unilateral work request is submitted; Whenever a user submits a set of RDMA unilateral work requests, if the credit point of the current thread is greater than the number of the work requests, submit the work requests; otherwise, postpone the submission of the work requests until the credit point of the current thread is greater than the number of the work requests before submitting; During the establishment of the connection between the local node and the remote node, allocate RDMA resources in a thread-aware manner, specifically including: The local node creates a completion queue CQ object; The local node allocates a doorbell register shared by multiple QP objects for each thread, and assigns the credit point Credit value Ct with the initial value CI; The local node creates a QP object for each thread respectively, and binds all QP objects of the same thread to the doorbell register and the queue CQ object allocated to this thread; The remote node creates the same number of QP objects, and the QP object can be bound to any doorbell register and CQ object; The local node and the remote node exchange QP object numbers and set the RC communication mode; If there are still remote nodes not connected by the current thread, reselect a remote node and repeat the above operations until all remote nodes are connected.
2. The adaptive optimization method for improving the scalability of RDMA unilateral operations according to claim 1, characterized in that, the allocation of RDMA resources includes initializing credit points, allocating QP objects, and allocating queue CQ objects.
3. The adaptive optimization method for improving the scalability of RDMA unilateral operations according to claim 1, characterized in that, whenever a user submits a set of RDMA unilateral work requests, if the credit point of the current thread is greater than the number of the work requests, submit the work requests; otherwise, postpone the submission of the work requests until the credit point of the current thread is greater than the number of the work requests before submitting, specifically including: Check whether the credit point Ct of the current thread is greater than the number N of this set of work requests; if not, suspend the execution of the current coroutine until the credit point Ct changes; If it meets the condition, allocate credit points, that is, set the credit point Ct of the current thread = Ct - N; Use the QP corresponding to the target remote node in the current thread to send a request to the RDMA device; When this set of work requests has been processed, recycle the credit points, that is, set the credit point Ct of the current thread = Ct + N.
4. The adaptive optimization method for improving the scalability of RDMA unilateral operations according to claim 1, characterized in that, the setting of a credit point at fixed intervals for determining whether an RDMA unilateral work request is submitted, specifically includes: Each fixed time period includes two stages: a regular execution mode and a sampling test mode; Among them, the maximum sending queue depth remains unchanged in the normal execution mode; it switches to the sampling test mode after a certain period of time; In the sampling test mode, it is necessary to test each element Di in the set of the maximum sending queue depth respectively, and obtain the total number f(Dk) of RDMA messages processed by all threads per unit time. f(Dk) reflects its message processing rate, where Dk is the maximum sending queue depth to be tested.
5. The adaptive optimization method for improving the scalability problem of RDMA one-sided operations according to claim 4, characterized in that, the total number f(Dk) of RDMA messages processed by all threads per unit time specifically includes: select the maximum sending queue depth Dk to be tested; obtain the current maximum sending queue depth D, and adjust the current credit points of each thread according to the formula Ct = Ct + (Dk - D); at the same time, modify the current maximum sending queue depth D to Dk; start measuring the number of RDMA messages processed by each thread, and stop measuring after a fixed time to obtain the total number f(Dk) of RDMA messages processed by all threads per unit time; if there is still a new maximum sending queue depth available for testing, re-select the maximum sending queue depth Dk to be tested and recalculate f(Dk); otherwise, select Di that maximizes f(Di), and obtain the current maximum sending queue depth D; adjust the current credit points of each thread according to the formula Ct = Ct + (Dk - D), and modify the current maximum sending queue depth D to Dk; switch to the normal execution mode after completion.
6. The adaptive optimization method for improving the scalability problem of RDMA one-sided operations according to claim 1, characterized in that, the credit points for judging whether an RDMA one-sided work request is submitted are set at fixed intervals, which is completed by a separate thread.
7. An adaptive optimization system for improving the scalability problem of RDMA one-sided operations, characterized in that, it includes an initialization module, a credit point dynamic adjustment module, and a credit point review module; The initialization module is used to allocate RDMA resources in a thread-aware manner during the establishment of a connection between the local node and the remote node; The credit point dynamic adjustment module is used to set the credit points for judging whether an RDMA one-sided work request is submitted at fixed intervals; The credit point review module is used to, whenever the user submits a set of RDMA one-sided work requests, if the credit points of the current thread are greater than the number of the work requests, allow the work requests and allocate credit points; otherwise, postpone the submission of the work requests until the credit points of the current thread are greater than the number of the work requests before submitting; the credit points are recycled after the work requests are executed; During the establishment of the connection between the local node and the remote node, allocating RDMA resources in a thread-aware manner specifically includes: the local node creates a completion queue CQ object; the local node allocates a doorbell register that can be shared by multiple QP objects for each thread, and assigns the credit point Credit value Ct as the initial value CI; The local node creates a QP object for each thread, and all QP objects of the same thread are bound to the doorbell register and the queue CQ object allocated to this thread; The remote node creates the same number of QP objects, and these QP objects can be bound to any doorbell register and CQ object; The local node and the remote node exchange QP object numbers and set the RC communication mode; If there are still remote nodes not connected by the current thread, then re-select a remote node and repeat the above operations until all remote nodes are connected.
8. The adaptive optimization system for improving the scalability problem of RDMA unilateral operations according to claim 7, wherein, The allocation of RDMA resources includes initializing credit points, allocating QP objects, and allocating queue CQ objects.
Citation Information
Patent Citations
Credit based flow control scheme over virtual interface architecture for system area networks
US6347337B1