Optimization method and device for RDMA comparison and exchange operation, electronic equipment and storage medium

Through real-time credit value management and target retry operations strategies, the comparison and exchange operation conflict caused by high concurrent operations in the RDMA network is solved, which improves the performance and scalability of application operations and reduces resource waste.

CN120104532APending Publication Date: 2025-06-06TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411259288.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In RDMA networks, high concurrent operations and uneven data distribution lead to frequent conflicts in RDMA comparison and exchange operations, resulting in waste of resources and affecting the throughput, latency and scalability of application operations.

Method used

Through real-time credit value management, determine the target retry operation, and perform retry operation when the real-time credit value meets the preset threshold to reduce the probability of conflict. The real-time credit value is used to indicate the number of running entities that can be run concurrently. When the credit value does not meet the threshold, the running entity will be adjusted to the pause state.

Benefits of technology

It effectively reduces the conflict frequency of RDMA comparison and exchange operations, reduces unsuccessful retry operations, improves the performance and scalability of application operations, and reduces the waste of network resources and processor resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104532A_ABST
    Figure CN120104532A_ABST
Patent Text Reader

Abstract

The invention relates to an optimization method and device for RDMA comparison and exchange operation, electronic equipment and a storage medium, and the method comprises the steps: determining a target retry operation and a real-time credit value corresponding to a target operation entity which executes the target retry operation, the target retry operation comprises at least one comparison and exchange operation, and the real-time credit value is corresponding to a target operation entity which executes the target retry operation; the real-time credit value is used for indicating the number of application operations executable by the target running entity; under the condition that the real-time credit value does not meet the preset threshold value, the target operation entity is adjusted to be in a pause state; and under the condition that the real-time credit value meets a preset threshold value, executing target retry operation, and determining a response result corresponding to each comparison and exchange operation in the target retry operation. According to the embodiment of the invention, under the condition that the number of concurrent operations is large, the running state of the running entity can be controlled by using the real-time credit value, the probability of conflict of comparison and exchange operations is reduced, the occurrence frequency of unsuccessful retry operations is reduced, and thus the waste of network resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of system software technology, and in particular to an optimization method, device, electronic device, and storage medium for RDMA compare and exchange operations. Background Art

[0002] Remote Direct Memory Access (RDMA) technology is a commonly used network communication technology that supports unilateral operations including memory read (write), write (read), compare and swap (CAS) and fetch and add (FAA) operations. When the number of concurrent operations is large or the data distribution is highly uneven, the application operation for a single data may need to issue more RDMA requests to complete, resulting in the waste of input / output per second (IOPS) resources or bandwidth resources, thereby affecting the throughput, latency and scalability of application operations. Summary of the invention

[0003] In view of this, the present disclosure proposes a technical solution for an optimization method, device, electronic device, and storage medium for RDMA compare and swap operations.

[0004] According to one aspect of the present disclosure, there is provided an optimization method for RDMA compare-and-exchange operations, comprising: determining a target retry operation, and a real-time credit value corresponding to a target running entity that executes the target retry operation, wherein the target retry operation includes at least one compare-and-exchange operation, and the real-time credit value is used to indicate the number of running entities that can currently run concurrently; if the real-time credit value does not meet a preset threshold, adjusting the target running entity to a paused state; if the real-time credit value meets the preset threshold, executing the target retry operation, and determining a response result corresponding to each compare-and-exchange operation in the target retry operation.

[0005] In a possible implementation, the determining the target retry operation includes: after completing the execution process of the previous retry operation, determining the target running entity and the target retry operation among a plurality of running entities in a paused state.

[0006] In a possible implementation, determining the real-time credit value corresponding to the target running entity executing the target retry operation includes: when the target retry operation is the first application operation executed, determining the real-time credit value according to a preset credit value.

[0007] In a possible implementation, when the real-time credit value satisfies the preset threshold, the target retry operation is executed, and a response result corresponding to each comparison and exchange operation in the target retry operation is determined, including: before executing the target retry operation, performing a self-decrement operation on the real-time credit value to determine a first updated credit value; for any comparison and exchange operation in the target retry operation, determining an operation request corresponding to the comparison and exchange operation; and sending the operation request to the target hardware corresponding to the comparison and exchange operation to determine a response result corresponding to the comparison and exchange operation.

[0008] In one possible implementation, the response result corresponding to any comparison and exchange operation is used to indicate whether the comparison and exchange operation is successfully executed; the method also includes: when the response result corresponding to any comparison and exchange operation in the target retry operation is unsuccessful, determining the waiting delay corresponding to the target retry operation according to the exponential backoff algorithm and the preset backoff period; according to the waiting delay, adjusting the target running entity to a paused state.

[0009] In a possible implementation, the method further includes: determining a retry rate based on a response result corresponding to each comparison and exchange operation within a preset time range; adjusting the preset credit value and the preset backoff period based on the retry rate, and determining the adjusted preset credit value and the adjusted preset backoff period.

[0010] In a possible implementation, the method further includes: after the target retry operation is completed, performing a self-increment operation on the first updated credit value to determine a second updated credit value.

[0011] According to another aspect of the present disclosure, there is provided an optimization device for RDMA compare-and-exchange operations, characterized in that it includes: a retry operation determination module, used to determine a target retry operation, and a real-time credit value corresponding to a target running entity that executes the target retry operation, wherein the target retry operation includes at least one compare-and-exchange operation, and the real-time credit value is used to indicate the number of running entities that can currently run concurrently; a first execution module, used to adjust the target running entity to a paused state if the real-time credit value does not meet a preset threshold; and a second execution module, used to execute the target retry operation if the real-time credit value meets the preset threshold, and determine a response result corresponding to each compare-and-exchange operation in the target retry operation.

[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0013] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.

[0014] The disclosed embodiments can determine a target retry operation and a real-time credit value corresponding to a target running entity that executes the target retry operation, wherein the target retry operation includes at least one compare-and-swap operation, and the real-time credit value is used to indicate the number of application operations that can be executed by the target running entity; when the real-time credit value does not meet a preset threshold, the target running entity can be adjusted to a paused state; when the real-time credit value meets the preset threshold, the target retry operation can be performed, and a response result corresponding to each compare-and-swap operation in the target retry operation can be determined, thereby realizing operation current limiting of the running entity using the real-time credit value, reducing the probability of conflict in RDMA compare-and-swap operations when the number of concurrent operations is large, thereby reducing the frequency of unsuccessful retry operations, improving the performance and scalability of application operations, and reducing the waste of network resources and processor resources.

[0015] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.

[0017] Figure 1 A schematic diagram showing a network topology structure according to an embodiment of the present disclosure;

[0018] Figure 2 A flowchart of an optimization method for RDMA compare-and-swap operations according to an embodiment of the present disclosure is shown;

[0019] Figure 3 A schematic diagram showing a code snippet corresponding to a retry operation according to the prior art;

[0020] Figure 4 A schematic diagram showing a code snippet corresponding to a retry operation according to an embodiment of the present disclosure;

[0021] Figure 5 A schematic diagram showing a method of concurrently running multiple applications according to the prior art;

[0022] Figure 6 A schematic diagram showing how the throughput of a split memory system varies with the number of concurrently running threads according to an embodiment of the present disclosure;

[0023] Figure 7 A schematic diagram showing a comparison of the number of single operation retries according to an embodiment of the present disclosure;

[0024] Figure 8 A block diagram of an optimization device for RDMA compare-and-swap operations according to an embodiment of the present disclosure is shown;

[0025] Fig. 9 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0026] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0027] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0028] The term "and / or" herein is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set consisting of A, B, and C.

[0029] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.

[0030] Remote Direct Memory Access (RDMA) technology is a commonly used network communication technology. Compared with TCP / IP protocol stack technology, it has the advantages of low latency and high bandwidth. For example, the RDMA network card Mellanox ConnectX-6 can provide 200Gbps bandwidth and less than 600ns latency. RDMA supports both bilateral and unilateral communication modes. Among them, the unilateral communication mode can include memory read (write), write (read), compare and swap (compare_and_swap, CAS) and fetch_and_add, FAA) operations, etc. These operations can directly operate the remote memory of the target object (including hardware devices or server nodes, etc.) instead of passing through the CPU corresponding to the target object (including hardware devices or server nodes, etc.).

[0031] Disaggregated Memory (DM) technology can analyze memory resources and computing resources in physical structure, and enable all computing nodes in the cluster to share and access memory resources. A common way to implement disaggregated memory in the prior art is to connect computing nodes and memory nodes in the cluster using RDMA networks and other methods, so that computing nodes can access memory resources based on RDMA unilateral communication, and replace the read and write operations used to access local memory resources with the unilateral operations included in the RDMA unilateral communication method.

[0032] Figure 1 A schematic diagram of a network topology structure according to an embodiment of the present disclosure is shown. Figure 1 As shown, the DM system includes a computing pool and a storage pool. The computing pool includes a computing node 1, and a memory decoupling application can run on the computing node 1, including at least one thread; the storage pool includes two memory nodes, namely, memory node 1 and memory node 2. The computing node 1 can access the memory node 1 and the memory node 2 respectively through a wireless broadband network (InfiniBand, IB) and / or a remote direct memory access network.

[0033] With the popularization of multi-core processors, the scalability of application operations in DM technology (for example, inserting and searching operations of data structures, etc.) is becoming more and more important. For various networks such as RDMA, the number of read and write operations that can be performed per second (Input / Output Per Second, IOPS) is a limited resource, and the available amount of IOPS resources will limit the scalability of application operations in DM technology. For example, for an RDMA network card, when the transmission data is less than or equal to 64B, the corresponding number of messages processed per second is limited by the maximum IOPS corresponding to the RDMA network card (about 110MOP / s); when the transmission data is greater than or equal to 1024B, the corresponding number of messages processed per second is limited by the network bandwidth corresponding to the RDMA network card (about 13GB / s). It can be seen that the number of RDMA reads and writes and the amount of read and write data generated by different application operations will cause different degrees of pressure on the limited network resources of the RDMA network. Therefore, when there is a waste of IOPS resources or network bandwidth resources, it may have a negative impact on the throughput, operation delay and scalability of application operations.

[0034] Specifically, since the maximum IOPS of the RDMA network is fixed, when the number of concurrent operations is large or the data distribution is highly uneven, the application operation for a single data may need to issue more RDMA requests to complete, resulting in a decrease in the throughput of the application operation.

[0035] On the other hand, in the prior art, the software program corresponding to the split memory system can ensure the consistency of data transmission through synchronization mechanisms such as lock mechanisms and / or lock-free mechanisms. However, for common skewed distributed loads (i.e., most memory resource access requests are concentrated on a few objects), when the number of threads executing concurrent application operations increases, the performance of update operations of lock-free data structures will decrease.

[0036] In addition, in order to avoid wasting the CPU working cycle while waiting for the RDMA response message (i.e., RDMAACK message), in the prior art, the separate memory system usually uses the coroutine technology to allow each thread to perform multiple tasks alternately to improve the CPU utilization. Generally, after the coroutine technology is enabled, the throughput of the application operation will be improved compared to when the coroutine technology is not enabled, but for skewed distributed loads, after the coroutine technology is enabled, the performance of the application operation may decrease.

[0037] In view of this, the present disclosure provides an optimization method for RDMA compare and swap operations, which can use real-time credit values ​​to implement coroutine flow control, reduce the frequency of unsuccessful retry operations when the number of concurrent operations is large, and reduce the probability of conflicts in RDMA compare and swap operations, thereby improving the performance and scalability of application operations and reducing the waste of network resources and processor resources such as IOPS. The optimization method for RDMA compare and swap operations provided by the present disclosure is described in detail below.

[0038] Figure 2 A flowchart of an optimization method for RDMA compare and exchange operations according to an embodiment of the present disclosure is shown. The optimization method for RDMA compare and exchange operations can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The optimization method for RDMA compare and exchange operations can be implemented by a processor calling computer-readable instructions stored in a memory. Alternatively, the optimization method for RDMA compare and exchange operations can be executed by a server. As Figure 2 As shown, the optimization method for RDMA compare-and-swap operations includes:

[0039] In step S21, a target retry operation and a real-time credit value corresponding to a target running entity that performs the target retry operation are determined, wherein the target retry operation includes at least one compare-and-swap operation, and the real-time credit value is used to indicate the number of running entities that can currently run concurrently.

[0040] The target retry operation here may represent an application operation that needs to be retried at least once when there is an operation conflict in the RDMA compare and swap operation. The specific form of the conflict in the RDMA compare and swap operation may refer to the implementation in the related art, for example, it may include that multiple different computing nodes in the RDMA network attempt to operate on the same address in the memory node, multiple different threads on the same computing node attempt to operate on the same address in the memory node, and different coroutines on the same thread attempt to operate on the same address in the memory node, etc. The present disclosure does not make specific limitations on this.

[0041] The specific method of determining the target retry operation can be flexibly set according to actual usage requirements, and this disclosure does not make specific limitations on this.

[0042] In one example, the code snippet corresponding to the target retry operation can be represented as a code snippet that includes retry logic, an entry and at least one exit. Therefore, the code corresponding to the application operation can be automatically identified based on the compiler, and the target retry operation can be determined based on the retry logic in the code.

[0043] Figure 3 FIG. 1 is a schematic diagram showing a code snippet corresponding to a retry operation according to the prior art. Figure 3 As shown, the code snippet can represent the update operation of an application in a split memory system including retry logic. Among them, the 5th line of the code snippet represents the operation of writing the key-value pair data to a new remote memory address and returning the remote memory address; the 6th line represents the computing node constructing new slot data (Slot), which can be a 64-bit unsigned integer, including the hash value of the key and the remote memory address returned by the 5th line operation; the 7th and 8th lines represent reading the existing slot data; the 10th line represents modifying the existing slot data to the new slot data constructed by the 6th line operation through a regular comparison and exchange operation.

[0044] In the case where the compare-and-exchange operation represented by line 10 is successful, the read, write and compare-and-exchange operations included in the code snippet only need to be executed once; in the case where the compare-and-exchange operation represented by line 10 fails, it is necessary to retry starting from line 5, that is, the read, write and compare-and-exchange operations included in the code snippet need to be retried. Therefore, when the number of retries of the code snippet is too many, more processor resources and network resources may be wasted.

[0045] In one example, the code corresponding to the application operation may be annotated in advance based on a preset annotation function, and the code snippet corresponding to the target retry operation may be determined according to the annotation function, thereby determining the target retry operation.

[0046] Figure 4 A schematic diagram showing a code snippet corresponding to a retry operation according to an embodiment of the present disclosure is shown. Figure 4 As shown, this code snippet can represent the following Figure 3 The code snippet shown is annotated with the code snippet shown in the figure. Figure 3 The code snippet shown includes a newly added BackoffGuard class and replaces the conventional compare-and-swap operation with a marked function cas_sync_backoff. By using the newly added BackoffGuard class and the marked function cas_sync_backoff, when the application operation corresponding to the code snippet needs to be retried, it can be directly determined as the target retry operation.

[0047] The specific form of the target running entity that executes the target retry operation can be flexibly set according to actual usage requirements. For example, it can include any computing node in the separate memory system, any thread on the computing node, any coroutine in the thread, etc. The present disclosure does not make specific limitations on this.

[0048] Figure 5 FIG. 2 is a schematic diagram showing a method of concurrently running multiple applications according to the prior art. Figure 5 As shown in (a) in FIG. 1 , thread 1 includes three coroutines, namely coroutine 1, coroutine 2 and coroutine 3; coroutine 1 is used to sequentially execute application operation 1 and application operation 4, coroutine 2 is used to sequentially execute application operation 2 and application operation 5, and coroutine 3 is used to sequentially execute application operation 3 and application operation 6. Figure 5 As shown in (b) in the figure, during the execution of application operation 1 by coroutine 1, coroutine 1 may release (yield) the control of thread 1 due to reasons such as waiting for the RDMAACK message, resulting in coroutine scheduling and switching to the process of coroutine 2 executing application operation 2, so that thread 1 can alternately execute application operation 1 and application operation 2. However, this will cause the execution process of application operation 1 and application operation 2 to overlap, resulting in an operation conflict of the compare and exchange operation.

[0049] For example, when both application operation 1 and application operation 2 need to update the key value k, and the target values ​​are different, coroutine 1 and coroutine 2 will respectively obtain the content of the target data structure corresponding to the key value k through RDMA read operations, and read the initial value of the key value k. Then, coroutine 1 and coroutine 2 will respectively issue RDMA compare-and-exchange requests to modify the target data structure. Since the compare-and-exchange request issued by coroutine 1 has the same target address, the same initial value, and a different target value as the compare-and-exchange request issued by coroutine 2. Therefore, at least one compare-and-exchange operation will fail, causing the corresponding application operation to be retried. Figure 5 As shown in (c), in order to avoid the above-mentioned conflict, in the prior art, the commonly used coroutine scheduling method will control coroutine 1 to release the control of thread 1 after application operation 1 is completed, so that coroutine 2 can execute application operation 2. In this case, although thread 1 includes multiple coroutines, thread 1 can only execute each application operation in sequence, resulting in that the CPU time consumed by each coroutine waiting for the RDMAACK message cannot be utilized, thereby causing a waste of processor resources and network.

[0050] In view of this, the embodiment of the present disclosure indicates the number of currently concurrently running running entities through the real-time credit value corresponding to the target running entity, so as to control the running state of the target running entity, thereby reducing the probability of overlapping when different running entities execute their corresponding application operations, resulting in application operation conflicts. Among them, the value of the real-time credit value depends on the number of running entities that are already in the running state, and the present disclosure does not make any specific restrictions.

[0051] In step S22, when the real-time credit value does not meet the preset threshold, the target operation entity is adjusted to a suspended state.

[0052] The specific value of the preset threshold value can be flexibly set according to actual usage requirements. For example, the preset threshold value can be set to 0, etc., and the present disclosure does not make specific limitations on this.

[0053] When the real-time credit value does not meet the preset threshold, it means that the number of running entities that can currently run concurrently has reached the preset upper limit. In this case, if the number of running entities running concurrently is increased, the probability of operation conflicts may increase, thereby increasing the number of retry operations and the number of retries, resulting in a waste of network resources and processor resources. Therefore, when the real-time credit value does not meet the preset threshold, the target running entity can be adjusted to a paused state until the retry operation being executed is completed and the target running entity is awakened. The preset upper limit here can be flexibly set according to actual usage requirements, depending on the separate memory system and the actual performance of the RDMA network, and the present disclosure does not make specific limitations on this.

[0054] In step S23, when the real-time credit value meets the preset threshold, the target retry operation is performed, and a response result corresponding to each comparison and exchange operation in the target retry operation is determined.

[0055] When the real-time credit value meets the preset threshold, it means that the number of running entities that can currently run concurrently has not reached the preset upper limit. In this case, the probability of operation conflicts is low, and the target running entity can directly execute the target retry operation and determine the response result corresponding to each comparison and exchange operation in the target retry operation. Among them, the specific form and content of the response result corresponding to any comparison and exchange can be flexibly set according to actual usage requirements, depending on the actual situation of the comparison and exchange operation. For example, the corresponding response result of any comparison and exchange operation may include whether the comparison and exchange operation is successful, the memory address and target value corresponding to the comparison and exchange operation, etc. The present disclosure does not make specific limitations on this.

[0056] The disclosed embodiments can determine a target retry operation and a real-time credit value corresponding to a target running entity that executes the target retry operation, wherein the target retry operation includes at least one compare-and-swap operation, and the real-time credit value is used to indicate the number of application operations that can be executed by the target running entity; when the real-time credit value does not meet a preset threshold, the target running entity can be adjusted to a paused state; when the real-time credit value meets the preset threshold, the target retry operation can be performed, and a response result corresponding to each compare-and-swap operation in the target retry operation can be determined, thereby realizing operation current limiting of the running entity using the real-time credit value, reducing the probability of conflict in RDMA compare-and-swap operations when the number of concurrent operations is large, thereby reducing the frequency of unsuccessful retry operations, improving the performance and scalability of application operations, and reducing the waste of network resources and processor resources.

[0057] In a possible implementation, determining a target retry operation includes: after completing the execution process of a previous retry operation, determining a target running entity and a target retry operation among a plurality of running entities in a paused state.

[0058] When the execution process of the last retry operation is completed, it means that at least one of the concurrently running entities has ended the running state, that is, the number of running entities that can currently run concurrently has increased by at least one. At this time, the target running entity can be determined from multiple running entities in a paused state, the application operation corresponding to the target running entity can be determined as the target retry operation, and the target running entity can be adjusted to the running state. Among them, the specific method of determining the target running entity among multiple running entities in a paused state can be flexibly set according to actual usage requirements, and the present disclosure does not make specific limitations on this.

[0059] In one example, a corresponding priority may be set in advance for each running entity, and the running entity with the highest priority among a plurality of running entities in a paused state may be determined as the target running entity.

[0060] In one example, a corresponding number may be set in advance for each running entity, and a queue data structure (wait_list) may be constructed. According to a first-in-first-out principle, a target running entity may be determined from the running entities in a paused state in the queue data structure.

[0061] In a possible implementation, determining a real-time credit value corresponding to a target running entity that executes a target retry operation includes: determining a real-time credit value according to a preset credit value when the target retry operation is the first application operation executed.

[0062] The preset credit value here may represent a preset upper limit corresponding to the number of running entities that can run concurrently, and its specific value may be flexibly set according to actual usage requirements, and the present disclosure does not make any specific limitation on this.

[0063] When the target retry operation is the first executed application operation, it means that no running entity is in the running state. Therefore, the real-time credit value can be set equal to the preset credit value to indicate the number of running entities that can currently run concurrently.

[0064] In a possible implementation, when the real-time credit value meets a preset threshold, a target retry operation is performed, and a response result corresponding to each comparison and exchange operation in the target retry operation is determined, including: before executing the target retry operation, performing a self-decrement operation on the real-time credit value to determine a first updated credit value; for any comparison and exchange operation in the target retry operation, determining an operation request corresponding to the comparison and exchange operation; sending the operation request to the target hardware corresponding to the comparison and exchange operation, and determining a response result corresponding to the comparison and exchange operation.

[0065] Before executing the target retry operation, the real-time credit value can be decremented to determine the first updated credit value to indicate that the target running entity is about to be adjusted from the paused state to the running state, the number of running entities that can run concurrently is reduced, and the target running entity is allowed to execute the target retry operation.

[0066] During the process of executing the target retry operation, for any comparison and exchange operation in the target retry operation, the operation request corresponding to the comparison and exchange operation can be determined, and the operation request can be sent to the target object corresponding to the comparison and exchange operation to obtain the response result corresponding to the comparison and exchange operation. Among them, the specific form and content of the operation request can be flexibly set according to actual usage requirements, for example, it can include the remote memory address corresponding to the target object, data read and write operations and target values, etc., and the present disclosure does not make specific limitations on this. The specific form and content of the response result can be flexibly set according to actual usage requirements, and the present disclosure does not make specific limitations on this; the specific form of the target object corresponding to any comparison and exchange operation can be flexibly set according to actual usage requirements, for example, it can include memory nodes corresponding to a separate memory system, and / or hardware devices, etc., and the present disclosure does not make specific limitations on this.

[0067] In one possible implementation, the response result corresponding to any comparison and exchange operation is used to indicate whether the comparison and exchange operation is successfully executed; the method also includes: when the response result corresponding to any comparison and exchange operation in the target retry operation is unsuccessful, determining the waiting delay corresponding to the target retry operation according to the exponential backoff algorithm and the preset backoff period; according to the waiting delay, adjusting the target running entity to a paused state.

[0068] The response result corresponding to any compare-and-swap operation in the target retry operation can be used to indicate whether the compare-and-swap operation is successfully executed, thereby determining whether the target retry operation needs to be retried again.

[0069] Specifically, when all response results corresponding to the comparison and exchange operations in the target retry operation are successfully executed, there is no need to retry the target retry operation again.

[0070] When the response result corresponding to any comparison-and-exchange operation in the target retry operation is an unsuccessful execution, the target retry operation needs to be retried again; however, since the response result of the comparison-and-exchange operation is an unsuccessful execution, it can be determined that at least one comparison-and-exchange operation is being executed at this time, and there is an operation conflict with the comparison-and-exchange operation. If the target retry operation is retried again immediately, it may still fail due to the operation conflict with the comparison-and-exchange operation, resulting in the target retry operation needing to be retried multiple times, increasing the number of retries, and causing a waste of network resources and processor resources.

[0071] Therefore, when the response result corresponding to any comparison and exchange operation in the target retry operation is an unsuccessful execution, the waiting delay corresponding to the target retry operation can be determined according to the exponential backoff algorithm and the preset backoff period, and during the waiting delay, the target running entity is controlled to remain in a paused state, so as to reduce the frequency of operation failures in the target retry operation, reduce the number of retries, and improve the scalability of the target retry operation. Moreover, by implementing the scheduling of the running entity in this way, there is no need to use a special scalable spin lock, and it can be applied to lock-free data update operations. During the waiting delay, other running entities can be adjusted to the running state to perform other application operations.

[0072] The specific value of the preset backoff period can be flexibly set according to actual usage requirements, depending on the performance of the separated memory system, the RDMA network, and the upper limit of the number of concurrently running entities, etc., and this disclosure does not make specific restrictions on this. The specific form of the exponential backoff algorithm can refer to the implementation methods in the relevant technology, and this disclosure does not make specific restrictions on this; preferably, a truncated exponential backoff algorithm can be used.

[0073] The specific method of determining the waiting delay corresponding to the target retry operation according to the exponential backoff algorithm and the preset backoff period can be flexibly set according to actual usage requirements, and the present disclosure does not make specific limitations on this.

[0074] In one example, based on the truncated exponential backoff algorithm, the waiting delay corresponding to the target retry operation after any retry failure can be determined according to the preset backoff period and the number of retries of the target retry operation. Specifically, the waiting delay can be expressed as formula (1):

[0075] t=min{to×2,tmax}+Rand(to)(1)

[0076] Where, t represents the waiting delay; t 0 Indicates the preset waiting delay when the target retry operation is retried for the first time. Its specific value can be flexibly set according to actual usage requirements. For example, you can set t 0 = 4096 CPU clock cycles, which is approximately equal to the time required for one RDMA read / write operation, and the present disclosure does not make any specific limitation on this; i represents the number of target retry operations; t max Indicates the preset backoff period; Rand(t 0 ) represents 0 to t 0 A random value between .

[0077] Through the above process, the operation flow limiting and scheduling of the running entity can be further realized based on the exponential backoff strategy and the preset backoff period, which can reduce the frequency of unsuccessful retry operations when the number of concurrently running application operations is high, reduce the number of retries, and improve the performance and scalability of application operations.

[0078] Figure 6 A schematic diagram showing how the throughput of a separate memory system according to an embodiment of the present disclosure varies with the number of concurrently running threads is shown. As shown in FIG. (6), the curve marked with a triangle represents how the throughput of the separate memory system varies with the number of concurrently running threads after the optimization method provided by the present disclosure is used; the curve marked with a diamond represents how the throughput of the separate memory system without optimization varies with the number of concurrently running threads.

[0079] like Figure 6 As shown in (a), when the separated memory system includes only a single computing node and its corresponding workload is a write-intensive application operation, after using the optimization method provided by the present disclosure, when the number of concurrently running threads is large, the throughput of the separated memory system will not drop significantly; while the throughput of the unoptimized separated memory system drops significantly to a lower value after the number of concurrently running threads is greater than about 10.

[0080] like Figure 6 As shown in (b), when the separated memory system includes only a single computing node and its corresponding workload is a read-intensive application operation, after using the optimization method provided by the present disclosure, when the number of concurrently running threads is large, the throughput of the separated memory system will not drop significantly; while the throughput of the unoptimized separated memory system drops significantly to a lower value after the number of concurrently running threads is greater than about 10.

[0081] like Figure 6 As shown in (c), when the separated memory system includes only a single computing node and its corresponding workload is a read-only application operation, after using the optimization method provided by the present invention, the throughput of the separated memory system will increase with the increase in the number of concurrently running threads, reach a stable value when the number of threads is about 70, and maintain a high value; while the throughput of the unoptimized separated memory system will fluctuate and decrease after the number of concurrently running threads is greater than about 20.

[0082] like Figure 6 As shown in (d), when the separated memory system only includes multiple computing nodes and its corresponding workload is a write-intensive application operation, after using the optimization method provided by the present invention, when the number of concurrently running threads is large, the throughput of the separated memory system can be maintained at a high value; while the throughput of the separated memory system that has not been optimized is lower.

[0083] like Figure 6 As shown in (e), when the separated memory system includes only a single computing node and its corresponding workload is a write-intensive application operation, after utilizing the optimization method provided by the present invention, when the number of concurrently running threads is large, the throughput of the separated memory system can be maintained at a high value; while the throughput of the separated memory system that has not been optimized is lower.

[0084] like Figure 6 As shown in (f) in the figure, when the separated memory system includes only a single computing node and its corresponding workload is a write-intensive application operation, after using the optimization method provided by the present disclosure, the throughput of the separated memory system will increase with the increase of the number of concurrently running threads, reach a stable value when the number of threads is about 500, and maintain a high value; while the throughput of the unoptimized separated memory system, although it increases with the increase of the number of concurrently running threads, is always lower than the throughput of the separated memory system after using the optimization method provided by the present disclosure.

[0085] Figure 7 FIG. 2 is a schematic diagram showing a comparison of the number of single operation retries according to an embodiment of the present disclosure. Figure 7As shown, the horizontal axis is the number of retries of a single application operation, and the vertical axis is the percentage of any application operation with a certain number of retries to all application operations. The yellow bar graph represents an unoptimized split memory system; the blue bar graph represents a split memory system optimized using the optimization method disclosed in the present invention.

[0086] like Figure 7 As shown, the average number of application operations retried by the separated memory system optimized by the optimization method disclosed in the present invention is about 1.1 times, and about 93.3% of application operations do not need multiple retries. In contrast, the average number of application operations in the unoptimized separated memory system is about 11.5 times, and there are many application operations that need to be retried more than 9 times.

[0087] In a possible implementation, the method further includes: determining a retry rate based on a response result corresponding to each comparison and exchange operation within a preset time range; adjusting a preset credit value and a preset backoff period based on the retry rate, and determining an adjusted preset credit value and an adjusted preset backoff period.

[0088] The specific value of the preset backoff period will affect the performance and scalability of the application operation. Specifically, if the value of the preset backoff period is small, the target retry operation may be frequently retried, increasing the number of retries and the probability of operation conflicts during each retry; if the value of the preset backoff period is large, the tail delay of the target retry operation will be increased, reducing the throughput of the running entity; therefore, the preset backoff period needs to be set reasonably.

[0089] The value of the preset backoff period will be affected by the upper limit of the number of concurrently running entities, that is, the preset credit value. Therefore, in order to improve the performance of application operations, the preset credit value and the preset backoff period can be dynamically adjusted. Specifically, the response results corresponding to each comparison and exchange operation within the preset time range can be counted to determine the retry rate, and the preset credit value and the preset backoff period can be adjusted according to the retry rate to determine the adjusted preset credit value and the adjusted preset backoff period. Among them, the specific value of the preset time range can be flexibly set according to actual usage requirements, and the present disclosure does not make specific limitations on this.

[0090] The specific method of adjusting the preset credit value and the preset backoff period according to the retry rate can be flexibly set according to actual usage requirements, and the present disclosure does not make specific limitations on this.

[0091] In a possible implementation, when the retry rate is greater than the first threshold, it indicates that the comparison and exchange operation in the application operation / retry operation has a large number of operation conflicts, and it is necessary to reduce the upper limit of the number of concurrently running entities, that is, reduce the preset credit value, and determine the preset credit value after reduction. The specific value of the first threshold can be flexibly set according to actual usage requirements. For example, the first threshold can be set to 0.5, and this disclosure does not specifically limit this.

[0092] In a possible implementation, when the retry rate is less than the second threshold, it means that the number of operation conflicts in the comparison and exchange operation in the application operation / retry operation is small, and the upper limit of the number of concurrently running running entities can be increased, that is, the preset credit value is increased, and the increased preset credit value is determined. The specific value of the second threshold can be flexibly set according to actual usage requirements. For example, the second threshold can be set to 0.1, and this disclosure does not specifically limit this.

[0093] In one possible implementation, when the retry rate is greater than the first threshold, it indicates that a large number of operation conflicts occur in the comparison and exchange operations in the application operation / retry operation, and it is necessary to extend the waiting delay of the running entity, that is, increase the preset backoff period, and determine the increased preset backoff period.

[0094] In one possible implementation, when the retry rate is less than the second threshold, it indicates that the comparison and exchange operations in the application operation / retry operation have fewer operation conflicts, and the waiting delay of the running entity can be reduced, that is, the preset credit value is reduced, and the reduced preset credit value is determined.

[0095] Through the above process, the real-time credit value and the preset backoff period can be adjusted in real time and dynamically according to the retry rate, and can adapt to different workloads without human intervention. It can reduce the impact of unreasonable parameter settings on application operating performance, and can enable the separated memory system to maintain high efficiency and stability under different operating conditions.

[0096] In a possible implementation, the method further includes: after the target retry operation is completed, performing a self-increment operation on the first updated credit value to determine a second updated credit value.

[0097] When the execution process of the target retry operation is completed, it indicates that the target running entity is about to end the running state, that is, the number of running entities that can currently run concurrently increases by one. Therefore, the first updated credit value can be incremented to determine the second updated credit value.

[0098] Further, after the target retry operation is executed, a new target running entity may be determined from a plurality of running entities in a paused state, and the second updated credit value may be determined as the real-time credit value corresponding to the new target running entity.

[0099] In an embodiment of the present disclosure, a target retry operation and a real-time credit value corresponding to a target running entity that executes the target retry operation can be determined, wherein the target retry operation includes at least one compare-and-exchange operation, and the real-time credit value is used to indicate the number of application operations that can be executed by the target running entity; when the real-time credit value does not meet a preset threshold, the target running entity can be adjusted to a paused state; when the real-time credit value meets the preset threshold, the target retry operation can be performed, and a response result corresponding to each compare-and-exchange operation in the target retry operation can be determined, thereby realizing operation current limiting and scheduling of the running entity using the real-time credit value and the exponential backoff strategy, reducing the probability of conflict in RDMA compare-and-exchange operations when the number of concurrent operations is large, thereby reducing the frequency of unsuccessful retry operations, improving the performance and scalability of application operations, and reducing the waste of network resources and processor resources; and, the embodiment of the present disclosure does not require the use of a special extensible spin lock, and therefore can be applied to scenarios of lock-free data update operations. Furthermore, parameters such as real-time credit value and preset backoff period can be adjusted in real time and dynamically according to the retry rate, and can adapt to different workloads without human intervention. This can reduce the impact of unreasonable parameter settings on application operating performance, and enable the separated memory system to maintain high efficiency and stability under different operating conditions.

[0100] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not repeat them. It can be understood by those skilled in the art that in the above-mentioned method of the specific implementation method, the specific execution order of each step should be determined according to its function and possible internal logic.

[0101] In addition, the present disclosure also provides an optimization device, an electronic device, and a computer-readable storage medium for RDMA compare-and-swap operations, all of which can be used to implement any optimization method for RDMA compare-and-swap operations provided by the present disclosure. The corresponding technical solutions and descriptions are referred to in the corresponding records of the method part and will not be repeated here.

[0102] Figure 8 FIG. 2 is a block diagram of an optimization device for RDMA compare and swap operations according to an embodiment of the present disclosure. Figure 8 As shown, the device 800 includes:

[0103] A retry operation determination module 801 is used to determine a target retry operation and a real-time credit value corresponding to a target running entity that performs the target retry operation, wherein the target retry operation includes at least one compare-and-swap operation, and the real-time credit value is used to indicate the number of running entities that can currently run concurrently;

[0104] The first execution module 802 is used to adjust the target operation entity to a suspended state when the real-time credit value does not meet a preset threshold;

[0105] The second execution module 803 is used to execute the target retry operation when the real-time credit value meets the preset threshold, and determine the response result corresponding to each comparison and exchange operation in the target retry operation.

[0106] In a possible implementation, the retry operation determination module 801 is specifically used to: after completing the execution process of the last retry operation, determine a target running entity and a target retry operation among a plurality of running entities in a paused state.

[0107] In a possible implementation, the retry operation determination module 801 is specifically configured to: determine the real-time credit value according to a preset credit value when the target retry operation is the first application operation executed.

[0108] In one possible implementation, the second execution module 803 is specifically used to: before executing the target retry operation, perform a self-decrement operation on the real-time credit value to determine a first updated credit value; for any comparison and exchange operation in the target retry operation, determine an operation request corresponding to the comparison and exchange operation; send the operation request to the target object corresponding to the comparison and exchange operation, and determine a response result corresponding to the comparison and exchange operation.

[0109] In one possible implementation, the response result corresponding to any comparison and exchange operation is used to indicate whether the comparison and exchange operation is successfully executed; the second execution module 803 is also used to: when the response result corresponding to any comparison and exchange operation in the target retry operation is an unsuccessful execution, determine the waiting delay corresponding to the target retry operation according to the exponential backoff algorithm and the preset backoff period; according to the waiting delay, adjust the target running entity to a paused state.

[0110] In one possible implementation, the second execution module 803 is also used to: determine the retry rate based on the response results corresponding to each comparison and exchange operation within a preset time range; adjust the preset credit value and the preset backoff period based on the retry rate, and determine the adjusted preset credit value and the adjusted preset backoff period.

[0111] In a possible implementation, the second execution module 803 is further configured to: after the target retry operation is completed, perform a self-increment operation on the first updated credit value to determine the second updated credit value.

[0112] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation thereof can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0113] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0114] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0115] Fig. 9 1900 is a block diagram of an electronic device according to an embodiment of the present disclosure. For example, the apparatus 1900 may be provided as a server or a terminal device. Fig. 9 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions, such as an application, that can be executed by the processing component 1922. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0116] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2000. TM , MacOS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.

[0117] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the device 1900 to perform the above method.

[0118] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0119] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0120] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0121] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0122] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0123] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0124] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0125] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.

[0126] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. An optimization method for RDMA compare-and-swap operations, characterized in that: include: Determine a target retry operation and a real-time credit value corresponding to a target running entity that performs the target retry operation, wherein the target retry operation includes at least one compare-and-swap operation, and the real-time credit value is used to indicate the number of running entities that can currently run concurrently; When the real-time credit value does not meet a preset threshold, adjusting the target operation entity to a suspended state; In a case where the real-time credit value meets the preset threshold, the target retry operation is performed, and a response result corresponding to each comparison and exchange operation in the target retry operation is determined.

2. The method according to claim 1, characterized in that The determining target retry operation comprises: After completing the execution process of the last retry operation, the target running entity and the target retry operation are determined among a plurality of running entities in a paused state.

3. The method according to claim 1 or 2, characterized in that: The determining a real-time credit value corresponding to a target running entity that executes the target retry operation includes: In a case where the target retry operation is the first application operation executed, the real-time credit value is determined according to a preset credit value.

4. The method according to claim 3, characterized in that When the real-time credit value meets the preset threshold, executing the target retry operation and determining a response result corresponding to each comparison and exchange operation in the target retry operation includes: Before executing the target retry operation, performing a self-decrement operation on the real-time credit value to determine a first updated credit value; For any compare-and-swap operation in the target retry operation, determining an operation request corresponding to the compare-and-swap operation; The operation request is sent to the target hardware corresponding to the comparison and exchange operation, and a response result corresponding to the comparison and exchange operation is determined.

5. The method according to claim 4, characterized in that The response result corresponding to any compare-and-swap operation is used to indicate whether the compare-and-swap operation is executed successfully; The method further comprises: When a response result corresponding to any comparison and exchange operation in the target retry operation is an unsuccessful execution, determining a waiting delay corresponding to the target retry operation according to an exponential backoff algorithm and a preset backoff period; According to the waiting delay, the target running entity is adjusted to a pause state.

6. The method according to claim 5, characterized in that The method further comprises: Determine a retry rate based on the response results corresponding to each comparison and exchange operation within a preset time range; The preset credit value and the preset backoff period are adjusted according to the retry rate to determine the adjusted preset credit value and the adjusted preset backoff period.

7. The method according to claim 4, characterized in that The method further comprises: After the target retry operation is completed, the first updated credit value is incremented to determine a second updated credit value.

8. An optimization device for RDMA compare and swap operations, characterized in that: include: A retry operation determination module, used to determine a target retry operation and a real-time credit value corresponding to a target running entity that performs the target retry operation, wherein the target retry operation includes at least one compare-and-swap operation, and the real-time credit value is used to indicate the number of running entities that can currently run concurrently; A first execution module, configured to adjust the target operation entity to a suspended state when the real-time credit value does not meet a preset threshold; The second execution module is used to execute the target retry operation when the real-time credit value meets the preset threshold, and determine the response result corresponding to each comparison and exchange operation in the target retry operation.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method described in any one of claims 1 to 7 when executing the instructions stored in the memory.

10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.