Remote Direct Memory Operation (RDMO) for Transactional Processing Systems

By combining data processing logic into RDMO and using FPGA to process RDMO, the problem of complex operation delay and efficiency of remote memory is solved, efficient and reliable remote data processing is achieved, and the scalability and performance of the database system are enhanced.

CN112771501BActive Publication Date: 2025-05-13ORACLE INT CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980063000.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-15
Filing Date
2019-08-16
Publication Date
2025-05-13
Estimated Expiration
2039-10-04

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle complex operations of remote memory when providing high availability, especially when a failed component is responsible for performing complex operations, and the problems of delay and efficiency are prominent.

Method used

By combining data processing logic into a logical unit called Remote Direct Memory Operation (RDMO), the logical unit can perform complex operations in a round trip, using the FPGA as an execution candidate to process RDMO.

Benefits of technology

It effectively hides the microsecond-level remote memory delay in the data center, improves the efficiency and reliability of the system, supports more complex remote operations, and enhances the scalability and performance of the database system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112771501B_ABST
    Figure CN112771501B_ABST
Patent Text Reader

Abstract

Techniques for offloading remote direct memory operations (RDMOs) to "execution candidates" are described. An execution candidate can be any hardware capable of performing the operations being offloaded. Thus, an execution candidate can be a network interface controller, a dedicated coprocessor, an FPGA, etc. The execution candidate can be located on a machine remote from the processor from which the operations are being offloaded, or it can be on the same machine as the processor from which the operations are being offloaded. Details are provided for certain specific RDMOs that are particularly useful in online transaction processing (OLTP) and hybrid transaction / analytics (HTAP) workloads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to remote direct memory operations. Background Art

[0002] One way to improve service availability is to design the service so that it can continue to function properly even if one or more of its components fail. For example, U.S. Patent Application No. 15 / 606,322, filed on May 26, 2017, describes a technique for enabling a requesting entity to retrieve data managed by a database server instance from a volatile memory of a remote server machine on which the database server instance is executing, without involving the database server instance in the retrieval operation.

[0003] Because the retrieval does not involve the database server instance, the retrieval operation can succeed even when the database server instance (or the operating system of the host server machine itself) has stopped or become unresponsive. In addition to increased availability, direct retrieval of data is often faster and more efficient than retrieving the same information through conventional interaction with a database server instance.

[0004] In the case of using RDMA technology, in order to retrieve the "target data" specified in the database command from the remote machine without involving the remote database server instance, the requesting entity first uses remote direct memory access (RDMA) to access information about the location where the target data resides in the server machine. Based on such target location information, the requesting entity uses RDMA to retrieve the target data from the host server machine without involving the database server instance. The RDMA read (data retrieval operation) issued by the requesting entity is a unilateral operation that does not require CPU interrupts or OS kernel participation on the host server machine (RDBMS server). For example, such RDMA operations can be handled by hardware on the network interface controller (NIC) of the remote device.

[0005] RDMA technology is well suited for operations that involve only retrieving data from the volatile memory of a crashed server machine. However, it is desirable to provide high availability even when the failed component is responsible for performing operations more complex than simple memory access. To meet this requirement, some systems provide a limited set of "verbs" for performing remote operations such as memory access and atomic operations (test and set, compare-and-swap) via a network interface controller. These operations can be completed as long as the system is powered on and the NIC can access the host memory. However, the types of operations they support are limited.

[0006] More complex operations on data residing in the memory of a remote machine are typically performed by making remote procedure calls (RPCs) to an application running on the remote machine. For example, a database client that needs the average of a set of numbers stored in a database can make an RPC to a remote database server instance that manages the database. In response to the RPC, the remote database server instance reads the set of numbers, calculates the average, and sends the average back to the database client.

[0007] In this example, if the remote database server fails, then the average calculation operation will fail. However, it may be possible to use RDMA to retrieve each number in the set from the remote server. Once the requesting entity has retrieved each number in the set of numbers, the requesting entity can perform the average calculation operation. However, using RDMA to retrieve each number in the set and then performing the calculation locally is much less efficient than having the application on the remote server retrieve the data and perform the average calculation operation.

[0008] While modern networks have impressive bandwidth, round-trip latency in RDMA-capable data centers remains stubbornly above 1 microsecond. As a result, improvements in latency have lagged behind improvements in network bandwidth. The widening gap between improvements in (a) bandwidth and (b) latency has exposed a “network wall” for data-intensive applications. Hardware techniques including simultaneous multithreading, prefetching, speculative loading, transactional memory, and configurable coherence domains are ineffective at hiding the fact that accesses to remote memory have 10 times higher latency than accesses to local memory.

[0009] This network wall poses a design challenge to database systems running transactional workloads. Scaling is challenging because remote manipulation of even the simplest data structures requires multiple round trips across the network. The essence of the problem is that the end-to-end latency for a series of operations on remote memory in either a "lock, write, unlock" or lock-free "fetch-and-add, then write" mode is in the 10us range.

[0010] The approaches described in this section are approaches that could be pursued, but are not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section are prior art merely by virtue of their inclusion in this section. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In the attached picture:

[0012] Figure 1 This is an example of a SlottedAppend RDMO in an OLTP workload that modifies the free pointer and then writes to the slot.

[0013] Figure 2 is an example of a ConditionalGather RDMO that traverses a linked list of versions, performs a visibility check, and returns matching tuples in one operation.

[0014] Figure 3 This is an example of a SignaledRead RDMO that elides locks when reading any lock-based data structure in one request.

[0015] Figure 4 is an example of the ScatterAndAccumulate RDMO that accumulates the elements in a scatter list at a user-defined offset from the baseaddr pointer.

[0016] Figure 5 is an example of a block diagram of a computer system upon which the techniques described herein may be implemented.

[0017] Figure 6 is a block diagram of a system including a machine in which an FPGA serves as an execution candidate for locally and / or remotely initiated RDMOs. DETAILED DESCRIPTION

[0018] In the following description, for the purpose of explanation, many specific details are set forth to provide a thorough understanding of the present invention. However, it will be apparent that the present invention can be practiced without these specific details. In other cases, well-known structures and devices are shown in block diagram form to avoid unnecessarily obscuring the present invention.

[0019] RDMOS

[0020] Database systems can extend this network wall for online transaction processing (OLTP) and hybrid transaction / analytic (HTAP) workloads through novel software / hardware co-design techniques that hide microsecond remote memory latencies in the data center. As described below, the network wall is extended by combining data processing logic into a logical "unit" that can be dispatched to a remote node and executed in a single round trip. This computational unit is referred to in this article as remote direct memory operation or "RDMO".

[0021] Unlike RDMA, which typically involves trivial, simple operations such as retrieving a single value from the remote machine's memory, RDMO can be arbitrarily complex. For example, an RDMO can cause a remote machine to calculate the average of a set of numbers, where those numbers reside in the remote machine's memory. As another example, an RDMO can be a short series of reads, writes, and atomic memory operations that will be transmitted and executed at the remote node without interrupting the remote CPU, similar to an RDMA read or RDMA write operation.

[0022] With RDMO, simple computations can be offloaded to the remote network interface controller (NIC) to improve the efficiency of the system. However, RDMO has many issues with system reliability, especially those related to data corruption. Corruption can be intentional, such as when an RDMA authentication key is sniffed on the line while an RDMA transfer is in progress and reused by a malicious actor. Corruption can also be unintentional, such as when a bit flip occurs in a high-density memory that is directly exposed via RDMA. The non-volatile nature of NVRAM makes it particularly vulnerable to this form of unintentional corruption.

[0023] For mission-critical use cases, efficiency and reliability challenges are barriers to deploying scalable and robust database systems. A reliable and efficient mechanism is provided for transaction processing systems. According to one embodiment, transaction processing is made more efficient and scalable by encoding a series of common operations as RDMOs. Reliability is improved through encrypted RDMA when data is in transit, and reliability is improved through fault-aware RDMA when data is at rest in DRAM and NVRAM.

[0024] Addressing efficiency and reliability issues allows for multi-tenant and cloud deployment of RDMA-capable transaction processing systems. These advances are critical to achieving enterprise-quality transaction processing.

[0025] Same machine RDMOS

[0026] In this document, an entity that is capable of performing a particular operation is referred to as an "execution candidate" for that operation. As described above, RDMOS can be used to offload operations to a remote execution candidate, such as a NIC of a remote host machine. Additionally, RDMOS can be used to offload operations to an execution candidate within the same machine as the process that is offloading the operation.

[0027] For example, a process executing on a main CPU of a computing device can offload operations to a NIC of the same computing device. The local NIC is just one example of the type of execution candidates available on a computing device.

[0028] In the case where the operation (such as RDMO) is a relatively simple operation with high I / O and low computational requirements, the operation can be performed more efficiently by an execution candidate running in an auxiliary processor than by an application running on a local processor. Therefore, it may be desirable to implement the auxiliary processor on a field programmable gate array FPGA configured with logic for performing the relatively simple operation. On the other hand, if the operation (such as RDMO) is a relatively complex operation with low I / O and high computational requirements, it may be more efficient to perform the operation using an application running on a local processor.

[0029] Using FPGA to process locally initiated RDMOS

[0030] Figure 6 6 is a block diagram of a system including a machine 600 that can offload operations to a field programmable gate array (FPGA). The machine 600 includes a processor 620 that executes an operating system 632 and any number of applications, such as application 634. Code for the operating system 632 and applications 634 can be stored on a persistent storage device 636 and loaded into volatile memory 630 as needed. The processor 620 can be one of any number of processors within the machine 600. The processor 620 itself can have many different computing units, shown as cores 622, 624, 622, and 628.

[0031] The machine 600 includes an FPGA 650 with an accelerator unit AU 652. The AU 652 has direct access to data 640 like the processor 620. The processor 620 includes circuitry (shown as uncore 642) that allows the AU to access data 640 in the volatile memory 630 of the machine 600 without involving the cores 622, 624, 622, and 628 of the processor 620.

[0032] The machine 600 includes a network interface controller NIC 602. The NIC 602 may be coupled to the FPGA 650, or may be integrated on the FPGA 650.

[0033] For example, if the operation requested from the machine 600 itself is a relatively simple operation with high I / O and low computational requirements, then the AU 652 may be able to perform the operation more efficiently than the application 634 running on the processor 620. Therefore, it may be desirable to have the AU 652 implemented on the FPGA 650 perform the relatively simple operation. Therefore, the application 634 running on the processor 620 offloads the execution of the operation to the AU 652. On the other hand, if the operation is a relatively complex operation with low I / O and high computational requirements, then it may be more efficient to use the application 634 running on the processor 620 to perform the operation.

[0034] Using FPGA to handle remotely initiated RDMOS

[0035] The above section describes the use of FPGA as an execution candidate to execute a locally initiated RDMO. In a similar manner, the FPGA can be used as an execution candidate to execute a remotely initiated RDMO. For example, assume that an external requesting entity 610 sends a RDMO request to the NIC 602. If the RDMO is an operation supported by the AU 652, then the AU 652 can perform the operation more efficiently than the NIC 602 itself or the application 634 running on the processor 620. Therefore, the NIC 602 will let the AU 652 implemented on the FGPA 650 execute the RDMO. On the other hand, if the operation is a relatively complex operation with low I / O and high computational requirements, then it can be more efficient to use the application 634 running on the processor 620 to perform the operation.

[0036] Support for new operations can be added by reprogramming the FPGA. Such reprogramming can be performed, for example, by loading a revised FPGA bitstream into the FPGA at power-up.

[0037] A set of common RDMOs for accelerating OLTP and HTAP

[0038] If the workload is not amenable to partitioning, scaling out transactions is challenging. RDMA has recently generated a lot of excitement in the research community as a mechanism to run multi-version concurrency control algorithms directly in remote memory. However, many simple manipulations cannot be performed in a single RDMA operation. For example, an update in lock-based concurrency control requires at least three RDMA requests to lock (RDMA CAS), write (RDMA write), and unlock (RDMA write). An update in an MVCC (multi-version concurrency control) algorithm using immutable versions and wait-free readers will require a write (RDMA write) to the new version, a read (RDMA read) to the old version, and a compare and exchange (RDMA CAS) to atomically modify the visibility of the old version. In either case, a single write requires three interdependent RDMA operations to be executed in sequence. As a result, transaction performance is severely affected by the latency of remote data access. The following describes a set of common RDMOs that perform a predefined short series of read / write / CAS operations in one network request to speed up transaction performance.

[0039] SLOTTEDAPPEND (slot append) RDMO

[0040] Many data structures in database systems, including disk pages, are divided into variable-length slots that are accessed through indirection. The simplest form of this indirection is an offset field that points to the first byte of an entry based on the position of the offset. The SlottedAppend RDMO safely manipulates such structures in a single network request. The pseudo code is as follows:

[0041]

[0042] SlottedAppend RDMO reads * position 110 (see Figure 1 ) and it atomically increments it by the length of the buffer to be written. If the entire length of the buffer fits maxsize bytes, then the buffer is copied at *old position 110a and the write position is returned. If writing to the buffer would exceed the allowable size maxsize, then *position 100 is decremented and the special value maxsize is returned to signal that there is insufficient space to complete the write.

[0043] SlottedAppend is designed to improve write contention. It reduces the conflict window and allows more concurrency under contention compared to RDMA-only implementations. It does not protect readers from incomplete reads or in-progress writes. In embodiments where reader / writer isolation is desired, other coordination mechanisms (such as locks or multi-version processing) are used to enforce this isolation.

[0044] Example of use: An example of the use of SlottedAppend in OLTP workload is to insert a tuple 108 into page 104, such as Figure 1 The corresponding page 104 is first located in memory. SlottedAppend then atomically increments the free pointer 102 and writes the data to the slotted page 104 using a fetch-and-add operation in one network operation.

[0045] CONDITIONALGATHER (conditional collection) RDMO

[0046] Traversing common data structures, including B-trees and hash tables, requires tracking pointers. Performing a series of such lookups over RDMA is inefficient because it requires selectively checking each element over RDMA. The ConditionalGather RDMO traverses pointer-based data structures in a single request and gathers elements that match a user-defined predicate in a buffer that is passed back.

[0047]

[0048] The ConditionalGather RDMO collects the output in the gatherbuf buffer 220 (see Figure 2) and will use at most gatherlen bytes. ConditionalGather reads count elements starting at position 202. Each pointed-to element is compared to parameter value 214 using operator COMP 212∈{<, ≤, =, ≥, >, ≠}. If predicate 210 matches, then ConditionalGather copies length bytes from the compared position into gatherbuf buffer 204. This process is repeated until gatherbuf buffer 220 is full or count entries have been processed. ConditionalGather returns the number of elements of size length in gatherbuf buffer 220.

[0049] The ConditionalGather RDMO has two interesting variants. Instead of accessing an indirect array in remote memory, the first variant accepts a local array and transfers it to the remote side. This variant is useful when the location is known from a previous RDMA operation or the gather operation retrieves at a fixed offset (the latter corresponds to a strided access pattern). The second variant does not require a fixed length parameter to be passed, but instead accepts a remote array with a record length. This variant allows gathering on structures of variable length.

[0050] Example use: Multi-version concurrency control can point snapshot transactions to older versions of records to achieve higher concurrency. Identifying the appropriate record to read in a page requires determining version visibility. Performing this visibility check on RDMA operations is inefficient: selectively checking each version through RDMA operations is too lengthy and makes the entire operation latency-bound. Figure 2 As shown in , ConditionalGather RDMO can return the visible version of a certain timestamp in one operation. In HTAP workloads involving joins, ConditionalGatherRDMO can retrieve all tuples matching a key in a hash bucket in one round trip.

[0051] SignaledRead (Signaled Read) RDMO

[0052] Many data structures have no known lock-free alternatives, or their lock-free alternatives impose strict constraints on the types of concurrent operations supported. These data structures often fall back to lock-based synchronization. However, exposing lock-based data structures over RDMA requires at least three RDMA requests: two RDMA operations targeting the lock, and at least one RDMA required to perform the intended operation. When reading in lock-based data structures, the SignaledRead RDMO can be used to "elide" the RDMA operations for the lock, similar to speculative lock elision in hardware. This avoids two round trips to acquire and release the lock.

[0053]

[0054] SignaledRead RDMO attempts to read cas position 300 (see Figure 3 ) performs a compare and swap operation and changes the value from expectedval to desiredval. If the compare and swap operation fails, the RDMO retries the compare and swap several times and then returns the last read value (the purpose is to amortize the round-trip network latency over several retries in the case of a spurious failure). If the compare and swap succeeds, the SignaledRead RDMO reads length bytes at read position 310 and then resets cas position 300 to expectedval.

[0055] Note that the first four bytes of read location 310 are copied last, after the memory fence from the previous memory copy. The purpose is to allow these first four bytes to serve as a test of whether the memory copy was completed successfully in the event that read location 310 points to a crash of non-volatile memory.

[0056] Example use: The SignaledRead RDMO can be used to omit explicit RDMA operations for locks when reading into lock-based data structures. This greatly increases concurrency: avoiding two round trips for acquiring and releasing a lock means that the lock is held for a much shorter duration. Using SignaledRead guarantees that the read is consistent, since incoming updates will not be able to acquire the lock.

[0057] WriteAndSeal (Write and Seal) RDMO

[0058] A variant of SignaledRead is the WriteThenSeal RDMO. This RDMO can be used in lock-based synchronization to update and release locks in one RDMA operation. If the seal location is in persistent storage, it is not possible to provide torn write protection because there are no atomicity or consistency guarantees for the sealing write. Using WriteThenSeal for persistence will require atomicity and persistence primitives for non-volatile storage, as discussed below.

[0059]

[0060] WriteAndSeal RDMO writes length bytes of data into the buffer at the write position, waits for the memory fence, and then writes the value at the seal position. Note that the write position and the seal position can overlap. The operation is guaranteed to be sequential, but may not be atomic.

[0061] SCATTERANDACCUMULATE (scatter and accumulate) RDMO

[0062] A common parallel operation is the scatter, where a dense array of elements is copied into a sparse array. A more useful primitive for database systems is a scatter operation that involves addressing the destination indirectly (such as through a hash function or an index). Furthermore, database systems often want to accumulate values ​​into values ​​that already exist in the destination, rather than overwriting that data. The ScatterAndAccumulate RDMO achieves this goal.

[0063]

[0064] The ScatterAndAccumulate RDMO accepts an array of element scatter lists 420 of length count. For each element, ScatterAndAccumulate accesses the base address baseaddr 430 in the offset list 410 array and the user defined offset. ScatterAndAccumulate then accumulates the results of the operations OP∈{FETCHANDADD, SUM, MAX, MIN, COUNT} between the incoming data and the data already at the destination location.

[0065] Example of use: In HTAP workloads, ScatterAndAccumulate RDMO can accelerate parallel hash-based aggregation: the sender performs local aggregation and then calculates the hash bucket destination for each element. ScatterAndAccumulate then transfers the data and performs aggregation on these buckets in one operation.

[0066] In addition, this RDMO can be used for min-hash sketching for set similarity estimation.

[0067] OLTP kernel with RDMO capabilities

[0068] To best utilize these RDMOs in regular transaction processing activities, RDMOs can be integrated into the storage engine of the database system and in-memory storage structures (such as indexes) can be directly exposed through RDMA. This design does not require a lot of changes to the database system architecture, but it does not use RDMO for transaction management modules and concurrency control. Therefore, the RDMO-aware storage engine will overcome inefficiencies at the physical data access level, such as optimizing or omitting latches (i.e., short-term locks that protect critical sections that are several instructions long).

[0069] Further efficiency gains can be achieved by integrating RDMO functionality into the concurrency control protocol itself. This is consistent with recent research that combines concurrency control and storage to achieve high performance in certain transaction processing systems. In one embodiment, RDMO is used to verify the serializability of distributed transactions by directly accessing remote storage rather than using a storage engine as an "intermediary" to access shared state. In addition, this design choice allows us to interweave transaction management aspects (such as admission control, transaction restart, and service level objectives) in the concurrency control protocol itself. The opportunity is to use RDMO to bypass or accelerate synchronization at the logical level, such as synchronization for write anomalies or phantoms. This logical synchronization often requires locks that are maintained for the entire duration of the transaction.

[0070] Reliable communication in data centers via RDMA and RDMO

[0071] Current RDMA-based OLTP solutions ignore the possibility of data corruption. However, a single bit flip in NVRAM triggers a hardware error that breaks the RDMA connection (in the best case) or crashes the entire database system (in the worst case). The transactional capabilities of database systems present interesting research opportunities: if data corruption is encapsulated within the concept of a transaction, then the database system can handle data corruption more gracefully than general non-transactional applications by relying on existing abort / retry / log replay actions.

[0072] Key Management and Encryption

[0073] RDMA has been designed for trusted components within a data center to communicate efficiently. The RDMA protocol has a coarse authentication mechanism in the form of 32-bit R-keys and Q-keys, which verify that the remote sender has access to the local memory window and receive queue, respectively. However, the payload transmitted over the network is not encrypted, as this would typically consume a lot of CPU resources on both sides. In addition, it is not possible to encrypt RDMA message headers without making it impossible for the HCA to interpret these headers.

[0074] However, enterprises are increasingly concerned about malicious actors being able to intercept information at the physical network level in the data center. In practice, RDMA exacerbates network security vulnerabilities: a malicious actor who intercepts a transmitted RDMA key can forge a valid RDMA request to read or modify the entire memory window associated with this key. Furthermore, because RDMA bypasses the CPU at the destination, malicious network activity is ignored by the OS and receiving applications, so malicious network activity will not be reflected in network statistics or audit logs on the target node.

[0075] In one embodiment, end-to-end encryption is provided for applications connecting to RDMA. This includes reliably forming connections and transferring data securely and privately. The current facilities will be used for authentication. RDMA queue pairs are created after a properly authenticated SSL connection has been created between the participants and the protocol and network to be used for RDMA have been negotiated. At some point in this negotiation, the application either creates a secure queue pair for communication only between the two processes, or accesses the appropriate Q key to access the endpoint. The queue pair can be destroyed by the operating system that created the queue pair or by a compromised application that holds a valid queue pair, R key, or Q key.

[0076] An alternative is to use DTLS, which works in any application on top of an unreliable datagram protocol, but this interferes with hardware acceleration. The goal is to investigate whether IPSec over RoCE can be modified to make new connections cheaper.

[0077] Data Corruption and NVRAM

[0078] Low-latency, high-capacity NVRAMs (such as Intel 3DXP) will soon be commercially available. Future NVRAMs will have higher density than DRAM. However, a key prerequisite for enterprise deployment is memory media corruption. In addition, since NVRAM modules will be distributed across memory channels and CPU sockets, the probability of corruption of each cache line may not be uniform. Therefore, the inevitable result is that some memory locations (whether due to wear or firmware errors) will cause subtle data corruption or become inaccessible. Such hardware errors will be reported to the OS through the Machine-Check-Architecture (MCA) facility. MCA provides a mechanism to handle memory corruption instead of crashing the entire machine or propagating incorrect values.

[0079] Enabling RDMA access to NVRAM is a necessary optimization in converged datacenter architectures, as accessing remote memory through existing OS primitives will overshadow the DRAM-like speed of NVRAM. Since low-latency NVRAM will be directly attached to the memory bus, this means that remote access from the HCA must sense hardware errors reported through the MCA and check the memory poison bit. However, the current RDMA framework has no facilities to detect, handle, and clear the poison bit. One possibility is that the RDMA client will silently receive corrupted data. This can be mitigated by a checksum (either in software or in hardware). However, a more serious error would be to lose the entire RDMA connection, which would make the entire node unavailable due to a single piece of data corruption. Our solution includes two approaches with different completion timelines.

[0080] In one embodiment, a pure software solution is used, which involves only the network device driver without any firmware or hardware changes. The mitigation mechanism couples an application-aware retry mechanism with a corruption detection facility.

[0081] When data corruption generates a machine check exception (MCE), the exception in turn triggers a kernel panic. This manifests as a disconnect from the RDMA client. Embodiments may also include transparent software-level retries that require a higher-level API for retryable data access, such as a key-value store API or a transactional interface for a DBMS.

[0082] It may also be the case that a node suppresses MCE and silently returns corrupted memory. In this case, a checksum-based corruption detection mechanism is used. This requires a higher level API to maintain the checksum and retry access. One embodiment involves hardware offload of the HCA.

[0083] An alternative embodiment involves hardware / software co-design. In particular, the stack is made aware of the memory poison bit so that user mode applications can choose to handle memory corruption in the application. This requires changes to the firmware and drivers of the network device, and also extends the standard RDMA verbs API with a callback registration mechanism for handling data corruption.

[0084] Hardware Overview

[0085] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing device may be hardwired to perform these techniques, or may include digital electronic devices (such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) permanently programmed to perform these techniques), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a special-purpose computing device may also combine customized hard-wired logic, ASICs, or FPGAs with customized programming to implement these techniques. The special-purpose computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device that combines hard-wiring and / or program logic to implement these techniques.

[0086] For example, Figure 5 5 is a block diagram illustrating a computer system 500 on which an embodiment of the present invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor. The computer system 500 also includes a main memory 506 coupled to the bus 502, such as a random access memory (RAM) or other dynamic storage device, for storing information and instructions to be executed by the processor 504. The main memory 506 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 504. When stored in a non-transitory storage medium accessible to the processor 504, these instructions make the computer system 500 a special-purpose machine customized to perform the operations specified in the instructions.

[0087] The computer system 500 also includes a read-only memory (ROM) 508 or other static storage device coupled to the bus 502 for storing static information and instructions for the processor 504. A storage device 510, such as a magnetic disk, optical disk, or solid-state drive, is provided and coupled to the bus 502 for storing information and instructions. The computer system 500 may be coupled to a display 512, such as a cathode ray tube (CRT), via the bus 502 for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to the bus 502 for communicating information and command selections to the processor 504. Another type of user input device is a cursor control 516, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 504 and for controlling cursor movement on the display 512. Such input devices typically have two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), which allows the device to specify a position in a plane.

[0088] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that is combined with a computer system to make or program computer system 500 into a special purpose machine. According to one embodiment, computer system 500 performs the techniques described herein in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. These instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the processing steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.

[0089] As used herein, the term "storage medium" refers to any non-transient medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical disks, magnetic disks, or solid-state drives, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, PROM and EPROM, FLASH-EPROM, NVRAM, any other memory chip, or cassette tape.

[0090] Storage media are distinct from but can be used in conjunction with transmission media. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wire, and optical fiber, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0091] Various forms of media may be involved in carrying one or more sequences of one or more instructions to the processor 504 for execution. For example, the instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 500 may receive the data on the telephone line and use an infrared transmitter to convert the data into an infrared signal. An infrared detector may receive the data carried in the infrared signal and appropriate circuitry may place the data on the bus 502. The bus 502 carries the data to the main memory 506 from which the processor 504 retrieves and executes the instructions. The instructions received by the main memory 506 may optionally be stored on the storage device 510 before or after execution by the processor 504.

[0092] The computer system 500 also includes a communication interface 518 coupled to the bus 502. The communication interface 518 provides a two-way data communication coupled to a network link 520, wherein the network link 520 is connected to a local network 522. For example, the communication interface 518 can be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, the communication interface 518 can be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link can also be implemented. In any such implementation, the communication interface 518 sends and receives electrical signals, electromagnetic signals, or optical signals that carry digital data streams representing various types of information.

[0093] The network link 520 typically provides data communication to other data devices through one or more networks. For example, the network link 520 can provide a connection to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526 through a local network 522. ISP 526 in turn provides data communication services through a global packet data communication network, now commonly referred to as the "Internet" 528. Both the local network 522 and the Internet 528 use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on the network link 520 and through the communication interface 518 (which carry the digital data to and from the computer system 500) are example forms of transmission media.

[0094] Computer system 500 may send messages and receive data, including program code, through the network(s), network link 520, and communication interface 518. In the Internet example, server 530 may send the requested code for an application program through Internet 528, ISP 526, local network 522, and communication interface 518. The received code may be executed by processor 504 as it is received, and / or stored in storage device 510 or other non-volatile storage for later execution.

[0095] cloud computing

[0096] The term "cloud computing" is used generally herein to describe a computing model that enables on-demand access to a shared pool of computing resources, such as computer networks, servers, software applications, and services, and allows resources to be rapidly provisioned and released with minimal management effort or service provider interaction.

[0097] In general, the cloud computing model enables some of those responsibilities that may have previously been provided by an organization's own information technology department to be delivered instead as a service layer within the cloud environment for use by consumers (either internally or externally to the organization, depending on the public / private nature of the cloud). Depending on the specific implementation, the precise definition of the components or features provided by or within each cloud service layer may vary, but common examples include: Software as a Service (SaaS), in which consumers use software applications running on the cloud infrastructure, while the SaaS provider manages or controls the underlying cloud infrastructure and applications. Platform as a Service (PaaS), in which consumers can use software programming languages ​​and development tools supported by the PaaS provider to develop, deploy, and otherwise control their own applications, while the PaaS provider manages or controls other aspects of the cloud environment (i.e., everything under the runtime execution environment). Infrastructure as a Service (IaaS), in which consumers can deploy and run arbitrary software applications, and / or provide processing, storage, networking, and other basic computing resources, while the IaaS provider manages or controls the underlying physical cloud infrastructure (i.e., everything below the operating system layer). Database as a Service (DBaaS), in which the consumer uses a database server or database management system running on a cloud infrastructure, while the DBaaS provider manages or controls the underlying cloud infrastructure, applications, and servers (including one or more database servers).

[0098] Extensions and replacements

[0099] In the foregoing specification, embodiments of the present invention have been described with reference to numerous specific details, which may vary from implementation to implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. The sole and exclusive indicator of the scope of the invention, and what the applicants intend as the scope of the invention, is the literal and equivalent range of the set of claims issuing from this application in the specific form in which such claims issue, including any subsequent corrections.

Claims

1. A computer-implemented method comprising: a process executing on a first processor causing a remote direct memory operation RDMO to be executed by an execution candidate that does not include the first processor; wherein the RDMO comprises a predefined sequence of two or more sub-operations; wherein each of the two or more sub-operations is one of: a read operation, a write operation, a compare and swap operation; and Wherein, causing the RDMO to be executed includes causing the execution candidate to execute one of the following: (a) Append the write value to the contents of the disk page by doing the following: reading a location value pointing to the location of an entry in the disk page; Based at least in part on the position value, determining whether the disk page includes free space of at least the size of the written value; and In response to determining that the disk page includes free space of at least the size of the written value: appending the write value to the contents of the disk page at the location of the entry in the disk page; (b) traversing the pointer-based data structure to identify one or more element values ​​stored at the pointer-based data structure that satisfy a predicate by: For each element value stored at the pointer-based data structure: determining whether each element value satisfies the predicate by comparing each element value with a parameter value using a comparison operator, Wherein, the comparison operator is one of the following: less than, less than or equal to, equal to, greater than or equal to, greater than, not equal to, and In response to determining whether each element value satisfies the predicate, including the value in a result buffer; and transmitting the result buffer to a remote machine; (c) reading from a lock-based data structure associated with the lock, wherein acquisition of the lock requires performance of operations in the lock-based data structure, the operations comprising: performing a compare and swap operation at a lock position of the lock; wherein, before performing the compare and swap operation, the lock position stores an original value; In response to determining success of performing the compare and swap operation on the lock position: reading from the lock-based data structure, and after reading from the lock-based data structure, resetting the lock position to the original value; (d) performing an accumulation operation to accumulate the value of the scatter list with the set of values ​​already present at the destination location by performing the following operation on each value in the scatter list: identifying an offset value corresponding to each of the values ​​in the scatter list; identifying a destination location for each value of the scatter list using the offset value and a base address; performing an accumulation operation between said each value of said scatter list and a value already existing at said destination location to produce an accumulated result value; and storing the accumulated result value at the destination location; The accumulation operation is one of FETCHANDADD, SUM, MAX, MIN, and COUNT.

2. The computer-implemented method of claim 1, wherein: Causing the RDMO to be executed involves causing the execution candidate to execute: appending the write value to the contents of the disk page.

3. The computer-implemented method of claim 2, wherein: Appending the write value to the contents of the disk page also includes atomically incrementing the location value by a length of the write value.

4. The computer-implemented method of claim 1, wherein: Causing the RDMO to be executed involves causing the execution candidate to execute: traversing the pointer-based data structure to identify one or more elements stored at the pointer-based data structure that satisfy a predicate.

5. The computer-implemented method of claim 1 , wherein: Causing the RDMO to be executed involves causing the execution candidate to execute: Reading from a lock-based data structure.

6. The computer-implemented method of claim 1, wherein: Causing the RDMO to be executed involves causing the execution candidate to: perform an accumulation operation to accumulate the values ​​of the scatter list with the values ​​already present at a set of destination locations.

7. The computer-implemented method of claim 1, wherein: The first processor is on a first computing device; The execution candidate is on a second computing device remotely located relative to the first computing device; and Causing the RDMO to be executed by the execution candidate includes sending the RDMO to the second computing device via a single network request.

8. The computer-implemented method of claim 1, wherein: The first processor and the execution candidate are within a single computing device.

9. The computer-implemented method of claim 1, wherein: The execution candidate is a field programmable gate array (FPGA) configured to execute the RDMO.

10. A computer-implemented method comprising: a process executing on a particular processor sending a single request to have a remote direct memory operation RDMO executed by an execution candidate that does not include the particular processor; wherein the specific processor and the execution candidate are implemented on a specific computing device; wherein the RDMO is defined as requiring execution of a predefined sequence of two or more different sub-operations, wherein the predefined sequence of the two or more different sub-operations of the RDMO are dispatched to the execution candidate, and results of executing the predefined sequence of the two or more different sub-operations are received by the process in a single round trip to reduce round trip latency caused by multiple round trips associated with remote direct memory access (RDMA) operations; wherein the two or more different sub-operations include a first sub-operation and one or more additional sub-operations performed according to the predefined sequence; Each of the two or more different sub-operations is one of the following: a read operation, a write operation, a compare and swap operation.

11. The computer-implemented method of claim 10, wherein: The execution candidate is an auxiliary processor of the particular computing device.

12. The computer-implemented method of claim 11, wherein: The auxiliary processor is implemented on a field programmable gate array configured with logic to execute the RDMO.

13. The computer-implemented method of claim 11, wherein: The RDMO requires access to data in a local volatile memory of the particular computing device, and wherein the auxiliary processor has direct access to the local volatile memory of the particular computing device.

14. The computer-implemented method of claim 13, wherein: The auxiliary processor has direct access to the local volatile memory of the particular computing device via circuitry of the particular processor that is distinct from one or more cores of the particular processor.

15. The computer-implemented method of claim 10, wherein: Causing the RDMO to be executed by the execution candidate is performed in response to a determination that execution of the RDMO by the execution candidate is more efficient than execution of the RDMO by the particular processor.

16. The computer-implemented method of claim 10, wherein: The particular processor is in a first reliability domain and the execution candidate is in a second reliability domain.

17. A computer-implemented method comprising: a process executing on a specific processor causing a remote direct memory operation RDMO to be executed by an execution candidate that does not include the specific processor; wherein the RDMO includes a predefined subsequence of two or more sub-operations; wherein each of the two or more sub-operations is one of: a read operation, a write operation, a compare and swap operation; and The causing the RDMO to be executed includes causing the execution candidate to execute the content of appending the write value to the disk page by the following operations: determining whether the disk page includes free space of at least the size of the written value based at least in part on the position value of the content of the disk page; and In response to determining that the disk page includes free space of at least the size of the written value: appending the written value to the end of the contents of the disk page.

18. The computer-implemented method of claim 17, wherein: Appending the write value to the contents of the disk page also includes atomically incrementing the location value by a length of the write value.

19. A system configured to perform the method of any one of claims 1-18.

20. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the method of any one of claims 1-18 to be performed.

21. A computer program product comprising instructions which, when executed by one or more processors, cause the method of any one of claims 1 to 18 to be performed.

Citation Information

Patent Citations

  • Method for efficient primary key based queries using atomic RDMA reads on cache friendly in-memory hash index

    US20180341653A1

  • Gateway for connecting clients and servers utilizing remote direct memory access controls to separate data path from control path

    US8527661B1