One-sided reliable remote direct memory operation
By providing multiple execution candidates in a remote server, each belonging to different reliability domains, the problem of complex operations unavailable in the event of remote server failure is solved, and the availability and stability of the system is improved.
Patent Information
- Application Number
- CN201980052298.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-06
- Filing Date
- 2019-03-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-06-25
AI Technical Summary
The prior art is difficult to continue to perform more complex operations from the remote server when a remote server fails, resulting in reduced system availability.
By providing multiple execution candidates in a remote server, each belonging to a different reliability domain, the requesting entity may attempt different execution candidates to perform remote direct memory operations (RDMO), thereby improving the availability of operations.
This enables the ability to continue to perform complex operations when a remote server fails, improving the availability and stability of the system.
Smart Images

Figure CN112534412B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer systems and, more particularly, to techniques for increasing the availability of remote access functionality. Background Art
[0002] One way to improve the availability of a service is to design the service in a way that allows the service to continue to function normally even when one or more of its components fail. For example, U.S. Patent Application No. 15 / 606,322, filed on May 26, 2017 (which is incorporated herein by reference) describes a technique for enabling a requesting entity to retrieve data managed by a database server instance from a volatile memory of a remote server machine executing the database server instance without involving the database server instance in the retrieval operation.
[0003] Because the retrieval does not involve the database server instance, the retrieval operation can succeed even when the database server instance (or the host server machine itself) has stalled or become unresponsive. In addition to increased availability, directly retrieving data will generally be faster and more efficient than retrieving the same information through conventional interaction with the database server instance.
[0004] In order to retrieve the "target data" specified in the database command from the remote machine without involving the remote database server instance, the requesting entity first uses Remote Direct Memory Access (RDMA) to access information about the location where the target data resides in the server machine. Based on such target location information, the requesting entity uses RDMA to retrieve the target data from the host server machine without involving the database server instance. The RDMA read (data retrieval operation) issued by the requesting entity is a unilateral operation and does not require CPU interrupts or OS kernel participation on the host server machine (RDBMS server).
[0005] RDMA technology works well for operations that only involve retrieving data from the volatile memory of a crashed server machine. However, it is desirable to provide high availability even when the failed component is responsible for performing operations more complex than just memory access. To meet this need, some systems provide a limited set of "verbs" for performing remote operations such as memory access and atomic operations (test and set, compare and swap) via a network interface controller. These operations can be completed as long as the system is powered on and the NIC can access the host memory. However, the types of operations they support are limited.
[0006] Typically, more complex operations are performed on data residing in the memory of a remote machine by making remote procedure calls (RPCs) to an application running on the remote machine. For example, a database client that desires the average of a set of numbers stored in a database can make an RPC to a remote database server instance that manages the database. In response to the RPC, the remote database server instance reads the set of numbers, calculates the average, and then sends the average back to the database client.
[0007] In this example, if the remote database server fails, the average calculation operation fails. However, it may be possible to use RDMA to retrieve each number in the group from the remote server. Once the requesting entity has retrieved each number in the group of numbers, the requesting entity can perform the average calculation operation. However, using RDMA to retrieve each number in the group and then perform the calculation locally is much less efficient than having an application on the remote server retrieve the data and perform the average calculation operation. Therefore, it is desirable to expand the range of operations that continue to be available from a remote server when the remote server is not fully functional.
[0008] The approaches described in this section are approaches that could be pursued, but are not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be admitted that any approach described in this section is prior art merely by virtue of its inclusion in this section. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In the figure:
[0010] Figure 1 is a block diagram of a requesting entity interacting with a remote machine including a number of execution candidates for a requested operation, wherein each execution candidate is implemented in a different reliability domain, according to an embodiment;
[0011] Figure 2 is a flow chart illustrating the use of backup execution candidates according to an embodiment;
[0012] Figure 3 is a flow chart illustrating the use of parallel execution candidates according to an embodiment; and
[0013] Figure 4 is a block diagram of a computer system that can be used to implement the techniques described herein. DETAILED DESCRIPTION
[0014] In the following description, for the purpose of explanation, many specific details are set forth to provide a thorough understanding of the present invention. However, it will be apparent that the present invention can be practiced without these specific details. In other examples, well-known structures and devices are shown in block diagram form to avoid unnecessary obscurity of the present invention.
[0015] General Overview
[0016] This article describes techniques for allowing machines that are not fully functional to perform more complex operations remotely. Operations that can be reliably performed by machines that have experienced hardware and / or software errors are referred to herein as remote direct memory operations or "RDMOs." Unlike RDMA, which typically involves very simple operations (such as retrieving a single value from the memory of a remote machine), RDMOs can be arbitrarily complex. For example, RDMO can enable a remote machine to calculate the average of a set of numbers, where the numbers reside in the memory of the remote machine. When there are software failures or glitches on the remote system with which the application interacts, the techniques described herein can help the application run without interruption.
[0017] Execution Candidates and Reliability Domains
[0018] According to one embodiment, within a single machine, multiple entities are provided to execute the same RDMO. Entities that are capable of executing a particular RDMO within a given machine are referred to herein as "execution candidates" for the RDMO.
[0019] According to one embodiment, although in the same machine, the execution candidates of the RDMO belong to separate reliability domains. The "reliability domain" of an execution candidate generally refers to the software / hardware that must function correctly on the machine in order for the execution candidate to correctly execute the RDMO. If one of the execution candidates is able to correctly execute the RDMO, and software and / or hardware errors have prevented the other execution candidate from correctly executing the RDMO, then the two execution candidates belong to different reliability domains.
[0020] Because multiple execution candidates for a RDMO belong to different reliability domains, it is possible that one of the execution candidates will execute the RDMO during a period of time when a hardware / software error within the machine prevents other execution candidates in the machine from executing the RDMO. The fact that multiple execution candidates are available for a particular RDMO increases the likelihood that the RDMO will succeed when requested by a requesting entity that is not resident on the machine.
[0021] Increase the availability of RDMO
[0022] When the remote server is executing multiple execution candidates for a specific RDMO, and these execution candidates are from different reliability domains, the availability of the RDMO increases. For example, in one embodiment, when a requesting entity requests a specific RDMO, an attempt is made to execute the specific RDMO using the first execution candidate on the machine. If the first execution candidate cannot execute the RDMO, an attempt is made to execute the specific RDMO using the second execution candidate on the machine. This process can continue until the specific RDMO succeeds, or all execution candidates have been tried.
[0023] In an alternative embodiment, when a requesting entity requests a specific RDMO, the specific RDMO may be tried simultaneously by two or more of the execution candidates. If any execution candidate succeeds, the success of the specific RDMO is reported to the requesting entity.
[0024] In one embodiment, the execution candidate of the RDMO can be an entity that has been dynamically programmed to execute the RDMO. For example, a computing unit in a network interface controller of a machine can execute an interpreter. In response to determining that a specific RDMO should be executed in a network controller rather than by an application running on the machine, data including instructions for executing the specific RDMO can be sent to the interpreter. In response to interpreting those instructions at the network controller, the specific RDMO can be executed even if the application itself may have crashed.
[0025] Reliability Domain
[0026] As described above, the "reliability domain" of an execution candidate generally refers to the software / hardware that must function correctly in order for the execution candidate to successfully execute a RDMO. Figure 1 1 is a block diagram of a system including a machine 100 in which execution candidates for RDMOs reside in multiple different reliability domains. Specifically, a requesting entity 110 can request a specific RDMO by sending a request to the machine 100 via a network 112. The machine 100 includes a network interface controller 102 having one or more computing units (represented as processing units 104), which is executing firmware and / or software loaded into its local volatile memory 106.
[0027] The machine 100 includes a processor 120 that executes an operating system 132 and any number of application programs, such as application 134. The code for the operating system 132 and application 134 may be stored in a persistent storage device 136 and loaded into volatile memory 130 as needed. The processor 120 may be one of any number of processors within the machine 100. The processor 120 itself may have many different computing units, shown as cores 122, 124, 126, and 128.
[0028] Processor 120 includes circuitry (shown as uncore 142 ) that allows an entity executing in network interface controller 102 to access data 140 in volatile memory 130 of machine 100 without involving cores 122 , 124 , 126 , and 128 .
[0029] In a remote server such as machine 100, an entity capable of executing the RDMO requested by requesting entity 110 may reside in any number of reliability domains, including but not limited to any of the following:
[0030] Hardware within the network interface controller 102
[0031] FPGA in the network interface controller 102
[0032] Firmware stored in the network interface controller 102
[0033] Software executed by the processing unit 104 within the network interface controller 102
[0034] Instructions in volatile memory 106 interpreted by an interpreter being executed by processing unit 104 within network interface controller 102
[0035] An operating system 132 loaded into the volatile memory 130 of the machine 100 and executed by one or more cores (122, 124, 126, and 128) of the processor 120
[0036] An application 134 loaded into the volatile memory 130 of the machine 100 and executed by one or more cores (122, 124, 126, and 128) of the processor 120
[0037] - Software executing in a privileged domain (eg, "domO"). Privileged domains and domO are described, for example, at en.wikipedia.org / wiki / Xen.
[0038] This reliability domain list is exemplary only, and the techniques described herein are not limited to execution candidates from these reliability domains. The examples given above serve as different reliability domains because entities within these domains fail under different conditions. For example, the operating system 132 may continue to function normally even when the application 134 has crashed or otherwise failed. Similarly, execution candidates running within the network interface controller 102 may continue to function normally even when all processes being executed by the processor 120 (including the operating system 132 and the application 134) have crashed.
[0039] Furthermore, because each core within processor 120 is a computing unit that can fail independently of other computing units within processor 120 , the execution candidates being executed by core 122 are in a different reliability domain than the execution candidates being executed by core 124 .
[0040] Add support for new RDMO
[0041] The machine 100 may include dedicated hardware in the network interface controller 102 or elsewhere to implement a particular RDMO. However, hardware-implemented execution candidates cannot be easily extended to support additional RDMOs. Therefore, according to one embodiment, a mechanism is provided for adding support for additional RDMOs to other types of execution candidates. For example, assume that the NIC 102 includes firmware for executing a specific set of RDMOs. Under these conditions, support for additional RDMOs can be added to the NIC 102 through conventional firmware update techniques.
[0042] On the other hand, if the execution candidate is implemented in an FPGA, support for the new RDMO can be added by reprogramming the FPGA. Such reprogramming can be performed, for example, by loading a modified FPGA bitstream into the FPGA at power-up.
[0043] Similarly, support for new RDMOs can be added to software-implemented execution candidates (e.g., software in NIC 102, operating system 132, and application 134) using conventional software update techniques. In embodiments involving execution of an interpreter within NIC 102, new RDMOs can be supported by sending code implementing the new RDMO to NIC 102. NIC 102 can store the code in volatile memory 106 and execute the new RDMO by interpreting the code. For example, NIC 102 may be executing a Java virtual machine, and requesting entity 110 can cause NIC 102 to execute a new RDMO (e.g., calculate the average of a set of numbers) by sending Java bytecode to NIC 102, which, when interpreted by the Java virtual machine, causes the target set of numbers to be retrieved from volatile memory 130 and its average value to be calculated.
[0044] Example RDMO
[0045] In the previous discussion, an example is given in which the RDMO in question calculates the average value of a set of values. This is just an example of an RDMO that can be supported by execution candidates from multiple different reliability domains. As mentioned above, an RDMO may be arbitrarily complex. However, the more complex the RDMO, the greater the resources required for efficient execution of the RDMO, and the greater the possibility that the execution candidate of the RDMO will encounter problems. In addition, complex RDMOs may be executed slowly when executed by execution candidates from reliability domains with limited resources. For example, a complex RDMO that is usually executed by an application 134 executed by all computing units using a processor 120 will take significantly longer time if executed by a relatively "lightweight" processing unit 104 on a NIC 102.
[0046] Examples of RDMOs that can be supported by multiple execution candidates from different reliability domains within the same machine include, but are not limited to:
[0047] ●Setting a number of flags in the remote server machine to indicate that the corresponding portion of the remote server machine's volatile memory is unavailable or corrupted
[0048] ● Adjusting state within the remote server machine so that the remote server machine responds to certain types of requests with error notifications. This can avoid the need to notify all other computing devices when the remote server machine cannot properly process certain requests.
[0049] Performing a batch of changes to a secondary copy of a data structure (stored in volatile and / or persistent memory at a remote server machine) so that the secondary copy reflects changes made to a primary copy of the data structure (stored in volatile and / or persistent memory of another machine).
[0050] - Any algorithm involving branching (eg, logic configured to test one or more conditions and determine which way to branch (and therefore which actions to perform) based on the test results).
[0051] ● Any algorithm that involves multiple reads and / or multiple writes to a remote server's volatile memory. Because each read and / or write does not involve RDMA communication between machines, performing multiple reads and / or writes in response to a single RDMA request is not a problem.
[0052] OR-write significantly reduces inter-machine communication overhead.
[0053] ● from a persistent storage device operatively coupled to a remote server machine (e.g.,
[0054] 136). The ability to use execution candidates in NIC 102 to perform this operation is particularly useful when the persistent storage is not directly accessible to the requesting entity, so that the data on the persistent storage 136 remains available to the requesting entity even when the software (potentially including the operating system 132 itself) running on the computing units (e.g., cores 122, 124, 126, and 128) of the remote machine may have crashed.
[0055] Backup execution candidates
[0056] As described above, the availability of RDMO can be increased by having RDMO execution candidates from one reliability domain serve as backups for RDMO execution candidates from another reliability domain. The selection of which execution candidate to be the primary candidate for RDMO and which to be the backup candidate may differ based on the characteristics of the candidates and the nature of the RDMO.
[0057] For example, in the case where the RDMO is used for relatively simple operations, it is possible that the execution candidate running in the NIC 102 can perform the operation more efficiently than the application 134 running on the processor 120. Therefore, for this particular RDMO, it may be desirable to have the execution candidate running in the NIC 102 become the primary execution candidate for the RDMO, and the application 134 become the backup execution candidate. On the other hand, if the RDMO is relatively complex, it may be more efficient to use the application 134 as the primary execution candidate for the RDMO and the execution candidate on the NIC 102 as the backup execution candidate.
[0058] Figure 22 is a flow chart showing the use of backup execution candidates according to an embodiment. At step 200, a requesting entity sends a request for RDMO to a remote computing device via a network. At step 202, a first execution candidate at the remote computing device attempts to execute RDMO. If RDMO is successfully executed (determined at step 204), a response indicating success (which may contain additional result data) is sent back to the requesting entity at step 206. Otherwise, at step 208, it is determined whether another execution candidate that has not yet been attempted can be used to execute RDMO. If no other execution candidates are available for RDMO, an error indication is sent back to the requesting entity at step 210. Otherwise, another execution candidate is attempted to execute RDMO.
[0059] This process can continue until the RDMO is successfully executed or all execution candidates of the RDMO have failed. In one embodiment, the entity at the remote server is responsible for traversing the execution candidates. In such an embodiment, the attempt to execute the RDMO using the backup execution candidate is performed without notifying any failure to the requesting entity until all possible execution candidates of the RDMO have failed. This embodiment reduces the machine-to-machine message traffic generated during the traversal process.
[0060] In an alternative embodiment, the requesting entity is notified each time an execution candidate fails to execute the RDMO. In response to the failure indication, the requesting entity determines which execution candidate (if any) to try next. In response to determining that a specific execution candidate should be tried next, the requesting entity sends another request to the remote computing device. The new request indicates that a backup execution candidate for the RDMO should then be attempted.
[0061] It should be noted that the failure of an execution candidate may be indicated implicitly rather than explicitly. For example, if the execution candidate has not confirmed success, such as after a certain amount of time has passed, it can be assumed that the execution candidate has failed. In such a case, the requesting entity may issue a request for execution of the RDMO by another execution candidate before receiving any explicit indication that the previously selected execution candidate has failed.
[0062] The actual execution candidates that are attempted for any given RDMO and the order in which the execution candidates are attempted may vary based on a variety of factors. In one embodiment, the execution candidates are attempted in an order based on their likelihood of failure, with the application running on the remote computing device (which may be the most likely candidate to crash) being attempted first (because it has more available resources) and the hardware-implemented candidate in the NIC (which may be the least likely candidate to crash) being attempted last (because it has the fewest available resources).
[0063] As another example, in the event that it appears that the processor of the remote computing device is overloaded or crashed, the requesting entity may first request that the RDMO be performed by an entity implemented in the network interface controller of the remote computing device. On the other hand, if there is no such indication, the requesting entity may first request that the RDMO be performed by an application running on the remote machine, which may take full advantage of the computing hardware available on the remote machine.
[0064] As another example, for a relatively simple RDMO, the requesting device may first request a relatively "lightweight" execution candidate implemented in the network interface controller to execute the RDMO. On the other hand, a request for a relatively complex RDMO may first be sent to an application on a remote computing device, and only sent to a lightweight execution candidate if the application fails to execute the RDMO.
[0065] Parallel execution candidates
[0066] Also as described above, the availability of the RDMO can be increased by having multiple execution candidates from different reliability domains attempt to execute the RDMO in parallel. For example, assume that the RDMO will determine the average of a set of numbers. In response to a single request from the requesting entity 110, both the execution candidate implemented in the NIC 102 and the application 134 can be called to execute the RDMO. In this example, both the application 134 and the execution candidate in the NIC 102 will read the same group of values (e.g., data 140) from the volatile memory 130, count the numbers in the group, sum the numbers in the group, and then divide the sum by the count to obtain the average. If neither execution candidate fails, both execution candidates can provide their responses to the requesting entity 110, which can simply discard the duplicate responses.
[0067] In the event that one execution candidate fails and one execution candidate succeeds, the non-failed execution candidate returns a response to the requesting entity 110. Thus, even though something within the machine 100 did not function correctly, the RDMO completed successfully.
[0068] In the event that all execution candidates that initially attempt the RDMO fail, a second set of execution candidates from a different reliability domain than the first set of execution candidates can attempt to execute the RDMO in parallel. In the event that at least one backup execution candidate succeeds, the RDMO succeeds. This process can continue until one of the execution candidates of the RDMO succeeds, or all execution candidates of the RDMO fail.
[0069] Some "coordinating entity" within the machine 100 may call a group of execution candidates to execute the RDMO in parallel, rather than all successful execution candidates returning responses to the requesting entity 110. If multiple candidates succeed, the successful candidates may provide responses back to the coordinating entity, which then returns a single response to the requesting entity 110. This technique simplifies the logic of the requesting entity 110, making it transparent to the requesting entity 110 how many execution candidates on the machine 100 are being asked to execute the RDMO and which of those execution candidates succeeded.
[0070] Figure 3 is a flow chart illustrating the use of parallel execution candidates according to an embodiment. Figure 3 , at step 300, the requesting entity sends a request for the RDMO to a remote computing device. At steps 302-1 to 302-N, each of the N execution candidates of the RDMO simultaneously attempts to execute the RDMO. At step 304, if any execution candidate succeeds, control passes to step 306, where an indication of success (which may include additional result data) is sent to the requesting entity. Otherwise, if all fail, an indication of failure is sent to the requesting entity at step 310.
[0071] Interpreters as NIC-based execution candidates
[0072] As mentioned above, one form of execution candidate of RDMO is interpreter. In order to make the interpreter act as the execution candidate of RDMO, code is provided to the interpreter, and the code performs the operation required by RDMO when interpreted. Such interpreter can be executed by processing unit 104 in NIC 102, by processor 120 or by a subset of the core of processor 120, for example.
[0073] According to one embodiment, the code for a specific RDMO is registered with an interpreter. Once registered, a requesting entity can call the code (e.g., via a remote procedure call) so that the code is interpreted by the interpreter. In one embodiment, the interpreter is a Java virtual machine, and the code is a Java bytecode. However, the technology used herein is not limited to any particular type of interpreter or code. Although the RDMO executed by the interpreter in NIC 102 may take much longer than the time spent by a compiled application (e.g., application 134) executed on processor 120 to execute the same RDMO, the interpreter in NIC 102 may be operable during a time period when some errors prevent the operation of the compiled application.
[0074] Hardware Overview
[0075] According to one embodiment, the technology described herein is implemented by one or more special-purpose computing devices. The special-purpose computing device can be hardwired to perform the technology, or can include digital electronic devices, such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are permanently programmed to perform the technology, or can include one or more general-purpose hardware processors that are programmed to perform the technology according to program instructions in firmware, memory, other storage devices, or a combination thereof. Such a special-purpose computing device can also combine customized hard-wired logic, ASICs or FPGAs with customized programming to implement the technology. The special-purpose computing device can be a desktop computer system, a portable computer system, a handheld device, a network device, or any other device that combines hard-wiring and / or program logic to implement the technology.
[0076] For example, Figure 4 4 is a block diagram illustrating a computer system 400 upon which embodiments of the present invention may be implemented. Computer system 400 includes a bus 402 or other communication mechanism for communicating information, and a hardware processor 404 coupled to bus 402 for processing information. Hardware processor 404 may be, for example, a general purpose microprocessor.
[0077] The computer system 400 also includes a main memory 406, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 402 for storing information and instructions to be executed by the processor 404. The main memory 406 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the processor 404. Such instructions, when stored in a non-transitory storage medium accessible to the processor 404, make the computer system 400 a special-purpose machine customized to perform the operations specified in the instructions.
[0078] Computer system 400 also includes a read only memory (ROM) 408 or other static storage device coupled to bus 402 for storing static information and instructions for processor 404. A storage device 410, such as a magnetic disk, optical disk, or solid state drive, is provided and coupled to bus 402 for storing information and instructions.
[0079] The computer system 400 may be coupled to a display 412, such as a cathode ray tube (CRT), via the bus 402 for displaying information to a computer user. An input device 414, including alphanumeric and other keys, is coupled to the bus 402 for communicating information and command selections to the processor 404. Another type of user input device is a cursor control 416, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 404 and for controlling cursor movement on the display 412. The input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allow the device to specify a position in a plane.
[0080] Computer system 400 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that is combined with the computer system to make or program the computer system 400 into a special purpose machine. According to one embodiment, the techniques herein are performed by computer system 400 in response to processor 404 executing one or more sequences of one or more instructions contained in main memory 406. Such instructions may be read into main memory 406 from another storage medium, such as storage device 410. Execution of the sequences of instructions contained in main memory 406 causes processor 404 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0081] The term "storage medium" as used herein refers to any non-temporary medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical disks, magnetic disks, or solid-state drives, such as storage device 410. Volatile media include dynamic memory, such as main memory 406. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with hole patterns, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips, or cassette tapes.
[0082] Storage media are distinct from, but may be used in conjunction with, transmission media. Transmission media participate in the transmission of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including the wires that comprise bus 402. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0083] Carrying one or more sequences of one or more instructions to the processor 404 for execution may involve various forms of media. For example, the instructions may initially be carried on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to the computer system 400 may receive the data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on the bus 402. The bus 402 carries the data to the main memory 406, from which the processor 404 retrieves the instructions and executes them. The instructions received by the main memory 406 may optionally be stored on the storage device 410 before or after being executed by the processor 404.
[0084] Computer system 400 also includes a communication interface 418 coupled to bus 402. Communication interface 418 provides bidirectional data communication coupled to a network link 420 connected to a local network 422. For example, communication interface 418 can be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a telephone line of a corresponding type. As another example, communication interface 418 can be a local area network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link can also be implemented. In any such implementation, communication interface 418 sends and receives electrical signals, electromagnetic signals, or optical signals that carry digital data streams representing various types of information.
[0085] The network link 420 typically provides data communication through one or more networks to other data devices. For example, the network link 420 may provide a connection to a host computer 424 or data equipment operated by an Internet Service Provider (ISP) 426 through a local network 422. The ISP 426, in turn, provides data communication services through the global packet data communication network now commonly referred to as the "Internet" 428. Both the local network 422 and the Internet 428 use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on the network link 420 and through the communication interface 418 that carry the digital data to and from the computer system 400 are example forms of transmission media.
[0086] Computer system 400 can send messages and receive data, including program code, through network(s), network link 420 and communication interface 418. In the Internet example, server 430 can send the requested code for an application through Internet 428, ISP 426, local network 422 and communication interface 418.
[0087] The received code may be executed by processor 404 as it is received, and / or stored in storage device 410 or other non-volatile storage for later execution.
[0088] cloud computing
[0089] The term "cloud computing" is generally used herein to describe a computing model that enables on-demand access to a shared pool of computing resources (such as computer networks, servers, software applications, and services) and allows resources to be quickly provisioned and released with minimal management effort or service provider interaction.
[0090] A cloud computing environment (sometimes referred to as a cloud environment or just a cloud) can be implemented in a number of different ways to best suit different needs. For example, in a public cloud environment, the underlying computing infrastructure is owned by an organization that makes its cloud services available to other organizations or the general public. In contrast, a private cloud environment is typically intended for use only by or within a single organization. A community cloud is intended to be shared by several organizations within a community; while a hybrid cloud includes two or more types of clouds (e.g., private, community, or public) bound together by data and application portability.
[0091] In general, the cloud computing model enables some of those responsibilities that may have previously been provided by an organization's own information technology department to be delivered instead as a service layer within the cloud environment for use by consumers (either within the organization or externally, depending on the public / private nature of the cloud). The precise definition of the components or features provided by or within each cloud service layer may vary depending on the specific implementation, but common examples include: Software as a Service (SaaS), in which consumers use software applications running on the cloud infrastructure, while the SaaS provider manages or controls the underlying cloud infrastructure and applications. Platform as a Service (PaaS), in which consumers can develop, deploy, and otherwise control their own applications using software programming languages and development tools supported by the PaaS provider, while the PaaS provider manages or controls other aspects of the cloud environment (i.e., everything below the runtime execution environment). Infrastructure as a Service (IaaS), in which consumers can deploy and run arbitrary software applications, and / or provide processing, storage, networking, and other underlying computing resources, while the IaaS provider manages or controls the underlying physical cloud infrastructure (i.e., everything below the operating system layer). Database as a Service (DBaaS), in which the consumer uses a database server or database management system running on a cloud infrastructure, while the DbaaS provider manages or controls the underlying cloud infrastructure, applications, and servers, including one or more database servers.
[0092] In the foregoing specification, embodiments of the present invention have been described with reference to many specific details that may vary from implementation to implementation. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive. The sole and exclusive indicator of the scope of the invention and what the applicant intends to be the scope of the invention is the literal and equivalent range of the set of claims issuing from this application in the specific form in which the claims issue, including any subsequent corrections.
Claims
1. A method for remote access of a volatile memory, comprising: Implemented on a first computing device including a local volatile memory: a first execution candidate capable of executing the operation on the first computing device; as well as a second execution candidate on the first computing device capable of performing the operation; wherein the first execution candidate and the second execution candidate are able to directly access the local volatile memory of the first computing device; wherein the operation requires access to data in the local volatile memory; receiving a request to perform the operation from a requesting entity executing on a second computing device that (a) is remote from the first computing device and (b) does not have direct access to the local volatile memory; In response to the request, attempting to perform the operation using the first execution candidate; determining that the first execution candidate did not successfully execute the operation; as well as In response to determining that the first execution candidate did not successfully execute the operation, attempting to execute the operation using the second execution candidate.
2. The method of claim 1, wherein attempting to use the second execution candidate to perform the operation comprises attempting to use the second execution candidate to perform the operation without notifying the requesting entity that the first execution candidate failed to perform the operation.
3. The method of claim 1 , wherein attempting to use the second execution candidate to perform the operation comprises: Notifying the requesting entity that the first execution candidate fails to execute the operation; receiving, from the requesting entity, a second request to perform the operation; as well as In response to the second request, attempting to perform the operation using the second execution candidate.
4. The method of claim 1, wherein the operations involve reading and / or writing data on a persistent storage device that is directly accessible to the first computing device and not directly accessible to the second computing device.
5. The method of claim 1, wherein: One of the first execution candidate and the second execution candidate is an application running on one or more processors of the first computing device; as well as The other of the first execution candidate and the second execution candidate is implemented in a network interface controller of the first computing device.
6. The method of claim 5, wherein the another execution candidate is implemented in firmware of the network interface controller.
7. The method of claim 5, wherein the another execution candidate is software executing on one or more processors within the network interface controller.
8. The method of claim 5, wherein the another execution candidate is an interpreter that performs the operation by interpreting instructions specified in data provided to the network interface controller.
9. The method of claim 1, wherein one of the first execution candidate and the second execution candidate is implemented within an operating system executing on the first computing device.
10. The method of claim 1, wherein the second execution candidate is associated with a different reliability domain than the first execution candidate, and wherein one of the first execution candidate and the second execution candidate is implemented within a privileged domain on the first computing device.
11. The method of claim 1, wherein: The first execution candidate is executed on a first set of one or more cores of a processor in the first computing device; The second execution candidate is executed on a second set of one or more cores of the processor in the first computing device; as well as The members of the second group of one or more cores are different from the members of the first group of one or more cores.
12. The method of claim 11, wherein the second set of one or more cores includes at least one core that is not part of the first set of one or more cores.
13. The method of claim 1, wherein one of the first execution candidate and the second execution candidate comprises an interpreter that interprets instructions that, when interpreted, cause the operation to be performed.
14. The method of claim 13, wherein the interpreter is implemented on a network interface controller, and the first computing device communicates with the second computing device via a network through the network interface controller.
15. A method for remote access of volatile memory, comprising: Executed on a first computing device including local volatile memory: a first execution candidate capable of performing the operation; as well as a second execution candidate capable of performing the operation; wherein the first execution candidate and the second execution candidate have access to the local volatile memory of the first computing device; wherein the second execution candidate is associated with a different reliability domain than the first execution candidate; as well as receiving a request to perform the operation from a requesting entity executing on a second computing device remote from the first computing device; In response to the request, the first execution candidate and the second execution candidate are simultaneously requested to perform the operation.
16. The method of claim 15, wherein the operation requires multiple accesses to data in the local volatile memory.
17. The method of claim 15, wherein: The first execution candidate is implemented on a network interface controller through which the remote computing device communicates via a network; and The second execution candidate is executed by one or more processors of the remote computing device.
18. A method for remote access of a volatile memory, the method comprising: executing an application capable of performing operations on a first computing device including a local volatile memory; wherein the first computing device includes a network interface controller having an execution candidate capable of interpreting instructions that, when interpreted, cause the operations to be performed; wherein the application and the execution candidate have access to the local volatile memory of the first computing device; selecting, by a requesting entity executing on a second computing device remote from the first computing device, a target from among the application and the execution candidates based on one or more factors; In response to the application being selected as the target, the requesting entity sends a request to the application to perform the operation; as well as In response to the execution candidate being selected as the target, the requesting entity causes the execution candidate to interpret instructions that, when interpreted by the execution candidate, cause the operation to be performed.
19. The method of claim 18, wherein the operation requires multiple accesses to data in the local volatile memory.
20. The method of claim 18, further comprising sending data specifying the instruction from the second computing device to the first computing device, the instruction, when interpreted by the execution candidate, causing the operation to be performed.
21. The method of claim 18, wherein the execution candidate is an interpreter, and the interpreter is a JAVA virtual machine, and the instruction is a bytecode.
22. A method for remote access of volatile memory, comprising: Execute simultaneously on a specific computing device: a first execution candidate implemented in a first reliability domain on the particular computing device that is capable of performing the operation, a second execution candidate implemented in a second reliability domain on the particular computing device that is capable of performing the operation, and a third execution candidate implemented in a third reliability domain on the particular computing device capable of performing the operation; wherein the first reliability domain, the second reliability domain and the third reliability domain are different from each other; wherein the operation requires access to data in a shared memory associated with the particular computing device; selecting a target execution candidate from among the first execution candidate, the second execution candidate, and the third execution candidate based on one or more factors; and An attempt is made to perform the operation using the target execution candidate.
23. A method for remote access of volatile memory, comprising: Execute simultaneously on a specific computing device: a first execution candidate implemented in a first reliability domain on the particular computing device that is capable of performing the operation, and a second execution candidate implemented in a second reliability domain on the particular computing device that is capable of performing the operation; wherein the first reliability domain includes a first set of one or more cores of a processor of the particular computing device; wherein the second reliability domain includes a second group of one or more cores of the processor of the particular computing device; wherein the first group of one or more cores is different from the second group of one or more cores; wherein the operation requires access to data in a shared memory of the particular computing device; selecting a target execution candidate from among the first execution candidates and the second execution candidates; and The target execution candidate is caused to perform the operation.
24. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the method of any one of claims 1-23 to be performed.
25. A system for remote access of volatile memory, comprising: one or more computing devices; as well as One or more non-transitory computer-readable media storing instructions that, when executed by the one or more computing devices, cause the method of any one of claims 1-23 to be performed.
26. A computer program product comprising instructions which, when executed by one or more processors, cause the method of any one of claims 1 to 23 to be performed.
Citation Information
Patent Citations
Method for efficient primary key based queries using atomic RDMA reads on cache friendly in-memory hash index
US20180341653A1
Storage cluster failure detection
US20160132411A1
Single-sided distributed cache system
US9164702B1