Fault-tolerant system and method using a shared memory configuration - Patents.com
A shared memory configuration with cache-coherent switches enables rapid CPU and memory state failover in computing systems, addressing the challenge of downtime by flushing caches to shared memory, ensuring fast and efficient system recovery.
Patent Information
- Application Number
- JP2025531681
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-30
- Filing Date
- 2023-11-22
- Publication Date
- 2025-12-23
AI Technical Summary
Existing computing systems with high availability requirements face challenges in achieving fast and efficient failover in the event of performance degradation or failure, particularly due to the time-consuming process of copying memory states between nodes, which can lead to significant downtime and disruption.
A fault-tolerant system utilizing a shared memory configuration with cache-coherent switches, such as CXL or CCIX, allows for rapid failover by flushing the cache of a failing node to shared memory, enabling a standby node to take over without the need for memory copying, thereby reducing downtime to milliseconds.
The solution significantly reduces failover time to 1-800 milliseconds, minimizing application interruption and improving system availability by eliminating the need for memory copying, thus meeting strict real-time responsiveness requirements.
Smart Images

Figure 2025541742000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to and benefit of U.S. Patent Application No. 18 / 072,297, filed November 30, 2022, the disclosure of which is incorporated by reference in its entirety.
[0002] (Field) The present disclosure relates generally to the field of fault tolerance in computing systems. [Background technology]
[0003] Modern computing systems with high availability requirements make effective use of resource redundancy and failover mechanisms for a variety of purposes. Summary of the Invention [Means for solving the problem]
[0004] In part, the present disclosure relates to a fault-tolerant system including one or more shared memory complexes, each memory complex comprising a group of M computer-readable memory storage devices, one or more cache coherent switches comprising two or more host ports and one or more downstream device ports, the cache coherent switches in electrical communication with the one or more shared memory storage devices, a first management processor in electrical communication with the cache coherent switch, and a first compute node comprising a first processor and a first cache, the first management processor being connected to the one or more cache coherent switches and one or more downstream device ports. The system may include a first computing node in electrical communication with one or more shared memory complexes, and a second computing node having a second processor and a second cache, the second computing node in electrical communication with the one or more cache coherent switches and the one or more shared memory complexes, whereby data stored by the first computing node in the one or more shared memory complexes and modified due to operations executing on the first computing node is available for use and modification by the second computing node in substantially real time when the second computing node takes over for the first computing node in response to the first computing node experiencing a performance degradation event.
[0005] In some embodiments, the one or more cache coherent switches are one or more CXL switches. In various embodiments, the one or more cache coherent switches are one or more Cache Coherent Interconnect for Accelerators (CCIX) switches. In many embodiments, each of the M computer-readable memory storage devices is a DDR5 RAM module. In some embodiments, M is an integer greater than or equal to 1. In some embodiments, M is an even integer greater than or equal to 2. In some embodiments, the second compute node takes over as the active node in response to the first node experiencing a performance degradation event during a failover time ranging from about 1 millisecond to about 800 milliseconds. In various embodiments, the one or more CXL switches are in electrical communication with one or more secondary devices selected from the group consisting of storage devices, I / O devices, and accelerators. In some embodiments, the system may further include an interconnect comprising one or more front-end interconnects and one or more back-end interconnects, wherein the one or more CXL switches are in electrical communication with the one or more back-end interconnects, and the first computing node and the second computing node are in electrical communication with the one or more front-end interconnects.
[0006] In some embodiments, the first cache comprises state information, and the first cache is configured to flush the state information to one or more of the M computer-readable memory storage devices in response to detecting a performance degradation event related to the first compute node. In many embodiments, the one or more shared memory complexes are protected by one or more RAS features, such as hardware, software, or firmware systems, implemented within the one or more cache coherent switches for memory recovery or error correction. In some embodiments, the state information is accessible by the second compute node.
[0007] In various embodiments, the first compute node is running an operating system and one or more customer applications. In many embodiments, data associated with the one or more customer applications is accessible by the first compute node and the second compute node from one or more shared memory complexes. In some embodiments, the second compute node takes over for the first compute node and continues to execute the one or more customer applications and modify the data associated with the one or more customer applications. In various embodiments, the first compute node is configured to create a non-transparent bridging (NTB) window between the local memory of the first compute node and the local memory of the second compute node.
[0008] The present disclosure relates, in part, to a method for reducing recovery time in a fault-tolerant system, the method including providing a shared memory complex, a cache coherent switch in electrical communication with the shared memory complex, a first CPU node, a second CPU node, and a primary management processor, requesting permission for the first CPU node to fail over to the second CPU node in response to an occurrence of a performance degradation event, signaling by the primary management processor that the second CPU node is acting as a standby node and is available to take over for the first CPU node, messaging between the first CPU node and the second CPU node to transfer or support state transfer from the failing first CPU node to the standby second CPU node, and flushing one or more caches of the first CPU node to the shared memory.
[0009] In some embodiments, the method may further include suspending direct memory access traffic from the IO device to the first CPU node. In various embodiments, the IO device is selected from the group consisting of an I / O, a storage device, and an accelerator. In some embodiments, the method may further include avoiding copying local memory from the failing first CPU node to the second CPU node, the second CPU node transitioning to become the active node. In many embodiments, permission is requested from a primary management processor. In various embodiments, the method may further include the second CPU node taking over for the first CPU node in response to the first CPU node experiencing a performance degradation event during a failover time ranging from about 1 millisecond to about 800 milliseconds.
[0010] While the present disclosure relates to different aspects and embodiments, it should be understood that the different aspects and embodiments disclosed herein may be integrated, combined, or used together, as combined systems, or in part, as separate components, devices, and systems, as appropriate. Thus, each embodiment disclosed herein may incorporate each of the aspects to varying degrees, as appropriate for a given implementation. Additionally, the various CPU nodes, PCIe devices, PCIe switches, Compute Express Link (CXL) devices, CXL switches, various complexes, CXL complexes, Cache Coherent Interconnect for Accelerators (CCIX) devices, memory complexes, random access memory, non-volatile flash memory, persistent memory, memory devices, accelerators, RAS (reliability, availability, and serviceability) systems, methods for memory protection and error correction, software, and hardware, backplanes, midplanes, interconnects, data paths, I / O devices, caches, management CPUs, bridges, buses, network devices, interfaces, NVMe devices, disks, and portions of the foregoing disclosed herein can be used and shared with each other in various combinations and in any other devices and systems, without limitation.
[0011] These and other features of applicants' teachings are set forth herein. [Brief explanation of the drawings]
[0012] Unless otherwise specified, the accompanying drawings illustrate aspects of the innovations described herein. Referring to the drawings, in which like numbers refer to the several views and like parts throughout this specification, several embodiments of the presently disclosed principles are illustrated by way of example, and not by way of limitation. The drawings are not intended to be to scale. A more complete understanding of the present disclosure may be realized by reference to the accompanying drawings.
[0013] [Figure 1] FIG. 1 is a high-level block diagram of a system implementing a shared memory complex accessible by multiple CPUs via a cache-coherent switch, such as a Compute Express Link (CXL) switch, where the switch enables rapid CPU and memory state failover, in accordance with an exemplary embodiment of the present disclosure.
[0014] [Figure 2] FIG. 2 is a high-level block diagram of a system implementing two memory complexes, each accessible by multiple CPUs via a cache coherent switch, such as a CXL switch, in accordance with an exemplary embodiment of the present disclosure.
[0015] [Figure 3] FIG. 3 is a high-level block diagram of a system implementing a complex of cache coherent switches, such as a CXL switch, for intermediating memory complexes, storage devices, I / O devices, and accelerators, according to an exemplary embodiment of the present disclosure.
[0016] [Figure 4] FIG. 4 is a flowchart depicting one type of failover process according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0017] (Detailed explanation) In part, this disclosure relates to a fault-tolerant computer system that implements CPU and memory state failover using a shared memory configuration. In one aspect, this disclosure relates to a shared memory architecture that uses a cache-coherent interconnect, switch, or fabric configured to electrically couple or connect processors, accelerators, network interface devices, and storage devices to one another, and / or a shared collection or complex of memory devices, such as an array of random access memory or other memory modules, as needed for a given use case.
[0018] Furthermore, systems with expensive components such as accelerators are typically not replicated, and migrating memory state for an accelerator requires significant time and consumes power. Additionally, even without an accelerator, migrating memory state for a conventional processor can itself require significant time and power consumption.
[0019] In some embodiments of the CPU and memory state failover methods disclosed herein, multiple CPU nodes share a memory complex so that when one CPU node experiences a correctable error, warning / failure indicator, or other performance-degrading event, the erroneous CPU node can be replaced with a new CPU node connected to the same shared memory. The memory state of the replaced CPU node already present in the shared memory does not need to be copied to the new memory complex. In various embodiments, the memory can be individual memory devices, pairs of memory devices, or other combinations of memory devices. In some embodiments, memory interleaving is supported.
[0020] In various embodiments disclosed herein, a fault-tolerant computer system uses at least one cache coherent device or switch, such as, for example, a Compute Express Link (CXL) switch or a Cache Coherent Interconnect for Accelerators (CCIX) switch, to manage a low-latency memory complex shared among CPU nodes. In some embodiments, low latency may refer to a delay or latency ranging from about 10 to about 100 nanoseconds. In other embodiments, low latency may refer to a delay ranging from about 4 to about 500 nanoseconds. Managing a shared low-latency memory complex using a CXL switch can include memory writes, memory reads, resolving memory address conflicts between compute nodes using the memory complex as their primary random access memory, and modifying which compute nodes may access memory modules within the memory complex.
[0021] In various embodiments, the memory complex facilitates the exchange of data between cooperating CPU nodes, providing a very fast CPU failover solution with N+1 redundancy protection against CPU failures. In various embodiments, CPU failover times are very fast, meaning that failover times can range from about 1 millisecond to about 800 milliseconds. Such a failover solution can improve application availability and eliminate the time required to copy one failing CPU's memory to a replacement CPU, which may be required in other failover strategies. The memory copy step can be avoided by the replacement CPU using the failing CPU's memory by directly using the contents of the memory module previously written by the failing CPU. In various embodiments, caches are flushed back to shared memory. In some embodiments, architectural internal registers that are part of the running operating system and application state are migrated off the processor. One or more save / restore functions available for the processor can be used in various embodiments to save state from the original active node, copy the state to the standby node, and / or restore the state on the standby node that is about to become the new active node.
[0022] For example, to recover from a CPU or other hardware failure, some applications may tolerate an application restart, with the loss of some data that was resident in memory at the time of the crash or other hardware or application failure. In such cases, an application restart may be part of a cost-effective failover solution. However, an application restart would incur significant time costs to reload the application data from storage to a replacement processor, accelerator, or memory complex, or to an available spare CPU node to replace the failed CPU node. The use of a shared memory complex and CXL switch dramatically improves upon traditional approaches by significantly reducing the time it takes to restart an application after an outage.
[0023] Other failover solutions in the event of a CPU failure migrate CPU state from one physical node to another and may support both "bare metal" and virtualized operating systems. Such solutions may require substantial time to copy memory and state from one CPU node to another, as well as a non-negligible period during which all applications are paused. The time during which applications are paused may be unacceptable for certain applications with strict sub-second real-time responsiveness requirements. As described herein, the use of a shared memory complex and CXL switch dramatically improves fault-tolerant systems that rely on pausing applications and have the time delays otherwise associated with memory copies.
[0024] In most embodiments of the present disclosure, the execution state of the active node, also referred to herein as the active CPU node or active computer node, is substantially stored within the shared memory complex. Therefore, if this node fails, the time required to copy the failed node's state to a new node is eliminated. The failing node simply flushes its cache so that the node's state is contained entirely within the shared memory complex. After the failed node's state is available in the shared memory, the failing node can take the final steps to transfer its role as the active node to a replacement node.
[0025] In some embodiments, the system is configured for selective use of unused or otherwise idle components, components on standby, or components available as replacements for failed components, which may be selectively controlled and used, or rested or otherwise taken offline, as applicable to a given application or failure or error event.
[0026] In some embodiments, by decoupling the processor and memory subsystem, various fault-tolerant system embodiments of the present disclosure enable / allow for faster power down of idle processors. In some embodiments, memory may be self-refreshed (known as the S3 state) before powering down the processor. In most embodiments, memory is not available to other processors when in the self-refresh state.
[0027] Finally, in most embodiments of the present disclosure, the failover architecture allows CPU and I / O nodes (such as hot-pluggable / swappable assemblies containing one or more I / O devices) to be independently replaced as an example of a serviceable node. This modularity, allowing components to be independently swapped or replaced, improves serviceability over other hyperconverged infrastructure (HCI) solutions, i.e., solutions that embody a seamless combination of compute, storage, networking, and data services within a single physical system running virtualized or containerized workloads.
[0028] Referring now to the exemplary embodiment of FIG. 1 , a fault-tolerant computer system includes N CPU nodes 110 connected via interconnect 115 to a single Compute Express Link (CXL) switch 130 that multiplexes multiple memory modules 120. In many embodiments, various data paths 112 may be used to electrically connect or otherwise couple components so that the components are in electrical communication. Data paths 112 may include, but are not limited to, buses, traces, cables, conductive pathways, electrical connections, optical connections, and other pathways suitable for transmitting data and / or electrical signals. In various embodiments, the various components depicted in the diagram may be connected by one or more data paths indicated by solid lines and / or lines with arrows.
[0029] In some embodiments, interconnect 115 is a midplane or midplane interconnect. In various embodiments, interconnect 115 includes boards that plug in from both sides, a backplane with boards that plug in from only one side, or one or more devices that support a set of cables or data paths. The first CPU node, CPU node 110A, is shown, and the last of the CPU nodes is CPU node N 110B. In various embodiments, N is a positive integer greater than or equal to 2.
[0030] In various embodiments, a given node of the N CPU nodes, such as node 110A and / or 110B, may optionally include a local memory, such as local memory 116A, 116B shown in Figure 1. In various embodiments, a non-transparent bridging (NTB) window or other memory transfer mechanism between local memory 116A of a first CPU node and local memory 116B of a second CPU node may be created from either node 110A or node 110B. The various computer nodes described herein may optionally have one or more local memories.
[0031] Interconnect 115 has a front end connected to CPU nodes 110A, 110B and a back end connected to CXL switch 130. Memory modules 120 may include RAM memory, computer-readable memory, double data rate memory modules such as double data rate synchronous dynamic random access memory, persistent memory, and other high-speed, low-power consumption memory. In various embodiments, any suitable memory standard suitable for the application and other components in the fault-tolerant system may be used within the memory modules. In various embodiments, a group of memory modules, such as the eight shown in FIG. 1, may be treated together as a group or series of subgroups and classified as a memory complex. In some embodiments, one or more memory complexes are used in conjunction with one or more CXL switches.
[0032] In some embodiments, the CXL switch may also connect storage devices 140, I / O devices 142, or optional accelerators 144 to the CPU nodes. In various embodiments, the systems and methods disclosed herein provide support for accelerators 144 connected to one or more CXL switches in addition to memory. Multiple CPU nodes 110A are in an active state, while at least one remaining node 110B is in a standby state. In addition to the shared memory module 120, each CPU node has a memory cache 111A, 111B in which a portion of the CPU node's state may be stored. Finally, the CXL switch 130 may include or communicate with a management CPU (MCPU) / management processor that may execute firmware that coordinates and assists with failover functions. In various embodiments, the CXL switch 130 is one of multiple CXL switches that may be connected in various configurations. In some embodiments of a cache-coherent switch, the management processor may be external to the switch. In various embodiments, the management processor is a functional unit within the switch hardware itself.
[0033] In various embodiments, collections of components or individual components may be replaceable like modular components to support serviceability. These various components may be integrated on a shared board or bus and configured to be removed as a collection, or have individual connections for interfacing with a given system. In some embodiments, such components or collections of components may be referred to as serviceable nodes or serviceable components. A given serviceable node may also include individual serviceable components in some embodiments. As an example, serviceable node 145 includes some, all, or a subset of the components shown below midplane 115 in FIG. 1. A given serviceable node may be replaceable, swappable, hot-swappable, hot-pluggable, and configurable. Another example of a serviceable node is shown in FIG. 2 by serviceable nodes 225a and 225b. In some embodiments, various memories, individual memory modules, or combinations of memory modules may be serviceable nodes, such as serviceable node 221 in FIG. 2.
[0034] In many embodiments, an active CPU node, such as CPU node 111A, experiencing a correctable error, failure indicator, or other performance-degrading event that could lead to a system outage can be quickly failed over to a healthy standby node, such as CPU node N 110B, with minimal application interruption and no user intervention. The faulty CPU node can be isolated, undergo diagnostics to determine its health, and either become the new standby node or remain isolated and marked for replacement.
[0035] In various embodiments, to replace a failed active CPU node, a replacement standby CPU node inherits the state of the failed active CPU node. The state of the failed active node consists essentially of memory accessible by that node, its cache 111A, and primary memory 120. In most embodiments, memory module 120 is accessible by both the failed active node 110A and the replacement node 110B via CXL switch 130; therefore, no additional time or process is required to transfer memory 120 between the failed active node and the replacement node. However, in many embodiments, the cache 111A of the failed active node contains a portion of the failed active node's state. Therefore, to complete the replacement of the failed node, the cache of the failed active node is first flushed to memory module 120, which is part of the shared memory accessed by CXL. As a result, the flushed cache information or state is immediately available to the standby or replacement node 110B.
[0036] 2, a fault-tolerant computer system includes N CPU nodes 210 connected to two CXL switches 230A, 230B via an interconnect 215. The two CXL switches 230A, 230B provide a separate collection of memory modules 220A, 220B to each CPU node via the interconnect 215. The interconnect 215 has a front end connected to the CPU nodes 210A, 210B and a back end connected to the switches 230A, 230B.
[0037] In some embodiments, each CXL switch contains its own MCPU 235A, 235B. In embodiments containing multiple CXL switches, one MCPU will be selected by software supporting a failover solution as the primary MCPU, i.e., the central point for communications between the active and standby compute nodes. In some embodiments, multiple CXL switches are used to support modular interchangeability during compute mode operation in the event that one of the CXL switches fails. In some embodiments, the memory contents of memory module 220A may be mirrored to memory module 220B via two switches 230A, 230B.
[0038] Referring now to the alternative embodiment of FIG. 3 , a fault-tolerant computer system includes N CPU nodes 310 connected to a CXL switching complex 320 via interconnect 315. The CXL switching complex 320 contains multiple CXL switches 330A, 330B accessible by multiple active CPU nodes and at least one standby CPU node. The switching complex 320 is also in electrical communication with a group of memory modules 350. The systems depicted and described in FIGS. 1 , 2 , and 3 are exemplary embodiments through which and in which the failover solutions disclosed herein may be implemented. The systems of FIGS. 1 and 2 may also, in various embodiments, use the CXL switching complex 320 in place of individual CXL switches. In some embodiments, the CXL switching complex 320 may also connect storage devices 340, I / O devices 342, or optional accelerators 344 to the CPU nodes 310.
[0039] In most embodiments, a compute node or CPU node in a fault-tolerant computer system is a compute complex containing an x86 or other architecture processor and an integrated or discrete platform controller with supporting logic including voltage regulators, clocks, etc.
[0040] In some embodiments, a fault-tolerant computer system of CPU nodes connected to at least one memory complex accessible through a CXL switch is equipped with a PCIe switching fabric. In some embodiments, the PCIe switching fabric may be used for I / O devices such as network controllers, storage devices, etc. The various I / O devices may or may not be directly connected to the cache-coherent CXL switch.
[0041] In many embodiments, a management CPU (MCPU) may coordinate or assist in various aspects of the failover, i.e., the swapping of the failing active CPU and the standby CPU. In one embodiment, the MCPU is a PCIe fabric switch processing agent capable of snooping, intercepting, redirecting, and merging PCIe transactions within the switch domain. In many embodiments, during a failover, the fault-tolerant computer system may enter system management mode (SMM) and suspend operating system control as the failing active CPU is replaced by the standby CPU.
[0042] In most embodiments, the original driver software for various devices in a fault-tolerant computer system can be modified to support operations that effectively pause and then resume I / O operations. Such pause and resume functionality provides a window during which I / O resources can be switched from an active node to a standby node, with the final memory copy being performed without further modification.
[0043] In some embodiments, driver software modifications may vary depending on the driver software base, operating system capabilities, or device capabilities. In some embodiments, driver software modifications may include handling of new input / output controls (IOCTLs) or incorporate existing device management capabilities. Driver software modifications may allow "in-flight" transactions to complete and prevent further operation, or allow hibernating devices to resume operation. In some embodiments, the driver will leverage processor or switch features such as downstream port containment (DPC) to signal the detection of an error according to industry standards and isolate the erroneous IO device from the rest of the system. Finally, the driver software supports dirty page tracking for all I / O DMA operations, although this functionality may not be required if all memory allocated for I / O device DMA is within a shared memory space.
[0044] In various embodiments, a CPU node in a standby state may run a limited but suitably robust operating environment, including, for example, a BIOS or a minimal Unified Extensible Firmware Interface (UEFI) (an API of low-level routines available for execution in a privileged context), a software stack, or a minimal Linux operating system. In one embodiment, this limited operating environment may provide local CPU diagnostics and error handling facilities.
[0045] In other embodiments, the limited operating environment may also include a synthetic PCI hierarchy (PCIe fabric switch technology that allows multiple hosts connected to the same switch to see a mutually exclusive set of PCI functions) provided by the PCIe fabric MCPU. In other embodiments, the limited operating environment may handle the creation and management of an NTB window space to enable inter-CPU communication and processing of commands passed through the NTB window. In yet other embodiments, the limited operating environment of a standby CPU node may provide ACPI management services, system management interrupt (SMI) handling, and system management mode (SMM) state management.
[0046] In some embodiments, the standby node may be hibernated, placed in an idle mode, turned off, and / or replaced. In other embodiments, the limited operating environment of the standby CPU node may utilize PCI management functions and / or hardware to hibernate the standby CPU, bring it out of hibernation, and resume operation as an active CPU. In most embodiments, the CPU node in the active state runs an OS and applications in either a virtualized or non-virtualized environment.
[0047] In many embodiments, such as those of Figure 1, 2, or 3, if a failure, the onset of a failure, or a prediction of a failure occurs in an active CPU node, referred to as the failing active CPU node, a CPU node in a standby state, referred to as the replacement standby CPU node, may transition to an active state and replace the failing active CPU node. This transition from the failing CPU to the standby CPU typically involves the failing CPU node flushing any cached memory locations back to shared physical memory, allowing access by the new active node even when the failing node is reset, powered down, or removed for maintenance operations.
[0048] In various embodiments, flushing the failing node's cache to the shared memory module or complex makes the failing node's existing processor state available to the standby node that will take over, using state information and other data stored in the shared memory. Various embodiments described herein avoid the use of software to track memory pages modified by the operating system or customer applications during memory copy operations (referred to as brownout copies to indicate that the application is still running, but performance may be affected by memory copies consuming some memory and I / O bandwidth), such as through dirty memory page tracking through the use of shared memory and CXL switches. In various embodiments, the cache will be updated as a natural result of reading memory. No special action needs to be taken on the standby node with respect to caching data from the shared memory. In some embodiments, the system resets or clears the cache before connecting the standby node to the shared memory. In various embodiments, one or more steps are performed to remove or replace any stale data in the cache.
[0049] Referring now to the exemplary embodiment of FIG. 4, a flowchart depicting the failover process is illustrated. Initially, an active CPU experiencing a correctable error or reduced capacity or other performance degradation event or indicator may request permission for a failover from the primary MCPU (405). Typically, at least one standby node performing diagnostics may be in a ready state to take over the role of the failing active CPU. The standby node may also be powered down, hibernated, suspended, or maintained in an offline state and then brought up or powered on and brought into a ready state to take over the role of the failing active CPU.
[0050] In various embodiments, the state of the failing active node may be reflected in shared memory multiplexed by the CXL switch, in various caches local to the active CPU, or in various registers local to the active CPU. In most embodiments, replicating the state of the failing active CPU to the replacement standby CPU may involve flushing the local CPU cache to shared memory and flushing registers local to the replacement standby CPU by saving internal registers from the failing active node, copying them to an area on the standby CPU or shared memory, and then restoring them to internal registers on the replacement CPU node. In some embodiments, flushing the cache of the failing active CPU node and subsequently swapping the roles of the failing active node and the replacement standby node generates a “blackout” period for application execution. In these embodiments, during the blackout, all system workloads are paused while the final step of state copying is completed.
[0051] In one embodiment, after a failover request occurs, a standby node ready to take over the role of the failing active CPU is signaled by the MCPU (410), and in response, a non-transparent bridging (NTB) window (a PCIe technology that creates a point-to-point connection between two memory systems through PCI memory-mapped I / O (MMIO) space) activates in its memory and begins polling the failing active node for commands to replicate the state of the failing active CPU to the replacement standby CPU. The replacement standby node signals its state to the failing active node, which activates a data path for a direct memory access DMA memory copy (415). In various embodiments, the memory copy includes a copy of internal registers. The internal register copy may be performed by DMA or by memory-mapped IO transactions (MMIO), which may be faster than setting up a DMA engine if the number of registers to be copied is small. In some embodiments, the system switches the main paths from IO devices and directs them to the replacement CPU node. In some embodiments, in its current state, the replacement standby CPU is unable to "initiate" read or write access to the memory of the failing active CPU.
[0052] The failing active CPU signals its hardened drivers to pause all DMA traffic (420), indicating the start of the blackout phase. All CPU threads are aggregated, paused, isolated, and / or fenced off to prevent further dirtying of cacheable memory pages; therefore, the most recent caches can be flushed. All processor internal caches on the failing active node are then flushed back to shared memory. Any memory written into shared memory by applications running on the failing active CPU remains in shared memory for use by the standby CPU node, which is about to take over the role of the active CPU node.
[0053] Upon completion of the cache flush, the failing active CPU may copy its state to the replacement standby CPU through a general-purpose register (GPR) protocol (425). In some embodiments, a reserved area of shared memory is used for the copy operation. In some embodiments, the copied CPU state may include model-specific registers (MSRs), local highly programmable interrupt controller (APIC) state data, high-precision event timer (HPET) state data, or other state data. The state is copied to corresponding registers on the standby CPU.
[0054] The failing active CPU then sets 430 a token in its own NTB window and the NTB window of the replacement standby CPU so that both nodes know their intended new state after the failover operation. At any point throughout this process, the failover can be aborted and operation will simply continue based on the original active CPU. In some embodiments, the token is an instruction, flag, state, or action that implements a role identification function so that each node knows its role at different steps in the failover process and allows the system to change or revert role assignments or operations if the system needs to abort the failover.
[0055] To complete the failover, if all steps up to this point have been completed successfully, the failing active CPU sends a command to the primary MCPU to swap all resource mappings between host ports for the two CPUs involved in the failover operation (435). In this step, the primary MCPU coordinates with any secondary MCPUs to handle the error case.
[0056] The primary MCPU signals the failing active CPU and the replacement standby CPU when the switch reconfiguration is complete (440). In some embodiments, the switch reconfiguration may include changing the routing of any devices in communication with the PCIe and CXL switches to connect them to the replacement node rather than the failing active node. In many embodiments, the reconfiguration operates to redirect all, substantially all, or some traffic to the host ports for the new active CPU node. Both CPUs then read a token from their GPR indicating their new respective states (445). Here, the CPU designated the "failing active" CPU is actually in standby, while the CPU designated the "replacement standby" CPU is in active, with the latter CPU assuming the former's role.
[0057] Finally, the software performs any final cleanup and post-processing required (450). Post-processing may include, for example, replaying the enumeration cycle and training the CXL switch how to map transactions from the new host. The software also restores register contents from the in-memory stack and executes a resume from system management mode (RSM) instruction to return control to the operating system. Drivers that were signaled to quiesce their devices before the failover are now signaled to resume operation. Customer applications then resume (455).
[0058] While several aspects and embodiments of the present technology have been described, it should be understood that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. Accordingly, it should be understood that the foregoing embodiments are presented by way of example only, and that, within the scope of the appended claims and their equivalents, the embodiments of the present invention may be practiced otherwise than as specifically described. In addition, any combination of two or more features, systems, articles, materials, and / or methods described herein is included within the scope of the present disclosure, provided that such features, systems, articles, materials, and / or methods are not mutually inconsistent.
[0059] In most embodiments, the processors may be physical or virtual processors, while in other embodiments, the virtual processors may be distributed across one or more portions of one or more physical processors.
[0060] Also, as described, some aspects may be embodied as one or more methods. The acts performed as part of a method may be ordered in any suitable manner. Thus, embodiments may be constructed in which acts are performed in an order different from that shown, and may include performing some acts simultaneously even though shown as sequential acts in the illustrative embodiments.
[0061] As used herein, the term "and / or," as used in the specification and in the claims, should be understood to mean "either or both" of the elements so conjoined, i.e., elements that are present conjunctively in some cases and disjunctively in other cases.
[0062] As used herein in the specification and claims, the phrase "at least one," referring to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but does not necessarily include at least one of each and every element specifically listed in the list of elements, and does not exclude any combinations of elements in the list of elements. This definition also allows for elements other than the specifically identified elements in the list of elements to which the phrase "at least one" refers, whether related or unrelated to those specifically identified elements, may optionally be present.
[0063] The terms "approximately" and "about" can be used to mean, in some embodiments, within ±20% of a target value, in some embodiments, within ±10% of a target value, in some embodiments, within ±5% of a target value, and even in some embodiments, within ±2% of a target value. The terms "approximately" and "about" can include the target value.
[0064] In the claims and the above specification, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like, are to be understood to be open-ended, i.e., meaning "including but not limited to." The transitional phrases "consisting of" and "consisting essentially of" are intended to be closed or semi-closed transitional phrases, respectively.
[0065] When a range or list of values is provided, each intervening value between the upper and lower limits of that range or list of values is individually contemplated and encompassed within the disclosure, just as if each value were specifically recited herein. In addition, smaller ranges between and including the upper and lower limits of a given range are also contemplated and encompassed within the disclosure. The listing of example values or ranges does not negate other values or ranges between and including the upper and lower limits of a given range.
[0066] The use of headings and sections in this application is not intended to limit the disclosure, and each section may apply to any aspect, embodiment, or feature of the disclosure. Only those claims using the words "means for" are intended to be interpreted under 35 U.S.C. 112, sixth paragraph. In the absence of "means for" recitation within a claim, such claim is not to be interpreted under 35 U.S.C. 112. No limitations herein are intended to be read into any claim unless such limitations are expressly included within the claim.
[0067] The embodiments disclosed herein may be embodied as a system, method, or computer program product. Accordingly, the embodiments may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects (which may all be referred to generally herein as a "circuit," "module," or "system"). Furthermore, the embodiments may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied thereon.
[0068] As will be apparent from the discussion below, unless specifically stated otherwise, throughout the description, terms such as "processing," or "computing," or "calculating," or "delaying," or "comparing," or "generating," or "determining," or "forwarding," or "deferring," or "committing," or "interrupting," or "handling," or "receipt" may be used. It should be understood that discussions utilizing terms such as "receiving," or "buffering," or "allocating," or "displaying," or "flagging," or Boolean logic or other set related operations or equivalents, refer to the actions and processes of a computer system or electronic device that manipulate and convert data represented as physical (electronic) quantities in the registers and memory of the computer system or electronic device into other data similarly represented as physical quantities in electronic memory or registers, or other such information storage, transmission, or display device.
[0069] The algorithms presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the above description. In addition, the present disclosure is not described with reference to any particular programming language, and various embodiments may, therefore, be implemented using a variety of programming languages.
[0070] Several implementations have been described. Nevertheless, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. For example, various forms of the flows shown above may be used, with steps rearranged, added, or removed. Accordingly, other implementations are within the scope of the following claims.
[0071] The examples presented herein are intended to illustrate potential and specific implementations of the present disclosure. The examples are intended primarily for purposes of illustration of the present disclosure for those skilled in the art. No particular aspect or aspects of the examples are necessarily intended to limit the scope of the present disclosure.
[0072] The figures and descriptions of the present disclosure have been simplified for purposes of clarity to illustrate elements relevant for a clear understanding of the present disclosure while excluding other elements, however, those skilled in the art will recognize that these types of intensive discussions do not promote a deeper understanding of the present disclosure, and therefore, more detailed descriptions of such elements are not provided herein.
[0073] The processes associated with the present embodiment may be executed by a programmable device such as a computer. Software or other sets of instructions that may be employed to cause the programmable device to execute the processes may be stored in any storage device, such as a computer system (non-volatile) memory, an optical disk, a magnetic tape, or a magnetic disk. Furthermore, some of the processes may be programmed when the computer system is manufactured or via a computer-readable memory medium.
[0074] It should also be understood that certain process aspects described herein may be implemented using instructions stored on a computer-readable memory medium(s) that instruct a computer or computer system to perform the process steps. Computer-readable media may include, for example, memory devices such as diskettes, compact discs of both the read-only and read / write variety, optical disk drives, and hard disk drives. Computer-readable media may also include memory storage devices, which may be physical, virtual, permanent, temporary, semi-permanent, and / or semi-temporary.
[0075] The computer systems and computer-based devices disclosed herein may include memory for storing certain software applications used in acquiring, processing, and communicating information. It should be understood that such memory may be internal or external with respect to the operation of the disclosed embodiments. Memory may also include any means for storing software, including a hard disk, optical disk, floppy disk, ROM (read-only memory), RAM (random access memory), PROM (programmable ROM), EEPROM (electrically erasable programmable read-only memory), and / or other computer-readable memory medium. In various embodiments, a "host," "engine," "loader," "filter," "platform," or "component" may include various computers or computer systems or any reasonable combination of software, firmware, and / or hardware.
[0076] In various embodiments of the present disclosure, a single component may be replaced by multiple components, and multiple components may be replaced by a single component, to perform a given function or functions. Such substitutions are within the scope of the present disclosure, except to the extent that such substitutions would not be operable to practice embodiments of the present disclosure. Any of the servers may be replaced, for example, by a “server farm” or other group of networked servers (e.g., a group of server blades) located and configured for collaborative functions. It should be understood that a server farm may serve to distribute workloads among the individual components of the farm and expedite the computational process by utilizing the collective and collaborative capabilities of multiple servers. Such server farms may employ load balancing software to perform tasks such as, for example, tracking demand for processing power from different machines, prioritizing and scheduling tasks based on network demand, and / or providing backup contingency in the event of component failure or reduced operability.
[0077] In general, it may be apparent to those skilled in the art that the various embodiments described herein, or components or parts thereof, may be implemented in many different embodiments of software, firmware, and / or hardware, or modules thereof. The software code or specialized control hardware used to implement some of the present embodiments is not a limitation of this disclosure. Programming languages for computer software and other computer-implemented instructions may be converted into machine language by a compiler or assembler before execution, and / or may be converted directly at run time by an interpreter.
[0078] Examples of assembly languages include ARM, MIPS, and x86; examples of high-level languages include Ada, BASIC, C, C++, C#, COBOL, Fortran, Java, Lisp, Pascal, Object Pascal; and examples of scripting languages include Bourne script, JavaScript, Python, Ruby, PHP, and Perl. Various embodiments may be employed, for example, in a Lotus Notes environment. Such software may be stored on any type of suitable computer-readable medium or media, such as, for example, a magnetic or optical storage medium. Therefore, the operation and behavior of embodiments will be described without specific reference to actual software code or specialized hardware components. This lack of specific reference is understandable, as it is clearly understood that one skilled in the art would be able to design software and control hardware to implement embodiments of the present disclosure based on the description herein with only reasonable effort and without undue experimentation.
[0079] Various embodiments of the systems and methods described herein may employ one or more electronic computer networks to transfer data or share resources and information that facilitate communication between different components. Such computer networks can be categorized according to the hardware and software technologies used to interconnect devices within the network.
[0080] Computer networks may be characterized based on the functional relationships between the elements or components of the network, such as active networking, client-server, or peer-to-peer functional architectures. Computer networks may be classified according to their network topology, such as bus networks, star networks, ring networks, mesh networks, star-bus networks, or hierarchical topology networks. Computer networks may also be classified based on the method employed for data communication, such as digital and analog networks.
[0081] Embodiments of the methods, systems, and tools described herein may employ internetworking to connect two or more distinct electronic computer networks or network segments through common routing technologies. The type of internetwork employed may depend on the management and / or involvement within the internetwork. Non-limiting examples of internetworks include intranets, extranets, and the Internet. Intranets and extranets may or may not have connections to the Internet. When connected to the Internet, an intranet or extranet may be protected using appropriate authentication technology or other security measures. As applied herein, an intranet can be a collection of networks employing Internet protocols, web browsers, and / or file transfer applications under the collective control of an administrative entity. Such an administrative entity may restrict access to the intranet, for example, to only authorized users or to another internal work of an organization or commercial entity.
[0082] Unless otherwise indicated, all numbers expressing length, width, depth, or other dimensions used in this specification and claims are to be understood as indicating both the exact value as set forth in all instances and as modified by the term "about." As used herein, the term "about" refers to a ±10% variation from the nominal value. Thus, unless otherwise indicated, the numerical parameters set forth in this specification and the appended claims are approximations that may vary depending on the desired properties sought to be obtained. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Any specific value may vary by 20%.
[0083] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The foregoing embodiments are therefore to be considered in all respects as illustrative rather than restrictive with respect to the disclosure set forth herein. The scope of the invention is, therefore, indicated by the appended claims, rather than the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein.
[0084] It will be understood by those skilled in the art that various modifications and changes can be made without departing from the scope of the described technology. Such modifications and changes are intended to fall within the scope of the described embodiments. It will also be understood by those skilled in the art that features included in one embodiment can be interchanged with other embodiments, and that one or more features from a depicted embodiment can be included with other depicted embodiments in any combination. For example, any of the various components described herein and / or depicted in the figures can be combined, interchanged, or excluded from other embodiments.
Claims
1. 1. A fault-tolerant computer system, comprising: one or more shared memory complexes, each memory complex comprising a group of M computer-readable memory storage devices; one or more cache coherent switches, each comprising two or more host ports and one or more downstream device ports, in electrical communication with the one or more shared memory storage devices; a first management processor in electrical communication with the cache coherent switch; a first compute node comprising a first processor and a first cache, the first compute node in electrical communication with the one or more cache coherent switches and the one or more shared memory complexes; a second compute node comprising a second processor and a second cache, the second compute node in electrical communication with the one or more cache coherent switches and the one or more shared memory complexes; Equipped with 10. A fault-tolerant computer system, wherein data stored by the first computing node within the one or more shared memory complexes, and modified due to operations executing on the first computing node, is available for use and modification by the second computing node in substantially real time when the second computing node takes over for the first computing node in response to the first computing node experiencing a performance degradation event.
2. The system of claim 1 , wherein the one or more cache coherent switches are one or more CXL switches.
3. 10. The system of claim 1, wherein the one or more cache coherent switches are one or more Cache Coherent Interconnect for Accelerators (CCIX) switches.
4. 2. The system of claim 1, wherein each of the M computer-readable memory storage devices is a DDR5 RAM module.
5. 10. The system of claim 1, wherein the second compute node takes over for the active node in response to the first node experiencing a performance degradation event during a failover time ranging from about 1 millisecond to about 800 milliseconds.
6. 3. The system of claim 2, wherein the one or more CXL switches are in electrical communication with one or more secondary devices selected from the group consisting of storage devices, I / O devices, and accelerators.
7. 3. The system of claim 2, further comprising an interconnect comprising one or more front-end interconnects and one or more back-end interconnects, wherein one or more CXL switches are in electrical communication with the one or more back-end interconnects, and wherein the first computing node and the second computing node are in electrical communication with the one or more front-end interconnects.
8. 3. The system of claim 2, wherein the first cache comprises state information, and wherein the first cache is configured to flush the state information to one or more of the M computer-readable memory storage devices in response to detecting a performance degradation event for the first compute node.
9. 3. The system of claim 2, wherein the one or more shared memory complexes are protected by one or more RAS features, such as hardware, software, or firmware systems, implemented within the one or more cache coherent switches for memory recovery or error correction.
10. The system of claim 6 , wherein state information is accessible by the second compute node.
11. The system of claim 2 , wherein the first compute node is running an operating system and one or more customer applications.
12. 12. The system of claim 11, wherein the data associated with the one or more customer applications is accessible by the first compute node and the second compute node from the one or more shared memory complexes.
13. 10. The system of claim 9, wherein the second compute node takes over for the first compute node and continues to run the one or more customer applications and modify the data associated with the one or more customer applications.
14. The system of claim 2 , wherein the first computing node is configured to create a non-transparent bridging (NTB) window between a local memory of the first computing node and a local memory of the second computing node.
15. 1. A method for reducing recovery time in a fault tolerant system, the method comprising: providing a shared memory complex; providing a cache coherent switch in electrical communication with the shared memory complex, a first CPU node, a second CPU node, and a primary management processor; In response to an occurrence of a performance degradation event, the first CPU node requests permission to fail over to the second CPU node; signaling, by the primary management processor, that the second CPU node is available to act as the standby node and take over for the first CPU node; messaging between a first CPU node and a second CPU node to transfer or support state transfer from the failing first CPU node to the standby second CPU node; flushing one or more caches of the first CPU node to the shared memory; A method comprising:
16. 16. The method of claim 15, further comprising suspending direct memory access traffic from IO devices to the first CPU node.
17. 17. The method of claim 16, wherein the IO device is selected from the group consisting of an I / O, a storage device, and an accelerator.
18. 16. The method of claim 15, further comprising avoiding copying local memory from the failing first CPU node to a second CPU node, the second CPU node transitioning to become an active node.
19. The method of claim 15 , wherein the authorization is requested from the primary management processor.
20. 16. The method of claim 15, further comprising: the second CPU node taking over for the first CPU node in response to the first CPU node experiencing a performance degradation event during a failover time ranging from about 1 millisecond to about 800 milliseconds.
Citation Information
Patent Citations
Loossely coupled computer system with common memory
JP1998240556A
Server device, server system, and control method for server device
JP2011253290A
Information processing device, shared memory management method, and shared memory management program
JP2017111751A
Highly reliable fault-tolerant computer architecture
JP2021534524A
Failover for pooled memory
US20220004468A1