Highly reliable fault-tolerant computer architecture

The fault-tolerant computer system addresses inefficiencies in redundant components by using a switching fabric for seamless failover, ensuring high reliability and low downtime with efficient state and memory transfer.

JP7794884B2Active Publication Date: 2026-01-06STRATUS TECH IRELAND LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024065406
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-08-13
Filing Date
2024-04-15
Publication Date
2026-01-06
Estimated Expiration
2039-08-09

AI Technical Summary

Technical Problem

Existing fault-tolerant computer systems face high costs and performance slowdowns due to redundant components and periodic checkpointing methods, which are inefficient in maintaining high reliability and uptime.

Method used

A fault-tolerant computer system with modular redundant components interconnected by a switching fabric, enabling seamless failover of CPU and IO domains using a Non-Transparent Bridge (NTB) PCI Express switching fabric, and management processors to transfer state and memory to standby nodes, ensuring minimal downtime.

Benefits of technology

The system achieves high reliability with reduced downtime and cost by efficiently transferring state and memory between active and standby nodes, maintaining system functionality with minimal performance impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794884000001
    Figure 0007794884000001
  • Figure 0007794884000002
    Figure 0007794884000002
  • Figure 0007794884000003
    Figure 0007794884000003
Patent Text Reader

Abstract

To provide a suitable high reliability fault tolerant computer architecture.SOLUTION: A fault tolerant computer system and method are disclosed. The system may include: a plurality of CPU nodes, each including: a processor and a memory; at least two IO domains, wherein at least one of the IO domains is designated as an active IO domain performing communication functions for an active CPU node; and a switching fabric connecting each CPU node to each IO domain. One CPU node is designated as a standby CPU node and the rest is designated as the active CPU node. If a failure, a beginning of a failure, or a predicted failure occurs in an active node, the state and memory of the active CPU node are transferred to the standby CPU nod, which becomes a new active CPU node.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (Priority Claim) This application claims priority to U.S. Provisional Application No. 62 / 717,939, filed August 13, 2018, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates generally to architectures for highly reliable fault-tolerant computer systems, and more particularly to fault-tolerant computer systems with modular redundant components interconnected by a switching fabric. [Background technology]

[0003] A high-reliability fault-tolerant computer system is one that has at least "99.999%" reliability, meaning that the computer is functional at least 99.999% of the time and has a maximum of only about 5 minutes of unplanned downtime per year of continuous operation.

[0004] To achieve this high reliability, such fault-tolerant computer systems frequently have redundant components so that when one component fails, begins to fail, or is predicted to fail, programs using the failing computer component will instead use a similar but redundant component in the system. Generally, there are two ways in which this failover is implemented.

[0005] One method is to have two or more processor systems that each run the same application simultaneously and periodically compare their results to ensure they arrive at the same result. In such a system, when one system fails, the other can continue to operate as a simplex or single processor system until the failing system is replaced. When the failing system is replaced, the state of the non-failed system is copied to the replacement processor system, and both systems then continue to run the same application as the duplicated system.

[0006] A second method is to have two processor systems: one active processor system running an application, and another standby processor system. In this configuration, the standby processor system periodically receives updates from the active processor system along with the state and memory contents of the active processor system. These points in time are called checkpoints, and the transferred data is referred to as checkpoint data. When the active processor fails, begins to fail, or is predicted to fail, the active processor transfers its final state and memory contents to the standby processor system, which becomes the new active processor system and continues computation from the point where the active processor system previously transferred its final state.

[0007] Both of these methods have drawbacks: in the first method, the cost of continuously maintaining and running two duplicate computer systems is not insignificant; in the second method, the time required to periodically provide checkpoint state and memory data to the standby system slows down application processing by the active computer system.

[0008] The present disclosure addresses these shortcomings and others. Summary of the Invention [Means for solving the problem]

[0009] In one aspect, the present disclosure relates to a fault-tolerant computer system. In one embodiment, the system includes a plurality of CPU nodes, each CPU node including a processor and a memory, one of the CPU nodes designated as a standby CPU node and the remaining designated as active CPU nodes; at least two IO domains, at least one of the IO domains designated as an active IO domain performing communication functions for the active CPU node; and a switching fabric connecting each CPU node to each IO domain. In another embodiment, if one of a failure, an initiating failure, or a predicted failure occurs in the active node, the state and memory of the active CPU node are transferred through the switching fabric to the standby CPU node, and the standby CPU node becomes the new active CPU node and takes over for the previously failed node. In yet another embodiment, if one of a failure, an initiating failure, or a predicted failure occurs in an active IO domain, referred to as the failed active IO domain, the communication functions performed by the failed active IO domain are transferred to another IO domain.

[0010] In yet another embodiment, each CPU node further includes a communication interface that communicates with the switching fabric. In yet another embodiment, each IO domain includes at least two switching fabric control components, each switching fabric control component communicating with the switching fabric. In one embodiment, each IO domain further includes a management processor. In yet another embodiment, the management processor of the active IO domain controls communications through the switching fabric. In yet another embodiment, each IO domain communicates with other IO domains through serial links. In one embodiment, the switching fabric is a Non-Transparent Bridge (NTB) PCI Express (PCIe) switching fabric.

[0011] In one embodiment, each IO domain further includes a set of IO devices, each IO device including one or more physical and / or virtual functions, and one or more physical and / or virtual functions within an IO domain are shareable. In one embodiment, the one or more physical functions include one or more virtual functions. In one embodiment, the set of IO devices and the one or more physical and / or virtual functions are distributed among one or more CPU nodes and one or more of two switching fabric control components to define one or more sub-hierarchies assignable to ports of one or more CPU nodes. In one embodiment, one or more of the set of IO devices and the virtual functions are partitioned among a set of physical CPU nodes, the set of physical CPU nodes comprising an active CPU node and a standby CPU node. In one embodiment, the system further includes one or more management engine instances running on a management processor in each IO domain, each management engine querying switching fabric control components connected to the respective management engine to obtain an enumerated hierarchy of physical and / or virtual functions on a per control component basis, and each management engine merging the enumerated per-component hierarchy into a per-domain hierarchy of physical and / or virtual functions in the IO domain associated with each management engine.

[0012] In one embodiment, the system further includes one or more provisioning service instances, each provisioning service running on a management processor in each IO domain, each provisioning service querying a per-domain management engine instance for a per-domain hierarchy, and each per-domain instance of the provisioning service communicating with provisioning services in other IO domains to form a unified hierarchy of physical and / or virtual functions within the system. In one embodiment, any of the per-domain instances of the provisioning service can service requests from a composition interface, and any of the provisioning services can also validate the variability of one or more system composition requests with a view to ensuring redundancy across IO domains. In one embodiment, any of the per-domain instances of the provisioning service can service requests from an industry-standard system composition interface, and any of the per-domain instances of the provisioning service can also validate the variability of system composition requests with a view to ensuring redundancy across IO domains.

[0013] In another aspect, the present disclosure relates to a method for implementing CPU node failover in a fault-tolerant computer system. In one embodiment, the fault-tolerant computer system includes a plurality of CPU nodes, each including a processor and a memory, one of the CPU nodes designated as a standby CPU node and the remaining designated as active CPU nodes; at least two IO domains, at least one of the IO domains designated as an active IO domain performing communication functions for the active CPU node; and a switching fabric connecting each CPU node to each IO domain. In another embodiment, the method includes establishing a DMA data path between a memory of the active-but-failed CPU node and a memory of the standby CPU node. In yet another embodiment, the method includes transferring memory contents from the memory of the active-but-failed CPU node and the memory of the standby CPU node. In yet another embodiment, the method includes tracking addresses within the active-but-failed CPU node to which DMA accesses by the active-but-failed CPU node occur. In another embodiment, the method includes stopping accesses to memory on the active-but-failed CPU node and copying any memory data that has been accessed since DMA was initiated. In yet another embodiment, the method includes copying the state of the processors in the active-but-failed CPU node to a standby CPU node. In yet still another embodiment, the method includes exchanging all resource mappings from the active-but-failed CPU node to a standby CPU node and enabling a previously designated standby CPU node to be the new active CPU node.

[0014] In another embodiment, the method includes the step of the active-but-failed CPU node setting a flag in its own NTB window in the PCI-memory mapped IO space and in the NTB window of the standby CPU node so that each CPU node has its own intended new state after the failover operation. In yet another embodiment, the method includes the active-but-failed CPU node polling the standby CPU node regarding the status of the start routine.

[0015] In another aspect, the disclosure includes a method for implementing IO domain failover. In one embodiment, the method includes enabling a failure trigger for each switch fabric control component in each IO domain, the failure trigger including, but not limited to, a link down error, an uncorrectable fatal error, and a software trigger, and, in response to the failure trigger occurring, preventing drivers from using the failing IO domain.

[0016] While the present disclosure relates to different aspects and embodiments, it should be understood that the different aspects and embodiments disclosed herein may be integrated, combined, or used together, as combined systems, or in part, as separate components, devices, and systems, as appropriate. Thus, each embodiment disclosed herein may incorporate each of the aspects to varying degrees, as appropriate for a given implementation. The present specification also provides, for example, the following items: (Item 1) 1. A fault-tolerant computer system, comprising: a plurality of CPU nodes, each CPU node comprising a processor and a memory, one of the CPU nodes being designated as a standby CPU node and the remaining CPU nodes being designated as active CPU nodes; at least two IO domains, at least one of the IO domains designated as an active IO domain that performs communication functions for the active CPU node; A switching fabric that connects each CPU node to each IO domain Equipped with A fault-tolerant computer system, wherein if one of a failure, a failure initiation, and a predicted failure occurs in an active node, the state and memory of the active CPU node are transferred to the standby CPU node through a DMA data path, and the standby CPU node becomes the new active CPU node. (Item 2) Item 1. The fault-tolerant computer system of item 1, wherein each CPU node further comprises a communication interface for communicating with the switching fabric. (Item 3) 2. The fault-tolerant computer system of claim 1, wherein each IO domain comprises at least two switching fabric control components, each switching fabric control component communicating with the switching fabric. (Item 4) Item 1, a fault-tolerant computer system, wherein each IO domain further comprises a set of IO devices, each IO device comprises one or more physical functions and / or virtual functions, and one or more physical functions and / or virtual functions within an IO domain are sharable. (Item 5) Item 5. The fault-tolerant computer system of item 4, wherein the one or more physical functions comprise one or more virtual functions. (Item 6) 5. The fault-tolerant computer system of claim 4, wherein the set of IO devices and the one or more physical and / or virtual functions are distributed to one or more CPU nodes and one or more of two switching fabric control components to define one or more sub-hierarchies assignable to ports of the one or more CPU nodes. (Item 7) 5. The fault-tolerant computer system of claim 4, wherein one or more of the set of IO devices and the virtual functions are partitioned among a set of physical CPU nodes, the set of physical CPU nodes comprising the active CPU node and the standby CPU node. (Item 8) 7. The fault-tolerant computer system of claim 6, further comprising one or more management engine instances running on a management processor within each IO domain, each management engine querying the switching fabric control components connected to the respective management engine to obtain an enumerated hierarchy of physical and / or virtual functions on a per-control-component basis, and each management engine merging the enumerated per-component hierarchy into a per-domain hierarchy of physical and / or virtual functions within the IO domain associated with each management engine. (Item 9) 8. The fault-tolerant computer system of claim 7, further comprising one or more provisioning service instances, each provisioning service running on the management processor in each IO domain, each provisioning service querying a per-domain management engine instance on a per-domain hierarchy, each per-domain instance of each provisioning service communicating with the provisioning services in other I / O domains to form a unified hierarchy of physical and / or virtual functions within the system. (Item 10) 10. The fault-tolerant computer system of item 9, wherein any of the per-domain instances of the provisioning service are capable of servicing requests from a composition interface, and any of the provisioning services are also capable of validating the variability of one or more system composition requests with a view to ensuring redundancy across IO domains. (Item 11) Item 1. The fault-tolerant computer system of item 1, wherein each IO domain further comprises a management processor. (Item 12) Item 12. The fault-tolerant computer system of item 11, wherein a management processor of an active IO domain controls communications through the switching fabric. (Item 13) 2. The fault-tolerant computer system of claim 1, wherein each IO domain communicates with other IO domains through a serial link. (Item 14) Item 1. The fault-tolerant computer system of item 1, wherein the switching fabric is a Non-Transparent Bridge (NTB) PCI Express (PCIe) switching fabric. (Item 15) 1. A method for implementing CPU node failover in a fault-tolerant computer system, the fault-tolerant computer system having a plurality of CPU nodes, each CPU node comprising a processor and a memory, one of the CPU nodes designated as a standby CPU node and the remaining CPU nodes designated as active CPU nodes, at least two IO domains, at least one of the IO domains designated as an active IO domain that performs communication functions for the active CPU node, and a switching fabric connecting each CPU node to each IO domain, the method comprising: establishing a DMA data path between a memory of an active but failed CPU node and a memory of said standby CPU node; transferring memory contents from a memory of an active but failed CPU node and a memory of the standby CPU node through the DMA data path; tracking memory addresses within the active-but-failed CPU node to which DMA accesses occur by the active-but-failed CPU node; stopping access to memory on the active but failed CPU node and copying any memory data that has been accessed since DMA was initiated; copying the state of the processors in the active but failed CPU node to the standby CPU node; exchanging all resource mappings from the active but failed CPU node to the standby CPU node; enabling a previously designated standby CPU node to be the new active CPU node; A method comprising: (Item 16) Item 16. The method of item 15, further comprising the step of the active but failing CPU node setting a flag in its own NTB window in the PCI-memory mapped IO space and in the NTB window of the standby CPU node so that both CPU nodes have their own intended new state after the failover operation. (Item 17) Item 16. The method of item 15, further comprising the step of the active-but-failed CPU node polling the standby CPU node regarding the status of a start routine. (Item 18) 1. A method for implementing IO domain failover in a fault-tolerant computer system, the fault-tolerant computer system having a plurality of CPU nodes, each CPU node comprising a processor and a memory, one of the CPU nodes designated as a standby CPU node and the remaining CPU nodes designated as active CPU nodes, at least two IO domains, at least one of the IO domains designated as an active IO domain that performs communication functions for the active CPU node, and a switching fabric connecting each CPU node to each IO domain, the method comprising: enabling fault triggers for each switch fabric control component in each IO domain, the fault triggers comprising a link down error, an uncorrectable fatal error, and a software trigger; In response to a fault trigger occurring, stopping drivers that use the faulty IO domain; A method comprising: (Item 19) 1. A fault-tolerant computer system, comprising: a plurality of CPU nodes, each CPU node comprising a processor and a memory, one of the CPU nodes being designated as a standby CPU node and the remaining CPU nodes being designated as active CPU nodes; at least two IO domains, at least one of the IO domains designated as an active IO domain that performs communication functions for the active CPU node; A switching fabric that connects each CPU node to each IO domain Equipped with A fault-tolerant computer system, wherein if one of a failure, a failure initiation, and a predicted failure occurs in an active IO domain, referred to as a failing active IO domain, the communication function performed by the failing active IO domain is transferred to another IO domain. (Item 20) 20. The fault-tolerant computer system of claim 19, wherein each CPU node further comprises a communication interface for communicating with the switching fabric. (Item 21) 20. The fault-tolerant computer system of claim 19, wherein each IO domain comprises at least two switching fabric control components, each switching fabric control component communicating with the switching fabric. (Item 22) 20. The fault-tolerant computer system of item 19, wherein each IO domain further comprises a management processor. (Item 23) 23. The fault-tolerant computer system of claim 22, wherein a management processor of an active IO domain controls communications through the switching fabric. [Brief explanation of the drawings]

[0017] The structure and function of the present disclosure can best be understood from the description herein in conjunction with the accompanying drawings. The figures are not necessarily to scale, instead generally emphasizing illustrative principles. The figures are to be considered in all respects illustrative and not intended to limit the invention, the scope of which is defined solely by the claims.

[0018] [Figure 1] FIG. 1 is a block diagram of a reliable fault-tolerant computer system constructed in accordance with the present disclosure.

[0019] [Figure 2] FIG. 2 is a block diagram of a more detailed embodiment of the system of FIG.

[0020] [Figure 3] FIG. 3 is a flow diagram of an embodiment of the steps for CPU failover according to the present disclosure.

[0021] [Figure 4] FIG. 4 is a schematic diagram of an embodiment of a fault-tolerant system with active and standby computers of FIG.

[0022] [Figure 5A] 5A-5C are schematic diagrams of embodiments of the operating software and execution states of the OS, hypervisor, guest VMs, FT Virtual Machine Manager (FTVMM), and other layers within a fault-tolerant computer system during various stages of memory duplication or mirroring. [Figure 5B] 5A-5C are schematic diagrams of embodiments of the operating software and execution states of the OS, hypervisor, guest VMs, FT Virtual Machine Manager (FTVMM), and other layers within a fault-tolerant computer system during various stages of memory duplication or mirroring. [Figure 5C] 5A-5C are schematic diagrams of embodiments of the operating software and execution states of the OS, hypervisor, guest VMs, FT Virtual Machine Manager (FTVMM), and other layers within a fault-tolerant computer system during various stages of memory duplication or mirroring. DETAILED DESCRIPTION OF THE INVENTION

[0023] In summary, a highly reliable, fault-tolerant computer system 10 constructed in accordance with the present disclosure, in one embodiment, includes multiple CPU nodes (generally 14) interconnected into at least two IO domains (generally 26) through a mesh fabric network 30 as shown in FIG. 1. At least one of the nodes 14, 14C, is a standby node and does not run applications until one of the other CPU nodes 14, 14A, 14B, either begins to fail or actually fails. When a failure occurs, the standby CPU node 14C acquires the state of the failing CPU node (e.g., CPU node 14) and continues to run applications that were running on the failing CPU node 14.

[0024] Referring also to FIG. 2 , in normal operation, the CPU node 14 communicates with the outside world through one of two IO domains 26, 26A in one embodiment. The IO domains 26 are connected to the CPU node 14 by a switching fabric 30, with each CPU node 14 connected to each IO domain 26 by various buses, communication channels, interconnects, or links 31. In various embodiments, the buses, communication channels, interconnects, or links 31 may include multiple PCI Express xN interfaces, where N is 1, 2, 4, 8, or 16. The IO domains 26 are also redundant. In one embodiment, the IO domain 26 is a primary domain that controls communications from the CPU node 14, and a second IO domain 26A that acts as a standby domain in case the IO domain 26 fails. Each IO domain 26, 26A is also connected to the outside world through a network interface, generally 44, a storage controller and disks, generally 46, and additional IO devices.

[0025] Each IO domain 26, 26A may be connected to one or more storage disks 46. These storage disks are redundant and have duplicate or mirrored copies of the same data, as indicated by the dotted lines. Writes to an IO domain are sent to each internal disk in both IO domains 26 and written to a mirrored pair of disks D1 and D1A. Read requests for data are serviced from the primary domain 26, as long as it is functioning properly. In the event of a failure of the primary domain 26, read requests are serviced from disks in the secondary or standby IO domain 26A.

[0026] The present system will now be discussed generally in terms of its hardware architecture and its operation when a hardware failure occurs.

[0027] Hardware Implementation More specifically, and referring again to FIG. 2 , a system 10 constructed in accordance with one embodiment of the present disclosure includes several CPU nodes 14, 14A, 14B, 14C (generally, 14), each including a hardware CPU unit 18, 18A, 18B, 18C (generally, 18), and communication interfaces 22, 22A, 22B, 22C (generally, 22). System 10 also includes at least two IO domains 26, 26A, respectively, connected to each of the communication interfaces 22 through a Non-Transparent Bridge (NTB) PCI Express (PCIe) switching fabric 30. In various embodiments, the switching fabric is implemented with a variable number of channels. In these embodiments, various hardware configurations can be used to modify the bandwidth of those channels depending on the specific implementation or requirements of the system.

[0028] For example, in one embodiment, a CPU unit may connect to an IO domain using switches (34, 34A, 34B, 34C) that support the Gen 3 PCI Express protocol. In other embodiments, the switches (34, 34A, 34B, 34C) may support newer versions of the PCI Express protocol, such as Gen 4 or Gen 5. Each CPU hardware node 18 includes an operating system (OS), either a single system context OS or a virtual OS, including a hypervisor and multiple virtual machines. In general, the various systems and methods disclosed herein can be used with various application-specific integrated circuits, communication channels, buses, etc. that facilitate the use of PCIe (and other similar architectures) as a fabric for data centers, cloud computing, and other applications.

[0029] In use, in the depicted embodiment of system 10 of N CPU nodes 14, N-1 CPU nodes 14 are active nodes with running operating systems and application programs, and the Nth CPU node 14 is held in standby mode, typically running a minimal Unified Extensible Firmware Interface (UEFI) firmware program. Again, for a redundant system to function, at least one node must be held in reserve in case an active node fails, although the system may include multiple standby nodes. In one embodiment, the minimal UEFI is maintained on the Nth or standby CPU node 14 N The system includes a diagnostic program that provides information about the status of the hardware within the system. The status of the standby CPU node 14 is available to all of the active CPU nodes 14. When an active CPU node begins to fail and needs to fail over to transfer its computations to a non-failed CPU node 14, the failing CPU node 14 will reserve the standby CPU node 14 and initiate the failover process in cooperation with management processors (MP) 38 and 38A, as discussed below.

[0030] In one embodiment (used here as an example of system operation), N CPU nodes 14 are connected to at least two IO domains 26, but only one of the IO domains (IO0) functionally communicates with the N CPU nodes 14 and provides a communication link with an external or non-system network. Each IO board includes several embedded devices (e.g., network and storage controllers) and PCI Express slots that can be populated with a user-selected controller. Each IO domain 26 includes two switching fabric control components 34, 34A, 34B, 34C (generally, 34). Each switching fabric component 34 is configurable by firmware or by software through a set of API functions for internal switch management. In one embodiment, the switching fabric control components 34 are fabric-mode PCI Express switches with switching integrated circuits that, in conjunction with an embedded management processor (MP) 38, 38A (generally, 38), control connections to CPU nodes 14 in the IO domains 26 within the switching fabric 30 through an instantiation of software referred to as a manageability engine (ME).

[0031] Each ME communicates directly with the fabric control component 34 on board its respective MP 38. The MP 38 of a domain 26 communicates with the MP 38A of the other IO domain 26A through an Intelligent Platform Management Interface (IPMI) serial communication link 42. Each ME instance queries the fabric control switching component 34 to which it is connected by the switching component's firmware API for a list of physical and virtual functions within the switching component's hierarchy. A "virtual function" is a function defined under the PCI Express specification so that a single physical device can be shared among multiple guest virtual machines (VMs) without the overhead of placing additional demands on the hypervisor or host operating system that controls the functionality of individual guest VMs and their communication with the outside world. In addition, the ME also provides an API (ME-API) for viewing, configuring, and allocating IO devices 44, 46 that function within its IO domain 26, 26A. These I / O devices and functions are then allocated to the switching components 34, 34A and, again by the firmware API, to the CPUs 18 assigned to each switching component. This in turn creates N composite sub-hierarchies, each of which can be assigned to a host port, i.e., I / O devices 44, 46 are assigned to switching fabric components 34, 34A, which are then assigned to specific ports on CPU nodes 14, 14A, etc.

[0032] The allocation of nodes and devices is provided to the MEs by the user by a Management Service (MS) application, optionally including a user GUI, running on the host computer 50 and communicating with each ME over the network 52. The MS provides a set of function calls to the ME kernel driver. A Provisioning Service (PS) running on each ME receives, through ME-API calls from the host computer 50, a list of CPU, memory, VM, and IO resource requirements for each CPU node established by the user.

[0033] IO domains 26 and 26A are configured whereby each CPU node 14 has exclusive access to a subset of IO endpoint devices, such as disks 46, within IO domain 26 and 26A. The endpoint devices may provide physical or virtual functions within IO domain 26. In other embodiments, second IO domain 26A is also active, with a number of active and / or standby CPU nodes connected to it and communicating through it.

[0034] During operation, each CPU node 14 except one is active and executes its operating system and associated application program code. Each active CPU node 14's operating system (Windows® (Microsoft Corporation, Redmond, WA), ESXi) TM The IO domains (VMWare Inc., Palo Alto, CA, Linux® (Linus Torvalds, Dunthorpe, OR), etc.) may be the same as or different from the other active CPU nodes. Each active CPU node 14 communicates through one active IO domain (IO0) 26, while the other IO domain (IO1) 26A is available as a secondary IO domain. In some cases, the secondary IO domain maintains a mirrored copy of data, as done in a RAID1 configuration for disks. In other cases, the secondary domain may provide load balancing or hot standby services for multiple network ports treated as a single network port (called a "network bond" in Linux® or a "team" in Windows®).

[0035] Once the configuration of I / O domains and CPU nodes is established, if a CPU node 26, e.g., CPU node 142, begins to fail, and one or more active CPU nodes are running instances of the operating system, the operating system and application programs of that node will be able to continue running on standby CPU node 14. N The active IO domain (IO0) configures the switching fabric 30 to act as the previously standby node, now the active CPU node 14. N CPU node 14 N The previously active CPU node 142 that is diagnosed as faulty undergoes further diagnostics to determine whether it needs to be replaced. If the CPU node 142 now passes the diagnostics, the error is assumed to be either software-induced or transient, and the CPU node 142 becomes the new standby node.

[0036] The result is similar to when none of the CPU nodes 14 has failed, but an active IO domain (e.g., 26) is determined to have failed or is about to fail. In this case, connectivity information for the active CPU node 14 through the switching fabric 30 is passed by the MP 38 of the IO domain (e.g., 26) to the MP 38A of the IO domain 26A, and the IO domain 26A becomes the new active IO domain. The previously active IO domain 26 may then undergo further diagnostics or be removed and replaced. The IO domain 26 then becomes the new standby IO domain. Instructions and other data are passed between the MP 38 and MP 38A through the intelligent platform management interface serial communication link 42. Note that if both IO domains 26, 26A are active, if one of the IO domains (e.g., 26) fails, the fabric switch 30 is reconfigured so that the other IO domain 26A accepts all communications through the non-failed IO domain 26A.

[0037] In one embodiment using the Windows® operating system and zero-copy direct memory access (DMA), the switching fabric 30 can transfer processor state and memory contents from one CPU node 14 to another at approximately 56 GB / sec. The fact that the CPU nodes 14 and IO domains 26 are separate components has the added benefit of reducing the number of single points of failure and adding the ability to replace a failing component without affecting the corresponding non-failed component, e.g., the IO domain 26 but not the CPU node 14. Redundant IO domains 26 and CPU nodes 14 allow failing components to be dynamically replaced or even added without severely impacting applications running on the CPU nodes 14 and / or IO domains 26.

[0038] Referring also to FIG. 3, the operation of the system will now be considered in more detail.

[0039] System initialization The process of powering on a multi-node platform and provisioning IO resources can be performed for both physical IO devices and their PCIe functions, and for virtual functions for devices that support the features of the Single Root I / O Virtualization and Shared (SR-IOV) portion of the PCI Express specification. The initialization process is orchestrated by the hardware and software within each IO domain. During the initialization process, only the IO domain is released from reset, while all CPU nodes remain in a reset state.

[0040] Each switching component 34 completes its reset and initialization process and executes its internal switch management firmware, which enumerates its device hierarchy. The hierarchy is generally defined by the complete set of all switching components' primary bus reference numbers (how the switching component is accessed), secondary buses (the bus reference numbers for the other side of the bridge), and subordinate bus numbers (the highest bus numbers that exist anywhere below the bridge, along with all PCIe devices and functions below the switch and bridge). In one embodiment, a bridge refers to the bus address or bus number for the fabric controller. The present system and method can be implemented with various hierarchies, which may include bridges / switches below other bridges / switches. Simultaneously, each management processor 38 completes its reset and initialization and subsequently loads and executes an ME instance. Each ME instance queries the switching component 34, 34A to which it is connected by the switching component firmware API for a fully enumerated hierarchy. Each ME instance then merges the hierarchy from its connected switching component 34, 34A into a single list of physical and virtual functions within its IO domain.

[0041] Once the IO domain function lists are established, the ME instances communicate with each other to merge the domain-specific hierarchical lists into a unified list for the entire system, establishing one IO domain and associated ME as the "primary IO" for the system.

[0042] The previous steps allow users to configure the system according to their needs. However, for users who do not wish to take advantage of this capability, a default configuration of replicated resources is allocated to each active compute node 14 (active CPU), but the user can modify the allocation via the provisioning service if desired. The standby CPU node 14C (standby host) is given only a minimal set of IO devices so that it can run diagnostics and report its status to other nodes. In addition, if desired, the user can also override the allocation of IO resources to the various compute nodes. The new allocation data is then stored for the next cold boot.

[0043] Once resource provisioning is established for each CPU node 14, each ME instance deploys the desired resources to each switch component 34 for each associated CPU node port. The provisioning data is stored in flash or other non-volatile repository accessible to the ME 38, 38A for use the next time the entire platform or IO domain 26, 26A is reset or powered off and then powered on. In addition, each ME instance enables downstream port containment (DPC) triggers in the switch component 34, to which it is connected via the switch component firmware API, for each downstream port to detect and respond to events, including, but not limited to, “link down” errors, uncorrectable and fatal errors, and other software triggers. When any one IO device encounters an error, the switch component hardware 34 isolates that downstream link, and the firmware synthesizes responses / completions for any pending transactions to devices under that link. The firmware also signals the event to the ME, which in turn generates a platform interrupt to affected hosts, informing them that the IO device has become inaccessible.

[0044] Once the host-specific device hierarchy is established, each compute node 14 may be released from reset. The boot process for a given CPU 18 in a multi-node platform is identical to that which would be for any standard server with a standard BIOS and standard OS. Specifically, the BIOS on each CPU node uses a power-on self-test (POST process) to determine system health and enumerate the IO hierarchy exposed to that CPU node by the ME firmware. In one embodiment, the ME firmware boots first and allocates available IO resources among the CPU nodes that will be the active hosts. Once each host begins booting its own BIOS, each such host is unaware of the existence of the ME. Thus, each node boots in the normal manner. Once the OS boots and applicable software components are loaded, such software can again interact with the ME firmware. Despite this, from the BIOS and base OS frame of reference, neither is aware of nor interacts with the ME. The OS boot loader loads the OS image into memory and begins OS execution. The system then boots normally, with all IO domains 26 present and visible to the OS. The OS then loads a hardening driver for each instance of each IO device.

[0045] Network controller functions are bonded / teamed using standard OS features. Similarly, external storage controllers are duplexed when both instances of the controller are healthy and have connectivity to the external storage array. Internal storage controllers require additional consideration. When all IO domain devices other than legacy IO are duplexed, then the entire IO domain is safe to duplex / pull.

[0046] Once the system is fully operational, its operation under fault conditions is next considered.

[0047] CPU / Memory Failover The following steps are implemented by one embodiment of the present disclosure to avoid system failure when a CPU node (CPU processor 18 and / or memory) failure occurs or is predicted: Applications running on the failing CPU node will then be transferred to the standby CPU node 14C.

[0048] In overview, an active CPU node 14, suffering either a number of correctable errors above a predetermined threshold or other degraded capabilities, indicates to the MP 38 associated with the node's IO domain 26 that the node 14 has reached this degraded state and that a failover to a non-failed standby CPU node 14C should be initiated. The active CPU node 14, MP 38, and standby CPU node 14C then engage in a communication protocol to manage the failover process and the transfer of state from the active CPU node 14 to the standby CPU node 14C. The standby CPU node 14C, which is the target location for the failover operation, signals that it is ready to be removed from its diagnostic UEFI loop and begin the process of receiving memory contents and state information from the failed active CPU node 14. The active-but-failed CPU node 14 polls the standby CPU node 14C regarding the status of the standby CPU node's startup routine. The standby CPU node 14C activates an NTB window into its PCI-memory mapped IO space and begins polling for commands from the active but failing CPU node 14.

[0049] 3, at a high level, once the status from standby CPU node 14C is reported to active-but-failed CPU node 14A, the active-but-failed CPU node 14 enables the data path (step 300) to allow DMA memory copies from the memory of active-but-failed CPU node 14 to the memory of standby node 14C. Standby CPU node 14C at this point cannot access any IO domains 26 or initiate read or write accesses to the memory of active-but-failed CPU node 14.

[0050] An active but failing CPU node 14 signals all its drivers that are capable of tracking changes to memory to begin tracking addresses where DMA traffic is active (both DMA write buffers and DMA control structures).

[0051] All memory is copied from the active but failing CPU node 14 to the memory of the standby CPU node 14C (step 310) while DMA traffic continues and while the processor continues to execute instructions. The register state of each device physically located within the failing CPU node 14 is copied to the standby node 14C. This period of time during which memory is copied while DMA traffic is still occurring constitutes a power saving period.

[0052] The active but failing CPU node tracks pages modified by CPU accesses in addition to the driver tracking pages potentially modified by DMA traffic (step 320). During the save time, modified pages can be recopied while the driver and host software continue to track newly modified pages. This process is fully described in U.S. patent application Ser. No. 15 / 646,769, filed July 11, 2017, the contents of which are incorporated herein by reference in their entirety.

[0053] To understand how the power outage phase of this process operates, it is necessary to consider the operation of a fault-tolerant system in more detail. Referring now to Figure 4, a fault-tolerant computer system includes at least two identical computers or nodes 414 and 414A. One computer or node 414 is currently active, or the primary processor, receiving requests from clients or users and providing output data thereto. The other computer or node 414A is referred to as the standby or secondary computer or node.

[0054] Each computer or node (generally, 414) includes a CPU 422, 422-1, 422A, 422A-1, a memory 426, 426A, a switch 430, 430A, and an input / output (I / O) module 434, 434A. In one embodiment, the two physical processor subsystems 414 and 414A reside on and communicate with each other through the same switching fabric 438. The switching fabric controllers 430, 430A coordinate the transfer of data (arrows 440, 445, and 450) from the currently active memory 426 to the standby or mirrored memory 426A so that the fault-tolerant system can create identical memory contents in both (currently active and standby) subsystems 414, 414A. The I / O modules 434, 434A allow the two subsystems 414 and 414A to communicate with the outside world, such as disk storage 46 (FIG. 2) and a network, through a network interface (NI) 44 (FIG. 2).

[0055] Although the present discussion is in terms of an embodiment involving two processor subsystems, more than two processor subsystems can be used in a fault-tolerant computer system. In the case of a multiple processor subsystem, e.g., a three-processor (e.g., A, B, C) fault-tolerant computer system, mirroring of the three processor subsystems is performed in two steps. First, processor subsystems A and B are mirrored, then the resulting mirrored A, B processor subsystem is mirrored to the C processor subsystem, etc.

[0056] During power-saving and subsequent power-loss phases, modified memory must be tracked and subsequently copied when DMA traffic is stopped. The problem is that the server's native operating system may not provide a suitable interface for copying dirty pages from active memory 426 to mirror memory 426A, especially when a virtual machine (VM) system is used. For example, some physical processors, such as Intel Haswell and Broadwell processors (Intel Corporation, Santa Clara, CA USA), provide a set of hardware virtualization capabilities, including the VMX root operation, that allow multiple virtual operating systems to share the same physical processor while simultaneously taking full control of many aspects of system execution. Each virtual machine has its own operating system under the control of the host hypervisor. Such systems may not provide an interface for detecting and copying dirty pages for memory used by those virtual machines. To understand how the present disclosure addresses this limitation, Figures 5(A)-5(C) depict, as a layered software diagram, the state of a fault-tolerant computer system as it undergoes various operations.

[0057] Referring to FIG. 5A , in normal, non-mirrored operation, the layers in a fault-tolerant computer system include a hardware layer 500, including a DMA-enabled switch 430; a server firmware layer 504, including a system Universal Extensible Firmware Interface (UEFI) BIOS 508; and a layer zero reserved memory area 512, which is initialized to zero. The layer zero reserved memory 512 is reserved by the BIOS 508 at boot time. While most of the memory in a fault-tolerant computer system is available for use by the operating system and software, the reserved memory 512 is not. The size of the reserved memory area 512 provides sufficient space for the FTVMM and SLAT table, which is configured with a 4KB (four kilobyte) page granularity and a one-to-one mapping of all system memory. The FTVMM module allows all processors to run their programs as guests of the FTVMM module. A second level address translation table (SLAT) (also referred to by various manufacturers, i.e., Extended Page Table [EPT] by Intel and Rapid Virtualization Indexing [RVI] by AMD) within a reserved portion of memory is used to translate memory references to physical memory. In one embodiment, a four-level SLAT table provides a memory map with dirty and accessed bit settings that will identify all memory pages modified by the operating system kernel and other software. A four-level SLAT is sufficient to provide sufficient granularity to address each word of memory with four kilobyte granularity, although other page sizes and mappings are possible.

[0058] The next layer (L1) 520 includes the operating system and drivers for the fault-tolerant computer system, including FT kernel mode drivers 522, and the commonly used hypervisor host 524.

[0059] The last layer (L2) 530 includes non-virtualized server software components that are not controlled by the virtual machine control structure (VMCS) 550 during normal operation, such as processes, applications, and so forth 534, including any virtual machine guests (VMs) 538, 538A. The non-virtualized software components 534 include a FT management layer 536. Each virtual machine guest (VM) includes a VM guest operating system (VM OS) 542, 542A and a SLAT table associated with the VM (SLAT L2) 546, 546A. Also included within each VM 538, 538A are one or more virtual machine control structures (VMCS-N), generally 550, 550A, associated with the VM, one for each virtual processor 0-N allocated to that VM. In the diagram shown, the virtual processor VMCSs are labeled VMCS0 through VMCS-N. Each VMCS contains a control field to enable a SLAT table pointer (such as the Intel Extended Page Table Pointer EPTP) that can provide a mapping to translate guest physical addresses into system physical addresses.

[0060] 5B, at the initiation of mirroring, the fault-tolerant computer system is operating in a non-mirrored mode. FT management layer 536 causes FT kernel mode driver (FT driver) 522 to begin processing commands that enter mirrored execution. FT kernel mode driver 522 loads or writes the program and data code for FT virtual machine monitor (FTVMM) code 580, FTVMM data 584, SLAT L0 588, and VMCS-L0 array 592 into the reserved memory areas.

[0061] The FT driver initializes the VMCS L0 for each processor, installs the FTVMM, and runs as a hypervisor whose program code is executed directly by all VMEXIT events that occur within a fault-tolerant computer system (i.e., the processor mechanism that transfers execution from the guest L2 into the hypervisor that controls the guest). The FTVMM handles all VMEXITs and mimics the normal handling of the event that caused the VMEXIT, such that OS1, OS2, the OS's commonly used hypervisor L1, and the guest L2 will continue their processing in a functionally normal manner as if the FTVMM were not installed and active.

[0062] At this point, memory content transfer occurs under the two conditions discussed above: "power save" and "power outage." Mirroring during power save and power outage may occur once steady-state operation is reached, within a few minutes after the initial fault-tolerant computer system boot, or whenever a processor subsystem is brought back online after a hardware error on a running fault-tolerant computer system. As discussed above, during power save phases, normal system workloads are processed and processors continue to perform calculations and access and modify active memory. Dirty pages caused by memory writes during power save (while copying memory to the second subsystem) are tracked and copied in the next power save or power outage phase. The FTVMM provides a dirty page bitmap to identify memory pages modified in each phase. In power save phase 0, all memory is copied, tracking newly dirtied pages. In power save phase 1 and beyond, only pages dirtied during the previous phase are copied. In power outages, all processors except one are paused and interrupts are disabled. No system workload is processed during the power outage. Dirtyed pages from the previous (power saving) phase are copied and a final modified page range list is created. The remaining dirty pages and active processor state are then copied to standby computer memory. Once this is complete, the FT driver generates a system management interrupt and all processors running in the firmware UEFI BIOS and firmware SMM modules generate an SMI requesting MPs 38 and 38A to change host ports on switches 34, 34A, 34B, and 34C to standby CPU 14C, after which operation begins on CPU 14C, which is now the new online CPU and no longer the standby CPU.The firmware SMM performs a resume to the FT driver, which completes the power outage phase, unloads the FTVMM, releases the paused processors, enables interrupts, and completes its handling of the CPU failover request.

[0063] 5C, once the mirroring process is complete, the FTVMM code 580 in reserved memory 512 is unloaded and no longer executing. The FTVMM data 584, SLAT 588, and VMCS 192 are unused and the reserved memory is idle, waiting for the next error condition.

[0064] More specifically, during the first phase of power saving, the FT kernel-mode driver uses the VMCALL functional interface with the FTVMM to issue a command called enable memory page tracking to request that the FTVMM begin tracking all pages of modified memory in the system. A VMCALL processor instruction in the FT driver's functional interface to the FTVMM causes each logical processor to enter the FTVMM and process requests issued by the FT driver. The FTVMM performs a function on all processors to begin using its program code in the FTVMM hypervisor context in a manner that obtains a record of all newly modified system memory pages (dirty pages). The FTVMM searches the SLAT L0 and all SLAT L2 tables, sets the dirty bits in these tables to zero, and then invalidates the cached SLAT table mappings on each processor. When all processors have completed this function in the FTVMM, the FTVMM returns control to the FT driver by executing a VMRESUME instruction. The FT driver then copies all of the system memory to the second subsystem. The FT driver may perform a high-speed memory transfer operation to copy all system memory to a secondary or standby computer using a DMA controller or switch 430. The fault-tolerant computer system continues to perform its configured workload during this process.

[0065] Power saving memory copy stage 1 As part of power-saving memory copy phase 1, the FT driver retrieves the dirty page bitmap and copies newly dirtyed pages of memory to the second subsystem. The FT kernel-mode driver uses a functional interface to issue a command called enable memory page tracking on each processor. A VMCALL processor instruction in the FT driver's functional interface to the FTVMM causes each logical processor to enter the FTVMM and process requests issued by the FT driver. The FTVMM performs a function on all processors to begin using its program code in the FTVMM hypervisor context in a manner that obtains a record of all newly modified system memory pages (dirty pages). The FTVMM code on each processor then searches every 8-byte page table entry in the SLAT L0 table and each guest's SLAT L2 table and compares the dirty bit in each entry with the bit's TRUE value. When the comparison result is TRUE, the FTVMM sets a bit field in the dirty page bitmap at the bit field address that represents the address of the dirty or modified page in physical memory, and then clears the dirty bit in the page table entry. Because the memory mapping configured in SLAT L0 has a page size of 4 kilobytes, one bit in the dirty page bitmap is set for each dirty page found.

[0066] The memory mapping configured by the hypervisor L1 in the SLAT L2 table may be larger than 4 kilobytes, and the FTVMM sets a contiguous series of bit fields in the dirty page bitmap when this occurs, such as 512 contiguous bit field entries for a 2 megabyte page size. When this process is complete for the SLAT L0 and SLAT L2 tables, each processor performs a processor instruction (such as the Intel processor instruction INVEPT) to invalidate the processor's cached translations for the SLAT L0 and SLAT L2 tables, allowing the FTVMM to continue detecting new instances of dirtied pages that may be caused by system workload.

[0067] When all processors have completed this operation in the FTVMM, the FTVMM returns control to the FT driver by executing a VMRESUME instruction. The FT driver then issues another MCALL functional interface command, called a dirty page bitmap request. The FTVMM then provides a dirty page bitmap containing a record of recently modified pages and stores this data in a memory buffer located in the FT driver's data area. The FT driver then copies the set of physical memory pages identified in the dirty page bitmap to corresponding physical memory addresses in the secondary or standby computer. The FT driver may use a DMA controller or switch 430 to perform a fast memory transfer operation to copy the set of dirtied pages to the second subsystem.

[0068] Power saving memory copy step 2-N / iteration The procedure, Memory Copy Phase 1, may be repeated one or more times to obtain a smaller resulting set of dirtied pages that may be generated by the system workload during the final power-saving Memory Copy Phase N. For example, in one embodiment, the FT driver may repeat the same sequence, obtain another dirty page bitmap, and copy the newly dirtied pages to the second subsystem one or more times.

[0069] After the power-save copy phase is complete, the active-but-failed CPU 14 signals its driver tracking DMA memory accesses to suspend all DMA traffic (step 330). This is the beginning of the power-down phase. CPU threads are then all suspended to prevent further modification of memory pages. At this point, the final list of pages modified by either CPU or DMA accesses is copied to the standby CPU 14C.

[0070] More specifically, during a power outage, the FT driver executes driver code in parallel on all processors on the active but failing CPU 14 and copies the final set of dirtyed pages to the standby CPU 14C. The FT driver causes all processors on CPU 14 to disable system interrupt handling on each processor to prevent other programs in the fault-tolerant computer system from generating more dirty page bits. The FT driver uses the VMCALL functional interface to issue a command called power outage page tracking enable, which causes the FTVMM to identify a set of recently dirtyed memory pages and to identify certain volatile or frequently modified memory pages, such as VMCS-N and SLATL2, and include those pages in the set of dirty pages. The FTVMM may temporarily suspend all processors except for processor #0 in the FTVMM. The FT driver then issues another VMCALL functional interface command, a dirty page bitmap request, to obtain the bitmap of dirty pages. The FTVMM then provides a dirty page bitmap containing a record of recently modified pages and stores this data in a memory buffer located in the FT driver's data area.

[0071] In one embodiment, the FT driver then copies the set of physical memory pages identified in the dirty page bitmap to corresponding physical memory addresses in the second subsystem. The FT driver then creates a list of memory ranges that are assumed to be dirty or modified, including memory ranges for reserved memory regions, and stores this information in a data structure called the last blackout memory range list. This procedure is called a blackout memory copy because the system workload is not running and the workload undergoes a short server processing outage while the final set of dirtyed pages is copied to the standby CPU 14C.

[0072] Once all of the memory of the active-but-failed CPU node 14 has been copied, the active-but-failed CPU node 14 saves (step 340) the internal state of its processor (including its registers, local advanced programmable interrupt controller, high-precision event timer, etc.) to memory locations and copies that data to the standby CPU node 14C, which is then restored into the corresponding registers of the standby CPU node 14C. A server management interrupt (SMI) feedback stack is created on the standby CPU node 14C for the final set of registers (such as the program counter) that need to be restored on the standby CPU node 14C to avoid processing from the exact point the active-but-failed CPU node was turned off.

[0073] The active-but-failed CPU node 14 sets flags in its own NTB window in PCI-Memory Mapped IO (PCI-MMIO) space and in the NTB window of the standby CPU node 14C so that each CPU node 14, 14C has its own intended new state after the failover operation. Any time prior to the completion of this step, the failover can be aborted and operation simply continues on the original active-but-still-failed CPU node.

[0074] To complete the failover, once all steps up to this point have been completed successfully, the active but failed CPU sends a command to the primary management processor (which will coordinate with the secondary management processor to handle any error cases in this step) to exchange all of the resource mappings between the host ports for the two CPU nodes 14, 14A involved in the failover operation (step 350). Each management processor will then make a series of firmware API calls to its local switch to effect the resource mapping changes. The primary management processor then signals the two CPU nodes when the switch reconfiguration is complete.

[0075] Both CPU nodes 14, 14C read tokens from their mailbox mechanisms indicating their new individual states (exchanged from the initial active and standby designations). Software on the new active CPU node then performs any final cleanup, as required. For example, it may be necessary to discipline the switching fabric to map transactions from the new active CPU node (step 360), implement a Resume from System Management (RSM) command, and replay a PCI enumeration resume to return control to the operating system and resume interrupted instructions. The standby CPU node can reactivate previously quiesced devices and allow transactions to flow through the fabric to and from the standby CPU node.

[0076] In addition to the CPU / memory failover capabilities described above, the present disclosure also enables the transfer of an active IO domain, e.g., IO1, to another or standby IO domain, e.g., IO2.

[0077] IO Domain Failover In one embodiment, the following steps are implemented to provide IO domain failover: In normal operation, when system 10 boots, all IO domains 26 are present and visible to the operating system. Drivers are then loaded for each instance of each IO domain. Network controller functions are bonded or teamed between the two active IO domains, primary 26 and secondary 26A, using standard operating system commands. Any external storage controller is duplexed when both instances have connectivity to the external storage array. When all PCI domain devices other than legacy IO are duplexed, then the entire IO bay is duplexed / safe to pull, meaning the platform can tolerate the loss of an IO bay due to either a failure or a service action without affecting the normal operation of the system and therefore its availability.

[0078] Briefly, in overview, a failure of a device, such as a disk controller 46, will trigger a downstream port containment (DPC) event within the affected switch component 34. The switch component hardware 34 isolates the link between the switch component 34 and the device, and the switch component firmware 34 completes any pending transactions to the device. The firmware also signals the event to the ME 38, which in turn generates an interrupt to the affected CPU node 14, informing them that an IO device has failed. The failure may cause a single endpoint device to be isolated by the DPC logic if detected by the switch component immediately above the affected device, or the failure may result in the isolation of a larger set of devices within the IO domain if detected by a switch higher in the device hierarchy.

[0079] More specifically, the MP 38, 38A enables a DPC (downstream port containment) failover trigger within the fabric mode switching component (generally, 34) for each downstream port connecting a CPU node to a device. Potential failover triggers include, for example, link-down errors and uncorrectable fatal errors, in addition to intentional software triggers to shut down and remove an IO domain 26. When the failover trigger is enabled, if any IO domain 26 experiences an error, the fabric switching component 34 for that IO domain 26 isolates its communication link with the device and completes any pending transactions to the device utilizing that link. The MP 38, 38A then generates an interrupt to the CPU node 14 utilizing the device, informing them that the device or IO domain 26 has failed.

[0080] The CPU node OS receives a platform interrupt and interrogates the MPs 38, 38A regarding whether the failed IO device or the entire IO domain board 26 has failed. The CPU node OS will then initiate removal of drivers for the affected devices. The surviving duplex partner for each affected IO device will be marked as "simplex / primary." The MPs 38, 38A (Figure 2) on unaffected boards will be marked as primary.

[0081] The MP 38, 38A then reads the error registers from the failing IO domain 26 and from the switch component 34. The MP 38, 38A then attempts to bring the device back up by asserting an independent reset to that domain, followed by some diagnostics. If the diagnostics are successful, the MP 38, 38A will restart functions in the device, including any virtual functions that were created and stored in flash memory in the IO domain 26, 26A.

[0082] The MP38, 38A sends a message announcing the availability of a new hot-pluggable IO domain to each CPU node. Each CPU node then scans its PCIe hierarchy for devices, discovers the newly arrived device / function, and loads the appropriate driver.

[0083] When all IO devices are once again active in a duplex state, then the entire IO domain is safe to be duplexed / pulled again.

[0084] Unless otherwise specifically stated as will be apparent from the discussion below, throughout the description, discussions utilizing terms such as "processing," or "calculating," or "calculating," or "delaying," or "comparing," or "generating," or "determining," or "transferring," or "postponing," or "completing," or "suspending," or "handling," or "receiving," or "buffering," or "allocating," or "displaying," or "flag," or Boolean logic or other set related operations or equivalents, will be understood to refer to the actions and processes of a computer system or electronic device that manipulates and converts data represented as physical (electronic) quantities in registers and memory of the computer system or electronic device into other data that are similarly represented as physical quantities in electronic memory or registers, or other such information storage, transmission, or display device.

[0085] The algorithms presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the above description. In addition, the present disclosure is not described with reference to any particular programming language, and various embodiments may therefore be implemented using a variety of programming languages.

[0086] Several implementations have been described. Nevertheless, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. For example, various forms of the flows shown above may be used, in which steps are rearranged, added, or removed. Accordingly, other implementations are within the scope of the following claims.

[0087] The examples presented herein are intended to illustrate potential specific implementations of the present disclosure. The examples are intended primarily for purposes of illustration of the present disclosure for those skilled in the art. Any particular aspect or aspects of the examples are not necessarily intended to limit the scope of the present invention.

[0088] The figures and descriptions of the present disclosure have been simplified for purposes of clarity to illustrate elements relevant for a clear understanding of the present disclosure while excluding other elements. However, those skilled in the art will recognize that these types of intensive discussions do not facilitate a further understanding of the present disclosure, and therefore, more detailed descriptions of such elements are not provided herein.

[0089] The processes associated with the present embodiment may be executed by a programmable device such as a computer. Software or other sets of instructions that can be employed to cause the programmable device to execute the processes may be stored in any storage device, such as, for example, a computer system (non-volatile) memory, an optical disk, a magnetic tape, or a magnetic disk. Furthermore, some of the processes may be programmed when the computer system is manufactured or via a computer-readable memory medium.

[0090] It should also be understood that certain process aspects described herein may be implemented using instructions stored on one or more computer-readable memory media that direct a computer or computer system to perform the process steps. Computer-readable media may include, for example, memory devices such as diskettes, compact discs of both the read-only and read / write variety, optical disk drives, and hard disk drives. Computer-readable media may also include memory storage devices, which may be physical, virtual, permanent, temporary, semi-permanent, and / or semi-temporary.

[0091] The computer systems and computer-based devices disclosed herein may include memory for storing certain software applications used in acquiring, processing, and communicating information. It is understood that such memory may be internal or external with respect to the operation of the disclosed embodiments. Memory may also include any means for storing software, including hard disks, optical disks, floppy disks, ROM (read-only memory), RAM (random access memory), PROM (programmable ROM), EEPROM (electrically erasable programmable read-only memory), and / or other computer-readable memory media. In various embodiments, a "host," "engine," "loader," "filter," "platform," or "component" may include various computers or computer systems, or may include any reasonable combination of software, firmware, and / or hardware.

[0092] In various embodiments of the present disclosure, a single component may be replaced by multiple components, and multiple components may be replaced by a single component, to perform one or more given functions. Such substitutions are within the scope of the present disclosure, except to the extent that such substitutions would not function to practice embodiments of the present disclosure. Any of the servers may be replaced, for example, by a “server farm” or other group of networked servers (e.g., a group of server blades) located and configured for cooperative functions. It should be understood that a server farm may serve to distribute workload among the farm's individual components and expedite the computational process by utilizing the collective and cooperative capabilities of multiple servers. Such server farms may employ load balancing software to perform tasks such as, for example, tracking demand for processing power from different machines, prioritizing and scheduling tasks based on network demand, and / or providing backup contingency in the event of component failure or reduced operability.

[0093] In general, it may be apparent to those skilled in the art that the various embodiments described herein, or components or parts thereof, may be implemented in many different embodiments of software, firmware, and / or hardware, or modules thereof. The software code or specialized control hardware used to implement some of the present embodiments is not a limitation of the present disclosure. Programming languages ​​for computer software and other computer-implemented instructions may be converted into machine language by a compiler or assembler before execution, and / or may be converted directly at runtime by an interpreter.

[0094] Examples of assembly languages ​​include ARM, MIPS, and x86; examples of high-level languages ​​include Ada, BASIC, C, C++, C#, COBOL, Fortran, Java, Lisp, Pascal, Object Pascal; and examples of scripting languages ​​include Bourne script, JavaScript, Python, Ruby, PHP, and Perl. Various embodiments may be employed, for example, in a Lotus Notes environment. Such software may be stored on one or more suitable computer-readable media of any type, such as, for example, magnetic or optical storage media. Accordingly, the operation and behavior of embodiments will be described without specific reference to actual software code or specialized hardware components. This lack of specific reference is understandable, as it is clearly understood that one skilled in the art would be able to design software and control hardware to implement embodiments of the present disclosure based on the description herein with only reasonable effort and without undue experimentation.

[0095] Various embodiments of the systems and methods described herein may employ one or more electronic computer networks to facilitate communication between different components, transfer data, or share resources and information. Such computer networks can be categorized according to the hardware and software technologies used to interconnect devices in the network.

[0096] Computer networks may be characterized based on the functional relationships between the elements or components of the network, such as active networking, client-server, or peer-to-peer functional architectures. Computer networks may also be classified according to their network topology, such as bus networks, star networks, ring networks, mesh networks, star-bus networks, or hierarchical topology networks. Computer networks may also be classified based on the method employed for data communication, such as digital and analog networks.

[0097] Embodiments of the methods, systems, and tools described herein may employ internetworking to connect two or more distinct electronic computer networks or network segments through a common routing technology. The type of internetwork employed may depend on the management and / or involvement within the internetwork. Non-limiting examples of internetworks include intranets, extranets, and the Internet. Intranets and extranets may or may not have a connection to the Internet. If connected to the Internet, an intranet or extranet may be protected using appropriate authentication technology or other security measures. As applied herein, an intranet may be a group of networks employing Internet protocols, web browsers, and / or file transfer applications under common control by an administrative entity. Such an administrative entity may restrict access to the intranet, for example, to only authorized users or to another internal network of an organization or commercial entity.

[0098] Unless otherwise indicated, all numbers used in the specification and claims expressing lengths, widths, depths, or other dimensions, etc., are to be understood in all instances as indicating both the exact value as indicated and that which is modified by the term "about." As used herein, the term "about" refers to a ±10% variation from the nominal value. Thus, unless otherwise indicated, the numerical parameters set forth in this specification and the appended claims are approximations that may vary depending on the desired properties sought to be obtained. At the very least, and not as an attempt to limit the application of the principle to equivalents in the scope of the claims, each numerical parameter should be construed, at least in light of the number of reported significant digits and by applying ordinary rounding techniques. Any specific value may vary by 20%.

[0099] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The foregoing embodiments are, therefore, to be considered in all respects as illustrative and not limiting of the disclosure set forth herein. The scope of the invention is, therefore, indicated by the appended claims, rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein.

[0100] Those skilled in the art will understand that various modifications and changes can be made without departing from the scope of the described technology. Such modifications and changes are intended to fall within the scope of the described embodiments. Furthermore, those skilled in the art will understand that features included in one embodiment can be interchanged with other embodiments, and that one or more features from a depicted embodiment can be included with other depicted embodiments in any combination. For example, any of the various components described herein and / or depicted in the figures may be combined, interchanged, or excluded from other embodiments.

Claims

1. 1. A fault-tolerant computer system, comprising: a plurality of CPU nodes, each CPU node comprising a processor and a memory, one of the CPU nodes designated as a standby CPU node and the remaining CPU nodes designated as active CPU nodes, one or more of the active CPU nodes executing an operating system and one or more applications; at least two IO domains, at least one of which is designated an active IO domain that performs communication functions for the active CPU node; A switching fabric that connects each CPU node to each IO domain; Equipped with A fault-tolerant computer system, wherein when one of a failure, an onset of a failure, and a predicted failure occurs in an active node, the processor state and memory contents of the active CPU node are transferred to the standby CPU node through a DMA data path, the memory contents and processor state including the state of the operating system and the one or more applications such that the operating system and the one or more applications are transferred to the standby CPU node, and the standby CPU node becomes a new active CPU node and continues to run the transferred operating system and the one or more applications.

2. 2. The fault-tolerant computer system of claim 1, wherein each CPU node and IO domain is configured to be replaced without affecting applications running on one or more of the other CPU nodes and IO domains.

3. 2. The fault-tolerant computer system of claim 1, wherein each IO domain comprises at least two switching fabric control components, each switching fabric control component communicating with the switching fabric.

4. 2. The fault-tolerant computer system of claim 1, wherein each IO domain further comprises a set of IO devices, each IO device comprising one or more virtual functions, and wherein the one or more virtual functions within an IO domain are shareable.

5. 5. The fault-tolerant computer system of claim 4, wherein the set of IO devices and the one or more virtual functions are distributed to the one or more CPU nodes and one or more of the two switching fabric control components so as to define one or more sub-hierarchies assignable to ports of one or more CPU nodes.

6. 5. The fault-tolerant computer system of claim 4, wherein one or more of the set of IO devices and the virtual functions are partitioned among a set of physical CPU nodes, the set of physical CPU nodes comprising the active CPU node and the standby CPU node.

7. 6. The fault-tolerant computer system of claim 5, further comprising one or more management engine instances running on a management processor in each IO domain, each management engine querying the switching fabric control components connected to a respective management engine to obtain an enumerated hierarchy of the one or more virtual functions on a per control component basis, and each management engine merging the enumerated per-component hierarchy into a per-domain hierarchy of the one or more virtual functions in the IO domain associated with the respective management engine.

8. 7. The fault-tolerant computer system of claim 6, further comprising one or more provisioning service instances, each provisioning service running on the management processor in each IO domain, each provisioning service querying a per-domain management engine instance on a per-domain hierarchy, each per-domain instance of each provisioning service communicating with the provisioning services in other IO domains to form a unified hierarchy of the one or more virtual functions within the system.

9. 9. The fault-tolerant computer system of claim 8, wherein any of the per-domain instances of the provisioning service are capable of servicing requests from a composition interface, and any of the provisioning services are also capable of validating the variability of one or more system composition requests with a view to ensuring redundancy across IO domains.

10. 2. The fault-tolerant computer system of claim 1, wherein each IO domain further comprises a management processor and a first switching component, and each CPU node communicates with the switching fabric.

11. 11. The fault-tolerant computer system of claim 10, wherein the management processor of an active IO domain controls communications through the switching fabric.

12. 10. The fault-tolerant computer system of claim 1, wherein each IO domain communicates with other IO domains through a serial link.

13. 2. The fault-tolerant computer system of claim 1, wherein the switching fabric is a Non-Transparent Bridge (NTB) PCI Express (PCIe) switching fabric.

14. 1. A fault-tolerant computer system, comprising: a plurality of CPU nodes, each CPU node comprising a processor and a memory, one of the CPU nodes being designated as a standby CPU node and the remaining CPU nodes being designated as active CPU nodes; at least two IO domains, at least one of which is designated an active IO domain that performs communication functions for the active CPU node; A switching fabric that connects each CPU node to each IO domain; Equipped with one or more of the plurality of CPU nodes is configured to run an operating system and one or more applications; a DMA data path is established between the memory of the active but failed CPU node and the memory of the standby CPU node; memory contents and processor state are transferred from the memory of the active but failing CPU node to the memory of the standby CPU node through the DMA data path, the memory contents and processor state including the state of the operating system and one or more applications; the memory contents and processor state include the state of the operating system and the one or more applications, such that the operating system and the one or more applications are transferred to the standby CPU node, the standby CPU node becoming the new active CPU node and continuing to run the transferred operating system and the one or more applications; tracking memory addresses within the active-but-failed CPU node to which DMA accesses occur by the active-but-failed CPU node; Stopping access to memory on said active but failed CPU node and copying any memory data that has been accessed since DMA was initiated; The standby CPU node becomes the new active CPU node and continues to run the transferred operating system and the one or more applications.

15. 15. The fault-tolerant computer system of claim 14, wherein each CPU node further comprises a communication interface for communicating with the switching fabric.

16. 15. The fault-tolerant computer system of claim 14, wherein the active-but-failed CPU node is configured to poll the standby CPU node regarding the status of a start routine.

17. 15. The fault-tolerant computer system of claim 14, wherein each IO domain comprises at least two switching fabric control components, each switching fabric control component communicating with the switching fabric.

18. 15. The fault-tolerant computer system of claim 14, wherein each CPU node and each IO domain is a modular component, and each modular component, if faulty, can be replaced without affecting a corresponding non-faulty modular component.

19. 15. The fault-tolerant computer system of claim 14, wherein the switching fabric comprises a first switching component and a second switching component, the active IO domain comprises the first switching component, and the second IO domain comprises the second switching component.

Citation Information

Patent Citations

  • System hardware replacement

    JP2010510607A

  • Virtual computer, virtual computer system, and virtual computer control method

    JP2013084122A

  • Methods, apparatus, computer programs, and computer program products for operating a cluster of virtual machines.

    JP2014503904A

  • High Availability Receiving System

    JP2015528962A

  • Tracking modified pages on a computer system

    US20060117300A1