Fault-Tolerant Computer Midplane for Direct State Failover

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Contemporary computing systems with high availability requirements face challenges in efficiently transferring processor and memory state information and device hierarchy from a failing compute node to a standby node without complex software and hardware coordination.

Innovation Solution

A Smart Exchange protocol that enables direct transfer of CPU state, memory state, and device hierarchy from an active compute node to a standby node using modified switch firmware and DMA, allowing for a seamless failover process without external processor coordination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If complex software and hardware coordination is used to transfer processor and memory state information from a failing compute node to a standby node, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improvefault toleranceVSAvoidsoftware and hardware coordination
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The compute nodes autonomously perform state transfer and failover operations without requiring external processor coordination. The standby node automatically assumes the role of the failing node by directly receiving and applying state information, eliminating the need for complex centralized control software and hardware.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent extracts the coordination function from external processors and embeds it directly into the compute nodes themselves. Each node contains the necessary logic to independently initiate and complete failover operations, removing the dependency on complex external coordination mechanisms.

Inventive Principle:
Principle #2Taking out (Extraction)

2Device complexity

If direct transfer of CPU state and memory state is implemented without external processor coordination, then device complexity is reduced, but reliability may be compromised

Engineering Contradiction:
Improvesoftware and hardware coordinationVSAvoidfault tolerance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces an intermediary state transfer mechanism that enables direct communication between compute nodes. This intermediary layer facilitates reliable state transfer and failover operations while maintaining simplicity in the overall system architecture, avoiding the need for complex external processor coordination.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The standby node is pre-configured with the capability to assume the role of the active node. State information is transferred in advance or in real-time, ensuring that when failover is needed, the standby node is already prepared to take over immediately, maintaining reliability without complex coordination.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If comprehensive state information is transferred during failover, then reliability is improved, but loss of time increases

Engineering Contradiction:
Improvefault toleranceVSAvoidfailover time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs periodic state synchronization between active and standby nodes, ensuring that state information is continuously updated. This allows for faster failover times while maintaining comprehensive state transfer, as the standby node already has recent state information from periodic updates rather than requiring complete state transfer from scratch.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The state transfer process operates continuously or near-continuously, maintaining an up-to-date copy of system state on the standby node. This continuous action ensures that when failover occurs, the transition is rapid and seamless, minimizing time loss while ensuring complete state information is available.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12475006B2Cost reduced high reliability fault tolerant computer architecture
Publication Date: 2025.11.18 STRATUS TECH IRELAND LTD
  • US12475006B2 patent drawing
  • US12475006B2 patent drawing
  • US12475006B2 patent drawing

AI summary

In part, in one aspect, the disclosure relates to a first computer system including a first processor and first memory, a first IO storage subsystem including a first switch configured for one or more first storage devices, a first IO non-storage subsystem including a first which configured for one or more first non-storage devices, a second compute system including a second processor and second memory, a second storage IO subsystem including a second switch configured for one or more second storage devices, a second IO non-storage subsystem including a second switch configured for one or more second non-storage devices and a midplane including a power connector, a processor side and an IO side, wherein the processing side includes connectors in electrical communication with the computer systems, the IO side includes connectors in electrical communication with the storage and non-storage subsystems.