Fault-Tolerant Compute Node Failover via Smart State Exchange

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Contemporary computing systems with high availability requirements face challenges in efficiently transferring processor and memory state information and device hierarchy from a failing compute node to a standby node without complex software and hardware coordination.

Innovation Solution

A Smart Exchange protocol that enables the transfer of CPU state, memory state, and device hierarchy from an active to a standby compute node using direct memory access (DMA) and host-to-host messaging, with modified switch firmware facilitating device reprovisioning and synchronization of compute nodes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If complex software and hardware coordination is used to transfer processor and memory state information from a failing compute node to a standby node, then the reliability of the failover process is improved, but the device complexity and cost increase

Engineering Contradiction:
Improvefailover reliabilityVSAvoidsoftware and hardware coordination complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediate buffer memory and control logic that mediates the state transfer between compute nodes. The buffer memory serves as an intermediary storage element that temporarily holds processor and memory state information during the failover process, eliminating the need for complex real-time coordination software and hardware mechanisms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a copying mechanism where the state of the failing compute node (processor state, memory state, device hierarchy) is replicated to the standby compute node. This copying approach simplifies the failover process by creating a ready-to-use duplicate state rather than requiring complex state synchronization and coordination protocols.

Inventive Principle:
Principle #26Copying

2Reliability

If external processor coordination is used to manage failover between compute nodes, then the fault tolerance is improved, but the hardware complexity and cost increase

Engineering Contradiction:
Improvefault toleranceVSAvoidexternal processor coordination hardware
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent enables compute nodes to perform failover operations autonomously without requiring external processor coordination. Each compute node has self-contained control logic and buffer memory that allows it to detect failures, initiate failover sequences, and transfer state information independently, eliminating the need for external coordination hardware.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent divides the failover system into independent, self-sufficient compute nodes, each with its own buffer memory and control logic. This segmentation allows each node to operate autonomously and perform failover operations independently, reducing the need for complex external coordination hardware while maintaining fault tolerance.

Inventive Principle:
Principle #1Segmentation

3Reliability

If comprehensive state information (processor state, memory state, device hierarchy) is transferred during failover, then the reliability of service continuity is improved, but the time required for failover increases

Engineering Contradiction:
Improveservice continuityVSAvoidfailover time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-allocating buffer memory at each compute node and maintaining device hierarchy information in a ready-to-transfer format. This preparation ensures that when failover is needed, the state information can be quickly copied and transferred without delays associated with real-time data collection and formatting.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses efficient copying mechanisms to transfer processor state, memory state, and device hierarchy information. By maintaining pre-formatted state information and using direct memory copy operations, the system achieves fast state transfer that minimizes failover time while ensuring complete service continuity.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260056855A1Cost reduced high reliability fault tolerant computer architecture
Publication Date: 2026.02.26 STRATUS TECH IRELAND LTD
  • US20260056855A1 patent drawing
  • US20260056855A1 patent drawing
  • US20260056855A1 patent drawing

AI summary

In part, in one aspect, the disclosure relates to a first computer system including a first processor and first memory, a first IO storage subsystem including a first switch configured for one or more first storage devices, a first IO non-storage subsystem including a first witch configured for one or more first non-storage devices, a second compute system including a second processor and second memory, a second storage IO subsystem including a second switch configured for one or more second storage devices, a second IO non-storage subsystem including a second switch configured for one or more second non-storage devices and a midplane including a power connector, a processor side and an IO side, wherein the processing side includes connectors in electrical communication with the computer systems, the IO side includes connectors in electrical communication with the storage and non-storage subsystems.