Multi-node Compute Architecture for Fault-Tolerant Server Power

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-core socket systems lack fault tolerance, which is expensive to implement in Industry Standard Servers, and are not suitable for low-cost, low-power computing applications, limiting their use in cost-effective fault-tolerant servers.

Innovation Solution

A multi-node computing architecture with independent compute nodes, each equipped with a processor and memory controller, connected through dedicated I/O buses and a new Southbridge, providing system-level fault tolerance through power redundancy and redundant voltage regulators, allowing each node to access external devices independently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If fault tolerance is implemented in multi-core socket systems, then system reliability is improved, but cost and complexity increase significantly

Engineering Contradiction:
Improvesystem reliabilityVSAvoidhardware complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system is divided into multiple independent compute nodes, each with its own processor and memory controller. This segmentation allows fault isolation where a failure in one node does not affect other nodes, providing system-level fault tolerance without requiring complex redundant hardware across the entire system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each compute node is equipped with dedicated I/O bus connections and local resources, creating heterogeneous quality distribution. This allows each node to have tailored fault tolerance capabilities appropriate to its function, rather than uniformly complex protection across all system components.

Inventive Principle:
Principle #3Local quality

2Reliability

If traditional fault tolerant servers are used, then system reliability is improved, but cost increases

Engineering Contradiction:
Improvefault toleranceVSAvoidsystem cost
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates simplified copies of essential system components at the node level rather than using expensive full system redundancy. Each compute node contains a copy of critical elements (processor, memory controller, I/O interfaces) that can independently operate, providing fault tolerance through functional duplication rather than complete system mirroring.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The multi-node architecture uses individually replaceable compute nodes that can be independently managed and replaced. This allows for more economical fault tolerance compared to expensive traditional fault-tolerant server architectures, as individual nodes can be swapped without replacing entire system components or requiring complex redundant hardware.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Productivity

If multi-core sockets are used, then processing capability is improved, but fault tolerance capability deteriorates

Engineering Contradiction:
Improveprocessing capabilityVSAvoidfault tolerance capability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

Instead of relying on a single multi-core socket, the system segments processing across multiple independent compute nodes. Each node maintains its own processor with full fault tolerance capabilities, allowing the system to maintain high processing capability through parallel nodes while ensuring that a failure in one node does not compromise the entire system.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10108253B2Multiple compute nodes
Publication Date: 2018.10.23 HEWLETT PACKARD ENTERPRISE DEV LP
  • US10108253B2 patent drawing
  • US10108253B2 patent drawing
  • US10108253B2 patent drawing

AI summary

An example apparatus comprises a first compute node including a first processor; a second compute node including a second processor; an input/manta (I/O) interface to selectively couple the first and second compute nodes to a set of I/O resources; and a voltage regulator including a set of power phase circuits, the voltage regulator to operate in a fault tolerant mode to provide power from selected ones of a first portion of the set of power phase circuits to the first compute node and to provide power from selected ones of a second portion of the set of power phase circuits to the second compute node.