Multi-node Compute Architecture for Fault-Tolerant Server Power
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-core socket systems lack fault tolerance, which is expensive to implement in Industry Standard Servers, and are not suitable for low-cost, low-power computing applications, limiting their use in cost-effective fault-tolerant servers.
Innovation Solution
A multi-node computing architecture with independent compute nodes, each equipped with a processor and memory controller, connected through dedicated I/O buses and a new Southbridge, providing system-level fault tolerance through power redundancy and redundant voltage regulators, allowing each node to access external devices independently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If fault tolerance is implemented in multi-core socket systems, then system reliability is improved, but cost and complexity increase significantly
Solution Approach 1:
The system is divided into multiple independent compute nodes, each with its own processor and memory controller. This segmentation allows fault isolation where a failure in one node does not affect other nodes, providing system-level fault tolerance without requiring complex redundant hardware across the entire system.
Solution Approach 2:
Each compute node is equipped with dedicated I/O bus connections and local resources, creating heterogeneous quality distribution. This allows each node to have tailored fault tolerance capabilities appropriate to its function, rather than uniformly complex protection across all system components.
2Reliability
If traditional fault tolerant servers are used, then system reliability is improved, but cost increases
Solution Approach 1:
The patent creates simplified copies of essential system components at the node level rather than using expensive full system redundancy. Each compute node contains a copy of critical elements (processor, memory controller, I/O interfaces) that can independently operate, providing fault tolerance through functional duplication rather than complete system mirroring.
Solution Approach 2:
The multi-node architecture uses individually replaceable compute nodes that can be independently managed and replaced. This allows for more economical fault tolerance compared to expensive traditional fault-tolerant server architectures, as individual nodes can be swapped without replacing entire system components or requiring complex redundant hardware.
3Productivity
If multi-core sockets are used, then processing capability is improved, but fault tolerance capability deteriorates
Solution Approach 1:
Instead of relying on a single multi-core socket, the system segments processing across multiple independent compute nodes. Each node maintains its own processor with full fault tolerance capabilities, allowing the system to maintain high processing capability through parallel nodes while ensuring that a failure in one node does not compromise the entire system.
Data Source
AI summary
An example apparatus comprises a first compute node including a first processor; a second compute node including a second processor; an input/manta (I/O) interface to selectively couple the first and second compute nodes to a set of I/O resources; and a voltage regulator including a set of power phase circuits, the voltage regulator to operate in a fault tolerant mode to provide power from selected ones of a first portion of the set of power phase circuits to the first compute node and to provide power from selected ones of a second portion of the set of power phase circuits to the second compute node.


