Diagnostic Circuit for Reset Failure Detection in Multinode Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In multinode data processing systems, detecting and responding to server node failures during the reset process is challenging, especially when only one node executes startup code, as traditional diagnostic methods are ineffective, and manual user intervention is often required to manage failed primary nodes and reconfigure the system.
Innovation Solution
A diagnostic circuit is implemented in each server node, coupled to the code fetch chain, which provides diagnostic signals to detect problems before startup code retrieval, allowing for early failure detection and automatic signaling of failed nodes, enabling seamless system reconfiguration without user intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If only one server node executes startup code at reset, then system complexity is reduced and boot process is simplified, but the ability to detect failures in non-executing nodes is lost and manual intervention is required
Solution Approach 1:
The patent applies preliminary action by implementing diagnostic circuits that perform failure detection before the main startup code execution begins. The diagnostic circuits are triggered during the reset process to test code fetch chains in all nodes, including those not designated to execute startup code. This early detection mechanism identifies failures in non-executing nodes before the system proceeds with normal boot operations, eliminating the need for manual intervention while maintaining simplified boot complexity.
2Difficulty of detecting and measuring
If traditional diagnostic methods are used during reset, then failure detection is possible, but they are ineffective when only one node executes startup code
Solution Approach 1:
The patent implements universality by designing diagnostic circuits with multi-functionality that can detect failures in both executing and non-executing nodes. The same diagnostic circuit architecture is used across all server nodes regardless of their execution status. The diagnostic circuits can test code fetch chains, detect startup memory issues, and identify processor problems in any node, providing universal failure detection capability that works effectively whether one or multiple nodes are executing startup code, enabling complete automation of failure response.
3Ease of operation
If manual user intervention is required to manage failed primary nodes, then system reconfiguration can be performed, but system downtime increases and productivity decreases
Solution Approach 1:
The patent applies self-service by implementing automatic failure response mechanisms where the system reconfigures itself without human intervention. When diagnostic circuits detect a failed primary node, the system automatically identifies alternative nodes, promotes a new primary node, and updates system configuration. This self-service capability maintains full system reconfiguration functionality while eliminating manual intervention requirements, thereby maximizing system availability and productivity by reducing downtime during failure events.
Data Source
AI summary
Detection of a reset failure in a multinode data processing system is provided by a diagnostic circuit in each of a plurality of the server nodes of the system. Each diagnostic circuit is coupled to a code fetch chain of its corresponding node. At reset, prior to a node processor retrieving startup code from the code fetch chain, the diagnostic circuit provides diagnostic signals to the code fetch chain. A problem in the code fetch chain is detected from a response to the diagnostic signals. When a problem is detected, a node failure status for the problem node may be signaled to the other nodes. The multinode system may be configured in response to signaled node failure status, such as by dropping failed nodes and replacing a failed primary node with a secondary node if necessary.


