Management Node Failover via Detection and Reversal Device
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is a challenge in providing a cost-effective management node failover mechanism for high reliability systems, particularly in computing data centers, where ensuring seamless operation in case of management node failures is critical but not adequately addressed.
Innovation Solution
A system with two management nodes, one active and one passive, connected through a detection and reversal device that determines their roles during startup and switches roles when the active node fails, using heartbeat signals and probe signals to confirm status and initiate failover, with both nodes having identical hardware and software components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a management node failover mechanism is implemented, then system reliability is improved, but device complexity increases
Solution Approach 1:
The system segments management functions into two separate management nodes (primary and secondary), each capable of independent operation. This segmentation allows the system to maintain reliability through distribution while keeping each individual node's complexity manageable and identical to the other.
Solution Approach 2:
The failover mechanism performs preliminary actions by pre-configuring both management nodes with identical hardware and software components before any failure occurs. The secondary node is prepared in advance to immediately assume the primary node's responsibilities, eliminating the need for complex real-time reconfiguration during failure events.
2Quantity of substance
If identical hardware and software components are used for both nodes, then cost is reduced, but adaptability decreases
Solution Approach 1:
The system uses parameter changes to differentiate between identical hardware and software components. By changing operational parameters (primary/secondary role assignment, active/passive state) rather than physical characteristics, the system achieves adaptability while maintaining hardware uniformity and cost-effectiveness.
Solution Approach 2:
Both management nodes are designed with universal, identical hardware and software components that can perform any management function. This universality allows either node to assume any role, providing adaptability through configurability rather than hardware diversity, thus reducing costs.
3Reliability
If heartbeat signals and probe signals are used for status detection, then reliability of failover is improved, but use of energy increases
Solution Approach 1:
The system employs periodic heartbeat signals from the primary node to the secondary node, rather than continuous communication. This periodic action maintains reliable status detection while significantly reducing energy consumption compared to constant monitoring, as signals are transmitted only at scheduled intervals when status changes are expected.
Data Source
AI summary
Aspects of the disclosure relate to management node failover systems and methods. The system includes two management devices and a detection and reversal device. Each of the two management devices has a processor and a non-volatile memory storing computer executable code. The two management devices function respectively as an active node and a passive node. The detection and reversal device monitors status of the active node. When the active node fails, the detection and reversal device sends an activation signal to the passive node. The passive node, in response to receiving the active signal, switches from the passive node to the active node.


