Secondary Node Failover Detection via Application State Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtualization technologies face challenges in ensuring high availability of applications by efficiently detecting and managing failover events between virtual computing resources, particularly in scenarios where a single point of failure occurs, leading to potential service interruptions.
Innovation Solution
The system detects failover events at secondary nodes by evaluating application state information from a monitor, allowing for seamless performance switching without requiring a quorum or external arbitration, thus reducing costs and management complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a quorum or external arbitration mechanism is used to detect failover events, then the reliability of failover detection is improved, but the device complexity and management overhead increase
Solution Approach 1:
The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.
2Reliability
If a quorum or external arbitration mechanism is used to detect failover events, then the reliability of failover detection is improved, but the management overhead and costs increase
Solution Approach 1:
The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.
3Device complexity
If failover detection is performed without autonomous evaluation at secondary nodes, then the device complexity is reduced, but the productivity and response time of failover switching deteriorate
Solution Approach 1:
The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.
Solution Approach 2:
The system continuously monitors application state information in advance, so when a failover event occurs, the secondary node can immediately evaluate the pre-collected state information and switch performance without waiting for external arbitration, thereby improving response time while maintaining simplicity.
4Reliability
If external arbitration mechanisms are used to prevent split-brain scenarios, then the reliability is improved, but the device complexity and response time worsen
Solution Approach 1:
The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.
Solution Approach 2:
The monitor continuously provides application state information feedback to the secondary node, enabling the secondary node to independently detect failover events and prevent split-brain scenarios by evaluating real-time state information without waiting for external arbitration, thereby reducing detection time while maintaining reliability.
Data Source
AI summary
Secondary nodes may detect failover operations for applications performing at a primary node. Application state indications may be collected at a primary node and reported to a monitor for the primary node. A secondary node may obtain the state indications from the monitor in order to evaluate whether the application is performing correctly at the primary and if not, trigger a failover operation to switch performance of the application to the secondary node. In some embodiments, the state information may be encoded into a single metric that can be obtained by the secondary node and evaluated to detect failover events.


