Secondary Node Failover Detection via Application State Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtualization technologies face challenges in ensuring high availability of applications by efficiently detecting and managing failover events between virtual computing resources, particularly in scenarios where a single point of failure occurs, leading to potential service interruptions.

Innovation Solution

The system detects failover events at secondary nodes by evaluating application state information from a monitor, allowing for seamless performance switching without requiring a quorum or external arbitration, thus reducing costs and management complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a quorum or external arbitration mechanism is used to detect failover events, then the reliability of failover detection is improved, but the device complexity and management overhead increase

Engineering Contradiction:
Improvefailover detection reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.

Inventive Principle:
Principle #25Self-service

2Reliability

If a quorum or external arbitration mechanism is used to detect failover events, then the reliability of failover detection is improved, but the management overhead and costs increase

Engineering Contradiction:
Improvefailover detection reliabilityVSAvoidmanagement overhead
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.

Inventive Principle:
Principle #25Self-service

3Device complexity

If failover detection is performed without autonomous evaluation at secondary nodes, then the device complexity is reduced, but the productivity and response time of failover switching deteriorate

Engineering Contradiction:
Improvesystem complexityVSAvoidfailover switching speed
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously monitors application state information in advance, so when a failover event occurs, the secondary node can immediately evaluate the pre-collected state information and switch performance without waiting for external arbitration, thereby improving response time while maintaining simplicity.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If external arbitration mechanisms are used to prevent split-brain scenarios, then the reliability is improved, but the device complexity and response time worsen

Engineering Contradiction:
Improvesplit-brain preventionVSAvoidfailover detection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The secondary node autonomously detects failover events by evaluating application state information from the monitor without requiring external arbitration or quorum mechanisms. The secondary node independently determines whether to switch performance based on the monitored state, eliminating the need for complex coordination systems while maintaining reliable failover detection.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The monitor continuously provides application state information feedback to the secondary node, enabling the secondary node to independently detect failover events and prevent split-brain scenarios by evaluating real-time state information without waiting for external arbitration, thereby reducing detection time while maintaining reliability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10445197B1Detecting failover events at secondary nodes
Publication Date: 2019.10.15 AMAZON TECH INC
  • US10445197B1 patent drawing
  • US10445197B1 patent drawing
  • US10445197B1 patent drawing

AI summary

Secondary nodes may detect failover operations for applications performing at a primary node. Application state indications may be collected at a primary node and reported to a monitor for the primary node. A secondary node may obtain the state indications from the monitor in order to evaluate whether the application is performing correctly at the primary and if not, trigger a failover operation to switch performance of the application to the secondary node. In some embodiments, the state information may be encoded into a single metric that can be obtained by the secondary node and evaluated to detect failover events.