Arbitrated Failover Control for Multi-Node Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multi-node network clusters face challenges in efficient incident detection and recovery due to premature or improper failover processes, leading to resource wastage and potential catastrophic network failures, especially in two-node active/passive high availability clusters where split-brain scenarios and improper communication between nodes can occur.

Innovation Solution

The implementation of an arbitration and countermeasures system that monitors node health, applies automated remediation, and includes a failover countermeasure process, which involves external monitoring, alert creation, and approval-based remediation to ensure seamless transition of roles between nodes, thereby preventing split-brain scenarios and optimizing resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If failover is performed frequently to ensure network availability, then reliability is improved, but resource wastage increases due to premature failovers

Engineering Contradiction:
Improvenetwork availabilityVSAvoidresource wastage
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system performs preliminary health checks and incident detection before triggering failover. The arbitration system monitors node status and only initiates failover when actual failures are confirmed, preventing premature failovers while maintaining network availability when truly needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements continuous feedback mechanisms through health monitoring and status reporting between nodes. This feedback loop allows the arbitration system to make informed decisions about failover timing, ensuring reliability improvements without unnecessary resource consumption from premature failovers.

Inventive Principle:
Principle #23Feedback

2Reliability

If automated failover control is implemented to prevent split-brain scenarios, then reliability is improved, but device complexity increases due to arbitration mechanisms

Engineering Contradiction:
Improvefailure preventionVSAvoidcontrol mechanism complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an arbitration system as an intermediary component that mediates failover decisions between master and slave nodes. This intermediary prevents split-brain scenarios by centrally coordinating failover control, improving reliability while concentrating complexity in a dedicated management layer rather than distributing it across all nodes.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements self-service mechanisms where nodes automatically report their health status and the arbitration system automatically makes failover decisions based on monitored conditions. This automation prevents split-brain scenarios through self-coordinating behavior, improving reliability while the standardized self-service protocols keep implementation complexity manageable.

Inventive Principle:
Principle #25Self-service

3Reliability

If health monitoring and arbitration systems are added to prevent improper failover, then reliability is improved, but ease of operation deteriorates due to additional monitoring requirements

Engineering Contradiction:
Improvefailover accuracyVSAvoidsystem operation simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The health monitoring and arbitration systems operate autonomously without requiring manual intervention. Nodes automatically monitor their own health status, report to the arbitration system, and execute failover decisions automatically. This self-service approach improves failover accuracy while maintaining ease of operation by eliminating the need for operators to manually configure or trigger failover processes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12047226B2Systems and methods for arbitrated failover control using countermeasures
Publication Date: 2024.07.23 FORTINET INC
  • US12047226B2 patent drawing
  • US12047226B2 patent drawing
  • US12047226B2 patent drawing

AI summary

Various approaches for multi-node network cluster systems and methods. In some cases systems and methods for incident detection and/or recovery in multi-node processors are discussed.