Dynamic Node Promotion for Distributed System Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed computing environments, conventional approaches to specifying backup nodes are inflexible and do not provide customers with sufficient control over computing clusters, leading to suboptimal performance during failover scenarios due to potential mismatches in node configurations and incomplete control over operation and configuration.
Innovation Solution
Implementing a system where computing nodes are grouped into ranked subsets or tiers, allowing for the selection of candidate nodes for promotion based on ordinal values and additional provider criteria, such as capacity, location, and role, to ensure seamless failover with minimal performance degradation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If backup nodes are designated exclusively for failover purposes, then system reliability is improved, but resource utilization deteriorates as backup nodes remain unused during normal operation
Solution Approach 1:
The patent allows computing nodes to serve multiple roles dynamically. Nodes can function as primary nodes during normal operation and automatically transition to backup roles when needed, or serve as shared backup resources for multiple primaries. This multi-functionality resolves the contradiction by enabling backup nodes to contribute to resource pool utilization while maintaining failover capability.
Solution Approach 2:
The system implements dynamic role assignment where computing nodes can transition between primary, backup, and standalone roles based on operational conditions. This dynamic flexibility allows the system to optimize resource utilization by activating backup nodes for additional workloads when primaries are healthy, while ensuring immediate failover capability when needed.
2Ease of manufacture
If fixed or hard-wired approaches are used to specify backup nodes, then configuration simplicity is improved, but adaptability deteriorates in service-oriented computing environments where modifications occur over time
Solution Approach 1:
The patent replaces fixed backup node specifications with dynamic, software-defined role assignments. Computing nodes are assigned roles (primary, backup, standalone) that can be modified through configuration files or management interfaces without hardware changes. This enables the system to adapt to changing customer requirements and provider modifications while maintaining configuration simplicity through standardized role-based management.
Solution Approach 2:
The system uses configurable parameters and identifiers to define node roles and relationships. By changing software parameters rather than hard-wired connections, the system can easily adapt backup node assignments, promotion criteria, and failover behavior to match evolving service requirements while maintaining ease of configuration through parameter-based control.
3Ease of manufacture
If conventional fixed approaches are used for backup node specification, then implementation simplicity is improved, but customer control deteriorates over operation and configuration
Solution Approach 1:
The patent segments control authority between customers and providers through configurable role assignments. Customers can define their own promotion tiers, backup node selections, and failover criteria through configuration files, while providers implement the logic automatically. This segmentation enables customers to maintain control over their computing clusters' operation and configuration while keeping implementation simple through automated role-based management.
Solution Approach 2:
The system incorporates monitoring and feedback mechanisms that track node health, role assignments, and failover events. This feedback enables customers to observe and adjust their cluster configurations, ensuring they maintain appropriate control over operation and configuration while the system automatically manages the complexity of role transitions and failover execution.
Data Source
AI summary
A system and method for failover in a distributed system may comprise a computing device that receives client-provided information that groups computing nodes into ordered subsets. The subsets, or nodes in the subsets, may be associated with client-provided instructions for evaluating the health of a node. A node may be selected for failover based on executing the instructions and evaluating associated performance metrics. When a node is selected for failover, a replacement node may be selected based on the ordering of the subsets and the health of candidate nodes as determined based on executing the client-provided instructions.


