Failover Scopes for Cluster Node Resource Groups
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current clustering technologies lack flexibility in managing failover actions across geographically separated nodes, leading to inefficient resource allocation and increased costs due to indiscriminate failover mechanisms that do not distinguish between physically close and distant nodes, limiting administrators' ability to configure failover policies as desired.
Innovation Solution
The introduction of failover scopes, which are defined subsets of nodes within a cluster, allowing resource groups to be associated with ordered lists of scopes for controlled failover, enabling administrators to specify manual or automatic transitions between scopes and constrain node hosting for resource groups, thereby optimizing failover processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If resource groups automatically failover to any surviving node using random selection, then failover availability is improved, but node overload occurs and failover control flexibility deteriorates
Solution Approach 1:
The patent segments the cluster nodes into different geographic sites, creating distinct failover domains. Each site is treated as a separate segment with its own failover scope, allowing administrators to control failover behavior at the site level rather than treating the entire cluster uniformly. This segmentation enables differentiated failover policies for different geographic locations.
Solution Approach 2:
The patent applies local quality by allowing different failover scopes to be defined for different geographic sites. Each site can have its own failover scope configuration, enabling local customization of failover behavior. Administrators can specify which sites are included in each failover scope, creating location-specific failover policies that match geographic and network characteristics.
2Reliability
If resource groups failover across geographically distant nodes, then disaster protection is improved, but communication bandwidth and failover cost increase
Solution Approach 1:
The patent implements preliminary action by pre-defining failover scopes that include only specific geographic sites before failures occur. Administrators configure which sites belong to each failover scope in advance, establishing failover boundaries before disasters strike. This preliminary configuration prevents automatic failover to distant sites unless explicitly included in the scope, avoiding unnecessary cross-site failover costs.
Solution Approach 2:
The patent applies local quality by creating site-specific failover scopes that prioritize local failover within the same geographic location. Each scope can be configured to include only nearby sites, ensuring that failover occurs within acceptable bandwidth and cost parameters while still providing disaster protection through geographic distribution.
3Device complexity
If multiple resource groups failover to the same preferred node, then failover simplicity is improved, but node capacity is overwhelmed
Solution Approach 1:
The patent segments resource groups into different failover scopes based on geographic site membership. By organizing resource groups according to the sites they belong to, the system distributes failover targets across multiple sites rather than concentrating them on single nodes. This segmentation naturally load-balances failover traffic while maintaining simple scope-based configuration.
4Ease of operation
If administrators manually assess and fix failures before failover, then failover control precision is improved, but failover time increases
Solution Approach 1:
The patent implements dynamics by making failover scope configuration adjustable and adaptable. Administrators can dynamically modify which sites are included in failover scopes based on changing operational requirements, disaster scenarios, and site availability. This dynamic configuration allows precise control over failover behavior while adapting to different failure scenarios without requiring complete manual intervention.
Data Source
AI summary
A failover scope comprises a node collection in a computer cluster. A resource group (e.g., application program) is associated with one or more failover scopes. If a node fails, its hosted resource groups only failover to nodes identified in each resource group's associated failover scope(s), beginning with a first associated failover scope, in order, thereby defining an island of nodes within which a resource group can failover. If unable to failover to a node of a resource group's first failover scope, failover is attempted to a node represented in any next associated failover scope, which may require manual intervention. Failover scopes may represent geographic sites, whereby each resource group attempts to failover to nodes within its site before failing over to another site. Failover scopes may be managed by the cluster runtime automatically, e.g., an added node is detectable as belonging to a site represented by a failover scope.


