Compute Node Grouping for Stretched Cluster Fault Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current stretched cluster environments face challenges in fault recovery and high availability, particularly in geographically distant multisite recovery scenarios, as they lack automatic node grouping capabilities, leading to increased complexity and downtime due to manual configuration and dependency on a single virtualization management server.
Innovation Solution
The system automatically groups compute nodes based on user-configured network parameters and measured round-trip times using a virtual infrastructure management server, enabling load balancing, fault tolerance, and disaster recovery operations across multiple sites without administrative intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual configuration is used for stretched cluster environments, then fault recovery and high availability can be achieved, but system complexity and downtime increase
Solution Approach 1:
The system automatically groups compute nodes based on network parameters and round-trip time measurements without requiring manual administrative intervention. The virtual infrastructure management server autonomously performs the grouping operation, eliminating the need for manual configuration while maintaining fault recovery capabilities.
Solution Approach 2:
The system pre-configures compute node groupings based on network characteristics before failures occur. By establishing optimal groupings in advance through automated measurement and configuration, the system prepares the infrastructure for rapid fault recovery without requiring complex manual reconfiguration during incidents.
2Reliability
If manual configuration is used for compute node grouping, then fault recovery can be implemented, but downtime increases
Solution Approach 1:
The virtual infrastructure management server automatically performs compute node grouping without waiting for manual administrative actions. This self-service automation eliminates configuration delays and reduces downtime by immediately establishing optimal groupings when needed.
Solution Approach 2:
The system performs grouping operations in advance and maintains configured groupings ready for immediate use during failures. This preliminary configuration ensures that fault recovery can proceed without time-consuming manual setup, reducing overall downtime.
3Ease of operation
If a single virtualization management server is used, then stretched cluster operation is simplified, but fault tolerance capability is reduced
Solution Approach 1:
The system segments the management function by separating the virtual infrastructure management server from the compute node grouping operation. The management server provides centralized control and simplicity, while the automated grouping mechanism distributed across compute nodes provides fault tolerance through redundancy and independence from single-point failures.
4Ease of operation
If automatic node grouping is implemented, then operational complexity is reduced, but measurement and configuration difficulty increases
Solution Approach 1:
The system uses round-trip time measurements as feedback to automatically determine optimal compute node groupings. The virtual infrastructure management server measures network characteristics, processes this feedback information, and automatically configures groupings based on the measurements, eliminating manual complexity while handling the measurement burden automatically.
Data Source
AI summary
Techniques for automatic rule based grouping of compute nodes for a global optimal cluster are disclosed. In one embodiment, a virtual infrastructure management (VIM) server may obtain a list of operating, provisioned and/or about to be provisioned compute nodes in a cluster. The VIM server may then obtain user configured and run-time network parameters of a network interface card (NIC) of each compute node in the list. Further, the VIM server may then measure round-trip times (RTTs), using a ping, between each compute node and each of remaining compute nodes in the list. Furthermore, the VIM server may then group the compute nodes in the list based on the obtained user configured and run-time network parameters and/or the measured RTTs. In addition, the VIM server may perform a high availability (HA) operation in the cluster using the grouped compute nodes.


