Compute Node Grouping for Stretched Cluster Fault Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current stretched cluster environments face challenges in fault recovery and high availability, particularly in geographically distant multisite recovery scenarios, as they lack automatic node grouping capabilities, leading to increased complexity and downtime due to manual configuration and dependency on a single virtualization management server.

Innovation Solution

The system automatically groups compute nodes based on user-configured network parameters and measured round-trip times using a virtual infrastructure management server, enabling load balancing, fault tolerance, and disaster recovery operations across multiple sites without administrative intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual configuration is used for stretched cluster environments, then fault recovery and high availability can be achieved, but system complexity and downtime increase

Engineering Contradiction:
Improvefault recoveryVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically groups compute nodes based on network parameters and round-trip time measurements without requiring manual administrative intervention. The virtual infrastructure management server autonomously performs the grouping operation, eliminating the need for manual configuration while maintaining fault recovery capabilities.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-configures compute node groupings based on network characteristics before failures occur. By establishing optimal groupings in advance through automated measurement and configuration, the system prepares the infrastructure for rapid fault recovery without requiring complex manual reconfiguration during incidents.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual configuration is used for compute node grouping, then fault recovery can be implemented, but downtime increases

Engineering Contradiction:
Improvefault recoveryVSAvoiddowntime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The virtual infrastructure management server automatically performs compute node grouping without waiting for manual administrative actions. This self-service automation eliminates configuration delays and reduces downtime by immediately establishing optimal groupings when needed.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs grouping operations in advance and maintains configured groupings ready for immediate use during failures. This preliminary configuration ensures that fault recovery can proceed without time-consuming manual setup, reducing overall downtime.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If a single virtualization management server is used, then stretched cluster operation is simplified, but fault tolerance capability is reduced

Engineering Contradiction:
Improveoperation simplicityVSAvoidfault tolerance
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system segments the management function by separating the virtual infrastructure management server from the compute node grouping operation. The management server provides centralized control and simplicity, while the automated grouping mechanism distributed across compute nodes provides fault tolerance through redundancy and independence from single-point failures.

Inventive Principle:
Principle #1Segmentation

4Ease of operation

If automatic node grouping is implemented, then operational complexity is reduced, but measurement and configuration difficulty increases

Engineering Contradiction:
Improveoperational complexityVSAvoidmeasurement complexity
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The system uses round-trip time measurements as feedback to automatically determine optimal compute node groupings. The virtual infrastructure management server measures network characteristics, processes this feedback information, and automatically configures groupings based on the measurements, eliminating manual complexity while handling the measurement burden automatically.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10833918B2Automatic rule based grouping of compute nodes for a globally optimal cluster
Publication Date: 2020.11.10 VMWARE INC
  • US10833918B2 patent drawing
  • US10833918B2 patent drawing
  • US10833918B2 patent drawing

AI summary

Techniques for automatic rule based grouping of compute nodes for a global optimal cluster are disclosed. In one embodiment, a virtual infrastructure management (VIM) server may obtain a list of operating, provisioned and/or about to be provisioned compute nodes in a cluster. The VIM server may then obtain user configured and run-time network parameters of a network interface card (NIC) of each compute node in the list. Further, the VIM server may then measure round-trip times (RTTs), using a ping, between each compute node and each of remaining compute nodes in the list. Furthermore, the VIM server may then group the compute nodes in the list based on the obtained user configured and run-time network parameters and/or the measured RTTs. In addition, the VIM server may perform a high availability (HA) operation in the cluster using the grouped compute nodes.