Autonomous Cluster Control Plane for Virtualization Server Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtualization management servers in software-defined datacenters act as a single point of failure, leading to unplanned downtime and planned upgrades, which disrupt the availability of control and management planes for host clusters.
Innovation Solution
Implementing a cross-cluster control plane (xCCP) with an infravisor on a seed host to create an autonomous cluster, where the infravisor runs a cluster control plane (CCP) pod, applies configuration to create a single-node cluster, and extends it with additional nodes, ensuring redundancy and resilience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a centralized virtualization management server is used to manage host clusters, then infrastructure management and control is simplified and centralized, but the system becomes vulnerable to single point of failure and experiences downtime during server failures or upgrades
Solution Approach 1:
The patent segments the centralized control plane into distributed cluster control plane pods that run on individual host clusters. Each pod can independently manage its host cluster, eliminating the single point of failure. The segmentation allows control functions to be distributed across multiple nodes while maintaining centralized coordination through the virtualization management server when available.
Solution Approach 2:
The system changes the operational parameter of the control plane from centralized-only to distributed-state capable. The cluster control plane pods can operate in different modes: fully centralized when the virtualization management server is available, or autonomously distributed when it fails. This parameter change enables the system to maintain reliability while preserving ease of operation.
2Reliability
If the virtualization management server is upgraded or maintained, then system security and performance are improved, but control and management planes are lost causing planned downtime
Solution Approach 1:
The patent implements preliminary action by pre-configuring cluster control plane pods on host clusters before the virtualization management server needs to be upgraded. These pods are ready to take over control functions immediately when the server becomes unavailable, preventing downtime. The preliminary setup ensures that failover capability exists before it is needed.
Solution Approach 2:
The system provides beforehand cushioning by creating redundant control plane capacity in the form of cluster control plane pods. This redundant capacity acts as a cushion that absorbs the impact of virtualization management server unavailability during upgrades or maintenance, ensuring continuous operation without interruption to host cluster management.
3Reliability
If control plane functions are distributed across multiple hosts, then single point of failure is mitigated and uptime is improved, but system complexity increases
Solution Approach 1:
The cluster control plane pods are designed with universality to perform multiple functions: they can operate as standby control planes, actively manage host clusters, or coordinate with the virtualization management server. This multi-functionality reduces the need for separate specialized components, managing complexity while providing distributed control capability for improved reliability.
Solution Approach 2:
The cluster control plane pods are self-service capable, automatically detecting when the virtualization management server is unavailable and taking over control functions without manual intervention. They can also automatically fail back when the server returns. This self-service behavior reduces operational complexity despite the distributed architecture.
Data Source
AI summary
An example method of creating an autonomous cluster of hosts in a virtualized computing system includes: enabling, by a virtualization management server executing a cross cluster control plane (xCCP), an infravisor in a seed host of the hosts, the infravisor a component of a hypervisor executing on the seed host; running, by the infravisor, a cluster control plane (CCP) pod on the seed host executing a CCP; providing, by the infravisor, a CCP configuration to the CCP pod; applying, by an initialization script of the CCP pod, the CCP configuration to the CCP to create the autonomous cluster having the seed host as a single node thereof; and extending the autonomous cluster with remaining hosts of the hosts other than the seed host as additional nodes thereof, the CCP applying a cluster personality to each of the remaining hosts derived from the seed host.


