Reserving Failover Capacity in Cloud Data Centers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud computing architectures face challenges in ensuring fault tolerance and process availability across multiple data centers, as existing high availability modules fail to reserve adequate failover capacity across different availability zones, leading to difficulties in seamlessly failing over processes between data centers in case of failures.

Innovation Solution

A method is introduced to determine if a management process is executing at a data center, and if not, initiate a host at a different data center and execute the management process there, ensuring reserved failover capacity by allocating necessary resources, such as processing and storage, to maintain system stability and availability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If failover capacity is reserved across multiple data centers, then fault tolerance and process availability are improved, but device complexity and resource allocation complexity increase

Engineering Contradiction:
Improvefault toleranceVSAvoidresource allocation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically detects management process failures and initiates failover to backup data centers without human intervention. The automated detection and activation mechanisms enable the system to self-manage failover capacity, reducing operational complexity while maintaining improved fault tolerance across distributed data centers.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Backup management processes are pre-positioned in standby mode at alternative data centers before failures occur. This preliminary preparation ensures that failover capacity is already reserved and configured, eliminating the need for complex real-time resource allocation during actual failover events and simplifying the overall system architecture.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If resources are allocated for failover capacity, then process availability is improved, but loss of energy and operational costs increase

Engineering Contradiction:
Improveprocess availabilityVSAvoidoperational costs
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system dynamically adjusts the activation state of backup management processes based on real-time monitoring of primary process health. Backup processes remain in a low-power standby state during normal operation and only activate when failures are detected, optimizing the balance between process availability and energy consumption by avoiding continuous full operation of redundant resources.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system continuously monitors the execution status of management processes and uses this feedback to control the activation of backup processes. This feedback mechanism ensures that failover capacity is only fully utilized when actually needed, reducing unnecessary energy consumption and operational costs while maintaining process availability guarantees.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11755432B2Reserving failover capacity in cloud computing
Publication Date: 2023.09.12 VMWARE INC
  • US11755432B2 patent drawing
  • US11755432B2 patent drawing
  • US11755432B2 patent drawing

AI summary

Methods and devices for providing reserved failover capacity across a plurality of data centers are described herein. An exemplary method includes determining whether a management process is executing at a first data center corresponding to a first physical location. In accordance with a determination that the management process is not executing at the first data center corresponding to the first physical location a host is initiated at a second data center corresponding to a second physical location and the management process is executed on the initiated host at the second data center corresponding to the second physical location.