Cloud Deployment Failure Escalation Policy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cloud high availability solutions focus on passive monitoring and restarting virtual machines upon failure, lacking escalation mechanisms to higher level components in bare-metal systems, which limits recovery and availability.

Innovation Solution

Implementing a failure escalation policy in cloud computing environments, where health monitors detect failures in virtual machines, applications, and collections, and initiate relocation to a secondary cloud upon predefined failure criteria, enabling proactive recovery and high availability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If passive monitoring and restarting virtual machines is used, then simple failure recovery is achieved, but high availability is insufficient and service continuity is compromised

Engineering Contradiction:
Improvehigh availabilityVSAvoidfailure recovery mechanism
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The failure recovery mechanism is segmented into multiple hierarchical levels (application level, virtual machine level, cloud level) with distinct recovery actions at each level. This allows targeted recovery strategies that improve overall reliability without requiring complex unified management of all failure scenarios.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-defining escalation policies and maintaining backup clouds ready for migration. When failures occur, the predefined recovery paths and pre-positioned resources enable rapid response without requiring complex real-time decision-making, thus improving availability while keeping management manageable.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If continuous restarting of virtual machines is performed, then failure recovery is attempted, but service disruption increases and productivity decreases

Engineering Contradiction:
Improvefailure recoveryVSAvoidservice continuity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

Backup clouds are prepared in advance with pre-configured resources and escalation policies. When a virtual machine fails repeatedly, the system can immediately migrate to the pre-prepared backup cloud without continuous restarting, thereby recovering services faster and maintaining productivity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system skips the continuous restarting cycle by implementing a hard stop after a predefined number of failure attempts. Instead of endlessly restarting, the system escalates to higher-level recovery actions or permanent shutdown, reducing service disruption time and improving productivity.

Inventive Principle:
Principle #21Skipping (Rushing through)

3Reliability

If no escalation mechanism is implemented, then simple monitoring is maintained, but recovery options are limited and availability is insufficient

Engineering Contradiction:
Improverecovery optionsVSAvoidmonitoring system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The monitoring and recovery system is segmented into hierarchical levels (application, virtual machine, cloud) with dedicated monitors and recovery mechanisms at each level. This modular approach expands recovery options without requiring a monolithic complex system, as each level can operate independently with its own simplified monitoring and recovery logic.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback loops where failure detection at one level triggers escalation to higher levels. This feedback mechanism provides structured recovery options while keeping individual monitoring components relatively simple, as each level only needs to monitor its own state and respond according to predefined policies.

Inventive Principle:
Principle #23Feedback

4Reliability

If manual intervention is required for repeated failures, then operational control is maintained, but response time increases and productivity is reduced

Engineering Contradiction:
Improvefailure managementVSAvoidresponse time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-service by automatically executing escalation policies and initiating recovery actions without requiring manual intervention. When failures occur, the system autonomously migrates workloads to backup clouds or shuts down persistent failures according to predefined policies, thereby reducing response time and improving productivity while maintaining reliable failure management.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Escalation policies and recovery actions are predefined in advance, so when failures occur, the system can immediately execute the appropriate response without waiting for manual decisions. This preliminary configuration of recovery strategies eliminates delays associated with manual intervention while ensuring consistent and reliable failure management.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9081750B2Recovery escalation of cloud deployments
Publication Date: 2015.07.14 RED HAT INC
  • US9081750B2 patent drawing
  • US9081750B2 patent drawing
  • US9081750B2 patent drawing

AI summary

Methods and systems for escalating component failures in a cloud are provided. A cloud controller of a cloud receives an indication that a collection of virtual machines of the first cloud has failed based on a collection of virtual machines escalation policy. The cloud controller initiates relocating the collection of virtual machines to a second cloud.