Fault Recovery Manager for Cloud Service Priority Reallocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cloud computing systems lack effective policies for managing resource allocation during large and unusual failures, such as natural disasters or acts of terrorism, which exceed the reserve hardware capacity, leading to incomplete recovery of services and potential cascading failures.

Innovation Solution

Implementing a fault recovery manager that dynamically reassesses and reallocates computational resources from lower priority services to higher priority ones, ensuring minimal availability thresholds are maintained without sacrificing operability, and dedicating restored capacity to reserve against further failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If computational resources are reallocated from lower priority services to higher priority services during large failures, then recovery speed of critical services is improved, but resource availability for non-critical services deteriorates

Engineering Contradiction:
Improverecovery speedVSAvoidservice availability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies local quality by differentiating resource allocation based on service priority levels. During large failures, the system dynamically adjusts resource distribution to provide higher quality service to critical services while maintaining minimal service to non-critical services, rather than applying uniform resource allocation across all services

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system implements dynamic resource reallocation that adapts to changing failure conditions. The fault recovery manager continuously monitors service status and dynamically adjusts computational resource distribution between services based on current priority needs, transitioning from static to dynamic resource management during crisis periods

Inventive Principle:
Principle #15Dynamics

2Reliability

If reserve hardware capacity is increased to handle large failures, then system reliability is improved, but device complexity and cost increase

Engineering Contradiction:
Improvefault toleranceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes the parameter of resource allocation from fixed to variable based on failure detection. Instead of maintaining constant reserve capacity, the system dynamically adjusts resource allocation parameters in response to failure conditions, optimizing the balance between reliability and complexity by only activating additional resource management when needed

Inventive Principle:
Principle #35Parameter changes

3Reliability

If minimal availability thresholds are enforced for all services, then service operability is maintained, but recovery speed of high-priority services deteriorates

Engineering Contradiction:
Improveservice operabilityVSAvoidrecovery speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies asymmetry by implementing different availability threshold requirements for services based on their priority classification. High-priority services receive preferential treatment with lower threshold enforcement during reallocation, while low-priority services maintain stricter thresholds, creating an asymmetric resource management approach that accelerates critical service recovery

Inventive Principle:
Principle #4Asymmetry

Data Source

PatentUS10664348B2Fault recovery management in a cloud computing environment
Publication Date: 2020.05.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10664348B2 patent drawing
  • US10664348B2 patent drawing
  • US10664348B2 patent drawing

AI summary

Technologies for managing fault recovery in a cloud computing environment may be used after faults of various sizes, including faults which put total functioning capacity below subscribed capacity. Computing services have repair priorities. A fault recovery manager selects a higher priority service whose capacity is below a minimum availability, and chooses a lower priority service still above its minimal availability, and reassigns capacity from the lower priority service to the higher priority service without depriving the lower priority service of operability. Capacity reassignment continues at least until the higher priority service is at or above minimal availability, or the lower priority service is at minimal availability. Lower priority services may also be terminated entirely to free up resources for higher priority services. New deployments may be prevented until all services are at or above minimal availability. Spare capacity may be reserved against demand fluctuations or further faults.