Cloud Service Recovery via Customer Transaction Simulation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional recovery solutions for cloud computing systems are limited as they primarily address failures at individual components and lack efficient automated methods for restoring usability across distributed systems, making them costly and ineffective for widespread dispersed computing resources.
Innovation Solution
A management application simulates customer transactions to detect failures, maps them to recovery actions, and executes these actions to restore subsystems, with monitoring to determine success status, enhancing the recovery process for cloud-based services.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional recovery solutions monitor and address failures at individual components, then failure detection capability is maintained, but recovery effectiveness deteriorates due to inability to restore usability across distributed systems
Solution Approach 1:
The patent segments the cloud computing system into multiple hierarchical levels (individual components, clusters, and overall system). Each level has its own monitoring and recovery mechanisms, allowing failure detection at component level while enabling coordinated recovery across distributed systems through cluster-level orchestration.
Solution Approach 2:
The patent introduces cluster-level controllers as intermediary components between individual components and central management. These controllers coordinate recovery actions across multiple components within a cluster, enabling effective restoration of distributed systems while maintaining precise failure detection at the component level.
2Manufacturing precision
If manual installation and configuration support is provided to cloud computing assets, then installation accuracy is improved, but cost effectiveness deteriorates for widely dispersed computing resources
Solution Approach 1:
The patent implements automated self-service mechanisms where cloud computing assets perform their own installation and configuration through pre-configured templates and automated provisioning systems. This eliminates the need for manual intervention while maintaining high installation accuracy through standardized automated processes.
Solution Approach 2:
The patent uses parameter-based configuration where installation and deployment are controlled through configurable parameters and templates. This allows automated systems to adapt to different cloud computing assets by changing parameters rather than requiring manual configuration, achieving both accuracy and cost-effectiveness.
3Reliability
If recovery actions are transmitted and executed to a cluster to resolve failures, then system-wide recovery capability is improved, but response time deteriorates due to coordination overhead
Solution Approach 1:
The patent implements preliminary action by pre-configuring recovery policies, templates, and procedures at the cluster level before failures occur. When failures are detected, pre-planned recovery actions are automatically executed with minimal coordination overhead, reducing response time while maintaining system-wide recovery capability.
Solution Approach 2:
The patent introduces dynamic recovery mechanisms that adapt coordination intensity based on failure severity and cluster state. For minor failures, individual components can self-recover without cluster coordination. For major failures, dynamic coordination is activated to orchestrate system-wide recovery, optimizing response time across different failure scenarios.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Usability of a cloud based service is recovered from a system failure. A customer transaction associated with the customer experience is executed to simulate the customer experience in the cloud based service. A failure associated with a subsystem the cloud based service is detected from an output of the customer transaction. A recovery action is determined to be associated with the failure. The recovery action is executed on the subsystem and monitored to determine a success status.