Predictive Region Switching for Shared Cloud Component Failures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Failures in cloud computing environments, such as device or service outages, can lead to application failures and resource over-allocation, causing user frustration and potential data loss, especially in complex architectures where multiple applications share infrastructure components.
Innovation Solution
Proactively identify shared and intermittent components by analyzing monitoring data, probing for responses, and using a prediction model to determine if region-switching criteria are met, thereby initiating a switch to a new data center region before failures occur.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If passive scanning of distributed computing environment is used, then system simplicity is maintained, but failure detection precision is insufficient and weak point failures are missed
Solution Approach 1:
The system performs preliminary actions by actively probing intermittent components before failures occur. The prediction model analyzes monitoring data and sends proactive probing messages to detect potential failures in shared intermittent components before they cause application failures, enabling preventive region-switching operations.
Solution Approach 2:
The patent introduces an intermediary prediction model that acts as a mediator between passive monitoring data collection and active failure detection. The prediction model processes monitoring data and determines when to send probing messages to intermittent components, bridging the gap between simple monitoring and complex active detection.
2Reliability
If reactive failover switching is performed after failure occurrence, then system complexity is minimized, but application reliability deteriorates due to client-detected failures
Solution Approach 1:
The system performs preliminary region-switching operations before failures occur by using the prediction model to identify at-risk intermittent components. When the model predicts a failure is imminent, the system proactively switches regions and provisions resources in advance, preventing application failures and maintaining high reliability.
Solution Approach 2:
The patent implements beforehand cushioning by provisioning backup infrastructure resources in advance based on prediction model alerts. When potential failures are detected in shared intermittent components, the system pre-provisions resources in alternative regions, creating a cushion that prevents service disruption when failures actually occur.
3Adaptability or versatility
If piecemeal failover of applications is performed, then resource allocation flexibility is improved, but severe over-allocation of resources occurs in destination data center region
Solution Approach 1:
The patent merges multiple applications that share intermittent components into a unified monitoring and failover strategy. By identifying shared intermittent components used by multiple applications, the system coordinates region-switching operations for all affected applications simultaneously, preventing duplicate resource provisioning and eliminating over-allocation in destination regions.
Solution Approach 2:
The prediction model serves a universal function by monitoring shared intermittent components that are used by multiple different applications. A single prediction alert can trigger coordinated failover for multiple applications sharing the same intermittent component, enabling the system to handle diverse application portfolios with a unified approach that optimizes resource allocation.
Data Source
AI summary
A method and related system for application resilience by proactively switching data center regions based on detected failures in shared intermittent components by determining shared components of a first data center region based on monitoring data associated with a set of deployed applications with or without requiring the occurrence of active traffic. The method includes determining an intermittent component of the shared components based on the monitoring data and an activity gap threshold. The method further includes probing the intermittent component, obtaining a set of responses from the intermittent component, and determining a combined resource value based on performance data associated with the set of deployed applications. The method further includes, in response to a determination that the set of responses satisfies a set of region-switching criteria, provisioning a second set of infrastructure resources of a second data center region based on the combined resource value.


