FMEA Engine for Cloud Application Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for managing failure events in cloud computing applications, especially in dynamic cloud continuum environments, are time-intensive, expensive, and ineffective due to their complexity and heterogeneity, making it difficult to accurately detect and recover from failure modes.
Innovation Solution
A Failure Mode Effect Analysis (FMEA) engine is employed to collect historical metadata, train machine learning models, and dynamically monitor and evaluate failure modes, recommending efficient recovery processes based on context-specific and deployment-specific parameters to minimize computational load and latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional IT solutions are used to manage failure events, then failure analysis can be performed, but the process becomes time-intensive and expensive
Solution Approach 1:
The system performs preliminary actions by collecting and storing historical metadata about failure modes and recovery processes before actual failures occur. This pre-prepared data enables the FMEA engine to quickly identify failure modes and recommend recovery processes without time-consuming analysis during incident response.
Solution Approach 2:
The system creates a virtual copy of the complex cloud continuum application topology and uses this copy for FMEA analysis. The FMEA engine processes metadata representing the application structure, allowing failure analysis without disrupting the actual complex system operations.
2Measurement precision
If conventional manual analysis methods are used, then failure causes can be identified, but the complexity increases with heterogeneous cloud continuum applications
Solution Approach 1:
The system segments the complex cloud continuum application into hierarchical components (cloud services, data centers, networks, end devices) and analyzes failure modes at each level separately. This segmentation simplifies the analysis of heterogeneous components while maintaining comprehensive coverage of the complex topology.
Solution Approach 2:
The FMEA engine transforms complex system state information into standardized failure mode parameters and recovery process parameters. By normalizing metadata from diverse heterogeneous components into unified parameter formats, the system simplifies analysis while preserving measurement precision.
3Reliability
If recovery processes are implemented without optimization, then system reliability improves, but computational load and latency increase
Solution Approach 1:
The system implements partial recovery actions tailored to the specific failure mode and its impact on service level objectives. Rather than implementing comprehensive recovery for all possible failure scenarios, the system selects and executes only the necessary recovery processes that address the actual failure condition, minimizing unnecessary computational overhead.
Solution Approach 2:
The FMEA engine continuously monitors system metadata and provides feedback to adjust recovery process recommendations in real-time. This feedback mechanism allows the system to optimize recovery actions based on actual system state, selecting recovery processes that restore reliability while minimizing computational load and latency.
Data Source
AI summary
Aspects of the present disclosure provide methods, devices, and computer-readable storage media that support detection, effect monitoring, and recovery from failure modes in cloud computing application using a failure mode effect analysis (FMEA) engine. Historical metadata related to operation of a hierarchy of devices may be used as training data to train the FMEA engine to identify failure modes experienced by the hierarchy of devices. After training the FMEA engine, metadata from the hierarchy of devices may be input to the FMEA engine to identify a failure mode that may have occurred, and the FMEA engine may select a recovery process to recommend for addressing or mitigating the identified failure mode. In some implementations, the FMEA engine may output an indication of the recommended recovery process and/or initiate performance of one or more operations at the hierarchy of devices to recover from the failure event.


