FMEA Engine for Cloud Application Failure Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for managing failure events in cloud computing applications, especially in dynamic cloud continuum environments, are time-intensive, expensive, and ineffective due to their complexity and heterogeneity, making it difficult to accurately detect and recover from failure modes.

Innovation Solution

A Failure Mode Effect Analysis (FMEA) engine is employed to collect historical metadata, train machine learning models, and dynamically monitor and evaluate failure modes, recommending efficient recovery processes based on context-specific and deployment-specific parameters to minimize computational load and latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional IT solutions are used to manage failure events, then failure analysis can be performed, but the process becomes time-intensive and expensive

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidtime to analyze and recover
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by collecting and storing historical metadata about failure modes and recovery processes before actual failures occur. This pre-prepared data enables the FMEA engine to quickly identify failure modes and recommend recovery processes without time-consuming analysis during incident response.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a virtual copy of the complex cloud continuum application topology and uses this copy for FMEA analysis. The FMEA engine processes metadata representing the application structure, allowing failure analysis without disrupting the actual complex system operations.

Inventive Principle:
Principle #26Copying

2Measurement precision

If conventional manual analysis methods are used, then failure causes can be identified, but the complexity increases with heterogeneous cloud continuum applications

Engineering Contradiction:
Improvefailure mode identification accuracyVSAvoidsystem topology complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex cloud continuum application into hierarchical components (cloud services, data centers, networks, end devices) and analyzes failure modes at each level separately. This segmentation simplifies the analysis of heterogeneous components while maintaining comprehensive coverage of the complex topology.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The FMEA engine transforms complex system state information into standardized failure mode parameters and recovery process parameters. By normalizing metadata from diverse heterogeneous components into unified parameter formats, the system simplifies analysis while preserving measurement precision.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If recovery processes are implemented without optimization, then system reliability improves, but computational load and latency increase

Engineering Contradiction:
Improvesystem availabilityVSAvoidcomputational load
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The system implements partial recovery actions tailored to the specific failure mode and its impact on service level objectives. Rather than implementing comprehensive recovery for all possible failure scenarios, the system selects and executes only the necessary recovery processes that address the actual failure condition, minimizing unnecessary computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The FMEA engine continuously monitors system metadata and provides feedback to adjust recovery process recommendations in real-time. This feedback mechanism allows the system to optimize recovery actions based on actual system state, selecting recovery processes that restore reliability while minimizing computational load and latency.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230315954A1Method and device for dynamic failure mode effect analysis and recovery process recommendation for cloud computing applications
Publication Date: 2023.10.05 ACCENTURE GLOBAL SOLUTIONS LTD
  • US20230315954A1 patent drawing
  • US20230315954A1 patent drawing
  • US20230315954A1 patent drawing

AI summary

Aspects of the present disclosure provide methods, devices, and computer-readable storage media that support detection, effect monitoring, and recovery from failure modes in cloud computing application using a failure mode effect analysis (FMEA) engine. Historical metadata related to operation of a hierarchy of devices may be used as training data to train the FMEA engine to identify failure modes experienced by the hierarchy of devices. After training the FMEA engine, metadata from the hierarchy of devices may be input to the FMEA engine to identify a failure mode that may have occurred, and the FMEA engine may select a recovery process to recommend for addressing or mitigating the identified failure mode. In some implementations, the FMEA engine may output an indication of the recommended recovery process and/or initiate performance of one or more operations at the hierarchy of devices to recover from the failure event.