Cloud Fault Tolerance via Probabilistic Failure Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud infrastructure environments face challenges in maintaining fault tolerance due to the probabilistic nature of underlying infrastructure failures, which differs from the deterministic failure model of on-premise systems, leading to inefficiencies in redundancy and service continuity.

Innovation Solution

A fault tolerance system is implemented in cloud environments using a monitoring component and a pattern recognition component to predict infrastructure failures, allowing for proactive spinning up of new component pieces to compensate for failures, thereby ensuring continuous service.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional deterministic redundancy models are used in cloud environments, then fault tolerance is provided, but resource allocation inefficiency occurs due to the probabilistic nature of cloud infrastructure failures

Engineering Contradiction:
Improvefault toleranceVSAvoidresource allocation efficiency
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent changes the fundamental parameter of failure prediction from deterministic to probabilistic/statistical. The system uses statistical analysis of historical failure data to generate probability scores for potential failures, allowing dynamic adjustment of redundancy levels based on actual risk assessment rather than fixed deterministic models. This resolves the contradiction by optimizing resource allocation according to probabilistic failure likelihood while maintaining appropriate fault tolerance.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements dynamic redundancy allocation that adapts to changing failure probabilities in real-time. Instead of static deterministic redundancy, the system continuously monitors infrastructure components, updates failure probability assessments, and dynamically adjusts the level of redundancy provided. This dynamic approach ensures fault tolerance is maintained while avoiding over-provisioning resources for low-risk components.

Inventive Principle:
Principle #15Dynamics

2Reliability

If proactive failure prediction and mitigation is implemented, then service continuity is improved, but system complexity increases due to monitoring and pattern recognition requirements

Engineering Contradiction:
Improveservice continuityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements feedback loops where failure data from monitoring components is continuously fed into pattern recognition algorithms. These algorithms analyze patterns in the feedback data to predict future failures and adjust redundancy allocation accordingly. The feedback mechanism enables proactive failure mitigation while managing complexity through automated closed-loop control rather than manual intervention.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system employs self-service mechanisms where the pattern recognition component autonomously analyzes failure patterns and makes decisions about redundancy allocation without requiring external intervention. The system serves itself by automatically detecting, predicting, and responding to potential failures, reducing the operational complexity burden on users while maintaining high service continuity.

Inventive Principle:
Principle #25Self-service

3Quantity of substance

If statistical failure prediction is used instead of deterministic models, then resource allocation is optimized, but measurement and detection difficulty increases

Engineering Contradiction:
Improveresource allocation optimizationVSAvoidfailure pattern detection
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces pattern recognition components as intermediaries between raw failure data and statistical analysis. These intermediaries collect, preprocess, and structure failure data from multiple sources, making it suitable for statistical pattern recognition. The intermediary layer simplifies the detection and measurement process by transforming complex raw data into structured patterns that can be more easily analyzed for failure prediction and resource optimization.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11544149B2System and method for improved fault tolerance in a network cloud environment
Publication Date: 2023.01.03 ORACLE INT CORP
  • US11544149B2 patent drawing
  • US11544149B2 patent drawing
  • US11544149B2 patent drawing

AI summary

Described herein are systems and methods for fault tolerance in a network cloud environment. In accordance with various embodiments, the present disclosure provides an improved fault tolerance solution, and improvement in the fault tolerance of systems, by way of failure prediction, or prediction of when an underlying infrastructure will fail, and using the predictions to counteract the failure by spinning up or otherwise providing new component pieces to compensate for the failure.