Cloud Network Fault Tolerance Framework
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud networks face challenges in providing fault tolerance and resiliency due to the diversity of network devices from various vendors, which makes it difficult to handle hardware or software failures across a wide range of devices, potentially disrupting mission-critical applications.
Innovation Solution
A fault tolerance and resiliency framework that monitors network devices, performs recovery operations, and alerts administrators for non-recoverable errors, including automated means to replace faulty components and generate alerts, ensuring the cloud network's continuity and high availability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If network devices from multiple third-party vendors are utilized to build a cloud network, then device diversity and vendor options increase, but the difficulty of providing resiliency and handling hardware or software failures across varied devices increases
Solution Approach 1:
The patent implements a universal monitoring and management framework that can handle multiple types of network devices from different vendors through a single system. The framework uses standardized monitoring protocols and abstraction layers to provide consistent fault detection, analysis, and remediation capabilities across heterogeneous device types, eliminating the need for vendor-specific management approaches.
2Reliability
If comprehensive monitoring and recovery operations are implemented across all network devices, then fault detection capability improves, but system complexity and resource consumption increase
Solution Approach 1:
The patent segments the monitoring system into modular components including device agents, central monitoring server, analysis modules, and remediation systems. Each component has a specific function and can operate independently, allowing the system to scale and manage complexity through modular architecture while maintaining comprehensive monitoring capabilities across the network.
3Reliability
If automated recovery operations are performed to restore cloud network capacity, then service continuity improves, but the risk of improper recovery actions affecting device operation increases
Solution Approach 1:
The patent implements preliminary validation and verification steps before executing recovery operations. The system analyzes device status, determines appropriate recovery actions based on pre-defined policies, and performs verification checks to ensure safety before applying remediation. This preliminary action framework reduces the risk of harmful recovery operations while maintaining service continuity.
Data Source
AI summary
In accordance with an embodiment, described herein is a system and method for providing fault tolerance and resiliency within a cloud network. A cloud computing environment provides access, via the cloud network, to software applications executing within the cloud environment. The cloud network can include a plurality of network devices, of which various network devices can be configured as virtual chassis devices, cluster members, or standalone devices. A fault tolerance and resiliency framework can monitor the network devices, to receive status information associated with the devices. In the event the system determines a failure or error associated with a network device, it can attempt to perform recovery operations to restore the cloud network to its original capacity or state. If the system determines that a particular network device cannot recover from the failure or error, it can alert an administrator for further action.


