Localized Service Resiliency via Edge Resiliency Controller
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern data centers face challenges in providing localized service resiliency, especially when disconnected from remote management consoles, leading to potential network failures and service disruptions during natural disasters or equipment failures, with existing solutions failing to quickly respond to faults and correlate physical and virtual domains effectively.
Innovation Solution
A local resiliency controller is implemented on a local hardware platform to automatically restore NFV-based services using preloaded resiliency policies, capable of handling physical and virtual faults, and providing immediate corrective actions based on local resource knowledge, even in the absence of a remote management system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a remote management console is used to manage NFV services, then centralized control and coordination are improved, but service resiliency during disconnection events deteriorates
Solution Approach 1:
The system divides resiliency management into two segments: centralized policy management at the remote console and localized autonomous execution at the edge device. The resiliency controller at the edge can independently execute pre-configured policies without continuous remote management console connectivity, while still being part of the centralized system architecture.
Solution Approach 2:
Resiliency policies are pre-configured and stored locally at the edge device before disconnection events occur. When the remote management console is accessible, it uploads and stores remediation policies, resource definitions, and correlation rules locally. This preliminary action ensures that when disconnection occurs, the edge device can immediately execute pre-loaded policies without waiting for remote instructions.
2Productivity
If traditional fault response mechanisms are used, then simple faults can be handled, but rapid response to complex physical-virtual correlated faults deteriorates
Solution Approach 1:
The system implements bidirectional feedback mechanisms: the resiliency controller monitors virtual domain faults and automatically correlates them with physical domain events using pre-configured correlation rules. When a fault is detected, the system feeds back the correlated physical-virtual fault information to the policy execution engine, which then selects and executes the appropriate remediation policy. This closed-loop feedback enables rapid automated response to complex correlated faults without manual intervention.
Solution Approach 2:
The resiliency controller acts as an intermediary layer between the virtual domain (NFV services) and physical domain (hardware events). It correlates faults across these domains using pre-configured rules and triggers appropriate remediation actions. This intermediary mechanism enables rapid response to complex faults by automatically bridging the gap between virtual service failures and physical causes without requiring manual correlation analysis.
3Adaptability or versatility
If manual fault remediation procedures are used, then flexibility in handling diverse faults is improved, but response speed and service uptime deteriorate
Solution Approach 1:
The system uses parameter-driven policy execution where remediation actions are selected based on changing fault parameters and conditions. Pre-configured policies contain conditional logic that evaluates fault parameters (type, severity, correlation) and automatically selects the appropriate remediation parameters and actions. This parameter-based approach provides flexibility comparable to manual procedures while enabling automated execution at machine speed, thus maintaining service uptime.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
There is disclosed in one example a computing apparatus, including: a local platform including a hardware platform; a management interface to communicatively couple the local platform to a management controller; a virtualization infrastructure to operate on the hardware platform and to provide a local virtualized function; and a resiliency controller to operate on the hardware platform, and configured to: receive a resiliency policy from the management controller via the management interface, the resiliency policy including information to handle a fault in the virtualized function; detect a fault in the local virtualized function; and effect a resiliency action responsive to detecting the fault.