Localized Service Resiliency via Edge Resiliency Controller

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern data centers face challenges in providing localized service resiliency, especially when disconnected from remote management consoles, leading to potential network failures and service disruptions during natural disasters or equipment failures, with existing solutions failing to quickly respond to faults and correlate physical and virtual domains effectively.

Innovation Solution

A local resiliency controller is implemented on a local hardware platform to automatically restore NFV-based services using preloaded resiliency policies, capable of handling physical and virtual faults, and providing immediate corrective actions based on local resource knowledge, even in the absence of a remote management system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a remote management console is used to manage NFV services, then centralized control and coordination are improved, but service resiliency during disconnection events deteriorates

Engineering Contradiction:
Improvecentralized controlVSAvoidservice resiliency during disconnection
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system divides resiliency management into two segments: centralized policy management at the remote console and localized autonomous execution at the edge device. The resiliency controller at the edge can independently execute pre-configured policies without continuous remote management console connectivity, while still being part of the centralized system architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Resiliency policies are pre-configured and stored locally at the edge device before disconnection events occur. When the remote management console is accessible, it uploads and stores remediation policies, resource definitions, and correlation rules locally. This preliminary action ensures that when disconnection occurs, the edge device can immediately execute pre-loaded policies without waiting for remote instructions.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If traditional fault response mechanisms are used, then simple faults can be handled, but rapid response to complex physical-virtual correlated faults deteriorates

Engineering Contradiction:
Improvefault handling capabilityVSAvoidresponse time to complex faults
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system implements bidirectional feedback mechanisms: the resiliency controller monitors virtual domain faults and automatically correlates them with physical domain events using pre-configured correlation rules. When a fault is detected, the system feeds back the correlated physical-virtual fault information to the policy execution engine, which then selects and executes the appropriate remediation policy. This closed-loop feedback enables rapid automated response to complex correlated faults without manual intervention.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The resiliency controller acts as an intermediary layer between the virtual domain (NFV services) and physical domain (hardware events). It correlates faults across these domains using pre-configured rules and triggers appropriate remediation actions. This intermediary mechanism enables rapid response to complex faults by automatically bridging the gap between virtual service failures and physical causes without requiring manual correlation analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If manual fault remediation procedures are used, then flexibility in handling diverse faults is improved, but response speed and service uptime deteriorate

Engineering Contradiction:
Improvefault handling flexibilityVSAvoidservice uptime
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system uses parameter-driven policy execution where remediation actions are selected based on changing fault parameters and conditions. Pre-configured policies contain conditional logic that evaluates fault parameters (type, severity, correlation) and automatically selects the appropriate remediation parameters and actions. This parameter-based approach provides flexibility comparable to manual procedures while enabling automated execution at machine speed, thus maintaining service uptime.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3588855B1Localized service resiliency
Publication Date: 2024.12.25 INTEL CORP
  • EP3588855B1 patent drawingFigure 1
  • EP3588855B1 patent drawingFigure 2
  • EP3588855B1 patent drawingFigure 3

AI summary

There is disclosed in one example a computing apparatus, including: a local platform including a hardware platform; a management interface to communicatively couple the local platform to a management controller; a virtualization infrastructure to operate on the hardware platform and to provide a local virtualized function; and a resiliency controller to operate on the hardware platform, and configured to: receive a resiliency policy from the management controller via the management interface, the resiliency policy including information to handle a fault in the virtualized function; detect a fault in the local virtualized function; and effect a resiliency action responsive to detecting the fault.