Local Recovery Controller for NFV Service Survivability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network management systems, such as MANO, face challenges in minimizing downtime during outages by redeploying all NFV services regardless of their operational state, leading to additional service interruptions and failing to support survivability during management outages, which affects service level agreements and scalability, especially during widespread faults or disasters.
Innovation Solution
Implementing a local recovery controller that can operate independently of central management systems, preloaded with recovery policies, to detect running VNFs and services, allowing them to remain active or apply targeted recovery policies instead of full redeployment, and utilizing a remote recovery manager to distribute recovery actions across multiple nodes and platforms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the management system redeploys all NFV services after connection recovery, then service availability is improved, but service interruption time increases and service level agreements are violated
Solution Approach 1:
The local recovery controller performs preliminary actions by detecting the operational state of VNFs and services during the connection outage period. Before the management system reconnects, the local controller has already identified which services are running and which need recovery, preparing the system for targeted rather than blanket redeployment.
Solution Approach 2:
The recovery process is segmented into different categories: services that are already running and operational, services that failed and need recovery, and services that never started. The local recovery controller applies different recovery strategies to each segment, avoiding unnecessary redeployment of services that are already functioning properly.
2Reliability
If the management system attempts full redeployment of all services, then complete service recovery is achieved, but system complexity increases and scalability is reduced
Solution Approach 1:
The local recovery controller enables the system to serve itself during management outages. When the connection to the central management system is lost, the local controller autonomously detects service states, determines which services need recovery, and executes recovery actions without external intervention. This self-service capability reduces the complexity of centralized control while maintaining complete service recovery.
3Ease of operation
If the management system redeploys services without checking operational state, then recovery process is simplified, but service disruptions increase and mean time to recover increases
Solution Approach 1:
The local recovery controller implements a feedback mechanism that continuously monitors and detects the operational state of VNFs and services during the connection outage. This feedback information is used to determine which services actually need recovery actions, allowing the system to maintain simplicity while avoiding unnecessary redeployment that would increase recovery time.
4Extent of automation
If the system uses centralized management control, then service management is centralized, but service survivability during management outages is reduced
Solution Approach 1:
The control architecture is segmented into centralized management functions and distributed local recovery functions. The local recovery controller handles service detection and recovery actions locally, while the central management system retains oversight for non-critical services. This segmentation enables service survivability during outages while maintaining centralized management control for normal operations.
Data Source
AI summary
Examples described herein relate to a management system that determines which services to redeploy on one or more platforms. A platform can receive a configuration to perform during a failure of connectivity with a management system. The platform can monitor activity of one or more services. The platform can, based on failure of connectivity with the management system and recovery of connectivity with the management system, provide the monitored activity of one or more services to the management system to influence services re-deployed by the management system. In some examples, based on failure to re-establish a connection with the management system within an amount of time, the platform can connect with the management system using a secondary management interface.


