Virtual Network Function Resilience via Dual Hypervisor Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing network function virtualization (NFV) systems face challenges in ensuring resilience against hardware and software failures, particularly in customer premises equipment where shared storage and clustering may be unimplementable due to cost or design constraints, leading to potential cascading failures of virtual network functions.
Innovation Solution
A host platform with dual processor and memory blocks (PMBs) connected to an I/O switch, where one PMB runs a primary hypervisor for normal operation and a secondary PMB with less processing power runs an emergency hypervisor to maintain network function virtualization (VNF) functionality in case of primary hypervisor failure, with a system controller monitoring and switching traffic to the secondary PMB to ensure continuity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a primary hypervisor is used for normal VNF operation, then network function performance is improved, but system reliability deteriorates due to single point of failure
Solution Approach 1:
An emergency hypervisor is pre-configured and maintained in a standby state on a separate PMB before any failure occurs. This emergency hypervisor includes all necessary VNF configurations and is ready to immediately take over if the primary hypervisor fails, thus preventing single point of failure while maintaining normal performance characteristics of the primary system.
Solution Approach 2:
The system differentiates between primary and emergency hypervisors with distinct functional roles. The primary hypervisor handles normal high-performance operations, while the emergency hypervisor is specifically designed for failover scenarios. This local differentiation allows each component to be optimized for its specific purpose, improving both performance and reliability.
2Reliability
If dual PMBs with hypervisors are implemented for failover capability, then system reliability is improved, but device complexity increases
Solution Approach 1:
Both the primary and emergency hypervisors are integrated within a single host platform, sharing common I/O resources, storage, and management infrastructure. This merging approach allows failover capability to be achieved without duplicating entire separate systems, thereby improving reliability while controlling device complexity through resource sharing.
Solution Approach 2:
The second PMB serves multiple functions: it hosts the emergency hypervisor for failover scenarios, shares I/O resources with the primary PMB through the I/O switch, and can potentially host additional VNFs. This multi-functionality reduces the need for dedicated separate components, thereby improving reliability without proportionally increasing device complexity.
3Quantity of substance
If emergency hypervisor with less processing power is used, then resource utilization is improved, but network function performance during emergency operation deteriorates
Solution Approach 1:
The emergency hypervisor is designed to handle critical network functions at reduced capacity rather than attempting to maintain full performance. During emergency operation, it provides sufficient functionality to maintain essential services and enable graceful degradation, accepting reduced performance as a trade-off for achieving failover capability with available resources.
Data Source
AI summary
There is provided a method of resilient operation of a virtual network function (VNF) and a host platform configured to host VNF. The host platform comprises: first PMB configured to host first virtual platform (VP) with the help of first hypervisor, the first VP is configured, when the first hypervisor is in operational mode, to execute the VNF under primary VNF configuration; second PMB configured to host second VP with the help of second hypervisor with substantially less processing power than processing power of the first hypervisor, the second VP is configured, when the first hypervisor has failed, to execute the VNF under emergency VNF configuration. Second PMB is further configured to execute system controller operatively connected to first hypervisor, second hypervisor and I/O switch, the system controller configured to monitor status of first hypervisor and, responsive to a failure of first hypervisor, enable executing the VNF on second VP.


