Virtual Switching Stack Performance Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Troubleshooting performance issues in virtual computing environments, such as NFV, requires significant human intervention, leading to increased capital and operational expenditures and reduced efficiency, as existing methods rely heavily on manual analysis of telemetry data.
Innovation Solution
The implementation of an automated system that includes a virtual switching monitor and controller to detect performance issues and remediate them automatically, reducing human involvement by generating graphical topology views and implementing remedial actions based on telemetry data from platform resources and switching stacks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual analysis of telemetry data is used to detect and remediate performance issues, then human control and flexibility are maintained, but human involvement increases capital and operational expenditures and reduces efficiency
Solution Approach 1:
The system enables self-service automation where the virtual environment automatically detects performance issues through telemetry analysis and applies remedial actions without human intervention. The automated system monitors itself and performs self-diagnosis and self-healing, eliminating the need for manual troubleshooting while maintaining operational control.
Solution Approach 2:
An automated intermediary system is introduced between the telemetry data and the remediation actions. This intermediary automatically processes telemetry information, identifies performance issues, determines appropriate remedial actions, and executes them, thereby reducing direct human involvement while maintaining system control and efficiency.
2Loss of time
If manual analysis of telemetry information is performed, then flexibility in decision-making is maintained, but the complexity and time required for troubleshooting increases
Solution Approach 1:
The system performs preliminary actions by pre-configuring remedial actions for known performance issues and maintaining a library of remediation strategies. When a performance issue is detected, the system can immediately apply pre-prepared remedial actions, significantly reducing troubleshooting time and simplifying the process.
Solution Approach 2:
The system implements continuous feedback loops where telemetry data is constantly monitored, analyzed, and used to trigger automated remedial actions. The system learns from past performance issues and remediation outcomes, improving its ability to quickly identify and resolve problems while reducing overall process complexity.
3Reliability
If automated remediation is implemented, then efficiency and error reduction are improved, but service providers may lose manual control over their environment
Solution Approach 1:
The system implements dynamic control where the level of automation can be adjusted based on service provider preferences and specific operational contexts. Service providers can configure the system to operate in fully automated mode, semi-automated mode with human approval required, or manual mode, allowing them to balance reliability and control according to their needs.
Solution Approach 2:
The system allows service providers to change operational parameters such as automation thresholds, approval requirements, and monitoring sensitivity. By adjusting these parameters, service providers can optimize the balance between automated reliability and manual control, enabling flexible deployment that adapts to different operational requirements.
Data Source
Figure 1
Figure 2
Figure 3A~3B
AI summary
Telemetry information provided by a computing device includes switching key performance indicators (KPIs), platform KPIs, and topology information. The telemetry information is used to identify performance issues at the computing device, such as packets being dropped in a virtual switching stack or misconfiguration errors. A virtual switching monitor can identify which layers in the switching stack have errors and whether the errors occur along a transmit or receive path in the switching stack. A virtual switching controller can identify remedial actions that can be taken at the computing device to remedy a performance issue. A remedial action can be taken automatically, subject to user approval, or automatically after additional criteria are met.