Network Congestion Management via Virtual Switch Reaction Point
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data center network congestion management and resource isolation techniques are sub-optimal, particularly in virtualized environments, as they rely on protocol stacks within guest virtual machines, leading to longer control loops and inefficiencies due to mismatched reaction points between resource management and congestion control.
Innovation Solution
Implementing a network architecture that brings the reaction point closer to the network ports by using software-based virtual switches associated with a hypervisor and leveraging direct VM to NIC data movement, with scheduler and shaper modules in the NIC managing transmit and receive queues to dynamically adjust packet scheduling and shaping based on congestion feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If protocol stacks within guest virtual machines are used for congestion management, then resource isolation can be provided, but control loops become longer and efficiency decreases
Solution Approach 1:
The congestion management reaction point is extracted from the guest virtual machine protocol stack and relocated to the virtual switch. This separates the congestion control function from the VM layer, eliminating the long control loops while maintaining resource isolation through the virtual switch's ability to enforce per-tenant policies and SLAs.
Solution Approach 2:
The virtual switch acts as an intermediary between the physical network and guest virtual machines. It provides the reaction point for congestion management while maintaining resource isolation, serving as a mediator that enforces per-tenant policies and manages network resources without requiring protocol stack changes in the VMs.
2Adaptability or versatility
If reaction point is located in guest virtual machines, then per-tenant policy enforcement is possible, but mismatch between resource management and congestion control reaction points occurs
Solution Approach 1:
The virtual switch merges resource management and congestion control reaction points into a single location. This unified approach eliminates the mismatch between resource management and congestion control while maintaining per-tenant policy enforcement capabilities through integrated scheduling and shaping functions.
Solution Approach 2:
The virtual switch provides multiple functions including resource isolation, per-tenant policy enforcement, congestion management, and scheduling. By consolidating these functions in a single component, the system achieves both adaptability for policy enforcement and reduced complexity through unified reaction point management.
3Productivity
If buffer congestion occurs in data center network, then packet drops increase, but average and tail latency are affected
Solution Approach 1:
The virtual switch performs preliminary congestion management actions by monitoring buffer levels and applying scheduling policies before packets are dropped. This proactive approach prevents buffer congestion from escalating to packet drops, thereby protecting both throughput and latency performance.
Solution Approach 2:
The system implements feedback mechanisms where the virtual switch monitors network conditions and buffer congestion levels, then dynamically adjusts scheduling and shaping policies. This feedback loop enables the system to respond to congestion conditions in real-time, minimizing packet drops and maintaining optimal latency and throughput.
Data Source
AI summary
System, method and apparatus for network congestion management and network resource isolation. A high level network usage and device architecture is provided that can satisfy buffering and network bandwidth resource management for data center networks. The congestion management can be defined to bring the reaction point closer to the network ports. In one embodiment, the reaction point is resident in a network interface card (NIC).


