Fabric QoS Manager for Scalable Point-to-Point Bandwidth Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data center architectures face challenges in meeting quality of service (QoS) requirements for point-to-point connections due to limitations in virtual channels and traffic classes, which restrict the ability to provide dedicated bandwidth to hundreds or thousands of applications hosted on compute nodes, leading to inadequate QoS for service level agreements (SLAs).
Innovation Solution
A system comprising a QoS manager and host fabric interface (HFI) logic that orchestrates resource provisioning and manages point-to-point connections by allocating virtual channels and bandwidth dynamically, using process address space identifiers (PASIDs) to track and manage SLA requests, ensuring that each application meets its QoS requirements through efficient bandwidth allocation and monitoring.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If virtual channels and traffic classes are used to manage fabric connections, then QoS requirements can be met for a limited number of applications, but the system cannot scale to support hundreds or thousands of applications with dedicated bandwidth
Solution Approach 1:
The patent segments the fabric connection resources by introducing virtual channel groups (VCGs) that can be independently allocated to different applications. Each VCG contains multiple virtual channels that can be dynamically assigned, allowing the system to support many applications without increasing overall fabric complexity. This segmentation enables fine-grained resource allocation at the application level while maintaining manageable complexity at the fabric level.
Solution Approach 2:
The patent adds a new dimension of abstraction by introducing the concept of virtual channel groups as an intermediate layer between physical fabric resources and applications. This dimensional addition allows multiple applications to share fabric resources through organized groups, enabling scaling from dozens to thousands of applications without proportionally increasing management complexity.
2Reliability
If dedicated bandwidth is allocated to each application through virtual channels, then QoS requirements can be met, but the number of available virtual channels is limited to a few dozen
Solution Approach 1:
The patent merges multiple virtual channels into virtual channel groups that can be collectively allocated to applications. By combining several VC lanes into a single allocatable unit (VCG), the system effectively multiplies the number of available dedicated bandwidth allocations without requiring a proportional increase in individual virtual channel counts. This merging approach allows hundreds or thousands of applications to receive dedicated bandwidth guarantees.
Solution Approach 2:
Virtual channel groups serve multiple functions: they provide dedicated bandwidth allocation, enable application-specific QoS management, and allow dynamic resource sharing. A single VCG can be allocated to one application for dedicated bandwidth or shared among multiple applications with different QoS requirements, making the fabric resource management system universally applicable to diverse workloads.
3Ease of operation
If greedy ad-hoc provisioning of virtual channels is used, then applications can request bandwidth independently, but the system cannot provide sufficient dedicated channels for hundreds or thousands of applications
Solution Approach 1:
The patent introduces dynamic allocation of virtual channel groups where the size and composition of VCGs can change based on application requirements. Applications can request bandwidth dynamically, and the system can allocate or de-allocate VCGs in real-time. This dynamic approach maintains ease of operation for individual applications while enabling the system to scale to thousands of applications through efficient resource orchestration.
Data Source
AI summary
Examples include techniques to meet quality of service (QoS) requirements for a fabric point to point connection. Examples include an application hosted by a compute node coupled with a fabric requesting bandwidth for a point to point connection through the fabric and the request being granted or not granted based at least partially on whether bandwidth is available for allocation to meet one or more QoS requirements.


