RDMA Latency-Based SLA Management via Dynamic Priority Scheduling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
RDMA networks in data centers face challenges in managing lossless quality of service (QoS), particularly due to device ports using IEEE priority flow control (PFC) or data center bridging (DCB), which can lead to performance slowdowns and head of line blocking.
Innovation Solution
Implementing a system that manages RDMA latency-based service level agreements through a controller device and host devices, using flexible buffer management and self-correcting scheduling algorithms to prioritize RDMA requests based on transaction latency, rather than static priorities, thereby improving throughput and reducing latency deviations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If IEEE priority flow control (PFC) or data center bridging (DCB) is used for QoS management, then lossless quality of service is achieved, but performance slows down and head of line blocking occurs
Solution Approach 1:
The patent implements dynamic priority scheduling based on real-time latency measurements rather than static priorities. The system continuously monitors RDMA transaction latency and adjusts queue priorities dynamically, allowing the system to adapt to changing network conditions and avoid fixed priority issues like head of line blocking while maintaining lossless QoS.
Solution Approach 2:
The system employs feedback mechanisms by monitoring actual RDMA transaction latency and using this information to adjust scheduling decisions. The latency monitor provides continuous feedback to the scheduler, which then modifies queue priorities and buffer allocations based on measured performance, creating a closed-loop control system that maintains QoS without performance penalty.
2Device complexity
If static priority scheduling is used for RDMA requests, then implementation is simple, but latency deviations increase and throughput is limited
Solution Approach 1:
The system implements self-service through autonomous latency measurement and self-adjusting priority scheduling. The RDMA endpoints automatically monitor their own transaction latency and trigger priority adjustments without external intervention, allowing the system to self-optimize performance while maintaining manageable complexity through standardized monitoring interfaces.
Solution Approach 2:
The patent changes the scheduling parameter from static priority values to dynamic priority values based on measured latency. By transitioning from fixed priority assignments to variable priorities that respond to actual transaction characteristics, the system achieves higher throughput and reduced latency deviations while keeping the core scheduling mechanism relatively simple.
3Productivity
If flexible buffer management is implemented for latency-based scheduling, then throughput improves and latency deviations reduce, but system complexity increases
Solution Approach 1:
The system segments buffer management by creating separate latency-based queues for different RDMA traffic classes. Instead of managing a single monolithic buffer pool, the system divides buffers into multiple segmented queues that can be independently managed and prioritized, reducing overall complexity through modularization while enabling flexible latency-based scheduling.
Solution Approach 2:
The implementation uses partial buffer allocation strategies where only the necessary portion of buffer resources is allocated to each queue based on actual latency requirements. This partial action approach avoids over-provisioning and reduces management complexity by allocating buffers only when and where needed, rather than maintaining full buffer availability for all queues.
Data Source
AI summary
Technologies for latency based service level agreement (SLA) management in remote direct memory access (RDMA) networks include multiple compute devices in communication via a network switch. A compute device determines a service level objective (SLO) indicative of a guaranteed maximum latency for a percentage of RDMA requests of an RDMA session. The compute device receives latency data indicative of latency of an RDMA request from a host device. The compute device determines a priority associated with the RDMA request as a function of the SLO and the latency data. The compute device schedules the RDMA request based on the priority. The network switch may allocate queue resources to the RDMA request based on the priority, reclaim the queue resources after the RDMA request is scheduled, and then return the queue resources to a free pool. Other embodiments are described and claimed.


