DRL-Based Wireless Packet Scheduling With Preference Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern cellular base station schedulers face challenges in balancing competing performance metrics such as throughput and packet delay, especially with diverse Quality of Service (QoS) requirements and dynamic network conditions, which existing heuristic algorithms struggle to efficiently manage.
Innovation Solution
The implementation of a Deep Reinforcement Learning (DRL) based scheduling method that uses a preference vector to jointly optimize multiple performance metrics, including packet size, packet delay, and QoS requirements, by training a DRL policy with composite reward functions to determine optimal weight values for scheduling decisions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If heuristic scheduling algorithms (round robin, proportional fair, exponential rule) are used, then implementation simplicity is maintained, but ability to handle diverse QoS requirements and dynamic network conditions deteriorates
Solution Approach 1:
The patent replaces traditional heuristic scheduling algorithms (mechanical rule-based systems) with a Deep Reinforcement Learning model. The DRL agent learns optimal scheduling decisions through training with composite reward functions that incorporate multiple QoS metrics, enabling adaptive handling of diverse traffic requirements without manual configuration of complex heuristic rules.
Solution Approach 2:
The patent changes the fundamental parameters of the scheduling system by introducing a preference vector that dynamically weights multiple performance metrics (throughput, delay, fairness, energy efficiency). This allows the system to adapt to different network conditions and QoS requirements by adjusting parameter weights rather than switching between fixed heuristic algorithms.
2Productivity
If multiple performance metrics are jointly optimized using DRL with preference vectors, then network performance improves, but computational complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-training the DRL model offline with composite reward functions that capture multiple performance metrics. The trained model and preference vectors are stored and deployed for real-time scheduling decisions, moving the computationally intensive training phase to beforehand and enabling simple inference during actual network operation.
Solution Approach 2:
The preference vector serves as an intermediary that bridges multiple performance metrics and the DRL decision-making process. By consolidating multiple QoS requirements into a single weighted preference vector, the system simplifies the optimization problem while still jointly considering throughput, delay, fairness, and energy efficiency metrics.
3Ease of operation
If traditional scheduling algorithms prioritize single metric optimization, then algorithm simplicity is maintained, but ability to balance competing performance metrics deteriorates
Solution Approach 1:
The patent implements universality by designing a unified DRL-based scheduler that simultaneously optimizes multiple performance metrics (throughput, delay, fairness, energy efficiency) through a single preference vector. This multi-functional approach replaces the need for separate specialized algorithms for each metric, achieving both simplicity and comprehensive performance balancing.
Solution Approach 2:
The patent incorporates feedback mechanisms through composite reward functions that provide continuous performance evaluation across multiple metrics. The DRL agent receives feedback on throughput, delay, fairness, and energy efficiency simultaneously, enabling it to learn and balance competing objectives through iterative training and adaptation.
Data Source
AI summary
Systems and methods are disclosed herein for Deep Reinforcement Learning (DRL) based packet scheduling. In one embodiment, a method performed by a network node for DRB-based scheduling comprises performing a DRL-based scheduling procedure using a preference vector for a plurality of network performance metrics correlated to one of a plurality of desired network performance behaviors, the preference vector defining weights for the plurality of network performance metrics correlated to the one of the plurality of desired network performance behaviors. In this manner, DRL-based scheduling is provided in a manner in which multiple performance metrics are jointly optimized.


