DRL-Based Wireless Packet Scheduling With Preference Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern cellular base station schedulers face challenges in balancing competing performance metrics such as throughput and packet delay, especially with diverse Quality of Service (QoS) requirements and dynamic network conditions, which existing heuristic algorithms struggle to efficiently manage.

Innovation Solution

The implementation of a Deep Reinforcement Learning (DRL) based scheduling method that uses a preference vector to jointly optimize multiple performance metrics, including packet size, packet delay, and QoS requirements, by training a DRL policy with composite reward functions to determine optimal weight values for scheduling decisions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If heuristic scheduling algorithms (round robin, proportional fair, exponential rule) are used, then implementation simplicity is maintained, but ability to handle diverse QoS requirements and dynamic network conditions deteriorates

Engineering Contradiction:
Improveimplementation simplicityVSAvoidability to handle diverse QoS requirements
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent replaces traditional heuristic scheduling algorithms (mechanical rule-based systems) with a Deep Reinforcement Learning model. The DRL agent learns optimal scheduling decisions through training with composite reward functions that incorporate multiple QoS metrics, enabling adaptive handling of diverse traffic requirements without manual configuration of complex heuristic rules.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the scheduling system by introducing a preference vector that dynamically weights multiple performance metrics (throughput, delay, fairness, energy efficiency). This allows the system to adapt to different network conditions and QoS requirements by adjusting parameter weights rather than switching between fixed heuristic algorithms.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If multiple performance metrics are jointly optimized using DRL with preference vectors, then network performance improves, but computational complexity increases

Engineering Contradiction:
Improvenetwork performanceVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-training the DRL model offline with composite reward functions that capture multiple performance metrics. The trained model and preference vectors are stored and deployed for real-time scheduling decisions, moving the computationally intensive training phase to beforehand and enabling simple inference during actual network operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The preference vector serves as an intermediary that bridges multiple performance metrics and the DRL decision-making process. By consolidating multiple QoS requirements into a single weighted preference vector, the system simplifies the optimization problem while still jointly considering throughput, delay, fairness, and energy efficiency metrics.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If traditional scheduling algorithms prioritize single metric optimization, then algorithm simplicity is maintained, but ability to balance competing performance metrics deteriorates

Engineering Contradiction:
Improvealgorithm simplicityVSAvoidability to balance competing performance metrics
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements universality by designing a unified DRL-based scheduler that simultaneously optimizes multiple performance metrics (throughput, delay, fairness, energy efficiency) through a single preference vector. This multi-functional approach replaces the need for separate specialized algorithms for each metric, achieving both simplicity and comprehensive performance balancing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent incorporates feedback mechanisms through composite reward functions that provide continuous performance evaluation across multiple metrics. The DRL agent receives feedback on throughput, delay, fairness, and energy efficiency simultaneously, enabling it to learn and balance competing objectives through iterative training and adaptation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230262683A1Method and system for deep reinforcement learning (DRL) based scheduling in a wireless system
Publication Date: 2023.08.17 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20230262683A1 patent drawing
  • US20230262683A1 patent drawing
  • US20230262683A1 patent drawing

AI summary

Systems and methods are disclosed herein for Deep Reinforcement Learning (DRL) based packet scheduling. In one embodiment, a method performed by a network node for DRB-based scheduling comprises performing a DRL-based scheduling procedure using a preference vector for a plurality of network performance metrics correlated to one of a plurality of desired network performance behaviors, the preference vector defining weights for the plurality of network performance metrics correlated to the one of the plurality of desired network performance behaviors. In this manner, DRL-based scheduling is provided in a manner in which multiple performance metrics are jointly optimized.