Multi-Agent Q-Learning for 5G Power and Resource Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Next-generation wireless networks face challenges in balancing the diverse quality of service requirements for heterogeneous services like Ultra-Reliable Low-Latency (URLLC) and enhanced Mobile Broadband (eMBB) over shared 5G channels, particularly in coexistence scenarios where latency and reliability trade-offs lead to suboptimal performance.

Innovation Solution

A multi-agent Q-learning algorithm for joint power and resource allocation is employed, utilizing the 5G-NR time-frequency grid to optimize resource allocation based on demand, with a carefully crafted reward function to address latency, reliability, and queuing delays, and a decentralized learning approach to maximize performance across different service categories.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If Hybrid Automatic Repeat Request (HARQ) re-transmissions with link adaptation are used to achieve high reliability, then reliability is improved, but latency increases due to multiple re-transmissions

Engineering Contradiction:
ImprovereliabilityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements dynamic adjustment of HARQ retransmission parameters based on real-time channel conditions and service requirements. The system dynamically modifies retransmission attempts, timing, and resource allocation to balance reliability improvements against latency penalties, allowing adaptive optimization rather than fixed conservative settings

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key parameters including retransmission thresholds, timing advance values, and power allocation levels based on service category (URLLC vs eMBB) and channel state. By adjusting these parameters dynamically, the system achieves high reliability for URLLC while minimizing latency impact through reduced retransmission rounds when channel conditions permit

Inventive Principle:
Principle #35Parameter changes

2Reliability

If scheduling mechanisms prioritize URLLC traffic by allocating mini-slots and modifying re-transmission mechanisms, then latency and reliability of URLLC are improved, but throughput of eMBB users deteriorates

Engineering Contradiction:
ImproveURLLC reliabilityVSAvoideMBB throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies different scheduling qualities and resource allocation strategies to different service categories within the same network. URLLC users receive prioritized mini-slot allocations and modified HARQ parameters, while eMBB users receive optimized resource allocation that accounts for URLLC reservations. This local differentiation ensures URLLC performance guarantees without completely sacrificing eMBB throughput

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system implements partial prioritization where URLLC traffic receives enhanced scheduling treatment only when needed, rather than always occupying premium resources. The scheduling mechanism dynamically determines the degree of URLLC prioritization based on traffic load and channel conditions, allowing eMBB users to utilize remaining resources efficiently

Inventive Principle:
Principle #16Partial or excessive action

3Loss of time

If resource allocation focuses on minimizing latency for URLLC users, then latency is reduced, but packet drop rate increases

Engineering Contradiction:
ImprovelatencyVSAvoidpacket drop rate
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system performs preliminary actions by pre-allocating resources and preparing transmission parameters before actual data arrival. For URLLC traffic, the scheduler pre-configures mini-slots and anticipates retransmission needs, allowing rapid response without compromising reliability. This proactive resource preparation reduces both latency and packet drops by eliminating decision delays

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11678272B2System and method for joint power and resource allocation using reinforcement learning
Publication Date: 2023.06.13 UNIVERSITY OF OTTAWA
  • US11678272B2 patent drawing
  • US11678272B2 patent drawing
  • US11678272B2 patent drawing

AI summary

Systems and methods for joint power and resource allocation on a shared 5G channel. The method selects one of a group of grouped actions and implements this selected group of actions. The effects of the actions on the environment and/or the users are then assessed. Based on the result, a reward is allocated for the system. Multiple iterations are then executed with a view to maximizing the reward. Each of the grouped actions comprises joint power and resource allocation actions.