Frequency Domain Wireless Scheduling With Policy-Based RL Allocation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional schedulers in wireless communication systems, such as those used in 6G networks, rely on heuristic metrics like Proportional Fairness (PF) for resource allocation, which may not optimize network performance effectively.

Innovation Solution

Implementing a policy-based reinforcement learning agent to determine frequency domain scheduling decisions based on a feature vector of devices and frequency resource units, allowing simultaneous allocation across the entire set of available resources, enhancing throughput and reducing latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional heuristic-based schedulers (e.g., Proportional Fairness) are used for resource allocation, then the scheduling operation is simple to implement, but network performance optimization is insufficient

Engineering Contradiction:
Improvenetwork performanceVSAvoidscheduling operation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces conventional heuristic-based scheduling mechanisms with a reinforcement learning-based scheduling operation. The RL agent learns optimal scheduling policies through interaction with the wireless communication environment, substituting traditional mechanical scheduling rules with adaptive intelligent decision-making. This enables the system to optimize network performance while the scheduling operation itself remains encapsulated as a unified black-box process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If frequency domain scheduling decisions are made sequentially or in subsets, then processing complexity is reduced, but spectral efficiency and resource utilization are suboptimal

Engineering Contradiction:
Improvespectral efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent merges the scheduling decisions for all frequency resources into a single unified scheduling operation performed by the reinforcement learning agent. Instead of making sequential or subset-based decisions, the RL agent processes the entire frequency domain scheduling problem simultaneously, considering all available frequency resources and user equipment together. This holistic approach maximizes spectral efficiency and resource utilization while the complex processing is encapsulated within the unified RL-based scheduling operation.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4687367A1Policy-based scheduling of wireless resources
Publication Date: 2026.02.04 NOKIA SOLUTIONS & NETWORKS OY
  • EP4687367A1 patent drawingFigure 1~2
  • EP4687367A1 patent drawingFigure 3
  • EP4687367A1 patent drawingFigure 4

AI summary

The present subject matter relates to a method for performing for a time unit a scheduling operation comprising: receiving (301) a time domain scheduling decision for the time unit, the time domain scheduling decision indicating a set of devices of the wireless communication system; determining (303) a feature vector descriptive of the set of devices and an available set of frequency resource units of the wireless communication system; inputting (305) the feature vector to a policy-based reinforcement learning agent for receiving an output, the output comprising a distribution between the set of devices and the set of frequency resource units; and using (307) the distribution to determine a frequency domain scheduling decision.