RL Beam Management via Side Link Data Exchange

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning (RL) models for beam management in wireless communication networks face challenges in effectively utilizing real data for training while minimizing the impact on system quality of service (QoS) and efficiently utilizing radio resources.

Innovation Solution

The proposed solution utilizes side links to enable real data traffic exchange for RL explorative training, allowing the RL model to take multiple actions for the same input and transmit the same user data over different beams, thereby increasing the chances of receiving data via an exploitive beam while testing exploration beams.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If reinforcement learning models use real data for training beam management, then the accuracy of the RL model is improved, but the training time and impact on system QoS increase

Engineering Contradiction:
ImproveRL model accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the training process into explorative training (using multiple beams for the same data) and exploitive training (using best beam). By separating exploration and exploitation phases, the system can efficiently utilize real data for training while controlling training time and system impact through structured training protocols.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts training parameters including the number of explorative beams, training data selection, and real-time performance monitoring. The system adapts the training process based on observed performance metrics, allowing efficient use of real data while controlling training duration and system QoS impact through dynamic parameter adjustment.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If reinforcement learning models test multiple exploration beams, then the system can find optimal beams, but the radio resources are consumed

Engineering Contradiction:
Improvebeam selection capabilityVSAvoidradio resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by limiting the number of exploration beams tested simultaneously rather than exhaustively testing all possible beams. The system selects a manageable subset of beams for explorative training, achieving sufficient beam selection capability while controlling radio resource consumption through deliberate limitation of exploration scope.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent ensures continuous useful action by maintaining both explorative and exploitive training activities running concurrently. The system continuously monitors performance and adjusts training protocols to ensure ongoing learning while optimizing resource usage, allowing the RL model to progressively improve beam selection capability without exhausting radio resources.

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If the RL model transmits data via multiple beams for training, then real data can be used for training purposes, but the system quality of service is impacted

Engineering Contradiction:
Improvetraining data qualityVSAvoidsystem QoS
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent implements feedback mechanisms that continuously monitor system performance during training and adjust training protocols accordingly. By observing real-time QoS metrics and comparing them against performance thresholds, the system can modify training parameters to maintain data quality for RL training while preserving acceptable system QoS levels through adaptive control.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes training parameters such as the proportion of data transmitted via exploration versus exploitation beams, the selection criteria for training data, and the timing of training operations. By dynamically adjusting these parameters, the system optimizes the balance between obtaining high-quality training data and maintaining system QoS, allowing real data usage without excessive QoS degradation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250056519A1Mechanism for reinforcement learning on beam management
Publication Date: 2025.02.13 NOKIA TECHNOLOGIES OY
  • US20250056519A1 patent drawing
  • US20250056519A1 patent drawing
  • US20250056519A1 patent drawing

AI summary

The present disclosure relates to a reinforcement learning (RL) on beam management. In particular, it utilizes side links capabilities for enabling real data/traffic exchange for RL explorative training step. In this way, it enables a radio system performance friendly RL learning training operation from one side and will utilize the available radio air resources on the other side.