RL Beam Management via Side Link Data Exchange
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current reinforcement learning (RL) models for beam management in wireless communication networks face challenges in effectively utilizing real data for training while minimizing the impact on system quality of service (QoS) and efficiently utilizing radio resources.
Innovation Solution
The proposed solution utilizes side links to enable real data traffic exchange for RL explorative training, allowing the RL model to take multiple actions for the same input and transmit the same user data over different beams, thereby increasing the chances of receiving data via an exploitive beam while testing exploration beams.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If reinforcement learning models use real data for training beam management, then the accuracy of the RL model is improved, but the training time and impact on system QoS increase
Solution Approach 1:
The patent segments the training process into explorative training (using multiple beams for the same data) and exploitive training (using best beam). By separating exploration and exploitation phases, the system can efficiently utilize real data for training while controlling training time and system impact through structured training protocols.
Solution Approach 2:
The patent dynamically adjusts training parameters including the number of explorative beams, training data selection, and real-time performance monitoring. The system adapts the training process based on observed performance metrics, allowing efficient use of real data while controlling training duration and system QoS impact through dynamic parameter adjustment.
2Adaptability or versatility
If reinforcement learning models test multiple exploration beams, then the system can find optimal beams, but the radio resources are consumed
Solution Approach 1:
The patent applies partial action by limiting the number of exploration beams tested simultaneously rather than exhaustively testing all possible beams. The system selects a manageable subset of beams for explorative training, achieving sufficient beam selection capability while controlling radio resource consumption through deliberate limitation of exploration scope.
Solution Approach 2:
The patent ensures continuous useful action by maintaining both explorative and exploitive training activities running concurrently. The system continuously monitors performance and adjusts training protocols to ensure ongoing learning while optimizing resource usage, allowing the RL model to progressively improve beam selection capability without exhausting radio resources.
3Measurement precision
If the RL model transmits data via multiple beams for training, then real data can be used for training purposes, but the system quality of service is impacted
Solution Approach 1:
The patent implements feedback mechanisms that continuously monitor system performance during training and adjust training protocols accordingly. By observing real-time QoS metrics and comparing them against performance thresholds, the system can modify training parameters to maintain data quality for RL training while preserving acceptable system QoS levels through adaptive control.
Solution Approach 2:
The patent changes training parameters such as the proportion of data transmitted via exploration versus exploitation beams, the selection criteria for training data, and the timing of training operations. By dynamically adjusting these parameters, the system optimizes the balance between obtaining high-quality training data and maintaining system QoS, allowing real data usage without excessive QoS degradation.
Data Source
AI summary
The present disclosure relates to a reinforcement learning (RL) on beam management. In particular, it utilizes side links capabilities for enabling real data/traffic exchange for RL explorative training step. In this way, it enables a radio system performance friendly RL learning training operation from one side and will utilize the available radio air resources on the other side.


