RL-Based Downlink Data Splitting in 5G Multi-Connectivity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions for splitting downlink data in 5G NR radio access networks are reactive and prone to congestion, as they only redirect packets after detecting high delays without considering potential future congestion.
Innovation Solution
The implementation of reinforcement learning agents at the first base station to proactively split downlink data between a first path and at least one second path, using a common neural network policy that is updated based on observed performance metrics and rewards, to optimize path selection and avoid congestion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If reactive delay estimation methods (such as Little's Law or supervised learning) are used to redirect packets, then packet routing decisions can be made based on observed delays, but the system becomes congestion-prone because paths are avoided only after congestion is detected rather than prevented
Solution Approach 1:
The patent applies preliminary action by using reinforcement learning agents to predict future congestion conditions and proactively redirect packets before actual congestion occurs. The RL agents continuously learn from network state observations and reward signals, enabling them to anticipate congestion-prone paths and switch to alternative paths in advance, rather than reacting after delay is detected.
Solution Approach 2:
The patent implements feedback through the reinforcement learning framework where agents receive reward signals based on path performance metrics. This feedback loop enables continuous learning and adaptation, allowing the system to refine its path selection decisions over time based on actual network conditions and congestion patterns.
2Ease of manufacture
If known splitting solutions redirect packets only after measuring high delay, then implementation is simpler using analytical methods, but the system cannot prevent future congestion and performance deteriorates over time
Solution Approach 1:
The patent replaces traditional mechanical/analytical delay estimation methods with a reinforcement learning-based intelligent system. The RL agents use neural networks to process network state information and make adaptive routing decisions, substituting simple analytical calculations with a more sophisticated learning-based mechanism that can predict and prevent congestion.
Solution Approach 2:
The patent changes the parameter used for path selection from static delay measurements to dynamic predictions based on reinforcement learning models. The system transitions from using observed delay values to using predicted future congestion probabilities, enabling proactive rather than reactive path selection.
3Speed
If a single path is selected based on current shortest delay, then immediate packet routing is optimized, but the system lacks far-sightedness to avoid future congestion in the overall system
Solution Approach 1:
The reinforcement learning agents perform preliminary analysis of network conditions and predict future congestion patterns before making routing decisions. This enables the system to balance immediate speed optimization with long-term reliability by selecting paths that appear short-term optimal but may lead to future congestion, versus paths that are slightly longer now but will remain reliable in the future.
Solution Approach 2:
The patent introduces dynamics to path selection through continuous learning and adaptation. The RL agents update their policies based on changing network conditions, allowing the system to dynamically adjust its routing strategy to balance immediate performance with long-term congestion avoidance as network conditions evolve over time.
Data Source
AI summary
A first base station and a central entity for use in a radio access network implementing a dual/multiconnectivity scheme where the first base station is caused to split downlink data associated with a plurality of user devices between a first path through the first base station and a second path through the at least one second base station. The splitting comprises a plurality of reinforcement learning agents associated with the plurality of user devices. The reinforcement learning agents apply a current common neural network policy to select a path amongst the first and the second paths based on observed performance metrics of the first and second paths and to generate a reward associated with the selected path and send their experiences to the central entity. The central entity updates the common neural network policy based on the received experiences. The central unit can be implemented in a RIC.


