Adaptive RRC Timer Control Using Dual DRL State Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional static timer configurations for Radio Resource Control (RRC) state transitions in user equipment (UE) fail to adapt to dynamic user behavior and network conditions, leading to inefficient power consumption, increased signaling overhead, and connection establishment delays.
Innovation Solution
Implementing a dual Deep Reinforcement Learning (DRL) system architecture that optimizes RRC state transitions by learning from usage patterns and network conditions to minimize power consumption, reduce latency, and decrease signaling overhead, using two parallel DRL systems tailored for RRC_ACTIVE and RRC_INACTIVE/IDLE states.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If static timer configurations are used for RRC state transitions, then device complexity is reduced and ease of operation is improved, but power consumption efficiency deteriorates and connection establishment delays increase
Solution Approach 1:
The patent applies dynamics by transitioning from static timer configurations to dynamic, adaptive timer configurations. The system continuously learns optimal timer values through reinforcement learning and adjusts RRC state transition timers in real-time based on changing network conditions and user behavior patterns, thereby improving power consumption efficiency while maintaining operational simplicity
Solution Approach 2:
The patent implements feedback mechanisms where the system monitors network conditions, user behavior, and power consumption metrics, then uses this feedback to continuously refine timer configurations through machine learning algorithms. This closed-loop feedback enables the system to adaptively optimize power efficiency without increasing operational complexity for end users
2Device complexity
If static timer configurations are used for RRC state transitions, then device complexity is reduced, but signaling overhead increases and connection establishment delays worsen
Solution Approach 1:
The system dynamically adjusts timer configurations based on real-time network conditions and historical data analysis. By making timer values adaptive rather than static, the system reduces connection establishment delays without requiring complex manual configuration, as the complexity is shifted to automated learning algorithms that optimize performance
3Use of energy by moving object
If dynamic adaptive timer configurations are implemented, then power consumption efficiency is improved and connection establishment delays are reduced, but device complexity increases
Solution Approach 1:
The system applies self-service by implementing autonomous self-optimization through machine learning algorithms. The network automatically learns optimal timer configurations and adjusts them without manual intervention, thereby improving power efficiency while keeping the user interface and operational complexity low. The complexity is contained within the automated learning system rather than requiring complex user-side configurations
4Loss of substance
If conventional static timer configurations are used, then signaling overhead is reduced, but power consumption increases and connection delays increase
Solution Approach 1:
The patent applies parameter changes by dynamically modifying timer configuration parameters based on learned patterns and current network conditions. This enables the system to optimize the balance between signaling overhead and power consumption by adjusting timer values to match actual usage patterns, thereby reducing unnecessary state transitions and associated signaling while lowering power consumption
Data Source
AI summary
A system can train and maintain a first deep reinforcement learning model, wherein the first deep reinforcement learning model was generated according to a first objective to improve timing of transitions from a radio resource control active state. The system can train and maintain a second deep reinforcement learning model, wherein the second deep reinforcement learning model was generated according to a second objective to improve timing of transitions from a radio resource control inactive state or a radio resource control idle state, and wherein the first deep reinforcement learning model and the second deep reinforcement learning model share an objective function. The system can determine respective timers for respective radio resource control states based on a first result of the training of the first deep reinforcement learning model and a second result of the training of the second deep reinforcement learning model.


