Central Node RL Exploration Control in Radio Access Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge in radio access networks (RAN) is to minimize performance degradation due to exploration in Reinforcement Learning (RL) techniques, which can lead to coverage holes and suboptimal user experiences, as RL agents may take random actions during exploration, affecting system availability, accessibility, reliability, and user experience.

Innovation Solution

A central node evaluates the cost of actions and performance of RL modules in distributed nodes, determining and configuring exploration parameters to control the exploration strategy, thereby reducing the impact of performance degradation, especially in the presence of high-importance services, by optimizing exploration and training parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If RL agents perform exploration by taking random actions, then learning efficiency and adaptability are improved, but system reliability and user experience deteriorate due to coverage holes and performance degradation

Engineering Contradiction:
Improvelearning efficiencyVSAvoidsystem reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies local quality by differentiating exploration behavior across different service types. Critical services (e.g., emergency calls, low-latency applications) are excluded from exploration actions, while non-critical services allow exploration. This creates localized exploration policies where the same RL agent adapts its exploration intensity based on the QoS requirements of individual services, resolving the contradiction between learning efficiency and system reliability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamics by making the exploration rate a time-varying parameter that decreases as the RL agent gains experience. The exploration rate is dynamically adjusted based on confidence levels, performance metrics, and service criticality. This dynamic approach allows high exploration initially for learning, then progressively reduces exploration to maintain reliability, resolving the contradiction between adaptability and reliability.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If exploration rate is increased to improve learning, then adaptability improves, but performance degradation and coverage holes increase

Engineering Contradiction:
Improvelearning capabilityVSAvoidperformance degradation
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary action by pre-identifying and protecting critical services from exploration actions before exploration occurs. The system pre-classifies services based on QoS requirements and establishes protection rules that prevent exploration-induced performance degradation in critical services. This preliminary protective measure allows aggressive exploration in non-critical services while maintaining performance in critical ones.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements parameter changes by dynamically adjusting the exploration rate parameter based on multiple factors including service criticality, current performance metrics, and agent confidence levels. The exploration rate is transformed from a fixed parameter to a dynamic one that changes continuously, allowing the system to optimize learning capability while minimizing performance degradation through real-time parameter adaptation.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If centralized control is implemented to manage exploration strategies, then system reliability improves, but network complexity and signaling overhead increase

Engineering Contradiction:
Improvesystem reliabilityVSAvoidnetwork complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the exploration control function into modular components: service classification module, exploration rate determination module, and protection rule enforcement module. Each module handles a specific aspect of exploration management, making the centralized control more manageable and less complex. The segmentation allows independent optimization of each function while maintaining overall system reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements universality by designing the centralized exploration controller to handle multiple functions: service classification, exploration rate optimization, performance monitoring, and dynamic rule adjustment. This multi-functional controller consolidates what could be multiple separate systems into one universal management entity, reducing overall network complexity while maintaining comprehensive reliability control.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230403574A1Central node and a method for reinforcement learning in a radio access network
Publication Date: 2023.12.14 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20230403574A1 patent drawing
  • US20230403574A1 patent drawing
  • US20230403574A1 patent drawing

AI summary

A method performed by a central node for controlling an exploration strategy associated to Reinforcement Learning, RL, in one or more RL modules in a distributed node in a Radio Access Network, RAN, is provided. The central node evaluates a cost of actions performed for explorations in the one or more RL modules, and a performance of the one or more RL modules. Based on the evaluation, the central node determines one or more exploration parameters associated to the exploration strategy. The central node controls the exploration strategy by configuring the one or more RL modules with the determined one or more exploration parameters to update its exploration strategy, enforcing the respective one or more RL modules to act according to the updated exploration strategy to produce data samples for the one or more RL modules in the distributed node.