RL-Based RET Optimization for 5G Cell Interference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Optimizing Remote Electrical Tilt (RET) in 5G networks is challenging due to the complex interactions between cell performance and neighboring cells, making it difficult to determine optimal RET values that balance individual cell performance and network-wide performance.

Innovation Solution

The use of reinforcement learning (RL) agents to adjust operational parameters, such as RET, by determining reward metric values based on measurements from both individual cells and neighbor cells, and selecting actions that maximize these reward values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a cell's RET is modified to improve downlink SINR, then the cell's downlink SINR is improved, but the SINR of neighbor cells deteriorates

Engineering Contradiction:
Improvedownlink SINRVSAvoidinterference to neighbor cells
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent combines local reward information (from individual cells) and global reward information (from the entire network) into a unified reward function. This allows the RL agent to simultaneously consider both the improvement of downlink SINR in the target cell and the interference impact on neighbor cells, resolving the contradiction by merging local and global perspectives into a single decision-making framework.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements a feedback mechanism where the RL agent receives reward signals that reflect both local cell performance and global network performance. The reward function incorporates SINR measurements from the target cell and interference measurements from neighbor cells, providing feedback that guides the agent to adjust RET values that balance local improvement with global harm reduction.

Inventive Principle:
Principle #23Feedback

2Ease of manufacture

If conventional reward functions are used for RL-based optimization, then the implementation is simple, but the capture of RET change impacts on neighbor cells is insufficient

Engineering Contradiction:
Improveimplementation simplicityVSAvoidimpact information on neighbor cells
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent extends the reward function from a single-dimension (local cell reward) to a multi-dimensional structure that includes both local reward components and global reward components. By adding the global dimension that captures neighbor cell impacts, the system loses neither simplicity nor information, as the extended reward function maintains a clear structure while incorporating additional impact dimensions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If static rules defined by domain experts are used for network optimization, then the rules are universal and easy to implement, but they cannot adapt to specific network cases

Engineering Contradiction:
Improveadaptability to specific network casesVSAvoidoptimization system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a self-service mechanism where the RL agent autonomously learns and adapts optimization strategies for specific network cases without requiring manual rule modifications. The agent interacts with the network environment, receives feedback through the reward function, and automatically adjusts RET values to optimize performance for each specific network configuration, eliminating the need for expert intervention while adapting to local conditions.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250142356A1Reward for tilt optimization based on reinforcement learning (RL)
Publication Date: 2025.05.01 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20250142356A1 patent drawing
  • US20250142356A1 patent drawing
  • US20250142356A1 patent drawing

AI summary

Embodiments include computer-implemented methods for adjusting one or more operational parameters for a first cell of a communication network based on reinforcement learning (RL). Such methods include determining a plurality of reward metric values based on measurements representative of conditions in the first cell and in one or more neighbor cells of the first cell at a corresponding plurality of time instances. Such methods include determining a plurality of reward values based on differences between reward metric values at successive time instances and associating each of the reward values with a corresponding previous action that changed the one or more operational parameters. Such methods include selecting the previous action associated with a highest reward value as an action to change the one or more operational parameters. Other embodiments include RL agents and RL systems configured to perform such methods.