Antenna Tilt Policy Learning Without Live RL Exploration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing rule-based optimization strategies for controlling configurable parameters in telecommunications networks, such as Remote Electrical Tilt (RET) in 4G and 5G cellular networks, are becoming increasingly complex and time-consuming, leading to sub-optimal performance, while reinforcement learning (RL) with exploratory random actions is not applicable due to deployment constraints, and inverse propensity scoring (IPS) is difficult with continuous-valued Key Performance Indicators (KPIs.
Innovation Solution
A policy model is trained offline using a baseline dataset and inverse propensity scoring (IPS) on continuous-valued KPIs to optimize configurable parameters, enabling improved learning and deployment without exploratory random actions, utilizing a neural network to adapt weights based on reward and loss values for controlling parameters like antenna tilt.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If rule-based optimization strategies are used for controlling configurable parameters, then network performance can be improved, but the complexity and time consumption of the optimization process increases
Solution Approach 1:
The patent replaces traditional rule-based mechanical optimization systems with a neural network-based machine learning system. The neural network learns optimal parameter configurations from historical data and automatically applies them, eliminating the need for complex manual rule formulation and reducing optimization process complexity while maintaining or improving network performance.
Solution Approach 2:
The system enables self-service optimization by training the neural network to autonomously determine optimal parameter settings based on historical network data and performance metrics. The network automatically adjusts configurable parameters without requiring external intervention or complex rule-based decision-making processes, thereby reducing operational complexity.
2Adaptability or versatility
If reinforcement learning with exploratory random actions is applied, then learning capability is enhanced, but deployment becomes inapplicable due to deployment constraints
Solution Approach 1:
The patent applies preliminary action by training the neural network offline using historical network data before actual deployment. The model learns optimal parameter configurations in advance from past experiences and patterns, eliminating the need for exploratory random actions during live deployment. This pre-training approach maintains learning capability while ensuring deployment feasibility by avoiding disruptive random explorations in the production environment.
3Adaptability or versatility
If inverse propensity scoring is applied to discrete actions, then policy learning is improved, but it becomes difficult when dealing with continuous-valued Key Performance Indicators
Solution Approach 1:
The patent applies parameter changes by transforming the neural network's output layer to directly predict continuous parameter values instead of discrete actions. The network learns to output optimal parameter configurations continuously, and inverse propensity scoring is adapted to work with these continuous predictions by comparing predicted versus actual parameter values and their outcomes, thereby maintaining policy learning effectiveness while handling continuous-valued KPIs.
Data Source
AI summary
A method performed by a computer system for a telecommunications network. The computer system can access a network metrics repository to retrieve a baseline dataset collected from a baseline policy deployed in the telecommunications network for controlling a configurable parameter of the telecommunications network. The configurable parameter includes an antenna tilt degree. The baseline dataset includes key performance indicators (K PIs) that include K PIs having a continuous value and a plurality of historical changes made to the configurable parameter. The computer system can train a policy model while offline the telecommunications network using the baseline dataset and inverse propensity scoring on the input K PIs having continuous values to output from the policy model a probability of actions for controlling the configurable parameter. A method performed by network node or network nodes is also provided for using a trained policy model to control the configuration parameter.


