Antenna Tilt Policy Control Using Offline IPS Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for optimizing configurable parameters in telecommunications networks, such as Remote Electrical Tilt (RET) antenna angle control, are often based on rule-based policies that can lead to sub-optimal performance due to increasing network complexity, and reinforcement learning approaches are not deployable in customer networks due to the need for exploratory random actions.
Innovation Solution
A method involving offline training of a policy model using a baseline dataset and inverse propensity scoring (IPS) on continuous-valued Key Performance Indicators (KPIs) to determine the probability of actions for controlling configurable parameters, such as antenna tilt, without requiring exploratory actions in live networks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If rule-based policies are used for optimizing configurable parameters, then the control method is simple and easy to implement, but the performance becomes sub-optimal due to increasing network complexity
Solution Approach 1:
The system performs preliminary action by collecting historical configuration data and KPI data offline before deployment, training the policy model in advance using inverse propensity scoring to eliminate selection bias, and preparing the model for deployment without requiring exploratory actions in the live network
Solution Approach 2:
The patent replaces the mechanical rule-based control system with an intelligent policy model based on inverse propensity scoring that can adapt to increasing network complexity while maintaining automated operation, substituting fixed rules with a data-driven decision-making system
2Reliability
If reinforcement learning approaches are used for optimizing configurable parameters, then the optimization performance can be improved, but the method cannot be deployed in customer networks due to the need for exploratory random actions
Solution Approach 1:
The system performs all learning and training actions preliminarily offline using historical data collected from the live network during normal operation with the baseline policy, eliminating the need for exploratory random actions in the deployed system while still achieving improved optimization performance
Solution Approach 2:
The patent introduces an intermediary offline training phase that acts as a mediator between the baseline policy data collection and the final policy model deployment, allowing the system to learn optimal policies without disrupting live network operations or requiring exploratory actions in the deployed system
3Ease of operation
If offline training with inverse propensity scoring is used, then the policy model can be deployed without exploratory actions, but the training process becomes more complex
Solution Approach 1:
The system creates a copy of the baseline policy's historical data and operates on this replicated dataset during offline training, allowing the inverse propensity scoring training process to proceed without affecting the live network while using the same data structure and format
Data Source
AI summary
A method performed by a computer system for a telecommunications network. The computer system can access a network metrics repository to retrieve a baseline dataset collected from a baseline policy deployed in the telecommunications network for controlling a configurable parameter of the telecommunications network. The configurable parameter includes an antenna tilt degree. The baseline dataset includes key performance indicators (KPIs) that include KPIs having a continuous value and a plurality of historical changes made to the configurable parameter. The computer system can train a policy model while offline the telecommunications network using the baseline dataset and inverse propensity scoring on the input KPIs having continuous values to output from the policy model a probability of actions for controlling the configurable parameter. A method performed by network node or network nodes is also provided for using a trained policy model to control the configuration parameter.


