Transfer Deep Reinforcement Learning for Base Station Energy Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for wireless data services and the deployment of new cellular base stations lead to significant energy consumption and greenhouse gas emissions in 5G wireless communication networks, with existing reinforcement learning techniques requiring a large number of interactions, limiting their applicability in real-world scenarios.
Innovation Solution
An unsupervised transfer deep reinforcement learning framework is applied to select and implement an energy-saving control policy for target base stations, using pre-trained policies from source base stations to reduce energy consumption dynamically, without the need for excessive training iterations or data samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If reinforcement learning is applied to optimize network control policies for energy saving, then energy consumption can be reduced, but a large number of interactions with the network system are required which limits real-world applicability
Solution Approach 1:
The patent applies transfer learning to pre-train reinforcement learning models on source base stations before deploying them to target base stations. This preliminary training action on source domains enables the model to acquire useful policies in advance, reducing the number of interactions needed when applied to target base stations in real-world scenarios.
Solution Approach 2:
The patent introduces an unsupervised transfer reinforcement learning framework as an intermediary between source base stations and target base stations. This framework transfers knowledge and policies from source domains to target domains, enabling efficient energy optimization on target base stations without requiring extensive direct interactions with them.
2Productivity
If deep reinforcement learning is used to learn network control policies, then network performance can be improved, but a large number of data samples and interactions are required which constrains practical deployment
Solution Approach 1:
The patent copies trained reinforcement learning policies from source base stations to target base stations through the transfer learning framework. Instead of collecting large amounts of data samples from each target base station, the system replicates successful policies from source domains, maintaining high network performance while drastically reducing data sample requirements.
3Productivity
If more base stations are deployed to meet increasing wireless data demand, then service capacity increases, but energy consumption and greenhouse gas emissions increase significantly
Solution Approach 1:
The patent dynamically changes network control parameters such as cell activation states, resource allocation, and transmission power levels using reinforcement learning policies. These parameter adjustments enable the network to maintain high service capacity while optimizing energy consumption by adapting parameters to current network conditions and traffic demands.
Data Source
AI summary
The present disclosure provides methods, apparatuses, systems, and computer-readable mediums for operating a target base station by an apparatus. A method includes collecting a plurality of trajectories corresponding to the target base station and a plurality of source base stations, clustering, using an unsupervised reinforcement learning model, the plurality of trajectories into a plurality of clusters including a target cluster, selecting, as a target trajectory, a selected trajectory from the target cluster that maximizes an energy-saving parameter of the target base station, and applying, to the target base station, an energy-saving control policy corresponding to the target trajectory. The target cluster corresponds to the target base station and at least one source base station from among the plurality of source base stations.


