Transfer Deep Reinforcement Learning for Base Station Energy Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing demand for wireless data services and the deployment of new cellular base stations lead to significant energy consumption and greenhouse gas emissions in 5G wireless communication networks, with existing reinforcement learning techniques requiring a large number of interactions, limiting their applicability in real-world scenarios.

Innovation Solution

An unsupervised transfer deep reinforcement learning framework is applied to select and implement an energy-saving control policy for target base stations, using pre-trained policies from source base stations to reduce energy consumption dynamically, without the need for excessive training iterations or data samples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If reinforcement learning is applied to optimize network control policies for energy saving, then energy consumption can be reduced, but a large number of interactions with the network system are required which limits real-world applicability

Engineering Contradiction:
Improveenergy consumptionVSAvoidtraining interactions
Core Design Contradiction:
Loss of energyVSLoss of time

Solution Approach 1:

The patent applies transfer learning to pre-train reinforcement learning models on source base stations before deploying them to target base stations. This preliminary training action on source domains enables the model to acquire useful policies in advance, reducing the number of interactions needed when applied to target base stations in real-world scenarios.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an unsupervised transfer reinforcement learning framework as an intermediary between source base stations and target base stations. This framework transfers knowledge and policies from source domains to target domains, enabling efficient energy optimization on target base stations without requiring extensive direct interactions with them.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If deep reinforcement learning is used to learn network control policies, then network performance can be improved, but a large number of data samples and interactions are required which constrains practical deployment

Engineering Contradiction:
Improvenetwork performanceVSAvoiddata samples
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent copies trained reinforcement learning policies from source base stations to target base stations through the transfer learning framework. Instead of collecting large amounts of data samples from each target base station, the system replicates successful policies from source domains, maintaining high network performance while drastically reducing data sample requirements.

Inventive Principle:
Principle #26Copying

3Productivity

If more base stations are deployed to meet increasing wireless data demand, then service capacity increases, but energy consumption and greenhouse gas emissions increase significantly

Engineering Contradiction:
Improveservice capacityVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent dynamically changes network control parameters such as cell activation states, resource allocation, and transmission power levels using reinforcement learning policies. These parameter adjustments enable the network to maintain high service capacity while optimizing energy consumption by adapting parameters to current network conditions and traffic demands.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240406861A1Energy saving in cellular wireless networks via transfer deep reinforcement learning
Publication Date: 2024.12.05 SAMSUNG ELECTRONICS CO LTD
  • US20240406861A1 patent drawing
  • US20240406861A1 patent drawing
  • US20240406861A1 patent drawing

AI summary

The present disclosure provides methods, apparatuses, systems, and computer-readable mediums for operating a target base station by an apparatus. A method includes collecting a plurality of trajectories corresponding to the target base station and a plurality of source base stations, clustering, using an unsupervised reinforcement learning model, the plurality of trajectories into a plurality of clusters including a target cluster, selecting, as a target trajectory, a selected trajectory from the target cluster that maximizes an energy-saving parameter of the target base station, and applying, to the target base station, an energy-saving control policy corresponding to the target trajectory. The target cluster corresponds to the target base station and at least one source base station from among the plurality of source base stations.