Network Energy Storage Control Using Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for managing energy storages in cellular networks do not effectively consider network load, RBS locations, or load-dependent energy consumption, leading to inefficiencies in energy utilization and increased carbon footprint and operational costs.
Innovation Solution
A computer-implemented method using reinforcement learning to manage energy storages across multiple sites in a network, where a simulated environment is created based on power consumption data, and a reinforcement learning system is trained to optimize energy storage actions such as charging and discharging, considering constraints and aiming to maximize reward based on energy cost and carbon footprint reduction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If traditional energy management methods are used in cellular networks, then network operations can be maintained, but energy utilization efficiency is low and carbon footprint increases
Solution Approach 1:
The patent implements dynamic energy management by training reinforcement learning agents that continuously adapt charging and discharging decisions based on real-time network load conditions, traffic patterns, and energy storage states. This dynamic approach replaces static traditional methods, optimizing energy utilization efficiency while reducing carbon footprint through load-dependent energy consumption management.
Solution Approach 2:
The system employs feedback mechanisms where reinforcement learning agents receive continuous monitoring data on network load, energy storage levels, and power consumption, then adjust their charging and discharging actions accordingly. This closed-loop feedback enables the system to learn from past decisions and improve energy management performance, directly addressing the contradiction between energy efficiency and carbon emissions.
2Reliability
If energy storages are managed independently at each site, then local power supply stability is maintained, but overall network energy utilization efficiency is reduced
Solution Approach 1:
The patent merges independent site-level energy management into a coordinated network-wide system by deploying reinforcement learning agents that consider inter-site energy transfer opportunities. The system combines local reliability requirements with network-wide optimization, allowing energy storages at different sites to be charged or discharged based on overall network conditions rather than isolated local conditions, thereby improving both reliability and energy utilization efficiency.
Solution Approach 2:
The reinforcement learning framework provides a universal energy management approach that can be applied across multiple sites with different local conditions. The same RL agent architecture adapts to various site configurations, load patterns, and energy storage capacities, enabling coordinated management that maintains local reliability while optimizing network-wide energy utilization through multi-functional decision-making.
3Measurement precision
If reinforcement learning is trained on historical power consumption data, then accurate predictions can be achieved, but training time and computational resources increase
Solution Approach 1:
The patent applies partial action by training reinforcement learning agents on selectively relevant historical power consumption data rather than complete datasets. The system identifies and prioritizes key features and time periods that most influence energy management decisions, achieving sufficient prediction accuracy for practical deployment while reducing training time and computational resource requirements through focused data selection.
Data Source
AI summary
There is provided a computer-implemented method for managing a plurality of energy storages at a plurality of sites in a network, the method comprising: acquiring a first dataset including power consumption data of at least a subset of the plurality of energy storages over a predetermined amount of time; generating a first simulated environment of the network based on the acquired first dataset; and training a first reinforcement learning system by performing the following steps iteratively until a termination condition is met; selecting an action from a set of feasible actions, wherein each action in the set of feasible action is bounded by a set of constraints; calculating a reward of the selected action based on the generated first simulated environment of the network; and training the first reinforcement learning system to maximise reward for a given state of the network, based on the calculated reward for the selected action.


