Two-Stage Deep Reinforcement Learning for Reactive Voltage Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current power grid reactive voltage control methods face challenges due to model incompleteness, leading to low credibility of power grid model parameters and frequent changes, which result in sub-optimal control and increased network losses, necessitating the development of a more efficient and safer data-driven approach.
Innovation Solution
A two-stage deep reinforcement learning method is employed, involving offline training of a reactive voltage control model using a Soft Actor-Critic algorithm within an interactive training environment based on Markov decision processes, followed by online deployment and continuous updating to adapt to changing grid conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional model-based reactive voltage control is used, then control can be performed with approximate models, but control optimality cannot be guaranteed and voltage violations may worsen
Solution Approach 1:
The patent creates a virtual copy of the power grid system through digital twin technology, building a virtual power grid model that mirrors the physical system. This virtual model is used for training the reinforcement learning agent, allowing the control strategy to be optimized in a virtual environment before deployment to the actual physical system, thus avoiding direct trial-and-error on the real grid.
Solution Approach 2:
The patent performs preliminary training of the reinforcement learning agent in advance using historical data and virtual environment simulations. The agent learns optimal control strategies beforehand through extensive training in the virtual power grid, accumulating experience and optimizing its policy before being deployed to the actual physical system, thereby avoiding suboptimal control during initial operation.
2Reliability
If deep reinforcement learning is used for online training, then optimal reactive voltage control can be achieved, but training efficiency and safety are reduced
Solution Approach 1:
The patent performs the majority of training work in advance offline using historical data and virtual environment simulations. The reinforcement learning agent is pre-trained extensively before deployment, so that when deployed to the physical system, it requires minimal online training and can immediately provide optimal control, thus achieving both optimality and high training efficiency.
Solution Approach 2:
The patent uses a virtual copy of the power grid for training purposes, allowing the reinforcement learning agent to learn from simulated experiences without affecting the actual physical system. This virtual training environment enables efficient exploration of different control strategies and states without safety risks or operational disruptions to the real grid.
3Reliability
If deep reinforcement learning is used for online training, then optimal reactive voltage control can be achieved, but safety is reduced
Solution Approach 1:
The patent prepares compensatory measures in advance by training the reinforcement learning agent extensively in a virtual environment with various simulated fault conditions and edge cases. The agent learns to handle abnormal situations and voltage violations through pre-training, so that when deployed to the physical system, it is already prepared to safely handle unexpected conditions without causing harm.
Solution Approach 2:
The patent uses a virtual copy of the power grid for training purposes, allowing the reinforcement learning agent to learn from simulated experiences without affecting the actual physical system. This virtual training environment enables safe exploration of potentially dangerous control actions and system states, eliminating safety risks associated with online training on the real grid.
4Measurement precision
If model parameters are frequently updated to reflect grid changes, then model accuracy improves, but model maintenance difficulty increases
Solution Approach 1:
The patent implements a dynamic model update mechanism where the virtual power grid model automatically adapts to changes in the physical system. The model parameters are dynamically adjusted based on real-time data from the physical grid, allowing the virtual model to continuously track and reflect the actual system state without requiring manual intervention for frequent updates.
Solution Approach 2:
The patent enables the virtual model to self-update and self-correct by automatically ingesting data from the physical system and adjusting its parameters accordingly. The system performs self-validation and self-calibration, reducing the need for manual model maintenance and parameter adjustment while maintaining high accuracy.
Data Source
AI summary
A power grid reactive voltage control method and control system based on two-stage deep reinforcement learning, comprising steps of: building interactive training environment based on Markov decision process, according to a regional power grid simulation model and a reactive voltage optimization model; training a reactive voltage control model offline by using a SAC algorithm, in the interactive training environment based on Markov decision process; deploying the reactive voltage control model to a regional power grid online system; and acquiring operating state information of the regional power grid, updating the reactive voltage control model, and generating an optimal reactive voltage control policy. As compared with the existing power grid optimizing method based on reinforcement learning, the online control training according to the present disclosure has costs and safety hazards greatly reduced, and is more suitable for deployment in an actual power system.


