Two-Stage Deep Reinforcement Learning for Reactive Voltage Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current power grid reactive voltage control methods face challenges due to model incompleteness, leading to low credibility of power grid model parameters and frequent changes, which result in sub-optimal control and increased network losses, necessitating the development of a more efficient and safer data-driven approach.

Innovation Solution

A two-stage deep reinforcement learning method is employed, involving offline training of a reactive voltage control model using a Soft Actor-Critic algorithm within an interactive training environment based on Markov decision processes, followed by online deployment and continuous updating to adapt to changing grid conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional model-based reactive voltage control is used, then control can be performed with approximate models, but control optimality cannot be guaranteed and voltage violations may worsen

Engineering Contradiction:
Improveease of control implementationVSAvoidcontrol optimality
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent creates a virtual copy of the power grid system through digital twin technology, building a virtual power grid model that mirrors the physical system. This virtual model is used for training the reinforcement learning agent, allowing the control strategy to be optimized in a virtual environment before deployment to the actual physical system, thus avoiding direct trial-and-error on the real grid.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary training of the reinforcement learning agent in advance using historical data and virtual environment simulations. The agent learns optimal control strategies beforehand through extensive training in the virtual power grid, accumulating experience and optimizing its policy before being deployed to the actual physical system, thereby avoiding suboptimal control during initial operation.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If deep reinforcement learning is used for online training, then optimal reactive voltage control can be achieved, but training efficiency and safety are reduced

Engineering Contradiction:
Improvecontrol optimalityVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs the majority of training work in advance offline using historical data and virtual environment simulations. The reinforcement learning agent is pre-trained extensively before deployment, so that when deployed to the physical system, it requires minimal online training and can immediately provide optimal control, thus achieving both optimality and high training efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses a virtual copy of the power grid for training purposes, allowing the reinforcement learning agent to learn from simulated experiences without affecting the actual physical system. This virtual training environment enables efficient exploration of different control strategies and states without safety risks or operational disruptions to the real grid.

Inventive Principle:
Principle #26Copying

3Reliability

If deep reinforcement learning is used for online training, then optimal reactive voltage control can be achieved, but safety is reduced

Engineering Contradiction:
Improvecontrol optimalityVSAvoidsafety risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent prepares compensatory measures in advance by training the reinforcement learning agent extensively in a virtual environment with various simulated fault conditions and edge cases. The agent learns to handle abnormal situations and voltage violations through pre-training, so that when deployed to the physical system, it is already prepared to safely handle unexpected conditions without causing harm.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Solution Approach 2:

The patent uses a virtual copy of the power grid for training purposes, allowing the reinforcement learning agent to learn from simulated experiences without affecting the actual physical system. This virtual training environment enables safe exploration of potentially dangerous control actions and system states, eliminating safety risks associated with online training on the real grid.

Inventive Principle:
Principle #26Copying

4Measurement precision

If model parameters are frequently updated to reflect grid changes, then model accuracy improves, but model maintenance difficulty increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidmodel maintenance difficulty
Core Design Contradiction:
Measurement precisionVSEase of repair

Solution Approach 1:

The patent implements a dynamic model update mechanism where the virtual power grid model automatically adapts to changes in the physical system. The model parameters are dynamically adjusted based on real-time data from the physical grid, allowing the virtual model to continuously track and reflect the actual system state without requiring manual intervention for frequent updates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent enables the virtual model to self-update and self-correct by automatically ingesting data from the physical system and adjusting its parameters accordingly. The system performs self-validation and self-calibration, reducing the need for manual model maintenance and parameter adjustment while maintaining high accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11442420B2Power grid reactive voltage control method based on two-stage deep reinforcement learning
Publication Date: 2022.09.13 TSINGHUA UNIVERSITY
  • US11442420B2 patent drawing
  • US11442420B2 patent drawing
  • US11442420B2 patent drawing

AI summary

A power grid reactive voltage control method and control system based on two-stage deep reinforcement learning, comprising steps of: building interactive training environment based on Markov decision process, according to a regional power grid simulation model and a reactive voltage optimization model; training a reactive voltage control model offline by using a SAC algorithm, in the interactive training environment based on Markov decision process; deploying the reactive voltage control model to a regional power grid online system; and acquiring operating state information of the regional power grid, updating the reactive voltage control model, and generating an optimal reactive voltage control policy. As compared with the existing power grid optimizing method based on reinforcement learning, the online control training according to the present disclosure has costs and safety hazards greatly reduced, and is more suitable for deployment in an actual power system.