Power Grid Reactive Voltage Control Model Training via Adversarial Markov Decision Process

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing power grid reactive voltage control methods face challenges with high training costs and safety risks due to the inefficiency of online training and model incompleteness, particularly with the lack of transferability of deep reinforcement learning models from offline to online systems, leading to sub-optimal control effects and operational inefficiencies.

Innovation Solution

A power grid reactive voltage control model training method is developed, which establishes a simulation model, builds an interactive training environment based on Adversarial Markov Decision Process, and uses a joint adversarial training algorithm to train a transferable model that can be safely and efficiently applied online, reducing the need for extensive online training and minimizing control deviations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data-driven deep reinforcement learning methods are used for power grid reactive voltage control, then control flexibility and adaptability are improved, but online training costs increase and security risks arise

Engineering Contradiction:
Improvecontrol adaptabilityVSAvoidsecurity risk
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies preliminary action by conducting offline training on a power grid simulation model before deploying the deep reinforcement learning model to the online system. This pre-training phase allows the model to learn optimal control strategies in advance, so that when deployed online, it can directly apply learned knowledge without requiring extensive online training, thereby reducing security risks and training costs while maintaining adaptability

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If offline training on simulation model is performed, then online training costs are reduced, but model deviation occurs and transferability is lost

Engineering Contradiction:
Improveonline training timeVSAvoidmodel transferability
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The patent implements feedback by establishing a feedback mechanism between the offline simulation model and the online power grid system. The model learns from simulation environments and continuously adapts to real-world conditions through feedback loops, ensuring that knowledge transferred from offline training remains accurate and applicable online, thus maintaining both efficiency and transferability

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies parameter changes by dynamically adjusting model parameters and training configurations based on the specific characteristics of the power grid system. This allows the model to adapt its behavior to match real-world conditions while retaining the benefits of offline training, resolving the contradiction between training efficiency and model transferability

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If traditional model-based optimizing method is used, then control stability is maintained, but control effect deteriorates due to model incompleteness

Engineering Contradiction:
Improvecontrol stabilityVSAvoidcontrol precision
Core Design Contradiction:
Stability of the object's compositionVSManufacturing precision

Solution Approach 1:

The patent applies copying by creating a detailed simulation model that replicates the power grid system's behavior. This virtual copy allows the deep reinforcement learning model to learn optimal control strategies without affecting the actual power grid, combining the stability of simulation-based training with the precision of data-driven approaches when deployed online

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11689021B2Power grid reactive voltage control model training method and system
Publication Date: 2023.06.27 TSINGHUA UNIVERSITY
  • US11689021B2 patent drawing
  • US11689021B2 patent drawing
  • US11689021B2 patent drawing

AI summary

A power grid reactive voltage control model training method. The method comprises: establishing a power grid simulation model; establishing a reactive voltage optimization model, according to a power grid reactive voltage control target; building interactive training environment based on Adversarial Markov Decision Process, in combination with the power grid simulation model and the reactive voltage optimization model; training the power grid reactive voltage control model through a joint adversarial training algorithm; and transferring the trained power grid reactive voltage control model to an online system. The power grid reactive voltage control model trained by using the method according to the present disclosure has transferability as compared with the traditional method, and may be directly used for online power grid reactive voltage control.