Power Grid Reactive Voltage Control Model Training via Adversarial Markov Decision Process
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing power grid reactive voltage control methods face challenges with high training costs and safety risks due to the inefficiency of online training and model incompleteness, particularly with the lack of transferability of deep reinforcement learning models from offline to online systems, leading to sub-optimal control effects and operational inefficiencies.
Innovation Solution
A power grid reactive voltage control model training method is developed, which establishes a simulation model, builds an interactive training environment based on Adversarial Markov Decision Process, and uses a joint adversarial training algorithm to train a transferable model that can be safely and efficiently applied online, reducing the need for extensive online training and minimizing control deviations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data-driven deep reinforcement learning methods are used for power grid reactive voltage control, then control flexibility and adaptability are improved, but online training costs increase and security risks arise
Solution Approach 1:
The patent applies preliminary action by conducting offline training on a power grid simulation model before deploying the deep reinforcement learning model to the online system. This pre-training phase allows the model to learn optimal control strategies in advance, so that when deployed online, it can directly apply learned knowledge without requiring extensive online training, thereby reducing security risks and training costs while maintaining adaptability
2Loss of time
If offline training on simulation model is performed, then online training costs are reduced, but model deviation occurs and transferability is lost
Solution Approach 1:
The patent implements feedback by establishing a feedback mechanism between the offline simulation model and the online power grid system. The model learns from simulation environments and continuously adapts to real-world conditions through feedback loops, ensuring that knowledge transferred from offline training remains accurate and applicable online, thus maintaining both efficiency and transferability
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting model parameters and training configurations based on the specific characteristics of the power grid system. This allows the model to adapt its behavior to match real-world conditions while retaining the benefits of offline training, resolving the contradiction between training efficiency and model transferability
3Stability of the object's composition
If traditional model-based optimizing method is used, then control stability is maintained, but control effect deteriorates due to model incompleteness
Solution Approach 1:
The patent applies copying by creating a detailed simulation model that replicates the power grid system's behavior. This virtual copy allows the deep reinforcement learning model to learn optimal control strategies without affecting the actual power grid, combining the stability of simulation-based training with the precision of data-driven approaches when deployed online
Data Source
AI summary
A power grid reactive voltage control model training method. The method comprises: establishing a power grid simulation model; establishing a reactive voltage optimization model, according to a power grid reactive voltage control target; building interactive training environment based on Adversarial Markov Decision Process, in combination with the power grid simulation model and the reactive voltage optimization model; training the power grid reactive voltage control model through a joint adversarial training algorithm; and transferring the trained power grid reactive voltage control model to an online system. The power grid reactive voltage control model trained by using the method according to the present disclosure has transferability as compared with the traditional method, and may be directly used for online power grid reactive voltage control.


