Deep Reinforcement Learning Agent for Power Grid Voltage and Flow Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in power systems is to derive real-time, system-wide optimal controls for active and reactive power flow while ensuring security constraints, which is complicated by the non-convex and dynamic nature of the ACOPF problem, leading to suboptimal solutions and reliance on simplified models like DC-based OPF.
Innovation Solution
The implementation of an autonomous multi-objective control model using Deep Reinforcement Learning (DRL) agents, trained with a Markov decision process (MDP) to regulate voltage profiles, line flows, and transmission losses, enabling data-driven, real-time control strategies and optimizing power controller actions in dynamic environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional optimization methods (nonlinear programming, quadratic programming, Lagrangian relaxation) are used to solve ACOPF problem, then solution accuracy is improved, but computation time increases significantly making real-time control infeasible
Solution Approach 1:
The system performs preliminary training of deep reinforcement learning agents offline using massive representative operating conditions and contingencies. This preliminary action pre-computes optimal control strategies for various scenarios, enabling real-time deployment without requiring complex online optimization computations.
Solution Approach 2:
The patent replaces traditional mechanical optimization algorithms (nonlinear programming, quadratic programming, interior point method) with an intelligent agent-based system using deep reinforcement learning. This substitution transforms the computational approach from deterministic mathematical programming to data-driven autonomous decision-making, achieving both speed and accuracy.
2Productivity
If DC-based OPF models are used to obtain fast solutions, then computation speed is improved, but solution optimality deteriorates due to model simplification
Solution Approach 1:
The system changes the operational parameters and training conditions by exposing the reinforcement learning agent to massive representative operating conditions including various contingencies during offline training. This enables the agent to learn accurate AC power flow characteristics without requiring simplified DC models during real-time operation.
Solution Approach 2:
The deep reinforcement learning agent autonomously learns optimal control strategies through self-service training processes, exploring the solution space independently without relying on simplified mathematical models. The agent develops its own understanding of AC power flow dynamics through interaction with simulation environments.
3Reliability
If system-wide optimal control is achieved by coordinating many controllers, then control effectiveness is improved, but system complexity increases making real-time coordination challenging
Solution Approach 1:
The patent merges multiple distributed controllers into a unified deep reinforcement learning agent that performs system-wide optimization. Instead of coordinating many independent controllers, the merged agent directly computes coordinated control actions for all controllable resources, simplifying the architectural complexity while maintaining system-wide optimality.
Solution Approach 2:
The deep reinforcement learning agent serves multiple functions simultaneously: it performs system-state assessment, identifies optimal control actions, coordinates distributed resources, and ensures constraint compliance all within a single unified framework, replacing multiple specialized components.
4Adaptability or versatility
If deep reinforcement learning agents are trained with massive representative operating conditions and contingencies, then control robustness is improved, but training complexity and computational resources increase
Solution Approach 1:
The system performs the computationally intensive training process with massive representative operating conditions and contingencies as a preliminary offline action. This shifts the computational burden from real-time operation to offline preparation, allowing robust training without impacting real-time control performance or operational complexity.
Data Source
AI summary
Systems and methods are disclosed for control voltage profiles, line flows and transmission losses of a power grid by forming an autonomous multi-objective control model with one or more neural networks as a Deep Reinforcement Learning (DRL) agent; training the DRL agent to provide data-driven, real-time and autonomous grid control strategies; and coordinating and optimizing power controllers to regulate voltage profiles, line flows and transmission losses in the power grid with a Markov decision process (MDP) operating with reinforcement learning to control problems in dynamic and stochastic environments.


