Deep Reinforcement Learning Agent for Power Grid Voltage and Flow Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge in power systems is to derive real-time, system-wide optimal controls for active and reactive power flow while ensuring security constraints, which is complicated by the non-convex and dynamic nature of the ACOPF problem, leading to suboptimal solutions and reliance on simplified models like DC-based OPF.

Innovation Solution

The implementation of an autonomous multi-objective control model using Deep Reinforcement Learning (DRL) agents, trained with a Markov decision process (MDP) to regulate voltage profiles, line flows, and transmission losses, enabling data-driven, real-time control strategies and optimizing power controller actions in dynamic environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional optimization methods (nonlinear programming, quadratic programming, Lagrangian relaxation) are used to solve ACOPF problem, then solution accuracy is improved, but computation time increases significantly making real-time control infeasible

Engineering Contradiction:
Improvesolution accuracyVSAvoidcomputation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary training of deep reinforcement learning agents offline using massive representative operating conditions and contingencies. This preliminary action pre-computes optimal control strategies for various scenarios, enabling real-time deployment without requiring complex online optimization computations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical optimization algorithms (nonlinear programming, quadratic programming, interior point method) with an intelligent agent-based system using deep reinforcement learning. This substitution transforms the computational approach from deterministic mathematical programming to data-driven autonomous decision-making, achieving both speed and accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If DC-based OPF models are used to obtain fast solutions, then computation speed is improved, but solution optimality deteriorates due to model simplification

Engineering Contradiction:
Improvecomputation speedVSAvoidsolution optimality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system changes the operational parameters and training conditions by exposing the reinforcement learning agent to massive representative operating conditions including various contingencies during offline training. This enables the agent to learn accurate AC power flow characteristics without requiring simplified DC models during real-time operation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The deep reinforcement learning agent autonomously learns optimal control strategies through self-service training processes, exploring the solution space independently without relying on simplified mathematical models. The agent develops its own understanding of AC power flow dynamics through interaction with simulation environments.

Inventive Principle:
Principle #25Self-service

3Reliability

If system-wide optimal control is achieved by coordinating many controllers, then control effectiveness is improved, but system complexity increases making real-time coordination challenging

Engineering Contradiction:
Improvecontrol effectivenessVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple distributed controllers into a unified deep reinforcement learning agent that performs system-wide optimization. Instead of coordinating many independent controllers, the merged agent directly computes coordinated control actions for all controllable resources, simplifying the architectural complexity while maintaining system-wide optimality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The deep reinforcement learning agent serves multiple functions simultaneously: it performs system-state assessment, identifies optimal control actions, coordinates distributed resources, and ensures constraint compliance all within a single unified framework, replacing multiple specialized components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If deep reinforcement learning agents are trained with massive representative operating conditions and contingencies, then control robustness is improved, but training complexity and computational resources increase

Engineering Contradiction:
Improvecontrol robustnessVSAvoidtraining complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs the computationally intensive training process with massive representative operating conditions and contingencies as a preliminary offline action. This shifts the computational burden from real-time operation to offline preparation, allowing robust training without impacting real-time control performance or operational complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11336092B2Multi-objective real-time power flow control method using soft actor-critic
Publication Date: 2022.05.17 DIAO RUISHENG
  • US11336092B2 patent drawing
  • US11336092B2 patent drawing
  • US11336092B2 patent drawing

AI summary

Systems and methods are disclosed for control voltage profiles, line flows and transmission losses of a power grid by forming an autonomous multi-objective control model with one or more neural networks as a Deep Reinforcement Learning (DRL) agent; training the DRL agent to provide data-driven, real-time and autonomous grid control strategies; and coordinating and optimizing power controllers to regulate voltage profiles, line flows and transmission losses in the power grid with a Markov decision process (MDP) operating with reinforcement learning to control problems in dynamic and stochastic environments.