Microgrid RL Agent Consensus for Wildfire Outage Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for handling power outages caused by wildfires are inefficient and unable to manage severe outages in modern power systems, leading to increased computing, networking, and transportation resource consumption, as well as safety violations and legal issues.

Innovation Solution

A reinforcement learning (RL) system manages RL agents using multi-criteria group consensus in a localized microgrid cluster, modeling the network as a spatiotemporal representation to identify consensus master RL agents and control the microgrid environment, thereby reducing wildfire risk and improving power continuity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current techniques are used to handle power outages caused by wildfires, then power outage management is performed, but computing, networking, and transportation resource consumption increases and response efficiency decreases

Engineering Contradiction:
Improveresponse efficiencyVSAvoidresource consumption
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent divides the microgrid into multiple localized clusters, each managed by its own RL agents. This segmentation allows autonomous decision-making at the cluster level, reducing the computational burden on centralized systems and improving response efficiency while lowering overall resource consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts the behavior and communication of RL agents based on real-time grid conditions, wildfire risks, and resource availability. This dynamic adaptation enables the system to optimize resource consumption and response efficiency under varying operational scenarios.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If a centralized RL system is used to manage the microgrid, then control decisions are made, but communication delays and system complexity increase

Engineering Contradiction:
Improvecontrol decision-makingVSAvoidsystem architecture complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The centralized RL system is segmented into distributed RL agents operating at the cluster level. Each agent independently makes control decisions for its local cluster, eliminating the need for complex centralized communication and decision-making, thereby reducing system complexity and communication delays.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each RL agent autonomously manages its local microgrid cluster without requiring constant centralized coordination. The agents self-organize and make independent control decisions, simplifying the overall system architecture while maintaining effective control.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If RL agents communicate frequently to achieve consensus, then consensus accuracy improves, but communication overhead and time consumption increase

Engineering Contradiction:
Improveconsensus accuracyVSAvoidcommunication time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system implements a threshold-based consensus mechanism where RL agents communicate only when necessary to achieve sufficient agreement. This partial communication approach maintains adequate consensus accuracy while significantly reducing communication overhead and time consumption compared to continuous communication protocols.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4258172A1Managing reinforcement learning agents using multi-criteria group consensus in a localized microgrid cluster
Publication Date: 2023.10.11 ACCENTURE GLOBAL SOLUTIONS LTD
  • EP4258172A1 patent drawingFigure 1A
  • EP4258172A1 patent drawingFigure 1B
  • EP4258172A1 patent drawingFigure 1C

AI summary

A device may receive state data, actions, and rewards associated with a network of RL agents monitoring a microgrid environment, and may model the network of RL agents as a spatiotemporal representation. The device may represent interactions of the RL agents as edge attributes in the spatiotemporal representation, and may determine edge attributes, transmissibility, connectedness, and communication delay for each of the RL agents in the spatiotemporal representation. The device may determine, based on the transmissibility, the connectedness, and the communication delay, localized clusters of the RL agents, and may process the localized clusters, with a first machine learning model, to identify consensus master RL agents. The device may process the consensus master RL agents, with a second machine learning model, to identify a final master RL agent for the network of RL agents, and cause the final master RL agent to control the microgrid environment.