Multi-agent reinforcement learning peak shaving method for ultra-large power grid source load dynamic game
By employing a multi-agent reinforcement learning approach, combined with hierarchical training and causal structure analysis, the dynamic game problem involving diverse participants in ultra-large power grids was solved, enabling efficient and reliable operation of the power grid.
Patent Information
- Application Number
- CN202510877659.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
Existing power grid peak-shaving methods struggle to effectively coordinate the dynamic game relationships among diverse stakeholders when dealing with ultra-large power grids, leading to global optimization challenges. In particular, their adaptability and robustness are significantly reduced in scenarios with highly dynamic load changes, with traditional methods exhibiting a failure rate as high as 30-45%.
By employing a multi-agent reinforcement learning approach, a collaborative-competitive hybrid training mechanism is constructed through a hierarchical training mechanism and causal structure analysis. Combined with entropy embedding technology, this enables precise control of complex power grid systems.
It significantly improves the accuracy and robustness of power grid peak shaving, reduces system energy consumption, enhances the adaptability and stability of the power grid under extreme scenarios, and achieves balanced optimization of multi-dimensional objectives.
Smart Images

Figure CN120996408A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power system regulation, in particular, to a multi-agent reinforcement learning peak regulation method for super-large power grid source-load dynamic game. BACKGROUND
[0002] With the acceleration of urbanization and the promotion of new infrastructure construction, the scale of modern super-large urban power grids continues to expand, and the electricity load presents the characteristics of high density and high volatility. Especially during the peak period of weekdays, seasonal peak electricity consumption and special events, the power grid faces tremendous peak regulation pressure. How to effectively manage the peak-valley difference of power grid load and improve the stability and economy of the power grid has become a key challenge in modern power grid management. According to the research of Liu et al., the daily load peak-valley difference of super-large urban power grids can reach 30-50%, which is significantly higher than the 15-25% of traditional small and medium-sized urban power grids. This difference is mainly due to the complex load structure of large cities and the aggregation effect of extreme electricity consumption behavior [Liu et al., IEEE Transactions on Smart Grid, 2023, Smart contract assisted privacy-preserving data aggregation and management scheme for smart grid ] .
[0003] In traditional power grid peak regulation methods, centralized control strategies such as demand response, peak-valley electricity price and auxiliary services are mainly used for adjustment. However, with the access of distributed energy, energy storage systems, controllable loads and other diversified participants, the power grid system presents a complex dynamic game relationship. Each participant has its own interest demand and decision logic, while pursuing the maximization of its own interests, which may lead to the reduction of overall system efficiency. Zhang et al. pointed out through big data analysis that in the absence of effective coordination mechanisms, the self-interest behavior of distributed resources may reduce the efficiency of power grid peak regulation by 15-30%, and even cause new "secondary peak" problems [Zhang et al., Energy Economics, 2022, An integrated optimization and multi-scale input–output model for interaction mechanism analysis of energy–economic–environmental policy in a typical fossil-energy-dependent region].
[0004] Currently, the solutions to the grid peak shaving problem mainly include the following methods: one is the scheduling optimization method based on prediction, which formulates the scheduling plan in advance by predicting the load and renewable energy output; two is the demand response method based on market mechanism, which guides users to adjust their electricity consumption behavior through price signals; three is the adaptive control method based on artificial intelligence, which optimizes the control strategy using machine learning algorithms. Chen et al. evaluated the application effects of these three methods in actual projects in detail and found that the method based on artificial intelligence has obvious advantages in dealing with uncertainty and complexity, but still has limitations in model generalization ability and computational efficiency [Chen et al., Applied Energy, 2023, Hierarchical optimal scheduling method for regional integrated energy systems considering electricity-hydrogen shared energy].
[0005] However, these methods have obvious shortcomings when faced with super-large power grids: first, traditional prediction methods are difficult to accurately grasp the dynamic change characteristics of massive load nodes, especially in extreme weather and emergency situations. As Wang et al. study shows, the average prediction error of traditional prediction methods in extreme weather events can reach 15-25%, significantly affecting the peak shaving effect [Wang et al., Energy AI, 2024, Does artificial intelligence promote energy transition and curb carbon emissions?The role of trade openness]; second, simple market mechanisms are difficult to coordinate the conflict of interests of multiple parties, especially when the number of participants is large and there are complex game relationships. Johnson et al. study points out that in the electricity market with large-scale distributed resources, the game equilibrium point may deviate from the maximum social welfare by 20-35% [Johnson et al., Energy Policy, 2023, Global energy policy analysis to achieve near-term climate goals in the United States]; third, existing artificial intelligence methods often face dimension disaster, low training efficiency and other problems when dealing with high-dimensional, strongly coupled power grid systems, making it difficult to meet the demand for real-time regulation. Li et al. points out that the training time of conventional reinforcement learning algorithms when dealing with thousand-node-level power grid problems can be as long as several weeks, which cannot meet the demand of practical application [Li et al., IEEE Transactions on Neural Networks and Learning Systems, 2023, A Fast Feedforward Small-World Neural Network for Nonlinear System Modeling].
[0006] Furthermore, existing methods generally lack effective modeling of complex interaction relationships among multiple participants, failing to fully capture the causal relationships and potential conflicts between different nodes, resulting in difficulty in achieving globally optimal regulation effects in practical applications. Brown et al. through empirical research showed that without considering the causal relationship between nodes, even advanced distributed control algorithms can only achieve 65-80% of the theoretical optimal effect [Brown et al., Nature Energy, 2023, High-throughput Li plating quantification for fast-charging battery design]. Especially in the scenario of highly dynamic load changes, the adaptability and robustness of traditional methods are significantly reduced. Smith and Zhang's analysis of 10 large city power grid cases found that the failure rate of traditional peak shaving methods in the case of load mutation was as high as 30-45% [Smith and Zhang, Energy, 2024, A high-energy-density NASICON-type Na3V1.25Ga0.75(PO4)3cathode with reversible V4+ / V5+redox couple for sodium ion batteries].
[0007] In recent years, multi-agent reinforcement learning has gradually attracted attention in the field of power systems. Wilson et al. first applied multi-agent reinforcement learning to distribution network regulation and control, achieving preliminary success, but it is still limited to small-scale systems [Wilson et al., Applied AI Letters, 2023, Artificial Intelligence and Human-Induced Seismicity: Initial Observations of ChatGPT]; Rodriguez et al. proposed a hierarchical control structure based on deep Q network, but failed to effectively solve the problem of game between agents [Rodriguez et al., Energy Conversion and Management, 2023, Involving energy security and a water–energy-environment nexus framework in the optimal integration of rural water–energy supply systems]; Park et al. tried to combine causal inference with reinforcement learning for power grid regulation, but lacked verification for large-scale systems [Park et al., Electric Power Systems Research, 2024, Enhancing target benefits of power system stakeholders: A columns and constraint generation-based power-to-gas linked economic dispatch]. These studies provide an important theoretical and technical basis for the present invention, but none of them can systematically solve the problem of super large power grid source and load dynamic game. SUMMARY
[0008] In view of the problems existing in the prior art, the purpose of the present application is to provide a multi-agent reinforcement learning peak shaving method for ultra-large power grid source-load dynamic game; the present application introduces a multi-agent framework to model various participating subjects in the power grid, and through a hierarchical training mechanism and a reinforcement learning algorithm embedded with causal structure analysis, precise regulation and control of the complex power grid system are realized. Specifically, the present application designs a cooperation-competition hybrid training mechanism, so that each agent can take into account the overall system benefit while pursuing its own interests, effectively solving the global optimization problem faced by traditional methods. Under the background of smart grid transformation, the present application introduces advanced artificial intelligence technology, especially the multi-agent reinforcement learning method, and establishes a new paradigm for power grid peak shaving in complex dynamic environments, providing an innovative solution for the economic, efficient and reliable operation of the power grid. The present application can be applied to distributed energy management systems, demand response mechanisms, power market transaction strategy optimization, energy management systems and advanced power distribution and utilization systems, and has wide technical applicability and practical value.
[0009] The technical solutions of the present application are specifically introduced as follows.
[0010] The present application provides a multi-agent reinforcement learning peak shaving method for ultra-large power grid source-load dynamic game, comprising the following steps:
[0011] Step one, multi-agent modeling
[0012] The participating subjects in the power grid are abstracted as independent agents, and the specific state space, action space and reward function are defined for each agent, by regarding the power grid operating environment as a Markov decision process; according to the geographical location and electrical connection relationship, the agents are organized into a hierarchical structure including intra-regional groups, inter-regional layers and global dispatching layers;
[0013] Step two, hierarchical training mechanism implementation
[0014] A three-layer training mechanism of intra-regional cooperation, inter-regional competition and global coordination is adopted to perform multi-level progressive training on the multi-agent, the upper layer decision provides constraint conditions and optimization objectives for the lower layer, and the lower layer execution result provides feedback information for the upper layer, forming a closed-loop optimization mechanism;
[0015] Step three, causal structure analysis and entropy embedding enhancement
[0016] Based on the historical agent interaction data, a causal relationship diagram between agent behaviors is constructed to provide structured knowledge for strategy optimization to guide the exploration direction and strategy update of the agent; the state transition entropy and the strategy entropy are included in the agent learning goal to realize the attention to both the strategy diversity and the environmental influence, enhancing the exploration ability of the model;
[0017] Step four, multi-objective optimization decision deployment
[0018] For the multi-dimensional objectives of peak regulation effect, economic cost and system stability, the strategy set balancing the interests of all parties is generated by adopting the Pareto optimization principle; the trained strategy is gradually deployed in the actual environment, and closed-loop optimization is maintained through continuous real-time monitoring and feedback.
[0019] In the present application, in step one, the participating subjects include load aggregators, distributed power sources and energy storage systems; load regulation, distributed power output scheduling and energy storage charging and discharging control operations are unified into a dynamic game framework, and the three types of intelligent agents learn in parallel under the architecture of shared value network and independent strategy network to meet the differentiated needs of load reduction, distributed power regulation and energy storage charging and discharging collaborative control in the peak regulation process of super large power grid.
[0020] In the present application, in step two, the training is divided into three stages: intra-regional collaborative training stage, inter-regional competitive game stage and global coordination optimization stage.
[0021] Firstly, in the intra-regional collaborative training stage, multi-agent reinforcement learning based on Actor-Critic is used to make the intelligent agents in the same region learn local optimal strategies under the architecture of shared value network and independent strategy network, and each intelligent agent receives local and public observation information and realizes regional cooperation through a reward function.
[0022] Then, in the inter-regional competitive game stage, different regions are regarded as independent decision-making subjects, and the meta-game and iterative best response method are adopted to search for Nash equilibrium, taking into account resource constraints and safety constraints.
[0023] Finally, in the global coordination optimization stage, a high-level controller is introduced to coordinate resource quotas and priorities, and hierarchical reinforcement learning is used to ensure that each region executes its own strategy while obeying the overall optimal strategy.
[0024] In the present application, in step three, the collected state-action-reward sequences are time-aligned and normalized; a combination method of two-stage causal discovery is used to construct a preliminary causal graph; key causal paths are verified through intervention experiments, false associations are eliminated, causal structure information discovered is encoded as prior knowledge and integrated into the state representation of intelligent agents, and a causal relationship graph between intelligent agent behaviors is constructed.
[0025] In the present application, the combination method of two-stage causal discovery first uses a causal discovery algorithm based on conditional independence constraints to quickly contract the causal skeleton, and then uses a causal discovery algorithm based on fractional / continuous optimization to accurately direct and weight in the continuous relaxation of the adjacency matrix space, obtaining a preliminary causal graph satisfying the directed acyclic graph DAG constraint.
[0026] In the present application, in step three, the state transition entropy and the policy entropy are included in the learning goal of the agent, and the method for enhancing the exploration ability of the model comprises:
[0027] A policy entropy term is added to the policy objective function to encourage policy diversity.
[0028] The influence of the action of the agent on the uncertainty of the state transition of the environment is analyzed, and is quantified as state transition entropy.
[0029] According to the training phase and the change of the environment, the entropy weight of the policy entropy is dynamically adjusted, the area with high state transition entropy is taken as the priority exploration target, and the learning of the key state space is accelerated.
[0030] In the present application, in step four, the peak regulation effect, economic cost and system stability goal are decomposed into multi-dimensional quantifiable indexes including peak-valley difference reduction ratio, load curve smoothness, total regulation cost, subject income distribution, voltage deviation, frequency fluctuation and standby capacity, and according to the real-time state of the power grid or external decision preference, the weight of each sub-goal is periodically or event-triggered corrected, so that the system can flexibly focus on the most critical optimization direction.
[0031] In the present application, in step four, a multi-objective reinforcement learning method is used to construct and continuously update the Pareto frontier during the training process, to provide a candidate solution set for the compromise combination between different goals.
[0032] Compared with the prior art, the present application has the beneficial effects that:
[0033] Multi-agent modeling framework: By modeling the power grid participating subjects as independent agents, the dynamic game relationship in the system is accurately captured, and a more realistic abstract expression is provided for the complex power grid system.
[0034] Causal structure analysis technology: It breaks through the "black box" understanding limitation of traditional reinforcement learning on the internal mechanism of the system, identifies the causal relationship between agents, provides structured knowledge for strategy optimization, and enhances the interpretability of the model.
[0035] Entropy embedding learning method: The information theory concept is introduced into the reinforcement learning framework, the robustness and exploration efficiency of the model in the uncertain environment are significantly improved by maximizing information acquisition and policy diversity.
[0036] Multi-objective balance mechanism: By designing a comprehensive reward function and a Pareto optimal solution search algorithm, the balance optimization of multi-dimensional goals such as peak regulation effect, economic cost and system stability is realized, and the complex requirements of actual power grid operation are met. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 This is a diagram illustrating the overall architecture of the multi-agent reinforcement learning peak-shaving method described in this invention.
[0038] Figure 2 This is a flowchart of the hierarchical training mechanism of the present invention.
[0039] Figure 3 This is a schematic diagram of the causal structure analysis and entropy embedding reinforcement learning module of the present invention.
[0040] Figure 4 This is a flowchart of the multi-objective optimization decision-making process of the present invention. Detailed Implementation
[0041] The technical solution of the present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0042] This invention provides a peak-shaving method for ultra-large power grids based on multi-agent reinforcement learning, aiming to address the inefficiency of existing power grid control methods in high-dimensional, strongly coupled, and dynamic game-theoretic environments. By dividing the power grid control task into a multi-level collaborative decision-making process, the system can achieve globally optimal peak-shaving operations while ensuring the interests of all participating entities. Furthermore, the causal structure analysis and entropy embedding techniques introduced in this invention significantly improve the model's ability to understand high-order interaction patterns. This method not only improves the accuracy of power grid peak-shaving but also reduces system energy consumption and enhances the robustness of the power grid in extreme scenarios.
[0043] The core principle of this invention is based on the integration of a multi-agent modeling framework, a hierarchical training mechanism, causal structure analysis technology, entropy embedding learning method, and a multi-objective balance mechanism, aiming to solve the complex game problem in peak shaving of ultra-large power grids. The technical solution proposed in this invention includes the following key modules:
[0044] I. Multi-agent modeling module
[0045] The multi-agent modeling framework is the basis of this invention. By abstracting each participating entity in the power grid into an intelligent agent with autonomous decision-making capabilities, it accurately captures the game relationship in the system.
[0046] Intelligent agent abstraction and classification: Based on function and characteristics, the participating entities are divided into three types of intelligent agents: load aggregators, distributed power sources and energy storage systems. Each type of intelligent agent has a unique state observation, action selection and reward evaluation mechanism.
[0047] State-Action-Reward Design: The state space, action space, and reward function are carefully designed for each type of agent to ensure that the model can fully capture the key characteristics of power grid regulation while maintaining controllable computational complexity.
[0048] This module abstracts various participants in the power grid (such as load aggregators, distributed power sources, and energy storage systems) into independent agents, and defines a specific state space, action space, and reward function for each agent. By treating the power grid operating environment as a Markov Decision Process (MDP), this invention unifies operations such as load regulation, distributed power source output scheduling, and energy storage charging and discharging control into a dynamic game framework.
[0049] II. Hierarchical Training Mechanism
[0050] The hierarchical training mechanism is a key innovation in solving the dimensional disaster problem of ultra-large power grids. By decomposing global optimization into a multi-level progressive optimization process, it significantly improves training efficiency and strategy quality.
[0051] A three-tiered architecture of regional collaboration, inter-regional competition, and global coordination: This design aligns with the physical characteristics and management structure of the power grid, making the model training process more consistent with the actual system operation logic.
[0052] Inter-layer information flow and decision flow design: Upper-layer decisions provide constraints and optimization objectives for lower layers, and lower-layer execution results provide feedback information to upper layers, forming a closed-loop optimization mechanism.
[0053] To overcome the "curse of dimensionality" problem that easily occurs in multi-agent reinforcement learning in ultra-large urban power grids, this invention designs a three-layer training mechanism. Within a relatively small local power grid or zone, load aggregators and distributed power sources first learn basic collaborative strategies to achieve local peak shaving, energy storage utilization, and supply-demand balance. When multiple regions are interconnected and may influence each other in terms of network flow and transmission constraints, a competition and game mechanism is introduced, allowing agents to adapt to resource competition and collusion on a larger scale, thereby improving the diversity of training samples and the robustness of strategies. Building on the first two layers, iterative global comprehensive scheduling is performed on all regions, enabling the overall power grid to achieve a higher optimal balance in terms of renewable energy utilization, security constraints, and economic efficiency. This layered process gradually reduces the search and optimization space, significantly improves training efficiency, and provides a scalable zoned scheduling solution for urban power grids.
[0054] III. Causal Structure Analysis Module
[0055] Causal structure analysis technology breaks through the limitations of traditional reinforcement learning's "black box" understanding of the internal mechanisms of the system. By mining the causal relationships between agent behaviors, it provides structured knowledge for policy optimization.
[0056] Causal discovery and verification methods: Combining constraint-based causal discovery algorithms and fraction-based causal discovery algorithms to construct accurate causal graphs.
[0057] Integrating causal knowledge into strategy optimization: Encoding the discovered causal structures as prior knowledge guides the agent's exploration direction and strategy updates, significantly improving learning efficiency.
[0058] In multi-agent game theory, due to the high coupling between power grid operation characteristics and the environment, complex dependencies or potential conflicts often exist between the actions of various agents. This invention introduces a causal inference tool based on directed acyclic graphs (DAGs), utilizing causal discovery algorithms to identify key causal paths and influencing factors. By performing causal structure analysis on time-series data generated by agent interactions, a better understanding of the contribution and risks of different actions to network security, peak load, and energy storage lifetime can be achieved.
[0059] IV. Entropy Embedding Reinforcement Learning Module
[0060] This invention introduces the concept of entropy from information theory into the reinforcement learning framework. By adding an entropy term to the objective function, it addresses both policy diversity and environmental influence, effectively improving the model's robustness and adaptability in uncertain environments. Specifically, it mainly includes the following two aspects:
[0061] Dual guidance from policy entropy and state transition entropy: A policy entropy term is added to the traditional reinforcement learning objective function to encourage the agent to retain the exploration of multiple possible actions in the early stages of training, avoiding premature convergence to a suboptimal solution; a state transition entropy index is introduced to quantify the uncertainty brought about by the agent's actions to the evolution of the environment; when an action has a high impact on the environment, the system will correspondingly increase the exploration weight of that action or related states, prompting the agent to pay more attention to these "high-impact" decision points.
[0062] Adaptive entropy weight adjustment: The entropy coefficient is adjusted in real time according to different training stages, environmental fluctuation levels and agent learning progress to ensure a balance between sufficient exploration and stable convergence. In scenarios with frequent random events such as wind and solar power output fluctuations and sudden load increases, the agent can automatically increase the entropy weight to increase the exploration of uncertain states, and gradually reduce the entropy weight after the risk buffer, so as to ensure the feasibility and economy of the scheduling strategy.
[0063] This invention introduces the concept of entropy from information theory into the reinforcement learning framework, embedding it into the policy network and state transition process, aiming to enhance the model's exploratory ability and adaptability to uncertainty. By maximizing policy entropy, the system can maintain its willingness to explore multiple potential actions, avoiding getting trapped in local optima; at the same time, guided by state transition entropy, the agent can maintain high fault tolerance and resilience when encountering sudden load increases or large fluctuations in wind and solar power output.
[0064] V. Multi-objective optimization decision-making module
[0065] To balance multiple objectives such as peak-shaving effectiveness, economic cost, and grid operation safety, this invention proposes a multi-objective balancing mechanism. This mechanism achieves flexible trade-offs and dynamic scheduling among these objectives through a combination of comprehensive reward function design and Pareto optimal solution search. This mechanism mainly involves the following two key aspects:
[0066] Multi-objective decomposition and weight design: Complex objectives such as peak load reduction, electricity cost reduction, and network security assurance are decomposed into multi-dimensional quantifiable indicators, such as load peak-to-valley difference, dispatching costs, line utilization, and frequency stability. Based on the real-time state of the power grid or external decision preferences, the weights of each sub-objective are periodically or event-triggered, enabling the system to flexibly focus on the most critical optimization direction (e.g., increasing the weight of peak shaving objectives during load surges; prioritizing economic costs during frequent start-ups and shutdowns of conventional units).
[0067] Pareto optimal solution search: Utilizing multi-objective reinforcement learning (MORL), a Pareto front is constructed and continuously updated during training to provide a candidate solution set for compromise combinations among different objectives. After obtaining a series of non-dominated solutions balancing different objectives, the optimal (or suboptimal) solution can be selected from the Pareto front by combining it with actual operational constraints, scheduling policy priorities, and decision-maker preferences. This allows for meeting the needs of various operational scenarios and facilitating rapid switching of scheduling schemes in emergency situations or special periods.
[0068] This module unifies the modeling of multiple objectives, including the economic efficiency, stability, reliability, and environmental protection of power grid operation. It achieves a coordinated balance among these objectives by designing a comprehensive reward function or employing a Pareto optimal solution search algorithm. To ensure a trade-off between peak load suppression, generator start-up and shutdown frequency control, energy storage lifespan protection, and carbon emission targets, this module assigns appropriate weights to each objective or uses Pareto Front non-dominated solution retrieval.
[0069] Based on the above modules, the operation flow of the multi-agent reinforcement learning-based peak shaving method for ultra-large power grids of the present invention is as follows:
[0070] (1) System initialization
[0071] By collecting fundamental data such as power grid topology, load characteristics, and distributed resources, a digital twin model of the power grid is constructed, and the state and action spaces and reward functions of each agent are initialized. First, data cleaning and missing value handling are performed to ensure data quality. Then, the power grid topology is analyzed to identify key nodes and potential weak points. Next, loads are clustered and typical feature patterns are extracted. Simultaneously, mathematical models of output constraints and operational boundaries are established for distributed resources such as photovoltaic, wind power, and energy storage. Finally, the above information is integrated to form the digital twin model of the power grid, and initial state spaces, action spaces, and reward functions are designed for each agent based on preset objectives.
[0072] (2) Hierarchical training
[0073] In the regional collaborative training phase, multi-agent reinforcement learning based on Actor-Critic is used to enable agents within the same region to learn locally optimal policies within an architecture of shared value network and independent policy network. Each agent receives local and public observation information and achieves regional collaboration by designing an appropriate reward function. Next, in the inter-regional competitive game phase, different regions are treated as independent decision-making agents, and a meta-game and iterative best response method is adopted to search for Nash equilibrium, taking into account both resource and security constraints. Finally, in the global coordination and optimization phase, a high-level controller is introduced to coordinate resource allocation and priorities, and hierarchical reinforcement learning ensures that each region executes its own policy while conforming to the overall optimum.
[0074] (3) Causal Structure Analysis
[0075] Based on historical interaction data, a causal graph of agent behavior is constructed to uncover key influence paths and potential conflict points, providing structured knowledge for policy optimization. This process first performs time-series alignment, normalization, and noise reduction. Then, a two-stage causal discovery method is used to progressively optimize the causal framework and edge directions. Intervention experiments are then conducted to verify the effectiveness of causality and eliminate spurious associations. Finally, a causal impact index is used to identify key nodes or weak links within the system. The verified causal graph is then integrated into the agent's state representation and policy updates in a priori form, and is dynamically updated periodically according to the power grid's operational status.
[0076] (4) Entropy Embedding Reinforcement Learning
[0077] By incorporating state transition entropy and policy entropy into the reinforcement learning objective, the model is enhanced to have stronger exploration and adaptation capabilities in uncertain environments. First, a policy entropy term is added to the objective function to encourage diversity in different actions and prevent premature convergence. Simultaneously, a state transition entropy measurement framework is established to quantify the uncertain impact of agent actions on environmental evolution, and an adaptive entropy weight tuning mechanism is employed to dynamically adjust entropy weights based on factors such as the training stage and the degree of environmental fluctuation. Combined with an entropy-based focused exploration method, the system prioritizes learning high-entropy regions and integrates with traditional reinforcement learning algorithms to achieve rapid adaptation and efficient exploration capabilities for random events.
[0078] (5) Multi-objective optimization decision
[0079] To address multi-dimensional objectives such as peak-shaving effectiveness, economic cost, and system stability, the Pareto optimality principle is employed to generate a strategy set that balances the interests of all parties. A multi-dimensional evaluation index system is used to quantify technical indicators (such as peak-to-valley difference and load factor), economic indicators (cost and benefit), and safety indicators (voltage and frequency). These are then integrated into a dynamically weighted multi-objective reward function, and a Pareto front is searched using multi-objective reinforcement learning (MORL). Finally, based on decision-makers' preferences and grid operation constraints, the final solution is selected from the front, and a strategy set management mechanism is established to combine pre-generated strategy sets with online optimization.
[0080] (6) Online deployment and adaptive adjustment
[0081] The trained strategies are gradually deployed in real-world environments, and closed-loop optimization is maintained through continuous real-time monitoring and feedback. Initially, a smooth transition is achieved through human-machine collaboration, with the automation ratio gradually increased. Performance metrics are evaluated in real time, and a correction process is triggered if deviations exceed thresholds. Incremental learning and online adaptation methods are used to continuously update the model to cope with environmental dynamics. If a decision may cause serious risks, a fallback to a conservative strategy is implemented. Comprehensive offline retraining or policy library updates are performed regularly to ensure the long-term stability and safety of the overall performance.
[0082] To better explain the implementation method of this invention, the following will describe the implementation process in detail with reference to specific steps and key parameters in a peak-shaving scenario for a mega-city power grid. Through the integrated operation of multiple modules, the system can achieve power grid source-load coordinated peak shaving based on multi-agent reinforcement learning.
[0083] Example 1: Peak shaving of urban core area power grid based on three-layer architecture
[0084] Implementation Environment Setup
[0085] In this embodiment, the system is applied to a power grid peak-shaving scenario in the core area of a megacity, and depends on the following implementation environment:
[0086] The power grid includes 500 medium-voltage distribution network nodes and 3,000 low-voltage distribution network nodes, with a total peak load of 2.5 GW.
[0087] Distributed energy includes rooftop photovoltaics (total installed capacity of 300MW), commercial building energy storage systems (total capacity of 150MWh), and electric vehicle charging stations (50 stations with a total charging power of 100MW).
[0088] There are 10 load aggregators, each managing different types of controllable loads, including commercial building air conditioning systems, large data centers, and industrial interruptible loads.
[0089] Implementation Step 1: Multi-agent Modeling and Initialization
[0090] (1) Definition and abstraction of intelligent agents
[0091] First, the participants in the power grid are abstracted into three types of intelligent agents: load aggregator intelligent agents, distributed power generation intelligent agents, and energy storage system intelligent agents.
[0092] 1) Load Aggregator Agent (LAA)
[0093] • State space: The managed load's current power, adjustable capacity, user comfort indicators, time characteristics (hour / date / season), real-time electricity price signals, etc.
[0094] • Action space: the amount of load reduction or increase, the duration of adjustment, and the compensation price paid to users;
[0095] • Reward function: r_LAA = w1·Peak shaving revenue - w2·User dissatisfaction - w3·Operation cost, where w1, w2, and w3 are weighting coefficients.
[0096] 2) Distributed Generation Agent (DGA)
[0097] • State space: Current active / reactive power output, predicted available power for the next 15 minutes, local meteorological information (sunlight, wind speed), grid connection point voltage, real-time electricity price, and operation and maintenance status indicators;
[0098] • Action space: output adjustment ratio ΔP, power factor setpoint pf_set, start / stop control commands;
[0099] • Reward function: r_DGA = g1·Electricity sales revenue - g2·Wind and solar curtailment penalties - g3·Voltage over-limit penalties, where g1, g2, and g3 are weighting coefficients.
[0100] 3) Energy Storage Agent (ESA)
[0101] • State space: current state of charge (SOC), upper limit of charging / discharging power, predicted electricity price curve λ(t), battery health status (SOH), line load factor and bus voltage;
[0102] • Action space: Charge / discharge power command P_cmd (positive value for charging, negative value for discharging), working mode switching (constant power / constant voltage / standby, etc.);
[0103] • Reward function: r_ESA = h1·(Electricity sales revenue - Electricity purchase cost) - h2·Battery degradation cost - h3·Power default penalty, where h1, h2, and h3 are weighting coefficients.
[0104] The three types of intelligent agents learn in parallel under the architecture of a shared value network and their own independent policy networks to meet the differentiated needs of load reduction, distributed generation regulation and energy storage charging and discharging coordination during the peak shaving process of ultra-large power grids.
[0105] (2) Intelligent agent organizational structure design
[0106] Based on geographical location and electrical connections, the agents are organized into a hierarchical structure:
[0107] Intra-regional group: Organize electrically similar and strongly coupled intelligent agents into a group, such as loads, power sources and energy storage under the same distribution network.
[0108] Inter-regional layer: Different regional groups form a competitive relationship, vying for system resources and peak-shaving benefits.
[0109] Global scheduling layer: responsible for coordinating resource allocation and strategy adjustments among different regions to ensure global optimization.
[0110] The initialization parameters of the agent need to be pre-trained based on historical data to improve the quality of the initial policy. The state space design should consider dimensionality balance to avoid excessive dimensionality leading to low training efficiency.
[0111] Implementation Step 2: Implementation of Hierarchical Training Mechanism
[0112] (1) Regional collaborative training
[0113] First, agent-based collaborative training is conducted within the region using the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm, based on an Actor-Critic architecture. The training environment is a simulation built using real-world power grid data, including a power flow calculation module and a load response model. The initial learning rate is 0.001, with an adaptive learning rate adjustment strategy; the discount factor γ = 0.95; and the replay buffer size is 100,000 samples. Agents within each region make decisions based on shared observations but maintain independent policy and value networks. The average reward fluctuates by no more than 3% over 50 consecutive training rounds, or reaches the preset maximum number of training steps (200,000 steps).
[0114] (2) Inter-regional competitive game training
[0115] After training within the region, the agent policies within the region are kept fixed, and inter-regional competitive game training is conducted. A meta-game-based training method is adopted, in which each region is regarded as a player in the meta-game; an inter-regional competitive reward function is designed to encourage efficient use of resources while punishing excessive competition; the training objective is to find the Nash equilibrium point between regions, so that each region cannot obtain higher gains by unilaterally changing its strategy in the equilibrium state.
[0116] (3) Global Coordination Optimization
[0117] Following this, global coordination and optimization are performed to ensure maximum overall system performance. A global coordinator is introduced to adjust resource quotas and priorities in each region; a hierarchical reinforcement learning framework is adopted, with the global coordinator acting as a high-level agent to learn how to optimally allocate system resources; a global reward function is designed, focusing on the reduction of peak-to-valley differences, total system cost, and stability metrics; iterative training continues until the global performance metrics reach the preset target or converge. During inter-region training, policy oscillations must be prevented using a soft update mechanism; an appropriate exploration rate should be maintained during the global optimization phase to avoid getting trapped in local optima.
[0118] (4) Reward function design
[0119] The following describes the three-layer training phase and the typical reward function for entropy embedding. For three types of agents within the same region: ① Load Aggregator Agent (LAA), ② Distributed Generation Agent (DGA), and ③ Energy Storage Agent (ESA), the reward functions are defined as follows:
[0120] r_LAA=β1·ΔP_peak-β2·Discomfort-β3·CompCost
[0121] r_DGA=γ1·AGR-γ2·Curtail
[0122] r_ESA=η1·(E_sell-E_buy)-η2·Degradation
[0123] in:
[0124] ΔP_peak — The peak power reduced during this period;
[0125] Discomfort—a quantitative value for the loss of user comfort;
[0126] CompCost – (Compensation Cost) – The fee paid to the user;
[0127] AGR—Accuracy Gain Ratio, the degree of agreement between predicted and actual renewable energy output.
[0128] Curtail – Curtailed wind / solar power;
[0129] E_sell, E_buy — Energy storage electricity sold and purchased;
[0130] Degradation – A quantitative value indicating battery life degradation;
[0131] β, γ, and η are weighting coefficients, which can be obtained through offline parameter tuning or online adaptive tuning.
[0132] Treating each region as a higher-level "player," the reward function is:
[0133] R_region=λ1·Profit_net-λ2·TieLineOverload-λ3·VoltageViol
[0134] in:
[0135] Profit_net — Regional net profit;
[0136] TieLineOverload – Penalty for cross-regional line overload;
[0137] VoltageViol — Penalty for exceeding voltage limits;
[0138] λ1, λ2, and λ3 are the weights in the regional game.
[0139] The inter-regional equilibrium search uses the "Meta-Game with Iterated BestResponse" method until an approximate Nash equilibrium is reached.
[0140] Global coordinator reward is defined as
[0141] R_global=α1·(PeakValleyRatio) -1 +α2·(TotalCost) -1 -α3·InstabIndex
[0142] in:
[0143] PeakValleyRatio – The ratio of peak to valley differences;
[0144] TotalCost — Total peak-shaving cost of the system;
[0145] InstabIndex is a system instability index obtained by normalizing and weighting indicators such as voltage, frequency, and reserve margin.
[0146] α1, α2, and α3 are the global layer weights.
[0147] In addition to all the above rewards, policy entropy H(π_θ) and state transition entropy H_trans are added:
[0148]
[0149] π_θ — Policy Network;
[0150] H(π_θ) — Policy entropy, used to encourage action diversity;
[0151] H_trans — State transition entropy, which measures the impact of an action on environmental uncertainty;
[0152] κ1, κ2 — Entropy weights, which are dynamically adjusted according to the training phase and environmental fluctuations.
[0153] Implementation Step 3: Causal Structure Analysis and Entropy Embedding Enhancement
[0154] (1) Implementation of causal structure analysis
[0155] Based on the collected agent interaction data, the state-action-reward sequence is first time-aligned and normalized. Then, the PC algorithm (Peter-Clark, a causal discovery algorithm based on conditional independence constraints) is used to quickly shrink the causal skeleton. Next, a fractional / continuous optimization-based causal discovery algorithm is used to accurately orient and weight the causal path in a continuously relaxed adjacency matrix space, resulting in a preliminary causal graph that satisfies the constraints of a directed acyclic graph (DAG). After that, intervention experiments are used to verify key causal paths and eliminate false associations. Finally, the confirmed causal structure information is embedded as prior knowledge into the agent's state representation to guide policy updates and exploration directions.
[0156] (2) Implementation of Entropy Embedding Reinforcement Learning
[0157] Embedding the concept of information entropy into the reinforcement learning framework enhances the model's exploratory capabilities. A policy entropy term is added to the policy objective function to encourage policy diversity, with the formula: L(θ)=E[Σr_t]+α·H(π_θ), where Σr_t is the cumulative reward, H(π_θ) is the policy entropy, and α is the entropy weight. The uncertainty of the agent's actions on the environmental state transitions is analyzed and quantified as state transition entropy. The entropy weight α is dynamically adjusted based on the training phase and environmental changes, prioritizing regions with high state transition entropy to accelerate learning of key state spaces. The causal structure needs to be updated periodically to adapt to dynamic system changes; excessively large entropy weights can lead to overexploration and must be adjusted in a timely manner based on convergence.
[0158] Implementation Step 4: Multi-objective Optimization Decision Deployment
[0159] (1) Design of multi-objective reward function
[0160] Design a reward function that comprehensively considers multi-dimensional objectives:
[0161] Peak shaving performance indicators: reduction in peak-to-valley difference and smoothness of load curve;
[0162] Economic cost indicators: total regulatory costs and distribution of benefits among various stakeholders;
[0163] System stability indicators: voltage deviation, frequency fluctuation, and reserve capacity;
[0164] Comprehensive reward calculation: R = w_peak·R_peak + w_econ·R_econ + w_stab·R_stab, where w_peak, w_econ, and w_stab are the weight coefficients of peak shaving and valley filling performance, economic performance, and system stability performance, respectively; R_peak is the peak shaving and valley filling score; R_econ is the economic score; and R_stab is the system stability score.
[0165] (2) Pareto optimality search and decision
[0166] Based on the Pareto optimality principle, a set of peak-shaving strategies that balance the interests of all parties is generated: the Pareto front is constructed using the multi-objective reinforcement learning algorithm MORL; the final decision scheme is selected from the Pareto front using the reference point method; in emergency situations, the preference vector can be adjusted to prioritize the system's security.
[0167] (3) Online deployment and closed-loop control
[0168] Deploy the trained strategy to the actual power grid operation environment: design a smooth migration strategy from the model to the actual system, initially using a human-machine collaborative approach; establish a real-time monitoring and correction mechanism to promptly capture deviations between the model and the actual system; implement a progressive application strategy, starting from non-critical periods and gradually expanding to critical peak-shaving periods; regularly collect actual operation data for offline retraining and strategy updates.
[0169] Implementation effect
[0170] Through this embodiment, the system has achieved precise peak-shaving control of the power grid in the core area of a mega-city. The measured results show that the peak-to-valley difference has been reduced by 23.5%, the system peak-shaving cost has been reduced by 18.7%, and the reliability of power grid dispatch has been improved by 12.3%. It can better cope with load changes and extreme weather conditions, and is conducive to being further expanded to larger-scale power grid systems in the future.
[0171] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the inventive concept without creative effort. For example, the abstract granularity of the agent can be adjusted according to the specific power grid scale and characteristics; the reward function design can be optimized for different application scenarios; or other advanced reinforcement learning algorithms can be integrated to improve model performance. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A multi-agent reinforcement learning peak-shaving method for dynamic game theory of source and load in ultra-large power grids, characterized in that, Includes the following steps: Step 1: Multi-agent modeling The participants in the power grid are abstracted into independent intelligent agents, and a specific state space, action space and reward function are defined for each intelligent agent. The power grid operating environment is regarded as a Markov decision process. Based on geographical location and electrical connection relationship, the intelligent agents are organized into a hierarchical structure including intra-regional groups, inter-regional layers and global scheduling layer. Step Two: Implementation of the Hierarchical Training Mechanism A three-layer training mechanism of regional collaboration, inter-regional competition, and global coordination is adopted to carry out progressive training of multi-agents at multiple levels. The upper layer decision provides constraints and optimization objectives for the lower layer, and the execution results of the lower layer provide feedback information for the upper layer, forming a closed-loop optimization mechanism. Step 3: Causal Structure Analysis and Entropy Embedding Enhancement Based on historical agent interaction data, a causal relationship graph between agent behaviors is constructed to provide structured knowledge to guide the agent's exploration direction and policy updates for policy optimization; state transition entropy and policy entropy are incorporated into the agent's learning objectives to achieve attention to both policy diversity and environmental influence, thereby enhancing the model's exploration capabilities. Step 4: Multi-objective optimization decision-making and deployment To address the multi-dimensional objectives of peak shaving effect, economic cost, and system stability, the Pareto optimality principle is adopted to generate a strategy set that balances the interests of all parties. The trained strategies are gradually deployed in a real-world environment, and closed-loop optimization is maintained through continuous real-time monitoring and feedback.
2. The multi-agent reinforcement learning peak-shaving method according to claim 1, characterized in that, In step one, the participating entities include load aggregators, distributed generation sources, and energy storage systems. Load regulation, distributed generation output scheduling, and energy storage charging and discharging control operations are uniformly incorporated into a dynamic game framework. The three types of intelligent agents learn in parallel under the architecture of a shared value network and their respective independent policy networks to meet the differentiated needs of load reduction, distributed generation regulation, and coordinated control of energy storage charging and discharging during the peak shaving process of ultra-large power grids.
3. The multi-agent reinforcement learning peak-shaving method for dynamic game theory of source and load in ultra-large power grids according to claim 1, characterized in that, In step two, the training is divided into three stages: regional collaborative training, inter-regional competitive game, and global coordination optimization. First, in the regional collaborative training phase, multi-agent reinforcement learning based on Actor-Critic is used to enable agents in the same region to learn locally optimal policies under the architecture of shared value network and independent policy network. Each agent receives local and public observation information and achieves regional collaboration through reward function. Next, in the inter-regional competition game stage, different regions are treated as independent decision-making entities, and the meta-game and iterative best response methods are used to search for Nash equilibrium, taking into account both resource constraints and security constraints. Finally, in the global coordination and optimization phase, a high-level controller is introduced to coordinate resource allocation and priorities, and hierarchical reinforcement learning is used to ensure that each region executes its own strategy while conforming to the overall optimality.
4. The multi-agent reinforcement learning peak-shaving method for dynamic game theory of source and load in ultra-large power grids according to claim 1, characterized in that, In step three, the collected state-action-reward sequences are time-aligned and normalized; a two-stage causal discovery method is used to construct a preliminary causal graph; key causal paths are verified through intervention experiments, spurious associations are eliminated, and the discovered causal structure information is encoded as prior knowledge and integrated into the agent's state representation to realize the construction of a causal relationship graph between agent behaviors.
5. The multi-agent reinforcement learning peak-shaving method for dynamic game theory of source and load in ultra-large power grids according to claim 4, characterized in that, The two-stage causal discovery method first uses a causal discovery algorithm based on conditional independence constraints to quickly shrink the causal skeleton, and then uses a causal discovery algorithm based on fractional / continuous optimization to accurately orient and weight the causal graph in a continuously relaxed adjacency matrix space, thereby obtaining a preliminary causal graph that satisfies the DAG constraints.
6. The multi-agent reinforcement learning peak-shaving method for dynamic game theory of source and load in ultra-large power grids according to claim 1, characterized in that, In step three, methods for incorporating state transition entropy and policy entropy into the agent's learning objectives include: Add a policy entropy term to the policy objective function to encourage policy diversity; The uncertainty of the agent's actions on the environmental state transition is analyzed and quantified as state transition entropy; Based on the training phase and environmental changes, the entropy weight of the policy entropy is dynamically adjusted, and regions with high state transition entropy are prioritized for exploration to accelerate the learning of the key state space.
7. The multi-agent reinforcement learning peak-shaving method for dynamic game theory of source and load in ultra-large power grids according to claim 1, characterized in that, In step four, the peak-shaving effect, economic cost, and system stability target are decomposed into multi-dimensional quantifiable indicators, including the reduction ratio of peak-to-valley difference, load curve smoothness, total control cost, distribution of benefits among various entities, voltage deviation, frequency fluctuation, and reserve capacity. Based on the real-time status of the power grid or external decision preferences, the weights of each sub-target are periodically or event-triggered, enabling the system to flexibly focus on the most critical optimization direction at present.
8. The multi-agent reinforcement learning peak-shaving method for dynamic game theory of source and load in ultra-large power grids according to claim 1, characterized in that, In step four, a multi-objective reinforcement learning method is used to construct and continuously update the Pareto front during the training process, providing a candidate solution set for compromise combinations among different objectives. After obtaining a series of non-dominated solutions that balance different objectives, the optimal or suboptimal solution is selected from the Pareto front by combining them with actual operational constraints, scheduling policy priorities, and decision-maker preferences.
Citation Information
Cited By
Power distribution network optimization scheduling method and system based on multi-target deep reinforcement learning
CN121906489A