Power outage planning method and system based on multi-agent interpretable reinforcement learning

Through multi-agents, reinforcement learning methods can be explained, combined with the grid simulation environment and distributed trend computing, the problem of multi-objective optimization in the grid outage plan is solved, efficient and transparent power outage plan arrangement is achieved, and the safety and economicality of the power grid is improved.

CN120197915BActive Publication Date: 2025-08-22BEIJING KEDONG ELECTRIC POWER CONTROL SYST CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510669075.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-22
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to take into account the multi-objective optimization of ensuring safety, supply and consumption in the optimization of power grid power outage plans. The traditional heuristic algorithm has a high computing burden, is difficult to meet real-time decision-making requirements, and is low in intelligence.

Method used

The method based on multi-agent interpretable reinforcement learning is adopted. By establishing a grid simulation environment, multi-objective reward function is designed, and multi-intelligent collaboration against AC reinforcement learning algorithm is used to train the agent's action strategy, combining distributed trend computing and interpretable reinforcement learning algorithm to achieve multi-objective collaborative optimization.

Benefits of technology

Efficient and accurate power outage planning is achieved, the safety, reliability and economicality of power grid operation is improved, and the transparency and real-time decision-making is ensured through interpretable models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197915B_ABST
    Figure CN120197915B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for scheduling power outage plans based on multi-agent interpretable reinforcement learning. The method includes: establishing a power grid simulation environment as an interactive environment for training multiple agents; constructing the state space and action space of the agents based on the current scheduling options; designing the reward function of each corresponding agent based on the three optimization goals of ensuring safety, ensuring supply, and ensuring consumption; using a multi-intelligence collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​and combining it with the power grid simulation environment to train the action strategy of each agent; integrating the action strategies of each trained agent according to the set proportional weights, deciding the final collaborative action strategy, and thus generating the optimal power outage scheduling solution. The present invention achieves a multi-objective power outage scheduling problem through the collaboration of multiple agents, and can make the best strategy solution more efficiently and accurately, thereby improving the reliability of power grid operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power systems, and in particular relates to a method and system for scheduling power outage plans based on multi-agent interpretable reinforcement learning. Background Art

[0002] As the power grid continues to expand, the number of planned power outages is increasing. While the need for routine equipment maintenance increases, the main grid also faces various outage demands. The safety requirements of on-site operations will significantly increase the number of outages, resulting in high-dimensional and highly nonlinear variables in the optimization and planning of power outage plans. Furthermore, during the planning process, conflicts between the fundamental principles of grid operational safety, guaranteed power supply, and renewable energy integration are becoming increasingly prominent, making it difficult to simultaneously address the three control objectives of ensuring safety, power supply, and renewable energy integration.

[0003] Currently, solutions to power grid outage scheduling optimization models are primarily limited to algorithms based on swarm intelligence optimization frameworks, typically heuristic methods such as particle swarm optimization, dragonfly algorithms, and genetic algorithms. While these algorithms have the ability to avoid local extrema by simulating the behavioral characteristics of biological swarms, they are still essentially static optimization models, lacking real-time interaction and feedback mechanisms with dynamic environments. Furthermore, with the accelerated development of new power systems, grid topologies are becoming increasingly complex, leading to a superlinear expansion of the solution space for outage scheduling models. When the number of device nodes reaches the thousands, the computational burden of traditional heuristic algorithms increases exponentially, making it difficult for optimization efficiency to meet the time requirements of real-time power dispatch decision-making. Consequently, these algorithms exhibit significant limitations in terms of dynamic adaptability and online optimization efficiency. Furthermore, prior art discloses intelligent scheduling devices and methods for power grid outage scheduling based on deep reinforcement learning, which employ reinforcement learning to address multi-objective optimization. However, these approaches are primarily based on the scheduling experience of dispatch experts, resulting in low efficiency, limited content, and a low level of intelligence.

[0004] Therefore, there is an urgent need for an outage planning optimization strategy that can take into account the multi-objective outage planning problems of ensuring safety, supply and consumption, so as to adapt to today's increasingly complex new power systems. Summary of the Invention

[0005] In order to address the deficiencies in the prior art, the present invention provides a method and system for scheduling power outage plans based on multi-agent interpretable reinforcement learning. Multi-agents are used to implement collaborative optimization decisions for the three goals of ensuring safety, ensuring supply, and ensuring consumption that need to be considered in power grid outage plans. Multi-agents are constructed to perform reinforcement learning training to respond to different goals, which can more efficiently and accurately make the best strategic plans, minimize the scope and duration of equipment power outages, and improve the reliability and economy of power grid operation.

[0006] The present invention adopts the following technical solutions.

[0007] In a first aspect, the present invention provides a method for scheduling power outages based on multi-agent interpretable reinforcement learning, the method comprising:

[0008] According to the topology and operating parameters of the target power grid, a power grid simulation environment is established as an interactive environment for training multi-agents;

[0009] Construct the state space and action space of the agent based on the current orchestration plan options;

[0010] Design the reward function for each corresponding agent based on the three optimization goals of ensuring safety, ensuring supply, and ensuring consumption;

[0011] Based on the established environmental constraints, the agent's state space, action space, and various reward functions, a multi-agent collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​is used in conjunction with a power grid simulation environment to train the action strategies of each agent.

[0012] The trained action strategies of each intelligent agent are integrated according to the set proportional weights to decide the final collaborative action strategy, thereby generating the optimal power outage plan scheduling scheme.

[0013] In conjunction with the first aspect, optionally, the step of establishing a power grid simulation environment as an interactive environment for training multiple agents includes:

[0014] Based on the current agent collaborative action strategy, the corresponding parameters on the BPA file are modified, and distributed power flow calculation is performed based on the modified parameters to simulate the power grid operation, so as to support the use of historical power grid data for iterative training during the agent reinforcement learning process.

[0015] In combination with the first aspect, optionally, the distributed power flow calculation includes:

[0016] Based on topological structure and sensitivity analysis, the overall power grid model is tailored according to the supply area relationship;

[0017] Deploy the tailored power grid models of each supply area on each slave node of a distributed cluster containing multiple computing nodes;

[0018] Parallel calculation of the boundary flow of the power grid model in each slave node, and transmission of the boundary flow results of the power grid calculated by each slave node to the master node of the cluster;

[0019] The master node verifies the boundary flow results of each supply area and integrates the flow data of each supply area to finally form the flow calculation results of the overall power grid model.

[0020] In combination with the first aspect, optionally, constructing the state space of the intelligent agent includes: integrating the state information of the power grid equipment flow, the power grid load, the output of each unit and the initial requirements of the power outage plan in the state space; and updating the state space in real time through the data bus to ensure that multiple intelligent agents can make decisions based on the latest power outage plan data at each time step.

[0021] In combination with the first aspect, optionally, constructing the action space of the intelligent agent includes: defining a set of actions that the intelligent agent can take, including the specific time of power outage of the equipment, under the premise of complying with the physical constraints of the power grid system and the power outage operation requirements.

[0022] In combination with the first aspect, optionally, the environmental constraints include: line flow safety constraints, section flow safety constraints, power outage plan unchangeable constraints, simultaneous power outage constraints and / or power outage mutual exclusion constraints.

[0023] In conjunction with the first aspect, optionally, the expressions of the reward functions of the corresponding intelligent agents designed based on the three optimization objectives of ensuring safety, ensuring supply, and ensuring consumption are as follows:

[0024]

[0025] Where, is the reward function based on the safety objective, and Represent the coefficient weights of the current load term and the voltage load rate term, n and m Respectively represent the total number of AC lines and buses in the power grid under the current strategy, Indicates time line The current, Indicates line The current limit, Indicates time busbar The voltage, Indicates busbar Voltage limit is the set minimum value; is the reward function based on the supply guarantee objective, The quantitative values ​​of the impact of power outage frequency on users , Quantified value of the impact of power supply on users , Power outage for user maintenance and power supply section safety bonus value The weight coefficient of is the reward function based on the guaranteed consumption objective, and Respectively represent the reward value of new energy power outage units and power transmission section safety bonus value The weight coefficient of .

[0026] In conjunction with the first aspect, optionally, the step of adopting a multi-intelligence collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​and training the action strategy of each agent in combination with a power grid simulation environment includes:

[0027] Step 1: Randomly select an environment state from the current space state;

[0028] Step 2: Use the three Actor networks updated currently to represent the three corresponding agents. i 1. i 2 and i 3. Determine the action strategy to be taken under the current environmental conditions;

[0029] Step 3: Integrate the current action strategies of each agent according to the set proportional weights to obtain the current collaborative action strategy;

[0030] Step 4: Modify the corresponding parameters in the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraints are met. If not, return to step 1; if so, update the spatial state based on the current collaborative action strategy and execute step 5;

[0031] Step 5: Calculate the agents separately i 1. i 2 and i 3 The corresponding reward function values ​​under its current action policy 、 and , and use the gradient descent method to update the critic network corresponding to each agent, and use the currently updated critic networks to fit the Shapley Q value of each agent respectively;

[0032] Step 6: Based on the current Shapley Q value of each agent, use the gradient ascent method to update each Actor network, and return to step 1 for iterative training until the set number of iterations is met, and output the final collaborative action strategy.

[0033] In conjunction with the first aspect, optionally, the step of integrating the current action strategies of the agents according to the set proportional weights includes:

[0034] The decision actions of each agent in the action strategy for each device are normalized to the range of (-1, 1), and the execution actions of each device are integrated using the following formula:

[0035]

[0036] Where, represents the collaborative action strategy for device z, Represents the current agent i 1. i 2 and i The decision action for device z in the action strategy of 3; Represented as agents i 1. i 2 and i 3 sets the proportional weight, ;in, .

[0037] In conjunction with the first aspect, optionally, the expressions for fitting the Shapley Q value of each agent are as follows:

[0038]

[0039] Where, Representing an agent i In the environmental state S Execute action strategy a i Shapley's Q value; Representing an agent i In the environmental state S Execute action strategy a i The corresponding Q value is obtained by fitting the reward value corresponding to the agent based on the Bell optimal equation of reinforcement learning and using the Critic network; the subscript C represents the set of all agents, C / i Indicates exclusion of agents i The rest of the agents, Indicates that the remaining agents are in the environment state S Execute respective action strategies The sum of the corresponding Q values.

[0040] In combination with the first aspect, optionally, the method further includes:

[0041] Based on the decision tree framework in the interpretable reinforcement learning algorithm, the relationship between the constraints and power outage actions in the power outage plan is explained, and the state, action selection, constraints and final results of the intelligent agent when making decisions are recorded. The decision-making process is traced through the log module of the data reading and writing engine to facilitate the understanding of the behavioral logic of the intelligent agent. At the same time, charts and data flow diagrams are used to represent the real-time state and decision-making process of the intelligent agent, so as to display the state changes, action selection and constraints of the intelligent agent to the user.

[0042] In conjunction with the first aspect, optionally, the step of explaining the relationship between the constraints and the power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm includes:

[0043] Use the decision tree algorithm to classify the action data determined by the agent during iterative training to identify the device actions that cause the collaborative action strategy to fail to meet various constraints.

[0044] By building an interpretable model, we can analyze the importance of each device's decision-making action in affecting each constraint condition.

[0045] The interpretable model includes a recurrent neural network (RNN) encoder, a multi-layer perceptron (MLP) encoder, a self-explanatory model, and a linear regression unit connected in sequence.

[0046] In conjunction with the first aspect, optionally, the step of analyzing the importance of each device decision action affecting each constraint condition by constructing an interpretable model includes:

[0047] The environment state, action strategy, reward function parameters, and constraint condition parameters that the agent confronts during training are used as the input parameters of the model. The RNN encoder is used to encode the input parameters to capture the decision-making characteristics of the agent.

[0048] Use the MLP encoder to learn features that the RNN encoder fails to capture to obtain the overall adversarial features of the agent;

[0049] Establishing the correlation between the decision step features and the overall adversarial features through the Gaussian process in the self-explanatory model, and obtaining the correlation features between the decision step and the adversarial round;

[0050] The correlation features obtained each time during the training process are input into the linear regression unit for linear fitting to obtain the regression coefficient of each device decision action, which is used to represent the degree of influence of each device decision action on each constraint condition and the final reward, thus completing the explanation of the training environment, training process and decision action.

[0051] In a second aspect, the present invention provides a power outage scheduling system based on multi-agent interpretable reinforcement learning, which executes the steps of any method described in the first aspect of the present invention, and the system includes:

[0052] The environment establishment module is used to establish a power grid simulation environment as an interactive environment for training multi-agents based on the topology and operating parameters of the target power grid;

[0053] The space construction module is used to construct the state space and action space of the agent based on the current orchestration plan options;

[0054] The reward function design module is used to design the reward function for each corresponding agent based on the three optimization objectives of ensuring safety, ensuring supply, and ensuring consumption;

[0055] The agent training module is used to train the action strategies of each agent based on the specified environmental constraints, the agent's state space, action space, and various reward functions, using a multi-agent collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​and combined with a power grid simulation environment;

[0056] The collaborative action decision module is used to integrate the trained action strategies of each intelligent agent according to the set proportional weights, decide the final collaborative action strategy, and thus generate the optimal power outage plan scheduling scheme.

[0057] In conjunction with the second aspect, optionally, when the power grid simulation environment in the agent training module trains the action strategy of each agent:

[0058] The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current agent collaborative action strategy, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic results after the power grid state changes, thereby completing the simulation of power grid operation to support the use of historical power grid data for iterative training during the agent reinforcement learning process.

[0059] In conjunction with the second aspect, optionally, the distributed power flow calculation includes:

[0060] Based on topological structure and sensitivity analysis, the overall power grid model is tailored according to the supply area relationship;

[0061] Deploy the tailored power grid models of each supply area on each slave node of a distributed cluster containing multiple computing nodes;

[0062] Parallel calculation of the boundary flow of the power grid model in each slave node, and transmission of the boundary flow results of the power grid calculated by each slave node to the master node of the cluster;

[0063] The master node verifies the boundary flow results of each supply area and integrates the flow data of each supply area to finally form the flow calculation results of the overall power grid model.

[0064] In combination with the second aspect, optionally, in the space construction module, constructing the state space of the intelligent agent includes: integrating the state information of the power grid equipment flow, the power grid load, the output of each unit and the initial requirements of the power outage plan in the state space; and updating the state space in real time through the data bus to ensure that multiple intelligent agents can make decisions based on the latest power outage plan data at each time step.

[0065] In combination with the second aspect, optionally, in the space construction module, constructing the action space of the intelligent agent includes: defining the action space that the intelligent agent can take, including the specific time of power outage of the equipment, under the premise of complying with the physical constraints of the power grid system and the power outage operation requirements.

[0066] In combination with the second aspect, optionally, in the intelligent agent training module, the environmental constraints include: line flow safety constraints, section flow safety constraints, power outage plan unchangeable constraints, simultaneous power outage constraints and / or power outage mutual exclusion constraints.

[0067] In conjunction with the second aspect, optionally, in the reward function design module, the expressions of the reward functions for each corresponding agent are designed based on the three optimization objectives of ensuring safety, ensuring supply, and ensuring consumption, respectively, as follows:

[0068]

[0069] Where, is the reward function based on the safety objective, and Represent the coefficient weights of the current load term and the voltage load rate term, n and m Respectively represent the total number of AC lines and buses in the power grid under the current strategy, Indicates time line The current, Indicates line The current limit, Indicates time busbar The voltage, Indicates busbar Voltage limit is the set minimum value; is the reward function based on the supply guarantee objective, The quantitative values ​​of the impact of power outage frequency on users , Quantified value of the impact of power supply on users , Power outage for user maintenance and power supply section safety bonus value The weight coefficient of is the reward function based on the guaranteed consumption objective, and Respectively represent the reward value of new energy power outage units and power transmission section safety bonus value The weight coefficient of .

[0070] In conjunction with the second aspect, optionally, the agent training module includes:

[0071] A state selection unit is used to randomly select an environment state from the current space state;

[0072] The action decision unit is used to use the three Actor networks currently updated to determine the three corresponding agents. i 1. i 2 and i 3. Determine the action strategy to be taken under the current environmental conditions;

[0073] The action integration unit is used to integrate the current action strategies of each agent according to the set proportional weight to obtain the current collaborative action strategy;

[0074] A simulation judgment unit is used to modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraints are met. If not, the state selection unit is instructed to reselect an environmental state; if so, the current collaborative action strategy update space state is fed back to the reward calculation unit;

[0075] Reward calculation unit, used to calculate the agent i 1. i 2 and i 3 The corresponding reward function values ​​under its current action policy 、 and , and use the gradient descent method to update the critic network corresponding to each agent, and use the currently updated critic networks to fit the Shapley Q value of each agent respectively;

[0076] The update iteration unit is used to update each Actor network using the gradient ascent method based on the current Shapley Q value of each intelligent agent, and feed it back to the state selection unit for iterative training until the set number of iterations is met and the final collaborative action strategy is output.

[0077] In conjunction with the second aspect, optionally, the action integration unit includes:

[0078] The normalization subunit is used to normalize the decision actions of each device in each agent's action strategy to the range of (-1, 1);

[0079] The integration subunit is used to integrate the execution actions of each device through the following formula:

[0080]

[0081] Where, represents the collaborative action strategy for device z, Represents the current agent i 1.i 2 and i The decision action for device z in the action strategy of 3; Represented as agents i 1. i 2 and i 3 sets the proportional weight, ;in, .

[0082] In conjunction with the second aspect, optionally, in the reward calculation unit, the expressions for fitting the Shapley Q value of each agent are as follows:

[0083]

[0084] Where, Representing an agent i In the environmental state S Execute action strategy a i Shapley's Q value; Representing an agent i In the environmental state S Execute action strategy a i The corresponding Q value is obtained by fitting the reward value corresponding to the agent based on the Bell optimal equation of reinforcement learning and using the Critic network; the subscript C represents the set of all agents, C / i Indicates exclusion of agents i The rest of the agents, Indicates that the remaining agents are in the environment state S Execute respective action strategies The sum of the corresponding Q values.

[0085] In conjunction with the second aspect, optionally, the system further includes:

[0086] The explanation module is used to explain the relationship between constraints and outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, record the state, action selection, constraints and final results of the intelligent agent when making decisions, and trace the decision-making process through the log module of the data reading and writing engine to facilitate understanding of the intelligent agent's behavioral logic; at the same time, charts and data flow diagrams are used to represent the real-time state and decision-making process of the intelligent agent, so as to display the state changes, action selection and constraints of the intelligent agent to the user.

[0087] In conjunction with the second aspect, optionally, the interpretation module includes:

[0088] The decision classification unit is used to classify the action data determined by the agent during iterative training using a decision tree algorithm, so as to identify the device decision actions that cause the collaborative action strategy to fail to meet various constraints.

[0089] A construction unit is used to analyze the importance of each device's decision-making action in affecting each constraint condition by building an interpretable model;

[0090] The interpretable model includes a recurrent neural network (RNN) encoder, a multi-layer perceptron (MLP) encoder, a self-explanatory model, and a linear regression unit connected in sequence.

[0091] In conjunction with the second aspect, optionally, the construction unit includes:

[0092] The encoding subunit is used to take the environment state, action strategy, reward function parameters and constraint condition parameters that the agent confronts during training as the input parameters of the model, and encode the input parameters using the RNN encoder to capture the decision-step characteristics of the agent;

[0093] Learning subunits, using MLP encoders to learn features that RNN encoders fail to capture, to obtain the overall adversarial features of the agent;

[0094] A self-explanatory subunit, configured to establish a correlation between the decision-step feature and the overall adversarial feature through a Gaussian process in a self-explanatory model, thereby obtaining a correlation feature between the decision-step and the adversarial round;

[0095] The linear fitting unit is used to input the correlation features obtained each time during the training process into the linear regression unit for linear fitting, and obtain the regression coefficient of each device decision action, which is used to express the degree of influence of each device decision action on each constraint condition and the final reward, that is, to complete the explanation of the training environment, training process and decision action.

[0096] In a third aspect, the present invention provides a terminal including a processor and a storage medium;

[0097] The storage medium is used to store instructions;

[0098] The processor is configured to operate according to the instructions to execute the steps of any one of the methods described in the first aspect of the present invention.

[0099] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect of the present invention.

[0100] The beneficial effect of the present invention is that, compared with the prior art,

[0101] (1) The present invention establishes a power grid simulation environment, formulates environmental constraints related to various power outage plans, introduces a multi-intelligence collaborative confrontation AC reinforcement learning algorithm based on Shapley value, and calculates the reward value of the agent's action according to the designed multi-objective reward function to feed back to the agent. Through continuous training iteration, the optimal decision of the power outage plan is finally obtained, which ensures the real-time and collaborative nature of the agent's decision. Through the mutual cooperation and confrontation between each agent, the multi-objective power outage plan scheduling problem that takes into account safety, supply and consumption is achieved. By constructing multi-agent reinforcement learning training to deal with different goals, the decision-making results are more efficient and accurate.

[0102] (2) The present invention takes safety as the basic premise. When deciding the collaborative action strategy of each intelligent agent, the decision-making action of the intelligent agent that ensures safety is given the highest weight. Through the scheme of the main target intelligent agent providing a safety guarantee, the final power outage plan scheduling decision can focus on the realization of safety goals and improve the safety and reliability of power grid operation.

[0103] (3) The present invention establishes a power grid simulation environment through distributed accelerated power flow calculation to support the rapid and effective training of reinforcement learning agents, so that the agents can make the best power outage plan optimization and scheduling scheme, minimize the scope and duration of equipment power outages, and improve the reliability and economy of power grid operation.

[0104] (4) The present invention explains the relationship between the constraints and the power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, establishes interpretable decision rules for the power outage plan using the decision tree algorithm, and constructs an interpretable model. From the three aspects of the interpretability of the training environment, the interpretability of the decision actions, and the interpretability of the training process, the process of multi-agent action decision-making in the comprehensive optimization of the power outage plan can be clearly explained, ensuring the transparency of the agent's decision-making, which is conducive to the user's understanding of the agent's behavior. At the same time, the interpretable model can help users analyze the decision-making basis of the agent in a specific environmental state, thereby facilitating the discovery of unreasonable decisions and timely adjustment of strategies. The optimization and scheduling of the power grid outage plan is more stable and has higher interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 1 is a flow chart of a method for scheduling power outages based on multi-agent interpretable reinforcement learning in the present invention;

[0106] Figure 2 is a schematic diagram of an interpretable model in the present invention;

[0107] Figure 3 It is a schematic diagram of the process flow of distributed power flow calculation in the present invention;

[0108] Figure 4This is an experimental data graph showing the relationship between the comprehensive reward value of the agent and the number of iterations during the agent training process of the present invention;

[0109] Figure 5 This is a block diagram of the structural principles of the power outage planning system based on multi-agent interpretable reinforcement learning in the present invention. DETAILED DESCRIPTION

[0110] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. The embodiments described in the present invention are only part of the embodiments of the present invention, not all of the embodiments. Based on the spirit of the present invention, other embodiments obtained by ordinary technicians in this field without making any creative efforts are all within the scope of protection of the present invention.

[0111] Example 1:

[0112] Reference Figure 1 The embodiment of the present invention provides a method for scheduling power outage plans based on multi-agent interpretable reinforcement learning, which specifically includes the following steps:

[0113] S1. Establish a power grid simulation environment as an interactive environment for training multi-agents based on the topology and operating parameters of the target power grid;

[0114] It should be noted that during the power outage planning decision-making training process, the power grid simulation environment needs to realize the power flow calculation and solution function based on the power grid operation data. Depending on different needs, centralized or distributed power flow calculation methods can be selected to support the rapid and effective training of reinforcement learning agents.

[0115] A preferred but non-limiting embodiment, referring to Figure 3 As shown, the power grid simulation environment of the present invention adopts a distributed power flow calculation method. During training interaction, the simulation environment modifies the corresponding parameters in the BPA file based on the current agent collaborative action strategy, and performs distributed power flow calculation based on the modified parameters, thereby simulating the power grid operation to support the use of historical power grid data for iterative training during the agent reinforcement learning process. Among them, the distributed power flow calculation specifically includes:

[0116] S1.2: Based on the grid operation BPA file, read the grid operation data and establish the grid model;

[0117] S1.3: Based on topology and sensitivity analysis, the overall grid model is tailored according to the supply area relationship to form multiple small grid models;

[0118] S1.4: Deploy the tailored power grid models of each supply area on each slave node of a distributed cluster containing multiple computing nodes;

[0119] S1.5: Parallel calculation of the boundary flow of the supply area power grid model in each slave node, and transmission of the boundary flow results of the supply area calculated by each slave node to the master node of the cluster;

[0120] S1.6: The master node verifies the boundary flow results of each supply area, reduces the transmission consumption of cluster data volume, integrates the flow data of each supply area, and finally forms the flow calculation results of the overall power grid model.

[0121] This embodiment greatly accelerates the data processing speed of the simulation by performing parallel power flow calculation on the trimmed partition model.

[0122] S2, construct the state space and action space of the agent based on the current choreography plan options;

[0123] Specifically, the state space is composed of the current environmental state of the intelligent agent and its own state. Constructing the state space of the intelligent agent includes: integrating the state information of the power grid equipment flow and the environmental parameters (such as the power grid load, the output of each unit, the power outage equipment and time obtained according to the initial requirements of the power outage plan, weather data, and information on major events affecting the power grid) into the state space; for the initial requirements of the power outage plan, since they are unstructured data, the UIE framework is also required to identify and extract the power outage equipment entities; at the same time, the state space is updated in real time through the data bus to ensure that multiple intelligent agents can make decisions based on the latest power outage plan data at each time step.

[0124] Constructing the agent's action space involves defining the set of actions the agent can take, including the specific time of the device outage, while adhering to the physical constraints of the power grid system and the power outage operation requirements. These actions can range from single-step actions (such as turning a device on or off) to compound actions (such as collaboratively assigning tasks to multiple devices). A collaborative multi-agent action space is then constructed, allowing an agent to consider the states and potential actions of other agents when making online decisions. For example, this can be expressed as follows: The parameters in brackets represent the power outage time of AC line, busbar, transformer, switch and unit respectively.

[0125] S3. Design the reward function for each corresponding agent based on the three optimization goals of ensuring safety, ensuring supply, and ensuring consumption;

[0126] The specific process of designing the reward functions required by the reinforcement learning agent to ensure safety, supply, and consumption for the optimized outage schedule is as follows:

[0127] (1) Safety objectives:

[0128] Considering the line current load factor and bus voltage load factor, the higher the load factor, the greater the safety risk, and thus the worse the safety. Therefore, the safety reward value is obtained by multiplying the current load factor and voltage load factor by a negative weight coefficient. Furthermore, since the system is already under the constraints of power flow and voltage, and to achieve data normalization, the load factor is limited to a maximum of 1. The reward function is as follows:

[0129]

[0130] Where, is the reward function based on the safety objective, and Represent the coefficient weights of the current load term and the voltage load rate term, n and m Respectively represent the total number of AC lines and buses in the power grid under the current strategy, Indicates time line The current, Indicates line The current limit, Indicates time busbar The voltage, Indicates busbar Voltage limit It is the minimum value set to avoid the denominator being zero.

[0131] (2) Supply guarantee target:

[0132] a: Comprehensively consider the impact of power outages on users. It is necessary to calculate based on the frequency of power outages in the past to minimize the impact of power outages on users. The quantitative value of the impact of power outage frequency on users is calculated using the following formula:

[0133]

[0134] Where, For equipment Frequency of power outages (calculated year by year); For equipment Number of users affected by the power outage; is the total number of users; is the set of power-off devices; this value represents the impact of the power outage plan on users. The greater the impact on users, the worse the power supply guarantee effect. The inverse of this value can be used as the reward value for the power supply guarantee target.

[0135] b: Calculate the quantitative value of the impact of power supply on users by the following formula To quantify the reliability of power supply:

[0136]

[0137] Where, is the total number of power grid nodes in the area, i is the index, The duration of power outage for users in the area is used to quantify the power supply situation using the average node power outage duration. The inverse of this value represents the power supply reliability caused by the power outage plan, which can be used as the reward value for the supply guarantee target.

[0138] c. The implementation of maintenance outage plans often results in power outages for users, which reduces power supply reliability. To reduce the number of power outages caused by equipment maintenance, reduce the amount of power outages, and improve power supply efficiency, the amount of power outages for users should be minimized:

[0139]

[0140] Where, is the user number, is the total number of users, 、 For the Total power outage power and total power outage electricity after the power outage plan for each user is implemented; is the power outage time; the supply target is described by the power outage amount of the user for maintenance. The lower the power outage amount of the user for maintenance, the higher the power supply reliability. The inverse of this value is used as the reward function for ensuring supply. The higher the reward value, the better the supply guarantee effect.

[0141] d: Considering the impact of power supply section, power supply section safety reward function The more restricted the power supply section is, the lower the reward value is:

[0142]

[0143] Where, is the total number of current nodes, l is the index, Power supply section, A constant with a minimum value to avoid the denominator being zero.

[0144] Finally, by comprehensively considering the impact of power outages on users, quantifying the reliability of power supply, the amount of power outages required for user maintenance, and the power supply sections, we obtain the final reward function for the power supply guarantee goal:

[0145]

[0146] Where, is the reward function based on the supply guarantee objective, The quantitative values ​​of the impact of power outage frequency on users , Quantified value of the impact of power supply on users , Power outage for user maintenance and the weight coefficient of the power supply section safety bonus value.

[0147] (3) Guaranteeing consumption targets:

[0148] Renewable energy consumption means minimizing the waste of renewable energy and fully utilizing green energy sources such as wind and solar power. This can be achieved by encouraging the timely dispatch of renewable energy production capacity to areas or times with higher load demand.

[0149] Considering reducing the outage of renewable energy units as one of the goals of ensuring power consumption, we first add a reward function for renewable energy units with power outages. The greater the power generation capacity of the renewable energy units that are out of power, the lower the reward value:

[0150]

[0151] Where, It represents the output of renewable energy units lost due to planned power outages; i Indicates the i The new energy units that were shut down due to power outage plan, Indicates the total number of renewable energy units shut down due to power outage plans.

[0152] Secondly, a new power transmission section safety reward function is added The higher the degree of power transmission section restriction, the lower the reward value:

[0153]

[0154] Where, represents the total number of current nodes, Power transmission section, A constant with a minimum value to avoid the denominator being zero.

[0155] Finally, taking into account the total load, renewable energy outage units, and transmission sections, the final reward function for ensuring power consumption is obtained:

[0156]

[0157] Where, is the reward function based on the guaranteed consumption objective, and Respectively represent the reward value of new energy power outage units and power transmission section safety bonus value The weight coefficient of .

[0158] S4. Based on the established environmental constraints, the agent's state space, action space, and reward functions, a multi-agent collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​is used in conjunction with a power grid simulation environment to train the action strategies of each agent.

[0159] The environmental constraints formulated in this embodiment include: line flow safety constraints, section flow safety constraints, power outage plan non-change constraints, simultaneous power outage constraints, and power outage mutual exclusion constraints; the details are as follows:

[0160] 1) Line power flow safety constraints:

[0161]

[0162] Where, For the line l Safety and stability constraints; For the crew i Node to line l The generator output power transfer distribution factor; For the crew i exist t Always make contributions; K is the number of nodes in the system; For nodes k Line l The generator output power transfer distribution factor; For nodes k exist t Bus load value for the time period; and Line l The forward and reverse power flow slack variables.

[0163] 2) Sectional tidal flow safety constraints:

[0164]

[0165] Where, and Section s tidal constraints; For the crew i Node section s The generator output power transfer distribution factor; For nodes k Cross-section s The generator output power transfer distribution factor; and Section s The forward and reverse power flow slack variables.

[0166] 3) Power outage plan cannot be changed:

[0167]

[0168] Where, Indicates the power outage object is in t There was a power outage.

[0169] 4) Simultaneous power outage constraints:

[0170]

[0171] Where, and They represent the scheduled execution start time and end time of the power outage plan respectively.

[0172] 5) The power outage mutual exclusion constraint can be expressed as:

[0173]

[0174] Where, T Indicates the length of the entire power outage scheduling cycle.

[0175] Specifically, in this embodiment S4, the steps of using the multi-intelligence collaborative adversarial AC reinforcement learning algorithm based on Shapley value and combining it with the power grid simulation environment to train the action strategy of each agent include:

[0176] Step 1: Randomly select an environment state from the current space state;

[0177] Step 2: Use the three Actor networks updated currently to represent the three corresponding agents. i 1. i 2 and i 3. Determine the action strategy to be taken under the current environmental conditions;

[0178] Step 3: Integrate the current action strategies of each agent according to the set proportional weights to obtain the current collaborative action strategy;

[0179] In a preferred but non-limiting embodiment, the step of integrating the current action strategies of each agent according to the set proportional weight in step 2 includes normalizing the decision actions for each device in the action strategy of each agent to the range of (-1, 1), and integrating the execution actions of each device respectively using the following formula:

[0180]

[0181] Where, represents the collaborative action strategy for device z, Represents the current agent i 1. i2 and i The decision action for device z in the action strategy of 3; Represented as agents i 1. i 2 and i 3 sets the proportional weight, ;in, .

[0182] This embodiment takes safety as the primary goal, so the action weight of the intelligent agent that achieves the safety goal is greater than the action weight of the supply and consumption goals. is 0.5, and Both are 0.25.

[0183] Step 4: Modify the corresponding parameters in the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraints are met. If not, return to step 1; if so, update the spatial state based on the current collaborative action strategy and execute step 5;

[0184] It is further explained that when the power grid operation simulation is performed in step 4 of this embodiment to determine whether the environmental constraints are met, the collaborative action is first substituted into the existing unchangeable plan for verification. If the power outage time of the unchangeable plan is not met, it means that this constraint is not met, and the process returns to step 1; if it is met, the corresponding parameters on the BPA file in the power grid simulation environment are modified according to the current collaborative action strategy, and a flow calculation is performed to determine whether the result of the flow calculation meets the flow constraints. Based on the topological results obtained by the flow calculation and the correlation relationship of the power outage plan, it is determined whether the scheme meets the same-outage constraint and the mutual exclusion constraint. If not, the process returns to step 1.

[0185] Step 5: Calculate the agents separately i 1. i 2 and i 3 The corresponding reward function values ​​under its current action policy 、 and , and use the gradient descent method to update the critic network corresponding to each agent, and use the currently updated critic networks to fit the Shapley Q value of each agent respectively;

[0186] In this embodiment, the expressions for fitting the Shapley Q value of each agent are as follows:

[0187]

[0188] Where, Representing an agent iIn the environmental state S Execute action strategy a i Shapley's Q value; Representing an agent i In the environmental state S Execute action strategy a i The corresponding Q value is obtained by fitting the reward value corresponding to the agent based on the Bell optimal equation of reinforcement learning and using the Critic network; the subscript C represents the set of all agents, C / i Indicates exclusion of agents i The rest of the agents, Indicates that the remaining agents are in the environment state S Execute respective action strategies The sum of the corresponding Q values ​​under The principle and The same, no further details here.

[0189] Step 6: Based on the current Shapley Q value of each agent, use the gradient ascent method to update each Actor network, and return to step 1 for iterative training until the set number of iterations is met, and output the final collaborative action strategy.

[0190] In a preferred but non-limiting embodiment, the method for scheduling a power outage based on multi-agent interpretable reinforcement learning provided in this embodiment further includes:

[0191] Based on the decision tree framework in the interpretable reinforcement learning algorithm, the relationship between the constraints and power outage actions in the power outage plan is explained, and the state, action selection, constraints and final results of the intelligent agent when making decisions are recorded. The decision-making process is traced through the log module of the data reading and writing engine to facilitate the understanding of the behavioral logic of the intelligent agent. At the same time, charts and data flow diagrams are used to represent the real-time state and decision-making process of the intelligent agent, so as to display the state changes, action selection and constraints of the intelligent agent to the user.

[0192] The steps of explaining the relationship between constraints and outage actions in the outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm include:

[0193] (1) Use the decision tree algorithm to classify the action data determined by the agent during iterative training to identify the device decision actions that make the collaborative action strategy fail to meet various constraints;

[0194] Specifically, according to the decision tree algorithm, each node can be recursively split. By calculating the conditional Gini index of each power outage action during the training process, the attribute with the smallest conditional Gini index is selected as the split node. The calculation formula of the Gini index is:

[0195]

[0196] Where, D It is a power outage action dataset; Action for power outage D The Gini index; k is the number of constraint categories; The action sample for the selected device decision belongs to the category k The decision tree intuitively shows the decision path from the constraints to the power outage action through its tree structure. Each internal node represents a decision point, which is divided based on the threshold of a constraint feature. Each leaf node of the decision tree represents a specific action or action category. The path from the root node to the leaf node forms a series of interpretable decision rules. These rules describe how the agent chooses the power outage action under different constraints.

[0197] (2) Analyze the importance of each device's decision-making action in affecting each constraint by building an interpretable model;

[0198] Reference Figure 2 As shown in the figure, the interpretable model includes a recurrent neural network RNN ​​encoder, a multi-layer perceptron network MLP encoder, a self-explanatory model, and a linear regression unit connected in sequence.

[0199] In this embodiment, the steps of analyzing the importance of each device decision action in affecting each constraint condition by constructing an interpretable model include:

[0200] A: The environment state, action strategy, reward function parameters, and constraint parameters that the agent faces during training are used as the model's input parameters. The RNN encoder is used to encode the input parameters to capture the agent's decision-step characteristics.

[0201] B: Use the MLP encoder to learn features that the RNN encoder fails to capture to obtain the overall adversarial features of the agent;

[0202] C: Establishing the correlation between the decision-step features and the overall adversarial features through the Gaussian process in the self-explanatory model to obtain the correlation features between the decision-step and the adversarial round;

[0203] D: Input each correlation feature obtained during the training process into the linear regression unit for linear fitting to obtain the regression coefficient of each device's decision action. This is used to represent the degree of influence of each device's decision action on each constraint and the final reward, thus completing the explanation of the training environment, training process, and decision action.

[0204] In summary, the multi-agent interpretable reinforcement learning-based outage scheduling method provided in this embodiment uses a reward function designed to optimize safety, supply, and consumption, while simultaneously satisfying line flow safety constraints, section flow safety constraints, outage schedule immutability constraints, simultaneous outage constraints, and outage mutual exclusion constraints. This model constructs an interpretable model for multi-agent action decisions, which accurately and clearly explains the results of multi-agent decisions. While satisfying these constraints, the agents aim to minimize flow overshoots, outage frequency, and wind / solar curtailment, ultimately resulting in a final outage strategy. To address the difficulty in interpreting the outputs of multi-agent strategies, a policy-level interpretation method for deep reinforcement learning agents is employed. The state space S and action space A during the agent confrontation process are used as input. The outputs of the recurrent neural network (RNN) encoder and the multi-layer perceptron (MLP) encoder are superimposed on a Gaussian process through an interpretation model. A linear regression unit is used to fit the output features of the self-interpretable model. The importance and correlation of decision steps are obtained through the linear regression process. This approach is used to analyze the decision-making process of multi-agent actions in outage scheduling and explain the policy outputs.

[0205] The power outage scheduling method based on multi-agent interpretable reinforcement learning provided in this embodiment is applied in detail in the training instance as follows:

[0206] 1. Based on the BPA file of power grid operation data and the original demand file of power outage plan, the state space and action space of the power outage plan scheduling reinforcement learning task are constructed;

[0207] 2. The power grid simulation environment service program reads the power grid operation data BPA file and the power outage plan initial requirement file, builds a power grid model, extracts the current power grid state and sends it to the multi-agent established using the three target reward functions designed by the present invention;

[0208] 3. Based on the read grid status, each agent obtains its own decision action, normalizes the action output by each agent and multiplies it by the weight of the corresponding target, and then adds them together to obtain the comprehensive decision action of the multi-agent. The action is then transmitted to the grid simulation environment service program;

[0209] 4. Using the line flow safety constraints, section flow safety constraints, power outage plan immutability constraints, simultaneous power outage constraints, and power outage mutual exclusion constraints formulated by the present invention, in a power grid simulation environment, the agent's decision-making behavior is calculated based on the constraint formulas to determine whether it meets the relevant constraints, and actions that do not meet the constraints are eliminated;

[0210] 5. Based on the grid power flow safety constraints, the grid state (including unit output status, load status, and line power flow values) and the original requirements of the power outage plan are used as the state space, and the power outage time of the power outage equipment is used as the action space to establish an interpretable decision tree. The decision action of the intelligent agent, as well as the constraint condition parameters and the target reward function parameters, are transmitted to the trained interpretable model to obtain the regression coefficient corresponding to these input values, reflecting the importance of the decision action.

[0211] 6. The power grid simulation environment service program modifies the corresponding parameters in the BPA file based on the comprehensive decision-making actions and performs power flow calculations to conduct power grid simulation deduction. When performing power flow calculations, the power grid simulation environment first tailors the overall power grid model according to the partition relationship and deploys each tailored model in a cluster of multiple computing nodes. Finally, the slave nodes only need to transmit the various small-scale supply zone boundary power flow results to the master node. The master node verifies and integrates the boundary power flow results of the supply zone to achieve distributed power flow calculation.

[0212] 7. The power grid simulation environment then calculates the reward function values ​​for ensuring safety, ensuring supply, and ensuring consumption based on the deduction results. After normalization, the values ​​are calculated using the following formula: The calculated comprehensive reward value is then fed back to the agent. The critic network is used to fit the Shapley Q value, and the actor network is improved using gradient ascent. Through continuous training iterations, the actor is guided to adjust its parameters and make the best decision.

[0213] In this embodiment, the experiment takes the adjustment of the monthly power outage plan as an example. The daily decision is used as a step in the reinforcement learning process. 30 steps form an episode. The average agent behavior reward value per episode is used as the vertical axis, and the number of training episodes is used as the horizontal axis. The experimental data of the relationship between the comprehensive reward value of the agent trained using the method proposed in this invention and the number of iterations is obtained as follows: Figure 4 shown.

[0214] The beneficial effect of the present invention is that, compared with the prior art,

[0215] (1) The present invention establishes a power grid simulation environment, formulates environmental constraints related to various power outage plans, introduces a multi-intelligence collaborative confrontation AC reinforcement learning algorithm based on Shapley value, and calculates the reward value of the agent's action according to the designed multi-objective reward function to feed back to the agent. Through continuous training iteration, the optimal decision of the power outage plan is finally obtained, which ensures the real-time and collaborative nature of the agent's decision. Through the mutual cooperation and confrontation between each agent, the multi-objective power outage plan scheduling problem that takes into account safety, supply and consumption is achieved. By constructing multi-agent reinforcement learning training to deal with different goals, the decision-making results are more efficient and accurate.

[0216] (2) The present invention takes safety as the basic premise. When deciding the collaborative action strategy of each intelligent agent, the decision-making action of the intelligent agent that ensures safety is given the highest weight. Through the scheme of the main target intelligent agent providing a safety guarantee, the final power outage plan scheduling decision can focus on the realization of safety goals and improve the safety and reliability of power grid operation.

[0217] (3) The present invention establishes a power grid simulation environment through distributed accelerated power flow calculation to support the rapid and effective training of reinforcement learning agents, so that the agents can make the best power outage plan optimization and scheduling scheme, minimize the scope and duration of equipment power outages, and improve the reliability and economy of power grid operation.

[0218] (4) The present invention explains the relationship between the constraints and the power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, establishes interpretable decision rules for the power outage plan using the decision tree algorithm, and constructs an interpretable model. From the three aspects of the interpretability of the training environment, the interpretability of the decision actions, and the interpretability of the training process, the process of multi-agent action decision-making in the comprehensive optimization of the power outage plan can be clearly explained, ensuring the transparency of the agent's decision-making, which is conducive to the user's understanding of the agent's behavior. At the same time, the interpretable model can help users analyze the decision-making basis of the agent in a specific environmental state, thereby facilitating the discovery of unreasonable decisions and timely adjustment of strategies. The optimization and scheduling of the power grid outage plan is more stable and has higher interpretability.

[0219] Example 2:

[0220] like Figure 5 As shown, the present invention provides a power outage scheduling system based on multi-agent interpretable reinforcement learning. The system is used to implement the steps of the method in the above embodiment 1, and the system specifically includes:

[0221] The environment establishment module is used to establish a power grid simulation environment as an interactive environment for training multiple intelligent agents based on the topology and operating parameters of the target power grid;

[0222] The space construction module is used to construct the state space and action space of the agent based on the current orchestration plan options;

[0223] The reward function design module is used to design the reward function for each corresponding agent based on the three optimization objectives of ensuring safety, ensuring supply, and ensuring consumption;

[0224] The agent training module is used to train the action strategies of each agent based on the specified environmental constraints, the agent's state space, action space, and various reward functions, using a multi-agent collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​and combined with a power grid simulation environment;

[0225] The collaborative action decision module is used to integrate the trained action strategies of each intelligent agent according to the set proportional weights, decide the final collaborative action strategy, and thus generate the optimal power outage plan scheduling scheme.

[0226] Specifically, when the power grid simulation environment in the agent training module trains the action strategy of each agent:

[0227] The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current agent collaborative action strategy, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic results after the power grid state changes, thereby completing the simulation of power grid operation to support the use of historical power grid data for iterative training during the agent reinforcement learning process.

[0228] In this embodiment, the distributed power flow calculation includes: tailoring the overall power grid model according to the supply area relationship based on the topology structure and sensitivity analysis; deploying the tailored power grid models of each supply area on each slave node of a distributed cluster containing multiple computing nodes; parallelly calculating the boundary power flow of the supply area power grid model in each slave node, and transmitting the supply area boundary power flow results calculated by each slave node to the master node of the cluster; the master node verifies the boundary power flow results of each supply area, and integrates the power flow data of each supply area, and finally forms the power flow calculation results of the overall power grid model.

[0229] The spatial construction module constructs the agent's state space, integrating information about power flow patterns, grid load, unit output, and the initial requirements of the outage plan. This state space is updated in real time via a data bus to ensure that multiple agents make decisions based on the latest outage plan data at each time step. Constructing the agent's action space involves defining the range of possible actions the agent can take, including the specific timing of equipment outages, while adhering to the physical constraints of the grid system and the operational requirements of the outage.

[0230] In this embodiment, the environmental constraints include: line flow safety constraints, section flow safety constraints, power outage plan non-change constraints, simultaneous power outage constraints, and power outage mutual exclusion constraints.

[0231] In the reward function design module, the expressions of the reward functions for each corresponding agent are designed based on the three optimization objectives of ensuring safety, ensuring supply, and ensuring consumption, as follows:

[0232]

[0233] Where, is the reward function based on the safety objective, and Represent the coefficient weights of the current load term and the voltage load rate term,n and m Respectively represent the total number of AC lines and buses in the power grid under the current strategy, Indicates time line The current, Indicates line The current limit, Indicates time busbar The voltage, Indicates busbar Voltage limit is the set minimum value; is the reward function based on the supply guarantee objective, The quantitative values ​​of the impact of power outage frequency on users , Quantified value of the impact of power supply on users , Power outage for user maintenance and power supply section safety bonus value The weight coefficient of is the reward function based on the guaranteed consumption objective, and Respectively represent the reward value of new energy power outage units and power transmission section safety bonus value The weight coefficient of .

[0234] As an embodiment of the present invention, the agent training module includes:

[0235] A state selection unit is used to randomly select an environment state from the current space state;

[0236] The action decision unit is used to use the three Actor networks currently updated to determine the three corresponding agents. i 1. i 2 and i 3. Determine the action strategy to be taken under the current environmental conditions;

[0237] The action integration unit is used to integrate the current action strategies of each agent according to the set proportional weight to obtain the current collaborative action strategy;

[0238] A simulation judgment unit is used to modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraints are met. If not, the state selection unit is instructed to reselect an environmental state; if so, the current collaborative action strategy update space state is fed back to the reward calculation unit;

[0239] Reward calculation unit, used to calculate the agenti 1. i 2 and i 3 The corresponding reward function values ​​under its current action policy 、 and , and use the gradient descent method to update the critic network corresponding to each agent, and use the currently updated critic networks to fit the Shapley Q value of each agent respectively;

[0240] The update iteration unit is used to update each Actor network using the gradient ascent method based on the current Shapley Q value of each intelligent agent, and feed it back to the state selection unit for iterative training until the set number of iterations is met and the final collaborative action strategy is output.

[0241] Furthermore, the action integration unit includes:

[0242] The normalization subunit is used to normalize the decision actions of each device in each agent's action strategy to the range of (-1, 1);

[0243] The integration subunit is used to integrate the execution actions of each device through the following formula:

[0244]

[0245] Where, represents the collaborative action strategy for device z, Represents the current agent i 1. i 2 and i The decision action for device z in the action strategy of 3; Represented as agents i 1. i 2 and i 3 sets the proportional weight, ;in, .

[0246] Furthermore, in the reward calculation unit, the expressions for fitting the Shapley Q value of each agent are as follows:

[0247]

[0248] Where, Representing an agent i In the environmental state S Execute action strategy a i Shapley's Q value; Representing an agent i In the environmental state S Execute action strategya i The corresponding Q value is obtained by fitting the reward value corresponding to the agent based on the Bell optimal equation of reinforcement learning and using the Critic network; the subscript C represents the set of all agents, C / i Indicates exclusion of agents i The rest of the agents, Indicates that the remaining agents are in the environment state S Execute respective action strategies The sum of the corresponding Q values.

[0249] In a preferred but non-limiting embodiment, the system further comprises:

[0250] The explanation module is used to explain the relationship between constraints and outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, record the state, action selection, constraints and final results of the intelligent agent when making decisions, and trace the decision-making process through the log module of the data reading and writing engine to facilitate understanding of the intelligent agent's behavioral logic; at the same time, charts and data flow diagrams are used to represent the real-time state and decision-making process of the intelligent agent, so as to display the state changes, action selection and constraints of the intelligent agent to the user.

[0251] In this embodiment, the interpretation module includes:

[0252] The decision classification unit is used to classify the action data determined by the agent during iterative training using a decision tree algorithm, so as to identify the device decision actions that cause the collaborative action strategy to fail to meet various constraints.

[0253] A construction unit is used to analyze the importance of each device's decision-making action in affecting each constraint condition by building an interpretable model;

[0254] The interpretable model includes a recurrent neural network (RNN) encoder, a multi-layer perceptron (MLP) encoder, a self-explanatory model, and a linear regression unit connected in sequence.

[0255] Specifically, the building blocks include:

[0256] The encoding subunit is used to take the environment state, action strategy, reward function parameters and constraint condition parameters that the agent confronts during training as the input parameters of the model, and encode the input parameters using the RNN encoder to capture the decision-step characteristics of the agent;

[0257] Learning subunits, using MLP encoders to learn features that RNN encoders fail to capture, to obtain the overall adversarial features of the agent;

[0258] A self-explanatory subunit, configured to establish a correlation between the decision-step feature and the overall adversarial feature through a Gaussian process in a self-explanatory model, thereby obtaining a correlation feature between the decision-step and the adversarial round;

[0259] The linear fitting unit is used to input the correlation features obtained each time during the training process into the linear regression unit for linear fitting, and obtain the regression coefficient of each device decision action, which is used to express the degree of influence of each device decision action on each constraint condition and the final reward, that is, to complete the explanation of the training environment, training process and decision action.

[0260] It is further explained that this embodiment dynamically obtains and analyzes the system status, action space and constraints through the data service bus of the control cloud platform, and constructs a power outage plan scheduling system based on multi-agent interpretable reinforcement learning. The data reading and writing engine based on the control cloud platform enables multiple agents to obtain information related to the power outage plan in real time, ensuring the accuracy and timeliness of decision-making. At the same time, the interpretation module introduces an interpretable reinforcement learning analysis method to increase the transparency of decision-making, making it easier for power grid control personnel to understand the decision-making behavior of multiple agents; the collaborative action decision module introduces a multi-agent action collaborative confrontation method to achieve optimal decision-making based on multiple objectives. At the same time, the system proposed in the present invention can also support single-agent decision training to adapt to reinforcement learning algorithms with different numbers of agents.

[0261] In actual application, the specific process of building a power outage planning system based on multi-agent interpretable reinforcement learning in this embodiment is as follows:

[0262] (1) Define constraints. The flow safety constraints, immutable plan constraints, simultaneous outage constraints, and mutual exclusion constraints defined in the interpretable reinforcement learning analysis method are used as training constraints. These constraints will be transmitted by the data service bus as boundary conditions for action selection. This ensures that the online decision-making of multiple agents complies with the business logic of the power grid outage plan, while making it easier for users and decision makers to understand the decision-making behavior of the agents. By using the data service bus to receive changing data in real time and update the parameters of the constraints, the agents can adapt to changing needs under different states of the power grid control system, thereby making more accurate decisions.

[0263] (2) Establish an interpretability enhancement mechanism. Based on the decision tree framework in the interpretable reinforcement learning analysis method, the relationship between the constraints and the power outage actions in the power outage plan is explained. The state, action selection, constraints and final results of each agent when making a decision are recorded. The decision process is traced through the log module of the data reading and writing engine to facilitate the understanding of the agent's behavioral logic. Based on the decision data, an explanation module is added, which can be based on rule-based explanations or attention-based explanations, so that the system can explain why each agent chooses a certain action or avoids a certain constraint. The real-time state and decision-making process of the agent are represented by charts and data flow diagrams, making the system operation transparent and controllable, thereby displaying the state changes, action selections and constraints of the agent to the user.

[0264] (3) A multi-agent collaborative confrontation AC algorithm based on Shapley value is introduced to enable each agent to cooperate and confront each other with the goal of ensuring safety, supply and consumption. With safety as the primary goal, each agent determines its own actions and obtains the optimal strategy for power outage scheduling.

[0265] (4) Build a data reading and writing engine. The data reading and writing engine encapsulates functions such as data reading, parsing, storage, and updating. It accesses the state, action, and constraint data of multiple agents through the API interface of the data bus, conveniently reads the current state and constraints, and selects the corresponding action. In addition, to support multi-agent parallel reading and writing, an efficient concurrency control mechanism is designed, using distributed locks and timestamp-based access control strategies to ensure the security and consistency of data access and avoid data competition or conflicts among multiple agents for the same resource.

[0266] (5) Construct a power grid simulation environment, run the BPA file based on the power grid, extract and model the power grid data modified by the intelligent agent decision, and further tailor the power grid model using a method based on topology structure and sensitivity analysis. The Newton-Pull iterative solution method is used to calculate the power flow for each part of the power grid model after tailoring. The architecture proposed in this embodiment will use a distributed computing method to perform parallel power flow calculation on the tailored model to speed up the processing. At the same time, in order to ensure flexibility, the traditional centralized power flow calculation method can also be used. The distributed power flow calculation process is as follows: Figure 3 As shown:

[0267] (6) Based on the designed multi-objective reward function, the reward value of the agent's action is calculated and fed back to the agent. The algorithm finally obtains the optimal decision through continuous training iteration.

[0268] Through the above process, a power outage planning system based on multi-agent interpretable reinforcement learning is constructed, which ensures the real-time, collaborative and transparent decision-making of the agents, making the optimized planning of power grid outages more stable and more interpretable.

[0269] The system proposed in the present invention can also be compatible with single-agent algorithms according to actual application conditions. The constructed power grid simulation environment can choose centralized traditional computing or distributed parallel computing power flow solution methods according to user needs to support the rapid and effective training of reinforcement learning agents. The system architecture is flexible and applicable to various power grid simulation scenarios.

[0270] The power outage plan scheduling system based on multi-agent interpretable reinforcement learning provided in the embodiment of the present invention and the power outage plan scheduling method based on multi-agent interpretable reinforcement learning provided in Example 1 are based on the same technical concept and can produce the beneficial effects as described in Example 1. For the contents not described in detail in this embodiment, please refer to Example 1.

[0271] Example 3:

[0272] An embodiment of the present invention provides a terminal comprising a processor and a storage medium, the terminal being an embedded computer system device. The terminal's storage medium is used to store instructions, and the memory comprises a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database; the internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium; and the database stores instruction data. The terminal's processor is configured to operate according to the instructions provided by the storage medium to execute the steps of the method for scheduling power outages based on multi-agent interpretable reinforcement learning as described in any one of the first embodiments.

[0273] Example 4:

[0274] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the method described in any one of the first embodiments are implemented.

[0275] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0276] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0277] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0278] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, the state information of the computer-readable program instructions is used to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), so that the electronic circuit can execute the computer-readable program instructions, thereby implementing various aspects of the present disclosure.

[0279] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for scheduling power outages based on multi-agent interpretable reinforcement learning, characterized in that: Methods include: According to the topology and operating parameters of the target power grid, a power grid simulation environment is established as an interactive environment for training multi-agents; Construct the state space and action space of the agent based on the current orchestration plan options; Constructing the state space of the intelligent agents involves integrating the state information of power flow of power grid equipment, grid load, output of each unit, and the initial requirements of the power outage plan into the state space. Simultaneously, the state space is updated in real time via the data bus to ensure that multiple intelligent agents can make decisions based on the latest power outage plan data at each time step. Constructing the agent's action space involves defining the space of actions the agent can take, including the specific time of the equipment outage, while complying with the physical constraints of the grid system and the power outage operation requirements; Based on the three optimization goals of ensuring safety, ensuring supply, and ensuring consumption, the reward functions of each corresponding agent are designed, and their expressions are as follows: Where, is the reward function based on the safety objective, and Represent the coefficient weights of the current load term and the voltage load rate term, n and m Respectively represent the total number of AC lines and buses in the power grid under the current strategy, Indicates time line The current, Indicates line The current limit, Indicates time busbar The voltage, Indicates busbar Voltage limit is the set minimum value; is the reward function based on the supply guarantee objective, The quantitative values ​​of the impact of power outage frequency on users , Quantified value of the impact of power supply on users , Power outage for user maintenance and power supply section safety bonus value The weight coefficient of is the reward function based on the guaranteed consumption objective, and Respectively represent the reward value of new energy power outage units and power transmission section safety bonus value The weight coefficient of Based on the established environmental constraints, the agent's state space, action space, and various reward functions, a multi-agent collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​is used in conjunction with a power grid simulation environment to train the action strategies of each agent. The trained action strategies of each intelligent agent are integrated according to the set proportional weights to decide the final collaborative action strategy, thereby generating the optimal power outage plan scheduling scheme.

2. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that: When training the action strategies of each agent in the power grid simulation environment: The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current agent collaborative action strategy, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic results after the power grid state changes, thereby completing the simulation of power grid operation to support the use of historical power grid data for iterative training during the agent reinforcement learning process.

3. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 2, characterized in that: The distributed power flow calculation includes: Based on topological structure and sensitivity analysis, the overall power grid model is tailored according to the supply area relationship; Deploy the tailored power grid models of each supply area on each slave node of a distributed cluster containing multiple computing nodes; Parallel calculation of the boundary flow of the power grid model in each slave node, and transmission of the boundary flow results of the power grid calculated by each slave node to the master node of the cluster; The master node verifies the boundary flow results of each supply area and integrates the flow data of each supply area to finally form the flow calculation results of the overall power grid model.

4. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that: The environmental constraints include: line flow safety constraints, section flow safety constraints, power outage plan unchangeable constraints, simultaneous power outage constraints and / or power outage mutual exclusion constraints.

5. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that: The steps of adopting the multi-intelligence cooperative countermeasure AC reinforcement learning algorithm based on Shapley value and combining it with the power grid simulation environment to train the action strategy of each intelligent agent include: Step 1: Randomly select an environment state from the current space state; Step 2: Use the three Actor networks updated currently to represent the three corresponding agents. i 1. i 2 and i 3. Determine the action strategy to be taken under the current environmental conditions; Step 3: Integrate the current action strategies of each agent according to the set proportional weights to obtain the current collaborative action strategy; Step 4: Modify the corresponding parameters in the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraints are met. If not, return to step 1; if so, update the spatial state based on the current collaborative action strategy and execute step 5; Step 5: Calculate the agents separately i 1. i 2 and i 3 The corresponding reward function values ​​under its current action policy 、 and , and use the gradient descent method to update the critic network corresponding to each agent, and use the currently updated critic networks to fit the Shapley Q value of each agent respectively; Step 6: Based on the current Shapley Q value of each agent, use the gradient ascent method to update each Actor network, and return to step 1 for iterative training until the set number of iterations is met, and output the final collaborative action strategy.

6. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 5, characterized in that: The step of integrating the current action strategies of each agent according to the set proportional weights includes: The decision actions of each agent in the action strategy for each device are normalized to the range of (-1, 1), and the execution actions of each device are integrated using the following formula: Where, represents the collaborative action strategy for device z, Represents the current agent i 1. i 2 and i The decision action for device z in the action strategy of 3; Represented as agents i 1. i 2 and i 3 sets the proportional weight, ;in, .

7. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 5, characterized in that: The expressions for fitting the Shapley Q value of each agent are as follows: Where, Representing an agent i In the environmental state S Execute action strategy a i Shapley's Q value; Representing an agent i In the environmental state S Execute action strategy a i The corresponding Q value is obtained by fitting the reward value corresponding to the agent based on the Bell optimal equation of reinforcement learning and using the Critic network; the subscript C represents the set of all agents, C\ i Indicates exclusion of agents i The rest of the agents, Indicates that the remaining agents are in the environment state S Execute respective action strategies The sum of the corresponding Q values.

8. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that: The method further comprises: Based on the decision tree framework in the interpretable reinforcement learning algorithm, the relationship between the constraints and power outage actions in the power outage plan is explained, and the state, action selection, constraints and final results of the intelligent agent when making decisions are recorded. The decision-making process is traced through the log module of the data reading and writing engine to facilitate the understanding of the behavioral logic of the intelligent agent. At the same time, charts and data flow diagrams are used to represent the real-time state and decision-making process of the intelligent agent, so as to display the state changes, action selection and constraints of the intelligent agent to the user.

9. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 8, characterized in that: The step of explaining the relationship between the constraints and the blackout actions in the blackout plan based on the decision tree framework in the interpretable reinforcement learning algorithm includes: Use the decision tree algorithm to classify the action data determined by the agent during iterative training to identify the device actions that cause the collaborative action strategy to fail to meet various constraints. By building an interpretable model, we can analyze the importance of each device's decision-making action in affecting each constraint condition. The interpretable model includes a recurrent neural network (RNN) encoder, a multi-layer perceptron (MLP) encoder, a self-explanatory model, and a linear regression unit connected in sequence.

10. The method for scheduling power outages based on multi-agent interpretable reinforcement learning according to claim 9, characterized in that: The steps of analyzing the importance of each device decision action in affecting each constraint condition by constructing an interpretable model include: The environment state, action strategy, reward function parameters, and constraint condition parameters that the agent confronts during training are used as the input parameters of the model. The RNN encoder is used to encode the input parameters to capture the decision-making characteristics of the agent. Use the MLP encoder to learn features that the RNN encoder fails to capture to obtain the overall adversarial features of the agent; Establishing the correlation between the decision step features and the overall adversarial features through the Gaussian process in the self-explanatory model, and obtaining the correlation features between the decision step and the adversarial round; The correlation features obtained each time during the training process are input into the linear regression unit for linear fitting to obtain the regression coefficient of each device decision action, which is used to represent the degree of influence of each device decision action on each constraint condition and the final reward, thus completing the explanation of the training environment, training process and decision action.

11. A power outage scheduling system based on multi-agent interpretable reinforcement learning, running the power outage scheduling method based on multi-agent interpretable reinforcement learning according to any one of claims 1 to 10, characterized in that: The system includes: The environment establishment module is used to establish a power grid simulation environment as an interactive environment for training multiple intelligent agents based on the topology and operating parameters of the target power grid; The space construction module is used to construct the state space and action space of the agent based on the current orchestration plan options; Constructing the state space of the intelligent agents involves integrating the state information of power flow of power grid equipment, grid load, output of each unit, and the initial requirements of the power outage plan into the state space. Simultaneously, the state space is updated in real time via the data bus to ensure that multiple intelligent agents can make decisions based on the latest power outage plan data at each time step. Constructing the agent's action space involves defining the space of actions the agent can take, including the specific time of the equipment outage, while complying with the physical constraints of the grid system and the power outage operation requirements; The reward function design module is used to design the reward function for each corresponding agent based on the three optimization objectives of ensuring safety, ensuring supply, and ensuring consumption. The expressions are as follows: Where, is the reward function based on the safety objective, and Represent the coefficient weights of the current load term and the voltage load rate term, n and m Respectively represent the total number of AC lines and buses in the power grid under the current strategy, Indicates time line The current, Indicates line The current limit, Indicates time busbar The voltage, Indicates busbar Voltage limit is the set minimum value; is the reward function based on the supply guarantee objective, The quantitative values ​​of the impact of power outage frequency on users , Quantified value of the impact of power supply on users , Power outage for user maintenance and power supply section safety bonus value The weight coefficient of is the reward function based on the guaranteed consumption objective, and Respectively represent the reward value of new energy power outage units and power transmission section safety bonus value The weight coefficient of The agent training module is used to train the action strategies of each agent based on the specified environmental constraints, the agent's state space, action space, and various reward functions, using a multi-agent collaborative adversarial AC reinforcement learning algorithm based on Shapley values ​​and combined with a power grid simulation environment; The collaborative action decision module is used to integrate the trained action strategies of each intelligent agent according to the set proportional weights, decide the final collaborative action strategy, and thus generate the optimal power outage plan scheduling scheme.

12. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 11, characterized in that: When the power grid simulation environment in the agent training module trains the action strategy of each agent: The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current agent collaborative action strategy, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic results after the power grid state changes, thereby completing the simulation of power grid operation to support the use of historical power grid data for iterative training during the agent reinforcement learning process.

13. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 12, characterized in that: The distributed power flow calculation includes: Based on topological structure and sensitivity analysis, the overall power grid model is tailored according to the supply area relationship; Deploy the tailored power grid models of each supply area on each slave node of a distributed cluster containing multiple computing nodes; Parallel calculation of the boundary flow of the power grid model in each slave node, and transmission of the boundary flow results of the power grid calculated by each slave node to the master node of the cluster; The master node verifies the boundary flow results of each supply area and integrates the flow data of each supply area to finally form the flow calculation results of the overall power grid model.

14. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 11, characterized in that: In the agent training module, environmental constraints include: line flow safety constraints, section flow safety constraints, power outage plan unchangeable constraints, simultaneous power outage constraints and / or power outage mutual exclusion constraints.

15. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 11, characterized in that: The agent training module includes: A state selection unit is used to randomly select an environment state from the current space state; The action decision unit is used to use the three Actor networks currently updated to determine the three corresponding agents. i 1. i 2 and i 3. Determine the action strategy to be taken under the current environmental conditions; The action integration unit is used to integrate the current action strategies of each agent according to the set proportional weight to obtain the current collaborative action strategy; A simulation judgment unit is used to modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraints are met. If not, the state selection unit is instructed to reselect an environmental state; if so, the current collaborative action strategy update space state is fed back to the reward calculation unit; Reward calculation unit, used to calculate the agent i 1. i 2 and i 3 The corresponding reward function values ​​under its current action policy 、 and , and use the gradient descent method to update the critic network corresponding to each agent, and use the currently updated critic networks to fit the Shapley Q value of each agent respectively; The update iteration unit is used to update each Actor network using the gradient ascent method based on the current Shapley Q value of each intelligent agent, and feed it back to the state selection unit for iterative training until the set number of iterations is met and the final collaborative action strategy is output.

16. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 15, characterized in that: The action integration unit includes: The normalization subunit is used to normalize the decision actions of each device in each agent's action strategy to the range of (-1, 1); The integration subunit is used to integrate the execution actions of each device through the following formula: Where, represents the collaborative action strategy for device z, Represents the current agent i 1. i 2 and i The decision action for device z in the action strategy of 3; Represented as agents i 1. i 2 and i 3 sets the proportional weight, ;in, .

17. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 15, characterized in that: In the reward calculation unit, the expressions for fitting the Shapley Q value of each agent are as follows: Where, Representing an agent i In the environmental state S Execute action strategy a i Shapley's Q value; Representing an agent i In the environmental state S Execute action strategy a i The corresponding Q value is obtained by fitting the reward value corresponding to the agent based on the Bell optimal equation of reinforcement learning and using the Critic network; the subscript C represents the set of all agents, C\ i Indicates exclusion of agents i The rest of the agents, Indicates that the remaining agents are in the environment state S Execute respective action strategies The sum of the corresponding Q values.

18. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 11, characterized in that: The system further comprises: The explanation module is used to explain the relationship between constraints and outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, record the state, action selection, constraints and final results of the intelligent agent when making decisions, and trace the decision-making process through the log module of the data reading and writing engine to facilitate understanding of the intelligent agent's behavioral logic; at the same time, charts and data flow diagrams are used to represent the real-time state and decision-making process of the intelligent agent, so as to display the state changes, action selection and constraints of the intelligent agent to the user.

19. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 18, characterized in that: The interpretation module includes: The decision classification unit is used to classify the action data determined by the agent during iterative training using a decision tree algorithm, so as to identify the device decision actions that cause the collaborative action strategy to fail to meet various constraints. A construction unit is used to analyze the importance of each device's decision-making action in affecting each constraint condition by building an interpretable model; The interpretable model includes a recurrent neural network (RNN) encoder, a multi-layer perceptron (MLP) encoder, a self-explanatory model, and a linear regression unit connected in sequence.

20. The power outage planning system based on multi-agent interpretable reinforcement learning according to claim 19, characterized in that: The building block comprises: The encoding subunit is used to take the environment state, action strategy, reward function parameters and constraint condition parameters that the agent confronts during training as the input parameters of the model, and encode the input parameters using the RNN encoder to capture the decision-step characteristics of the agent; Learning subunits, using MLP encoders to learn features that RNN encoders fail to capture, to obtain the overall adversarial features of the agent; A self-explanatory subunit, configured to establish a correlation between the decision-step feature and the overall adversarial feature through a Gaussian process in a self-explanatory model, thereby obtaining a correlation feature between the decision-step and the adversarial round; The linear fitting unit is used to input the correlation features obtained each time during the training process into the linear regression unit for linear fitting, and obtain the regression coefficient of each device decision action, which is used to express the degree of influence of each device decision action on each constraint condition and the final reward, that is, to complete the explanation of the training environment, training process and decision action.

21. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 10.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Multi-regional power grid collaborative optimization method, system and device and readable storage medium

    CN115333111A

  • Multi-agent deep reinforcement learning-based power grid power failure arrangement system and method, and medium

    CN119784018A