Power-cut plan arrangement method and system based on multi-agent interpretable reinforcement learning

Through multi-agents, reinforcement learning methods can be explained, and the agents can be trained to coordinately optimize the power outage plan in the grid simulation environment, solving the problem of difficulty in taking into account the grid safety, supply and consumption goals in the existing technology, and achieving efficient, accurate and transparent power outage plan decisions.

CN120197915AActive Publication Date: 2025-06-24BEIJING KEDONG ELECTRIC POWER CONTROL SYST CO LTD +2

Patent Information

Application Number
CN202510669075.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively take into account the three control goals of safety, supply and consumption in power grid operation, resulting in low efficiency and low intelligence on the optimization and orchestration of power outage plans.

Method used

Using a method based on multi-agents to interpret reinforcement learning, by establishing a grid simulation environment and designing multi-objective reward function, the agent is trained using Shapley's multi-intelligent collaboration against AC reinforcement learning algorithm to achieve collaborative optimization decisions for power outage plans.

Benefits of technology

It realizes efficient and accurate decision-making of power outage plans, shortens the range and time of equipment power outages, improves the reliability and economicality of power grid operation, and improves the transparency of decision-making through interpretability enhancement mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197915A_ABST
    Figure CN120197915A_ABST
Patent Text Reader

Abstract

The invention discloses a power-cut plan arrangement method and system based on multi-agent interpretable reinforcement learning. The method comprises the following steps: establishing a power grid simulation environment as an interaction environment for training multiple agents; constructing a state space and an action space of the intelligent agent according to the current arrangement plan selectable scheme; designing a reward function of each corresponding agent based on three optimization objectives of guaranteed safety, guaranteed supply and guaranteed consumption; adopting a multi-intelligent cooperative confrontation AC reinforcement learning algorithm based on a Shapley value and combining with a power grid simulation environment to train an action strategy of each agent; and integrating the action strategies of the trained intelligent agents according to a set proportion weight, and deciding a final cooperative action strategy, thereby generating an optimal power-cut plan arrangement scheme. According to the method, the multi-target power-cut plan arrangement problem is considered through cooperation of multiple agents, the optimal strategy scheme can be made more efficiently and accurately, and the operation reliability of a power grid is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power systems, and particularly relates to a power outage plan scheduling method and system based on multi-agent interpretable reinforcement learning. Background Art

[0002] With the continuous expansion of the power grid scale, the number of power outage plans is increasing day by day. On the premise that the daily maintenance requirements of equipment increase correspondingly, the main grid also faces various power outage demands, and the safety requirements for on-site operations will also significantly increase the number of power outages, making the optimization and scheduling of power grid outage plans exhibit characteristics such as high variable dimensions and strong non-linearity. Moreover, during the scheduling process, the contradictions among the basic principles such as power grid operation safety, power supply guarantee, and new energy consumption are becoming increasingly prominent, and it is difficult to simultaneously take into account the three control objectives of ensuring safety, supply, and consumption.

[0003] Currently, for the problem of solving the power grid outage plan optimization model, it is mainly limited to algorithms based on the swarm intelligence optimization framework. Typical representatives include heuristic methods such as particle swarm optimization algorithm, dragonfly algorithm, genetic algorithm, etc.; although these algorithms have the ability to avoid local extrema to a certain extent by simulating the behavioral characteristics of biological groups, they still essentially belong to the static optimization paradigm and lack real-time interaction and feedback mechanisms with the dynamic environment. At the same time, with the acceleration of the construction process of the new power system, the power grid topology is becoming increasingly complex, which makes the solution space dimension of the outage plan model show a superlinear expansion trend. When the number of equipment nodes reaches the thousands level, the computational burden of traditional heuristic algorithms will increase exponentially, resulting in the optimization efficiency being difficult to meet the time requirements of power dispatching real-time decision-making; therefore, such algorithms show significant limitations in terms of dynamic adaptability and online optimization efficiency. In addition, the existing power grid outage plan intelligent scheduling devices and methods based on deep reinforcement learning, although they propose to use the reinforcement learning method to try to take into account multi-objective optimization, mainly explore based on the scheduling experts' scheduling experience, and have problems such as low efficiency, limited content, and low intelligence level.

[0004] Therefore, there is an urgent need for a power outage plan scheduling optimization strategy for the multi-objective power outage plan scheduling problem that can take into account safety guarantee, supply guarantee, and consumption guarantee, so as to adapt to the increasingly complex new power system today. Summary of the Invention

[0005] To solve the deficiencies in the prior art, the present invention provides a power outage plan scheduling method and system based on multi-agent interpretable reinforcement learning, which uses multi-agents to achieve collaborative optimization decisions for the three objectives of safety guarantee, supply guarantee, and consumption guarantee that need to be considered in the power grid outage plan, constructs multi-agents to perform reinforcement learning training for different objectives, and can make the best strategy plan more efficiently and accurately, minimizing the power outage range and time of equipment, and improving the reliability and economy of power grid operation.

[0006] The present invention adopts the following technical solutions.

[0007] In a first aspect, the present invention provides a power outage plan scheduling method based on multi-agent interpretable reinforcement learning. The method includes: Based on the topological structure and operating parameters of the target power grid, establish a power grid simulation environment as the interaction environment for training multi-agents; Construct the state space and action space of the agent according to the optional solutions of the current scheduling plan; Design the reward functions of the corresponding agents based on three optimization goals of ensuring safety, ensuring power supply, and ensuring power consumption; Based on the formulated environmental constraints, the state space and action space of the agents, and each reward function, adopt the multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value and combine it with the power grid simulation environment to train the action strategies of each agent; Integrate the action strategies of the trained agents according to the set proportional weights, and decide the final collaborative action strategy, so as to generate the optimal power outage plan scheduling scheme.

[0008] In combination with the first aspect, optionally, the step of establishing a power grid simulation environment as the interaction environment for training multi-agents includes: Modify the corresponding parameters on the BPA file based on the current agent collaborative action strategy, and perform distributed power flow calculation based on the modified parameters, so as to simulate the operation of the power grid to support the iterative training using historical power grid data during the agent reinforcement learning process.

[0009] In combination with the first aspect, optionally, the distributed power flow calculation includes: Crop the overall power grid model according to the supply area relationship based on the topological structure and sensitivity analysis; Deploy the cropped power grid models of each supply area to each slave node of a distributed cluster including multiple computing nodes; Parallel-compute the boundary power flows of the power grid models of each supply area in each slave node, and transmit the boundary power flow results calculated by each slave node to the master node of the cluster; The master node checks the boundary power flow results of each supply area, and integrates the power flow data of each supply area to finally form the power flow calculation result of the overall power grid model.

[0010] In combination with the first aspect, optionally, constructing the state space of the agent includes: integrating the state information of the power grid equipment power flow, the power grid load, the output of each unit, and the initial demand of the power outage plan into the state space; at the same time, updating the state space in real time through the data bus to ensure that multi-agents can make decisions according to the latest power outage plan data at each time step.

[0011] In combination with the first aspect, optionally, constructing the action space of the agent includes: defining the set of actions that the agent can take, including the specific time of equipment power outage, on the premise of meeting the physical constraints of the power grid system and the requirements of power outage operations.

[0012] In combination with the first aspect, optionally, the environmental constraint conditions include: line power flow security constraints, section power flow security constraints, power outage non-change plan constraints, simultaneous power outage constraints, and / or power outage mutual exclusion constraints.

[0013] In combination with the first aspect, optionally, the expressions for designing the reward functions of the corresponding agents based on the three optimization goals of ensuring safety, ensuring supply, and ensuring consumption are as follows:

[0014] In the formula, is the reward function based on the goal of ensuring safety, and respectively represent the coefficient weights of the current load term and the voltage load rate term, n and m respectively represent the total number of AC lines and the total number of busbars in the power grid under the current strategy, represents the time the current of line , represents line the current limit of, represents the time the voltage of bus , represents bus the voltage limit of is a set minimum value; is the reward function based on the goal of ensuring supply, respectively represent the quantified value of the impact of the power outage frequency on users , the quantified value of the impact of power supply on users , the power consumption of user maintenance power outage and the power supply section safety reward value the weight coefficients of; is the reward function based on the goal of ensuring consumption, and respectively represent the reward value of new energy shutdown units and the power transmission section safety reward value the weight coefficients of.

[0015] In combination with the first aspect, optionally, the steps of using the multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value and training the action strategies of each agent in combination with the power grid simulation environment include: Step 1: Randomly select an environmental state from the current spatial state; Step 2: Use the three Actor networks updated in the current iteration to determine the action strategies that should be taken by their corresponding three agents i 1, i 2, and i 3 in the current environmental state; Step 3: Integrate the action strategies of the current agents according to the set proportional weights to obtain the current collaborative action strategy; Step 4: Modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraint conditions are met. If not, return to Step 1; if so, update the spatial state based on the current collaborative action strategy and execute Step 5; Step 5: Calculate the respective corresponding reward function values i 1, i 2, and i 3 of the agents under their current action strategies , , and , and update the Critic network corresponding to each agent by using gradient descent. At the same time, use the currently updated Critic networks to fit the Shapley Q values of each agent respectively; Step 6: Update each Actor network by using gradient ascent based on the current Shapley Q values of each agent, and return to execute Step 1 for iterative training until the set number of iterations is met, and output the final collaborative action strategy.

[0016] Combined with the first aspect, optionally, the step of integrating the action strategies of the current agents according to the set proportional weights includes: Normalize the decision actions of each device in the action strategies of each agent to the range of (-1, 1), and integrate the execution actions of each device through the following formula:

[0017] In the formula, represents the collaborative action strategy for device z, respectively represent the decision actions for device z in the action strategies of the current agents i 1, i 2, and i 3; respectively represent the proportional weights set for the agents i 1, i 2, and i 3, ; where, .

[0018] In combination with the first aspect, optionally, the expressions for separately fitting the Shapley Q values of each agent are as follows:

[0019] In the formula, represents the Shapley Q value of agent i executing the action policy S under the environmental state a i ; represents the Q value corresponding to agent i executing the action policy S under the environmental state a i The Q value is obtained by substituting the reward value corresponding to the agent into the Bell optimal equation of reinforcement learning and using the Critic network for fitting; the subscript C represents the set of all agents, C / i represents the remaining agents after excluding agent i , represents the sum of the Q values corresponding to the remaining agents executing their respective action policies S under the environmental state .

[0020] In combination with the first aspect, optionally, the method further includes: Based on the decision tree framework in the interpretable reinforcement learning algorithm, explain the relationship between the constraint conditions and outage actions in the outage plan, record the state, action selection, constraint conditions, and final results when the agent makes a decision, and trace the decision-making process through the log module of the data reading and writing engine to facilitate understanding of the agent's behavior logic; at the same time, use charts and data flow diagrams to represent the real-time state and decision-making process of the agent, so as to display the state changes, action selections, and constraint conditions of the agent to the user.

[0021] In combination with the first aspect, optionally, the step of explaining the relationship between the constraint conditions and outage actions in the outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm includes: Use the decision tree algorithm to classify each action data decided during the iterative training process of the agent to identify which device decision actions make the collaborative action policy not satisfy various constraint conditions; Analyze the importance of each device decision action affecting each constraint condition by constructing an interpretable model; Among them, the interpretable model includes a recurrent neural network RNN encoder, a multi-layer perceptron network MLP encoder, a self-explaining model, and a linear regression unit connected in sequence.

[0022] In combination with the first aspect, optionally, the step of analyzing the importance of the decision-making actions of each device on each constraint condition by constructing an interpretable model includes: Taking the environmental state, action strategy, reward function parameters, and constraint condition parameters that the agent confronts during training as the input parameters of the model, and using an RNN encoder to encode the input parameters to capture the decision-making step features of the agent; Using an MLP encoder to learn the features that the RNN encoder fails to capture to obtain the overall confrontation features of the agent; Establishing the correlation between the decision-making step features and the overall confrontation features through the Gaussian process in the self-explanation model to obtain the association features between the decision-making step and the confrontation round; Inputting the association features obtained each time during the training process into a linear regression unit for linear fitting to obtain the regression coefficients of the decision-making actions of each device, which are used to represent the influence degrees of the decision-making actions of each device on each constraint condition and the final reward, that is, the explanation of the training environment, training process, and decision-making actions is completed.

[0023] In the second aspect, the present invention provides a power outage plan scheduling system based on multi-agent interpretable reinforcement learning, which runs the steps of any one of the methods in the first aspect of the present invention. The system includes: An environment establishment module for establishing a power grid simulation environment as an interactive environment for training multi-agents according to the topological structure and operating parameters of the target power grid; A space construction module for constructing the state space and action space of the agent according to the optional schemes of the current scheduling plan; A reward function design module for designing the reward functions of each corresponding agent based on three optimization objectives of ensuring safety, ensuring power supply, and ensuring consumption; An agent training module for training the action strategies of each agent by using the multi-agent cooperative confrontation AC reinforcement learning algorithm based on the Shapley value and combining with the power grid simulation environment based on the formulated environmental constraint conditions, the state space, action space, and each reward function of the agent; A cooperative action decision-making module for integrating the action strategies of the trained agents according to the set proportional weights, and making a decision on the final cooperative action strategy, thereby generating an optimal power outage plan scheduling scheme.

[0024] In combination with the second aspect, optionally, when the power grid simulation environment in the agent training module trains the action strategies of each agent: The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current agent collaborative action strategy, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic results after the power grid state changes, thereby completing the simulation of power grid operation to support the use of historical power grid data for iterative training in the agent reinforcement learning process.

[0025] In conjunction with the second aspect, optionally, the distributed power flow calculation includes: Based on topological structure and sensitivity analysis, the overall model of the power grid is tailored according to the supply area relationship; Deploy the tailored power grid models of each supply area on each slave node of a distributed cluster including multiple computing nodes; Parallel calculation of the boundary flow of the supply area power grid model in each slave node, and transmission of the supply area boundary flow results calculated by each slave node to the master node of the cluster; The master node verifies the boundary flow results of each supply area and integrates the flow data of each supply area to finally form the flow calculation results of the overall power grid model.

[0026] In combination with the second aspect, optionally, in the space construction module, constructing the state space of the intelligent agent includes: integrating the state information of the power grid equipment flow, the power grid load, the output of each unit and the initial demand of the power outage plan in the state space; and updating the state space in real time through the data bus to ensure that multiple agents can make decisions based on the latest power outage plan data at each time step.

[0027] In combination with the second aspect, optionally, in the space construction module, constructing the action space of the intelligent agent includes: defining the action space that the intelligent agent can take, including the specific time of power outage of the equipment, under the premise of complying with the physical constraints of the power grid system and the power outage operation requirements.

[0028] In combination with the second aspect, optionally, in the intelligent agent training module, the environmental constraints include: line flow safety constraints, section flow safety constraints, power outage plan cannot be changed constraints, simultaneous power outage constraints and / or power outage mutual exclusion constraints.

[0029] In combination with the second aspect, optionally, in the reward function design module, the expressions of the reward functions of the corresponding agents are designed based on the three optimization goals of ensuring safety, ensuring supply, and ensuring consumption as follows:

[0030] In the formula, is the reward function based on the safety objective, and Represent the coefficient weights of the current load term and the voltage load rate term, n and mrespectively represent the total number of AC lines and the total number of buses in the power grid under the current strategy, represents the time line current of, represents line current limit value of, represents the time bus voltage of, represents bus voltage limit value of is the set minimum value; is the reward function based on the power supply guarantee goal, respectively represent the quantization values of the impact degree of power outage frequency on users , the quantization value of the impact degree of power supply on users , the power consumption of user maintenance power outage and the safety reward value of the power supply section weight coefficients of; is the reward function based on the power consumption guarantee goal, and respectively represent the reward value of new energy shutdown units and the safety reward value of the power transmission section weight coefficients of.

[0031] Combined with the second aspect, optionally, the intelligent agent training module includes: A state selection unit for randomly selecting an environmental state from the current spatial state; An action decision unit for using the three Actor networks updated this time to respectively determine the action strategies that should be taken by its corresponding three intelligent agents i 1, i 2 and i 3 in the current environmental state; An action integration unit for integrating the action strategies of the current intelligent agents according to the set proportional weights to obtain the current collaborative action strategy; A simulation judgment unit for modifying the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy and performing power grid operation simulation to judge whether the environmental constraint conditions are met. If not, the state selection unit is made to re-select an environmental state; if so, the current collaborative action strategy is updated to the spatial state and fed back to the reward calculation unit; A reward calculation unit for respectively calculating the corresponding reward function values of the intelligent agents i 1, i 2 and i 3 under their current action strategies , and , and update the Critic network corresponding to each agent in the way of gradient descent. At the same time, use the currently updated Critic networks to respectively fit the Shapley Q-values of each agent; An update iteration unit, which is used to update each Actor network in the way of gradient ascent based on the current Shapley Q-values of each agent, and feed it back to the state selection unit for iterative training until the set number of iterations is satisfied, and output the final collaborative action strategy.

[0032] Combined with the second aspect, optionally, the action integration unit includes: A normalization subunit, which is used to normalize the decision-making actions for each device in the action strategy of each agent to the range of (-1, 1); An integration subunit, which is used to integrate the execution actions for each device respectively through the following formula:

[0033] In the formula, represents the collaborative action strategy for device z, respectively represent the decision-making actions for device z in the action strategies of the current agent i 1, i 2 and i 3; respectively represent the proportional weights set for agents i 1, i 2 and i 3, ; where, .

[0034] Combined with the second aspect, optionally, in the reward calculation unit, the expressions for respectively fitting the Shapley Q-values of each agent are as follows:

[0035] In the formula, represents the Shapley Q-value of agent i executing the action strategy S under the environmental state a i ; represents the Q-value corresponding to agent i executing the action strategy S under the environmental state a i , and the Q-value is obtained by substituting the reward value corresponding to the agent into the Bell optimal equation of reinforcement learning and using the Critic network for fitting; the subscript C represents the set of all agents, C / i represents the remaining agents after excluding agent i , indicating the remaining agents in the environmental state S executing each automatic action policy and the sum of the corresponding Q values in the following.

[0036] Combined with the second aspect, optionally, the system further includes: An explanation module, which is used to explain the relationship between the constraint conditions and the power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, record the state, action selection, constraint conditions and final results of the agent when making decisions, and trace the decision-making process through the log module of the data reading and writing engine, so as to facilitate understanding the behavior logic of the agent; at the same time, use charts and data flow diagrams to represent the real-time state and decision-making process of the agent, so as to display the state changes, action selections and constraint conditions of the agent to the user.

[0037] Combined with the second aspect, optionally, the explanation module includes: A decision classification unit, which is used to classify each action data decided during the iterative training process of the agent by using the decision tree algorithm, so as to identify which device decision actions make the collaborative action policy not meet various constraint conditions; A construction unit, which is used to analyze the importance of each device decision action affecting each constraint condition by constructing an interpretable model; Among them, the interpretable model includes a recurrent neural network RNN encoder, a multi-layer perceptron network MLP encoder, a self-explanation model and a linear regression unit connected in sequence.

[0038] Combined with the second aspect, optionally, the construction unit includes: An encoding subunit, which is used to use the environmental state, action policy, reward function parameters and constraint condition parameters that the agent confronts during the training process as the input parameters of the model, and use the RNN encoder to encode the input parameters to capture the decision step features of the agent; A learning subunit, which uses the MLP encoder to learn the features that the RNN encoder fails to capture to obtain the overall confrontation features of the agent; A self-explanation subunit, which is used to establish the correlation between the decision step features and the overall confrontation features through the Gaussian process in the self-explanation model to obtain the correlation features between the decision step and the confrontation round; A linear fitting unit, which is used to input each correlation feature obtained during the training process into the linear regression unit for linear fitting to obtain the regression coefficients of each device decision action, which are used to represent the influence degree of each device decision action on each constraint condition and the final reward, that is, to complete the explanation of the training environment, training process and decision actions.

[0039] In a third aspect, the present invention provides a terminal, including a processor and a storage medium; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method described in any one of the first aspects of the present invention.

[0040] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the first aspects of the present invention are implemented.

[0041] The beneficial effects of the present invention are as follows. Compared with the prior art, (1) By establishing a power grid simulation environment, formulating environmental constraint conditions related to various power outage plans, introducing a multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value, and calculating the reward value of the agent's actions according to the designed multi-objective reward function and feeding it back to the agent, the optimal decision of the power outage plan is finally obtained through continuous training and iteration, ensuring the real-time performance and collaboration of the agent's decision-making. Through the mutual collaboration and confrontation between each agent, the multi-objective power outage plan scheduling problem that takes into account ensuring safety, supply, and consumption is realized. By constructing multi-agents to conduct reinforcement learning training for different objectives, the decision-making results are more efficient and accurate.

[0042] (2) Based on safety as the basic premise, when deciding the collaborative action strategies of each agent, the highest weight is given to the decision-making actions of the agent for the safety guarantee goal. Through the solution of the main objective agent taking responsibility, the final power outage plan scheduling decision can focus on the realization of the safety goal, improving the safety and reliability of power grid operation.

[0043] (3) The present invention establishes a power grid simulation environment through distributed accelerated power flow calculation to support the rapid and effective training of reinforcement learning agents, so that the agents can make the best optimized scheduling plan for power outages, minimizing the scope and time of equipment power outages and improving the reliability and economy of power grid operation.

[0044] (4) The present invention explains the relationship between the constraint conditions and power outage actions in the power outage plan through the decision tree framework in the interpretable reinforcement learning algorithm, uses the decision tree algorithm to establish interpretable decision rules for the power outage plan, and constructs an interpretable model. Starting from the interpretability of the training environment, the interpretability of decision-making actions, and the interpretability of the training process, the process of multi-agent action decision-making in the comprehensive optimization of the power outage plan can be clearly explained, ensuring the transparency of the agent's decision-making, which is beneficial for users to understand the behavior actions of the agent. At the same time, the interpretable model can help users analyze the decision-making basis of the agent in a specific environmental state, so as to facilitate the discovery of unreasonable decisions and timely adjust the strategy, making the optimization scheduling of the power grid power outage plan more stable and with higher interpretability. Description of the Drawings

[0045] Figure 1 It is a schematic flowchart of the power outage plan scheduling method based on multi-agent interpretable reinforcement learning in the present invention; Figure 2 It is a schematic diagram of the interpretable model in the present invention; Figure 3 It is a schematic flowchart of the distributed power flow calculation in the present invention; Figure 4 It is an experimental data graph of the relationship between the comprehensive reward value of the agent and the number of iterations during the agent training process in the present invention; Figure 5 It is a structural principle block diagram of the power outage plan scheduling system based on multi-agent interpretable reinforcement learning in the present invention. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in the present invention are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0047] Embodiment 1: Referring to Figure 1 , the embodiment of the present invention provides a power outage plan scheduling method based on multi-agent interpretable reinforcement learning, which specifically includes the following steps: S1. According to the topological structure and operating parameters of the target power grid, establish a power grid simulation environment as the interaction environment for training multi-agents; It should be noted that during the training process of making power outage plan scheduling decisions, the power grid simulation environment needs to implement the power flow calculation and solution function based on the power grid operation data. According to different requirements, a centralized or distributed power flow calculation method can be selected to support the rapid and effective training of the reinforcement learning agent.

[0048] In a preferred but non-limiting embodiment, referring to Figure 3 as shown, the power grid simulation environment of the present invention selects the distributed power flow calculation method. When the simulation environment conducts training interactions, it modifies the corresponding parameters on the BPA file based on the current agent cooperation action strategy, and executes the distributed power flow calculation based on the modified parameters, so as to simulate the power grid operation, so as to support the agent to use historical power grid data for iterative training during the reinforcement learning process. Among them, the distributed power flow calculation specifically includes: S1.2: Based on the power grid operation BPA file, read the power grid operation data and establish a power grid model; S1.3: Prune the overall power grid model according to the supply area relationship based on the topological structure and sensitivity analysis to form multiple small power grid models; S1.4: Deploy the pruned power grid models of each supply area to each slave node of a distributed cluster containing multiple computing nodes; S1.5: Parallelly calculate the boundary power flows of the power grid models of each supply area in the slave nodes, and transmit the boundary power flow results calculated by each slave node to the master node of the cluster; S1.6: The master node checks the boundary power flow results of each supply area, reduces the transmission consumption of the cluster data volume, integrates the power flow data of each supply area, and finally forms the power flow calculation result of the overall power grid model.

[0049] In this embodiment, by performing parallel power flow calculations on the pruned partition models, the data processing speed of the simulation is greatly accelerated.

[0050] S2: Construct the state space and action space of the intelligent agent according to the optional schemes of the current scheduling plan; Specifically, the state space consists of the environmental state and the self-state where the intelligent agent is currently located. Constructing the state space of the intelligent agent includes: integrating the state information of the power flow of power grid equipment and environmental parameters (such as power grid load, output of each unit, power-off equipment and time obtained from the initial demand of the power-off plan, weather data, major event information affecting the power grid, etc.) into the state space; for the initial demand of the power-off plan, since it belongs to unstructured data, the UIE framework is also required to identify and extract the power-off equipment entities; at the same time, the state space is updated in real time through the data bus to ensure that multiple intelligent agents can make decisions based on the latest power-off plan data at each time step.

[0051] Constructing the action space of the intelligent agent includes: defining the set of actions that the intelligent agent can take, including the specific time of equipment power-off, on the premise of meeting the physical constraints of the power grid system and the requirements of power-off operations. Among them, the actions that can be taken can be single-step actions (such as turning on / off a certain device) or composite actions (such as collaboratively allocating multi-device tasks). Then construct the collaborative action space of multiple intelligent agents so that when an intelligent agent makes an online decision, it considers the states and potential actions of other intelligent agents. For example, the specific form can be expressed as follows: , where each parameter in the parentheses represents the power-off time of the AC line, bus, transformer, switch, and unit respectively.

[0052] S3: Design the reward function of each corresponding intelligent agent based on three optimization goals of ensuring safety, ensuring supply, and ensuring consumption; The specific processes of designing the reward functions for the goals of ensuring safety, ensuring supply, and ensuring consumption required by the reinforcement learning intelligent agent for the optimal scheduling of the power-off plan are as follows: (1) Goal of ensuring safety: Considering the line current load rate and the bus voltage load rate, the higher the load rate, the greater the safety risk and the worse the security. Therefore, multiplying the current load rate and the voltage load rate by a negative weight coefficient can obtain the safety reward value. And since it is already under the constraints of power flow and voltage, and at the same time to achieve data normalization, the maximum load rate is limited to 1. The reward function is as follows:

[0053] In the formula, is the reward function based on the goal of ensuring safety, and respectively represent the coefficient weights of the current load term and the voltage load rate term, n and m respectively represent the total number of AC lines and the total number of buses in the power grid under the current strategy, represents the time Line Current of, Represents the line Current limit of, Represents the time Bus Voltage of, Represents the bus Voltage limit of Is a set minimum value to avoid a zero denominator situation.

[0054] (2)Goal of ensuring supply: a: Considering the impact of power outages on users comprehensively, it is necessary to calculate based on the past power outage frequency to minimize the impact of power outages on users. The quantification value of the impact of power outage frequency on users is calculated by the following formula:

[0055] In the formula, Is the power outage frequency of equipment (Calculated year by year); Is the number of users affected by the power outage of equipment ; Is the total number of users; Is the set of power outage equipment; This value represents the impact of the power outage plan on users. The greater the impact on users, the worse the supply guarantee effect on the side. It can be negated as the reward value of the supply guarantee goal.

[0056] b: The quantification value of the impact of power supply on users is calculated by the following formula To quantify the reliability of power supply:

[0057] In the formula, is the total number of power grid nodes in this area, i is the index, is the power outage duration of users in this area. The average node power outage duration is used to quantify the power supply situation. Taking the inverse of this value represents the power supply reliability caused by the power outage plan and can be used as the reward value for the power supply guarantee target.

[0058] c: The implementation of the maintenance power outage plan often leads to power outages of users, resulting in a decrease in power supply reliability; to reduce the number of user power outages caused by equipment maintenance, reduce the power outage electricity, and improve the power supply efficiency, the user maintenance power outage electricity should be minimized:

[0059] In the formula, is the user number, is the total number of users, and are the total power outage power and total power outage electricity after the implementation of the th user power outage plan; is the power outage time; the user maintenance power outage electricity is used to describe the supply target. The lower the user maintenance power outage electricity, the higher the power supply reliability; taking the inverse of this value is used as the reward function for power supply guarantee. The higher the reward value, the better the power supply guarantee effect.

[0060] d: Considering the influence of the power supply section, the power supply section safety reward function , the higher the degree of over-limit of the power supply section, the lower the reward value:

[0061] In the formula, is the total number of current nodes, l is the index, power supply section, is a constant of the minimum value to avoid the situation of the denominator being zero.

[0062] Finally, comprehensively considering the impact of power outages on users, quantifying the reliability of power supply, user maintenance power outage electricity, and power supply section, the final reward function for the power supply guarantee target is obtained:

[0063] In the formula, is the reward function based on the power supply guarantee target, respectively represent the quantified values of the influence degree of power outage frequency on users , the quantified value of the influence degree of power supply on users , user maintenance power outage electricity and the weight coefficients of the power supply section safety reward value.

[0064] (3)Power consumption guarantee target: Consuming renewable energy means minimizing the waste of renewable energy and making full use of green energy such as wind energy and solar energy. This can be achieved by encouraging the timely scheduling of renewable energy production capacity to areas or periods with high load demand.

[0065] Considering reducing the power outage of new energy units as one of the power consumption guarantee targets, first, a reward function for new energy units with power outages is added. The larger the available power generation of the new energy units with power outages, the lower the reward value:

[0066] In the formula, represents the output loss of new energy units due to planned power outages; i represents the i th new energy unit with a planned power outage, represents the total number of new energy units with planned power outages.

[0067] Secondly, a safety reward function for power transmission sections is added. The higher the degree of over-limit of the power transmission section, the lower the reward value:

[0068] In the formula, represents the total number of current nodes, power transmission section, is a constant with a minimum value to avoid the denominator being zero.

[0069] Finally, considering the total load, new energy units with power outages, and power transmission sections comprehensively, the reward function for the final power consumption guarantee target is obtained:

[0070] In the formula, is the reward function based on the power consumption guarantee target, and represent the reward values of new energy units with power outages and the safety reward value of the power transmission section weight coefficients respectively.

[0071] S4. Based on the formulated environmental constraints, the state space, action space of the agent, and each reward function, use the multi-agent cooperative adversarial AC reinforcement learning algorithm based on the Shapley value and combine it with the power grid simulation environment to train the action strategies of each agent; The environmental constraints formulated in this embodiment include: line power flow safety constraints, section power flow safety constraints, non-changeable power outage plan constraints, simultaneous power outage constraints, and power outage mutual exclusion constraints; specifically as follows: 1) Line power flow security constraint:

[0072] In the formula, is the security and stability constraint of line l ; is the generator output power transfer distribution factor of the node where unit i is located with respect to line l ; is the active power output of unit i at t time; K is the number of nodes in the system; is the generator output power transfer distribution factor of node k with respect to line l ; is the bus load value of node k at t period; and are the positive and reverse power flow relaxation variables of line l respectively.

[0073] 2) Section power flow security constraint:

[0074] In the formula, and are the power flow constraints of section s respectively; is the generator output power transfer distribution factor of the node where unit i is located with respect to section s ; is the generator output power transfer distribution factor of node k with respect to section s ; and are the forward and reverse power flow relaxation variables of section s respectively.

[0075] 3) Power outage non-changeable plan constraint:

[0076] In the formula, indicates that the power outage object is in the power outage state on the t th day.

[0077] 4) Simultaneous power outage constraint:

[0078] In the formula, and They respectively represent the scheduled execution start time and end time of the power outage plan.

[0079] 5) The power outage mutual exclusion constraint can be expressed as:

[0080] In the formula, T represents the length of the entire power outage plan scheduling period.

[0081] Specifically, the steps of training the action strategies of each agent by using the multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value and combining with the power grid simulation environment in step S4 of this embodiment include: Step 1: Randomly select an environmental state from the current state space; Step 2: Use the three Actor networks updated this time to respectively determine the action strategies that should be taken by their corresponding three agents i 1, i 2 and i 3 in the current environmental state; Step 3: Integrate the action strategies of the current agents according to the set proportional weights to obtain the current collaborative action strategy; In a preferred but non-limiting embodiment, the step of integrating the action strategies of the current agents according to the set proportional weights in step 2 includes: normalizing the decision actions of each device in the action strategies of each agent to the range of (-1, 1), and integrating the execution actions of each device respectively through the following formula:

[0082] In the formula, represents the collaborative action strategy for device z, respectively represent the decision actions for device z in the action strategies of the current agents i 1, i 2 and i 3; respectively represent the proportional weights set for agents i 1, i 2 and i 3, ; where, .

[0083] In this embodiment, the goal of ensuring safety is the primary goal, so the action weight value of the agent that achieves the goal of ensuring safety is greater than the action weights of the goals of ensuring power supply and ensuring consumption. Preferably, in this embodiment, is set to 0.5, and are both 0.25.

[0084] Step 4: Modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraint conditions are met. If not, return to Step 1; if so, update the spatial state based on the current collaborative action strategy and execute Step 5; Furthermore, when performing power grid operation simulation in Step 4 of this embodiment to determine whether the environmental constraint conditions are met, first substitute the collaborative action into the existing non-changeable plan for verification. If the power outage time of the non-changeable plan is not met, it means that this constraint is not met, and return to Step 1; if it is met, then modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power flow calculation to determine whether the result of the power flow calculation meets the power flow constraint conditions. And according to the topological result obtained from the power flow calculation, combined with the power outage plan association relationship, determine whether the plan meets the co-outage constraint and the mutual exclusion constraint. If not, return to Step 1.

[0085] Step 5: Calculate the respective corresponding reward function values of agents i 1, i 2, and i 3 under their current action strategies , , and , and update the corresponding Critic network of each agent in a gradient descent manner. At the same time, use the currently updated Critic networks to respectively fit the Shapley Q values of each agent; In this embodiment, the expressions for respectively fitting the Shapley Q values of each agent are as follows:

[0086] In the formula, represents the Shapley Q value of agent i executing the action strategy S under the environmental state a i ; represents the Q value of agent i executing the action strategy S under the environmental state a i corresponding thereto. The Q value is obtained by substituting the reward value corresponding to the agent into the Bell optimal equation of reinforcement learning and using the Critic network for fitting; the subscript C represents the set of all agents, C / i represents the remaining agents excluding agent i , represents the sum of the Q values corresponding to the remaining agents executing their respective action strategies S under the environmental state ; among them, the principle of calculating is the same as that of The same as above and will not be elaborated here.

[0087] Step 6: Based on the Shapley Q-values of the current agents, update each Actor network using the gradient ascent method, and return to execute Step 1 for iterative training until the set number of iterations is satisfied, and output the final collaborative action strategy.

[0088] A preferred but non-limiting embodiment. The power outage plan scheduling method based on multi-agent interpretable reinforcement learning provided in this embodiment further includes: Explain the relationship between the constraint conditions and power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, record the state, action selection, constraint conditions, and final results when the agent makes a decision, and trace the decision-making process through the log module of the data reading and writing engine to facilitate understanding of the agent's behavior logic; at the same time, use charts and data flow diagrams to represent the real-time state and decision-making process of the agent, so as to display the state changes, action selections, and constraint conditions of the agent to the user.

[0089] Among them, the steps of explaining the relationship between the constraint conditions and power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm include: (1) Use the decision tree algorithm to classify each action data determined during the iterative training of the agent to identify which device decision actions make the collaborative action strategy not meet various constraint conditions. Specifically, according to the decision tree algorithm, each node can be recursively split. By calculating the conditional Gini index of each power outage action during the training process, and then selecting the attribute with the smallest conditional Gini index as the splitting node. The calculation formula of the Gini index is:

[0090] In the formula, D is the power outage action data set; is the power outage action D of the Gini index; k is the number of constraint condition categories; is the probability that the selected device decision action sample belongs to category k ; The decision tree intuitively shows the decision path from the constraint conditions to the power outage actions through its tree structure; that is, each internal node represents 1 decision point, which is divided based on the threshold of a certain constraint condition feature. Each leaf node of the decision tree represents 1 specific action or action category; the path from the root node to the leaf node forms a series of interpretable decision rules; these rules describe how the agent selects power outage actions under different constraint conditions.

[0091] (2) Analyze the importance of the decision-making actions of each device on each constraint condition by constructing an interpretable model; Refer to Figure 2 As shown, the interpretable model includes a recurrent neural network (RNN) encoder, a multi-layer perceptron (MLP) encoder, a self-explanatory model, and a linear regression unit connected in sequence.

[0092] In this embodiment, the steps of analyzing the importance of the decision-making actions of each device on each constraint condition by constructing an interpretable model include: A: Use the environmental state, action strategy, reward function parameters, and constraint condition parameters that the agent confronts during training as the input parameters of the model, and use the RNN encoder to encode the input parameters to capture the decision-making step features of the agent; B: Use the MLP encoder to learn the features that the RNN encoder fails to capture to obtain the overall confrontation features of the agent; C: Establish the correlation between the decision-making step features and the overall confrontation features through the Gaussian process in the self-explanatory model to obtain the association features between the decision-making step and the confrontation round; D: Input the association features obtained each time during the training process into the linear regression unit for linear fitting to obtain the regression coefficients of the decision-making actions of each device, which are used to represent the influence degrees of the decision-making actions of each device on each constraint condition and the final reward, that is, the explanation of the training environment, training process, and decision-making actions is completed.

[0093] In summary, the power outage plan scheduling method based on multi-agent interpretable reinforcement learning provided in this embodiment is based on the reward function designed with the optimization goals of ensuring safety, supply, and consumption. It simultaneously satisfies the line power flow safety constraint, section power flow safety constraint, power outage non-changeable plan constraint, simultaneous power outage constraint, and power outage mutual exclusion constraint, and constructs a multi-agent action decision interpretable model to correctly and clearly explain the results of multi-agent decisions. While satisfying the constraints, the agent aims to minimize the power flow over-limit, power outage frequency, and wind / solar curtailment to obtain the final power outage strategy. Aiming at the problem that the results of multi-agent policy output are difficult to interpret, based on the policy-level interpretation method for deep reinforcement learning agents, taking the state space S and action space A in the agent confrontation process as inputs, and superimposing the output results of the recurrent neural network (RNN) encoder and the multi-layer perceptron (MLP) encoder through the interpretation model with the Gaussian process, and using the linear regression unit to fit the output features of the self-explanatory model, and solving through the linear regression process to obtain the importance and correlation of the decision-making step, so as to analyze the decision-making process of multi-agent actions in the power outage plan and explain the results of its policy output.

[0094] The power outage plan scheduling method based on multi-agent interpretable reinforcement learning provided in this embodiment is applied in detail in the training examples as follows: 1. Based on the BPA file of grid operation data and the original demand file of power outage plan, construct the state space and action space of the power outage plan scheduling reinforcement learning task; 2. The grid simulation environment service program reads the BPA file of grid operation data and the initial demand file of power outage plan, constructs a grid model, and extracts the grid state at this time and sends it to the multi-agent established by using the three target reward functions designed by the present invention; 3. Based on the read grid state respectively, obtain their respective decision-making actions, normalize the actions output by each agent, multiply by the weight of the corresponding target, and then add them up to obtain the comprehensive decision-making action of the multi-agent, and transmit the action to the grid simulation environment service program; 4. Use the line power flow security constraint, section power flow security constraint, power outage non-changeable plan constraint, simultaneous power outage constraint and power outage mutual exclusion constraint formulated by the present invention, that is, in the grid simulation environment, calculate whether the decision-making behavior of the agent meets the relevant constraints according to the formula of each constraint condition, and exclude the actions that do not meet the constraints; 5. Based on the grid power flow security constraint, take the grid state (including generator output state, load state, line power flow value, etc.) and the original demand of the power outage plan as the state space, and the power outage time of the power outage equipment as the action space, and establish an interpretable decision tree; transmit the decision-making action of the agent, as well as the constraint condition parameters and target reward function parameters to the trained interpretable model, and obtain the regression coefficients corresponding to these input values, reflecting the importance of the decision-making action; 6. The grid simulation environment service program modifies the corresponding parameters on the BPA file according to the comprehensive decision-making action, and executes the power flow calculation for grid simulation deduction. When performing the power flow calculation, the grid simulation environment first cuts the overall grid model according to the partition relationship, deploys the cut models of each part in the cluster of multi-computing nodes, and finally the slave nodes only need to transmit the small-scale supply area boundary power flow results of various calculations to the master node, and the master node checks and integrates the boundary power flow results of the supply area to achieve distributed power flow calculation.

[0095] 7. The grid simulation environment then calculates the values of the security protection, supply guarantee and consumption guarantee reward functions according to the deduction results, normalizes them respectively, and through the calculation formula: Calculate the comprehensive reward value, and then feedback it to the agent. Use the critic network to fit the Shapley Q value, and use the gradient ascent method to improve the actor network. After continuous training and iteration, guide it to adjust the parameters and make the best decision.

[0096] In this embodiment, taking the adjustment of the monthly power outage plan in the experiment as an example, each day's decision is regarded as one step in the reinforcement learning process. 30 steps form an episode. Taking the average agent behavior reward value per episode as the ordinate and the number of training episodes as the abscissa, the experimental data of the relationship between the comprehensive reward value of the agent trained by the method proposed in the present invention and the number of iterations are as Figure 4 shown.

[0097] The beneficial effects of the present invention are as follows. Compared with the prior art, (1) By establishing a power grid simulation environment, formulating environmental constraint conditions related to various power outage plans, introducing a multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value, and calculating the reward value of the agent's actions according to the designed multi-objective reward function and feeding it back to the agent, the optimal decision of the power outage plan is finally obtained through continuous training and iteration, ensuring the real-time and collaborative nature of the agent's decision-making. Through the mutual cooperation and confrontation between each agent, the multi-objective power outage plan scheduling problem that takes into account ensuring safety, ensuring power supply, and ensuring power consumption is realized. By constructing multi-agents to conduct reinforcement learning training for different objectives, the decision-making results are more efficient and accurate.

[0098] (2) Based on safety as the basic premise, when deciding the collaborative action strategies of each agent, the highest weight is given to the decision-making actions of the agent for the safety guarantee goal. Through the solution of the main goal agent's backup guarantee, the final power outage plan scheduling decision can focus on the realization of the safety goal, improving the safety and reliability of the power grid operation.

[0099] (3) By establishing a power grid simulation environment through distributed accelerated power flow calculation, it supports the rapid and effective training of the reinforcement learning agent, enabling the agent to make the best power outage plan optimization and scheduling plan, minimizing the power outage scope and time of equipment, and improving the reliability and economy of the power grid operation.

[0100] (4) By explaining the relationship between the constraint conditions and power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, using the decision tree algorithm to establish the interpretable decision rules for the power outage plan, and constructing an interpretable model, starting from the interpretability of the training environment, the interpretability of the decision-making actions, and the interpretability of the training process, the process of the multi-agent action decision-making in the comprehensive optimization of the power outage plan can be clearly explained, ensuring the transparency of the agent's decision-making, which is beneficial for users to understand the agent's behavior actions. At the same time, the interpretable model can help users analyze the decision-making basis of the agent in a specific environmental state, so as to facilitate the discovery of unreasonable decisions and timely adjust the strategy, making the power grid power outage plan optimization and scheduling more stable and highly interpretable.

[0101] Embodiment 2: As Figure 5As shown in the figure, the present invention provides a power outage plan scheduling system based on multi-agent interpretable reinforcement learning. The system is used to implement the steps of the method in the first embodiment above. Specifically, the system includes: An environment establishment module, configured to establish a power grid simulation environment as an interaction environment for training multi-agents according to the topological structure and operation parameters of the target power grid; A space construction module, configured to construct a state space and an action space of the agent according to the optional solutions of the current scheduling plan; A reward function design module, configured to design the reward functions of the corresponding agents based on three optimization objectives of ensuring safety, ensuring power supply, and ensuring consumption; An agent training module, configured to adopt a multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value and combine with the power grid simulation environment to train the action strategies of each agent based on the formulated environmental constraints, the state space, the action space, and each reward function of the agent; A collaborative action decision-making module, configured to integrate the action strategies of the trained agents according to the set proportional weights, and make a decision on the final collaborative action strategy, so as to generate an optimal power outage plan scheduling scheme.

[0102] Specifically, when the power grid simulation environment in the agent training module trains the action strategies of each agent: The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current agent collaborative action strategy, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic result after the power grid state changes, so as to complete the simulation of the power grid operation, so as to support the iterative training using historical power grid data during the agent reinforcement learning process.

[0103] In this embodiment, the distributed power flow calculation includes: cutting the overall power grid model according to the supply area relationship based on the topological structure and sensitivity analysis; respectively deploying the cut power grid models of each supply area on each slave node of a distributed cluster including multiple computing nodes; parallel calculating the boundary power flow of the power grid models of each supply area in each slave node, and transmitting the boundary power flow results calculated by each slave node to the master node of the cluster; the master node checks the boundary power flow results of each supply area, and integrates the power flow data of each supply area to finally form the power flow calculation result of the overall power grid model.

[0104] The state space construction module constructs the state space of the agent, including integrating the state information of the power flow of grid equipment, grid load, the output of each unit, and the initial demand of the power outage plan into the state space. At the same time, the state space is updated in real time through the data bus to ensure that multi-agents can make decisions based on the latest power outage plan data at each time step. The action space construction of the agent includes defining the action space that the agent can take, including the specific time of equipment power outage, on the premise of meeting the physical constraints of the power grid system and the requirements of power outage operations.

[0105] In this embodiment, the environmental constraint conditions include: line power flow safety constraints, section power flow safety constraints, power outage non-change plan constraints, simultaneous power outage constraints, power outage mutual exclusion constraints, etc.

[0106] In the reward function design module, the expressions of the reward functions for each corresponding agent are designed based on three optimization goals: ensuring safety, ensuring supply, and ensuring consumption, as follows:

[0107] In the formula, is the reward function based on the goal of ensuring safety, and respectively represent the coefficient weights of the current load item and the voltage load rate item, n and m respectively represent the total number of AC lines and the total number of busbars in the power grid under the current strategy, represents the time The current of line , represents line The current limit of the line, represents the time The voltage of bus , represents bus The voltage limit of the bus is a set minimum value; is the reward function based on the goal of ensuring supply, respectively represent the quantization value of the impact of the power outage frequency on users , the quantization value of the impact of power supply on users , the power consumption of user maintenance power outage and the safety reward value of the power supply section The weight coefficient; is the reward function based on the goal of ensuring consumption, and respectively represent the reward value of new energy shutdown units and the safety reward value of the power transmission section The weight coefficient.

[0108] As an embodiment of the present invention, the agent training module includes: A state selection unit, configured to randomly select an environmental state from the current spatial state; An action decision unit, configured to use the three Actor networks updated in the current iteration to determine the action strategies to be taken by the corresponding three agents i 1, i 2, and i 3 in the current environmental state; An action integration unit, configured to integrate the action strategies of the current agents according to the set proportional weights to obtain the current collaborative action strategy; A simulation judgment unit, configured to modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to determine whether the environmental constraint conditions are met. If not, the state selection unit is instructed to re-select an environmental state; if so, the current collaborative action strategy is updated and the spatial state is fed back to the reward calculation unit; A reward calculation unit, configured to calculate the respective corresponding reward function values of agents i 1, i 2, and i 3 under their current action strategies , , and , and update the Critic networks corresponding to the respective agents by means of gradient descent, and at the same time use the currently updated Critic networks to respectively fit the Shapley Q values of the respective agents; An update iteration unit, configured to update the respective Actor networks by means of gradient ascent based on the current Shapley Q values of the respective agents, and feed back to the state selection unit for iterative training until the set number of iterations is met, and output the final collaborative action strategy.

[0109] Further, the action integration unit includes: A normalization subunit, configured to normalize the decision actions on each device in the action strategies of the respective agents to the range of (-1, 1); An integration subunit, configured to integrate the execution actions on each device respectively through the following formula:

[0110] In the formula, represents the collaborative action strategy for device z, respectively represent the decision actions on device z in the action strategies of the current agents i 1, i 2, and i 3; respectively represent for agent i1. i The proportional weights set for 2 and i 3, wherein, .

[0111] Furthermore, in the reward calculation unit, the expressions for fitting the Shapley Q-values of each agent are as follows:

[0112] In the formula, represents the Shapley Q-value of agent i executing the action policy S under the environmental state a i ; represents the Q-value corresponding to agent i executing the action policy S under the environmental state a i The Q-value is obtained by substituting the reward value corresponding to the agent into the Bell optimal equation of reinforcement learning and fitting it using the Critic network; the subscript C represents the set of all agents, C / i represents the remaining agents after excluding agent i , represents the sum of the Q-values corresponding to each of the remaining agents executing their respective action policies S under the environmental state .

[0113] A preferred but non-limiting embodiment, the system further includes: An explanation module for explaining the relationship between the constraint conditions and the power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, recording the state, action selection, constraint conditions, and final result of the agent when making a decision, and tracing the decision-making process through the log module of the data reading and writing engine, so as to facilitate understanding the behavior logic of the agent; at the same time, using charts and data flow diagrams to represent the real-time state and decision-making process of the agent, so as to display the state changes, action selections, and constraint conditions of the agent to the user.

[0114] In this embodiment, the explanation module includes: A decision classification unit for classifying each action data decided during the iterative training process of the agent using the decision tree algorithm to identify which device decision actions make the collaborative action policy not satisfy various constraint conditions; A construction unit for analyzing the importance of each device decision action affecting each constraint condition by constructing an interpretable model; Among them, the interpretable model includes a recurrent neural network (RNN) encoder, a multi-layer perceptron (MLP) encoder, an interpretable model, and a linear regression unit connected in sequence.

[0115] Specifically, the construction unit includes: An encoding subunit, configured to use the RNN encoder to encode the input parameters such as the environmental state, action policy, reward function parameters, and constraint condition parameters that the agent confronts during the training process, so as to capture the decision step features of the agent; A learning subunit, configured to use the MLP encoder to learn the features that the RNN encoder fails to capture, so as to obtain the overall confrontation features of the agent; An interpretable subunit, configured to establish the correlation between the decision step features and the overall confrontation features through the Gaussian process in the interpretable model, and obtain the correlation features between the decision step and the confrontation round; A linear fitting unit, configured to input each of the obtained correlation features during each training process into the linear regression unit for linear fitting, and obtain the regression coefficients of each device's decision action, which are used to represent the influence degree of each device's decision action on each constraint condition and the final reward, that is, to complete the explanation of the training environment, training process, and decision action.

[0116] Furthermore, in this embodiment, by regulating the data service bus of the cloud platform, the system state, action space, and constraint conditions are dynamically obtained and parsed, and a power outage plan scheduling system based on multi-agent interpretable reinforcement learning is constructed. Based on the data reading and writing engine of the regulated cloud platform, multi-agents can obtain the power outage plan-related information in real time, ensuring the accuracy and timeliness of decision-making. At the same time, the explanation module introduces an interpretable reinforcement learning analysis method to increase the transparency of decision-making, facilitating the power grid dispatcher to understand the decision-making behavior of multi-agents; the collaborative action decision module introduces a multi-agent action collaboration and confrontation method to achieve optimal decision-making based on multiple objectives. At the same time, the system proposed in the present invention can also support the decision-making training of single agents to adapt to the reinforcement learning algorithms with different numbers of agents.

[0117] In practical applications, the specific process of constructing a power outage plan scheduling system based on multi-agent interpretable reinforcement learning in this embodiment is as follows: (1)Define constraint conditions. The power flow security constraints, immutable plan constraints, co-outage constraints, and mutual exclusion constraints defined in the interpretable reinforcement learning analysis method are used as training constraint conditions; these constraints will be transmitted by the data service bus as the boundary conditions for action selection. This enables multi-agent online decision-making to conform to the business logic of the power grid outage plan, and at the same time, it is easier for users and decision-makers to understand the decision-making behavior of the agents. The data service bus is used to receive changing data in real time and update the parameters of the constraint conditions, enabling the agents to adapt to changing requirements in different states of the power grid regulation system and thus make more accurate decisions.

[0118] (2)Establish an interpretability enhancement mechanism. Based on the decision tree framework in the interpretable reinforcement learning analysis method, explain the relationship between the constraint conditions and outage actions in the outage plan, record the state, action selection, constraint conditions, and final results of each agent when making decisions, and trace the decision-making process through the log module of the data reading and writing engine to facilitate understanding of the agent's behavior logic. On the basis of the decision-making data, add an explanation module, either rule-based explanation or attention mechanism-based explanation, to enable the system to explain why each agent chooses a certain action or avoids a certain constraint. Use charts and data flow diagrams to represent the real-time state and decision-making process of the agents, making the system operation transparent and controllable, and thus presenting the state changes, action selections, and constraint conditions of the agents to the users.

[0119] (3)Introduce a multi-agent collaborative adversarial AC algorithm based on the Shapley value, enabling each agent to collaborate and compete with each other, aiming to ensure security - ensure supply - ensure consumption, and taking ensuring security as the primary goal, to determine its own actions and obtain the optimal strategy for outage plan scheduling.

[0120] (4)Construct a data reading and writing engine. The data reading and writing engine encapsulates functions such as data reading, parsing, storage, and update, accesses the state, actions, and constraint data of multi-agents through the API interface of the data bus, and conveniently reads the current state and constraints and selects corresponding actions. In addition, to support parallel reading and writing of multi-agents, an efficient concurrency control mechanism is designed, adopting a distributed lock and a timestamp-based access control strategy to ensure the security and consistency of data access and avoid data competition or conflicts among multiple agents for the same resource.

[0121] (5) Construct a power grid simulation environment. Based on the BPA file of power grid operation, extract and model the power grid data modified by the agent's decision, and further trim the power grid model using the method based on topological structure and sensitivity analysis. For each part of the trimmed power grid model, use the Newton-Raphson iterative solution method to perform power flow calculation. The architecture proposed in this embodiment will adopt a distributed computing method to perform parallel power flow calculation on the trimmed model to accelerate the processing speed. At the same time, to ensure flexibility, a traditional centralized power flow calculation method can also be adopted. The distributed power flow calculation process is as Figure 3 shown: (6) According to the designed multi-objective reward function, calculate the reward value of the agent's action and feedback it to the agent. The algorithm finally obtains the optimal decision through continuous training and iteration.

[0122] Through the above process, a power outage plan scheduling system based on multi-agent interpretable reinforcement learning is constructed, which ensures the real-time, collaborative and transparent decision-making of the agent, making the optimization and scheduling of the power grid outage plan more stable and more interpretable.

[0123] The system proposed in the present invention can also be compatible with single-agent algorithms according to the actual application situation. The constructed power grid simulation environment can select the power flow solution method of centralized traditional calculation or distributed parallel calculation according to the user's needs, support the rapid and effective training of the reinforcement learning agent, and the system architecture is flexible and applicable to various power grid simulation scenarios.

[0124] The power outage plan scheduling system based on multi-agent interpretable reinforcement learning provided in the embodiment of the present invention and the power outage plan scheduling method based on multi-agent interpretable reinforcement learning provided in Embodiment 1 are based on the same technical concept, and can produce the beneficial effects described in Embodiment 1. The content not described in detail in this embodiment can be referred to in Embodiment 1.

[0125] Embodiment 3: A terminal provided in an embodiment of the present invention includes a processor and a storage medium, and this terminal is an embedded computer system device. The storage medium of the terminal is used to store instructions. The memory includes a non-volatile storage medium and an internal memory; among them, the non-volatile storage medium stores an operating system, a computer program and a database, and the internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium, and the database is used to store instruction data. The processor of the terminal is used to operate according to the instructions provided by the storage medium to execute the steps of the power outage plan scheduling method based on multi-agent interpretable reinforcement learning according to any one of the embodiments in Embodiment 1.

[0126] Embodiment 4: A computer-readable storage medium provided by an embodiment of the present invention stores a computer program thereon, and when the program is executed by a processor, the steps of the method described in any one of Embodiment 1 are implemented.

[0127] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present disclosure.

[0128] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0129] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0130] Computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A power outage plan scheduling method based on multi-agent interpretable reinforcement learning, characterized in that The method includes: Based on the topological structure and operation parameters of the target power grid, establish a power grid simulation environment as the interaction environment for training multi-agent; Construct the state space and action space of the agent according to the current alternative scheduling plans; Design the reward functions of the corresponding agents based on three optimization objectives of ensuring security, ensuring power supply, and ensuring consumption; Based on the formulated environmental constraints, the state space and action space of the agent, and each reward function, adopt the multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value and combine with the power grid simulation environment to train the action strategies of each agent; Integrate the action strategies of the trained agents according to the set proportional weights, and make a decision on the final collaborative action strategy, so as to generate the optimal power outage plan scheduling scheme.

2. The outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 1, wherein When the power grid simulation environment trains the action strategies of each agent: The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current collaborative action strategy of the agent, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic results after the power grid state changes, so as to complete the simulation of the power grid operation, so as to support the iterative training using historical power grid data in the agent reinforcement learning process.

3. The outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 2, wherein, The said distributed power flow calculation includes: Based on the topological structure and sensitivity analysis, cut the overall power grid model according to the supply area relationship; Deploy the cut power grid models of each supply area to each slave node of the distributed cluster containing multiple computing nodes; Parallelly calculate the boundary power flows of the power grid models of each supply area in each slave node, and transmit the boundary power flow results calculated by each slave node to the master node of the cluster; The master node checks the boundary power flow results of each supply area, and integrates the power flow data of each supply area to finally form the power flow calculation results of the overall power grid model.

4. The power outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 1, wherein Constructing the state space of the agent includes: integrating the state information of the power flow of power grid equipment, the power grid load, the output of each unit, and the initial demand of the power outage plan into the state space; at the same time, updating the state space in real time through the data bus to ensure that the multi-agent can make decisions according to the latest power outage plan data at each time step.

5. The power outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that Constructing the action space of the agent includes: defining the action space that the agent can take on the premise of meeting the physical constraints of the power grid system and the requirements of power outage operations, including the specific time of equipment power outage.

6. The outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that The said environmental constraints include: line power flow security constraints, section power flow security constraints, power outage non-changeable plan constraints, simultaneous power outage constraints, and / or power outage mutual exclusion constraints.

7. The power outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that The expressions of the reward functions of the corresponding agents designed based on the three optimization objectives of ensuring security, ensuring power supply, and ensuring consumption are as follows: Wherein, is the reward function based on the security guarantee objective, and respectively represent the coefficient weights of the current load item and the voltage load rate item, n and m respectively represent the total number of AC lines and the total number of buses in the power grid under the current strategy, represents the time the current of line , represents the current limit of line , represents the voltage of bus at time , represents the voltage limit of bus ; is a set minimum value; is the reward function based on the supply guarantee objective, respectively represent the quantization values of the impact degree of power outage frequency on users , the quantization value of the impact degree of power supply on users , the power consumption of user maintenance power outage and the safety reward value of the power supply section 's weight coefficients; is the reward function based on the consumption guarantee objective, and respectively represent the reward value of new energy shutdown units and the safety reward value of the power transmission section 's weight coefficients.

8. The outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 7, characterized in that, The steps of adopting the multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value and combining with the power grid simulation environment to train the action strategies of each agent include: Step 1: Randomly select an environmental state from the current space state; Step 2: Use the three Actor networks updated in the current iteration to determine the action strategies that should be taken by their corresponding three agents i 1, i 2, and i 3 in the current environmental state; Step 3: Integrate the action strategies of the current agents according to the set proportional weights to obtain the current collaborative action strategy; Step 4: Modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and conduct power grid operation simulation to determine whether the environmental constraint conditions are met. If not, return to Step 1; if so, update the spatial state based on the current collaborative action strategy and execute Step 5; Step 5: Calculate the agent i 1, i 2, and i 3 under their current action policies for the corresponding reward function values , , and , and update the Critic network corresponding to each agent in a gradient descent manner. At the same time, use the currently updated Critic networks to fit the Shapley Q-values of each agent respectively; Step 6: Based on the current Shapley Q-values of each agent, update each Actor network using the gradient ascent method, and return to execute Step 1 for iterative training until the set number of iterations is met, and output the final collaborative action strategy.

9. The power outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 8, wherein The step of integrating the action strategies of the current agents according to the set proportional weights includes: Normalize the decision actions of each device in the action strategies of each agent to the range of (-1, 1), and integrate the execution actions of each device respectively through the following formula: In the formula, represents the collaborative action strategy for device z, respectively represent the decision-making actions for device z in the action strategies of the current agents i 1, i 2, and i 3; respectively represent the proportional weights set for agents i 1, i 2, and i 3, ; where .

10. The outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 8, wherein The expressions for fitting the Shapley Q-values of each agent are as follows: In the formula, represents the agent i executing the action policy S under the environmental state a i of the Shapley Q value; represents the agent i executing the action policy S under the environmental state a i corresponding Q value, and the Q value is obtained by substituting the reward value corresponding to the agent into the Bell optimal equation of reinforcement learning and fitting it using the Critic network; the subscript C represents the set of all agents, C / i represents excluding the agent i after which the remaining agents, represents the sum of the Q values corresponding to each of the remaining agents executing their respective action policies S under the environmental state i .

11. The outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 1, characterized in that The method further includes: Based on the decision tree framework in the interpretable reinforcement learning algorithm, explain the relationship between the constraint conditions and the power outage actions in the power outage plan, record the state, action selection, constraint conditions, and final results when the agent makes a decision, and trace the decision-making process through the log module of the data reading and writing engine to facilitate understanding of the agent's behavior logic; at the same time, use charts and data flow diagrams to represent the real-time state and decision-making process of the agent, so as to display the state changes, action selections, and constraint conditions of the agent to the user.

12. The outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 11, characterized in that The step of explaining the relationship between the constraint conditions and the power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm includes: Use the decision tree algorithm to classify each action data decided during the iterative training of the agent to identify the device decision actions that make the collaborative action strategy not meet various constraint conditions; Analyze the importance of the decision actions of each device on each constraint condition by constructing an interpretable model; Among them, the interpretable model includes a recurrent neural network RNN encoder, a multi-layer perceptron network MLP encoder, a self-explanatory model, and a linear regression unit connected in sequence.

13. The power outage plan scheduling method based on multi-agent interpretable reinforcement learning according to claim 12, wherein The step of analyzing the importance of the decision actions of each device on each constraint condition by constructing an interpretable model includes: Use the environmental state, action strategy, reward function parameters, and constraint condition parameters that the agent confronts during the training process as the input parameters of the model, and use the RNN encoder to encode the input parameters to capture the decision step features of the agent; Use the MLP encoder to learn the features that the RNN encoder fails to capture to obtain the overall confrontation features of the agent; Establish the correlation between the decision step features and the overall confrontation features through the Gaussian process in the self-explanatory model to obtain the correlation features between the decision step and the confrontation round; Input each correlation feature obtained each time during the training process into the linear regression unit for linear fitting to obtain the regression coefficients of the decision actions of each device, which are used to represent the influence degrees of the decision actions of each device on each constraint condition and the final reward, that is, complete the explanation of the training environment, training process, and decision actions.

14. A power outage plan scheduling system based on multi-agent interpretable reinforcement learning, which runs the power outage plan scheduling method based on multi-agent interpretable reinforcement learning according to any one of claims 1-13, characterized in that, The system includes: An environment establishment module, which is used to establish a power grid simulation environment as an interaction environment for training multi-agent according to the topological structure and operation parameters of the target power grid; A space construction module, which is used to construct the state space and action space of the agent according to the current alternative plans of the scheduling plan; A reward function design module, which is used to design the reward functions of the corresponding agents based on three optimization objectives of ensuring safety, ensuring power supply, and ensuring consumption; An agent training module, which is used to adopt the multi-agent collaborative adversarial AC reinforcement learning algorithm based on the Shapley value and combine the power grid simulation environment to train the action strategies of each agent based on the formulated environmental constraints, the state space, action space of the agent, and each reward function; A collaborative action decision-making module, which is used to integrate the action strategies of the trained agents according to the set proportional weights, and decide the final collaborative action strategy, so as to generate the optimal blackout plan scheduling scheme.

15. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 14, characterized in that, When the power grid simulation environment in the agent training module trains the action strategies of each agent: The power grid simulation environment dynamically modifies the corresponding parameters on the original power grid operation data BPA file based on the current collaborative action strategy of the agent, and then performs distributed power flow calculation based on the modified parameters to obtain the simulation deduction dynamic results after the power grid state changes, so as to complete the simulation of the power grid operation, so as to support the iterative training of the agent reinforcement learning process using historical power grid data.

16. The power outage plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 15, characterized in that, The distributed power flow calculation includes: Cutting the overall power grid model according to the supply area relationship based on the topological structure and sensitivity analysis; Deploying the cut power grid models of each supply area on each slave node of a distributed cluster containing multiple computing nodes; Parallelly calculating the boundary power flow of the power grid models of each supply area in each slave node, and transmitting the boundary power flow results calculated by each slave node to the master node of the cluster; The master node checks the boundary power flow results of each supply area, integrates the power flow data of each supply area, and finally forms the power flow calculation results of the overall power grid model.

17. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 14, characterized in that In the space construction module, constructing the state space of the agent includes: integrating the state information of the power flow of power grid equipment, the power grid load, the output of each unit, and the initial demand of the blackout plan into the state space; at the same time, the state space is updated in real time through the data bus to ensure that the multi-agent can make decisions according to the latest blackout plan data at each time step.

18. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 14, characterized in that, In the space construction module, constructing the action space of the agent includes: defining the action space that the agent can take on the premise of meeting the physical constraints of the power grid system and the requirements of blackout operations, including the specific time of equipment blackout.

19. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 14, characterized in that, In the agent training module, the environmental constraints include: line power flow safety constraints, section power flow safety constraints, blackout non-changeable plan constraints, simultaneous blackout constraints, and / or blackout mutual exclusion constraints.

20. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 14, characterized in that, In the reward function design module, the expressions of the reward functions of the corresponding agents designed based on the three optimization objectives of ensuring safety, ensuring power supply, and ensuring consumption are as follows: In the formula, is the reward function based on the security guarantee objective, and respectively represent the coefficient weights of the current load item and the voltage load rate item, n and m respectively represent the total number of AC lines and the total number of buses in the power grid under the current strategy, represents the time the current of line , represents the current limit of line , represents the time the voltage of bus , represents the voltage limit of bus ; is the set minimum value; is the reward function based on the supply guarantee objective, respectively represent the quantization value of the impact degree of power outage frequency on users , the quantization value of the impact degree of power supply on users , the power consumption of user maintenance power outage and the security reward value of the power supply section 's weight coefficient; is the reward function based on the accommodation guarantee objective, and respectively represent the reward value of new energy shutdown units and the security reward value of the power transmission section 's weight coefficient.

21. The power outage plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 20, characterized in that, The agent training module includes: A state selection unit, which is used to randomly select an environmental state from the current space state; An action decision-making unit for using the three Actor networks updated in the current iteration to determine the action strategies to be taken by their corresponding three agents respectively i 1、 i 2 and i 3 in the current environmental state An action integration unit, which is used to integrate the action strategies of current agents according to the set proportional weights to obtain the current collaborative action strategy; A simulation judgment unit, which is used to modify the corresponding parameters on the BPA file in the power grid simulation environment according to the current collaborative action strategy, and perform power grid operation simulation to judge whether the environmental constraint conditions are met. If not, the state selection unit is required to re-select an environmental state; if so, the updated space state of the current collaborative action strategy is fed back to the reward calculation unit; A reward calculation unit for separately calculating the i 1, i 2, and i 3 corresponding reward function values under their current action policies , , and , and updating the Critic network corresponding to each agent by means of gradient descent, and simultaneously using the currently updated Critic networks to respectively fit the Shapley Q values of each agent; An update and iteration unit, which is used to update each Actor network in a gradient ascent manner based on the current Shapley Q values of each agent, and feed it back to the state selection unit for iterative training until the set number of iterations is met, and output the final collaborative action strategy.

22. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 21, characterized in that, The action integration unit includes: A normalization sub-unit, which is used to normalize the decision-making actions of each device in the action strategies of each agent to the range of (-1, 1); An integration sub-unit, which is used to integrate the execution actions of each device respectively through the following formula: In the formula, represents the collaborative action strategy for device z, respectively represent the current agent i 1, i 2 and i 3's decision-making actions for device z in their action strategies; respectively represent the proportional weights set for agents i 1, i 2 and i 3, ; where .

23. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 21, wherein In the reward calculation unit, the expressions for fitting the Shapley Q values of each agent are as follows: wherein, represents the agent i executing the action policy S under the environmental state a i of the Shapley Q-value; represents the agent i executing the action policy S under the environmental state a i corresponding Q-value, and the Q-value is obtained by substituting the reward value corresponding to the agent into the Bell optimal equation of reinforcement learning and using the Critic network for fitting; the subscript C represents the set of all agents, C / i represents the remaining agents excluding the agent i ; represents the sum of the Q-values corresponding to the remaining agents executing their respective action policies S under the environmental state i .

24. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 14, wherein The system further includes: An explanation module, which is used to explain the relationship between the constraint conditions and the power outage actions in the power outage plan based on the decision tree framework in the interpretable reinforcement learning algorithm, record the state, action selection, constraint conditions and final results of the agent when making decisions, and trace the decision-making process through the log module of the data reading and writing engine, so as to facilitate understanding the behavior logic of the agent; at the same time, use charts and data flow diagrams to represent the real-time state and decision-making process of the agent, so as to display the state changes, action selections and constraint conditions of the agent to the user.

25. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 24, wherein The explanation module includes: A decision classification unit, which is used to classify each action data decided during the iterative training of the agent by using the decision tree algorithm to identify which device decision-making actions make the collaborative action strategy not meet various constraint conditions; A construction unit, which is used to analyze the importance of the decision-making actions of each device affecting each constraint condition by constructing an interpretable model; Among them, the interpretable model includes a recurrent neural network RNN encoder, a multi-layer perceptron network MLP encoder, a self-explanation model, and a linear regression unit connected in sequence.

26. The blackout plan scheduling system based on multi-agent interpretable reinforcement learning according to claim 25, wherein The construction unit includes: An encoding sub-unit, which is used to use the RNN encoder to encode the input parameters of the environmental state, action strategy, reward function parameters and constraint condition parameters that the agent confronts during the training process to capture the decision-making step features of the agent; A learning sub-unit, which uses the MLP encoder to learn the features that the RNN encoder fails to capture to obtain the overall confrontation features of the agent; A self-explanation sub-unit, which is used to establish the correlation between the decision-making step features and the overall confrontation features through the Gaussian process in the self-explanation model to obtain the correlation features between the decision-making step and the confrontation round; A linear fitting unit is configured to input each of the obtained correlation features during each training process into a linear regression unit for linear fitting, so as to obtain regression coefficients of each device decision action, which are used to represent the influence degrees of each device decision action on each constraint condition and the final reward, that is, to complete the interpretation of the training environment, the training process, and the decision action.

27. A terminal, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-13.

28. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1-13 are implemented.

Citation Information

Patent Citations

  • Multi-regional power grid collaborative optimization method, system and device and readable storage medium

    CN115333111A

  • Power grid power-cut plan intelligent arrangement device and method based on deep reinforcement learning

    CN117610869A

  • Multi-agent deep reinforcement learning-based power grid power failure arrangement system and method, and medium

    CN119784018A

  • Fault self-healing multi-strategy collaborative optimization method and system for power distribution network

    CN119940582A

  • Architecture for explainable reinforcement learning

    US20220147876A1

Cited By

  • Power grid maintenance schedule optimization method based on interpretable deep reinforcement learning

    CN121436297A

  • Power grid equipment power-cut plan scheduling method under multi-agent cooperation

    CN121660358A

  • Building operation and maintenance agent strategy updating method and storage medium

    CN121920409A