Multi-Agent Reinforcement Learning Rolling Scheduling Method, Device, Equipment and Storage Medium
By constructing a distributed intraday rolling scheduling algorithm for multi-agent reinforcement learning, the problems of slow speed and difficulty of training in grid scheduling are solved, efficient scheduling of a high proportion of new energy grids is achieved, and the accuracy and speed of scheduling are improved.
Patent Information
- Application Number
- CN202210828178.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-07-13
AI Technical Summary
The single-agent reinforcement learning modeling and solving in the field of power grid scheduling in the prior art is slow, the training is difficult, and it is difficult to effectively cope with the complexity and uncertainty of a high proportion of new energy grids.
Build a rolling scheduling model for a high-proportion new energy power system, use the decentralized parts of the multi-agent, and model the Marcolf decision-making process, obtain the attention network that improves the regional feature aggregation graph, and build a training architecture of the distributed intraday rolling scheduling algorithm through the multi-agent reinforcement learning algorithm, including the design of optimization objective functions and constraints.
The accuracy and efficiency of multi-agent reinforcement learning rolling scheduling are improved, and the optimization of multi-agent joint strategy can be achieved through coordinated training, which is in line with the actual application scenarios of power grid scheduling, and the modeling solution speed and simplicity of the training process are improved.
Smart Images

Figure CN115310775B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence power intersection, and particularly to a multi-agent reinforcement learning rolling scheduling method, device, equipment and storage medium. Background Art
[0002] Achieving a high proportion of new energy access and consumption plays a very important role in alleviating energy shortages and optimizing the energy structure; however, the increase in the access proportion of new energy units also brings great challenges to the safe operation of the power system; the high-proportion new energy power grid has the characteristics of complex operation scenarios and many emergencies. Relying solely on traditional dispatching regulations has limited means, and it is difficult to ensure the effectiveness of traditional solutions in the face of a large number of new operation modes; in addition, the energy ratio in the high-proportion new energy power grid is inappropriate, and the dispatching resources lack coordination. Therefore, when there is a large deviation in the output of new energy, the dispatching measures within the system are limited and the adjustment margin is insufficient; these reasons ultimately lead to frequent accidents in the high-proportion new energy system.
[0003] Currently, the academic community's optimization dispatching methods for power systems under high-proportion new energy access can be mainly divided into several categories such as deterministic modeling, stochastic optimization modeling, robust optimization modeling and their combined methods, but all of them have the problem of relying on prior mathematical models and new energy prediction values; reinforcement learning has no prior requirements for mathematical models and takes the maximum expected discounted reward as the optimization goal, which can solve the problems of environmental uncertainty, safety of dispatching schemes and long-term economic optimality that need to be considered in power grid dispatching; the academic community has carried out extensive and in-depth research on reinforcement learning in the field of power system optimization dispatching.
[0004] However, the current research in the field of power grid dispatching mainly focuses on the centralized single-agent reinforcement learning method, which requires the agent to be observable and controllable for the entire system. However, in the actual dispatching scenario, the formulation of decisions often involves the game of multiple agents and multiple regions. Modeling and solving only by a single agent have problems such as large training difficulty and being divorced from the actual application scenario. Summary of the Invention
[0005] The main purpose of the present invention is to provide a multi-agent reinforcement learning rolling scheduling method, device, equipment and storage medium, aiming to solve the technical problems in the prior art of single-agent reinforcement learning in the field of power grid dispatching, such as slow modeling and solving speed, large training difficulty, and being divorced from the actual application scenario.
[0006] In the first aspect, the present invention provides a multi-agent reinforcement learning rolling scheduling method, and the multi-agent reinforcement learning rolling scheduling method includes the following steps:
[0007] Construct a rolling scheduling model for the active power within a day corresponding to a high-proportion new energy power system;
[0008] Model the rolling scheduling model with a multi-agent decentralized partially observable Markov decision process to obtain a multi-agent scheduling architecture;
[0009] Obtain the attention network of the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtain a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. Construct a training architecture for a distributed intraday rolling scheduling algorithm based on multi-agent reinforcement learning according to the attention network and the multi-agent reinforcement learning algorithm.
[0010] Optionally, the construction of the rolling scheduling model for the intraday active power corresponding to a high-proportion new energy power system includes:
[0011] Select the action quantities of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and construct an optimization objective function for the rolling scheduling model of the intraday active power;
[0012] Establish the constraint conditions of the rolling scheduling model, and construct a rolling scheduling model for the intraday active power corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions.
[0013] Optionally, the selection of the action quantities of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and the construction of an optimization objective function for the rolling scheduling model of the intraday active power includes:
[0014] Select the action quantities of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and construct an optimization objective function for the rolling scheduling model of the intraday active power through the following formula:
[0015]
[0016] Where, respectively represent the starting time of the rolling scheduling, the total scheduling duration, and the system collapse time, is the number of adjustable units and energy storage devices in the system, and respectively represent the cost coefficients of the rescheduled unit power generation and rescheduling, the action cost coefficient of the energy storage, the network loss cost, and the system collapse penalty coefficient, respectively represent the power generation and power generation rescheduling volume of the conventional unit, represents the energy storage device i at t the state of charge at the moment.
[0017] Optionally, the establishment of the constraint conditions of the rolling scheduling model, and the construction of a rolling scheduling model for the intraday active power corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions includes:
[0018] The power balance constraint is determined by the following formula:
[0019]
[0020] Wherein, respectively represent the fluctuating power of the energy storage scheduling power, the new energy output and the load demand;
[0021] The upper and lower limits of the adjustable unit output are determined by the following formula:
[0022]
[0023] Wherein, are respectively the i minimum and maximum output powers of the adjustable unit;
[0024] The output ramp constraint of the adjustable unit is determined by the following formula:
[0025]
[0026] Wherein, represents the maximum ramp power of the adjustable unit i per unit time;
[0027] The upper and lower limits of the state of charge of the energy storage device are determined by the following formula:
[0028]
[0029] Wherein, are respectively the i minimum and maximum state of charge, the maximum capacity and the maximum ramp power per unit time of the energy storage device;
[0030] The output ramp constraint per unit time of the energy storage device is determined by the following formula:
[0031]
[0032] Wherein, are respectively the i minimum and maximum state of charge of the energy storage device;
[0033] The head and tail constraints of the state of charge of the energy storage device are determined by the following formula:
[0034]
[0035] Wherein, represents the maximum step size of the active rolling scheduling, represents the allowable deviation of the initial state of charge and the state of charge at the last moment;
[0036] Take the power balance constraint, the upper and lower limits of the adjustable unit output constraint, the adjustable unit output ramp constraint, the upper and lower limits of the energy storage device state of charge constraint, the energy storage device unit time output ramp constraint, and the energy storage device state of charge start and end constraint as the constraint conditions of the rolling dispatch model;
[0037] Construct a rolling dispatch model for the active power corresponding to the high-proportion new energy power system according to the optimization objective function and the constraint conditions.
[0038] Optionally, perform multi-agent decentralized partially observable Markov decision process modeling on the rolling dispatch model to obtain a multi-agent dispatch architecture, including:
[0039] Perform modeling on the agent, state, observation value, action, and reward parts of the rolling dispatch model;
[0040] Design the reward mechanism of the reinforcement learning agent and derive the state equation in the intraday rolling dispatch scenario to obtain a multi-agent dispatch architecture.
[0041] Optionally, the design of the reward mechanism of the reinforcement learning agent and the derivation of the state equation in the intraday rolling dispatch scenario to obtain a multi-agent dispatch architecture includes:
[0042] Obtain the action costs of the overloaded lines and heavily loaded lines, controlled units, and energy storage devices in the intraday rolling dispatch scenario, and obtain the start and end constraints of the energy storage device SoC;
[0043] Determine the action reward of the multi-agent system at t time according to the action cost and the start and end constraints through the following statement:
[0044]
[0045] where, is the load rate of the transmission line l , defined as the ratio of the current transmission power to the long-term maximum allowable transmission power of the line; is the total number of system transmission lines; are the penalty coefficients for overloaded and heavily loaded lines respectively; are the active power adjustment amounts of the controlled generator and the energy storage device respectively; are the quadratic and primary coefficients corresponding to the generator action cost function. Considering that the generator is not penalized when following the day-ahead plan, the constant term is omitted; is the number of adjustable units in the system; is an indicator function used to count the number of actions of the energy storage device; represent the penalty coefficients for the number of actions of the energy storage and the SoC offset respectively; is the variable weight coefficient of the SoC offset; is the number of system energy storage devices;
[0046] Obtain the immediate reward and long-term expected reward of the multi-agent, derive the Bellman equation for reinforcement learning of the multi-agent, and obtain a multi-agent scheduling architecture.
[0047] Optionally, obtain the attention network for the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtain a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. Construct a training architecture for a distributed intraday rolling scheduling algorithm based on multi-agent reinforcement learning according to the attention network and the multi-agent reinforcement learning algorithm, including:
[0048] Construct a power grid graph attention network layer suitable for the power grid graph structure of the multi-agent, and construct an improved graph attention network layer based on regional feature information aggregation;
[0049] Construct a rolling scheduling agent time series decision-making module that can extract power grid time series decision-making feature information, and construct a value mixing network that coordinates joint policy training among multi-agents;
[0050] Construct a training architecture for a distributed intraday rolling scheduling algorithm of multi-agent reinforcement learning suitable for power grid rolling scheduling scenarios according to the power grid graph attention network layer, the improved graph attention network layer, the rolling scheduling agent time series decision-making module, and the value mixing network, and perform training iteration on the model parameters in the training architecture until a specified number of times is reached.
[0051] In a second aspect, to achieve the above object, the present invention also proposes a multi-agent reinforcement learning rolling scheduling device, and the multi-agent reinforcement learning rolling scheduling device includes:
[0052] A construction module, configured to construct a rolling scheduling model for the active power corresponding to the intraday of a high-proportion new energy power system;
[0053] A modeling module, configured to perform multi-agent decentralized partially observable Markov decision process modeling on the rolling scheduling model to obtain a multi-agent scheduling architecture;
[0054] A training architecture module, configured to obtain the attention network for the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtain a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. Construct a training architecture for a distributed intraday rolling scheduling algorithm based on multi-agent reinforcement learning according to the attention network and the multi-agent reinforcement learning algorithm.
[0055] Third aspect, to achieve the above object, the present invention further provides a multi-agent reinforcement learning rolling scheduling device, and the multi-agent reinforcement learning rolling scheduling device includes: a memory, a processor, and a multi-agent reinforcement learning rolling scheduling program stored on the memory and executable on the processor, and the multi-agent reinforcement learning rolling scheduling program is configured to implement the steps of the multi-agent reinforcement learning rolling scheduling method as described above.
[0056] Fourth aspect, to achieve the above object, the present invention further provides a storage medium, on which a multi-agent reinforcement learning rolling scheduling program is stored, and when the multi-agent reinforcement learning rolling scheduling program is executed by a processor, it implements the steps of the multi-agent reinforcement learning rolling scheduling method as described above.
[0057] The multi-agent reinforcement learning rolling scheduling method proposed by the present invention constructs a rolling scheduling model for the active power within a day corresponding to a high-proportion new energy power system; performs decentralized partially observable Markov decision process modeling of multiple agents on the rolling scheduling model to obtain a multi-agent scheduling architecture; obtains an attention network for the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtains a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. According to the attention network and the multi-agent reinforcement learning algorithm, a training architecture for a distributed intra-day rolling scheduling algorithm based on multi-agent reinforcement learning is constructed. The modeling and solution speed is fast, the training process is simple, it conforms to the actual application scenario of power grid scheduling, improves the accuracy of multi-agent reinforcement learning rolling scheduling, and can optimize the joint strategy of multiple agents through coordinated training, improving the speed and efficiency of multi-agent reinforcement learning rolling scheduling. Description of the Drawings
[0058] Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment solution of the present invention;
[0059] Figure 2 It is a schematic flowchart of the first embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention;
[0060] Figure 3 It is a schematic flowchart of the second embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention;
[0061] Figure 4 It is a schematic flowchart of the third embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention;
[0062] Figure 5 It is a schematic flowchart of the fourth embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention;
[0063] Figure 6It is a schematic diagram of the overall architecture of the RGAT-QMIX algorithm in the multi-agent reinforcement learning rolling scheduling method of the present invention;
[0064] Figure 7 It is a functional module diagram of the first embodiment of the multi-agent reinforcement learning rolling scheduling device of the present invention.
[0065] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments
[0066] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0067] The solution of the embodiment of the present invention is mainly: by constructing a rolling scheduling model for the active power within a day corresponding to a high-proportion new energy power system; performing multi-agent decentralized partially observable Markov decision process modeling on the rolling scheduling model to obtain a multi-agent scheduling architecture; obtaining the attention network of the improved regional feature aggregation graph of the multi-agent scheduling architecture, and obtaining a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation, and constructing a training architecture of a distributed intra-day rolling scheduling algorithm based on multi-agent reinforcement learning according to the attention network and the multi-agent reinforcement learning algorithm, with fast modeling and solving speed and simple training process, meeting the actual application scenario of power grid scheduling, improving the accuracy of multi-agent reinforcement learning rolling scheduling, being able to optimize the multi-agent joint strategy through coordinated training, enhancing the speed and efficiency of multi-agent reinforcement learning rolling scheduling, and solving the technical problems of single-agent reinforcement learning in the field of power grid scheduling in the prior art, such as slow modeling and solving speed, large training difficulty, and being divorced from the actual application scenario.
[0068] Refer to Figure 1 , Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the present invention.
[0069] Such as Figure 1As shown in the figure, the device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a stable memory (Non-Volatile Memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0070] Those skilled in the art can understand that Figure 1 the device structure shown in the figure does not constitute a limitation on the device, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0071] As Figure 1 shown, the memory 1005, as a storage medium, may include an operating device, a network communication module, a user interface module, and a multi-agent reinforcement learning rolling scheduling program.
[0072] The device of the present invention calls the multi-agent reinforcement learning rolling scheduling program stored in the memory 1005 through the processor 1001 and executes the operations in the embodiments of the multi-agent reinforcement learning rolling scheduling method described below.
[0073] Based on the above hardware structure, embodiments of the multi-agent reinforcement learning rolling scheduling method of the present invention are proposed.
[0074] Referring to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention.
[0075] In the first embodiment, the multi-agent reinforcement learning rolling scheduling method includes the following steps:
[0076] Step S10: Construct a rolling scheduling model for the active power within a day corresponding to a high-proportion new energy power system.
[0077] It should be noted that by modeling the power system under high-proportion new energy access, the optimal scheduling of the power system can be realized. First, a rolling scheduling model for the active power within a day corresponding to a high-proportion new energy power system can be constructed.
[0078] Step S20: Model the rolling scheduling model through a decentralized partially observable Markov decision process for multi-agent to obtain a multi-agent scheduling architecture.
[0079] It can be understood that modeling the decision-making process of the rolling scheduling model specifically means modeling through a decentralized partially observable Markov decision process for multi-agent, and a multi-agent scheduling architecture can be obtained.
[0080] Step S30: Obtain the attention network of the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtain a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. Based on the attention network and the multi-agent reinforcement learning algorithm, construct a training architecture for a distributed intra-day rolling scheduling algorithm based on multi-agent reinforcement learning.
[0081] It should be understood that after obtaining the multi-agent scheduling architecture, an attention network for the improved region feature aggregation graph suitable for the multi-agent scheduling architecture can be designed, and a multi-agent reinforcement learning algorithm RGAT-QMIX that supports spatio-temporal multi-dimensional feature aggregation can be designed. Furthermore, a training architecture for a distributed intra-day rolling scheduling algorithm based on multi-agent reinforcement learning can be constructed.
[0082] Through the above solutions in this embodiment, by constructing a rolling scheduling model for the intra-day active power corresponding to a high-proportion new energy power system; modeling the rolling scheduling model through a decentralized partially observable Markov decision process for multi-agent to obtain a multi-agent scheduling architecture; obtaining the attention network of the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtaining a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. Based on the attention network and the multi-agent reinforcement learning algorithm, construct a training architecture for a distributed intra-day rolling scheduling algorithm based on multi-agent reinforcement learning, the modeling and solution speed is fast, the training process is simple, it conforms to the actual application scenario of power grid scheduling, improves the accuracy of multi-agent reinforcement learning rolling scheduling, can realize the optimization of the multi-agent joint strategy through coordinated training, and improves the speed and efficiency of multi-agent reinforcement learning rolling scheduling.
[0083] Furthermore, Figure 3 is a schematic flowchart of the second embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention. As Figure 3 shown, the second embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention is proposed based on the first embodiment. In this embodiment, the step S10 specifically includes the following steps:
[0084] Step S11: Select the action amounts of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and construct an optimization objective function for the rolling scheduling model of the intra-day active power.
[0085] It should be noted that there are partially adjustable units with strong adjustable capabilities in the multi-agent system, and the action amount of the energy storage device is used as a decision variable to participate in the rolling dispatch, and an optimization objective function of the rolling dispatch model for intra-day active power can be constructed.
[0086] Further, the step S11 specifically includes the following steps:
[0087] Select the action amounts of some adjustable units and energy storage devices in the system as decision variables to participate in the rolling dispatch, and construct an optimization objective function of the rolling dispatch model for intra-day active power through the following formula:
[0088]
[0089] Where, respectively represent the starting time of the rolling dispatch, the total dispatch duration, and the system collapse time, is the number of adjustable units and energy storage devices in the system, and respectively represent the generation cost coefficient and rescheduling cost coefficient of the rescheduled unit, the action cost coefficient of the energy storage, the network loss cost, and the system collapse penalty coefficient, respectively represent the power generation and power generation rescheduling volume of the conventional unit, represents the energy storage device i at t the state of charge at the moment.
[0090] It should be understood that the intra-day active power rolling dispatch strategy of the power system aims to use the day-ahead active power dispatch plan as the power reference point, and according to the correction of short-term hourly new energy unit output and load demand data, conduct comprehensive economic and safety rescheduling for some adjustable units and flexible dispatchable resources in the system such as energy storage or flexible loads. When the renewable energy penetration rate is low, the impact of unit and load fluctuations on the overall power supply and demand balance of the system is small, and the uncertainty can be handled only through spinning reserve. The actual output of the unit will not deviate too much from the day-ahead dispatch plan, and the prediction error terms of renewable units and loads are usually ignored when calculating the supply and demand balance in the dispatch model.
[0091] Step S12: Establish the constraint conditions of the rolling dispatch model, and construct a rolling dispatch model for intra-day active power corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions.
[0092] It can be understood that after establishing the constraint conditions of the intra-day active power rolling dispatch model, a rolling dispatch model for intra-day active power corresponding to a high-proportion new energy power system can be constructed according to the optimization objective function and the constraint conditions.
[0093] Further, the step S12 specifically includes the following steps:
[0094] Determine the power balance constraint by the following formula:
[0095]
[0096] where respectively represent the fluctuating power of the energy storage dispatch power, new energy output, and load demand;
[0097] Determine the upper and lower limits of the adjustable unit output by the following formula:
[0098]
[0099] where are respectively the i minimum and maximum output powers of the adjustable unit;
[0100] Determine the ramping constraint of the adjustable unit output by the following formula:
[0101]
[0102] where represents the maximum ramping power of the adjustable unit i per unit time;
[0103] Determine the upper and lower limits of the state of charge of the energy storage device by the following formula:
[0104]
[0105] where are respectively the i minimum and maximum state of charge, maximum capacity, and maximum ramping power per unit time of the energy storage device;
[0106] Determine the ramping constraint of the energy storage device output per unit time by the following formula:
[0107]
[0108] where are respectively the i minimum and maximum state of charge of the energy storage device;
[0109] Determine the head and tail constraints of the state of charge of the energy storage device by the following formula:
[0110]
[0111] where represents the maximum step size of the active rolling dispatch, and represents the allowable deviation of the initial state of charge and the state of charge at the last moment;
[0112] Take the power balance constraint, the upper and lower limits of the adjustable unit output, the ramp constraint of the adjustable unit output, the upper and lower limits of the state of charge of the energy storage device, the ramp constraint of the energy storage device output per unit time, and the head and tail constraints of the state of charge of the energy storage device as the constraint conditions of the rolling dispatch model;
[0113] Construct a rolling dispatch model for the active power within a day corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions.
[0114] It can be understood that the constraint conditions of the intra-day rolling dispatch model can be established through the above formula, that is, taking the power balance constraint, the upper and lower limits of the adjustable unit output, the ramp constraint of the adjustable unit output, the upper and lower limits of the state of charge of the energy storage device, the ramp constraint of the energy storage device output per unit time, and the head and tail constraints of the state of charge of the energy storage device as the constraint conditions of the rolling dispatch model. Furthermore, a rolling dispatch model for the active power within a day corresponding to a high-proportion new energy power system can be constructed according to the optimization objective function and the constraint conditions.
[0115] It should be understood that the power balance constraint for intra-day rolling dispatch, the upper and lower limits of the adjustable unit output and the ramp constraint per unit time, and the upper and lower limits of the state of charge of the energy storage device and the ramp constraint of the output per unit time are as shown in the above formula. In addition, considering that continuous decisions need to be made for a period of time in the active power rolling dispatch, the initial SoC and the final SoC of the energy storage system should be kept within a certain difference range in each dispatch cycle.
[0116] Through the above solution in this embodiment, by selecting the action quantities of some adjustable units and energy storage devices in the system as the decision variables participating in the rolling dispatch, an optimization objective function for the intra-day active power rolling dispatch model is constructed; the constraint conditions of the rolling dispatch model are established, and a rolling dispatch model for the active power within a day corresponding to a high-proportion new energy power system is constructed according to the optimization objective function and the constraint conditions, which can achieve a fast modeling and solving speed, a simple training process, conform to the actual application scenario of power grid dispatch, and improve the accuracy of multi-agent reinforcement learning rolling dispatch.
[0117] Furthermore, Figure 4 is a schematic flow chart of the third embodiment of the multi-agent reinforcement learning rolling dispatch method of the present invention. As Figure 4 shown, the third embodiment of the multi-agent reinforcement learning rolling dispatch method of the present invention is proposed based on the first embodiment. In this embodiment, the step S20 specifically includes the following steps:
[0118] Step S21: Model the agent, state, observation value, action, and reward parts of the rolling dispatch model.
[0119] It should be noted that the decentralized partially observable Markov decision process modeling applicable to the multi-agent reinforcement learning framework mainly includes agents, states, observations, actions, and rewards.
[0120] In the specific implementation, the decentralized Markov decision process of the multi-agent system can be described in the form of a cooperative stochastic game, that is, the main components of this game include: Agent i: The agent used to be deployed on the controlled node to participate in the decision-making of the intraday rolling optimization scheduling of the power system, where represents the set of agents; State : All indicators reflecting the system operation state at time t, including the actual output values of conventional units, wind farms, and photovoltaic power plants, as well as the predicted output values at time t + 1, the load demands of each node and the load prediction values at time t + 1, and the voltage and power angle of each node in the system, the active and reactive power transmission power of the line, and the line loss, etc.; , represents the feature space. Observation : Considering that the information that Agent i can obtain in the actual environment will be restricted by the physical communication system and data privacy, the observation of Agent i is set to only include the information of this node and its K-hop neighbor nodes, as well as the operation state information of the included transmission lines; Action : The decision-making action of Agent i at time t, which is divided into the rescheduling action of the controlled generator and the charge and discharge action of the energy storage according to the type of the agent control node; In this invention, considering the discrete action space, after normalizing the action spaces of the generator and the energy storage to [-1, 1] according to their single-time ramp-up capabilities, they are equally divided into 21 actions based on a granularity of 0.1. Therefore, assuming that the system contains n agents, the joint action space of the entire power grid will contain 21n scheduling actions; , represents the action space of Agent i, and the joint action ; The action space scale that a single-agent algorithm applicable to the discrete action space can handle is basically between dozens and hundreds. Therefore, when n is large, the single-agent algorithm cannot be trained; Reward : The reward value feedback to the multi-agent system by the environment after accepting the joint action ; Considering that the reward needs to represent the operation state of the power system, it includes three parts: the transmission line load rate, the action cost of the controlled unit, and the action cost of the energy storage device. The reward mechanism is an important part of the reinforcement learning framework, directly affecting the policy goal and convergence performance of the agent, and involving professional domain knowledge.
[0121] Step S22, Design the reward mechanism of the reinforcement learning agent and derive the state equation in the intraday rolling scheduling scenario to obtain a multi-agent scheduling architecture.
[0122] It is understandable that by designing the reward mechanism and deriving the state equation of the reinforcement learning agent in the intraday rolling scheduling scenario, a multi-agent scheduling architecture can be obtained, which can achieve the coordination of the joint scheduling strategies of multiple agents in a large-scale power grid; in the multi-agent reinforcement learning framework, each agent follows the basic learning paradigm of reinforcement learning and formulates strategies based on the decentralized partially observable Markov decision process model.
[0123] Further, the step S22 specifically includes the following steps:
[0124] Obtain the action costs of overloaded lines, heavily loaded lines, controlled units and energy storage devices in the intraday rolling scheduling scenario, and obtain the head and tail constraints of the SoC of the energy storage device;
[0125] According to the action costs and the head and tail constraints, determine the action reward of the multi-agent system at t time through the following description:
[0126]
[0127] Among them, is the load rate of the transmission line l , defined as the ratio of the current transmission power to the long-term maximum allowable transmission power of the line; is the total number of system transmission lines; are the penalty coefficients for overloaded and heavily loaded lines respectively; are the active regulation amounts of the controlled generator and the energy storage device respectively; are the quadratic and linear coefficients corresponding to the generator action cost function. Considering that the generator is not penalized when following the day-ahead plan, the constant term is omitted; is the number of adjustable units in the system; is an indicator function used to count the number of actions of the energy storage device; represent the penalty coefficients for the number of actions and the SoC offset of the energy storage respectively; is the variable weight coefficient of the SoC offset; is the number of system energy storage devices;
[0128] Obtain the immediate reward and long-term expected reward of the multi-agent, and derive the Bellman equation of the multi-agent for reinforcement learning to obtain the multi-agent scheduling architecture.
[0129] It should be understood that the reward mechanism simultaneously considers the overloaded lines and heavily loaded lines in the system, the action costs of the controlled units and the energy storage devices, and the head and tail constraint problems of the SoC of the energy storage device; the action reward of the multi-agent system at time t can be expressed by the above formula.
[0130] In specific implementation, considering any moment t, the uncertainty of the output of renewable energy units may lead to imbalance between system supply and demand or overload of transmission line power flow. The purpose of the reward mechanism is to enable the agent to balance system supply and demand based on the existing operating state and prediction data, control the line load rate within a reasonable range, and ensure the safe operation of the system. Therefore, the reward mechanism penalizes the overloaded lines and heavily loaded lines in the system respectively, and also considers the action costs of controlled units and energy storage devices. The action cost of the generator is considered as a quadratic function of the rescheduling regulation power quantity, and the action cost of the energy storage is linearly related to the number of actions. In addition, it is necessary to ensure that the SoC at the end of the scenario does not deviate too much from the initial SoC for the energy storage device. Therefore, the control performance of the multi-agent system at time t can be expressed by the above formula.
[0131] is the variable weight coefficient of the SoC offset. Considering that the head and tail constraints of SoC only restrict its SoC level at the end of the scenario to be as close as possible to that at the initial moment, so it is very small at the start of the scenario, but will increase significantly in the last few time steps to guide the agent to adjust the SoC level of the energy storage as much as possible at the last moment. According to the requirements, piecewise functions, exponential functions, and sigmoid functions can be used. In this example, the sigmoid function is adopted and described as follows:
[0132]
[0133] When the power system is operating normally, the multi-agent system will obtain corresponding rewards according to its scheduling actions. When the multi-agent outputs wrong actions and causes the system to collapse, a large penalty value will be given. The multi-agent system will conduct a fully cooperative stochastic game under the above reward mechanism based on their respective observation results to reach a mixed-strategy Nash equilibrium.
[0134] When the agent i at t the immediate reward and the long-term expected reward k up to can be expressed by the following formula, where , and are the reward weights for heavily loaded and overloaded lines, rescheduling of controlled generators, and actions of energy storage devices respectively.
[0135]
[0136]
[0137] Among them, is the penalty constant given when the system collapses, The reward discount factor.
[0138] In a specific implementation, the derivation of the Bellman equation for multi-agent reinforcement learning can be the following steps:
[0139] To evaluate the long-term reward of the agent's policy, the value function, i.e., the long-term expected discounted reward, is introduced into the decentralized partially observable Markov decision process. k The state at time The agent i The value function can be expressed as the following formula:
[0140]
[0141] Among them, is the agent i Interacts with the environment k The trajectory obtained after steps, expressed as ; Represents the starting state of the trajectory; Represents the agent i 's policy; Represents at state After executing the action The probability that the environment transfers to state is
[0142] Assume that the rolling optimization scheduling policy of the entire system is expressed as the following formula:
[0143]
[0144] Then the optimal mixed Nash equilibrium solution can be expressed as the following formula:
[0145]
[0146] Among them, Represents the optimal joint policy, Represents the agent i 's optimal policy; Represents the agent i Any policy other than the optimal policy ; Represents the number of agents;
[0147] The value function of each agent is related to its policy, so The term can be expressed as The overall objective of the entire decentralized partially observable Markov decision process can be expressed as making the joint policy of the entire multi-agent system converge to the optimal while the policies of each agent within it converge to their optimal states. Any change in the policy of any one agent will lead to the deterioration of the joint policy. Therefore, the result of this stochastic game can be expressed as a mixed Nash equilibrium, that is, the final result of the game is that no agent will gain unilaterally by adjusting its policy.
[0148] State-action pair value function Can be expressed as the following formula:
[0149]
[0150] Value function Is the state-action pair value function Regarding the action The expected value of, through the above formula, the Bellman equation of the state-action pair can be expressed as the following formula:
[0151]
[0152] It can be understood that the optimal policy of an agent can be found by maximizing the state-action pair function of the agent Considering that the distributed rolling optimization scheduling problem is a completely cooperative game model, all agents cooperate with each other through local information and output reasonable scheduling actions to maintain the safe and economic operation of the system. Therefore, its value function should have the same monotonicity. The mixed Nash equilibrium of the multi-agent system game can be reached after all agents maximize their respective state-action value functions To solve the game problem of the multi-agent system, model-based schemes have been proposed in previous literature to find the mixed Nash equilibrium of the stochastic game. However, the performance of such methods mainly depends on the accuracy of the model itself, and the power grid with a high proportion of new energy access itself has a high degree of nonlinearity and uncertainty, making it difficult to establish an accurate model using mathematical methods. Therefore, this example adopts a model-free multi-agent reinforcement learning algorithm based on value decomposition to search for the Nash equilibrium strategy of the multi-agent game.
[0153] In this embodiment, through the above scheme, by switching the main and auxiliary fuel tank switch of the fuel quantity sensor to the main fuel tank of the target vehicle, reading the remaining fuel quantity of the main fuel tank, the main fuel tank fuel quantity data is obtained; by switching the main and auxiliary fuel tank switch of the fuel quantity sensor to the auxiliary fuel tank of the target vehicle, reading the remaining fuel quantity of the auxiliary fuel tank of the target vehicle, the auxiliary fuel tank fuel quantity data is obtained; it can quickly and accurately obtain the remaining fuel quantity of the fuel tank specifically, improving the accuracy of the multi-agent reinforcement learning rolling scheduling.
[0154] Furthermore, Figure 5This is a schematic flowchart of the fourth embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention. As Figure 5 shown, based on the first embodiment, the fourth embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention is proposed. In this embodiment, step S30 specifically includes the following steps:
[0155] Step S31: Construct a power grid graph attention network layer suitable for the power grid graph structure of the multi-agent, and construct an improved graph attention network layer based on regional feature information aggregation.
[0156] It should be noted that after constructing the power grid graph attention network layer suitable for the power grid graph structure of the multi-agent, an improved graph attention network layer based on regional feature information aggregation can also be constructed.
[0157] Construct a graph attention network layer suitable for the power grid graph structure.
[0158] Let the power grid graph structure feature be represented by the following formula:
[0159]
[0160] where F is the number of features of each node; after passing through a single graph attention layer, a new set of node features different from the original feature set is generated as the output of the network layer, that is:
[0161]
[0162] To more fully represent the high-level features of the graph, the graph attention layer uses a learnable linear attention weight matrix with a shared parameter for each node, which can be expressed as , and then, a self-attention mechanism is deployed on each node, thus forming a shared attention mechanism.
[0163] Then, a self-attention mechanism is deployed on each node, thus forming a shared attention mechanism. To avoid discarding all graph structure information, the graph attention network layer uses a masked attention mechanism to only calculate the attention coefficients of the neighborhood of each node; considering that the attention weight matrix needs to be shared among nodes, the attention coefficients of each node need to be normalized using the softmax function, which can be expressed by the following formula:
[0164]
[0165] Construct an improved graph attention network layer based on regional feature information aggregation:
[0166] For a given node, the contribution degree of neighbor node information should decay with the increase of distance. If the contribution degrees of different - order neighbors are the same, it may interfere with the agent's aggregation of key information. Therefore, this example introduces a spatial discount factor that can comprehensively consider the information and contribution degrees of multi - order neighbor nodes , and constructs an improved regional - aware graph attention network to achieve the aggregation of multi - order neighbor information and the automatic allocation of contribution degrees:
[0167]
[0168] where represents the set of i -th order neighbors of node K . represents the spatial discount factor. For a given node i , the contribution degree of its neighbor node j to it is inversely proportional to the length of the shortest distance i , j between ; can be described in various forms. In this paper, the exponential form is adopted, that is . Among them, is a hyper - parameter used to calculate the spatial discount factor, represents the shortest - path hop count from any neighbor node j in the power - grid topology to node i .
[0169] The multi - head attention mechanism is introduced, and the region - aggregated feature finally obtained by the agent can be expressed as the following formula:
[0170]
[0171] where is a non - linear mapping function, usually in the form of softmax or logistic sigmoid. represents the k -th sub - attention head, is the weight parameter of the k -th attention head, K is the number of attention heads.
[0172] Step S32: Construct a rolling - scheduling agent time - series decision - making module that can extract power - grid time - series decision - making feature information, and construct a value - mixing network for coordinating the joint - policy training among multiple agents.
[0173] It should be understood that the rolling scheduling intelligent agent time-series decision-making module DRQN for extracting grid time-series decision-making feature information is constructed by introducing a Gated Recurrent Unit (GRU) network into the Deep Q Network (DQN) network of the reinforcement learning intelligent agent to construct the DRQN module, so as to realize the feature aggregation of the intelligent agent for the information at historical moments. Let the high-dimensional feature information of the grid obtained by the intelligent agent after being extracted by the underlying network be, then the state update formula of the GRU is as follows:
[0174] Where, represents the update gate, represents the reset gate, represent the historical state information and the candidate state information respectively, and finally output the updated state information .
[0175] In the specific implementation, a value mixing network for coordinating the joint policy training among multiple intelligent agents is constructed. The value mixing network needs to be observable to the global information of the system, receive the actions and state-action value functions of each sub-intelligent agent, and output the global state-joint action value function for calculating the model update parameters. Set the joint Q value of the system as the monotonic non-linear mapping of the utility functions of all intelligent agents, as follows:
[0176]
[0177] It can be understood that each intelligent agent in the RGAT-QMIX algorithm has its own utility function. By adding a value mixing network (Mixing Network) layer and combining the global state s , set the joint Q value of the system as the monotonic non-linear mapping of the utility functions of all intelligent agents to ensure that the policies among intelligent agents can be coordinated and optimized through training. Therefore, RGAT-QMIX is more guaranteed for the convergence of the global policy; for the multi-intelligent agent system in the power system scheduling environment, the relationship among each intelligent agent is a completely cooperative game relationship, and the decision-making purpose of each intelligent agent is to maintain the safe and economic operation of the system. Therefore, the power grid scheduling scenario fully satisfies the constraints of the value mixing network on the joint Q value and the sub Q value. The relationship between the total Q value in RGAT-QMIX and the sub Q values output by each intelligent agent is as shown in the above formula.
[0178] Step S33: Construct a training architecture for the distributed intraday rolling scheduling algorithm of multi-agent reinforcement learning applicable to the power grid rolling scheduling scenario based on the power grid graph attention network layer, the improved graph attention network layer, the rolling scheduling agent time series decision-making module, and the value mixing network, and perform training iterations on the model parameters in the training architecture until the specified number of times is reached.
[0179] It can be understood that a training architecture RGAT-QMIX for the distributed intraday rolling scheduling algorithm of multi-agent reinforcement learning applicable to the power grid rolling scheduling scenario is constructed based on the power grid graph attention network layer, the improved graph attention network layer, the rolling scheduling agent time series decision-making module, and the value mixing network.
[0180] For any agent i it maintains a main network for finding the action with the maximum value at the current moment Q and at the same time maintains a target network with the same network structure for evaluating the action value of the agent during the training phase; the main network updates the parameters by minimizing the error function formula:
[0181]
[0182] where y is the target Q value calculated by the target network, p is the probability distribution of the state-action pair. The parameter update amount of the main network can be obtained through the stochastic gradient of the loss function; the target network parameters adopt a soft update method, as shown in the following formula:
[0183]
[0184] where and represent the parameters of the main network and the target network respectively, is the weight coefficient of the soft update, used to control the speed at which the target network approaches the main network.
[0185] The constraint of the above formula can ensure that the Q function of the agent in the RGAT-QMIX network has the same monotonicity, so the maximization of the joint Q function is equal to the maximization of the local Q function of each agent. Therefore, the optimal joint policy is as shown in the following formula:
[0186]
[0187] In the specific implementation, at time t , for any agent i it maintains a main network To find the current time Q The action with the largest value, while maintaining a target network with the same network structure , which is used to evaluate the action value of the intelligent agent during the training phase; the parameter update method of the target network is not through the back propagation of the loss gradient, but is delayed to the direction of the main network update, including hard update and soft update. The hard update copies the parameters of the main network to the target network after a long fixed number of steps, while the soft update takes a sliding average of the parameters of the target network and the parameters of the main network to update the target network after a short number of steps or after each training; compared with the hard update, the soft update can improve the stability of the training process to a certain extent, so this example uses the soft update method.
[0188] Since the parameter values of the target network are delayed from those of the main network, the actions selected by the main network may not be the same as those in the target network. Q This design can effectively avoid the training phase. Q The problem of over-estimation of values; after extracting sample data from the experience replay pool, the main network updates the parameters by minimizing the error function.
[0189] The RGAT-QMIX network loss function is shown as follows:
[0190]
[0191]
[0192] The improved regional feature aggregation graph attention network layer is embedded into the agent's neural network, and the multi-agent joint strategy is trained based on the RGAT-QMIX algorithm framework of value decomposition; considering that the underlying action logic of similar devices is similar, in order to learn a more generalized agent and reduce the number of parameters of the entire model, parameter sharing is adopted for the controlled generator and energy storage device respectively, that is, different generators or energy storage devices share the same network during training; in addition, compared with generators, energy storage devices need to consider the head and tail constraints of SoC, so the input features of the energy storage agent will additionally splice the current time step information.
[0193] Accordingly, Figure 6 Schematic diagram of the overall architecture of the RGAT-QMIX algorithm in the multi-agent reinforcement learning rolling scheduling method of the present invention, as shown in Figure 6 As shown, at time t , Agent i Get the local observation of the node As the input of the MLP layer, the system topology feature matrix Will pass through the graph attention layer to aggregate the agent nodes KInformation of first-order neighbors. The global state mixes the input values in the network to calculate two layers of MLP layers and and accepts the local observation-action pair values output by each agent network to calculate the global state-joint action value function ; where represents the interaction trajectory between the multi-agent system and the environment, represents a single agent i 's interaction sub-trajectory, represents the joint action of multi-agents; and respectively represent the policy of agent i and its local observation-action pair value function; it can be seen that the local observation-action value function is only related to the sub-trajectory obtained by the interaction of agents, and has nothing to do with the global state
[0194] In the above-mentioned distributed rolling scheduling method for power systems under high proportion of new energy access, the specific implementation of constructing the training architecture of the distributed intraday rolling scheduling algorithm based on multi-agent reinforcement learning according to the attention network and the multi-agent reinforcement learning algorithm includes:
[0195] Initialize the neural network parameters and hyperparameters of each agent, and reset the system environment PSE, and return the initial state after the rolling scheduling scenario is reset.
[0196] Each agent makes decisions based on the local observation information it can obtain, and considering the boundary constraints, filters out illegal actions in the action space.
[0197] After all agents output valid actions, input the joint action into the environment to obtain the next state and reward, and repeat this process until the entire scenario is completed or the system crashes.
[0198] After a fixed number of interaction steps, the experience composed of multiple trajectories will be stored in the experience replay pool; after the number of experiences in the experience replay pool reaches the threshold, sample batch samples and calculate the loss to update the model parameters of RGAT-QMIX; the training process will iterate the above operations until the specified number of times is reached.
[0199] In the specific implementation, initialize the network parameters of the edge scheduling agents of the controlled generators and energy storage and the value mixing agent of the dispatching center , initialize the hidden states of the controlled generator and energy storage networks , and initialize the experience replay pool;
[0200] Reset the rolling scheduling environment and obtain the initial environmental state ;
[0201] Each agent, based on the state obtains the local observations of each agent , as well as the topological feature matrix of the system ;
[0202] Each agent checks the set of legal actions at the current time step and outputs a scheduling action from the set of legal actions based on the policy ;
[0203] Concatenate the actions of all agents to obtain a joint action , execute the joint action in the environment and obtain the state at the next moment , action reward and , update the state ;
[0204] Store the trajectory in the experience replay pool;
[0205] After meeting the training start condition, sample random mini-batch samples from the experience replay pool;
[0206] Based on Equation calculate the loss and update the parameters of each main network, and slide-update the parameters of each target network based on Equation ;
[0207] In a specific implementation, the distributed intraday rolling scheduling algorithm includes a neural network for controlling controlled generators and energy storage devices and a hybrid network for calculating the joint value function ; after the algorithm interacts with the environment, trajectory data containing information such as state, action, value, etc. is obtained and stored in the experience buffer through an "insert" operation; in the training phase, data is sampled from the experience buffer through a "sampling" operation for training.
[0208] The value hybrid network is only used in the training phase. After centralized training, the agents will adopt a distributed execution framework during actual testing; since the hybrid network is no longer needed, the multi-agent system does not need to obtain the global state, and each agent only needs to obtain local observations and k the feature information of the -th order neighbors to make local decisions; compared with the centralized method, distributed execution can greatly improve the computational speed of decision-making, and the decision-making method based on local observation information is closer to the actual scenario, because in the actual large power grid scenario, considering communication costs and data privacy, it is often difficult for a single node to obtain the operating state of the entire power grid.
[0209] In this embodiment, through the above solution, a power grid graph attention network layer suitable for the power grid graph structure of the multi-agent is constructed, and an improved graph attention network layer based on regional feature information aggregation is constructed; a rolling scheduling agent time series decision-making module capable of extracting power grid time series decision-making feature information is constructed, and a value mixing network for coordinating joint policy training among multiple agents is constructed; according to the power grid graph attention network layer, the improved graph attention network layer, the rolling scheduling agent time series decision-making module and the value mixing network, a training architecture of a distributed intraday rolling scheduling algorithm for multi-agent reinforcement learning suitable for the power grid rolling scheduling scenario is constructed, and the model parameters in the training architecture are trained iteratively until the specified number of times is reached, which can realize the optimization of the multi-agent joint policy through coordinated training, has a fast modeling and solving speed, a simple training process, conforms to the actual application scenario of power grid scheduling, improves the accuracy of multi-agent reinforcement learning rolling scheduling, and enhances the speed and efficiency of multi-agent reinforcement learning rolling scheduling.
[0210] Correspondingly, the present invention further provides a multi-agent reinforcement learning rolling scheduling device.
[0211] Referring to Figure 7 , Figure 7 is the functional module diagram of the first embodiment of the multi-agent reinforcement learning rolling scheduling device of the present invention.
[0212] In the first embodiment of the multi-agent reinforcement learning rolling scheduling device of the present invention, the multi-agent reinforcement learning rolling scheduling device includes:
[0213] A construction module 10, configured to construct a rolling scheduling model for the active power corresponding to the high-proportion new energy power system within a day.
[0214] A modeling module 20, configured to perform multi-agent decentralized partially observable Markov decision process modeling on the rolling scheduling model to obtain a multi-agent scheduling architecture.
[0215] A training architecture module 30, configured to obtain an attention network of an improved regional feature aggregation graph of the multi-agent scheduling architecture, and obtain a multi-agent reinforcement learning algorithm supporting spatio-temporal multi-dimensional feature aggregation, and construct a training architecture of a distributed intraday rolling scheduling algorithm based on multi-agent reinforcement learning according to the attention network and the multi-agent reinforcement learning algorithm.
[0216] Wherein, the steps implemented by each functional module of the multi-agent reinforcement learning rolling scheduling device can refer to each embodiment of the multi-agent reinforcement learning rolling scheduling method of the present invention, and will not be elaborated here.
[0217] In addition, an embodiment of the present invention further provides a storage medium, on which a multi-agent reinforcement learning rolling scheduling program is stored. When the multi-agent reinforcement learning rolling scheduling program is executed by a processor, it implements the operations in the embodiment of the multi-agent reinforcement learning rolling scheduling method described above.
[0218] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the presence of additional identical elements in the process, method, article or device including such element.
[0219] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0220] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A multi-agent reinforcement learning rolling scheduling method, characterized in that The multi-agent reinforcement learning rolling scheduling method includes: Construct a rolling scheduling model for the intra-day active power corresponding to a high-proportion new energy power system; Perform decentralized partially observable Markov decision process modeling of the rolling scheduling model for multiple agents to obtain a multi-agent scheduling architecture; Obtain the attention network of the improved regional feature aggregation graph of the multi-agent scheduling architecture, and obtain a multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. Construct a training architecture for a distributed intra-day rolling scheduling algorithm based on multi-agent reinforcement learning according to the attention network and the multi-agent reinforcement learning algorithm; Among them, the construction of the rolling scheduling model for the intra-day active power corresponding to a high-proportion new energy power system includes: Select the action amounts of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and construct an optimization objective function for the rolling scheduling model of the intra-day active power; Establish the constraint conditions of the rolling scheduling model, and construct a rolling scheduling model for the intra-day active power corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions; Among them, the selection of the action amounts of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and the construction of the optimization objective function for the rolling scheduling model of the intra-day active power includes: Select the action amounts of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and construct an optimization objective function for the rolling scheduling model of the intra-day active power through the following formula: Among them, respectively represent the starting time of rolling scheduling, the total scheduling duration, and the system crash time, is the number of adjustable units and energy storage devices in the system, and respectively represent the cost coefficients of generation and rescheduling of rescheduled units, the operation cost coefficient of energy storage, the network loss cost, and the system crash penalty coefficient, respectively represent the power generation and power generation rescheduling volume of conventional units, represents the energy storage device i at t the state of charge at the moment.
2. The multi-agent reinforcement learning rolling scheduling method according to claim 1, wherein The establishment of the constraint conditions of the rolling scheduling model, and the construction of a rolling scheduling model for the intra-day active power corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions includes: Determine the power balance constraint through the following formula: Among them, respectively represent the fluctuating power of energy storage dispatching power, new energy output, and load demand; Determine the upper and lower limits of the output of adjustable units through the following formula: Among them, are respectively the minimum and maximum output powers of the adjustable unit i ; Determine the output ramp constraint of adjustable units through the following formula: Among them, represents the adjustable unit i the maximum ramp-up power per unit time; Determine the upper and lower limits of the state of charge of energy storage devices through the following formula: wherein, are respectively the minimum and maximum state of charge, the maximum capacity, and the maximum ramp rate per unit time of the energy storage device; i Determine the output ramp constraint per unit time of energy storage devices through the following formula: Among them, are the minimum and maximum state of charge of the energy storage device i respectively; Determine the head and tail constraints of the state of charge of energy storage devices through the following formula: Among them, represents the maximum step size of the active rolling schedule, represents the allowable deviation between the initial state of charge and the state of charge at the last moment; Take the power balance constraint, the upper and lower limits of the output of adjustable units, the output ramp constraint of adjustable units, the upper and lower limits of the state of charge of energy storage devices, the output ramp constraint per unit time of energy storage devices, and the head and tail constraints of the state of charge of energy storage devices as the constraint conditions of the rolling scheduling model; Construct a rolling scheduling model for the intra-day active power corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions.
3. The multi-agent reinforcement learning rolling scheduling method according to claim 1, characterized in that The modeling of the decentralized partially observable Markov decision process of multiple agents for the rolling scheduling model to obtain a multi-agent scheduling architecture includes: Model the agents, states, observations, actions, and rewards of the rolling scheduling model; Design the reward mechanism and derive the state equation of the reinforcement learning agent in the intra-day rolling scheduling scenario to obtain a multi-agent scheduling architecture.
4. The multi-agent reinforcement learning rolling scheduling method according to claim 3, characterized in that, The design of the reward mechanism and the derivation of the state equation of the reinforcement learning agent in the intra-day rolling scheduling scenario to obtain a multi-agent scheduling architecture includes: Under the intraday rolling scheduling scenario, obtain the action costs of overloaded lines, heavily loaded lines, controlled units, and energy storage devices, and obtain the head and tail constraints of the state of charge (SoC) of the energy storage devices. Determine the action reward of the multi-agent system at t the moment according to the action cost and the head and tail constraints through the following description: Among them, is the load rate of the transmission line l , defined as the ratio of the current transmission power to the long-term maximum allowable transmission power of the line; is the total number of the system's transmission lines; are the penalty coefficients for overloaded and heavily loaded lines respectively; are the active regulation amounts of the controlled generators and energy storage devices respectively; are the quadratic and primary coefficients corresponding to the generator action cost function. Considering that there is no penalty when the generator operates according to the day-ahead plan, the constant term is omitted; is the number of adjustable units in the system; is an indicator function used to count the action times of the energy storage device; represent the penalty coefficients for the action times and SoC offset of the energy storage respectively; is the variable weight coefficient of the SoC offset; is the number of the system's energy storage devices; Obtain the immediate rewards and long-term expected rewards of the multi-agent, perform the derivation of the Bellman equation for the reinforcement learning of the multi-agent, and obtain the multi-agent scheduling architecture.
5. The multi-agent reinforcement learning rolling scheduling method according to claim 1, characterized in that Obtain the attention network for the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtain the multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. According to the attention network and the multi-agent reinforcement learning algorithm, construct a training architecture for the distributed intraday rolling scheduling algorithm based on multi-agent reinforcement learning, including: Construct a power grid graph attention network layer suitable for the power grid graph structure of the multi-agent, and construct an improved graph attention network layer based on the aggregation of regional feature information. Construct a rolling scheduling agent time-series decision-making module that can extract power grid time-series decision feature information, and construct a value mixing network that coordinates the joint policy training among multi-agents. According to the power grid graph attention network layer, the improved graph attention network layer, the rolling scheduling agent time-series decision-making module, and the value mixing network, construct a training architecture for the distributed intraday rolling scheduling algorithm of multi-agent reinforcement learning suitable for the power grid rolling scheduling scenario, and perform training iteration on the model parameters in the training architecture until the specified number of times is reached.
6. A multi-agent reinforcement learning rolling scheduling device, characterized in that, The multi-agent reinforcement learning rolling scheduling device includes: A construction module, configured to construct a rolling scheduling model for the intraday active power corresponding to a high-proportion new energy power system. A modeling module, configured to perform a decentralized partially observable Markov decision process modeling on the rolling scheduling model to obtain a multi-agent scheduling architecture. A training architecture module, configured to obtain the attention network for the improved region feature aggregation graph of the multi-agent scheduling architecture, and obtain the multi-agent reinforcement learning algorithm that supports spatio-temporal multi-dimensional feature aggregation. According to the attention network and the multi-agent reinforcement learning algorithm, construct a training architecture for the distributed intraday rolling scheduling algorithm based on multi-agent reinforcement learning. The construction module is further configured to select the action amounts of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, construct an optimization objective function for the rolling scheduling model of the intraday active power; establish the constraint conditions of the rolling scheduling model, and construct a rolling scheduling model for the intraday active power corresponding to a high-proportion new energy power system according to the optimization objective function and the constraint conditions. The construction module is further configured to select the action amounts of some adjustable units and energy storage devices in the system as decision variables participating in the rolling scheduling, and construct an optimization objective function for the rolling scheduling model of the intraday active power through the following formula: Among them, respectively represent the starting time of rolling scheduling, the total scheduling duration, and the system crash time, is the number of adjustable units and energy storage devices in the system, and respectively represent the cost coefficients of generation and rescheduling of rescheduled units, the operation cost coefficient of energy storage, the network loss cost, and the system crash penalty coefficient, respectively represent the power generation and power generation rescheduling volume of conventional units, represents the energy storage device i at t the state of charge at the moment.
7. A multi-agent reinforcement learning rolling scheduling device, characterized in that, The multi-agent reinforcement learning rolling scheduling device includes: a memory, a processor, and a multi-agent reinforcement learning rolling scheduling program stored on the memory and executable on the processor. The multi-agent reinforcement learning rolling scheduling program is configured to implement the steps of the multi-agent reinforcement learning rolling scheduling method according to any one of claims 1 to 5.
8. A storage medium, characterized in that, The storage medium stores a multi-agent reinforcement learning rolling scheduler, and when the multi-agent reinforcement learning rolling scheduler is executed by a processor, it implements the steps of the multi-agent reinforcement learning rolling scheduling method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Enemy-friend deep deterministic strategy method and system based on reinforcement learning
CN112215364A
Multi-park energy scheduling method and system based on deep reinforcement learning
CN114091879A
Cited By
Emergency risk early warning and co-processing method based on multi-agent reinforcement learning
CN122048054A
An emergency risk early warning and collaborative disposal method based on multi-agent reinforcement learning
CN122048054B