Power grid active optimization scheduling method and system based on deep reinforcement learning multi-agent
Through the multi-agent method based on deep reinforcement learning, the Transformer network is used to capture the long-range dependence between the multi-agents in the power grid, and dynamically adjust the generator output, solving the stability and efficiency problems of traditional scheduling methods in a dynamic environment, and achieving the accuracy and stability of active optimization scheduling of the power grid.
Patent Information
- Application Number
- CN202510508042.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
The traditional active power scheduling method is difficult to flexibly adapt to actual operational needs in an environment where new energy is large-scale access to new energy and the state of power grids continues to change, resulting in the interaction between the actions and strategies between the agents, the accuracy of the return function is reduced, the difficulty of algorithm convergence increases, and the stability is reduced, making it difficult to maintain stable learning and effectively optimize the active grid scheduling in a dynamic non-stationary environment.
The multi-agent method based on deep reinforcement learning is adopted, and the power grid state is encoded using the Transformer network to capture the long-range dependence relationship between the multiple agents, generate the agent's actions through the multi-layer decoding module, and dynamically adjust the generator's output to achieve dynamic scheduling of the power grid.
It improves the accuracy and stability of grid scheduling decisions, reduces power generation costs and transmission losses, ensures the balance between power grid supply and demand, and improves system operation efficiency and safety.
Smart Images

Figure CN120377397A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent power system dispatching, and particularly to a grid active power optimal dispatching method and system based on multi-agent deep reinforcement learning. Background Art
[0002] With the continuous expansion of the scale of the power grid, the structure of the power system becomes more complex, the distributed characteristics are increasingly prominent, and it faces challenges of dynamic uncertainty and collaborative optimization. Traditional active power dispatching methods rely on accurate system models, but in the environment of large-scale access of new energy and continuous change of grid states, such methods are difficult to flexibly adapt to the actual operation requirements and difficult to achieve efficient optimal dispatching. In this context, in recent years, deep reinforcement learning has received extensive attention due to its excellent performance in solving complex optimal control problems. It learns the optimal strategy through the interaction between the agent and the environment, and can achieve system adaptive dispatching without the premise of accurate modeling. In a multi-agent environment, each agent not only considers its own actions and rewards, but also comprehensively considers the behaviors of other agents. This intricate interaction and connection process makes the environment constantly change dynamically. In a non-stationary environment, the actions and strategies between agents will affect each other, resulting in a decrease in the accuracy of the reward function, greatly increasing the convergence difficulty of the algorithm, reducing the stability, and breaking the balance between exploration and exploitation.
[0003] Therefore, there is an urgent need for an intelligent decision-making method that can maintain stable learning in a dynamic non-stationary environment and effectively optimize the grid active power dispatching, so as to construct a grid active power optimal dispatching decision-making model in a multi-agent environment. Summary of the Invention
[0004] Object of the Invention: To propose a grid active power optimal dispatching method based on multi-agent deep reinforcement learning, and further propose a system, an electronic device and a storage medium that can run the dispatching method, aiming to dynamically adjust the output of each generator under the conditions of intermittency and volatility of the grid multi-agent output, ensure the balance between power supply and demand of the grid; reduce the generation cost and transmission loss, improve the operation efficiency and safety of the system; utilize the global sequence modeling ability of Transformer and the collaborative optimization mechanism of MARL to capture the long-range dependence relationship of the states between grid multi-agents, and improve the accuracy and stability of the dispatching decision-making, so as to effectively solve the above problems existing in the prior art.
[0005] In the first aspect of the present invention, a grid active power optimal dispatching method based on multi-agent deep reinforcement learning is proposed, and the steps are as follows:
[0006] Obtain the key information of the grid operation to form a state space, and use the Transformer network to encode it to obtain the observation sequence of the grid state;
[0007] Input the observation sequence of the power grid state into the encoder to reconstruct the observation encoding of each agent;
[0008] Input the observation encoding of the agent into the decoder to decode and generate the agent action;
[0009] Generate corresponding scheduling decisions based on the agent action. The agent action realizes the dynamic scheduling of the entire power grid by adjusting the output of each generator, so as to meet the load demand.
[0010] In a further embodiment of the first aspect, the key information of the power grid operation includes the active power output P of each generator gen , the current system load demand P load , the current new energy output P renew , the wind power and photovoltaic output Line, and the system transmission loss Loss;
[0011] The above key information forms the state space o i :
[0012] o i ={P gen ,P load ,P renew ,Line,Loss}.
[0013] In a further embodiment of the first aspect, use the Transformer network to encode the state space under multiple time series to obtain the observation sequence of the power grid state where represents the state space under the nth time series.
[0014] In a further embodiment of the first aspect, input the observation sequence of the power grid state into the encoder to reconstruct the observation encoding of each agent. The observation encoding is expressed as
[0015] In a further embodiment of the first aspect, by training the encoder parameter φ, the encoder is made to fit the reward function, and the goal is to minimize the empirical Bellman error. The formula is as follows:
[0016]
[0017] In the formula, is the parameter of the target network; T represents the total number of simulation or training time steps; n represents the current simulation or training round; o t represents the environmental state at time step t; a t represents the agent action; R(o t ,a trepresents the immediate reward function; γ represents the discount factor; represents the environmental state observed by agent m at time step t+1; represents the state value function generated by the encoder at time step t; represents the state value function generated by the target network at time step t+1.
[0018] In a further embodiment of the first aspect, the observed encoding is input to the decoder, and the agent action a is output t = {ΔP1, ΔP2, …, ΔP n};
[0019] where ΔP i represents the output adjustment amount of the i-th generator at the current moment, and n is the total number of generators.
[0020] In a further embodiment of the first aspect, by adjusting the output of each generator, the dynamic scheduling of the entire power grid is realized, so as to meet the load demand;
[0021] The action design satisfies the following conditions:
[0022] where, P i is the output of the current generator i; P i min and P i max are the minimum and maximum outputs of the generator respectively, ensuring that the actual output of the generator is always within the safe operating range.
[0023] In a further embodiment of the first aspect, the encoder first transmits the encoded joint action to the multi-layer decoding module. Each layer of the decoding module adopts the masked self-attention mechanism to ensure that for each agent i j only calculates to the attention between the action heads, where r < j;
[0024] Subsequently, a second masked attention calculation is performed to calculate the attention between the action head and the observed feature sequence;
[0025] Finally, the decoding module ends with a multi-layer perceptron and a skip connection. The output of the last decoding module is the joint action representation, and the joint action representation is fed into the multi-layer perceptron to obtain the probability distribution of the action of agent i m of where $\pi_{\theta}^{i,m}$ represents the policy function of agent $i_m$; $a_{t}^{i,m}$ represents the action executed by agent $i_m$ at time step $t$; $o_{t}^{i,1:m}$ represents the observations of the first to the $m$-th agents at time step $t$; and $a_{t}^{i,1:m - 1}$ represents the actions selected by the first to the $(m - 1)$-th agents at time step $t$.
[0026] In a further embodiment of the first aspect, by training the parameters $\theta$ of the encoder, the goal is to minimize the clipped PPO objective function $L$ Decoder (\theta):
[0027]
[0028] where is the estimate of the joint advantage function; $T$ represents the total number of time steps of simulation or training; $n$ represents the current round of simulation or training; represents the advantage ratio obtained by agent $i$ at time step $t$ m based on the current parameter $\theta$; $\epsilon$ represents a hyperparameter in the PPO algorithm, and $clip(*)$ is a truncation function in reinforcement learning used to limit the magnitude of policy updates.
[0029] In the second aspect of the present invention, a power grid active power optimal scheduling system is proposed. The scheduling system includes four components: a collection module, an encoder, a decoder, and a decision output module. Among them:
[0030] The collection module is used to obtain key information of the power grid operation to form a state space, and encodes it using a Transformer network to obtain an observation sequence of the power grid state;
[0031] The observation sequence of the power grid state is input into the encoder to reconstruct the observation encoding of each agent;
[0032] The observation encoding of the agent is input into the decoder to decode and generate agent actions;
[0033] The decision output module generates corresponding scheduling decisions based on the agent actions, and realizes the dynamic scheduling of the entire power grid by adjusting the output of each generator.
[0034] In the third aspect of the present invention, an electronic device is proposed. The electronic device includes a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the power grid active power optimal scheduling method based on deep reinforcement learning multi-agent as described in the first aspect and its further embodiments.
[0035] In a fourth aspect of the present invention, a computer-readable storage medium is proposed. At least one executable instruction is stored in the storage medium. When the executable instruction runs on an electronic device, the electronic device is caused to execute the grid active power optimal scheduling method based on deep reinforcement learning multi-agent as described in the first aspect and its further embodiments.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] 1. Construct an optimization problem for decision-making dependencies between multi-agent sequences based on Transformer: By using the global sequence modeling ability of Transformer, capture the long-range dependence relationship of the states among the grid multi-agents, and improve the accuracy and stability of scheduling decisions.
[0038] 2. Efficiently implement parallel computing for multi-agent action generation: Significantly improve the training speed and computing efficiency, achieve collaborative optimization of multiple generating units, reduce the power generation cost, transmission loss, and improve the stability of the scheduling system.
[0039] 3. Dynamically adjust the output of each generator to ensure the balance between power grid supply and demand: Under the conditions of intermittency and volatility of the output of the grid multi-agents, dynamically adjust the output of each generator to ensure the balance between power grid supply and demand.
[0040] 4. Reduce the power generation cost and transmission loss, and improve the operation efficiency and safety of the system: Through optimizing the scheduling decision, reduce the power generation cost and transmission loss, and improve the operation efficiency and safety of the system. Description of the Drawings
[0041] Figure 1 It is a schematic diagram of the basic framework of deep reinforcement learning in power grid scheduling in the embodiment.
[0042] Figure 2 It is a schematic diagram of the multi-agent Transformer encoder-decoder architecture in the embodiment.
[0043] Figure 3 It is a diagram of the multi-agent training framework based on the MAPPO algorithm in the embodiment. Detailed Embodiments
[0044] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features known to the public are not described.
[0045] Embodiment 1:
[0046] The applicant's research found that traditional active power scheduling methods are difficult to flexibly adapt to the actual operation requirements in the environment of large-scale access of new energy and continuous changes in the grid state. In a multi-agent environment, the actions and strategies among agents affect each other, resulting in a decrease in the accuracy of the reward function, an increase in the difficulty of algorithm convergence, and a decrease in stability. Existing methods are difficult to maintain stable learning and effectively optimize the active power scheduling of the power grid in a dynamic non-stationary environment.
[0047] Therefore, this embodiment discloses a method for optimizing the active power scheduling of a power grid based on multi-agent deep reinforcement learning. Using a multi-agent reinforcement learning (MARL) framework, through the interactive learning between the agent and the power grid environment, the output of each generator is dynamically adjusted to optimize the scheduling efficiency of the power grid. The technical solution is as follows:
[0048] Step 1: Grid state encoding
[0049] Obtain the grid state (such as load, power generation, frequency, voltage, etc.) from the power grid simulation environment, and process the grid state through the input embedding layer to map the data into a high-dimensional vector space, thereby retaining the rich features of the grid state. Since the Transformer model does not have a natural time series structure, position encoding is added after the input embedding layer. By using sine and cosine functions to add time order information to the input data, the model can capture the dependency relationship between the grid state and the generator state. Through state space modeling in reinforcement learning, the grid state is represented as the state in the Markov decision process (MDP), providing a basis for subsequent policy optimization.
[0050] Step 2: Multi-agent state encoding
[0051] After the grid state encoding is completed, each multi-agent obtains its own state encoding. However, the decision-making of each agent is affected by its related agents. Therefore, each agent needs to reconstruct its own state representation according to the states of its related agents. During the reconstruction process, the empirical Bellman error between the original state and the reconstructed state is minimized. Through the value function approximation method in reinforcement learning, the state encoding of the agent can effectively capture the dynamic changes in the environment and provide accurate input for subsequent action generation.
[0052] Step 3: Multi-agent action decoding and generation
[0053] Based on the state encoding results of the agents, each agent needs to generate corresponding scheduling decisions. Based on the sequence encoding ability of Transformer, the decoder receives the multi-agent state representation and the previous agent action information, and decodes to generate the current agent action. During the decoding process, based on the clipped PPO loss function, the dynamic characteristics of the system are captured and the optimal action strategy is generated. Through the policy gradient method in reinforcement learning, the agent can gradually optimize its action strategy to maximize the long-term cumulative reward. Specifically, the PPO algorithm ensures the stability of the training process by restricting the magnitude of policy updates, avoiding performance degradation caused by policy mutations.
[0054] Step Four: Simulation Environment Verification
[0055] The scheduling decisions generated by the model are fed into the power grid simulation environment to obtain the next state of the power grid under the current action decision. Through the feedback of the simulation environment, the effectiveness of the scheduling decisions is evaluated, and the model is optimized according to the reward function (such as generation cost, transmission loss, etc.). The feedback information of the simulation environment is used to update the policy network of the agent, ensuring that the model can continuously optimize the scheduling decisions in the dynamically changing power grid environment. Through the experience replay mechanism in reinforcement learning, the agent can learn from historical data, avoid overfitting and improve generalization ability. In addition, through multi-agent collaborative training, the agent can learn the globally optimal strategy in a complex environment.
[0056] Embodiment 2:
[0057] Based on Embodiment 1 and in combination with the attached drawings, the specific details of each step are further disclosed in this embodiment. Among them Figure 1 is a schematic diagram of the basic framework of deep reinforcement learning in power grid scheduling in the present invention, which shows the process of an agent encoding the power grid state through a deep neural network and generating corresponding scheduling actions. Figure 2 is a schematic diagram of a multi-agent Transformer encoder-decoder architecture, showing the process of using the self-attention mechanism to capture the dependencies between multi-agent observation sequences and generate the joint action distribution. Figure 3 is a multi-agent training framework diagram based on the MAPPO algorithm, showing the overall process of agents interacting, policy updating and experience replay in the simulation environment.
[0058] Step 1: Power Grid State Encoding
[0059] Function: Obtain the power grid state from the power grid simulation environment and encode it into a high-dimensional vector.
[0060] Technical Means:
[0061] The state space includes key information on the operation of the power grid, covering dynamic variables such as generator status, load demand, and new energy output. The specific design is as follows: o i ={P gen ,P load ,P renew ,Line,Loss}, where represents the active power output of each generator, P load represents the current system load demand, P renew represents the current new energy output, Line represents the wind power and photovoltaic output, and Loss represents the system transmission loss.
[0062] To capture the sequential decision-making relationship of multi-agents based on load and new energy output, a Transformer network is used to encode the observed state sequences of the load and new energy output of each agent. After encoding, we get Through state space modeling in reinforcement learning, the power grid state is represented as the state in a Markov decision process (MDP), providing a basis for subsequent policy optimization. Specifically, state encoding maps the power grid state to a high-dimensional vector space through an input embedding layer and adds temporal order information through positional encoding, enabling the model to capture the long-range dependencies between the power grid state and the generator state.
[0063] Step 2: Multi-agent state encoding
[0064] Function: After the power grid state encoding is completed, each multi-agent obtains its own state encoding. However, the decision-making of each agent is affected by its related agents. Therefore, each agent needs to reconstruct its own state representation based on the states of its related agents.
[0065] Technical means:
[0066] The parameters of the encoder are represented by φ. It receives the observed sequence of the power grid state and contains multiple layers of computational modules. Each computational block consists of a self-attention mechanism and a multi-layer perceptron (MLP), and contains residual connections to prevent gradient vanishing and network degradation with increasing depth. We represent the output of the observed encoding as It not only encodes the respective observed information of agents i1,…,i n but also encodes the high-level relationships between agents. To obtain a more diverse representation, during the training phase, we let the encoder fit the reward function, and its goal is to minimize the empirical Bellman error through the following formula:
[0067]
[0068] In the formula, are the parameters of the target network; T represents the total number of time steps for simulation or training; n represents the current simulation or training round; o t represents the environmental state at time step t; a t represents the agent's action; R(o t ,a t ) represents the immediate reward function; γ represents the discount factor; represents the environmental state observed by agent m at time step t + 1; represents the state value function generated by the encoder at time step t; represents the state value function generated by the target network at time step t + 1.
[0069] Through the Bellman equation in reinforcement learning, the agent can estimate the value of the current state and optimize the value function by minimizing the Bellman error.
[0070] Step 3: Multi-agent action decoding and generation
[0071] Function: Based on the state encoding results of the agents, generate corresponding scheduling decisions.
[0072] Technical means:
[0073] The agent's action is: a t ={ΔP1, ΔP2, …, ΔP n},where ΔP i represents the output adjustment amount of the i-th generator at the current moment, and n is the total number of generators. By adjusting the outputs of each generator, the agent can achieve dynamic scheduling of the entire power grid to meet the load demand. At the same time, in the actual operation of the power grid, the output of the generator has physical constraints and economic constraints. Therefore, the action design needs to meet the following conditions: where and are the minimum and maximum outputs of the generator respectively, and P i is the output of the current generator i. This constraint ensures that the actual output of the generator is always within the safe operating range.
[0074] The parameters of the decoder are θ, which passes the encoded joint action to the multi-layer decoding module. Each decoding module uses the masked self-attention mechanism to ensure that for each agent i j , only calculate to Attention is calculated between the action heads, where r < j, to maintain an ordered update scheme. After this, a second masked attention calculation is performed, which calculates the attention between the action heads and the observed representation sequence. Finally, the decoding block ends with a multi-layer perceptron and skip connections. The output of the last decoding block is the representation of the joint action, which is fed into a multi-layer perceptron to obtain the probability distribution of the actions of agent i m of the action To train the decoder, we minimize the following clipped PPO objective function:
[0075]
[0076] where is the estimate of the joint advantage function, which can be obtained through Generalized Advantage Estimation as a robust estimate of the joint value function. Through the policy gradient method in reinforcement learning, the agent can gradually optimize its action policy to maximize the long-term cumulative reward.
[0077] Step 4: Simulation Environment Verification
[0078] Function: Input the scheduling decision into the power grid simulation environment, evaluate its effectiveness, and optimize the model.
[0079] Technical means:
[0080] In the active power scheduling problem of the power grid, the objective of the reward function design is to minimize the system generation cost, transmission loss, and meet the constraint conditions. The specific definition is as follows:
[0081] R t = -(C gen (P) + λ1 × Loss + λ2 × Penalty)
[0082] where C gen (P) is the generation cost, Loss is the system transmission loss, and Penalty is the constraint violation penalty, including the power balance constraint, that is, the power grid must maintain the balance between supply and demand, otherwise it will cause frequency fluctuations; and the line load constraint, exceeding the maximum load of the transmission line will lead to system safety risks. By setting the reward function, the agent can learn the optimal scheduling strategy to minimize the system cost and ensure the safe operation of the power grid.
[0083] The scheduling decisions generated by the model are input into the power grid simulation environment. The simulation environment calculates the next state of the power grid based on the current action decision and returns the corresponding feedback information. Based on this feedback, the system can evaluate the effectiveness of the scheduling decisions and optimize the model according to a preset reward function (such as generation cost, transmission loss, etc.). The feedback information of the simulation environment is not only used to update the policy network of the agent, but also through the policy iteration mechanism in reinforcement learning to ensure that the model can continuously optimize the scheduling decisions in a dynamically changing power grid environment.
[0084] Embodiment 3:
[0085] In this embodiment, a power grid active power optimization scheduling system is built to automatically run the technical process of the power grid active power optimization scheduling method based on deep reinforcement learning multi-agent disclosed in Embodiment 1 or 2 above. In this embodiment, the power grid active power optimization scheduling system can be composed of a collection module, an encoder, a decoder, and a decision output module. The collection module is used to obtain the key information of the power grid operation to form a state space, and encodes it using a Transformer network to obtain the observation sequence of the power grid state. The observation sequence of the power grid state is input into the encoder to reconstruct the observation encoding of each agent. The observation encoding of the agent is input into the decoder to decode and generate the agent action. The decision output module generates the corresponding scheduling decision based on the agent action, and realizes the dynamic scheduling of the entire power grid by adjusting the output of each generator.
[0086] Specifically, the operation processes of the above collection module, encoder, decoder, and decision output module can refer to the specific details disclosed in Embodiment 1 and Embodiment 2, and will not be elaborated here.
[0087] Embodiment 4:
[0088] The technical process of the power grid active power optimization scheduling method based on deep reinforcement learning multi-agent disclosed in Embodiment 1 or 2 above can be implemented in whole or in part by software, hardware, firmware, or any other combination.
[0089] When implemented using hardware, the above embodiments can run the working logic and calculation process on an electronic device after being compiled by software in whole or in part. The electronic device includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete the communication with each other through the communication bus. The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the technical process of the power grid active power optimization scheduling method based on deep reinforcement learning multi-agent disclosed in the above embodiments.
[0090] When implemented using software, the above-described embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. If the above method is implemented in the form of software functional modules and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the related art may be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0091] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation of the present invention itself. Various changes may be made to it in form and detail without departing from the spirit and scope of the present invention as defined by the appended claims.
Claims
1. A method for active power optimal scheduling of power grid based on multi-agent of deep reinforcement learning, characterized in that It includes the following steps: Obtain key information of power grid operation to form a state space, and use a Transformer network to encode it to obtain an observation sequence of the power grid state; Input the observation sequence of the power grid state into the encoder to reconstruct the observation encoding of each agent; Input the observation encoding of the agent into the decoder to decode and generate agent actions; Generate corresponding scheduling decisions based on the agent actions, and realize dynamic scheduling of the entire power grid by adjusting the output of each generator.
2. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 1, characterized in that, The key information of the power grid operation includes the active power output P of each generator gen , the current system load demand P load , the current new energy output P renew , the wind power and photovoltaic output Line, and the system transmission loss Loss; The above key information forms the state space o i : o i = {P gen , P load , P renew , Line, Loss}.
3. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 2, wherein Use a Transformer network to encode the state space under multiple time series to obtain an observation sequence of the power grid state where represents the state space under the nth time series.
4. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 1, characterized in that Input the observation sequence of the power grid state into the encoder, and reconstruct the observation encoding of each agent, where the observation encoding is represented as where represents the observation encoding at the n-th time series.
5. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 1, characterized in that It also includes: Train the encoder parameter φ so that the encoder fits the reward function, and the goal is to minimize the empirical Bellman error. The formula is as follows: In the formula, are the parameters of the target network; T represents the total number of time steps for simulation or training; n represents the current simulation or training round; o t represents the environmental state at time step t; a t represents the agent's action; R(o t ,a t ) represents the immediate reward function; γ represents the discount factor; represents the environmental state observed by agent m at time step t + 1; represents the state value function generated by the encoder at time step t; represents the state value function generated by the target network at time step t + 1.
6. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 1, characterized in that Encode the observations Input them into the decoder and output the agent action a t ={ΔP1, ΔP2, …, ΔP n}; where ΔP i represents the output adjustment amount of the i-th generator at the current moment, and n is the total number of generators.
7. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 6, characterized in that Realize dynamic scheduling of the entire power grid by adjusting the output of each generator, so as to meet the load demand; The motion design meets the following conditions: Among them, P i is the output of the current generator i; and are the minimum and maximum outputs of the generator respectively, ensuring that the actual output of the generator is always within the safe operating range.
8. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 1, characterized in that The encoder first transmits the encoded joint actions to a multi-layer decoding module, and each layer of the decoding module employs a masked self-attention mechanism to ensure that for each agent i j only calculates attention between the action heads up to where r < j; Subsequently, perform the second masked attention calculation to calculate the attention between the action head and the observation representation sequence; Finally, the decoding module ends with a multi-layer perceptron and skip connections. The output of the last decoding module is the joint action representation. The joint action representation is fed into the multi-layer perceptron to obtain the probability distribution of the actions of agent i m where πθi,m represents the policy function of agent im; ati,m represents the action executed by agent im at time step t; oti,1:m represents the observations of the first to the m-th agents at time step t; ati,1:m-1 represents the actions selected by the first to the (m-1)-th agents at time step t. 9. The grid active power optimal scheduling method based on deep reinforcement learning multi-agent according to claim 1, characterized in that It also includes: Train the parameters θ of the encoder, with the goal of minimizing the clipped PPO objective function L Decoder (θ): wherein, is the estimate of the combined advantage function; T represents the total number of time steps for simulation or training; n represents the current simulation or training round; represents the advantage ratio obtained by agent i m at time step t based on the current parameter θ; ∈ represents the hyperparameter in the PPO algorithm, and clip(*) is the truncation function in reinforcement learning, which is used to limit the amplitude of policy update.
10. A power grid active power optimization dispatching system, characterized in that, It includes: An acquisition module; The acquisition module is used to obtain key information of power grid operation to form a state space, and use a Transformer network to encode it to obtain an observation sequence of the power grid state; An encoder; Input the observation sequence of the power grid state into the encoder to reconstruct the observation encoding of each agent; A decoder; input the observation encoding of the agent into the decoder to decode and generate agent actions; A decision output module; the decision output module generates corresponding scheduling decisions based on the agent actions, and realizes dynamic scheduling of the entire power grid by adjusting the output of each generator.
11. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the method for active power optimization scheduling of a power grid based on deep reinforcement learning multi-agent as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, At least one executable instruction is stored in the storage medium. When the executable instruction runs on the electronic device, the electronic device executes the method for active power optimization scheduling of a power grid based on deep reinforcement learning multi-agent as described in any one of claims 1 to 9.