A power system real-time low-carbon economic dispatch method and device based on a parallel distributed flexible actor-critic, an electronic device, and a storage medium

By constructing a Markov decision process model and training an agent using a parallel distributed flexible actor-evaluator algorithm, the efficiency and accuracy issues of real-time low-carbon economic dispatching in power systems were solved, thus achieving efficient low-carbon economic dispatching of power systems.

CN119651634BActive Publication Date: 2025-10-17GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411780753.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-10-17
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

Existing technologies have low efficiency and accuracy in real-time low-carbon economic scheduling, and are unable to effectively deal with the uncertainty of power system operation and the complexity of carbon emissions and economic optimization brought about by the increase in new energy penetration.

Method used

A real-time low-carbon economic dispatching method for power systems based on parallel distributed flexible agent-evaluator is adopted. By constructing a Markov decision process model, the agent is trained offline using the parallel distributed flexible agent-evaluator algorithm and integrated into the power system dispatching equipment to generate a real-time dispatching scheme.

Benefits of technology

It has improved the efficiency and accuracy of real-time dispatching of the power system, optimized carbon emissions and economy, and achieved the coordinated development of low-carbon and economic goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119651634B_ABST
    Figure CN119651634B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, electronic device, and storage medium for real-time low-carbon economic dispatch of power systems based on a parallel distributed flexible actor-evaluator. The method comprises: constructing a real-time low-carbon economic dispatch model for power systems based on a Markov decision process, based on information on the joint participation of large-scale renewable energy and carbon trading mechanisms in real-time optimized dispatch of power systems. Based on the model, a parallel distributed flexible actor-evaluator algorithm is proposed to offline train an intelligent agent to learn low-carbon economic dispatch strategies; the trained intelligent agent is integrated into a power system dispatch device and put into use, and the dispatch plan for each generator unit is determined based on the real-time status information of the system. Based on the demand for real-time low-carbon economic dispatch of power systems, the present invention processes the continuous action space of the dispatch problem through a flexible actor-evaluator; and utilizes a distributed flexible strategy iteration framework and parallel technology to improve the flexible actor-evaluator, thereby increasing the speed and accuracy of the solution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid dispatching, and in particular to a power system real-time low-carbon economic dispatching method and device based on parallel distributed flexible actor-critic, an electronic device, and a storage medium. BACKGROUND

[0002] With the increasing global climate change problem, carbon emission reduction has become a core strategic goal of national development. In the process of promoting energy transformation, China is gradually transitioning from the traditional energy consumption double control policy to the carbon emission double control policy, and the energy saving and carbon reduction of the power system has become a key link to achieve low-carbon development. In particular, under the guidance of the carbon peak and carbon neutralization targets, the power industry is shifting from relying solely on new energy generation technology to a comprehensive consideration of the "electricity-carbon" coordination mechanism for the entire chain of carbon trading market and energy production and consumption. By effectively reducing carbon emissions in power system operation and optimizing dispatching economy, the coordinated development of low-carbon and economic goals is an important research direction at present.

[0003] However, the existing technology still faces many challenges in real-time low-carbon economic dispatching. The increase in new energy penetration rate has brought uncertainty to the operation of the power system, exacerbating the complexity of carbon emission and economic optimization. Traditional dispatching methods often have low efficiency and accuracy in dealing with randomness, nonlinearity, and dynamic changes in the system. SUMMARY

[0004] The present application provides a power system real-time low-carbon economic dispatching method and device based on parallel distributed flexible actor-critic, an electronic device, and a storage medium. By implementing the present application, the efficiency and accuracy of power system real-time dispatching can be improved.

[0005] An embodiment of the present application provides a power system real-time low-carbon economic dispatching method based on parallel distributed flexible actor-critic, comprising:

[0006] A power system real-time low-carbon economic dispatching model is constructed based on a Markov decision process; wherein the state space of the Markov decision process is constructed according to the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. the active power output of the thermal power unit at the time, the reactive power output of the thermal power unit at the time, the voltage amplitude of all nodes at the time, the phase angle of all nodes at the time, the predicted active power output of the new energy unit at the time, the predicted reactive power output of the new energy unit at the time, the predicted active load of the load node at the time, and the predicted reactive load of the load node at the time. Output increment at the moment, PV node Voltage amplitude at the moment and new energy units The action space is constructed based on the amount of wind and solar curtailment at the moment; given the current state , actions taken Down, and transfer to a new state Construct a state transfer function; construct a reward function based on the power generation cost of thermal power units, the cost of wind and solar power curtailment, the carbon emission rights trading cost, and the cost penalty;

[0007] Offline training of an intelligent agent corresponding to the real-time low-carbon economic dispatch model of the power system according to a parallel distributed flexible actor-evaluator algorithm;

[0008] The offline trained intelligent agent is integrated into the power system dispatching equipment so that the power system dispatching equipment generates a dispatching plan for each generator set in the power system based on real-time status information, and dispatches the generator sets in the power system according to the dispatching plan.

[0009] Furthermore, the state space includes:

[0010]

[0011] in, for Status information at all times; For thermal power units Always make contributions; For thermal power units Reactive power output at all times, For all nodes The voltage amplitude at the moment For all nodes The phase angle of time, For new energy units The predicted active power output at the time For new energy units Predicted reactive power output at any given moment, Load node The predicted active load at the time, Load node The predicted reactive load at the moment.

[0012] Furthermore, the action space includes:

[0013]

[0014] in, for Moment-by-moment action information; For thermal power units Output increment at each moment, PV node The voltage amplitude at the moment For new energy units The amount of wind and solar power curtailment at each moment.

[0015] Furthermore, the state transfer function includes:

[0016]

[0017]

[0018]

[0019] in, For new energy units Actual output at any moment; For new energy units The effort of every moment; For new energy units The amount of wind and solar power curtailment at each moment; Thermal power generation units for PV nodes Active power at the moment; For thermal power units Output increment at each moment; For nodes Thermal power units Active power injected at all times; For nodes New energy units Active power injected at all times; For nodes Load Active power injected at all times; and Node 、 The voltage amplitude at , forming a vector ; and and are the node admittance matrices in the first Row, No. the real and imaginary parts of the elements of the column; for Time Node and The phase angle difference between them forms a vector ; For nodes Thermal power units Reactive power injected at all times; For nodes New energy units reactive power injected at time instant t; for node i at time instant t reactive power injected at time instant t.

[0020] Further, the reward function comprises:

[0021]

[0022] wherein, for taking action a in current state s , the system moves to new state s' and the reward; is a cost scaling factor; is a penalty scaling factor; is the total cost of the power system; is the cost penalty term;

[0023] The total cost of the power system is calculated by the following formula:

[0024]

[0025] wherein, is the thermal power generation cost; is the active power output of the thermal power unit in time period t; is the wind or solar curtailment cost; is the wind or solar curtailment amount of the new energy unit in time period t; is the carbon emission trading cost; is the active power injected by the new energy unit at time instant t;

[0026] The cost penalty term is calculated by the following formula:

[0027]

[0028] wherein, is the limited system variable; and correspond to the lower limit and the upper limit, respectively;

[0029] The thermal power generation cost is calculated by the following formula:

[0030]

[0031] wherein, is the set of thermal power units; ,​​​ and is the power generation cost coefficient of the thermal power unit ; is the decision time interval;

[0032] The wind and light abandoned cost is calculated by the following formula:

[0033]

[0034] wherein, is the set of new energy power generation units; is the unit wind and light abandoned cost;

[0035] The carbon emission right transaction cost is calculated by the following formula:

[0036]

[0037] wherein, is the carbon trading price; is the carbon quota; is the carbon emission;

[0038] The carbon quota is calculated by the following formula:

[0039]

[0040] wherein, is the unit electricity carbon emission allocation quota;

[0041] The carbon emission is calculated by the following formula:

[0042]

[0043] wherein, is the unit coal consumption of the thermal power unit; is the coal emission factor;

[0044] The unit coal consumption of the thermal power unit is calculated by the following formula:

[0045]

[0046] wherein, is the active power output upper limit of the thermal power unit ; , and are unit coal consumption coefficients.

[0047] The parallel distributed flexible actor-critic algorithm corresponds to a parallel technical architecture, which includes a global actor network, a main network composed of N local actor networks and N critic networks, and a target network composed of a target global actor network and a target critic network, and a global experience replay pool;

[0048] The off-line training of the agent corresponding to the real-time low-carbon economic dispatch model of the power system according to the parallel distributed flexible actor-critic algorithm comprises:

[0049] Step one, generating the same number of simulators of the low-carbon economic dispatch model of the power system as the local actor networks, and distributing the generated simulators to the local actor networks in a one-to-one ratio;

[0050] Step two, initializing the network parameters of the critic network with random network parameters , the network parameters of the target critic network with random network parameters , the network parameters of the global actor network with random network parameters , the network parameters of the target global actor network with random network parameters , and initializing the global experience replay pool ; wherein the network parameters of the N local actor networks are , , the network parameters of the target critic network are , and the network parameters of the target global actor network are .

[0051] Step three, setting the number of rounds =1;

[0052] Step four, for the first round, obtaining the current state information;

[0053] Step five, setting the time step =1;

[0054] Step six, the local actor network interacts with the corresponding simulator, performs parallel sampling, obtains a sample set, and stores the sample set in the global experience replay pool ;

[0055] Step seven, extracting samples from the global experience replay pool ;

[0056] Step eight, updating the critic network parameters according to the first loss function;

[0057] Step 9: Adjust the global actor network parameters according to the second loss function Make updates;

[0058] Step 10: Regularize the entropy coefficient according to the third loss function Make updates;

[0059] Step 11: Target judge network parameters and target global actor network parameters Make updates;

[0060] Step 12: Set the parameters of the global actor network Parameters synchronized to all local actor networks 、 、…… In order to update the parameters of the local actor network;

[0061] Step 13: Determine the current time Is it less than the preset time step? If so, execute And jump to step 6, if not, jump to step 14; wherein, is the preset time step;

[0062] Step 14: Determine the time Is it less than the preset number of rounds? If so, execute And jump to step 4. If not, the current global actor network is regarded as the trained intelligent agent.

[0063] The first loss function is:

[0064]

[0065] in, for Divergence function; is the Bellman operator; The target judge network outputs Mapping to the distribution of flexible state-action pairs of rewards; The output of the judge network is Mapping to the distribution of flexible state-action pairs of rewards; In state Execute an action Flexible state-action pair rewards; is a constant; Data collected in the past for the strategy; are the network parameters of the target global actor network; an action policy output by the target global actor network;

[0066] The critic network parameters are updated by the following formula:

[0067]

[0068] wherein, is the network parameter of the critic network; is the learning rate of the critic network;

[0069] The second loss function is:

[0070]

[0071] wherein, is the expected value of the state-action pair; is the action policy output by the global actor network; here, the action can be converted by using the reparameterization technique, specifically:

[0072]

[0073] wherein, and represent the mean and variance of the action policy output by the global actor network, respectively;

[0074] The global actor network parameters are updated by the following formula:

[0075]

[0076] wherein, is the network parameter of the global actor network; is the learning rate of the global actor network;

[0077] The third loss function is:

[0078]

[0079] wherein, is the target value;

[0080] The entropy regularization coefficient is updated by the following formula:

[0081]

[0082] wherein, is the learning rate of the entropy regularization term coefficient:

[0083] ​​​​The target critic network parameters are updated by the following formula:

[0084]

[0085] wherein, is the update parameter of the target critic network; is the network parameter of the target critic network;

[0086] The target global actor network parameters are updated by the following formula:

[0087]

[0088] wherein, is the update parameter of the target global actor; is the network parameter of the global actor network.

[0089] On the basis of the above method embodiment, the application provides a device embodiment.

[0090] An embodiment of the application provides a parallel distributed flexible actor-critic based real-time low-carbon economic dispatch device of a power system, which comprises a model construction module, an agent offline training module and an agent online dispatch module.

[0091] The model construction module is used for acquiring a real-time low-carbon economic dispatch model of the power system constructed based on a Markov decision process; wherein the state space of the Markov decision process is constructed according to the active power output of the thermal power unit at the current moment, the reactive power output of the thermal power unit at the current moment, the voltage amplitude of all nodes at the current moment, the phase angle of all nodes at the current moment, the predicted active power output of the new energy unit at the current moment, the predicted reactive power output of the new energy unit at the current moment, the predicted active load of the load node at the current moment and the predicted reactive load of the load node at the current moment; the action space is constructed according to the output increment of the thermal power unit at the current moment on the PV node, the voltage amplitude of the PV node at the current moment and the wind and light abandonment amount of the new energy unit at the current moment; the state transition function is constructed according to the current state, the action taken and the transition to the new state; and the reward function is constructed according to the generation cost of the thermal power unit, the wind and light abandonment cost, the carbon emission right transaction cost and the cost penalty term.

[0092] ​​​​​​​​​​​​​​The intelligent agent offline training module is used to perform offline training on the intelligent agent corresponding to the real-time low-carbon economic dispatch model of the power system according to the parallel distributed flexible actor-evaluator algorithm;

[0093] The intelligent agent is put into the online scheduling module, which is used to integrate the intelligent agent that has completed offline training into the power system scheduling equipment, so that the power system scheduling equipment can generate a scheduling plan for each generator set in the power system based on real-time status information, and schedule the generator sets in the power system according to the scheduling plan.

[0094] Based on the above method embodiment, the present invention provides a corresponding electronic device embodiment.

[0095] An embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it can implement the real-time low-carbon economic dispatch method for the power system based on the parallel distributed flexible actor-evaluator as described in any one of the above-mentioned method embodiments.

[0096] Based on the above method embodiment, the present invention provides a corresponding storage medium embodiment.

[0097] An embodiment of the present invention provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for real-time low-carbon economic dispatch of a power system based on a parallel distributed flexible actor-evaluator as described in any one of the above method embodiments can be implemented.

[0098] Compared with the prior art, the present invention has the following beneficial effects:

[0099] Embodiments of the present invention provide a method, apparatus, electronic device, and storage medium for real-time low-carbon economic dispatch of a power system. The method constructs a real-time low-carbon economic dispatch model for a power system based on a Markov decision process, using data from large-scale renewable energy sources and carbon trading mechanisms that jointly participate in the real-time optimized dispatch of the power system. Based on this model, a parallel distributed flexible actor-evaluator algorithm is proposed to train an intelligent agent offline to learn effective low-carbon economic dispatch strategies. The trained intelligent agent is then integrated into the power system dispatch device for online use, determining dispatch plans for each generator set based on real-time system status information.

[0100] Based on the actual needs of real-time low-carbon economic dispatch of power systems, the present invention efficiently processes the continuous action space of the dispatch problem through a flexible actor-evaluator; and improves the flexible actor-evaluator by using a distributed flexible strategy iteration framework and parallel technology, thereby increasing the solution speed and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0101] Figure 1 The present invention provides a flowchart of a method for real-time low-carbon economic dispatch of a power system based on a parallel distributed flexible actor-evaluator according to an embodiment of the present invention.

[0102] Figure 2 1 is a flowchart of a specific training iteration of a parallel distributed flexible actor-evaluator algorithm provided in one embodiment of the present invention.

[0103] Figure 3 This is a structural diagram of a real-time low-carbon economic dispatch of a power system based on a parallel distributed flexible actor-evaluator provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0104] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0105] like Figure 1 As shown, an embodiment of the present invention provides a real-time low-carbon economic dispatch method for a power system based on a parallel distributed flexible actor-evaluator, which includes at least the following steps:

[0106] Step S1: construct a real-time low-carbon economic dispatch model for the power system based on the Markov decision process; wherein, according to the thermal power units Active power output at the moment, thermal power units Reactive power output at all times, all nodes Voltage amplitude at the moment, all nodes Phase angle of time, new energy units Time prediction of active power output and new energy units Predicted reactive power output and load nodes at each moment Predicted active load and load nodes at the moment The predicted reactive load at each moment constructs the state space of the Markov decision process; according to the thermal power units on the PV nodes Output increment at the moment, PV node Voltage amplitude at the moment and new energy units The action space is constructed based on the amount of wind and solar curtailment at the moment; given the current state , actions taken Down, and transfer to a new state Construct a state transfer function; construct a reward function based on the power generation cost of thermal power units, the cost of wind and solar power curtailment, the carbon emission rights trading cost, and the cost penalty;

[0107] In a preferred embodiment, the state space includes:

[0108]

[0109] in, for Status information at all times; For thermal power units Always make contributions; For thermal power units Reactive power output at all times, For all nodes The voltage amplitude at the moment For all nodes The phase angle of time, For new energy units The predicted active power output at the time For new energy units Predicted reactive power output at any given moment, Load node The predicted active load at the time, Load node The predicted reactive load at the moment.

[0110] The action space includes:

[0111]

[0112] in, for Moment-by-moment action information; For thermal power units Output increment at each moment, PV node The voltage amplitude at the moment For new energy units The amount of wind and solar power curtailment at each moment.

[0113] The state transfer function includes:

[0114]

[0115]

[0116]

[0117] in, For new energy units Actual output at any given moment; For new energy units The effort of every moment; For new energy units The amount of wind and solar power curtailment at each moment; Thermal power generation units for PV nodes Active power at the moment; For thermal power units Output increment at each moment; For nodes Thermal power units Active power injected at all times; For nodes New energy units Active power injected at all times; For nodes Load Active power injected at all times; and Node 、 The voltage amplitude at , forming a vector ; and and are the node admittance matrices in the first Row, No. the real and imaginary parts of the elements of the column; for Time Node and The phase angle difference between them forms a vector ; For nodes Thermal power units Reactive power injected at all times; For nodes New energy units Reactive power injected at all times; For nodes Load The reactive power injected at any moment.

[0118] The reward function includes:

[0119]

[0120] in, For the current state Take action , the system transitions to a new state Returns; is the cost scaling factor; is the penalty scaling factor; is the total cost of the power system; is the cost penalty item;

[0121] The total cost of the power system is calculated using the following formula:

[0122]

[0123] wherein, is the cost of thermal power generation units; is the thermal power generation units in the time period; is the active power output; is the cost of abandoned wind and light; is the new energy generation units in the time period; is the amount of abandoned wind and light; is the carbon emission trading cost; is the new energy units injected at time;

[0124] The cost penalty term is calculated by the following formula:

[0125]

[0126] wherein, is the limited system variable; and correspond to the lower limit and the upper limit, respectively;

[0127] The cost of thermal power generation units is calculated by the following formula:

[0128]

[0129] wherein, is the set of thermal power generation units; , and are the cost coefficients of the thermal power generation units ; is the decision time interval;

[0130] The cost of abandoned wind and light is calculated by the following formula:

[0131]

[0132] wherein, is the set of new energy generation units; is the unit cost of abandoned wind and light;

[0133] The carbon emission trading cost is calculated by the following formula:

[0134]

[0135] wherein, is the carbon trading price; is the carbon quota; carbon emission;

[0136] The carbon quota is calculated by the following formula:

[0137]

[0138] The carbon quota is calculated by the following formula:

[0139] The carbon emission is calculated by the following formula:

[0140]

[0141] The carbon emission is calculated by the following formula:

[0142] The unit coal consumption of the thermal power unit is calculated by the following formula:

[0143]

[0144] The carbon emission is calculated by the following formula: The active power output upper limit of the thermal power unit , and The unit coal consumption coefficient.

[0145] Step S2, according to the parallel distributed flexible actor-critic algorithm, off-line training the agent corresponding to the real-time low-carbon economic dispatching model of the power system.

[0146] In one embodiment, the parallel distributed flexible actor-critic algorithm corresponds to a parallel technical architecture, and the parallel technical architecture includes a main network composed of a global actor network, a local actor network and a critic network, a target network composed of a target global actor network and a target critic network, and a global experience replay pool.

[0147] As shown in Figure 2 , an embodiment of the present application provides specific training iteration steps of the parallel distributed flexible actor-critic algorithm, and the off-line training of the agent corresponding to the real-time low-carbon economic dispatching model of the power system according to the parallel distributed flexible actor-critic algorithm includes:

[0148] Step one, generating the same number of simulators of the low-carbon economic dispatching model of the power system as the local actor network, and distributing the generated simulators to each local actor network in a one-to-one ratio;

[0149] ​​​​​Step two, initializing network parameters of the critic network and the target critic network with random network parameters Step two, initializing network parameters of the critic network and the target critic network with random network parameters Step two, initializing network parameters of the critic network and the target critic network with random network parameters Step two, initializing network parameters of the critic network and the target critic network with random network parameters Step two, initializing network parameters of the critic network and the target critic network with random network parameters

[0150] Step three, setting the number of rounds =1;

[0151] Step four, for the first round, obtaining the current state information Step four, for the first round, obtaining the current state information

[0152] Step five, setting the time step =1;

[0153] Step six, each local actor network interacts with the corresponding simulator to perform parallel sampling to obtain a sample set, and stores the sample set in the global experience replay pool ;

[0154] Step seven, extracting samples from the global experience replay pool ;

[0155] Step eight, updating the critic network parameters according to the first loss function

[0156] Step nine, updating the global actor network parameters according to the second loss function

[0157] Step ten, updating the entropy regularization coefficient according to the third loss function

[0158] Step eleven, updating the target critic network parameters and the target global actor network parameters

[0159] Step twelve, synchronizing the global actor network parameters to the parameters of all local actor networks to update the parameters of the local actor networks

[0160] Step thirteen, determining whether the current time is less than the preset time step , if so, executing and jumping to step six, if not, jumping to step fourteen; wherein, is the preset time step

[0161] Step fourteen, determining whether the current time is less than the preset number of rounds , if yes, performing and jumping to step four, if no, the current global actor network is the trained agent. It should be noted that the agent used for scheduling here only refers to the last trained global actor network, and other networks are mainly used to assist the update of the global actor network. The other networks are not involved in the actual scheduling task after being trained, and only the global actor network needs to be retained and integrated into the scheduling system for application.

[0162] It should be noted that the first loss function is:

[0163]

[0164] wherein, is a divergence function; is a Bellman operator; is a mapping from to a distribution of the flexible state-action pair return output by the target critic network; is a mapping from to a distribution of the flexible state-action pair return output by the critic network; is a flexible state-action pair return of performing action in state is a constant; is the data collected by the policy in the past; is the network parameter of the target global actor network; is the action policy output by the target global actor network;

[0165] The critic network parameter is updated by the following formula:

[0166]

[0167] wherein, is the network parameter of the critic network; is the learning rate of the critic network; is calculated by the following formula:

[0168] The second loss function is:

[0169]

[0170] wherein, is the state-action pair expected value; is the action policy output by the global actor network; Here, the action can be converted by using the reparameterization technique, specifically:​​​

[0171]

[0172] wherein, and represent the mean and variance of the global actor network output, respectively;

[0173] The parameters of the global actor network are updated by the following formula:

[0174]

[0175] wherein, is the network parameter of the global actor network; is the learning rate of the global actor network; is calculated by the following formula:

[0176] The third loss function is:

[0177]

[0178] wherein, is the target value;

[0179] The entropy regularization coefficient is updated by the following formula:

[0180] wherein,

[0181] is the learning rate of the entropy regularization term coefficient, is calculated by the following formula: The target critic network parameters are updated by the following formula:

[0182]

[0183]

[0184] wherein, is the updated parameter of the target critic network; is the network parameter of the target critic network;

[0185] The target global actor network parameters are updated by the following formula:

[0186]

[0187] wherein, is the updated parameter of the target global actor; is the network parameter of the global actor network.

[0188] ​​​​It should be noted that deep reinforcement learning generally has the shortcoming of slow training convergence speed. Therefore, the parallel technology is introduced to fully utilize the computing resources of multi-core CPU and GPU by sharing the experience collected by multiple agents exploring multiple environments, thereby further improving the training speed. The agent is a local actor network, and the environment is a simulation model of low-carbon economic dispatch of the power system. The network parameters of all local actor networks are the same, and the parameters of all simulation models are also the same. Due to the different random seeds set by the local actor networks, the output action strategies are different, so that the experience collected by multiple agents exploring multiple environments is shared, the computing resources of multi-core CPU and GPU are fully utilized, and the training speed is further improved.

[0189] In addition, in the flexible actor-critic algorithm, the optimal maximum entropy strategy is learned by alternately performing flexible policy evaluation and flexible policy improvement. The flexible policy evaluation process calculates the expected value of the current strategy after selecting an action in a state, and the flexible policy improvement process improves the strategy based on the value evaluation result of the state-action pair, so that the action selection under the current strategy is more optimal. The expected value of the state-action pair is defined as follows:

[0190]

[0191] wherein, represents the cumulative return of the increase in entropy from . is a discount factor.

[0192] Unlike the flexible actor-critic algorithm that only considers the expected value of the state-action pair, the distributed flexible policy iteration framework directly learns the distribution of the flexible state-action pair return, considers the randomness of the cumulative return, and thus can alleviate the state-action value overestimation problem of the algorithm, improving the training speed and decision accuracy.

[0193] The flexible state-action pair return under the policy can be expressed as follows:

[0194]

[0195] Due to the random policy, the reward function and the state transition function, the flexible state-action pair return is usually a random variable, and the relationship between it and the expected value of the state-action pair is as follows:

[0196]

[0197] Let be the cumulative return of the increase in entropy from The mapping of the distribution of the return of the flexible state-action pair is called the flexible state-action pair return distribution or distribution value function. In the maximum entropy framework, the distribution variable of the Bellman operator can be derived as

[0198]

[0199] where is the Bellman operator, denotes that two random variables follow the same probability law.

[0200] Suppose , where is the distribution of . Therefore, in the flexible policy improvement process, the policy can be updated by minimizing the distribution distance between the Bellman operator and the current distribution, as follows:

[0201]

[0202] where, denotes a measure that measures the distance between two distributions; in order to facilitate calculation, many practical distribution deep reinforcement learning algorithms use the Kullback-Leibler (KL) divergence as the measure, denoted as .

[0203] Step S3, integrating the agent trained offline into the power system dispatching device, so that the power system dispatching device generates a dispatching scheme for each generator set in the power system according to real-time state information, and dispatches the generator sets in the power system according to the dispatching scheme.

[0204] In actual operation, the trained agent is integrated into the power system dispatching device for online application. The agent reads the real-time state information of the system, wherein the state information includes the active power and reactive power of the thermal power generator set at the last time, the voltage amplitude of all nodes at the last time, the phase angle of all nodes at the last time, the ultra-short-term predicted active power of the new energy generator set at the current time, the ultra-short-term predicted active load of all nodes at the current time, and the ultra-short-term predicted reactive load of all nodes at the current time.

[0205] According to the real-time state information of the system, the agent decides the action information of each generator set, wherein the action information includes: the output increment of the thermal power generator set at the current time on the PV node, the voltage amplitude of the PV node at the current time, and the amount of abandoned wind and light of the new energy generator set at the current time.

[0206] According to the action information of each generator set, a power system dispatching scheme is formed.

[0207] Specifically, after the offline training phase, the trained global actor network is integrated into the power system's dispatching system and used online to solve low-carbon economic dispatch strategies in real time. First, state information is constructed based on ultra-short-term forecasts of renewable energy and load and input into the dispatching system. Subsequently, the intelligent agent in the dispatching system uses this state information to make online decisions about the actual output plan for each generator unit. The dispatch plan is then applied to the actual power system environment, and changes in the state variables are fed back to the dispatching system at the next moment. This process is repeated until full-day dispatching is complete.

[0208] Based on the above method embodiments, the present invention provides corresponding device embodiments.

[0209] like Figure 3 As shown, an embodiment of the present invention provides a real-time low-carbon economic dispatching device for a power system based on a parallel distributed flexible actor-evaluator, comprising: a model building module, an intelligent agent offline training module, and an intelligent agent online dispatching module;

[0210] The model building module is used to obtain a real-time low-carbon economic dispatch model for the power system based on the Markov decision process; wherein, according to the thermal power unit Active power output at the moment, thermal power units Reactive power output at all times, all nodes Voltage amplitude at the moment, all nodes Phase angle of time, new energy units Time prediction of active power output and new energy units Predicted reactive power output and load nodes at each moment Predicted active load and load nodes at the moment The predicted reactive load at each moment constructs the state space of the Markov decision process; according to the thermal power units on the PV nodes Output increment at the moment, PV node Voltage amplitude at the moment and new energy units The action space is constructed based on the amount of wind and solar curtailment at the moment; given the current state , actions taken Down, and transfer to a new state Construct a state transfer function; construct a reward function based on the power generation cost of thermal power units, the cost of wind and solar power curtailment, the carbon emission rights trading cost, and the cost penalty;

[0211] The intelligent agent offline training module is used to perform offline training on the intelligent agent corresponding to the real-time low-carbon economic dispatch model of the power system according to the parallel distributed flexible actor-evaluator algorithm;

[0212] The intelligent agent is input into an online scheduling module, which is configured to integrate the intelligent agent trained offline into a power system scheduling device, so that the power system scheduling device generates a scheduling scheme of each generator set in the power system according to real-time state information, and schedules the generator sets in the power system according to the scheduling scheme.

[0213] It should be noted that the above-described embodiments of the device correspond to the above-described embodiments of the application, and can implement any one of the above-described power system real-time low-carbon economic scheduling methods based on parallel distributed flexible actor-critic. In addition, the above-described embodiments of the device are only illustrative, and the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. In addition, in the device embodiment provided by the application, the connection relationship between the modules indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.

[0214] On the basis of the above-described method embodiments of the application, an electronic device embodiment is provided.

[0215] An embodiment of the application provides an electronic device, which comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the power system real-time low-carbon economic scheduling method based on parallel distributed flexible actor-critic is implemented, or the processor executes the computer program to implement the functions of the modules in the above-described device embodiments.

[0216] For example, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the application. The one or more modules can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the terminal device.

[0217] The terminal device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and other computing devices. The terminal device can include, but is not limited to, a processor and a memory.

[0218] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor is a control center of the terminal device, and connects all parts of the terminal device through various interfaces and lines.

[0219] The memory can be used to store the computer program and / or modules, and the processor realizes various functions of the terminal device by running or executing the computer program and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required by a function, etc. The data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a nonvolatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.

[0220] On the basis of the above-mentioned method embodiment, the application provides a storage medium embodiment;

[0221] Another embodiment of the application provides a storage medium, which comprises a stored computer program. When the computer program runs, the device where the storage medium is located performs the power system real-time low-carbon economic dispatch method based on a parallel distributed flexible actor-critic provided by the application.

[0222] The storage medium is a computer readable storage medium, and the computer program includes computer program code in the form of source code, object code, an executable file, or some intermediate form, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, a recording medium, a U disk, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in a jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include an electrical carrier signal and a telecommunication signal.

[0223] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.

[0224] The above is the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.

Claims

1. A real-time low-carbon economic dispatch method for power systems based on parallel distributed flexible actor-evaluator, characterized in that: include: A real-time low-carbon economic dispatch model for power systems is constructed based on the Markov decision process; Active power output at the moment, thermal power units Reactive power output at all times, all nodes Voltage amplitude at the moment, all nodes Phase angle of time, new energy units Time prediction of active power output and new energy units Predicted reactive power output and load nodes at each moment Predicted active load and load nodes at the moment The predicted reactive load at each moment constructs the state space of the Markov decision process; according to the thermal power units on the PV nodes Output increment at the moment, PV node Voltage amplitude at the moment and new energy units The action space is constructed based on the amount of wind and solar curtailment at the moment; given the current state , actions taken Down, and transfer to a new state Construct a state transfer function; construct a reward function based on the power generation cost of thermal power units, the cost of wind and solar power curtailment, the carbon emission rights trading cost, and the cost penalty; Offline training of an intelligent agent corresponding to the real-time low-carbon economic dispatch model of the power system according to a parallel distributed flexible actor-evaluator algorithm; Integrating the offline trained agent into a power system dispatching device, so that the power system dispatching device generates a dispatching plan for each generator set in the power system based on real-time status information, and dispatches the generator sets in the power system according to the dispatching plan; The parallel distributed flexible actor-evaluator algorithm corresponds to a parallel technology architecture, which includes: a global actor network, A main network consisting of a local actor network and a critic network, a target network consisting of a target global actor network and a target critic network, and a global experience replay pool; The offline training of the intelligent agent corresponding to the real-time low-carbon economic dispatch model of the power system according to the parallel distributed flexible actor-evaluator algorithm includes: Step 1: Generate the same number of simulators of the low-carbon economic dispatch model of the power system as the number of local actor networks, and distribute the generated simulators to each local actor network in a one-to-one ratio; Step 2: Use random network parameters Initialize the network parameters of the judge network and the target judge network; use random network parameters Initialize the network parameters of the global actor network, the target global actor network, and N local actor networks; initialize the global experience replay pool ; Step 3: Set the number of rounds =1; Step 4: Round, get the current status information; Step 5: Set the time step =1; Step 6: The local actor network interacts with the corresponding simulator to perform parallel sampling to obtain a sample set, and stores the sample set in the global experience replay pool. ; Step 7: From Extract samples; Step 8: Adjust the network parameters of the judge according to the first loss function Make updates; Step 9: Adjust the global actor network parameters according to the second loss function Make updates; Step 10: Regularize the entropy coefficient according to the third loss function Make updates; Step 11: Update the target evaluator network parameters and the target global actor network parameters; Step 12: Set the parameters of the global actor network Synchronize to the parameters of all local actor networks to update the parameters of the local actor networks; Step 13: Determine the current time Is it less than the preset time step? If so, execute And jump to step 6, if not, jump to step 14; wherein, is the preset time step; Step 14: Determine the time Is it less than the preset number of rounds? If so, execute And jump to step 4. If not, the current global actor network is regarded as the trained intelligent agent; The first loss function is: in, for Divergence function; is the Bellman operator; The target judge network outputs Mapping to the distribution of flexible state-action pairs of rewards; The output of the judge network is Mapping to the distribution of flexible state-action pairs of rewards; In state Execute an action Flexible state-action pair rewards; is a constant; Data collected in the past for the strategy; are the network parameters of the target global actor network; The action strategy output by the target global actor network; The network parameters of the judge are calculated by the following formula To update: in, are the network parameters of the judge network; is the learning rate of the judge network; The second loss function is: in, is the expected value of the state-action pair; The action strategy output by the global actor network; here we can use the reparameterization technique to adjust the action Perform the conversion, specifically: in, and Represent the mean and variance of the global actor network output respectively; The parameters of the global actor network are calculated by the following formula To update: in, are the network parameters of the global actor network; is the learning rate of the global actor network; The third loss function is: in, is the target value; The entropy regularization coefficient is calculated by the following formula: To update: in, is the learning rate of the entropy regularization coefficient: The target judge network parameters are updated using the following formula: in, is the updated parameter of the target judge network; are the network parameters of the target judge network; The target global actor network parameters are updated using the following formula: in, is the updated parameter of the target global actor; are the network parameters of the global actor network.

2. The method for real-time low-carbon economic dispatch of power systems based on parallel distributed flexible actor-evaluator according to claim 1, characterized in that: The state space includes: in, for Status information at all times; For thermal power units Always make contributions; For thermal power units Reactive power output at all times, For all nodes The voltage amplitude at the moment For all nodes The phase angle of time, For new energy units Time prediction of active output, For new energy units Predicted reactive power output at any given moment, Load node The predicted active load at the time, Load node The predicted reactive load at the moment.

3. The method for real-time low-carbon economic dispatch of power systems based on parallel distributed flexible actor-evaluator according to claim 2, characterized in that: The action space includes: in, for Moment-by-moment action information; For thermal power units Output increment at each moment, PV node The voltage amplitude at the moment For new energy units The amount of wind and solar power curtailment at each moment.

4. The method for real-time low-carbon economic dispatch of power systems based on parallel distributed flexible actor-evaluator according to claim 3, characterized in that: The state transfer function includes: in, For new energy units Actual output at any given moment; For new energy units The effort of every moment; For new energy units The amount of wind and solar power curtailment at each moment; Thermal power generation units for PV nodes Active power at the moment; For thermal power units Output increment at each moment; For nodes Thermal power units Active power injected at all times; For nodes New energy units Active power injected at all times; For nodes Load Active power injected at all times; and Node 、 The voltage amplitude at , forming a vector ; and The node admittance matrix is Row, No. the real and imaginary parts of the elements of the column; for Time Node and The phase angle difference between them forms a vector ; For nodes Thermal power units Reactive power injected at all times; For nodes New energy units Reactive power injected at all times; For nodes Load The reactive power injected at any moment.

5. The method for real-time low-carbon economic dispatch of power systems based on parallel distributed flexible actor-evaluator according to claim 4, characterized in that: The reward function includes: in, For the current state Take action , the system transitions to a new state Returns; is the cost scaling factor; is the penalty scaling factor; is the total cost of the power system; is the cost penalty item; The total cost of the power system is calculated using the following formula: in, The cost of electricity generation for thermal power units; For thermal power units exist Active power output during the time period; Cost of curtailing wind and solar power; For new energy generators exist The amount of wind and solar power curtailment during the period; The cost of carbon emission rights trading; For new energy units exist Active power injected at all times; The cost penalty term is calculated by the following formula: in, is a restricted system variable; and Corresponding to the lower and upper limits respectively; The power generation cost of the thermal power unit is calculated using the following formula: in, is a collection of thermal power units; 、 and For thermal power units The power generation cost coefficient; is the decision time interval; The cost of curtailing wind and solar power is calculated using the following formula: in, It is a collection of new energy generators; is the unit cost of wind and solar curtailment; The carbon emission rights transaction cost is calculated using the following formula: in, is the carbon trading price; for carbon quotas; is carbon emissions; The carbon quota is calculated using the following formula: in, Allocate a quota for carbon emissions per unit of electricity; The carbon emissions are calculated using the following formula: in, is the unit coal consumption of thermal power units; is the coal emission factor; The unit coal consumption of the thermal power unit is calculated by the following formula: in, For thermal power units The upper limit of active power output; 、 and is the unit coal consumption coefficient.

6. A real-time low-carbon economic dispatching device for power systems based on parallel distributed flexible actors and evaluators, characterized in that: include: Model building module, agent offline training module and agent online scheduling module; The model building module is used to obtain a real-time low-carbon economic dispatch model for the power system based on the Markov decision process; wherein, according to the thermal power unit Active power output at the moment, thermal power units Reactive power output at all times, all nodes Voltage amplitude at the moment, all nodes Phase angle of time, new energy units Time prediction of active power output and new energy units Predicted reactive power output and load nodes at each moment Predicted active load and load nodes at the moment The predicted reactive load at each moment constructs the state space of the Markov decision process; according to the thermal power units on the PV nodes Output increment at the moment, PV node Voltage amplitude at the moment and new energy units The action space is constructed based on the amount of wind and solar curtailment at the moment; given the current state , actions taken Down, and transfer to a new state Construct a state transfer function; construct a reward function based on the power generation cost of thermal power units, the cost of wind and solar power curtailment, the carbon emission rights trading cost, and the cost penalty; The intelligent agent offline training module is used to perform offline training on the intelligent agent corresponding to the real-time low-carbon economic dispatch model of the power system according to the parallel distributed flexible actor-evaluator algorithm; The agent is put into an online scheduling module, which is used to integrate the offline trained agent into the power system scheduling equipment, so that the power system scheduling equipment generates a scheduling plan for each generator set in the power system based on real-time status information, and schedules the generator sets in the power system according to the scheduling plan; The parallel distributed flexible actor-evaluator algorithm corresponds to a parallel technology architecture, which includes: a global actor network, A main network consisting of a local actor network and a critic network, a target network consisting of a target global actor network and a target critic network, and a global experience replay pool; The offline training of the intelligent agent corresponding to the real-time low-carbon economic dispatch model of the power system according to the parallel distributed flexible actor-evaluator algorithm includes: Step 1: Generate the same number of simulators of the low-carbon economic dispatch model of the power system as the number of local actor networks, and distribute the generated simulators to each local actor network in a one-to-one ratio; Step 2: Use random network parameters Initialize the network parameters of the judge network and the target judge network; use random network parameters Initialize the network parameters of the global actor network, the target global actor network, and N local actor networks; initialize the global experience replay pool ; Step 3: Set the number of rounds =1; Step 4: Round, get the current status information; Step 5: Set the time step =1; Step 6: The local actor network interacts with the corresponding simulator to perform parallel sampling, obtains the sample set, and stores the sample set in the global experience replay pool. ; Step 7: From Extract samples; Step 8: Adjust the network parameters of the judge according to the first loss function Make updates; Step 9: Adjust the global actor network parameters according to the second loss function Make updates; Step 10: Regularize the entropy coefficient according to the third loss function Make updates; Step 11: Update the target evaluator network parameters and the target global actor network parameters; Step 12: Set the parameters of the global actor network Synchronize to the parameters of all local actor networks to update the parameters of the local actor networks; Step 13: Determine the current time Is it less than the preset time step? If so, execute And jump to step 6, if not, jump to step 14; wherein, is the preset time step; Step 14: Determine the time Is it less than the preset number of rounds? If so, execute And jump to step 4. If not, the current global actor network is regarded as the trained intelligent agent; The first loss function is: in, for Divergence function; is the Bellman operator; The target judge network outputs Mapping to the distribution of flexible state-action pairs of rewards; The output of the judge network is Mapping to the distribution of flexible state-action pairs of rewards; In state Execute an action Flexible state-action pair rewards; is a constant; Data collected in the past for the strategy; are the network parameters of the target global actor network; The action strategy output by the target global actor network; The network parameters of the judge are calculated by the following formula To update: in, are the network parameters of the judge network; is the learning rate of the judge network; The second loss function is: in, is the expected value of the state-action pair; The action strategy output by the global actor network; here we can use the reparameterization technique to adjust the action Perform the conversion, specifically: in, and Represent the mean and variance of the global actor network output respectively; The parameters of the global actor network are calculated by the following formula To update: in, are the network parameters of the global actor network; is the learning rate of the global actor network; The third loss function is: in, is the target value; The entropy regularization coefficient is calculated by the following formula: To update: in, The learning rate of the entropy regularization coefficient: The target judge network parameters are updated using the following formula: in, is the updated parameter of the target judge network; are the network parameters of the target judge network; The target global actor network parameters are updated using the following formula: in, is the updated parameter of the target global actor; are the network parameters of the global actor network.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it can implement the real-time low-carbon economic dispatch method for the power system based on the parallel distributed flexible actor-evaluator as described in any one of claims 1 to 5.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can implement the real-time low-carbon economic dispatch method of the power system based on the parallel distributed flexible actor-evaluator as described in any one of claims 1 to 5.