A power grid scheduling method, device and computer equipment based on a DDPG algorithm
By using a grid dispatching method based on the DDPG algorithm, target actions are generated through policy networks and dual-value networks to adjust the active power of generator units and loads. This solves the problem that it is difficult to simultaneously minimize grid dispatching costs and maximize renewable energy consumption in traditional technologies, thus achieving efficient and accurate grid dispatching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN POWER SUPPLY BUREAU
- Filing Date
- 2023-12-01
- Publication Date
- 2026-07-31
AI Technical Summary
Traditional technologies struggle to simultaneously minimize the operating costs of power grid dispatch and maximize the absorption of new energy sources, thus failing to effectively optimize power grid dispatch.
A grid dispatching method based on the DDPG algorithm is adopted. By obtaining the active power and control range of new energy generating units, thermal power units and loads, and combining the objective functions of minimizing operating costs and maximizing new energy consumption, the target actions are generated by using a strategy network and a dual value network to adjust the active power of generating units and loads to achieve optimal dispatching.
It improves the efficiency and accuracy of power grid dispatch, achieves optimal solutions under constraints, and simultaneously minimizes operating costs and maximizes the absorption of new energy sources.
Smart Images

Figure CN117748513B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power grid dispatching technology, and in particular to a power grid dispatching method, apparatus, computer equipment, and storage medium based on the DDPG algorithm. Background Technology
[0002] In traditional technologies, the power generation of new energy sources is easily affected by weather conditions, which may lead to large fluctuations in the generated electricity, making it difficult to predict and adjust the amount of electricity generated. Therefore, in power systems, both new energy sources and traditional energy sources are usually used for power generation.
[0003] However, traditional technologies do not optimize grid dispatch based on the two dimensions of operating costs and renewable energy consumption, making it difficult to simultaneously achieve grid dispatch that minimizes operating costs and maximizes renewable energy consumption. Summary of the Invention
[0004] Therefore, it is necessary to provide a grid dispatching method, device, computer equipment, computer-readable storage medium, and computer program product based on the DDPG algorithm that can efficiently and synchronously perform grid dispatching for minimizing operating costs and maximizing renewable energy consumption, in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a power grid dispatching method based on the DDPG algorithm, comprising:
[0006] The active power and corresponding control range of new energy units, active power and corresponding control range of thermal power units, and active power and corresponding control range of loads are obtained respectively.
[0007] Obtain the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; wherein the objective function of minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units;
[0008] Obtain a set of constraints, which includes active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units.
[0009] Based on the strategy network corresponding to the DDPG algorithm, corresponding actions are generated according to the states of the new energy units, thermal power units, and loads respectively. Based on the dual-value network corresponding to the DDPG algorithm, corresponding expected values of returns are generated according to the states and actions. The dual-value network is updated based on the difference of the expected values of returns to obtain the target expected value of returns. The strategy network is updated based on the target expected value of returns.
[0010] Based on the updated policy network and the updated dual-value network, the maximum expected return value is generated, and the action corresponding to the maximum expected return value is taken as the target action; wherein the target action refers to the optimal solution result of the objective function of minimizing operating costs and the objective function of maximizing new energy consumption under the set of constraints.
[0011] Based on the target action, the active power of the new energy generating units is adjusted according to the corresponding control range, the active power of the thermal power units is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range.
[0012] In one embodiment, acquiring the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load respectively includes: acquiring different acquisition times with the same time interval, sequentially using the different acquisition times as target acquisition times, and acquiring the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load corresponding to the target acquisition time; adjusting the active power of the new energy generating unit based on the corresponding control range, adjusting the active power of the thermal power unit based on the corresponding control range, and adjusting the active power of the load based on the corresponding control range based on the target action includes: adjusting the active power of the new energy generating unit corresponding to the target acquisition time based on the corresponding control range, adjusting the active power of the thermal power unit based on the corresponding control range, and adjusting the active power of the load based on the corresponding control range based on the target action.
[0013] In one embodiment, before acquiring the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load, the method further includes: acquiring the current operation and maintenance status of the new energy generating unit, the thermal power unit, and the load; wherein the operation and maintenance status refers to the operating condition, health status, and performance; calculating the expected return value corresponding to different actions selected under a given operation and maintenance status based on a dual Q network; converting each expected return value into the probability of selecting different actions under a given operation and maintenance status, and determining the optimal strategy based on the probability corresponding to each action; wherein the optimal strategy refers to selecting the action corresponding to the optimal expected return value under a given operation and maintenance status; and updating the operation and maintenance status of the new energy generating unit, the thermal power unit, and the load based on the optimal strategy.
[0014] In one embodiment, obtaining the objective function for minimizing operating costs and the objective function for maximizing renewable energy consumption includes: obtaining a first parameter based on the cost coefficient corresponding to each target unit; obtaining a second parameter based on the start-up and shutdown cost value and the start-up and shutdown indication value corresponding to each target unit; determining the objective function for minimizing operating costs based on the first parameter, the second parameter, and the variable active power corresponding to each target unit; wherein the target units include renewable energy units and thermal power units; and determining the objective function for maximizing renewable energy consumption based on the maximum active power and variable active power corresponding to each renewable energy unit.
[0015] In one embodiment, obtaining the set of constraints includes: determining the active power constraints corresponding to the new energy unit, the thermal power unit, and the balanced unit based on the minimum and maximum active power corresponding to the new energy unit, the thermal power unit, and the balanced unit, respectively; and determining the ramp rate constraints corresponding to the thermal power unit based on the minimum ramp rate and the maximum ramp efficiency corresponding to the thermal power unit.
[0016] In one embodiment, before generating corresponding actions based on the states of the new energy generating units, thermal power units, and loads according to the DDPG algorithm, the strategy network further includes: determining a first rule guidance function based on the difference between the change in the variable active power of the target generating unit and the change in the variable active power of the load; wherein the first rule guidance function is a function used to characterize the degree of balance deviation between power generation and load power, and the target generating units include new energy generating units, thermal power units, and balancing units; determining the maximum upward adjustment amount of the thermal power unit based on the minimum value of the variable ramp rate and variable active power of the thermal power unit, and determining a second rule guidance function based on the maximum upward adjustment amount; wherein the second rule guidance function is a function used to characterize the reserve amount considered in response to uncertainties in the grid dispatching process; determining a reward signal based on the first rule guidance function and the second rule guidance function, and updating the strategy network based on the reward signal.
[0017] In one embodiment, the step of acquiring the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load respectively includes: acquiring the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load respectively through a unified communication protocol; wherein the unified communication protocol refers to the communication standard formulated for data interaction of different new energy generating units, thermal power units, and loads.
[0018] Secondly, this application also provides a power grid dispatching device based on the DDPG algorithm, comprising:
[0019] The first acquisition module is used to acquire the active power and corresponding control range of new energy units, the active power and corresponding control range of thermal power units, and the active power and corresponding control range of loads, respectively.
[0020] The second acquisition module is used to acquire the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; wherein the objective function of minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units;
[0021] The third acquisition module is used to acquire a set of constraints, which includes active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units.
[0022] The first calculation module is used to generate corresponding actions based on the policy network corresponding to the DDPG algorithm, according to the states of the new energy units, thermal power units, and loads respectively; generate corresponding expected return values based on the state and the actions according to the dual value network corresponding to the DDPG algorithm; update the dual value network based on the difference of the expected return values and obtain the target expected return value; and update the policy network based on the target expected return value.
[0023] The second calculation module is used to generate the maximum expected return value based on the updated policy network and the updated dual-value network, and to take the action corresponding to the maximum expected return value as the target action; wherein the target action refers to the optimal solution result of the objective function of minimizing operating costs and the objective function of maximizing new energy consumption under the set of constraints.
[0024] The adjustment module is used to adjust the active power of the new energy unit based on the target action, adjust the active power of the thermal power unit based on the corresponding control range, and adjust the active power of the load based on the corresponding control range.
[0025] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0026] The active power and corresponding control range of new energy units, active power and corresponding control range of thermal power units, and active power and corresponding control range of loads are obtained respectively.
[0027] Obtain the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; wherein the objective function of minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units;
[0028] Obtain a set of constraints, which includes active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units.
[0029] Based on the strategy network corresponding to the DDPG algorithm, corresponding actions are generated according to the states of the new energy units, thermal power units, and loads respectively. Based on the dual-value network corresponding to the DDPG algorithm, corresponding expected values of returns are generated according to the states and actions. The dual-value network is updated based on the difference of the expected values of returns to obtain the target expected value of returns. The strategy network is updated based on the target expected value of returns.
[0030] Based on the updated policy network and the updated dual-value network, the maximum expected return value is generated, and the action corresponding to the maximum expected return value is taken as the target action; wherein the target action refers to the optimal solution result of the objective function of minimizing operating costs and the objective function of maximizing new energy consumption under the set of constraints.
[0031] Based on the target action, the active power of the new energy generating units is adjusted according to the corresponding control range, the active power of the thermal power units is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range.
[0032] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0033] The active power and corresponding control range of new energy units, active power and corresponding control range of thermal power units, and active power and corresponding control range of loads are obtained respectively.
[0034] Obtain the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; wherein the objective function of minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units;
[0035] Obtain a set of constraints, which includes active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units.
[0036] Based on the strategy network corresponding to the DDPG algorithm, corresponding actions are generated according to the states of the new energy units, thermal power units, and loads respectively. Based on the dual-value network corresponding to the DDPG algorithm, corresponding expected values of returns are generated according to the states and actions. The dual-value network is updated based on the difference of the expected values of returns to obtain the target expected value of returns. The strategy network is updated based on the target expected value of returns.
[0037] Based on the updated policy network and the updated dual-value network, the maximum expected return value is generated, and the action corresponding to the maximum expected return value is taken as the target action; wherein the target action refers to the optimal solution result of the objective function of minimizing operating costs and the objective function of maximizing new energy consumption under the set of constraints.
[0038] Based on the target action, the active power of the new energy generating units is adjusted according to the corresponding control range, the active power of the thermal power units is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range.
[0039] The aforementioned grid dispatching method, device, computer equipment, storage medium, and computer program products based on the DDPG algorithm generate corresponding actions based on the strategy network corresponding to the DDPG algorithm and the states of generator sets and loads. Based on the dual-value network corresponding to the DDPG algorithm, they generate corresponding expected returns based on the states and actions. The dual-value network is updated based on the differences in expected returns to obtain a target expected return value. The strategy network is then updated based on the target expected return value. Based on the updated strategy network and the updated dual-value network, the maximum expected return value is generated, and the action corresponding to this maximum expected return value is taken as the target action. Based on this target action, the active power of generator sets and loads is adjusted accordingly. Therefore, by utilizing the dual-value network corresponding to the DDPG algorithm, the convergence and decision-making capabilities of the algorithm are improved. This allows for efficient and accurate determination of the optimal solution results for minimizing operating costs and maximizing renewable energy consumption under the constraint set within the DDPG algorithm, thereby enabling efficient and synchronous grid dispatching for minimizing operating costs and maximizing renewable energy consumption. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating a power grid dispatching method based on the DDPG algorithm in one embodiment;
[0042] Figure 2 This is a flowchart illustrating the process of updating the operation and maintenance status of power equipment in one embodiment;
[0043] Figure 3 This is a flowchart illustrating the process of determining the rule-guided function in one embodiment;
[0044] Figure 4 Here is a block diagram of the DDPG algorithm in one embodiment;
[0045] Figure 5 This is a structural block diagram of a power grid dispatching device based on the DDPG algorithm in one embodiment;
[0046] Figure 6 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0048] In one embodiment, such as Figure 1 As shown, a power grid dispatching method based on the DDPG algorithm is provided. This embodiment illustrates the method by applying it to a server. It is understood that this method can also be applied to terminals, and further to systems including terminals and servers, and implemented through interaction between the terminals and servers. In this embodiment, the method includes the following steps S102 to S112, wherein:
[0049] Step S102: Obtain the active power and corresponding control range of the new energy unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load.
[0050] Among them, new energy units refer to generator sets that generate electricity using renewable energy sources such as solar energy, wind energy, and bioenergy; thermal power units refer to generator sets that generate electricity using fossil fuels such as coal, gas, and oil; load refers to power-consuming equipment; and active power refers to the actual power generated by the generator set or the actual power consumed by the power-consuming equipment.
[0051] The active power control range refers to the adjustable range corresponding to the active power. It can be determined based on the standard specifications of the equipment, such as standard rated power, maximum rated power and other parameters; or it can be determined based on the actual operating conditions of the equipment, such as average power and other parameters.
[0052] Step S104: Obtain the objective function for minimizing operating costs and the objective function for maximizing renewable energy consumption; wherein the objective function for minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function for maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units.
[0053] The objective function for minimizing operating costs aims to minimize the overall operating cost of generator sets and loads by adjusting the active power of generator sets and loads. This involves a comprehensive consideration of power generation costs and load demand.
[0054] The objective function for maximizing renewable energy absorption is to maximize the absorption of renewable energy by adjusting the active power of renewable energy generating units. This means that the considerations are based on the generation, transmission, storage, and utilization of renewable energy.
[0055] Step S106: Obtain the set of constraints. The set of constraints includes the active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as the ramp rate constraints corresponding to thermal power units.
[0056] Among them, a balancing unit refers to a generator set used to regulate the power generation to maintain the power balance of the power system, that is, to maintain the balance between the generation and consumption of electrical energy.
[0057] Among them, the active power constraint condition refers to the restriction condition set for the magnitude of active power. The active power of each unit must meet the corresponding active power constraint condition.
[0058] The ramp rate refers to the rate of power change of a thermal power unit from one moment to another. The ramp rate constraint is a restriction imposed on the magnitude of the ramp rate; the ramp rate of a thermal power unit must satisfy the corresponding ramp rate constraint.
[0059] Step S108: Based on the strategy network corresponding to the DDPG algorithm, generate corresponding actions according to the states of new energy units, thermal power units, and loads respectively. Based on the dual-value network corresponding to the DDPG algorithm, generate corresponding expected return values according to the states and actions. Update the dual-value network based on the difference in expected return values and obtain the target expected return value. Update the strategy network based on the target expected return value.
[0060] Among them, the DDPG algorithm (Deep Deterministic Policy Gradient) is a deep learning algorithm applied to decision problems in a continuous action space; the policy network is a neural network used to represent the policy of choosing an action in a given state; and the dual-value network is a neural network used to represent the expected value of the reward obtained based on the policy of choosing an action in a given state.
[0061] The expected value of reward refers to the expected value of the cumulative reward obtained by taking a specific action in a given state. It is an estimate used to characterize the cumulative reward that can be obtained in the long run by taking a specific action.
[0062] For example, based on the policy network corresponding to the DDPG algorithm, corresponding actions are generated according to the states corresponding to the active power levels of new energy units, thermal power units, and loads, respectively. Based on the dual-value network corresponding to the DDPG algorithm, two expected return values are generated according to the state and the action. The parameters of the dual-value network are updated based on the difference between the two expected return values, and a target expected return value is obtained based on the two expected return values. The parameters of the policy network are updated based on the target expected return value.
[0063] Optionally, a dual-value network refers to two neural networks with the same input and network architecture but different initialization parameters. The two neural networks perform the same action in a given state to generate different expected returns. The minimum of the two expected returns can be used as the target expected return to reduce the variance of the estimated value and avoid overestimating the value of the action; alternatively, the average of the two expected returns can be used as the target expected return to improve estimation accuracy and reduce overestimation bias that might be introduced by a single estimate.
[0064] Optionally, the parameters of the dual-value network are updated based on the difference between the two expected returns. The goal is to minimize the difference between the two output expected returns. For example, a loss function can be set to measure the difference between the two expected returns. According to the gradient descent algorithm, the parameters of the dual-value network are adjusted to minimize the loss function, that is, to adjust the parameters in the direction of minimizing the difference between the two expected returns.
[0065] Optionally, the parameters of the policy network are updated based on the expected return of the target reward. The goal is to maximize the expected return of the output target reward, that is, to be more likely to choose an action that brings a higher expected return in a given state. For example, a loss function can be set to measure the magnitude of the expected return. According to the gradient ascent algorithm, the parameters of the policy network are adjusted to maximize the loss function, that is, to adjust the parameters in the direction of maximizing the expected return.
[0066] Step S110: Based on the updated policy network and the updated dual-value network, generate the maximum expected return value, and take the action corresponding to the maximum expected return value as the target action; where the target action refers to the optimal solution result of the objective function of minimizing operating cost and the objective function of maximizing new energy consumption under the constraint set.
[0067] For example, the updated policy network regenerates the corresponding action based on the given state, and the updated dual-value network regenerates the corresponding expected return value and target expected return value based on the state and the action. Based on the regenerated expected return value and target expected return value, the dual-value network and policy network are then updated.
[0068] Based on the continuously updated dual-value network and policy network, the maximum expected return value is finally obtained. The action corresponding to the maximum expected return value is taken as the target action. That is, the active power setpoint of the new energy unit, thermal power unit, and load corresponding to the target action is the optimal solution result for the objective function of minimizing operating cost and maximizing new energy consumption under the set of constraints.
[0069] Step S112: Based on the target action, adjust the active power of the new energy unit according to the corresponding control range, adjust the active power of the thermal power unit according to the corresponding control range, and adjust the active power of the load according to the corresponding control range.
[0070] For example, based on the active power setpoints of the new energy generating units, thermal power units, and loads corresponding to the target action, the current active power of the new energy generating units, thermal power units, and loads is adjusted accordingly within the control range.
[0071] In the aforementioned grid dispatching method based on the DDPG algorithm, actions are generated according to the states of generator units and loads based on the strategy network corresponding to the DDPG algorithm. A corresponding expected return value is generated based on the state and actions based on the dual-value network corresponding to the DDPG algorithm. The dual-value network is updated based on the difference in expected return values to obtain a target expected return value. The strategy network is then updated based on the target expected return value. The maximum expected return value is generated based on the updated strategy network and the updated dual-value network, and the action corresponding to this maximum expected return value is taken as the target action. Based on this target action, the active power of generator units and loads is adjusted accordingly. Therefore, by utilizing the dual-value network corresponding to the DDPG algorithm, the convergence and decision-making capabilities of the algorithm are improved. This allows for efficient and accurate determination of the optimal solution results for minimizing operating costs and maximizing renewable energy consumption under the constraint set within the DDPG algorithm, thereby enabling efficient and synchronous grid dispatching for minimizing operating costs and maximizing renewable energy consumption.
[0072] In an exemplary embodiment, the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load are obtained respectively, including step S202; based on the target action, the active power of the new energy generating unit is adjusted according to the corresponding control range, the active power of the thermal power unit is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range, including step S204, wherein:
[0073] Step S202: Obtain different acquisition times with the same time interval, sequentially use the different acquisition times as target acquisition times, and obtain the active power of the new energy unit and the corresponding control range, the active power of the thermal power unit and the corresponding control range, and the active power of the load and the corresponding control range corresponding to the target acquisition time.
[0074] Step S204: Based on the target action, adjust the active power of the new energy unit corresponding to the target acquisition time according to the corresponding control range, adjust the active power of the thermal power unit according to the corresponding control range, and adjust the active power of the load according to the corresponding control range.
[0075] For example, multiple data collection times with the same time interval are acquired. At each data collection time, the active power and control range of the corresponding renewable energy units, thermal power units, and loads are obtained. In the DDPG algorithm, the optimal solution results for minimizing the operating cost objective function and maximizing the renewable energy consumption objective function under the set of constraints are obtained. Based on the optimal solution results, the active power of the renewable energy units, thermal power units, and loads corresponding to the data collection time is adjusted. After a preset time interval, i.e., at the next data collection time, the above steps are repeated.
[0076] Optionally, the data collection period can be set to 24 hours, and the time interval between data collection times can be set to 15 minutes. That is, the active power data of new energy units, thermal power units, and loads can be acquired every 15 minutes, and the active power of new energy units, thermal power units, and loads can be updated based on the calculation results.
[0077] Optionally, historical data can be analyzed to identify periods that may have high power generation or high power consumption, and data collection can be scheduled more frequently during these periods.
[0078] In this embodiment, the active power of new energy units, thermal power units, and loads is adjusted periodically and continuously through periodic data acquisition tasks to ensure the safe operation of new energy units, thermal power units, and loads within the acquisition period.
[0079] In one exemplary embodiment, such as Figure 2 As shown, before obtaining the active power and corresponding control range of the new energy unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load, steps S302 to S308 are included, wherein:
[0080] Step S302: Obtain the current operation and maintenance status of new energy units, thermal power units and loads; where operation and maintenance status refers to operating status, health status and performance.
[0081] Operation and maintenance status refers to the operating status, health status, and performance of power equipment or systems.
[0082] Among them, operating status refers to the operating state of power equipment or system, such as the state of being turned on, off, or running. It can also represent the basic operating parameters of power equipment or system, such as current and voltage. Monitoring operating status helps to understand the working status of power equipment or system in real time.
[0083] Health status refers to the overall health condition of electrical equipment or systems, such as the health of mechanical structures, electrical components, cooling systems, etc. Monitoring health status helps identify potential problems in electrical equipment or systems, such as wear and tear, insulation condition, heat distribution, etc., so as to carry out preventive maintenance.
[0084] Among them, performance refers to the performance characteristics and performance of power equipment or systems, such as power conversion efficiency, transmission loss, response time, etc. Monitoring performance helps to evaluate the efficiency and reliability of power equipment or systems.
[0085] For example, obtaining the current operation and maintenance status of new energy units, thermal power units and loads refers to obtaining indicator parameters of new energy units, thermal power units and loads regarding their operating status, health status and performance.
[0086] Step S304: Calculate the expected return value corresponding to different actions selected under a given operation and maintenance state based on the dual Q network.
[0087] Among them, the dual Q network refers to an algorithm used for reinforcement learning, which is typically used to address the estimation bias problem in Q-learning.
[0088] For example, in a dual-Q network, two Q-networks are introduced. One Q-network evaluates the current expected return of choosing a specific action under a given operational state, while the other Q-network evaluates the target expected return of choosing that action under the same operational state. The target expected return remains constant over a period of time, which reduces the correlation between the current and target expected returns to some extent, thus improving the algorithm's stability. Based on the difference between the current and target expected returns, a loss function is constructed. By minimizing this loss function, the parameters of the dual-Q network are optimized. The optimized dual-Q network then outputs the expected return for each possible action under a given operational state.
[0089] Step S306: Convert each expected return value into the probability of choosing different actions under a given operation and maintenance state, and determine the optimal strategy based on the probability corresponding to each action; where the optimal strategy refers to choosing the action corresponding to the optimal expected return value under a given operation and maintenance state.
[0090] For example, each expected return value output by the dual Q network is converted into the probability of selecting the corresponding action under a given operational state. That is, the selected action is associated with the probability value, and the action corresponding to the highest probability is represented as the action selected based on the optimal expected return value under a given operational state.
[0091] Optionally, based on a deterministic policy, an action is understood as a behavioral policy through probability values, thereby achieving intelligent decision-making.
[0092] Optionally, the expected value of each reward output by the dual Q network can be converted into the probability of choosing the corresponding action under a given operational state through a probability distribution function.
[0093] Step S308: Update the operation and maintenance status of new energy units, thermal power units and loads based on the optimal strategy.
[0094] For example, based on the action corresponding to the highest probability, the operation and maintenance status of new energy units, thermal power units and loads are updated accordingly.
[0095] Optionally, the collection and update cycles for operation and maintenance status can be set, thereby periodically adjusting the operation and maintenance status of new energy units, thermal power units, and loads.
[0096] Optionally, in the monitoring of the operation and maintenance status of new energy units, thermal power units and loads, a multi-agent mechanism is introduced, treating the multi-dimensional equipment involved as intelligent agents with autonomous perception and decision-making capabilities. While updating their own operation and maintenance status, they actively share their own operation and maintenance status information with the outside world, thereby achieving a panoramic, autonomous and accurate reproduction of the operation and maintenance status of multi-dimensional equipment.
[0097] In this embodiment, the expected return value corresponding to different actions selected under a given operation and maintenance state is calculated, and each expected return value is converted into the probability of selecting different actions under a given operation and maintenance state. The optimal strategy is determined based on the probability value, thereby accurately and efficiently adjusting the operation and maintenance state of new energy units, thermal power units and loads.
[0098] In an exemplary embodiment, obtaining the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption includes steps S402 to S404, wherein:
[0099] Step S402: Based on the cost coefficient corresponding to each target unit, obtain the corresponding first parameter; based on the start-up and shutdown cost value and the start-up and shutdown indication value corresponding to each target unit, obtain the corresponding second parameter; based on the first parameter, the second parameter, and the variable active power corresponding to each target unit, determine the objective function for minimizing operating costs; wherein the target units include new energy units and thermal power units.
[0100] The cost coefficient refers to the weighting factor that measures the operating costs of different generator sets. The start-up and shutdown cost value is used to characterize the costs incurred when the unit is started or stopped. The start-up and shutdown indication value is used to characterize whether the unit changes from an on state to a off state or vice versa.
[0101] For example, the first parameter of the objective function for minimizing operating costs, namely the cost coefficient, is used to weight the variable active power of each unit; the second parameter of the objective function for minimizing operating costs, namely the parameter determined by the start-up and shutdown cost value and the start-up and shutdown indication value, is used to calculate the start-up and shutdown cost of each unit; the objective function for minimizing operating costs is formed based on the first parameter, the second parameter, and the variable active power corresponding to each unit.
[0102] Alternatively, the objective function for minimizing the running cost can be expressed as follows:
[0103]
[0104] In equation (1), J cost,t P represents the operating cost of n generator units at time t, where the generator units include new energy units and thermal power units. i,t Let a represent the active power of the i-th generator unit at time t. i b i c i This represents the cost coefficient of the i-th generator set; This indicates that the active power of the i-th generator unit is weighted. d represents the start-stop cost value, and I represents the start-stop indication value. If the generator unit changes from an on state to an off state or from an off state to an on state at the current time compared to the previous time, it means that the generator unit has performed a start-stop operation, and I = 1; otherwise, I = 0. dI represents the determination of start-stop operation and the calculation of start-stop cost for the i-th generator unit.
[0105] Optionally, different cost coefficients can be set for different generator sets based on their corresponding power levels, or different cost coefficients can be set for different generator sets at different times. Different start-up and shutdown cost values can also be set for different generator sets based on their corresponding operation and maintenance status.
[0106] Step S404: Based on the maximum active power and variable active power corresponding to each new energy unit, determine the objective function for maximizing the absorption of new energy.
[0107] The maximum active power is expressed as the maximum active power value specified in the specifications, or it can be expressed as the maximum active power value recorded during historical operation.
[0108] For example, the objective function for maximizing the absorption of new energy is determined based on the numerical relationship between the sum of the active power of all new energy generating units and the sum of the maximum active power of all new energy generating units.
[0109] Optionally, the expression for maximizing the objective function of renewable energy consumption is as follows:
[0110]
[0111] In equation (2), J re,t Represents n at time t re The amount of renewable energy consumed by Taiwan's renewable energy units. i,t Let represent the active power of the i-th generator unit at time t. This indicates that n at time t re The active power of the new energy generating units is summed. P i,max This represents the maximum active power of the i-th generator set. For n at time t re The maximum active power of the new energy generating units is summed.
[0112] Furthermore, based on the consideration of safely carrying current on the load line, a target function for minimizing the average current load rate is also set. Its goal is to minimize the current load rate to ensure the safe operation of the power grid, and its expression is as follows:
[0113]
[0114] In equation (3), J rho,t Represents n at time t rho The current load factor corresponding to each load line; I i,t I represents the current flowing through the i-th load line at time t. i,max This represents the maximum current flowing through the i-th load line. Indicates in Choose the minimum value from 1, that is, if I i,t >I i,max If the current load factor of the i-th load line at time t is 1, then the current load factor of the i-th load line at time t is 1; otherwise, the current load factor of the i-th load line at time t is 1. This means summing the current load rates of all load lines and converting them into the average current load rate of each load line.
[0115] Equations (1), (2), and (3) are standardized, for example, by zero-mean and unit variance, so that these equations have similar scales. The processing procedure can be referred to in the following equation:
[0116]
[0117] In equation (4), C′ represents the standardized objective function, and C represents the objective function. This represents the mean of the objective function. This represents the standard deviation of the objective function.
[0118] Based on equation (4), the standardized objective function K for minimizing operating costs is obtained. c ′ ost,t The standardized objective function for maximizing the absorption of new energy sources is J. r ′ e,t Minimize the average current load factor of the standardized objective function J r ′ ho,t And obtain the comprehensive objective function J:
[0119] J = w cost J c ′ ost,t -w re J r ′ e,t +w rho J r ′ ho,t (5)
[0120] In equation (5), w cost w re w rho The standardized objective function J is to minimize the operating cost. c ′ ost,t The standardized objective function for maximizing the absorption of new energy sources is J. r ′ e,t Minimize the average current load factor of the standardized objective function J r ′ ho,t The weighting coefficients are 1, and the sum of the three is 1.
[0121] In this embodiment, by establishing an objective function that minimizes operating costs and an objective function that maximizes renewable energy consumption, the comprehensiveness and integrity of power grid dispatch are improved in both dimensions of minimizing operating costs and maximizing renewable energy consumption.
[0122] In an exemplary embodiment, obtaining a set of constraints includes steps S502 to S504, wherein:
[0123] Step S502: Based on the minimum and maximum active power of the new energy units, thermal power units, and balancing units respectively, determine the active power constraints for the new energy units, thermal power units, and balancing units respectively.
[0124] For example, if at any given moment the sum of the active power of thermal power units, new energy units, and balancing units equals the sum of the active power of the load, then the total power balance constraint is:
[0125]
[0126] In equation (6), O th,i,t This represents the active power of the i-th thermal power unit at time t. Indicates n at time t th The sum of the active power of the thermal power units in Taiwan; P re,i,t This represents the active power of the i-th renewable energy unit at time t. Indicates n at time t re The sum of the active power of the new energy generating units; P bal,t P represents the active power of the balancing unit at time t; d,i,t This represents the active power of the i-th load at time t. Indicates n at time t d The sum of the active power of the load.
[0127] For any thermal power unit, the active power at any given time should be between its minimum and maximum active power. Therefore, the active power constraint condition for a thermal power unit is:
[0128] P th,i,min ≤P th,i,t ≤P th,i,max (7)
[0129] In equation (7), P th,i,min P represents the minimum active power of the i-th thermal power unit. th,i,max This represents the maximum active power of the i-th thermal power unit.
[0130] For any new energy generating unit, the active power at any given time should be between 0 and its maximum active power. Therefore, the active power constraint condition for new energy generating units is:
[0131] 0≤P re,i,t ≤P re,i,max (8)
[0132] In equation (8), P re,i,max This represents the maximum active power of the i-th renewable energy unit.
[0133] For balancing units, the range of active power values can be determined based on preset ratios for minimum and maximum active power. For example, the active power constraint condition for a balancing unit can be expressed as:
[0134] 0.9P bal,min ≤P bal,t ≤1.1P bal,max (9)
[0135] In equation (9), P bal,min P represents the minimum active power of the balancing unit. bal,max This indicates the maximum active power of the balancing unit.
[0136] Step S504: Based on the minimum ramp rate and maximum ramp efficiency of the thermal power unit, determine the ramp rate constraint conditions corresponding to the thermal power unit.
[0137] For example, the ramp rate constraint for a thermal power unit can be expressed as:
[0138] D th,i ≤P th,i,t -P th,i,t-1 ≤U th,i (10)
[0139] In equation (10), D th,i U represents the minimum gradient required for the i-th thermal power unit. th,i P represents the maximum gradient of the i-th thermal power unit. th,i,t -P th,i,t-1 This represents the upward adjustment of the active power of the i-th thermal power unit at time t compared to the active power at the previous time t-1.
[0140] In this embodiment, by setting active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units, the objective function can be solved reliably and accurately within the constraints.
[0141] In one exemplary embodiment, such as Figure 3 As shown, before generating corresponding actions based on the policy network corresponding to the DDPG algorithm according to the states of new energy units, thermal power units, and loads, steps S602 to S606 are also included, wherein:
[0142] Step S602: Based on the difference between the change in the variable active power of the target unit and the change in the variable active power of the load, determine the first rule guiding function; wherein the first rule guiding function is a function used to characterize the degree of balance deviation between power generation and load power, and the target unit includes new energy units, thermal power units and balanced units.
[0143] For example, based on the total power balance constraint condition corresponding to equation (6), a first rule guidance function is constructed. This first rule guidance function is used to guide the algorithm to update and calculate parameters in the direction of balancing the change in active power of the generator set and the change in active power of the load. Its expression is:
[0144]
[0145] In equation (11), δ MSE Represents the mean squared error function. ΔP bal,t , These represent the changes in active power of all thermal power units, all new energy power units, balanced power units, and all loads, respectively.
[0146] Step S604: Based on the minimum value of the variable ramp rate and variable active power of the thermal power unit, determine the maximum upward adjustment amount of the thermal power unit, and determine the second rule guiding function based on the maximum upward adjustment amount; wherein the second rule guiding function is a function used to characterize the reserve amount considered in response to the uncertainties in the grid dispatching process.
[0147] For example, under the combined constraints of the active power constraint of the thermal power unit corresponding to equation (7) and the ramp rate constraint corresponding to equation (10), the actual maximum upward adjustment of the thermal power unit is determined. Based on the actual maximum upward adjustment and the theoretical actual maximum upward adjustment, a second rule guidance function is constructed. This second rule guidance function is used to guide the algorithm to update parameters and perform calculations in the direction of reserving a sufficiently large reserve for the thermal power unit to cope with the uncertainties at the generation end and the load end. The expression is:
[0148]
[0149] In equation (12), E represents the weighting coefficient; U th,i,t U represents the actual maximum upward adjustment of the i-th thermal power unit at time t. Specifically, at time t, the i-th thermal power unit obtains the corresponding first active power under the corresponding active power constraint condition and the corresponding second active power under the corresponding ramp rate constraint condition. The minimum value between the first and second active power is selected as the actual maximum upward adjustment. th,i This represents the theoretical maximum upward adjustment amount for the i-th thermal power unit.
[0150] Step S606: Determine the reward signal based on the first rule guidance function and the second rule guidance function, and update the policy network based on the reward signal.
[0151] For example, based on the rule guidance directions corresponding to the first rule guidance function and the second rule guidance function, a corresponding reward signal is determined. This reward signal is a signal used to guide the network learning. Based on this reward signal, the parameters of the policy network are updated accordingly.
[0152] In this embodiment, a rule-guided function is set to guide the update direction of the policy network, thereby enabling targeted updates of the policy network in the expected update direction and improving the convergence and decision-making capabilities of the algorithm.
[0153] In an exemplary embodiment, the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load are obtained respectively, including step S702, wherein:
[0154] Step S702: Obtain the active power and corresponding control range of new energy generating units, the active power and corresponding control range of thermal power units, and the active power and corresponding control range of loads through a unified communication protocol; wherein the unified communication protocol refers to the communication standard formulated for data interaction of different new energy generating units, thermal power units, and loads.
[0155] For example, a unified communication protocol is selected for different new energy units, thermal power units, and loads, and all are configured with communication parameters that conform to the unified communication protocol, such as device address, communication rate, and data format; a communication connection is established based on the unified communication protocol, and unified data reception and parsing operations are performed to ensure that data processing and application are carried out in accordance with the manner defined by the unified communication protocol.
[0156] In this embodiment, a communication connection is established based on a unified communication protocol to ensure the communication compatibility and consistency of each device, thereby efficiently and accurately obtaining data corresponding to devices from different manufacturers and models.
[0157] In one exemplary embodiment, such as Figure 4 As shown, an environmental model is set up. An environmental model is a computational model used to describe or simulate a power system. That is, it models the power system and power equipment in the external environment, aiming to capture the dynamic changes of the external environment.
[0158] The environmental model transmits the active power data of new energy units, thermal power units, balancing units, and loads to the DDPG algorithm for calculation to obtain the optimal solution for the comprehensive objective function corresponding to equation (5). The optimal solution must simultaneously satisfy the total power balance constraint condition corresponding to equation (6), the active power constraint condition of thermal power units corresponding to equation (7), the active power constraint condition of new energy units corresponding to equation (8), the active power constraint condition of balancing units corresponding to equation (9), and the ramp rate constraint condition corresponding to equation (10).
[0159] In the DDPG algorithm, there are corresponding policy networks (actor) and value networks (critic). There are also target policy networks and target value networks with the same structure as the policy networks and value networks, respectively, to improve the performance of the algorithm.
[0160] The environment model will define the state s at time t. t The reward value r at time t tThe target state s at time t+1 t+1 The data is transmitted to the policy network and the experience replay pool. The experience replay pool is a buffer used to store historical data. Historical data can be selected from the experience replay pool to update the network, thereby improving data utilization.
[0161] The policy network is based on the first action value function u and the state s at time t. t Generate the corresponding action a t The data is then transmitted to the value network; the target policy network is based on the second action value function u′ and the target state s at time t+1. t+1 Generate the corresponding action, which is used to estimate the action a at the next time step. t+1 The goal.
[0162] Value networks are based on action a t Generate the corresponding expected return value and transmit it to the policy network; the target value network is based on action a. t Generate the corresponding expected return value, which is used to estimate the target y of the expected return value corresponding to the action selected in the next time step. t .
[0163] based on Figure 4 The algorithm framework shown replaces the value network with a dual-value network and the target value network with a target dual-value network. The dual-value network is based on action a. t Two expected returns are generated, and the parameters θ of the dual-value network are updated based on minimizing the difference between these two expected returns. Q The minimum of the two expected returns is taken as the target expected return, and the parameters θ of the policy network are updated based on this target expected return. u .
[0164] Based on the updated policy network and the updated dual-value network, the maximum expected return value is generated. The action corresponding to the maximum expected return value is taken as the target action, which is transmitted to the environment model, and the active power of the circuit equipment in the power system is adjusted accordingly.
[0165] Optionally, the parameters θ of the value network or dual value network Q The update method can be achieved by minimizing the loss function L. Q To achieve, where:
[0166] L Q =E[(y t -Q(s t ,a t |θ Q )) 2 (13)
[0167] In equation (13), E represents the expectation, and y t The target value represents the expected return, Q(s). t ,a t |θ Q ) represents the expected return output by the value network.
[0168] Among them, the target y of the expected return t It can be represented as:
[0169] y t =r t +γQ′(s t+1 ,u′(s t+1 |θ u′ )|θ Q′ (14)
[0170] In equation (14), r t Let θ represent the reward value at time t, γ represent the discount factor, Q′ represent the output of the target value network, u′ represent the output of the target policy network, and θ represent the reward value at time t. Q′ The parameters represent the target value network.
[0171] Optionally, the parameters θ of the policy network u The update method can be achieved by minimizing the loss function L. u To achieve, where:
[0172] L u =-E[Q(s) t ,u(s t (15)
[0173] In equation (15), E represents expectation, and Q represents the expected return value output by the value network.
[0174] based on Figure 4 The algorithm framework shown introduces the first rule guidance function corresponding to Equation (11) and the second rule guidance function corresponding to Equation (12) between the environment model and the policy network. By combining the first rule guidance function and the second rule guidance function, a loss function L′ for another update method of the policy network is obtained. u :
[0175] L′ u =-E[Q(s) t ,u(s t ))]+ω1F1+ω2F2 (16)
[0176] In equation (16), ω1 and ω2 are the weighting coefficients of the first rule guiding function F1 and the second rule guiding function F2, respectively.
[0177] DDPG is a model-free reinforcement learning method that can complete the learning process without requiring a specific expression for the state transition function.
[0178] Furthermore, the reward value r t It can be expressed as the negative form of the comprehensive objective function corresponding to equation (5):
[0179] r t =-J=-w cost J′ cost,t +w re J′ re,t -w rho J′ rho,t (17)
[0180] Based on this, the objective function can be transformed into a form that obtains the maximum reward by optimizing the decision function.
[0181] In one exemplary embodiment, a stability control visualization system is provided, which has a centralized monitoring function to comprehensively display environmental model information and power system information through a graphical interface and client tools.
[0182] Optionally, it can monitor the communication status between the power system and the stability control visualization system, as well as the channel status of communication between power equipment, and provide channel message monitoring and debugging tools.
[0183] Optionally, it can monitor the operating status, abnormal signals, and setpoints of power equipment, as well as the operating information such as the load that can be cut off, the number of units that can be cut off, the maximum DC lift, and the pullback amount.
[0184] Optionally, it can display the line status, unit status, and operating data of power equipment in real time.
[0185] Optionally, it can promptly detect changes in the setpoints of power equipment, proactively perform numerical detection, and promptly push the discrepancy information to management personnel; it can also set an automatic inspection cycle to automatically detect the setpoints of power equipment at regular intervals.
[0186] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0187] Based on the same inventive concept, this application also provides a DDPG-based power grid dispatching device for implementing the DDPG-based power grid dispatching method described above. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more DDPG-based power grid dispatching device embodiments provided below can be found in the limitations of the DDPG-based power grid dispatching method described above, and will not be repeated here.
[0188] In one exemplary embodiment, such as Figure 5 As shown, a power grid dispatching device based on the DDPG algorithm is provided, comprising: a first acquisition module 802, a first acquisition module 804, a first acquisition module 806, a first calculation module 808, a second calculation module 810, and an adjustment module 812, wherein:
[0189] The first acquisition module 802 is used to acquire the active power and corresponding control range of new energy units, the active power and corresponding control range of thermal power units, and the active power and corresponding control range of loads, respectively.
[0190] The second acquisition module 804 is used to acquire the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; wherein the objective function of minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units.
[0191] The third acquisition module 806 is used to acquire a set of constraints, which includes active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units.
[0192] The first calculation module 808 is used to generate corresponding actions based on the policy network corresponding to the DDPG algorithm, according to the states of new energy units, thermal power units, and loads respectively; generate corresponding expected return values based on the state and actions using the dual value network corresponding to the DDPG algorithm; update the dual value network based on the difference in expected return values to obtain the target expected return value; and update the policy network based on the target expected return value.
[0193] The second calculation module 810 is used to generate the maximum expected return value based on the updated policy network and the updated dual value network, and to take the action corresponding to the maximum expected return value as the target action; where the target action refers to the optimal solution result of the objective function of minimizing operating cost and the objective function of maximizing new energy consumption under the constraint set.
[0194] The adjustment module 812 is used to adjust the active power of the new energy unit based on the corresponding control range, the active power of the thermal power unit based on the corresponding control range, and the active power of the load based on the corresponding control range, based on the target action.
[0195] In an exemplary embodiment, the first acquisition module 802 is further configured to acquire different acquisition times with the same time interval, sequentially use the different acquisition times as target acquisition times, and acquire the active power of the new energy unit and the corresponding control range, the active power of the thermal power unit and the corresponding control range, and the active power of the load and the corresponding control range corresponding to the target acquisition time; the adjustment module 812 is further configured to adjust the active power of the new energy unit corresponding to the target acquisition time based on the corresponding control range, adjust the active power of the thermal power unit based on the corresponding control range, and adjust the active power of the load based on the corresponding control range, based on the target action.
[0196] In an exemplary embodiment, the device further includes an operation and maintenance status adjustment module, which is used to: obtain the current operation and maintenance status of the new energy generating units, thermal power units, and loads; wherein the operation and maintenance status refers to the operating status, health status, and performance; calculate the expected return value corresponding to different actions selected under a given operation and maintenance status based on a dual Q network; convert each expected return value into the probability of selecting different actions under a given operation and maintenance status, and determine the optimal strategy based on the probability corresponding to each action; wherein the optimal strategy refers to selecting the action corresponding to the optimal expected return value under a given operation and maintenance status; and update the operation and maintenance status of the new energy generating units, thermal power units, and loads based on the optimal strategy.
[0197] In an exemplary embodiment, the second acquisition module 804 is further configured to obtain a corresponding first parameter based on the cost coefficient corresponding to each target unit, obtain a corresponding second parameter based on the start-up and shutdown cost value and the start-up and shutdown indication value corresponding to each target unit, and determine a minimum operating cost objective function based on the first parameter, the second parameter, and the variable active power corresponding to each target unit; wherein the target units include new energy units and thermal power units; and determine a maximum new energy consumption objective function based on the maximum active power and variable active power corresponding to each new energy unit.
[0198] In an exemplary embodiment, the third acquisition module 806 is further configured to determine the active power constraints corresponding to the new energy unit, the thermal power unit, and the balancing unit based on the minimum active power and maximum active power corresponding to the new energy unit, the thermal power unit, and the balancing unit, respectively; and to determine the ramp rate constraints corresponding to the thermal power unit based on the minimum ramp rate and maximum ramp efficiency corresponding to the thermal power unit.
[0199] In an exemplary embodiment, the device further includes a rule-guided function determination module. This module determines a first rule-guided function based on the difference between the change in the variable active power of the target unit and the change in the variable active power of the load. The first rule-guided function characterizes the degree of imbalance between power generation and load power. The target unit includes renewable energy units, thermal power units, and balancing units. Based on the minimum value of the variable ramp rate and variable active power of the thermal power unit, the maximum upward adjustment amount of the thermal power unit is determined. A second rule-guided function is determined based on this maximum upward adjustment amount. The second rule-guided function characterizes the reserve amount considered for uncertainties in the grid dispatching process. A reward signal is determined based on the first and second rule-guided functions, and the strategy network is updated based on the reward signal.
[0200] In an exemplary embodiment, the first acquisition module 802 is further configured to acquire the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load through a unified communication protocol; wherein the unified communication protocol refers to the communication standard formulated for data interaction of different new energy generating units, thermal power units, and loads.
[0201] The modules in the aforementioned power grid dispatching device based on the DDPG algorithm can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0202] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores the active power and corresponding control ranges of new energy generating units, thermal power units, and loads, as well as objective functions, constraints, network parameters of the DDPG algorithm, and calculation data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a power grid dispatching method based on the DDPG algorithm.
[0203] Those skilled in the art will understand that Figure 6 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0204] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0205] The active power and corresponding control range of new energy units, active power and corresponding control range of thermal power units, and active power and corresponding control range of loads are obtained respectively.
[0206] Obtain the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; where the objective function of minimizing operating costs is used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is used to characterize the maximization of renewable energy consumption by renewable energy units;
[0207] Obtain the set of constraints, which includes the active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as the ramp rate constraints corresponding to thermal power units.
[0208] Based on the policy network corresponding to the DDPG algorithm, corresponding actions are generated according to the states of new energy units, thermal power units, and loads respectively. Based on the dual value network corresponding to the DDPG algorithm, corresponding expected return values are generated according to the states and actions. The dual value network is updated based on the difference in expected return values to obtain the target expected return value. The policy network is updated based on the target expected return value.
[0209] Based on the updated policy network and the updated dual-value network, the maximum expected return value is generated, and the action corresponding to the maximum expected return value is taken as the target action; where the target action refers to the optimal solution result of the objective function of minimizing operating cost and the objective function of maximizing new energy consumption under the constraint set;
[0210] Based on the target action, the active power of the new energy generating units is adjusted according to the corresponding control range, the active power of the thermal power units is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range.
[0211] In one embodiment, when the processor executes the computer program, it further implements the following steps: acquiring different acquisition times with the same time interval, sequentially using the different acquisition times as target acquisition times, acquiring the active power of the new energy generating units and their corresponding control ranges, the active power of the thermal power units and their corresponding control ranges, and the active power of the load and their corresponding control ranges corresponding to the target acquisition times; based on the target action, adjusting the active power of the new energy generating units corresponding to the target acquisition times based on their corresponding control ranges, adjusting the active power of the thermal power units based on their corresponding control ranges, and adjusting the active power of the load based on their corresponding control ranges.
[0212] In one embodiment, when the processor executes the computer program, it further performs the following steps: obtaining the current operation and maintenance status of the new energy generating units, thermal power units, and loads; wherein the operation and maintenance status refers to the operating condition, health status, and performance; calculating the expected return value corresponding to different actions selected under a given operation and maintenance status based on a dual-Q network; converting each expected return value into the probability of selecting different actions under a given operation and maintenance status, and determining the optimal strategy based on the probability corresponding to each action; wherein the optimal strategy refers to selecting the action corresponding to the optimal expected return value under a given operation and maintenance status; and updating the operation and maintenance status of the new energy generating units, thermal power units, and loads based on the optimal strategy.
[0213] In one embodiment, when the processor executes the computer program, it further performs the following steps: obtaining a corresponding first parameter based on the cost coefficient corresponding to each target unit; obtaining a corresponding second parameter based on the start-up and shutdown cost value and the start-up and shutdown indication value corresponding to each target unit; determining a minimum operating cost objective function based on the first parameter, the second parameter, and the variable active power corresponding to each target unit; wherein the target units include new energy units and thermal power units; and determining a maximum new energy consumption objective function based on the maximum active power and variable active power corresponding to each new energy unit.
[0214] In one embodiment, when the processor executes the computer program, it further implements the following steps: determining the active power constraints corresponding to the new energy unit, the thermal power unit, and the balanced unit based on the minimum active power and maximum active power corresponding to the new energy unit, the thermal power unit, and the balanced unit, respectively; and determining the ramp rate constraints corresponding to the thermal power unit based on the minimum ramp rate and maximum ramp efficiency corresponding to the thermal power unit.
[0215] In one embodiment, when the processor executes the computer program, it further performs the following steps: determining a first rule-guided function based on the difference between the variable change in active power of the target unit and the variable change in active power of the load; wherein the first rule-guided function is a function used to characterize the degree of imbalance between power generation and load power, and the target unit includes new energy units, thermal power units, and balanced units; determining the maximum upward adjustment amount of the thermal power unit based on the minimum value of the variable ramp rate and variable active power of the thermal power unit, and determining a second rule-guided function based on the maximum upward adjustment amount; wherein the second rule-guided function is a function used to characterize the reserve amount considered for uncertainties in the grid dispatching process; determining a reward signal based on the first rule-guided function and the second rule-guided function, and updating the strategy network based on the reward signal.
[0216] In one embodiment, when the processor executes the computer program, it further implements the following steps: obtaining the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load through a unified communication protocol; wherein the unified communication protocol refers to the communication standard formulated for data interaction of different new energy generating units, thermal power units, and loads.
[0217] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0218] The active power and corresponding control range of new energy units, active power and corresponding control range of thermal power units, and active power and corresponding control range of loads are obtained respectively.
[0219] Obtain the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; where the objective function of minimizing operating costs is used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is used to characterize the maximization of renewable energy consumption by renewable energy units;
[0220] Obtain the set of constraints, which includes the active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as the ramp rate constraints corresponding to thermal power units.
[0221] Based on the policy network corresponding to the DDPG algorithm, corresponding actions are generated according to the states of new energy units, thermal power units, and loads respectively. Based on the dual value network corresponding to the DDPG algorithm, corresponding expected return values are generated according to the states and actions. The dual value network is updated based on the difference in expected return values to obtain the target expected return value. The policy network is updated based on the target expected return value.
[0222] Based on the updated policy network and the updated dual-value network, the maximum expected return value is generated, and the action corresponding to the maximum expected return value is taken as the target action; where the target action refers to the optimal solution result of the objective function of minimizing operating cost and the objective function of maximizing new energy consumption under the constraint set;
[0223] Based on the target action, the active power of the new energy generating units is adjusted according to the corresponding control range, the active power of the thermal power units is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range.
[0224] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: acquiring different acquisition times with the same time interval, sequentially using the different acquisition times as target acquisition times, acquiring the active power of the new energy generating units and their corresponding control ranges, the active power of the thermal power units and their corresponding control ranges, and the active power of the load and their corresponding control ranges corresponding to the target acquisition times; based on the target action, adjusting the active power of the new energy generating units corresponding to the target acquisition times based on their corresponding control ranges, adjusting the active power of the thermal power units based on their corresponding control ranges, and adjusting the active power of the load based on their corresponding control ranges.
[0225] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: obtaining the current operation and maintenance status of the new energy generating units, thermal power units, and loads; wherein the operation and maintenance status refers to the operating condition, health status, and performance; calculating the expected return value corresponding to different actions selected under a given operation and maintenance status based on a dual Q network; converting each expected return value into the probability of selecting different actions under a given operation and maintenance status, and determining the optimal strategy based on the probability corresponding to each action; wherein the optimal strategy refers to selecting the action corresponding to the optimal expected return value under a given operation and maintenance status; and updating the operation and maintenance status of the new energy generating units, thermal power units, and loads based on the optimal strategy.
[0226] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: obtaining a first parameter based on the cost coefficient corresponding to each target unit; obtaining a second parameter based on the start-up and shutdown cost value and the start-up and shutdown indication value corresponding to each target unit; determining a minimum operating cost objective function based on the first parameter, the second parameter, and the variable active power corresponding to each target unit; wherein the target units include new energy units and thermal power units; and determining a maximum new energy consumption objective function based on the maximum active power and variable active power corresponding to each new energy unit.
[0227] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: determining the active power constraints corresponding to the new energy unit, the thermal power unit, and the balanced unit based on the minimum active power and maximum active power corresponding to the new energy unit, the thermal power unit, and the balanced unit, respectively; and determining the ramp rate constraints corresponding to the thermal power unit based on the minimum ramp rate and maximum ramp efficiency corresponding to the thermal power unit.
[0228] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: determining a first rule-guided function based on the difference between the change in the variable active power of the target unit and the change in the variable active power of the load; wherein the first rule-guided function is a function used to characterize the degree of imbalance between power generation and load power, and the target unit includes new energy units, thermal power units, and balancing units; determining the maximum upward adjustment amount of the thermal power unit based on the minimum value of the variable ramp rate and variable active power of the thermal power unit, and determining a second rule-guided function based on the maximum upward adjustment amount; wherein the second rule-guided function is a function used to characterize the reserve amount considered for uncertainties in the grid dispatching process; determining a reward signal based on the first rule-guided function and the second rule-guided function, and updating the strategy network based on the reward signal.
[0229] In one embodiment, when the computer program is executed by the processor, it further performs the following steps: obtaining the active power and corresponding control range of the new energy generating unit, the active power and corresponding control range of the thermal power unit, and the active power and corresponding control range of the load through a unified communication protocol; wherein the unified communication protocol refers to the communication standard formulated for data interaction of different new energy generating units, thermal power units, and loads.
[0230] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0231] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0232] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A power grid dispatching method based on the DDPG algorithm, characterized in that, The method includes: Obtain the current operation and maintenance status of new energy units, thermal power units, and loads; wherein the operation and maintenance status refers to the operating status, health status, and performance. Based on the dual-Q network, calculate the expected return value corresponding to different actions selected under a given operational state; Each expected return value is converted into the probability of choosing different actions under a given operational state. Based on the probability corresponding to each action, the optimal strategy is determined. The optimal strategy refers to choosing the action corresponding to the optimal expected return value under a given operational state. The operation and maintenance status of the new energy units, thermal power units, and loads is updated based on the optimal strategy. The active power and corresponding control range of new energy units, active power and corresponding control range of thermal power units, and active power and corresponding control range of loads are obtained respectively. Obtain the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; wherein the objective function of minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units; Obtain a set of constraints, which includes active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units. Based on the strategy network corresponding to the DDPG algorithm, corresponding actions are generated according to the states of the new energy units, thermal power units, and loads respectively. Based on the dual-value network corresponding to the DDPG algorithm, corresponding expected values of returns are generated according to the states and actions. The dual-value network is updated based on the difference of the expected values of returns to obtain the target expected value of returns. The strategy network is updated based on the target expected value of returns. Based on the updated policy network and the updated dual-value network, the maximum expected return value is generated, and the action corresponding to the maximum expected return value is taken as the target action; wherein the target action refers to the optimal solution result of the objective function of minimizing operating costs and the objective function of maximizing new energy consumption under the set of constraints. Based on the target action, the active power of the new energy generating units is adjusted according to the corresponding control range, the active power of the thermal power units is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range.
2. The method according to claim 1, characterized in that, The acquisition of the active power and corresponding control range of new energy generating units, the active power and corresponding control range of thermal power units, and the active power and corresponding control range of loads includes: Different acquisition times with the same time interval are obtained, and the different acquisition times are sequentially used as target acquisition times. The active power of the new energy unit and the corresponding control range, the active power of the thermal power unit and the corresponding control range, and the active power of the load and the corresponding control range are obtained at the target acquisition times. The adjustment of the active power of new energy generating units, the active power of thermal power units, and the active power of the load based on the target action, according to the corresponding control range, includes: Based on the target action, the active power of the new energy unit corresponding to the target acquisition time is adjusted according to the corresponding control range, the active power of the thermal power unit is adjusted according to the corresponding control range, and the active power of the load is adjusted according to the corresponding control range.
3. The method according to claim 1, characterized in that, The process of obtaining the objective function for minimizing operating costs and the objective function for maximizing renewable energy consumption includes: Based on the cost coefficient corresponding to each target unit, a first parameter is obtained. Based on the start-up and shutdown cost value and the start-up and shutdown indication value corresponding to each target unit, a second parameter is obtained. Based on the first parameter, the second parameter, and the variable active power corresponding to each target unit, the objective function for minimizing operating costs is determined. The target units include new energy units and thermal power units. Based on the maximum active power and variable active power corresponding to each new energy unit, the objective function for maximizing the absorption of new energy is determined.
4. The method according to claim 1, characterized in that, The set of constraints to be obtained includes: Based on the minimum and maximum active power of new energy units, thermal power units, and balanced units respectively, the active power constraints of new energy units, thermal power units, and balanced units are determined respectively. Based on the minimum and maximum ramp rates corresponding to thermal power units, the ramp rate constraints corresponding to thermal power units are determined.
5. The method according to claim 1, characterized in that, Before generating corresponding actions based on the states of the new energy units, thermal power units, and loads, the policy network based on the DDPG algorithm further includes: Based on the difference between the variable active power of the target unit and the variable active power of the load, a first rule guiding function is determined; wherein the first rule guiding function is a function used to characterize the degree of balance deviation between power generation and load power, and the target unit includes new energy units, thermal power units and balanced units; Based on the minimum value of the variable ramp rate and variable active power of the thermal power unit, the maximum upward adjustment amount of the thermal power unit is determined, and a second rule guiding function is determined based on the maximum upward adjustment amount; wherein the second rule guiding function is a function used to characterize the reserve amount considered in response to the uncertainties in the grid dispatching process. Based on the first rule guidance function and the second rule guidance function, a reward signal is determined, and the policy network is updated based on the reward signal.
6. The method according to claim 1, characterized in that, The acquisition of the active power and corresponding control range of new energy generating units, the active power and corresponding control range of thermal power units, and the active power and corresponding control range of loads includes: The active power and corresponding control range of new energy generating units, thermal power units, and loads are obtained through a unified communication protocol. The unified communication protocol refers to the communication standard formulated for data interaction of different new energy generating units, thermal power units, and loads.
7. A power grid dispatching device based on the DDPG algorithm, characterized in that, The device includes: The operation and maintenance status adjustment module is used to obtain the current operation and maintenance status of new energy units, thermal power units, and loads; wherein the operation and maintenance status refers to the operating status, health status, and performance; based on a dual Q network, it calculates the expected return value corresponding to different actions selected under a given operation and maintenance status; it converts each expected return value into the probability of selecting different actions under a given operation and maintenance status, and determines the optimal strategy based on the probability corresponding to each action; wherein the optimal strategy refers to selecting the action corresponding to the optimal expected return value under a given operation and maintenance status; and it updates the operation and maintenance status of the new energy units, thermal power units, and loads based on the optimal strategy. The first acquisition module is used to acquire the active power and corresponding control range of new energy units, the active power and corresponding control range of thermal power units, and the active power and corresponding control range of loads, respectively. The second acquisition module is used to acquire the objective function of minimizing operating costs and the objective function of maximizing renewable energy consumption; wherein the objective function of minimizing operating costs is a function used to characterize the comprehensive minimization of operating costs of renewable energy units, thermal power units, and loads, and the objective function of maximizing renewable energy consumption is a function used to characterize the maximization of renewable energy consumption by renewable energy units; The third acquisition module is used to acquire a set of constraints, which includes active power constraints corresponding to new energy units, thermal power units, and balanced units, as well as ramp rate constraints corresponding to thermal power units. The first calculation module is used to generate corresponding actions based on the policy network corresponding to the DDPG algorithm, according to the states of the new energy units, thermal power units, and loads respectively; generate corresponding expected return values based on the state and the actions according to the dual value network corresponding to the DDPG algorithm; update the dual value network based on the difference of the expected return values and obtain the target expected return value; and update the policy network based on the target expected return value. The second calculation module is used to generate the maximum expected return value based on the updated policy network and the updated dual-value network, and to take the action corresponding to the maximum expected return value as the target action; wherein the target action refers to the optimal solution result of the objective function of minimizing operating costs and the objective function of maximizing new energy consumption under the set of constraints. The adjustment module is used to adjust the active power of the new energy unit based on the target action, adjust the active power of the thermal power unit based on the corresponding control range, and adjust the active power of the load based on the corresponding control range.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.