A multi-type energy storage climbing power second-level dynamic optimization distribution method

By constructing multi-type energy storage ramp power response models and transforming them into Markov decision processes, and combining them with the PPO algorithm for training and solving, the problem of poor power allocation optimization for multi-type energy storage in existing technologies has been solved. This has enabled second-level dynamic optimization and precise tracking, thereby improving the safety and stability of the power system.

CN120073801BActive Publication Date: 2026-03-31HUAZHONG UNIV OF SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing research on deep reinforcement learning combined with power systems has failed to effectively consider the impact of the multi-physical coupling process inside energy storage on external characteristics, resulting in poor performance of power optimization allocation for various types of energy storage, especially at the second-level time scale, where it is difficult to achieve accurate tracking and coordination.

Method used

A multi-type energy storage ramping power response model is constructed, which is transformed into a Markov decision process. Reinforcement learning algorithms, especially the PPO algorithm, are used for training and solving to obtain multi-type energy storage ramping power allocation strategies, taking into account the regulation characteristics and complex operating conditions of various types of energy storage.

Benefits of technology

It achieves second-level dynamic optimization allocation of ramp power for various types of energy storage, improves the accuracy of power command tracking, can better cope with the complex operating conditions of A-CAES and wind-storage systems, generates more reasonable and applicable allocation strategies, and reduces power response deviation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073801B_ABST
    Figure CN120073801B_ABST
Patent Text Reader

Abstract

The application discloses a multi-type energy storage climbing power second-level dynamic optimization distribution method, belongs to the field of energy storage supporting power system climbing and intelligent regulation and control, and constructs a multi-type energy storage climbing power distribution model and converts the model into a Markov decision process, and a reinforcement learning algorithm is used for training and solving, so that the power is quickly and coordinately distributed; compared with a traditional method, the method provided by the application can coordinately distribute the climbing power according to the regulation characteristics of various types of energy storage, accurately track the power instruction, and better cope with the complex operation conditions of A-CAES and wind storage systems; in addition, the simplified environment generally used in the prior art can depict the regulation characteristics of the multi-type energy storage to a certain extent, but it is difficult to reflect the complex working conditions of the energy storage, the application effect of the "simplified" strategy generated by the training is poor, and there are limitations; the fine training environment provided by the application can promote the intelligent agent to learn a distribution strategy that is more reasonable and applicable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of energy storage supporting power system ramping and its intelligent regulation, and more specifically, relates to a second-level dynamic optimization allocation method for ramping power of multiple types of energy storage. Background Technology

[0002] With the widespread integration of new energy power generation into the power grid, ramp-up events caused by their rapid, short-term fluctuations are becoming increasingly frequent, posing a serious challenge to the safe and stable operation of the system. Against this backdrop, novel energy storage technologies, with their superior regulation performance, have gained widespread favor from academia and industry. However, problems such as "high growth but low utilization" and "unclear functional positioning" still exist and are becoming increasingly prominent, urgently requiring research into key technologies such as intelligent regulation and coordinated operation of novel energy storage for system ramp-up.

[0003] Existing research on grid-side energy storage participation in power system ramping mainly focuses on the conventional dispatch time scale. At the second-level time scale, research on power optimization allocation for various types of energy storage during ramping is relatively limited. Deep reinforcement learning (DRL), through neural networks to evaluate states and generate policies, possesses powerful self-learning and generalization capabilities and has been applied to solve problems with complex models and high speed requirements. However, current research on combining DRL with power systems only builds simple transfer function models of each unit within the load frequency control (LFC) framework as the training environment for the agent, without considering the impact of the multi-physical coupling processes within energy storage on external characteristics. Summary of the Invention

[0004] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a second-level dynamic optimization allocation method for the ramping power of multiple types of energy storage, which fully considers the fine-tuning characteristics of energy storage and can realize the second-level dynamic optimization allocation of ramping power of multiple types of energy storage.

[0005] To achieve the above objectives, according to a first aspect of the present invention, a method for dynamic optimization allocation of ramping power for multiple types of energy storage at the second level is provided, comprising:

[0006] S1, Establish ramp-up power response models for multiple types of energy storage; wherein, the multiple types of energy storage include A-CAES, wind farm combined electrochemical energy storage and thermal power unit combined flywheel energy storage;

[0007] S2, the multi-type energy storage ramping power response model is transformed into a Markov decision process;

[0008] Wherein, the state space corresponding to the Markov decision process ΔP tLet t be the system unbalanced power at time step t. The maximum uphill ramp capability of each unit or energy storage at time step t, with subscripts 1:N a This indicates that the subscript takes values ​​from 1, 2, ..., N in sequence. a N a =3, where 1 represents a fire storage system, 2 represents A-CAES, and 3 represents a wind storage system. The maximum downhill ramp capability of each unit or energy storage at time step t. The actual output of each unit or energy storage at time step t. Let ω be the state of charge of each unit or energy storage at time step t, and ω be the wind turbine speed; the action space of the intelligent agent is a. t ={a 1,t ,a 2,t ,a 3,t}, and the action space satisfies the constraints: a 1,t ,a 2,t ,a 3,t These are the power allocation factors for the fire-storage system, A-CAES, and wind-storage system, respectively; the agent's reward function is r1 = c1(ΔP). TF +ΔP CAES +ΔP WT c1 is a constant, ΔP TF ,ΔP CAES ,ΔP WT These are the power response deviations for thermal storage, A-CAES, and wind storage, respectively.

[0009] S3, The Markov decision process is trained and solved using a reinforcement learning algorithm to obtain multi-type energy storage ramping power allocation strategies.

[0010] According to a second aspect of the present invention, an electronic device is provided, comprising: a computer-readable storage medium and a processor;

[0011] The computer-readable storage medium is used to store executable instructions;

[0012] The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in the first aspect.

[0013] According to a third aspect of the invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to perform the method as described in the first aspect.

[0014] According to a fourth aspect of the invention, a computer program product is provided, comprising a computer program or instructions that, when executed by a processor, implement the method described in the first aspect.

[0015] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0016] This invention provides a second-level dynamic optimization allocation method for ramping power of various types of energy storage. It constructs a ramping power allocation model for multiple types of energy storage and converts it into a Markov decision process, then uses a reinforcement learning algorithm for training and solving, achieving rapid and coordinated power allocation. Compared to traditional methods, the method provided by this invention can coordinate ramping power allocation based on the regulation characteristics of various types of energy storage, achieving precise tracking of power commands. It can also better handle the complex operating conditions of A-CAES and wind-storage systems. Furthermore, while existing technologies typically use simplified environments that can characterize the regulation characteristics of various types of energy storage to some extent, they are insufficient to reflect complex energy storage operating conditions. The "simplified" strategies generated through such training have poor application effects and limitations. The refined training environment provided by this invention enables the agent to learn more reasonable and applicable allocation strategies.

[0017] As a further preferred option, the reinforcement learning algorithm used in this invention is the PPO algorithm, which can improve the stability of training, achieve a better balance in policy exploration and utilization, and result in higher reward values.

[0018] As a further preferred embodiment, the present invention introduces training mechanisms such as dynamic learning rate decay and reward scaling during the training process, which can improve the training effect of the proximal policy optimization algorithm to a certain extent, and compared with other reinforcement learning algorithms, it can find a better balance between environment exploration and experience utilization. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the process for a second-level dynamic optimization allocation method for multi-type energy storage ramping power provided in an embodiment of the present invention;

[0020] Figure 2 The strategy and evaluation network structure diagram provided in the embodiments of the present invention;

[0021] Figure 3 A complete flowchart of the training process for the near-end policy optimization algorithm provided in this embodiment of the invention;

[0022] Figure 4 A schematic diagram illustrating the training process of the reinforcement learning algorithm for a training scenario provided in an embodiment of the present invention;

[0023] Figure 5 In the above, (a) and (b) are schematic diagrams of the allocation factors of the traditional strategy and the PPO strategy in the upward climbing scenario 1 provided by the embodiment of the present invention, respectively.

[0024] Figure 6(a) and (b) in the above are schematic diagrams of the climbing process of multi-element energy storage under the traditional strategy and the PPO strategy in the upward climbing of scenario 1 provided by the embodiment of the present invention, respectively.

[0025] Figure 7 In the above, (a) and (b) are schematic diagrams of allocation factors for the traditional strategy and the PPO strategy in the downward climbing process provided in the embodiments of the present invention, respectively.

[0026] Figure 8 (a) and (b) in the above are schematic diagrams of the ramping process of multi-element energy storage under the traditional strategy and the PPO strategy in the downward ramping of scenario 1 provided by the embodiment of the present invention.

[0027] Figure 9 A schematic diagram of the discontinuous adjustment range of A-CAES in scenario 2 provided by an embodiment of the present invention;

[0028] Figure 10 In the figures (a) and (b), respectively, the discontinuous adjustment interval for A-CAES in scenario 2 provided by the embodiment of the present invention, the allocation factor of the traditional strategy, and the performance of the ramping process are shown in the figure.

[0029] Figure 11 This is a schematic diagram of the downhill climbing process of multi-element energy storage under two strategies in scenario 2 provided by the embodiment of the present invention.

[0030] Figure 12 (a) and (b) in the figure are schematic diagrams of the allocation factor and power response of the PPO allocation strategy in scenario 2 provided by the embodiment of the present invention, respectively.

[0031] Figure 13 In the above, (a) and (b) are schematic diagrams of allocation factors for the traditional allocation strategy and the PPO allocation strategy when the climbing ability is insufficient in scenario 3 provided by the embodiment of the present invention.

[0032] Figure 14 (a) and (b) in the figure are schematic diagrams of the wind turbine speed and wind turbine output under the traditional strategy when the climbing ability is insufficient in scenario 3 provided by the embodiment of the present invention;

[0033] Figure 15 (a) and (b) in the figure are schematic diagrams of the power command of the traditional strategy, the PPO strategy and the power response of the multi-energy storage when the ramping ability of scenario 3 provided by the embodiment of the present invention is insufficient;

[0034] Figure 16 A schematic diagram of the linear model of the A-CAES and wind-storage system provided in the embodiments of the present invention.

[0035] Figure 17In the above, (a) and (b) are schematic diagrams of the allocation factor and power response of the simplified strategy in scenario 4 provided by the embodiment of the present invention, which have sufficient ramp capacity.

[0036] Figure 18 In the figures (a) and (b), respectively, the A-CAES discontinuous adjustment interval of scenario 4 provided by the embodiment of the present invention, the allocation factor of the simplified strategy, and the climbing process are schematic diagrams.

[0037] Figure 19 In the figures (a) and (b), respectively, the allocation factor and wind turbine output of the simplified strategy for scenario 4 provided by the embodiment of the present invention are shown. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0039] Different types of energy storage systems exhibit significant differences in performance, such as ramp rate, capacity, and duration. To achieve optimized power allocation during power ramping and fully leverage the synergistic and complementary potential of various energy storage systems, it is essential to clarify the performance differences of each energy storage type and characterize its power regulation characteristics.

[0040] Based on this, the present invention considers the regulation characteristics of multiple types of energy storage and proposes a fast and coordinated ramp power allocation strategy. Specifically, embodiments of the present invention provide a second-level dynamic optimization allocation method for ramp power of multiple types of energy storage, including:

[0041] S1. Establish ramp-up power response models for multiple types of energy storage; wherein, the multiple types of energy storage include A-CAES, wind farm combined electrochemical energy storage, and thermal power unit combined flywheel energy storage.

[0042] There are two main construction forms for grid-side energy storage: large-scale energy storage power stations operating independently and energy storage power stations integrated with other power sources. This invention, combining the mainstream trends in energy storage development, selects A-CAES (Advanced Compressed Air Energy Storage), wind farm-integrated electrochemical energy storage, and thermal power unit-integrated flywheel energy storage for modeling. That is, it constructs multi-type energy storage ramp-up power response models covering large-scale advanced adiabatic compressed air energy storage systems, wind farm-integrated electrochemical energy storage systems, and thermal power unit-integrated flywheel energy storage systems.

[0043] The power response model of a large-scale advanced adiabatic compressed air energy storage system is as follows:

[0044] (1) Expander

[0045]

[0046]

[0047] In the formula, n t P represents the number of stages in the expander. tm,i Let be the shaft power of the i-th stage expander. T is the air mass flow rate of the expander. td,i Let T be the inlet gas temperature of the i-th stage expander. tx,i Let π be the outlet gas temperature of the i-th stage expander. t,i Let k = c be the expansion ratio of the i-th stage expander. p / c v The specific heat ratio, also known as the adiabatic index, is c. p and c v These are the isobaric specific heat and isovolastic specific heat of air, respectively, η ti It is isentropic efficiency.

[0048] (2) Heat exchanger

[0049] T LTF,x =(1-ε)T LTF,d +εT HTF,d

[0050] T HTF,x =(1-ε)T HTF,d +εT LTF,d

[0051] In the formula, T LTF,x and T LTF,d T represents the inlet and outlet temperatures of the cold working fluid, respectively. HTF,x and T HTF,d These represent the inlet and outlet temperatures of the heat exchanger, respectively, and ε is the heat exchanger performance parameter.

[0052] (3) Compressor

[0053]

[0054] In the formula, T cd,j and T cx,j Let π be the inlet and outlet gas temperatures of the j-th stage compressor. c =p cx,j / p cd,j Let P be the compression ratio of the j-th stage compressor. cm n is the power supplied by the electric motor to the compressor. c The number of stages in the compressor. This refers to the air mass flow rate of the compressor.

[0055] (4) Thermal storage device

[0056]

[0057] In the formula, m HS0 T represents the initial mass of the thermal storage medium in the thermal storage device. HS (t) represents the temperature at time t in the thermal storage device. and T HTF,c,j (t) represents the mass flow rate and temperature of the heat storage medium in the heat exchanger after the j-th stage compressor at time t. and T HTF,t,i (t) represents the mass flow rate and temperature of the heat storage medium in the heat exchanger before the i-th stage expander at time t.

[0058] (5) Gas storage chamber

[0059]

[0060] In the formula, p s V s ,T s These represent the pressure, volume, and temperature of the gas storage chamber, respectively. s0 R is the initial pressure of the gas storage chamber, R is the gas constant, and T is the initial pressure of the gas storage chamber. s,d This refers to the temperature of the gas entering the storage chamber from the compressor.

[0061] The power response model for wind farms combined with electrochemical energy storage is as follows:

[0062] (1) Pneumatic module

[0063] P m =0.5ρπR 2 ν 3 C p (λ,β)

[0064]

[0065] In the formula, ρ, R, and v represent air density, rotor radius, and wind speed, respectively, and C P (λ,β) represents the wind energy utilization coefficient, where λ=ωR / v is the tip speed ratio, ω is the rotor speed, β is the blade pitch angle, and P m For mechanical power.

[0066] (2) Generator Module

[0067]

[0068] In the formula, P e P is the output power of the converter. setIt is the set power of the converter, T a is the converter time constant, and J is the rotor's moment of inertia.

[0069] (3) Based on speed regulation control strategy

[0070] P ref =0.5C p (θ=0,λ opt,θ=0 )ρπR 2 v0 3

[0071]

[0072] In the formula, v0 corresponds to the grid power command P. ref The reference wind speed. When ω ≥ ω up,lim When the wind speed is low, the pitch mechanism activates to prevent the turbine from overspeeding; when the wind speed is low, the turbine decelerates, and the rotor releases its stored kinetic energy to compensate for the power deficiency. However, when ω ≤ ω down,lim At this time, the wind turbine switches to maximum power point tracking mode to prevent it from continuing to decelerate. opt,θ=0 This represents the optimal tip speed ratio. ω up,lim ω down,lim These are the upper and lower limits of the wind turbine speed, respectively.

[0073] The power response model of the combined fire and energy storage system is as follows:

[0074]

[0075] In the formula, P d and P c These represent the maximum discharge power and maximum charging power of the flywheel energy storage, respectively. o For the rated power, K, P, r, and b are all constants.

[0076] S2, transform the multi-type energy storage ramping power response model into a Markov decision process.

[0077] The main elements of a Markov decision process (MDP) include the environment, state space, and action space.

[0078] (1) Environment

[0079] The environment is primarily responsible for determining the next state based on the current state and actions; this effect can be achieved through the model in step A. Additionally, the environment should provide reward feedback to the agent based on the reward function. The goal of this invention's hill-climb power optimization allocation is to quickly respond to the system's hill-climb power commands. Therefore, the reward function r... t The design takes into account the deviation between the power command and the actual power response of the unit:

[0080] r1=c1(ΔP TF +ΔP CAES +ΔP WT )

[0081] In the formula, c1 is a constant used to control the size of the reward; ΔP TF ,ΔP CAES ,ΔP WT The power response deviations for thermal power generation, A-CAES, and wind power generation, respectively, can be calculated from the unit power command and the actual unit power.

[0082] ΔP=|P ord -P|

[0083] (2) State space

[0084] state space s t Set to:

[0085]

[0086] In the formula, ΔP t Let t be the system unbalanced power at time step t. The maximum uphill ramp capacity of each unit or energy storage at time step t (1 represents the thermal energy storage system, 2 represents A-CAES, and 3 represents the wind energy storage system). The maximum downhill ramp capability of each unit or energy storage at time step t. The actual output of each unit or energy storage at time step t. ω represents the state of charge of the multi-element energy storage (1 represents flywheel energy storage, 2 represents A-CAES, and 3 represents electrochemical energy storage) at time step t, and ω represents the wind turbine rotation speed.

[0087] Subscript 1:N a This indicates that the subscript takes values ​​from 1, 2, ..., N in sequence. a N a =3 represents the dimension of the action space, for example That is, to represent three variables P 1,t P 2,t P 3,t , which represent the output power of the fire storage system, A-CAES, and wind storage system at time step t, respectively.

[0088] (3) Action space

[0089] a t ={a 1,t ,a 2,t ,a 3,t}

[0090] In the formula, a 1,t ,a2,t ,a 3,t These are the power allocation factors for the fire storage system, A-CAES, and wind storage system, respectively; the action space of the PPO agent should also satisfy the following constraints:

[0091]

[0092] S3, The Markov decision process is trained and solved using a reinforcement learning algorithm to obtain multi-type energy storage ramping power allocation strategies.

[0093] The reinforcement learning algorithm used in step S3 can be any reinforcement learning algorithm, such as DDPG, TD3, SAC, etc.

[0094] Furthermore, considering that DDPG and TD3 are both deterministic policy algorithms, their stable reward values ​​are similar, and their reward curves fluctuate little, but their exploration of the action space is insufficient, making them prone to convergence to local optima. The SAC algorithm is a probabilistic policy-based algorithm, affected by policy randomness, resulting in slower convergence and greater volatility. In contrast, while the PPO algorithm is also a probabilistic policy-based algorithm, it introduces probability ratio pruning into the objective function, which not only improves training stability but also achieves a better balance between policy exploration and utilization, resulting in higher reward values. Therefore, preferably, in step S3, the reinforcement learning algorithm used is the PPO algorithm. That is, preferably, a proximal policy optimization algorithm is used to obtain multi-type energy storage ramping power allocation strategies.

[0095] When designing the policy network, a common Gaussian distribution can be used. Furthermore, considering the non-negativity constraint of the power allocation factor, preferably, the output action is sampled using a Beta distribution, so the output of the policy network is the α and β parameters of the Beta distribution.

[0096] To ensure that the output α and β meet the requirements, the Softplus activation function is used to process the output of the policy network.

[0097] Furthermore, to avoid the activation function outputting zero on negative inputs, thus preventing "deactivation," tanh is used instead of ReLU as the pooling function. The network's output is evaluated as an estimate of the state value V(s), which is then used to calculate the generalized dominance estimate.

[0098] It is understandable that the input to both the policy network and the evaluation network is the state at time step t. The specific structures of the policy and evaluation networks are as follows: Figure 2 As shown.

[0099] like Figure 3 As shown, the training process for the above Markov decision process using the Proximal Policy Optimization (PPO) algorithm is as follows:

[0100] (1) Experience Acquisition: The agent interacts with the environment through a policy network. When the experience buffer is full, the interaction stops, and the agent calculates the experience at each time step in the experience buffer. V target ;

[0101]

[0102]

[0103] In the formula, For the estimation of the dominance function at time step t, r t Let γ be the instantaneous reward at time step t, λ be the discount parameter, λ be the trade-off parameter, V be the value function, and δ be the instantaneous reward at time step t. t Let be the temporal difference error at time step t (the difference between the agent's estimate of the value of the current state and its estimate based on the actual immediate reward and the value of the next state), where T is the time step at the end of the round. This indicates that before updating the evaluation network, the old evaluation network is used to estimate the value at s. t The state value under the given state.

[0104] (2) Network Update: Randomly shuffle the experience data and divide it into multiple mini-batches. Use the experience data in each mini-batch to estimate the policy, evaluate the network's loss value, and update the network once. Repeat the above process K. epoch Second-rate;

[0105]

[0106] In the formula, ε and c1 are both constants, and π θ For a policy network with θ as a parameter, a t ,s t These represent the action and state at time step t, respectively, and π. θ (a t |s t The meaning of ) is based on strategy π. θ In state s t Choose action a t The probability, H(π) θ (·|s t )) represents the policy entropy, N a Let a be the dimension of the action space. i For the i-th dimension action, V ω (s) represents the state value estimated by the current evaluation network, V. target (s) is the target value of the state value, which can be calculated based on GAE. target (s), r t (θ) is the cutting ratio coefficient.

[0107] Furthermore, this invention introduces dynamic decay of the learning rate and policy entropy H(π) θ (·|s t State normalization, to some extent, improves the training performance of the PPO algorithm.

[0108]

[0109]

[0110] In the formula, α0, α t Here, S represents the initial learning rate and the learning rate at time step t, respectively, where k is the decay rate controlling the learning rate. t ,S t,norm The current state value and the normalized current state value are given by μ and σ, respectively. μ and σ are the mean and standard deviation of the state values, and they are continuously updated during the training process.

[0111] (3) Clear the experience buffer, repeat (1) and (2), and stop when the maximum number of training steps is reached. The complete training process of step C above is as follows: Figure 2 As shown.

[0112] The method provided by this invention will be verified by simulation below.

[0113] (1) Training scenario

[0114] Step A: The training scenario parameters used in this invention are shown in Table 1. The training round time is set to 900 seconds; the sampling interval of the agent is set to 5 seconds.

[0115] Table 1

[0116]

[0117] The model parameters are shown in Table 2:

[0118] Table 2

[0119]

[0120]

[0121] Step B: A comparison of the training process using four reinforcement learning algorithms: DDPG, TD3, SAC, and PPO (used in this invention). Algorithm parameter settings are shown in Table 3.

[0122] Table 3

[0123]

[0124] Step C: Training results are as follows Figure 4As shown, DDPG and TD3 are both deterministic policy algorithms, with similar reward values ​​after stabilization and small fluctuations in their reward curves. However, they do not explore the action space sufficiently and are prone to converging to local optima. The SAC algorithm is a probabilistic policy-based algorithm, which is affected by policy randomness, resulting in slower convergence and greater volatility. In contrast, although the PPO algorithm is also a probabilistic policy-based algorithm, it improves training stability and achieves a better balance between policy exploration and utilization by introducing probability ratio pruning in the objective function, resulting in higher reward values. In addition, this invention improves the training effect of the PPO algorithm to a certain extent by introducing dynamic learning rate decay, policy entropy, and state normalization.

[0125] (2) Scene 1

[0126] Step A: Scene Parameter Settings

[0127] This embodiment analyzes the agent's decision-making under conditions of ample climbing capacity, comparing it with traditional methods. Scenario parameters are shown in Table 4.

[0128] Table 4

[0129]

[0130] Step B: Results Analysis

[0131] This invention analyzes the two strategies from the perspectives of changes in allocation factors and the power response process of multi-element energy storage. Figure 5 The changes in the allocation factors of multi-element energy storage during the ramping process, Figure 5 (a) in the text represents the traditional strategy. Figure 5 (b) in the table represents the PPO allocation strategy (the same applies below). Figure 5 It can be seen that the traditional strategy allocates power according to the proportion of ramp-up capacity. Initially, the ramp-up capacity of thermal power units is high, so the initial allocation factor is the largest. However, due to the fast response speed of the wind-storage system, it continuously fills the power response deviation caused by the response delay of the thermal-storage system and A-CAES, and its allocation factor gradually increases, and the ramp-up power it bears also increases. In the later stage of the ramp-up process, as the ramp-up power gradually increases, the available ramp-up capacity of the wind-storage system is exhausted, and the allocation factor decreases. The allocation factors of the thermal-storage system and A-CAES increase accordingly, eventually stabilizing. After stabilization, the thermal-storage system bears the main ramp-up power. In the PPO allocation strategy, the initial system ramp-up rate demand is high, and the allocation factor of the wind-storage system is the largest. However, as the ramp-up rate demand weakens, the ramp-up power gradually increases, and the allocation factors of A-CAES and thermal-storage system gradually increase, eventually stabilizing, with A-CAES bearing the main ramp-up power.

[0132] Figure 6This diagram illustrates the power response of multi-element energy storage systems under two allocation strategies in an "uphill scenario." The solid and dashed lines represent the actual power and power command of the multi-element energy storage system, respectively. The diagonal lines indicate the power response deviation, and the upper gray diagonal line indicates the total power response deviation. Figure 6 It is evident that the PPO strategy results in a smaller power response deviation, reduced by approximately 31.3% compared to the traditional strategy. Furthermore, the wind-storage system exhibits better regulation performance and can accurately track power commands, while thermal power units have poorer regulation performance. Excessive deployment of these units can lead to significant power response deviations. The PPO allocation strategy allocates a large amount of ramp-up power to A-CAES, fully leveraging the advantages of A-CAES's large capacity and low cost while effectively avoiding the ramp-up disadvantages of thermal power units, thereby reducing power response deviation.

[0133] Figure 7 and Figure 8 The distribution factors and power responses of the traditional distribution strategy and the PPO distribution strategy in the "downhill climbing scenario" are shown. They are basically similar to those in the "uphill climbing scenario" and will not be described in detail here.

[0134] (3) Scene 2

[0135] Step A: Set scene parameters as shown in Table 5.

[0136] Table 5

[0137]

[0138] Step B: Scenario Result Analysis

[0139] The A-CAES unit needs to undergo a switching process from compressed energy storage to expanded power generation, resulting in discontinuous adjustment ranges, such as... Figure 9 As shown, the allocation of climbing power should, as far as possible, avoid the influence of this discontinuous adjustment range. Therefore, this section analyzes the handling of this discontinuous adjustment range using two different strategies.

[0140] Traditional allocation strategies handle discontinuous adjustment intervals as follows: Figure 10 As shown. At approximately 310 seconds, A-CAES enters the discontinuous adjustment range, and its allocation factor increases significantly; and the total power P of the multi-element energy storage... Σ An upward step, significantly exceeding the power command P Σ,ord The operating status of A-CAES around this time is as follows: Figure 11 As shown, compression stops immediately after reaching the minimum compression power, and it can no longer track the power command P. A-CAES,ord Furthermore, due to the minimum expansion power limitation, it is difficult to switch operating conditions and remains in a shutdown state, greatly reducing the grid's flexible adjustment capability. Under the PPO allocation strategy, the allocation factor and power response of multi-energy storage are as follows: Figure 12As shown in the figure, the PPO allocation strategy differs from that in scenario 1, significantly reducing the allocation factor of A-CAES, thereby avoiding discontinuous adjustment intervals and achieving better overall power deficit control.

[0141] (4) Scene 3

[0142] Step A: Scene parameter settings are shown in Table 6.

[0143] Table 6

[0144]

[0145] Step B: Scenario Result Analysis

[0146] Depend on Figure 13 and Figure 14 It can be seen that under the traditional allocation strategy, the rotational speed ω of the wind-storage system decreases at approximately 200 seconds and 800 seconds, reaching the lower limit ω. down,lim At this point, the energy stored in the impeller is insufficient, and even relying on the electrochemical energy storage system cannot compensate for the mechanical power P. m The sharp drop in power caused the allocation factor to decrease rapidly and significantly, ultimately resulting in the wind-storage system being unable to track the grid power command P. e In contrast, the PPO allocation strategy gradually reduces reliance on the wind-storage system during the ramp-up process, increasing the allocation factors of the thermal-storage system and A-CAES. This avoids significant power deficits caused by insufficient ramp-up capability of the wind-storage system due to wind speed changes, effectively ensuring the safe and reliable operation of the system. Furthermore, as... Figure 15 As shown, the PPO allocation strategy has a lower power response deviation, is less affected by wind speed changes, and is more adaptable in this scenario.

[0147] (5) Scene 4

[0148] Step A: As Figure 16 As shown, the linear model (i.e., simplified model) of the A-CAES wind-storage system is as follows:

[0149]

[0150] In the formula, v in To cut off the wind speed, v out To cut off the wind speed, v o For the rated wind speed, P wo This refers to the rated power of the fan.

[0151] Step B: Scenario Result Analysis

[0152] contrast Figure 17 and Figure 5 (b) Figure 6As shown in (b) of the diagram, in scenario 1, the allocation trends of the "simplified" strategy and the "detailed" strategy are similar, but the error of the "simplified" strategy is 11.2% higher than that of the "detailed" strategy. This indicates that the simplified environment can characterize the regulation characteristics of multiple types of energy storage to some extent, but there are still differences in details, which leads to the slightly poor adaptability of the "simplified" strategy.

[0153] contrast Figure 18 and Figure 12 As can be seen, in scenario 2, the simplified environment cannot characterize the discontinuous adjustment range of A-CAES, and therefore tends to allocate a larger ramp power to A-CAES. At approximately 200 seconds, A-CAES enters the discontinuous adjustment range, causing a large power response deviation in the "simplified" strategy.

[0154] Furthermore, simplified wind-storage systems fail to consider the actual control conditions of wind turbines, resulting in more aggressive allocation strategies for wind-storage systems. For example... Figure 19 As shown, in Scenario 3, the "simplification" strategy caused the wind-storage system to fail to track grid power commands twice, resulting in significant fluctuations during the ramp-up process. Figure 15 The "detailed" strategy in the text can effectively avoid this problem and keep the power response deviation within a small range.

[0155] As can be seen from the above analysis, the simplified model has problems such as incomplete description of complex operating conditions, insufficient consideration of actual operation, and insufficient characterization of regulation characteristics. It is difficult to cope with the discontinuous regulation range of A-CAES and the large fluctuations in the output of the wind-storage system. The detailed model can solve the above problems, and the PPO allocation strategy obtained by training the detailed model is more flexible and accurate.

[0156] This invention provides an electronic device, including: a computer-readable storage medium and a processor;

[0157] The computer-readable storage medium is used to store executable instructions;

[0158] The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in any of the above embodiments.

[0159] This invention provides a computer-readable storage medium storing computer instructions that cause a processor to perform the method described in any of the above embodiments.

[0160] This invention provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the method described in any of the above embodiments.

[0161] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A multi-type energy storage hill power second-level dynamic optimization allocation method, characterized in that, The method comprises the following steps: S1, a multi-type energy storage climbing power response model is established; wherein the multi-type energy storage comprises advanced compressed air energy storage (A-CAES), wind farm combined with electrochemical energy storage, and thermal power unit combined with flywheel energy storage; S2, the multi-type energy storage climbing power response model is converted into a Markov decision process; wherein the state space corresponding to the Markov decision process , is the system unbalanced power at time step is the maximum upward ramping capability of each unit or energy storage at time step denotes the subscript takes 1, 2,..., in turn N a , N a = 3, 1 represents a fire storage system, 2 represents A-CAES, and 3 represents a wind storage system, is the maximum downward ramping capability of each unit or energy storage at time step is the actual output of each unit or energy storage at time step is the state of charge of each unit or energy storage at time step is the wind wheel speed; the action space of the agent , and the action space satisfies the constraint: , ; are respectively the power allocation factors of the fire storage system, A-CAES, and wind storage system; the reward function of the agent , is a constant, , , are respectively the power response deviations of the fire storage, A-CAES, and wind storage; S3, a reinforcement learning algorithm is used to train and solve the Markov decision process to obtain a multi-type energy storage climbing power distribution strategy.

2. The method of claim 1, wherein, The power response model of the A-CAES is: ; ; wherein, is the shaft power of the th stage expander, is the air mass flow of the expander, is the inlet gas temperature of the th stage expander, is the outlet gas temperature of the th stage expander, is the specific heat at constant pressure of air; is the power provided by the electric motor to the compressor, is the air mass flow of the compressor, and is the inlet and outlet gas temperature of the th stage compressor, is the number of stages of the compressor; The power response model of the wind farm combined with electrochemical energy storage is: wherein is the output power of the converter, is the set power of the converter, is the time constant of the converter, is the moment of inertia of the rotor, s is a complex frequency domain variable; The power response model of the thermal power unit combined with flywheel energy storage is: wherein, and are the maximum discharge power and the maximum charge power of the flywheel energy storage, respectively, is the rated power, are constants, is the maximum allowed state of charge, is the minimum allowed state of charge, is the state of charge.

3. The method of claim 1 or 2, wherein, In step S3, the reinforcement learning algorithm used is a PPO algorithm.

4. The method of claim 3, wherein, The output of the policy network is a Beta distribution and parameter.

5. The method of claim 1, wherein, During training, the time step is the learning rate at time t ; and state normalization is performed according to the formula ; wherein, , are the initial learning rate, the learning rate at time step , is the decay rate of the learning rate, , are the current state value and the normalized current state value, respectively, , are the mean and standard deviation of the state value, respectively.

6. An electronic device, comprising: The computer readable storage medium and the processor are included. The computer readable storage medium is used to store executable instructions. The processor is used to read the executable instructions stored in the computer readable storage medium and execute the method according to any one of claims 1-5. The computer readable storage medium stores computer instructions for causing the processor to execute the method according to any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The computer program or instructions, when executed by the processor, implement the method according to any one of claims 1-5.

8. A computer program product comprising computer programs or instructions, characterized in that, ​

Citation Information

Patent Citations

  • Wind-light-storage combined power generation optimization method and system based on deep reinforcement learning

    CN117175591A

  • Multi-energy virtual power plant optimization scheduling method and system

    CN118801386A