Second-level dynamic optimal distribution method for multi-type energy storage climbing power
By establishing a climbing power response model for multiple types of energy storage and applying reinforcement learning algorithms, the problem that the existing technology is difficult to optimize the power distribution of multiple types of energy storage climbing under the second-level time scale is solved, and the second-level dynamic optimization distribution of multiple types of energy storage climbing power is realized, improving the safety, stability and adaptability of the system.
Patent Information
- Application Number
- CN202510101012.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to effectively optimize the climbing power distribution of multiple types of energy storage in power systems under the second-level time scale, and does not fully consider the multi-physical coupling process within the energy storage.
By establishing a multi-type energy storage climbing power response model and converting it into a Markov decision-making process, reinforcement learning algorithms, especially PPO algorithms, are used to train and solve them to achieve second-level dynamic optimization allocation of multi-type energy storage climbing power.
It realizes the rapid and coordinated distribution of power of multiple types of energy storage climbing, and can accurately track power instructions, reduce power response deviations, adapt to complex operating conditions, and improve the safety and stability of the system.
Smart Images

Figure CN120073801A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of energy storage supporting the ramp of power systems and their intelligent regulation, and more specifically, relates to a method for second-level dynamic optimal allocation of ramp power of multi-type energy storage. Background Art
[0002] With the wide access of new energy power generation to the power grid, ramp events caused by its short-term large-scale changes are becoming increasingly frequent, posing severe challenges to the safe and stable operation of the system. In this context, new energy storage has been widely favored by the academic and industrial circles due to its excellent regulation performance. However, problems such as "high growth and low utilization" and "unclear functional positioning" still exist and are becoming increasingly prominent, and it is urgent to carry out key technology research on the intelligent regulation and coordinated operation of new energy storage participating in system ramps.
[0003] Existing research on grid-side energy storage participating in the ramp of power systems mainly focuses on the conventional scheduling time scale. For the second-level time scale, there is relatively little research on the optimal allocation of ramp power of multi-type energy storage. Deep reinforcement learning (DRL) evaluates states and generates strategies through neural networks, and has powerful self-learning and generalization abilities, and has been applied to solve problems with complex models and high speed requirements. However, existing research on the combination of deep reinforcement learning and power systems only builds a simple transfer function model of each unit as the training environment of the intelligent agent in the load frequency control (LFC) framework, and does not consider the influence of the multi-physical coupling process inside the energy storage on the external characteristics. Summary of the Invention
[0004] In view of the above defects or improvement requirements of the prior art, the present invention provides a method for second-level dynamic optimal allocation of ramp power of multi-type energy storage, which fully considers the refined regulation characteristics of energy storage and can realize the second-level dynamic optimal allocation of ramp power of multi-type energy storage.
[0005] To achieve the above object, according to the first aspect of the present invention, a method for second-level dynamic optimal allocation of ramp power of multi-type energy storage is provided, including:
[0006] S1, establishing a ramp power response model of multi-type energy storage; wherein, the multi-type energy storage includes A-CAES, wind farm combined with electrochemical energy storage, and thermal power unit combined with flywheel energy storage;
[0007] S2, converting the ramp power response model of multi-type energy storage into a Markov decision process;
[0008] Wherein, the state space corresponding to the Markov decision process ΔP tis the system unbalanced power at time step t, is the maximum upward ramping capacity of each unit or energy storage at time step t, with subscripts 1:N a indicates that the subscript takes values 1, 2,..., N in sequence a , N a = 3, 1 represents the thermal - energy storage system, 2 represents A - CAES, 3 represents the wind - energy storage system, is the maximum downward ramping capacity of each unit or energy storage at time step t, is the actual output of each unit or energy storage at time step t, is the state of charge of each unit or energy storage at time step t, ω is the wind turbine speed; the action space a of the agent t ={a 1,t , a 2,t , a 3,t}, and the action space satisfies the constraints: a 1,t , a 2,t , a 3,t are the power distribution factors of the thermal - energy storage system, A - CAES, and wind - energy storage system respectively; the reward function r of the agent 1 = c 1 (ΔP TF +ΔP CAES +ΔP WT ), c 1 is a constant, ΔP TF , ΔP CAES , ΔP WT are the power response deviations of the thermal - energy storage, A - CAES, and wind - energy storage respectively;
[0009] S3. The Markov decision process is trained and solved using a reinforcement learning algorithm to obtain a multi - type energy storage ramping power distribution strategy.
[0010] According to the second aspect of the present invention, there is provided an electronic device, including: a computer - readable storage medium and a processor;
[0011] The computer - readable storage medium is used to store executable instructions;
[0012] The processor is used to read the executable instructions stored in the computer - readable storage medium and execute the method as described in the first aspect.
[0013] According to the third aspect of the present invention, there is provided a computer - readable storage medium, where the computer - readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to execute the method as described in the first aspect.
[0014] According to the fourth aspect of the present invention, there is provided a computer program product, including a computer program or instruction, which when executed by a processor implements the method described in the first aspect.
[0015] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0016] The multi-type energy storage ramp power second-level dynamic optimization allocation method provided by the present invention constructs a multi-type energy storage ramp power allocation model and converts it into a Markov decision process, and uses a reinforcement learning algorithm for training and solving, realizing fast and coordinated allocation of power; compared with traditional methods, the method provided by the present invention can coordinately allocate ramp power according to the regulation characteristics of various energy storages, realize accurate tracking of power commands, and can better cope with the complex operating conditions of A-CAES and wind energy storage systems; in addition, the simplified environment usually adopted by the prior art can depict the regulation characteristics of multi-type energy storage to a certain extent, but it is difficult to reflect the complex working conditions of energy storage, and the application effect of the "simplified" strategy generated by its training is not good and there are limitations. The fine training environment provided by the present invention can prompt the intelligent agent to learn a more reasonable and applicable allocation strategy.
[0017] As a further preference, the reinforcement learning algorithm adopted by the present invention is the PPO algorithm, which can improve the stability of training, achieve a better balance in the exploration and utilization of strategies, and has a higher reward value.
[0018] As a further preference, the present invention introduces training mechanisms such as dynamic decay of the learning rate and reward scaling during the training process, which can improve the training effect of the proximal policy optimization algorithm to a certain extent, and can find a better balance point in environment exploration and experience utilization compared with other reinforcement learning algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flow chart of the multi-type energy storage ramp power second-level dynamic optimization allocation method provided by an embodiment of the present invention;
[0020] Figure 2 It is a structural diagram of the policy and evaluation network provided by an embodiment of the present invention;
[0021] Figure 3 It is a complete flow chart of the proximal policy optimization algorithm training provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic diagram of the training process of the reinforcement learning algorithm for the training scenario provided by an embodiment of the present invention;
[0023] Figure 5(a) and (b) in it are respectively the schematic diagrams of the distribution factors of the traditional strategy and the PPO strategy during the upward ramp of Scenario 1 provided by the embodiments of the present invention;
[0024] Figure 6 (a) and (b) in it are respectively the schematic diagrams of the ramp-up processes of the multi-energy storage under the traditional strategy and the PPO strategy during the upward ramp of Scenario 1 provided by the embodiments of the present invention;
[0025] Figure 7 (a) and (b) in it are respectively the schematic diagrams of the distribution factors of the traditional strategy and the PPO strategy during the downward ramp;
[0026] Figure 8 (a) and (b) in it are respectively the schematic diagrams of the ramp-down processes of the multi-energy storage under the traditional strategy and the PPO strategy in Scenario 1 provided by the embodiments of the present invention;
[0027] Figure 9 It is the schematic diagram of the discontinuous adjustment interval of A-CAES in Scenario 2 provided by the embodiments of the present invention;
[0028] Figure 10 (a) and (b) in it are respectively the schematic diagrams of the distribution factor and the ramp process performance of the traditional strategy for the discontinuous adjustment interval of A-CAES in Scenario 2 provided by the embodiments of the present invention;
[0029] Figure 11 It is the schematic diagram of the ramp-down processes of the multi-energy storage under two strategies in Scenario 2 provided by the embodiments of the present invention;
[0030] Figure 12 (a) and (b) in it are respectively the schematic diagrams of the distribution factor and the power response of the PPO distribution strategy in Scenario 2 provided by the embodiments of the present invention;
[0031] Figure 13 (a) and (b) in it are respectively the schematic diagrams of the distribution factors of the traditional distribution strategy and the PPO distribution strategy when the ramp-up ability is insufficient in Scenario 3 provided by the embodiments of the present invention;
[0032] Figure 14 (a) and (b) in it are respectively the schematic diagrams of the fan speed and the fan output under the traditional strategy when the ramp-up ability is insufficient in Scenario 3 provided by the embodiments of the present invention;
[0033] Figure 15 (a) and (b) in it are respectively the schematic diagrams of the power commands and the power responses of the multi-energy storage under the traditional strategy and the PPO strategy when the ramp-up ability is insufficient in Scenario 3 provided by the embodiments of the present invention;
[0034] Figure 16Schematic diagram of the linear model of the A-CAES and wind energy storage system provided by the embodiments of the present invention
[0035] Figure 17 In (a) and (b) of the figure, they are respectively the schematic diagrams of the allocation factor and power response of the simplified strategy with sufficient ramping capacity in Scenario 4 provided by the embodiments of the present invention;
[0036] Figure 18 In (a) and (b) of the figure, they are respectively the schematic diagrams of the allocation factor and ramping process of the simplified strategy in the discontinuous adjustment interval of A-CAES in Scenario 4 provided by the embodiments of the present invention;
[0037] Figure 19 In (a) and (b) of the figure, they are respectively the schematic diagrams of the allocation factor and wind turbine output of the simplified strategy with insufficient ramping capacity in Scenario 4 provided by the embodiments of the present invention. Detailed implementation manners
[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0039] There are significant differences in the performance of different types of energy storage systems in terms of ramping rate, capacity and duration. In order to realize the optimal allocation of the ramping power of the power system and fully explore the synergistic and complementary potential of multiple energy storages, it is necessary to clarify the performance differences of each energy storage type and characterize its power regulation characteristics.
[0040] Based on this, the present invention considers the regulation characteristics of multiple types of energy storages and proposes a fast and coordinated ramping power allocation strategy. Specifically, the embodiments of the present invention provide a multi-type energy storage ramping power second-level dynamic optimization allocation method, including:
[0041] S1, establishing a multi-type energy storage ramping power response model; wherein, the multi-type energy storage includes A-CAES, wind farm combined with electrochemical energy storage, and thermal power unit combined with flywheel energy storage.
[0042] There are mainly two construction forms for grid-side energy storage: large-scale energy storage power stations operating independently and built-in energy storage power stations combined with other power sources. Combining with the mainstream trend of energy storage development, this invention selects A-CAES (Advanced Compressed Air Energy Storage), combined wind farm and electrochemical energy storage, and combined thermal power unit and flywheel energy storage for modeling. That is, a multi-type energy storage ramp power response model covering large-scale advanced adiabatic compressed air energy storage system, combined wind farm and electrochemical energy storage system, and combined thermal power unit and flywheel energy storage system is constructed.
[0043] The power response model of the large-scale advanced adiabatic compressed air energy storage system is as follows:
[0044] (1) Expander
[0045]
[0046]
[0047] In the formula, n t is the number of stages of the expander, P tm,i is the shaft power of the i-th stage expander, is the air mass flow rate of the expander, T td,i is the inlet gas temperature of the i-th stage expander, T tx,i is the outlet gas temperature of the i-th stage expander, π t,i is the expansion ratio of the i-th stage expander, k = c p / c v is the specific heat ratio, also known as the adiabatic index, c p and c v are the specific heat at constant pressure and specific heat at constant volume of air respectively, η ti is the isentropic efficiency.
[0048] (2) Heat exchanger
[0049] T LTF,x =(1 - ε)T LTF,d + εT HTF,d
[0050] T HTF,x =(1 - ε)T HTF,d + εT LTF,d
[0051] In the formula, T LTF,x and T LTF,d represent the inlet and outlet temperatures of the cold working medium respectively, T HTF,x and T HTF,d represent the inlet and outlet temperatures of the hot working medium respectively, and ε is the heat exchanger effectiveness parameter.
[0052] (3) Compressor
[0053]
[0054] Wherein, T cd,j and T cx,j are the inlet and outlet gas temperatures of the j - th stage compressor, π c = p cx,j / p cd,j is the compression ratio of the j - th stage compressor, P cm is the power supplied by the motor to the compressor, n c is the number of stages of the compressor, is the air mass flow rate of the compressor.
[0055] (4) Heat storage device
[0056]
[0057] Wherein, m HS0 is the initial mass of the heat storage medium in the heat storage device, T HS (t) is the temperature of the heat storage device at time t, and T HTF,c,j (t) are the mass flow rate and temperature of the heat storage medium of the heat exchanger after the j - th stage compressor at time t, and T HTF,t,i (t) are the mass flow rate and temperature of the heat storage medium of the heat exchanger before the i - th stage expander at time t.
[0058] (5) Gas storage chamber
[0059]
[0060] Wherein, p s , V s , T s are the pressure, volume and temperature of the gas storage chamber respectively, p s0 is the initial pressure of the gas storage chamber, R is the gas constant, T s,d is the temperature of the gas entering the gas storage chamber from the compressor.
[0061] The power response model of the wind farm combined with electrochemical energy storage is as follows:
[0062] (1) Aerodynamic module
[0063] P m = 0.5ρπR 2 ν 3 C p (λ,β)
[0064]
[0065] Where ρ, R, and v are the air density, wind turbine radius, and wind speed respectively, and C P (λ,β) is the wind energy utilization coefficient, where λ = ωR / v is the tip speed ratio, ω is the wind turbine rotational speed, β is the pitch angle, and P m is the mechanical power.
[0066] (2) Generator module
[0067]
[0068] Where P e is the converter output power, P set is the set power of the converter, T a is the converter time constant, and J is the moment of inertia of the rotor.
[0069] (3) Control strategy based on speed regulation
[0070] P ref = 0.5C p (θ = 0, λ opt,θ=0 )ρπR 2 v 0 3
[0071]
[0072] Where v 0 is the reference wind speed corresponding to the grid power command P ref . When ω ≥ ω up,lim , the pitch mechanism starts to prevent the wind turbine from overspeed; when the wind speed is low, the wind turbine decelerates, and the rotor releases the stored kinetic energy to make up for the power shortage. However, when ω ≤ ω down,lim , the wind turbine switches to the maximum power tracking mode to prevent the wind turbine from continuing to decelerate. λ opt,θ=0 is the optimal tip speed ratio. ω up,lim , ω down,lim are the upper and lower limits of the wind turbine rotational speed respectively.
[0073] The power response model of the thermal energy storage combined system is as follows:
[0074]
[0075] Where P d and P c are the maximum discharge power and maximum charge power of the flywheel energy storage respectively, P o is the rated power, and K, P, r, b are all constants.
[0076] S2, convert the multi-type energy storage ramp power response model into a Markov decision process.
[0077] The main elements of a Markov decision process (MDP) include the environment, the state space, and the action space.
[0078] (1) Environment
[0079] The environment is mainly responsible for determining the next state based on the current state and action, and this effect can be achieved through the model in step A. Additionally, the environment should provide reward feedback to the agent according to the reward function. The goal of the ramp power optimization allocation in this invention is to quickly respond to the system ramp power command. Therefore, the reward function r t is designed considering the deviation between the power command and the actual power response of the unit:
[0080] r 1 = c 1 (ΔP TF + ΔP CAES + ΔP WT )
[0081] where c 1 is a constant used to control the magnitude of the reward; ΔP TF , ΔP CAES , ΔP WT are the power response deviations of the thermal energy storage, A-CAES, and wind energy storage respectively, which can be calculated from the unit power command and the actual power of the unit:
[0082] ΔP = |P ord - P|
[0083] (2) State space
[0084] The state space s t is set as:
[0085]
[0086] where ΔP t is the system imbalance power at time step t, is the maximum upward ramp capacity of each unit or energy storage at time step t (1 represents the thermal energy storage system, 2 represents A-CAES, 3 represents the wind energy storage system), is the maximum downward ramp capacity of each unit or energy storage at time step t, is the actual output of each unit or energy storage at time step t, is the state of charge of the multi-energy storage (1 represents flywheel energy storage, 2 represents A-CAES, 3 represents electrochemical energy storage) at time step t, and ω is the wind turbine speed.
[0087] The subscript 1:N in a means that the subscript takes the values 1, 2..., N in sequencea , N a = 3 is the dimension of the action space. For example it represents three variables P 1,t , P 2,t , P 3,t , which are the output powers of the thermal energy storage system, A-CAES, and wind energy storage system at time step t, respectively.
[0088] (3) Action space
[0089] a t = {a 1,t , a 2,t , a 3,t}
[0090] In the formula, a 1,t , a 2,t , a 3,t are the power distribution factors of the thermal energy storage system, A-CAES, and wind energy storage system, respectively; the action space of the PPO agent should also satisfy the following constraints:
[0091]
[0092] S3. Use a reinforcement learning algorithm to train and solve the Markov decision process to obtain a multi-type energy storage ramp power distribution strategy.
[0093] The reinforcement learning algorithm used in step S3 can be any reinforcement learning algorithm, such as DDPG, TD3, SAC, etc.
[0094] Furthermore, considering that both DDPG and TD3 are deterministic policy algorithms, the stabilized reward values are similar, the fluctuations of their reward curves are small, but the exploration of the action space is not sufficient enough, and it is easy to converge to a local optimum. The SAC algorithm is a probability-based policy algorithm. Affected by the randomness of the policy, its convergence speed is slow and the volatility is large; in contrast, although the PPO algorithm belongs to a probability-based policy algorithm, due to the introduction of probability ratio clipping in the objective function, it not only improves the stability of training, but also achieves a good balance between policy exploration and exploitation, and the reward value is higher. Based on this, preferably, in step S3, the reinforcement learning algorithm used is the PPO algorithm. That is, preferably, the proximal policy optimization algorithm is used to obtain a multi-type energy storage ramp power distribution strategy
[0095] When designing the policy network, a common Gaussian distribution can be used. Further, considering the non-negativity constraint of the power distribution factor, preferably, the output action is sampled using the Beta distribution. Therefore, the output of the policy network is the α and β parameters of the Beta distribution.
[0096] To ensure that the output α and β meet the requirements, the Softplus activation function is used to process the output of the policy network.
[0097] In addition, to avoid the output of the activation function being 0 for negative inputs, resulting in the "deactivation" phenomenon, tanh is used instead of relu as the pooling function. The output of the evaluation network is an estimate of the state value V(s), which is used to calculate the generalized advantage estimate.
[0098] It can be understood that the inputs of both the policy network and the evaluation network are the state at time step t. The specific structures of the policy and evaluation networks are as Figure 2 shown.
[0099] As Figure 3 shown, the process of training the above Markov decision process using the Proximal Policy Optimization algorithm (PPO) is as follows:
[0100] (1) Collecting experience: The agent interacts with the environment through the policy network. When the experience buffer is full, the interaction stops, and the V target ; at each time step in the experience buffer is calculated.
[0101]
[0102]
[0103] In the formula, is the estimated advantage function at time step t, r t is the immediate reward at time step t, γ is the discount parameter, λ is the trade-off parameter, V is the value function, δ t is the temporal difference error at time step t (the difference between the agent's value estimate of the current state and the value estimate based on the actual immediate reward and the value of the next state), T is the time step at the end of the episode, represents the state value estimated using the old evaluation network at state s t before the evaluation network is updated.
[0104] (2) Network update: Randomly shuffle the order of the experiences, divide them into multiple Mini-batches equally, use the experience data in each Mini-batch to estimate the loss values of the policy and evaluation networks, and update the network once; repeat the above process K epoch times;
[0105]
[0106] In the formula, ε and c 1 are both constants, π θ is the policy network with parameter θ, a t ,st They are the action and state at time step t, respectively, π θ (a t ∣s t ) means that according to the policy π θ , the probability of selecting action a t under state s t is H(π θ (·|s t )) is the policy entropy, N a is the dimension of the action space, a i is the action in the i-th dimension, V ω (s) is the state value estimated by the current evaluation network, V target (s) is the target value of the state value, and V target (s) can be calculated based on GAE, r t (θ) is the clipping ratio coefficient.
[0107] Furthermore, the present invention improves the training effect of the PPO algorithm to a certain extent by introducing dynamic decay of the learning rate, policy entropy H(π θ (·|s t )) and state normalization:
[0108]
[0109]
[0110] In the formula, α 0 , α t are the initial learning rate and the learning rate at time step t, respectively, k is the decay rate for controlling the learning rate, S t , S t,norm are the current state value and the current state value after normalization processing, μ, σ are the mean and standard deviation of the state value, and μ, σ are continuously updated during training.
[0111] (3) Clear the experience buffer, repeat (1) and (2), and stop when the maximum number of training steps is reached. The complete training process of step C is as Figure 2 shown.
[0112] Next, the method provided by the present invention is verified by simulation.
[0113] (1) Training scenario
[0114] Step A: The training scenario parameters adopted by the present invention are shown in Table 1. The training round time is set to 900 seconds; the sampling interval of the agent is set to 5 seconds.
[0115] Table 1
[0116]
[0117] The model parameters are shown in Table 2:
[0118] Table 2
[0119]
[0120]
[0121] Step B: Compare the training processes using four reinforcement learning algorithms: DDPG, TD3, SAC, and PPO (used in the present invention). The algorithm parameters are set as shown in Table 3:
[0122] Table 3
[0123]
[0124] Step C: The training results are as Figure 4 shown. Among them, both DDPG and TD3 are deterministic policy algorithms. The stabilized reward values are similar, and the fluctuations of their reward curves are small, but the exploration of the action space is not sufficient enough, and it is easy to converge to the local optimum; The SAC algorithm is a probability-based policy algorithm. Affected by the randomness of the policy, its convergence speed is slow and the volatility is large; In contrast, although the PPO algorithm belongs to the probability-based policy algorithm, due to the introduction of probability ratio clipping in the objective function, it not only improves the training stability, but also achieves a better balance between policy exploration and exploitation, and has a higher reward value. In addition, the present invention improves the training effect of the PPO algorithm to a certain extent by introducing dynamic decay of the learning rate, policy entropy, and state normalization.
[0125] (2) Scenario 1
[0126] Step A: Set the scenario parameters
[0127] In this embodiment, under the condition of sufficient climbing capacity, the decisions of the intelligent agent are analyzed by comparing with the traditional method. The scenario parameters are shown in Table 4:
[0128] Table 4
[0129]
[0130] Step B: Result analysis
[0131] The present invention analyzes the two strategies respectively from the change of the distribution factor and the multi-energy storage power response process. Figure 5 is the change of the distribution factor of the multi-energy storage during the climbing process, Figure 5 in which (a) is the traditional strategy, Figure 5 in which (b) is the PPO distribution strategy (the same below). From Figure 5It can be seen that the traditional strategy distributes according to the proportion of the ramping capacity. In the initial stage, the ramping capacity of thermal power units is high, so the initial distribution factor is the largest. However, due to the fast response speed of the wind-storage system, it continuously fills the power response deviation caused by the response delay of the thermal-storage system and A-CAES. Its distribution factor gradually increases, and the ramping power it undertakes also continuously rises. In the middle and late stages of the ramping process, due to the gradually increasing ramping power, the available ramping capacity of the wind-storage system is exhausted, and the distribution factor decreases. The distribution factors of the thermal-storage system and A-CAES increase accordingly and finally tend to be stable. After stabilization, the thermal-storage system undertakes the main ramping power. For the PPO distribution strategy, in the initial stage, the system's ramping rate requirement is high, and the distribution factor of the wind-storage system is the largest. However, as the ramping rate requirement weakens and the ramping power gradually increases, the distribution factors of A-CAES and thermal storage gradually increase and finally tend to be stable. A-CAES undertakes the main ramping power.
[0132] Figure 6 Fig. 4 shows the multi-energy storage power response of the two distribution strategies in the "upward ramping scenario". Among them, the solid line and the dashed line respectively represent the actual power and power command of the multi-energy storage. The slant-marked part is the power response deviation, and the upper gray slant is the total power response deviation. From Figure 6 It can be seen that the power response deviation corresponding to the PPO strategy is smaller, which is about 31.3% less than that of the traditional strategy. In addition, the wind-storage system has good regulation performance and can accurately track the power command. However, the regulation performance of thermal power units is poor. If a large amount is called, it will cause a large power response deviation. The PPO distribution strategy allocates a large amount of ramping power to A-CAES. While giving full play to the advantages of A-CAES such as large capacity and low cost, it can also effectively avoid the ramping disadvantages of thermal power units, thereby reducing the power response deviation.
[0133] Figure 7 and Figure 8 Figs. 5 and 6 show the distribution factors and power responses of the traditional distribution strategy and the PPO distribution strategy in the "downward ramping scenario", which are basically similar to the "upward ramping scenario" and will not be elaborated here.
[0134] (3) Scenario 2
[0135] Step A: Set the scenario parameters as shown in Table 5.
[0136] Table 5
[0137]
[0138] Step B: Analyze the scenario results
[0139] The A-CAES unit needs to go through a working condition conversion from the compressed energy storage state to the expanded power generation state, and there is a discontinuous regulation interval, such as Figure 9As shown in the figure. The distribution of the climbing power should avoid the influence of this discontinuous regulation interval as much as possible. For this reason, this section analyzes the handling of this discontinuous regulation interval by two strategies.
[0140] The handling of the discontinuous regulation interval by the traditional distribution strategy is as Figure 10 shown. At about 310 seconds, the A-CAES enters the discontinuous regulation interval, and its distribution factor rises sharply; and the total power P of the multi-energy storage Σ steps up, greatly exceeding the power command P Σ,ord . The operating state of the A-CAES near this time is as Figure 11 shown. After reaching the minimum compression power, it immediately stops compression and cannot continue to track the power command P A-CAES,ord , and limited by the minimum expansion power, it is difficult to perform condition conversion and remains in the shutdown state, greatly reducing the flexible regulation ability of the power grid. Under the PPO distribution strategy, the distribution factor and power response of the multi-energy storage are as Figure 12 shown. It can be seen from the figure that the PPO distribution strategy is different from its performance in Scenario 1, greatly reducing the distribution factor of the A-CAES, thus avoiding the discontinuous regulation interval, and the overall power deficit is better controlled.
[0141] (4) Scenario 3
[0142] Step A: The scenario parameters are set as shown in Table 6.
[0143] Table 6
[0144]
[0145] Step B: Analysis of the scenario results
[0146] From Figure 13 and Figure 14 , it can be seen that under the traditional distribution strategy, at about 200 seconds and 800 seconds, the rotational speed ω of the wind-storage system drops and reaches the lower limit ω down,lim , at this time, the energy stored in the impeller is insufficient, and even relying on the electrochemical energy storage system cannot compensate for the sharp drop of the mechanical power P m , resulting in a rapid and large drop in the distribution factor, and ultimately causing the wind-storage system to be unable to track the power command P e of the power grid. In contrast, the PPO distribution strategy will gradually reduce its dependence on the wind-storage system as the climbing process progresses, increase the distribution factors of the thermal energy storage system and the A-CAES, and avoid the phenomenon of a large power deficit due to insufficient climbing ability of the wind-storage system caused by wind speed changes, effectively ensuring the safe and reliable operation of the system. In addition, as Figure 15 shown, the power response deviation corresponding to the PPO distribution strategy is lower, less affected by wind speed changes, and has stronger adaptability in this scenario.
[0147] (5)Scenario 4
[0148] Step A: As shown in Figure 16 , the linear model (i.e., the simplified model) of A-CAES and the wind-storage system is:
[0149]
[0150] In the formula, v in is the cut-in wind speed, v out is the cut-out wind speed, v o is the rated wind speed, and P wo is the rated power of the wind turbine.
[0151] Step B: Analysis of scenario results
[0152] Comparing Figure 17 and Figure 5 in (b), and Figure 6 in (b), it can be seen that in Scenario 1, the allocation trends of the "simplified" strategy and the "detailed" strategy are similar. However, compared with the "detailed" strategy, the error of the "simplified" strategy has increased by 11.2%. This indicates that the simplified environment can, to a certain extent, characterize the regulation characteristics of multiple types of energy storage, but there are still differences in details, which leads to a slightly worse adaptability of the "simplified" strategy.
[0153] Comparing Figure 18 and Figure 12 it can be seen that in Scenario 2, the simplified environment cannot represent the discontinuous regulation interval of A-CAES. Therefore, it tends to allocate a larger ramp power to A-CAES. At about 200 seconds, A-CAES enters the discontinuous regulation interval, resulting in a large power response deviation in the "simplified" strategy.
[0154] In addition, the wind-storage system in the simplified environment cannot consider the actual control situation of the wind turbine, so the generated strategy is more aggressive in the allocation for the wind-storage system. As shown in Figure 19 , in Scenario 3, the "simplified" strategy causes the wind-storage system to be unable to track the grid power command twice, bringing large fluctuations to the ramp process. While Figure 15 the "detailed" strategy in
[0155] can effectively avoid this problem and control the power response deviation within a smaller range.
[0156] An embodiment of the present invention provides an electronic device, including: a computer-readable storage medium and a processor;
[0157] The computer-readable storage medium is used to store executable instructions;
[0158] The processor is configured to read the executable instructions stored in the computer-readable storage medium and execute the method as described in any of the above embodiments.
[0159] An embodiment of the present invention provides a computer-readable storage medium storing computer instructions for causing a processor to execute the method as described in any of the above embodiments.
[0160] An embodiment of the present invention provides a computer program product including a computer program or instructions which, when executed by a processor, implement the method as described in any of the above embodiments.
[0161] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for dynamically optimizing the allocation of multi-type energy storage climbing power in seconds, characterized in that: include: S1, establishing a multi-type energy storage ramp power response model; wherein the multi-type energy storage includes A-CAES, wind farm combined electrochemical energy storage and thermal power unit combined flywheel energy storage; S2, converting the multi-type energy storage ramp power response model into a Markov decision process; Among them, the state space corresponding to the Markov decision process is ΔP t is the system unbalanced power at time step t, is the maximum upward climbing capacity of each unit or energy storage at time step t, subscript 1:N a Indicates that the subscript is 1, 2..., N in sequence a , N a =3, 1 represents thermal energy storage system, 2 represents A-CAES, 3 represents wind energy storage system, is the maximum downward climbing capacity of each unit or energy storage at time step t, is the actual output of each unit or energy storage at time step t, is the state of charge of each unit or energy storage at time step t, ω is the wind wheel speed; the action space a of the intelligent agent t ={a 1,t ,a 2,t ,a 3,t }, and the action space satisfies the constraints: a 1,t ,a 2,t ,a 3,t are the power allocation factors of the thermal storage system, A-CAES, and wind storage system respectively; the reward function of the agent r1=c1(ΔP TF +ΔP CAES +ΔP WT ), c1 is a constant, ΔP TF ,ΔP CAES ,ΔP WT They are the power response deviations of thermal storage, A-CAES, and wind storage, respectively; S3, using a reinforcement learning algorithm to train and solve the Markov decision process to obtain a multi-type energy storage climbing power allocation strategy.
2. The method according to claim 1, characterized in that The power response model of the A-CAES is: Among them, P tm,i is the shaft power of the i-th stage expander, is the air mass flow rate of the expander, T td,i is the inlet gas temperature of the i-th stage expander, T tx,i is the outlet gas temperature of the i-th stage expander, c p is the constant pressure specific heat of air; P cm is the power supplied by the motor to the compressor, is the air mass flow rate of the compressor, T cd,j and T cx,j is the inlet and outlet gas temperature of the jth compressor, n c is the number of compressor stages; The power response model of the wind farm combined with electrochemical energy storage is: Among them, P e is the converter output power, P set is the set power of the converter, T a is the converter time constant, J is the moment of inertia of the rotor, and s is a complex frequency domain variable; The power response model of the thermal power unit combined with flywheel energy storage is: Among them, P d and P c are the maximum discharge power and maximum charging power of flywheel energy storage, P o is the rated power, K, P, r, b are all constants, SOC max is the maximum allowable state of charge, SOC min is the minimum allowable state of charge, SOC is the state of charge.
3. The method according to claim 1 or 2, characterized in that In step S3, the reinforcement learning algorithm adopted is the PPO algorithm.
4. The method according to claim 3, characterized in that The output of the policy network is the α and β parameters of the Beta distribution.
5. The method according to claim 1 or 3, characterized in that: During training, the learning rate at time step t is And according to the formula Normalize the state.
6. An electronic device, characterized in that: include: A computer readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the method according to any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the method according to any one of claims 1 to 5.
8. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Micro-grid group cooperative scheduling method and device, electronic equipment and storage medium
CN113962446A
Multi-regional power grid collaborative optimization method, system and device and readable storage medium
CN115333111A
Intelligent scheduling method and system for wind-water-fire integrated energy system
CN115986839A
Power distribution system power distribution method and system, computer equipment and storage medium
CN116436013A
Wind-light-storage combined power generation optimization method and system based on deep reinforcement learning
CN117175591A