Intelligent optimization control method and system for compressed air energy storage power station
By combining reinforcement learning and model predictive control and introducing intelligent agents, the problem of insufficient control in traditional compressed air energy storage power stations when facing the volatility of new energy power generation is solved, realizing intelligent dynamic optimization control and improving the stability of the power system and the operating efficiency of the energy storage power station.
Patent Information
- Application Number
- CN202511156419.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-12-26
AI Technical Summary
Traditional compressed air energy storage power station control methods are difficult to cope with the intermittency and volatility of new energy power generation, resulting in insufficient stability and robustness of the power system and an inability to adjust the operating status of the energy storage power station in a timely and effective manner.
By combining reinforcement learning and model predictive control, a reinforcement learning agent is introduced to learn and optimize control strategies based on historical data and real-time information, enabling rapid adaptation to changes in new energy power generation and electricity load.
It realizes intelligent automatic control of energy storage power stations, which can quickly adjust when the power generation of new energy fluctuates, improve the stability and reliability of the power system, reduce human intervention, and optimize the operating efficiency and economic benefits of energy storage power stations.
Smart Images

Figure CN121209255A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power grid technology, and specifically to an intelligent optimization control method and system for compressed air energy storage power stations. Background Technology
[0002] With the increasing global demand for clean energy and the growing challenges facing traditional energy structures, the application of new energy sources in power systems is becoming increasingly widespread. Solar and wind power generation technologies, with their advantages of being renewable and environmentally friendly, occupy an important position in the energy sector. However, the inherent characteristics of new energy power generation, such as intermittency and volatility, pose numerous challenges to the stable operation of power systems. The output power of new energy power generation fluctuates significantly with weather conditions, time, and other factors, making it difficult to maintain the supply-demand balance of the power system. This, in turn, affects the frequency and voltage stability of the power grid and even threatens the safe and reliable supply of electricity.
[0003] Among the technological means to address the volatility of renewable energy generation, energy storage technology has become a key solution. Compressed air energy storage (CAES), as a large-scale energy storage technology, boasts advantages such as large storage capacity, long storage period, and relatively low cost, making it a promising candidate for energy storage applications in power systems. Traditional CAES control methods are primarily based on deterministic models and rule-based control strategies. For example, the common model predictive control method establishes an accurate mathematical model to predict the future operating state of the energy storage station and optimizes the control strategy based on the prediction results. This method can, to some extent, consider the dynamic characteristics of the system and plan the charging and discharging operations of the energy storage station in advance to cope with changes in renewable energy generation and power load.
[0004] However, traditional control methods have significant drawbacks. On the one hand, the accuracy of their models highly depends on precise mathematical descriptions and parameter identification of the system. However, real-world power systems are complex and dynamic, with numerous uncertainties, such as prediction errors in renewable energy generation, uncertainties in load changes, and variations in equipment operating characteristics. These factors make it difficult for traditional models to accurately depict the system's real behavior, thus affecting control effectiveness. On the other hand, traditional control methods lack sufficient flexibility and adaptability when facing sudden situations or unmodeled dynamics, exhibiting poor robustness. For example, when extreme weather causes drastic changes in renewable energy generation or sudden grid failures, traditional control strategies may fail to adjust the operating status of energy storage stations in a timely and effective manner, threatening the stability of the power system.
[0005] In recent years, artificial intelligence (AI) technology has made groundbreaking progress in various fields, providing new ideas and methods for solving control problems of complex systems. Reinforcement learning, as an important branch of AI, learns optimal strategies through interaction between intelligent agents and the environment, exhibiting strong adaptability and learning capabilities. In the field of power system control, reinforcement learning has begun to be applied in some areas, such as energy management in microgrids and bidding strategies in the electricity market. However, combining reinforcement learning with model predictive control for intelligent optimization control of compressed air energy storage power stations is still in the exploratory stage. Existing technologies have not yet fully leveraged the advantages of both to achieve efficient, stable, and intelligent control of energy storage power stations.
[0006] Therefore, to meet practical needs, an intelligent optimization control technology for compressed air energy storage power stations is provided. Summary of the Invention
[0007] To address the shortcomings of existing technologies, the present invention aims to provide an intelligent optimization control method and system for compressed air energy storage power stations. By organically combining reinforcement learning and model predictive control, a reinforcement learning agent is introduced. This agent can continuously learn and optimize control strategies based on a large amount of historical data and real-time operating information. When faced with real-time fluctuations in new energy power generation, it can quickly make adaptive adjustments based on factors such as current power generation, power load, energy storage status, and environment.
[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0009] In a first aspect, this application provides an intelligent optimization control method for a compressed air energy storage power station, the method comprising the following steps:
[0010] S1: Establish the system model corresponding to the compressed air energy storage power station and collect data;
[0011] S2: Train the reinforcement learning agent to learn the optimal control strategy that can adapt to the operating environment of the compressed air energy storage power station;
[0012] S3: Use a predictive model to predict the future operating status of the compressed air energy storage power station and optimize the control sequence;
[0013] S4: Perform real-time control operations and adjust the control strategy based on feedback data to achieve dynamic optimization control.
[0014] Based on the above technical solution, step S1 involves establishing a system model corresponding to the compressed air energy storage power station, including the following steps:
[0015] Create models for the air compression process, air storage process, air expansion power generation process, interaction with the power grid, and system constraints.
[0016] Based on the above technical solution, the air compression process model is configured with:
[0017] m in =k1·n·P in ;
[0018] Among them, P comp m is the power consumed during the air compression process. in P is the intake air flow rate during the air compression process. in P is the intake pressure during the air compression process. out η is the exhaust pressure during the air compression process. comp C represents the compressor efficiency during the air compression process. p T is the specific heat capacity of air at constant pressure. in γ is the intake air temperature, n is the specific heat ratio of air, and k1 is the compressor speed.
[0019] Based on the above technical solution, the air storage process model is configured with:
[0020]
[0021] Among them, P tank The pressure inside the gas storage tank, m tank T represents the mass of the gas inside the storage tank. tank The temperature inside the gas storage tank is V, where R is the gas constant and V is the gas temperature. tank Let m be the volume of the gas storage tank. out This refers to the gas flow rate during the expansion process.
[0022] Based on the above technical solution, the air expansion power generation process model is configured with:
[0023] m out =k2·n exp ·P tank ;
[0024] P exp For the expander output power, m out For expander inlet air flow, P tank For expander inlet pressure, P exhaust η is the exhaust pressure of the expander. exp For the efficiency of the expander, n exp K is the speed of the expander, and k2 is the proportional coefficient.
[0025] Based on the above technical solution, the power grid interaction model is configured with:
[0026] P grid =P exp -P comp -L;
[0027] P grid -k p ·(f ref -f grid );
[0028] P grid P represents the power exchanged between the compressed air energy storage power station and the power grid. comp For the power consumed by the compressor, P exp Where L is the output power of the expander, f is the internal loss of the power plant, and f is the output power of the expander. grid k is the power grid frequency. p f is the frequency droop coefficient. ref This is the reference frequency for the power grid.
[0029] Based on the above technical solution, the system constraints include power limits for the compressor and expander, pressure limits for the gas storage tank, and speed limits for the equipment.
[0030] The power limits for the compressor and expander are as follows:
[0031] P comp,min ≤P comp ≤P comp,max and P exp,min ≤P exp ≤P exp,max ;
[0032] Among them, P comp,min P comp,max These are the minimum and maximum power of the compressor, P. exp,min P exp,max These are the minimum and maximum power of the expander, respectively;
[0033] The pressure limit of the gas storage tank is: P tank,min ≤P tank ≤P tank,max ;
[0034] Among them, P tank,min P tank,max These are the minimum and maximum allowable pressures for the gas storage tank, respectively.
[0035] The rotational speed of the device is limited to: n min ≤n≤n max and n exp,min ≤n exp ≤n exp,max ;
[0036] Where, n min n max n represents the minimum and maximum compressor speeds, respectively. exp,min n exp,max These represent the minimum and maximum speeds of the expander, respectively.
[0037] Based on the above technical solution, step S1, data collection includes the following steps:
[0038] Collect data on new energy power generation, power load, internal operation data of energy storage power stations, and environmental data; among which,
[0039] The new energy power generation data includes the real-time power output of new energy power generation equipment and meteorological data that affects new energy power generation;
[0040] The power load data includes real-time power load demand and load characteristics.
[0041] The internal operating data of the energy storage power station includes the operating status of the compressor, expander, gas storage tank and heat exchanger, as well as the pressure, temperature and energy storage status of the gas in the gas storage tank;
[0042] The environmental data includes the temperature, air pressure, and humidity around the energy storage power station.
[0043] Based on the above technical solution, the reinforcement learning agent consists of a policy network, a value network, and an environment interaction module;
[0044] The policy network is used to generate control actions based on the current state, and its output is a probability distribution for each possible action;
[0045] The value network is used to estimate the value of a given state, helping the agent to evaluate the merits of different states;
[0046] The environmental interaction module is responsible for applying the actions generated by the intelligent agent to the energy storage power station system environment, and receiving the next state and reward signal from the environment.
[0047] Secondly, this application also provides an intelligent optimization control system for compressed air energy storage power stations, the system comprising:
[0048] The model building module is used to build a system model for the compressed air energy storage power station and collect data.
[0049] The agent training module is used to train a reinforcement learning agent to learn the optimal control strategy that can adapt to the operating environment of the compressed air energy storage power station.
[0050] The prediction module is used to predict the future operating state of the compressed air energy storage power station using a prediction model and optimize the control sequence.
[0051] The dynamic control module is used to perform real-time control operations and adjust the control strategy based on feedback data to perform dynamic optimization control.
[0052] Compared with the prior art, the advantages of the present invention are as follows:
[0053] This invention introduces a reinforcement learning agent by organically combining reinforcement learning and model predictive control. This agent can continuously learn and optimize control strategies based on a large amount of historical data and real-time operating information. When faced with real-time fluctuations in the power generation of new energy sources, it can quickly make adaptive adjustments based on factors such as the current power generation, power load, energy storage status, and environment.
[0054] The innovative reward function designed in this invention, which comprehensively considers power system stability, energy storage power station operating efficiency, and economic benefits, is a key component of this invention. During the training process of the reinforcement learning agent, this reward function guides the agent to learn the optimal control strategy that balances multiple objectives.
[0055] This invention utilizes reinforcement learning and model predictive control technologies to achieve intelligent automatic control of energy storage power stations. By monitoring system operation data in real time, the intelligent agent and model predictive control module can automatically generate optimized control strategies based on the system state and execute control actions promptly, without frequent manual intervention. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 This is a flowchart illustrating the steps of an intelligent optimization control method for a compressed air energy storage power station according to an embodiment of the present invention.
[0058] Figure 2 This is a flowchart illustrating the principle of the intelligent optimization control method for compressed air energy storage power stations according to an embodiment of the present invention.
[0059] Figure 3 This is a schematic diagram of the reinforcement learning training process in the intelligent optimization control method for compressed air energy storage power stations according to an embodiment of the present invention.
[0060] Figure 4This is a technical schematic diagram of the intelligent optimization control system for compressed air energy storage power stations according to an embodiment of the present invention. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0062] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0063] This application provides an intelligent optimization control method and system for compressed air energy storage power stations. By organically combining reinforcement learning and model predictive control, a reinforcement learning agent is introduced. This agent can continuously learn and optimize control strategies based on a large amount of historical data and real-time operating information. When faced with real-time fluctuations in new energy power generation, it can quickly make adaptive adjustments based on factors such as current power generation, power load, energy storage status, and environment.
[0064] The embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0065] Firstly, see [the following] Figures 1-3 As shown in the figure, this application provides an intelligent optimization control method for a compressed air energy storage power station, which includes the following steps:
[0066] S1: Establish the system model corresponding to the compressed air energy storage power station and collect data;
[0067] S2: Train the reinforcement learning agent to learn the optimal control strategy that can adapt to the operating environment of the compressed air energy storage power station;
[0068] S3: Use predictive models to predict the future operating status of compressed air energy storage power stations and optimize the control sequence;
[0069] S4: Perform real-time control operations and adjust the control strategy based on feedback data to achieve dynamic optimization control.
[0070] The technical solutions of this application are intended to solve the following technical problems:
[0071] The intermittent and fluctuating nature of renewable energy generation leads to significant variations in its output power across different time scales. Traditional control methods struggle to accurately track these dynamic changes, hindering energy storage power stations from effectively and promptly balancing power supply and demand through charging and discharging operations. This invention aims to combine reinforcement learning and model predictive control to enable energy storage power stations to adaptively adjust their control strategies based on real-time fluctuations in renewable energy generation power and changes in power load. This ensures efficient and stable operation of the energy storage power station under various renewable energy generation conditions, thereby improving the power system's capacity to absorb renewable energy.
[0072] The operation of energy storage power stations involves energy conversion and equipment operation in multiple stages. Traditional control methods, when optimizing the operation of energy storage power stations, often only consider a single objective, such as power balance or equipment lifespan, making it difficult to comprehensively consider operational efficiency and economy. This invention designs a reward function that comprehensively considers power system stability, energy storage power station operating efficiency, and economic benefits, guiding a reinforcement learning agent to learn an optimal control strategy that takes into account multiple objectives.
[0073] As power systems continue to expand in scale and increase in complexity, the control requirements for energy storage power stations are also becoming more stringent. Traditional manual intervention or rule-based control methods are no longer sufficient to meet these demands. This invention utilizes reinforcement learning and model predictive control technologies to achieve intelligent automatic control of energy storage power stations.
[0074] In this embodiment, a reinforcement learning agent is introduced by organically combining reinforcement learning and model predictive control. This agent can continuously learn and optimize control strategies based on a large amount of historical data and real-time operating information. When faced with real-time fluctuations in the power generation of new energy sources, it can quickly make adaptive adjustments based on factors such as the current power generation, power load, energy storage status, and environment.
[0075] Furthermore, in step S1, establishing the system model corresponding to the compressed air energy storage power station includes the following steps:
[0076] Create models for the air compression process, air storage process, air expansion power generation process, interaction with the power grid, and system constraints.
[0077] Furthermore, the air compression process model is configured with:
[0078]
[0079] m in =k1·n·P in ;
[0080] Among them, P comp m is the power consumed during the air compression process. inP is the intake air flow rate during the air compression process. in P is the intake pressure during the air compression process. out η is the exhaust pressure during the air compression process. comp C represents the compressor efficiency during the air compression process. p T is the specific heat capacity of air at constant pressure. in γ is the intake air temperature, n is the specific heat ratio of air, and k1 is the compressor speed.
[0081]
[0082] Among them, P tank The pressure inside the gas storage tank, m tank T represents the mass of the gas inside the storage tank. tank The temperature inside the gas storage tank is V, where R is the gas constant and V is the gas temperature. tank Let m be the volume of the gas storage tank. out This refers to the gas flow rate during the expansion process.
[0083] Furthermore, the air expansion power generation process model is configured with:
[0084]
[0085] m out =k2·n exp ·P tank ;
[0086] P exp For the expander output power, m out For expander inlet air flow, P tank For expander inlet pressure, P exhaust η is the exhaust pressure of the expander. exp For the efficiency of the expander, n exp K is the speed of the expander, and k2 is the proportional coefficient.
[0087] Furthermore, the interaction model with the power grid is configured with:
[0088] P grid =P exp -P comp -L;
[0089] P grid -k p ·(f ref -f grid );
[0090] P grid P represents the power exchanged between the compressed air energy storage power station and the power grid. comp For the power consumed by the compressor, P expWhere L is the output power of the expander, f is the internal loss of the power plant, and f is the output power of the expander. grid k is the power grid frequency. p f is the frequency droop coefficient. ref This is the reference frequency for the power grid.
[0091] Furthermore, the system constraints include power limits for the compressor and expander, pressure limits for the gas storage tank, and speed limits for the equipment.
[0092] The power limits for the compressor and expander are as follows:
[0093] P comp,min ≤P comp ≤P comp,max and P exp,min ≤P exp ≤P exp,max ;
[0094] Among them, P comp,min P comp,max These are the minimum and maximum power of the compressor, P. exp,min P exp,max These are the minimum and maximum power of the expander, respectively;
[0095] The pressure limit of the gas storage tank is: P tank,min ≤P tank ≤P tank,max ;
[0096] Among them, P tank,min P tank,max These are the minimum and maximum allowable pressures for the gas storage tank, respectively.
[0097] The rotational speed of the device is limited to: n min ≤n≤n max and n exp,min ≤n exp ≤n exp,max ;
[0098] Where, n min n max n represents the minimum and maximum compressor speeds, respectively. exp,min n exp,max These represent the minimum and maximum speeds of the expander, respectively.
[0099] Furthermore, in step S1, data collection includes the following steps:
[0100] Collect data on new energy power generation, power load, internal operation data of energy storage power stations, and environmental data; among which,
[0101] The new energy power generation data includes the real-time power output of various new energy power generation equipment (such as photovoltaic and wind power), and meteorological data that affects new energy power generation, including solar radiation intensity (for photovoltaic), wind speed, wind direction, temperature, air pressure, etc.
[0102] The power load data includes real-time power load demand; load characteristics, such as load type, daily variation, weekly variation, and seasonal variation patterns.
[0103] The internal operating data of the energy storage power station includes the operating status of equipment such as compressors, expanders, gas storage tanks, and heat exchangers, including start-up and shutdown, operating time, temperature, speed, and pressure; and the pressure, temperature, and energy storage status (such as gas mass or equivalent energy) of the gas in the gas storage tank.
[0104] The environmental data includes the temperature and air pressure around the energy storage power station, as well as optional humidity data.
[0105] Furthermore, the reinforcement learning agent consists of a policy network, a value network, and an environment interaction module;
[0106] The policy network is used to generate control actions based on the current state, and its output is a probability distribution for each possible action;
[0107] The value network is used to estimate the value of a given state, helping the agent to evaluate the merits of different states;
[0108] The environmental interaction module is responsible for applying the actions generated by the intelligent agent to the energy storage power station system environment, and receiving the next state and reward signal from the environment.
[0109] Furthermore, in step S2, the state space S of the reinforcement learning agent contains multiple key pieces of information to comprehensively reflect the operating environment and its own state of the energy storage power station.
[0110] Let S = {s1, s2, s3, s4, s5, s6}, where s1 is the predicted value of new energy power generation. Its value range is determined based on statistical analysis of historical new energy power generation data, for example... in To predict the minimum power, s2 represents the predicted maximum power. s2 represents the current power load P. load The value range is based on statistical data from power grid load monitoring, such as [P]. load,min ,P load,max ] s3 represents the current energy state of the energy storage power station, which can be represented by the gas storage tank pressure P. tank It means that P tank The range of values for is [P] tank,min ,P tank,max ] s4 represents the ambient temperature Tenv The value range is based on local historical temperature data, such as [T min ,T max ] s5 represents the ambient air pressure P atm The range of values is based on local meteorological data statistics, such as [P] atm,min ,P atm,max ]. s6 is the state vector S of the energy storage power station equipment. eq For example, using binary encoding to represent the start / stop status of key equipment (compressors, expanders, etc.), if there are n devices, then S eq It is an n-dimensional vector, where each element takes the value 0 (device stopped) or 1 (device running).
[0111] Furthermore, in step S2, the action space A of the reinforcement learning agent is defined as the control operations that the energy storage power station can execute. Let A = {a1, a2, a3}, where a1 is the compressor control action, which can be represented as (P comp (on / off), that is, the compressor's power setting value P comp (The value range is [P) comp,min ,P comp,max ]) and start / stop commands (on means start, off means stop). a2 is the expander control action, similarly represented as (P exp (on / off), P exp The range of values is [P] exp,min ,P exp,max a3 represents the adjustment action ΔP for the gas storage tank pressure setpoint. tank The value range is determined based on the allowable pressure adjustment range of the gas storage tank, such as [-ΔP]. max ,ΔP max ], where ΔP max This is the maximum pressure adjustment amount.
[0112] Furthermore, in step S2, the reward function R of the reinforcement learning agent is designed by comprehensively considering the stability of the power system, the operating efficiency of the energy storage power station, and economic benefits:
[0113] R = w1·f stab (s,a)+w2·f eff (s,a)+w3·f eco (s,a);
[0114] Among them, ω1, ω2, and ω3 are weighting coefficients used to balance the importance of different objectives.
[0115] f stab (s,a) is the stability index function, defined as:
[0116] f stab (s,a)=-(Δf2 +ΔV 2 );
[0117] Where Δf is the grid frequency deviation, and ΔV is the grid voltage deviation. Frequency deviation Δf = ff ref f is the actual power grid frequency, f ref The reference frequency for the power grid; voltage deviation ΔV = VV ref V is the actual grid voltage. ref This is the reference voltage for the power grid.
[0118] f eff (s,a) is the operational efficiency index function, which can be expressed as: f eff (s,a)=-(L+C main ), where L represents the energy loss during the operation of the energy storage power station (including losses during compression, storage, and expansion), and C main This is an estimated value for equipment maintenance costs, which is related to factors such as equipment uptime and the number of start-ups and shutdowns. It can be calculated using empirical formulas, such as C. main =k1·t run +k2·n start Where k1 and k2 are coefficients, t run n represents the total operating time of the equipment. start This represents the number of times the equipment has been started and stopped.
[0119] f eco (s,a) is an economic benefit index function, defined as:
[0120] f eco (s,a)=I peak +I renew ;
[0121] Among them, I peak For the revenue obtained by energy storage power stations participating in grid peak shaving, I renew To improve the economic benefits of renewable energy consumption (such as reducing fines for wind and solar curtailment and increasing revenue from renewable energy power generation), its calculation is related to the grid pricing mechanism and renewable energy power generation policies.
[0122] Furthermore, in step S2, the selected reinforcement learning algorithm is a Deep Q-Network (DQN), and its training process is as follows:
[0123] ① Initialize the policy network θ and the target network θ - (initial θ) - =θ), and the experience replay buffer D.
[0124] ② In each training iteration, the agent adjusts its state based on the current state s. tAction a is generated using an ε-greedy policy of the policy network (which randomly selects an action with probability ε and selects the optimal action recommended by the policy network with probability 1-ε). t .
[0125] ③The action a t It acts on the environment, and the environment provides feedback on the next state s. t+1 and reward r t .
[0126] ④ will (s t ,a t ,r t ,s t+1 Stored in the experience playback buffer D.
[0127] ⑤ Randomly sample a batch of samples (s) from the experience playback buffer D. i ,a i ,r i ,s i+1 ), i = 1, 2, ..., m.
[0128] ⑥ Calculate the target value y i y i =r i +γ·max a′ Q(s i+1 ,a′;θ - ), where γ is the discount factor, max a′ Q(s i+1 ,a′;θ - ) represents the target network for the next state s i+1 The estimated value of the action a′ to be taken.
[0129] ⑦ Calculate the loss of the policy network using the mean squared error (MSE) loss function:
[0130]
[0131] ⑧ Update the policy network parameters θ using the backpropagation algorithm to minimize the loss function L(θ).
[0132] 9. Regularly update the target network parameters θ - For example, θ is updated once every C training iterations. - ←τ·θ+(1-τ)·θ - , where τ is the soft update coefficient.
[0133] ⑩ The training process continues until the stopping condition is met, such as reaching the maximum number of training iterations or the policy network converging (the loss function value is less than a set threshold). Through continuous training, the agent gradually learns the optimal control strategy to adapt to the operating environment of the energy storage power station, that is, to select actions that maximize long-term cumulative rewards under different states.
[0134] Furthermore, in step S3, model predictive control is used to optimize the control sequence output by reinforcement learning. The predictive model used is based on the compressed air energy storage power station system model established in step S1. This model describes the dynamic behavior of the energy storage power station system in future time periods, and its discrete-time state-space equation can be expressed as:
[0135] x k+1 =f(x) k ,u k );
[0136] Where, x k It is the system state vector at time k, containing key state variables of the energy storage power station, such as the gas tank pressure P. tank,k Equipment operating status (such as the start / stop status of compressors and expanders); u k It is the control input vector at time k, including the compressor power setpoint P. comp,k Expander power setting value P exp,k Adjustment amount ΔP of gas tank pressure setpoint tank,k etc.; f is the system's state transition function, which is based on the current state x. k and control input u k Calculate the system state x at the next time step. k+1 .
[0137] Furthermore, in step S3, model predictive control is used to optimize the control sequence output by reinforcement learning, with the prediction time domain and control time domain set as N, respectively. p and N c These factors determine the timeframe within which model predictive control will consider the future and the time step for optimizing control decisions. Prediction time domain N p The selection of N should be determined based on the volatility of new energy power generation, the cycle of power load changes, and the dynamic response characteristics of the energy storage power station. Typically, the value ranges from tens of minutes to several hours, for example, N. p = 60 minutes (when the time step is in minutes). Control time domain N c Generally less than or equal to the prediction time domain N p It represents the number of time steps in which the control input is actually adjusted within each optimization cycle, for example, N. c = 15 minutes.
[0138] Furthermore, in step S3, model predictive control is used to optimize the control sequence output by reinforcement learning, with the objective function being J, expressed as follows:
[0139]
[0140] ω1, ω2, and ω3 are weighting coefficients used to balance the importance of different objectives.
[0141] C op,k Let k be the operating cost of the energy storage power station at time k, including equipment operating energy consumption costs, equipment maintenance costs (related to equipment operating status), etc., which can be expressed as:
[0142] C op,k =c energy ·(P comp,k +P exp,k )+c main ·h(x k );
[0143] Where c energy It is the unit energy cost coefficient, c main It is the equipment maintenance cost coefficient, h(x) k ) is a maintenance cost function related to the equipment's operating status, calculated based on factors such as the equipment's cumulative operating time and the number of start-ups and shutdowns.
[0144] It is the actual renewable energy power generation at time k. This is the predicted value of new energy power generation at time k. This item is used to measure the absorption of new energy sources; the smaller the item, the better the absorption effect of new energy sources.
[0145] g(x k ,u k () is a penalty function used to penalize violations of system operating constraints, such as exceeding pressure limits in gas storage tanks or exceeding power limits in equipment. For example, the penalty function for gas storage tank pressure constraints can be expressed as:
[0146]
[0147] Where k p It is the penalty coefficient for exceeding the pressure limit.
[0148] Furthermore, in step S3, model predictive control is used to optimize the control sequence output by reinforcement learning. At each control time k, based on the current system state x... k And the prediction model, based on the set objective function and constraints, solves the optimization problem to obtain the future N c The optimal control sequence within one control step.
[0149] In actual operation, model predictive control employs a rolling optimization strategy. In each control cycle, only the first control action from the optimal control sequence is executed. Then in the next control cycle, based on the new system state x k+1 (Obtained through actual measurements or state estimation), the system then re-predicts and optimizes to obtain a new optimal control sequence. This continuous rolling optimization enables the energy storage power station to adjust its control strategy in a timely manner based on actual operating conditions, adapting to the dynamic changes in new energy generation and power load.
[0150] Furthermore, in step S4, the first control action of the optimal control sequence is... It is applied to the actual operation of energy storage power stations and collects real-time operating data, including the actual operating status of the equipment. Actual power output Actual pressure of the gas storage tank The system then feeds this data back to the agent and the model prediction and control module. The agent then adjusts the response based on the new state. (in To update the energy storage power station's energy based on actual operating data, the model prediction control module corrects the prediction model parameters based on feedback data, and re-predicts and optimizes, thereby achieving dynamic adjustment of the control strategy.
[0151] In summary, the technical solution of this application embodiment can significantly improve the stability and reliability of the power system during grid operation. By dynamically adjusting the operating status of the energy storage power station, it can cope with the intermittent fluctuations in new energy power generation and maintain the stability of grid frequency and voltage. For example, when sudden changes in wind speed cause large fluctuations in wind power output, the energy storage power station can respond quickly, balance supply and demand, and prevent grid frequency collapse. Its application involves real-time monitoring of the grid status and, based on reinforcement learning and model predictive control strategies, automatically optimizing the charging and discharging operations of the energy storage power station to ensure the safe and stable operation of the power system, reduce the risk of power outages, and improve power supply quality.
[0152] In the industrial and commercial sectors, the technical solutions of this application can be used to optimize energy management and reduce electricity costs. Enterprises can utilize energy storage power stations to store electrical energy during periods of low electricity prices and use the stored energy during periods of high electricity prices, achieving peak-valley arbitrage. Simultaneously, during industrial production, energy storage power stations can provide a stable power supply, avoiding production interruptions and equipment damage caused by grid fluctuations. Its application involves combining an enterprise's production plan and electricity consumption characteristics to formulate personalized energy storage control strategies, improving energy utilization efficiency, reducing operating costs, and enhancing the enterprise's competitiveness.
[0153] Secondly, see Figure 4As shown in the figure, this application provides an intelligent optimization control system for a compressed air energy storage power station, the system comprising:
[0154] The model building module is used to build a system model for the compressed air energy storage power station and collect data.
[0155] The agent training module is used to train a reinforcement learning agent to learn the optimal control strategy that can adapt to the operating environment of the compressed air energy storage power station.
[0156] The prediction module is used to predict the future operating state of the compressed air energy storage power station using a prediction model and optimize the control sequence.
[0157] The dynamic control module is used to perform real-time control operations and adjust the control strategy based on feedback data to perform dynamic optimization control.
[0158] In this embodiment of the application, a reinforcement learning agent is introduced by organically combining reinforcement learning and model predictive control. This agent can continuously learn and optimize control strategies based on a large amount of historical data and real-time operating information. When faced with real-time fluctuations in the power generation of new energy sources, it can quickly make adaptive adjustments based on factors such as the current power generation, power load, energy storage status, and environment.
[0159] It should be noted that the specific working content of each module in the intelligent optimization control system for compressed air energy storage power stations is as described in the first aspect of the intelligent optimization control method for compressed air energy storage power stations. Each module is used to execute the corresponding step and the sub-steps within that step.
[0160] It should be noted that the intelligent optimization control system for compressed air energy storage power stations mentioned in the second aspect is similar in technical principle to the intelligent optimization control method for compressed air energy storage power stations mentioned in the first aspect in terms of technical issues, technical means, and technical effects, and will not be elaborated here.
[0161] In the description of this application, it should be noted that the terms "upper," "lower," etc., indicating the orientation or positional relationship are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two elements. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.
[0162] It should be noted that in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0163] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for intelligent optimization control of compressed air energy storage power plants, characterized in that, The method comprises the following steps: S1: establishing a system model corresponding to the compressed air energy storage power station and collecting data; S2: training a reinforcement learning agent to learn an optimal control strategy that can adapt to the operating environment of the compressed air energy storage power station; S3: using a prediction model to predict the future operating state of the compressed air energy storage power station and optimizing the control sequence; S4: performing real-time control operations and adjusting the control strategy based on feedback data for dynamic optimization control.
2. The method of intelligent optimization control for compressed air energy storage power plants of claim 1, wherein, In step S1, the system model corresponding to the compressed air energy storage power station is established, comprising the following steps: Creating an air compression process model, an air storage process model, an air expansion power generation process model, an interaction model with the power grid, and system constraint conditions.
3. The intelligent optimization control method for a compressed air energy storage power station according to claim 2, wherein: The air compression process model is configured with: m in = k1 · n · P in ; where P comp is the power consumed in the air compression process, m in is the mass flow rate of air into the compressor, P in is the pressure of air into the compressor, P out is the pressure of air out of the compressor, η comp is the efficiency of the compressor, C p is the specific heat capacity of air at constant pressure, T in is the temperature of air into the compressor, γ is the ratio of specific heat capacities of air, n is the rotational speed of the compressor, and k1 is a proportional coefficient.
4. The intelligent optimization control method for a compressed air energy storage power station according to claim 3, wherein: The air storage process model is configured with: where P tank is the pressure inside the tank, m tank is the mass of gas inside the tank, T tank is the temperature inside the tank, R is the gas constant, V tank is the volume of the tank, m out is the outflow of gas during expansion.
5. The intelligent optimization control method for a compressed air energy storage power station according to claim 4, wherein: The air expansion power generation process model is configured with: m out = k2- n exp · P tank ; P exp is the output power of the expander, m out is the intake flow rate of the expander, P tank is the intake pressure of the expander, P exhaust is the exhaust pressure of the expander, η exp is the efficiency of the expander, n exp is the rotational speed of the expander, k2 is a proportional coefficient.
6. The intelligent optimization control method for a compressed air energy storage power station according to claim 5, wherein: In the interaction model with the power grid, the following are configured: P grid = P exp - P comp - L; P grid -k p ·(f ref -f grid ); P grid Pexch is the exchanged power between the compressed air energy storage plant and the grid, comp Pcomp is the compressor consumed power, exp Pout is the expander output power, L is the plant internal losses, grid f is the grid frequency, p k is the frequency droop coefficient, ref fref is the grid reference frequency.
7. The intelligent optimization control method for a compressed air energy storage power station according to claim 6, wherein: The system constraint conditions include power limits of the compressor and the expander, pressure limits of the gas tank, and device speed limits; The power limits of the compressor and the expander are: P comp,min ≤P comp ≤P comp,max and P exp,min ≤P exp ≤P exp,max ; where P comp,min , P comp,max are the minimum and maximum power of the compressor, respectively, and P exp,min , P exp,max are the minimum and maximum power of the expander, respectively. The gas tank pressure limit is: P tank,min ≤ P tank ≤ P tank,max ; where Pmin and Pmax are the minimum and maximum pressures allowed by the gas tank, respectively. tank,min Pmin and Pmax are the minimum and maximum pressures allowed by the gas tank, respectively. tank,max where Pmin and Pmax are the minimum and maximum pressures allowed by the gas tank, respectively The device rotational speed limit is: n min ≤ n ≤ n max and n exp,min ≤ n exp ≤ n exp,max ; where n min , n max are the minimum and maximum values of the compressor speed, respectively, and n exp,min , n exp,max are the minimum and maximum values of the expander speed, respectively.
8. The method for intelligent optimization control of compressed air energy storage power station of claim 1, wherein, In step S1, the data collection comprises the following steps: Collecting new energy generation data, power load data, internal operating data of the energy storage power station, and environmental data; wherein, The new energy generation data includes real-time power output of new energy generation equipment and meteorological data affecting new energy generation; The power load data includes real-time power load demand size and load characteristics; The internal operating data of the energy storage power station includes the operating state of the compressor, the expander, the gas tank, and the heat exchanger, and also includes the pressure, temperature, and energy storage state of the gas in the gas tank; The environmental data includes the temperature, pressure, and humidity around the energy storage power station.
9. The intelligent optimization control method for a compressed air energy storage power station according to claim 1, wherein: The reinforcement learning agent is composed of a policy network, a value network, and an environment interaction module; The policy network is used to generate control actions based on the current state, and its output is a probability distribution of each possible action; The value network is used to estimate the value of a given state, helping the agent evaluate the pros and cons of different states; The environment interaction module is responsible for applying the actions generated by the agent to the energy storage power station system environment and receiving the next state and reward signal from the environment feedback.
10. An intelligent optimization control system for compressed air energy storage power plants, characterized by, The system comprises: A model construction module for establishing a system model corresponding to the compressed air energy storage power station and collecting data; An agent training module for training a reinforcement learning agent to learn an optimal control strategy that can adapt to the operating environment of the compressed air energy storage power station; a running prediction module configured to predict future running states of the compressed air energy storage power station by using a prediction model and to optimize a control sequence; a dynamic control module configured to perform real-time control operations and to adjust a control strategy according to feedback data for dynamic optimization control.
Citation Information
Cited By
Collaborative optimization method and device of energy storage system, electronic equipment and storage medium
CN121923201A