A Method and Device for Intra-Day Dispatching of a Distribution Network with a Compressed Air Energy Storage System
By building a distribution network model of reinforced learning agents and optimizing the scheduling strategy of compressed air energy storage systems, the problem of insufficient distribution network stability caused by the intermittent and uncertainty of renewable energy is solved, and efficient and intelligent intraday scheduling is achieved.
Patent Information
- Application Number
- CN202510228445.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-02-28
AI Technical Summary
When the prior art deals with the intermittent and uncertainty of renewable energy, it leads to insufficient stability and reliability of distribution networks, and has high computational complexity, making it difficult to meet real-time scheduling requirements.
Using the method of strengthening the interaction between the learner and the environment, a distribution network model containing compressed air energy storage system is constructed, and the dispatching strategies of energy storage power stations and thermal power units are optimized to achieve intelligent scheduling through neural networks and near-end strategy optimization algorithms.
It improves the reliability and stability of the distribution network, reduces operating costs, adapts to complex constraints, and improves the intelligence level of scheduling.
Smart Images

Figure CN119726971B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distribution networks, and particularly to a method and device for intraday scheduling of a distribution network with a compressed air energy storage system. Background Art
[0002] Under the background of the global energy structure transformation, the proportion of renewable energy in the distribution network is increasing continuously. As important renewable energies, wind power and photovoltaic power have the advantages of cleanness and sustainability, etc. However, due to their intermittency and uncertainty, they bring huge challenges to the stable operation of the distribution network. In order to effectively cope with these characteristics of renewable energy and improve the reliability and stability of the distribution network, energy storage technology has become one of the key solutions.
[0003] Compressed air energy storage and electrochemical energy storage power stations are gradually playing important roles in the distribution network due to their respective advantages. Compressed air energy storage has the characteristics of high energy density, long service life, relatively low cost, etc.; electrochemical energy storage has the advantages of fast response speed and flexible adjustment. In a distribution network containing multiple energies, in order to achieve efficient intraday scheduling of the distribution network and give full play to the effectiveness of various energies and energy storage devices, the introduction of intelligent decision-making methods has become an inevitable trend. As an intelligent decision-making method, reinforcement learning can learn the optimal strategy through the interaction between the intelligent agent and the environment, providing new ideas and methods for the intraday scheduling of a distribution network with compressed air energy storage.
[0004] The intraday distribution network scheduling method based on model predictive control is a relatively common scheduling method at present. This method establishes a mathematical model of the distribution network, makes a very short-term prediction of the output of renewable energy and load demand in the future for a period of time, and based on the prediction results and the operation constraints of the distribution network, solves the optimal scheduling scheme through an optimization algorithm. Finally, the distribution network is controlled in real time according to the obtained scheduling scheme. The advantages of this method are as follows: it can fully consider future situations and make scheduling decisions in advance, thereby improving the stability and reliability of the distribution network. At the same time, through the solution of the optimization algorithm, the optimality of the scheduling scheme can be guaranteed to a certain extent. However, this method also has some deficiencies: one is that the accuracy of the prediction depends on the quality of historical data and the performance of the prediction algorithm, and the output of renewable energy and load demand have great uncertainties, which may lead to a large deviation between the prediction results and the actual situation. The other is that the computational complexity is relatively high. Especially for large-scale distribution network systems, a large amount of computational resources and time are required to solve the optimal scheduling scheme, making it difficult to meet the requirements of real-time scheduling.
[0005] Therefore, to meet the actual needs, a technology for intraday scheduling of a distribution network with a compressed air energy storage system is provided. Summary of the Invention
[0006] Aiming at the defects existing in the prior art, the purpose of the present invention is to provide a method and device for intraday scheduling of a distribution network with a compressed air energy storage system, aiming to overcome the deficiencies of the prior art, learn the optimal strategy through the interaction between the agent and the environment, better cope with the intermittency and uncertainty of renewable energy, improve the reliability and stability of the distribution network, and reduce the operating cost at the same time, providing new ideas and methods for the efficient operation of the distribution network.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] In the first aspect, the present application provides a method for intraday scheduling of a distribution network with a compressed air energy storage system, and the method includes the following steps:
[0009] S1, constructing a distribution network model with a compressed air energy storage system and a distribution network simulation environment;
[0010] S2, constructing a reinforcement learning agent based on a neural network according to the characteristics of thermal power units and energy storage power stations in the distribution network model;
[0011] S3, based on the proximal policy optimization algorithm, using the reinforcement learning agent to interact with the distribution network simulation environment, so that the reinforcement learning agent learns to obtain the optimal scheduling strategy and obtain the trained reinforcement learning agent;
[0012] S4, using the trained reinforcement learning agent to perform intraday scheduling of the actual distribution network, determining the charging and discharging power of the energy storage power station in each scheduling period, optimizing the output and standby of each thermal power unit, obtaining the scheduling strategy for each scheduling period, and completing the intraday scheduling for one day;
[0013] S5, after completing the intraday scheduling for one day, updating the scheduling strategy of the trained reinforcement learning agent, repeating S4, and performing the intraday scheduling for a new day.
[0014] On the basis of the above technical solution, the distribution network model includes multiple thermal power units, a wind power generation station, a photovoltaic power generation station, multiple electrochemical energy storage power stations, a compressed air energy storage power station, and a load.
[0015] On the basis of the above technical solution, the compression power constraint of the compressor in the compressed air energy storage power station is:
[0016] ;
[0017] Wherein, is the power of the compressor at time t, and and are the upper and lower limits of the compressor power respectively, is a binary variable representing the operating state of the compressor at time t, 1 indicates that the compressor is in the working state, 0 indicates that the compressor is in the shutdown state;
[0018] The expansion power constraint of the expander in the compressed air energy storage power station is:
[0019] ;
[0020] Among them, is the power of the expander in the t period, while and are the upper and lower limits of the expander power respectively, is a binary variable representing the operating state of the expander in the t period, 1 indicates that the expander is in the working state, 0 indicates that the expander is in the shutdown state;
[0021] Limited by the internal structure of the compressed air energy storage system, the operating state constraints of the compressor and the expander are: ;
[0022] The mathematical model and constraint conditions of the gas storage chamber in the compressed air energy storage power station are:
[0023] ;
[0024] ;
[0025] ;
[0026] ;
[0027] Among them, is the air pressure value of the gas storage chamber in the t period, which is respectively limited by the lower limit of the air pressure value and the upper limit of the air pressure value , while and are the air pressure rise rate and the air pressure drop rate of the gas storage chamber in the t period respectively, and they are respectively proportional to the compression power , the expansion power , and the proportionality coefficients are and respectively, is a scheduling period;
[0028] The mathematical model and constraint conditions of the heat storage tank in the compressed air energy storage power station are:
[0029] ;
[0030] ;
[0031] ;
[0032] ;
[0033] Among them, is the heat storage amount of the heat storage tank at time period t, which is respectively restricted by the lower limit of the heat storage amount and the upper limit of the heat storage amount , while is the heat generation power of the heat exchanger in the air compression stage, which is proportional to the compression power at the same time period, and the proportionality coefficient is , is the heat absorption power of the heat exchanger in the air expansion stage, which is proportional to the expansion power at the same time period, and the proportionality coefficient is ;
[0034] The spare constraints in the compressed air energy storage power station are as follows:
[0035] ;
[0036]
[0037] Among them, and are respectively the positive spare capacity and the negative spare capacity of the compressed air energy storage system in the compressed charging state, and are respectively the positive spare capacity and the negative spare capacity of the compressed air energy storage system in the expansion discharge state;
[0038] The mathematical model and constraints of the electrochemical energy storage power station are as follows:
[0039] ;
[0040] ;
[0041] ;
[0042] Among them, is the energy stored by the i-th electrochemical energy storage power station at time period t, while and are respectively its lower limit and upper limit, and are respectively the charging power and the discharging power of the i-th electrochemical energy storage power station at time period t, while and are respectively the upper limits of the power when the i-th electrochemical energy storage power station is in the charging or discharging state, and are the charging and discharging efficiencies of the $i$-th electrochemical energy storage power station, respectively. By using binary variables to represent the operating state of the $i$-th electrochemical energy storage power station at time period $t$, being 1 indicates the charging state, being 0 indicates the discharging state;
[0043] The power constraint of the thermal power unit is:
[0044] ;
[0045] ;
[0046] ;
[0047] where represents the power of thermal power unit $j$ at time period $t$, which is limited by the upper limit and the lower limit in the operating state, while is a binary variable representing the operating state of thermal power unit $j$ at time period $t$, being 1 indicates that thermal power unit $j$ is in the operating and running state, being 0 indicates the shutdown state, which is determined by the day-ahead dispatch plan, and are the upper and lower ramp rate limits of thermal power unit $j$, respectively;
[0048] The reserve constraint of the thermal power unit is:
[0049] ;
[0050] ;
[0051] where and represent the positive reserve capacity and negative reserve capacity of thermal power unit $j$ at time period $t$, respectively;
[0052] The output constraint of the renewable energy is as follows: ;
[0053] where is the actual dispatched output of the renewable energy at time period $t$, while is the maximum output of the renewable energy at time period $t$;
[0054] The power balance constraint of the distribution network system is:
[0055] ;
[0056] where is the load at time period t, is the number of electrochemical energy storage power stations, is the number of thermal power units;
[0057] The reserve constraint of the distribution network system is:
[0058] ;
[0059] The lower limit value of the reserve capacity of the distribution network system is linearly related to the load and the output of renewable energy , and the ratios are respectively and .
[0060] On the basis of the above technical solution, in step S2:
[0061] The constructed reinforcement learning agent outputs an action at each time period t according to the state of the distribution network at the current time period and the reward calculated through the reward function to adjust the charging and discharging power of the compressed air energy storage power station and each electrochemical energy storage power station;
[0062] The state of the distribution network at the current moment includes the renewable energy generation sequence and load sequence in the past 24 hours, the energy state of each electrochemical energy storage power station at the current time period, the air chamber pressure value of the compressed air energy storage power station at the current time period, the heat stored in the heat storage tank of the compressed air energy storage power station at the current time period, the start-stop plan of each thermal power unit's day-ahead dispatch, and the actual output of the previous time period;
[0063] The reward function for the reinforcement learning agent is set to the negative of the deviation between the output of the thermal power unit, the charging and discharging power of the compressed air energy storage, and the charging and discharging power of the electrochemical energy storage power station at the current time period t and the corresponding day-ahead dispatch plan, and the calculation formula is:
[0064] ;
[0065] Among them, , and are the corresponding deviation coefficients, while , , , , are the corresponding day-ahead dispatch plan values;
[0066] The action taken by the reinforcement learning agent for the state of the distribution network at the current moment Including the charging and discharging power of the compressed air energy storage power station and each electrochemical energy storage power station.
[0067] Based on the above technical solution, the reinforcement learning agent includes a policy neural network and a value network;
[0068] The policy neural network includes a gated recurrent unit and a fully connected layer. The processing process of the policy neural network for the state at each moment is as follows:
[0069] Input the renewable energy power generation sequence and load sequence in the past 24 hours into a gated recurrent unit to output temporal features;
[0070] Input the energy state of each electrochemical energy storage power station at the current time period, the air pressure value in the gas storage chamber of the compressed air energy storage power station, the heat stored in the heat storage tank, the start-stop plan of the thermal power unit's day-ahead scheduling, and the actual output in the previous time period into a fully connected layer network to output state features;
[0071] Input the temporal features and state features into another fully connected layer network, so as to output the charging and discharging states of the compressed air energy storage power station and each electrochemical energy storage power station at the current moment, as well as the mean and standard deviation of the Gaussian distribution of the charging and discharging power.
[0072] Based on the above technical solution, the value network of the reinforcement learning agent includes a fully connected layer, and inputs the state of the distribution network at time period t , and outputs the value in this state .
[0073] Based on the above technical solution, S3 includes the following steps:
[0074] S31. Model parameter initialization:
[0075] Initialize the parameters of the policy network and the value network, and set the learning rate , discount factor and truncation parameter , as well as the smoothing discount coefficient for calculating the advantage function by generalized advantage estimation;
[0076] S32. Data collection:
[0077] The reinforcement learning agent interacts with the distribution network simulation environment for multiple rounds to perform intraday scheduling;
[0078] At each time period t, after inputting the state of the distribution network into the policy neural network, output the charging and discharging states of the compressed air energy storage power station and the electrochemical energy storage power station at the current time period, as well as the mean and standard deviation of the Gaussian distribution of the charging and discharging power;
[0079] After determining the charging and discharging power of each energy storage power station in the current period through random sampling, the reinforcement learning agent clips the charging and discharging power according to the constraints of the energy storage power station to obtain the action , and the reinforcement learning agent adjusts the charging and discharging power of this period according to the action ;
[0080] After determining the charging and discharging power of each energy storage power station, the linear programming method is used to minimize the deviation between the intraday dispatch output of the thermal power unit and the day-ahead dispatch, and determine the output of each thermal power unit in this period and the positive and negative reserve capacities allocated to the thermal power unit and the compressed air energy storage power station;
[0081] Calculate the reward of this period through the reward function calculation expression , and calculate the distribution network state of the next period through the state equation ;
[0082] After multiple rounds of interaction to complete the intraday dispatch of a day, store the state , action , reward and the next state of each time period to form a complete trajectory;
[0083] S33. Policy update:
[0084] Smoothly estimate the advantage function of each time period t through generalized advantage estimation. The advantage function formula is: ;
[0085] Among them, ;
[0086] Calculate the loss function of the policy neural network. The advantage function formula is:
[0087] ;
[0088] Among them, and respectively represent the probabilities of taking action for state under the old and new policies;
[0089] Calculate the loss function of the value network. The loss function formula is as follows:
[0090] ;
[0091] Execute the backpropagation update of the policy network and the value network, and update the network parameters according to the gradient descent method;
[0092] S34. Multi-round iteration: Repeat S32 and S33 until a specific number of iterations is met.
[0093] Based on the above technical solution, S4 includes the following steps:
[0094] S41. At time period t, input the distribution network state into the policy neural network, output the charging and discharging states of each energy storage power station and the mean and standard deviation of the Gaussian distribution of the charging and discharging power, obtain a preliminary charging and discharging strategy through random sampling, and trim the charging and discharging strategy according to the constraint conditions of the energy storage power station in the distribution network model to determine the charging and discharging power of each energy storage power station.
[0095] S42. The start-stop state of the thermal power unit is determined by the day-ahead scheduling plan. With the goal of minimizing the deviation between the intra-day scheduling output of the thermal power unit and the corresponding day-ahead scheduling plan output within time period t, use linear programming method for optimization to determine the intra-day scheduling output of the thermal power unit and the positive and negative reserve capacities allocated to the thermal power unit and the compressed air energy storage power station.
[0096] S43. Execute the intra-day scheduling strategy for time period t and collect the distribution network state information of the next time period , and repeat S41 until the intra-day scheduling is completed.
[0097] Based on the above technical solution, S5 includes the following steps:
[0098] After completing the intra-day scheduling for one day, learn the state , action , reward and the next state for each time period within this day to optimize the scheduling strategy of the energy storage power station and participate in the intra-day scheduling of the distribution network for the next day.
[0099] In a second aspect, the present application also provides a distribution network intra-day scheduling device including a compressed air energy storage system, and the device includes:
[0100] A pre-construction module for constructing a distribution network model and a distribution network simulation environment including a compressed air energy storage system;
[0101] An intelligent agent construction module for constructing a reinforcement learning intelligent agent based on a neural network according to the characteristics of the thermal power unit and the energy storage power station in the distribution network model;
[0102] An intelligent agent training module for interacting with the distribution network simulation environment using the reinforcement learning intelligent agent based on the proximal policy optimization algorithm, so that the reinforcement learning intelligent agent learns to obtain an optimal scheduling strategy and obtain a trained reinforcement learning intelligent agent;
[0103] The day-ahead scheduling module is used to perform day-ahead scheduling of the actual distribution network by using the trained reinforcement learning agent, determine the charging and discharging power of the energy storage power station in each scheduling period, optimize the output and reserve of each thermal power unit, obtain the scheduling strategy for each scheduling period, and complete the day-ahead scheduling for one day;
[0104] The subsequent scheduling module is used to update the scheduling strategy of the trained reinforcement learning agent after completing the day-ahead scheduling for one day, and control the day-ahead scheduling module to perform day-ahead scheduling for a new day.
[0105] Compared with the prior art, the advantages of the present invention are as follows:
[0106] The present invention aims to overcome the deficiencies of the prior art, learn the optimal strategy through the interaction between the agent and the environment, better cope with the intermittency and uncertainty of renewable energy, improve the reliability and stability of the distribution network, and at the same time reduce the operating cost, providing new ideas and methods for the efficient operation of the distribution network. Brief Description of the Drawings
[0107] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0108] Figure 1 It is a flowchart of the steps of the method for day-ahead scheduling of a distribution network with a compressed air energy storage system according to an embodiment of the present invention;
[0109] Figure 2 It is a principle flowchart of the method for day-ahead scheduling of a distribution network with a compressed air energy storage system according to an embodiment of the present invention;
[0110] Figure 3 It is a strategy update process of the reinforcement learning agent in the method for day-ahead scheduling of a distribution network with a compressed air energy storage system according to an embodiment of the present invention;
[0111] Figure 4 It is a structural block diagram of the device for day-ahead scheduling of a distribution network with a compressed air energy storage system according to an embodiment of the present invention. Detailed Embodiments
[0112] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0113] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0114] The embodiments of this application provide a method and device for intraday scheduling of a distribution network with a compressed air energy storage system, aiming to overcome the deficiencies of the prior art. By the interaction between the agent and the environment to learn the optimal strategy, it can better cope with the intermittency and uncertainty of renewable energy, improve the reliability and stability of the distribution network, and at the same time reduce the operating cost, providing new ideas and methods for the efficient operation of the distribution network.
[0115] To achieve the above technical effects, the general idea of this application is as follows:
[0116] A method for intraday scheduling of a distribution network with a compressed air energy storage system, the method includes the steps of:
[0117] S1, constructing a distribution network model with a compressed air energy storage system and a distribution network simulation environment;
[0118] S2, based on the characteristics of thermal power units and energy storage power stations in the distribution network model, constructing a reinforcement learning agent based on a neural network;
[0119] S3, based on the proximal policy optimization algorithm, using the reinforcement learning agent to interact with the distribution network simulation environment so that the reinforcement learning agent learns to obtain the optimal scheduling strategy and obtain the trained reinforcement learning agent;
[0120] S4, using the trained reinforcement learning agent to perform intraday scheduling of the actual distribution network, determining the charge and discharge power of the energy storage power station in each scheduling period, optimizing the output and standby of each thermal power unit, obtaining the scheduling strategy for each scheduling period, and completing the intraday scheduling for one day;
[0121] S5, after completing the intraday scheduling for one day, updating the scheduling strategy of the trained reinforcement learning agent, repeating S4, and performing intraday scheduling for a new day.
[0122] The embodiments of this application will be further described in detail below with reference to the accompanying drawings.
[0123] In the first aspect, see Figures 1 to 3As shown in the figure, an embodiment of the present application provides a method for intraday scheduling of a distribution network with a compressed air energy storage system. The method includes the following steps:
[0124] S1. Build a distribution network model with a compressed air energy storage system and a distribution network simulation environment;
[0125] S2. Based on the characteristics of thermal power units and energy storage power stations in the distribution network model, build a reinforcement learning agent based on a neural network;
[0126] S3. Based on the proximal policy optimization algorithm, use the reinforcement learning agent to interact with the distribution network simulation environment so that the reinforcement learning agent learns to obtain an optimal scheduling strategy and obtain a trained reinforcement learning agent;
[0127] S4. Use the trained reinforcement learning agent to perform intraday scheduling of the actual distribution network, determine the charging and discharging power of the energy storage power station at each scheduling period, optimize the output and standby of each thermal power unit, obtain the scheduling strategy for each scheduling period, and complete the intraday scheduling for one day;
[0128] S5. After completing the intraday scheduling for one day, update the scheduling strategy of the trained reinforcement learning agent, repeat S4, and perform intraday scheduling for a new day.
[0129] In the embodiment of the present application, it aims to overcome the deficiencies of the prior art, learn the optimal strategy through the interaction between the agent and the environment, better cope with the intermittency and uncertainty of renewable energy, improve the reliability and stability of the distribution network, and at the same time reduce the operating cost, providing new ideas and methods for the efficient operation of the distribution network.
[0130] Further, the distribution network model includes multiple thermal power units, a wind power station, a photovoltaic power station, multiple electrochemical energy storage power stations, a compressed air energy storage power station, and a load.
[0131] Further, the compression power constraint of the compressor in the compressed air energy storage power station is: ;
[0132] Wherein, is the power of the compressor at time t, and and are the upper and lower limits of the compressor power respectively, is a binary variable representing the operating state of the compressor at time t, being 1 indicates that the compressor is in the working state, being 0 indicates that the compressor is in the shutdown state;
[0133] The expansion power constraint of the expander in the compressed air energy storage power station is:
[0134] ;
[0135] Among them, is the power of the expander in the t period, and and are the upper and lower limits of the expander power respectively, is a binary variable representing the operating state of the expander in the t period, being 1 indicates that the expander is in the working state, being 0 indicates that the expander is in the shutdown state;
[0136] Limited by the internal structure of the compressed air energy storage system, the operating state constraints of the compressor and the expander are: ;
[0137] The mathematical model and constraint conditions of the gas storage chamber in the compressed air energy storage power station are:
[0138] ;
[0139] ;
[0140] ;
[0141] ;
[0142] Among them, is the air pressure value of the gas storage chamber in the t period, which is respectively limited by the lower limit of the air pressure value and the upper limit of the air pressure value , and and are the air pressure rise rate and the air pressure drop rate of the gas storage chamber in the t period respectively, which are respectively proportional to the compression power , the expansion power , and their proportionality coefficients are and respectively, is a scheduling period;
[0143] The mathematical model and constraint conditions of the heat storage tank in the compressed air energy storage power station are:
[0144] ;
[0145] ;
[0146] ;
[0147] ;
[0148] Among them, The stored heat quantity of the heat storage tank in period t is respectively subject to the lower limit of the stored heat quantity and the upper limit of the stored heat quantity constraints, and is the heat generation power of the heat exchanger in the air compression stage, which is proportional to the compression power in the same period, and the proportionality coefficient is , is the heat absorption power of the heat exchanger in the air expansion stage, which is proportional to the expansion power in the same period, and the proportionality coefficient is ;
[0149] The reserve constraints in the compressed air energy storage power station are as follows:
[0150] ;
[0151]
[0152] wherein, and are respectively the positive reserve capacity and the negative reserve capacity of the compressed air energy storage system in the compressed charging state, and are respectively the positive reserve capacity and the negative reserve capacity of the compressed air energy storage system in the expansion discharge state;
[0153] The mathematical model and constraints of the electrochemical energy storage power station are as follows:
[0154] ;
[0155] ;
[0156] ;
[0157] wherein, is the energy stored by the i-th electrochemical energy storage power station in period t, and and are respectively its lower limit and upper limit, and are respectively the charging power and the discharging power of the i-th electrochemical energy storage power station in period t, and and are respectively the power upper limits when the i-th electrochemical energy storage power station is in the charging or discharging state, and are respectively the charging and discharging efficiencies of the i-th electrochemical energy storage power station. The binary variable is used to represent the working state of the i-th electrochemical energy storage power station in period t, being 1 indicates the charging state, being 0 indicates the discharging state;
[0158] The power constraint of the thermal power unit is as follows:
[0159] ;
[0160] ;
[0161] ;
[0162] where, represents the power of thermal power unit j at time period t, which is restricted by the upper limit and the lower limit in the operating state, while is a binary variable representing the operating state of thermal power unit j at time period t, being 1 indicates that thermal power unit j is in the operating and startup state, being 0 indicates the shutdown state, which is determined by the day-ahead dispatch plan, and are the limits of the upper and lower ramping rates of thermal power unit j respectively;
[0163] The reserve constraint of the thermal power unit is as follows:
[0164] ;
[0165] ;
[0166] where, and represent the positive reserve capacity and negative reserve capacity of thermal power unit j at time period t respectively;
[0167] The renewable energy output constraint is as follows: ;
[0168] where, is the actual dispatched output of renewable energy at time period t, while is the maximum output of renewable energy at time period t;
[0169] The power balance constraint of the distribution network system is as follows:
[0170] ;
[0171] where, is the load at time period t, is the number of electrochemical energy storage power stations, is the number of thermal power units;
[0172] The reserve constraint of the distribution network system is as follows:
[0173] ;
[0174] The lower limit value of the spare capacity of the distribution network system is linearly related to the load and the output of renewable energy , and their ratios are respectively and .
[0175] Furthermore, in step S2:
[0176] The constructed reinforcement learning agent outputs an action at each time period t according to the state of the distribution network in the current time period and the reward calculated through the reward function to adjust the charging and discharging power of the compressed air energy storage power station and each electrochemical energy storage power station;
[0177] The state of the distribution network at the current moment includes the renewable energy power generation sequence and load sequence in the past 24 hours, the energy state of each electrochemical energy storage power station in the current time period, the air chamber air pressure value of the compressed air energy storage power station in the current time period, the heat stored in the heat storage tank of the compressed air energy storage power station in the current time period, the start-stop plan of each thermal power unit's day-ahead scheduling, and the actual output in the previous time period;
[0178] The reward function for the reinforcement learning agent is set as the negative of the deviation between the output of the thermal power unit, the charging and discharging power of the compressed air energy storage, and the charging and discharging power of the electrochemical energy storage power station and the corresponding day-ahead scheduling plan in the current time period t, and the calculation formula is:
[0179] ;
[0180] wherein, , and are the corresponding deviation coefficients, while , , , , are the corresponding day-ahead scheduling plan values;
[0181] The action taken by the reinforcement learning agent for the state of the distribution network at the current moment includes the charging and discharging power of the compressed air energy storage power station and each electrochemical energy storage power station.
[0182] Furthermore, the reinforcement learning agent includes a policy neural network and a value network;
[0183] The policy neural network includes a gated recurrent unit and a fully connected layer. The processing process of the policy neural network for the state at each moment is as follows:
[0184] Input the renewable energy power generation sequence and load sequence in the past 24 hours into a gated recurrent unit, and output the time series features;
[0185] Input the energy state of each electrochemical energy storage power station, the air pressure value in the gas storage chamber of the compressed air energy storage power station, the heat stored in the heat storage tank, the start-stop plan of the thermal power unit's day-ahead scheduling, and the actual output in the previous period into a fully connected layer network, and output the state features;
[0186] Input the time series features and state features into another fully connected layer network, so as to output the charge-discharge states of the compressed air energy storage power station and each electrochemical energy storage power station at the current moment, as well as the mean and standard deviation of the Gaussian distribution of the charge-discharge power.
[0187] Furthermore, the value network of the reinforcement learning agent includes a fully connected layer, and inputs the state of the distribution network at time t , and outputs the value in this state .
[0188] Furthermore, the S3 includes the following steps:
[0189] S31. Model parameter initialization:
[0190] Initialize the parameters of the policy network and the value network, and set the learning rate , discount factor and truncation parameter , as well as the smoothing discount coefficient for calculating the advantage function used in generalized advantage estimation ;
[0191] S32. Data collection:
[0192] The reinforcement learning agent interacts with the distribution network simulation environment for multiple rounds to perform intraday scheduling;
[0193] At each time period t, after inputting the state of the distribution network into the policy neural network, output the charge-discharge states of the compressed air energy storage power station and the electrochemical energy storage power station at the current period, as well as the mean and standard deviation of the Gaussian distribution of the charge-discharge power;
[0194] After the reinforcement learning agent determines the charge-discharge power of each energy storage power station at the current period through random sampling, clip the charge-discharge power according to the constraint conditions of the energy storage power station, so as to obtain the action , and the reinforcement learning agent adjusts the charge-discharge power of this period according to the action ;
[0195] After determining the charge and discharge powers of each energy storage power station, a linear programming method is used to minimize the deviation between the intraday dispatch output of thermal power units and the day-ahead dispatch, and determine the output of each thermal power unit during this period and the positive and negative reserve capacities allocated to thermal power units and compressed air energy storage power stations;
[0196] Calculate the reward for this period through the reward function calculation expression and calculate the distribution network state of the next period through the state equation ;
[0197] After completing the intraday dispatch of a day through multiple rounds of interaction, store the state , action , reward and the next state of each time period to form a complete trajectory;
[0198] As shown in Figure 3 of the accompanying drawings of the specification, S33, Policy Update:
[0199] Smoothly estimate the advantage function of each time period t through generalized advantage estimation. The formula for the advantage function is:
[0200] ;
[0201] Among them, ;
[0202] Calculate the loss function of the policy neural network. The formula for the advantage function is:
[0203] ;
[0204] Among them, and respectively represent the probabilities of taking action for state under the old and new policies;
[0205] Calculate the loss function of the value network. The formula for the loss function is as follows:
[0206] ;
[0207] Execute the backpropagation update of the policy network and the value network, and update the network parameters according to the gradient descent method;
[0208] S34, Multiple Rounds of Iteration: Repeat S32 and S33 until a specific number of iterations is satisfied.
[0209] Furthermore, the S4 includes the following steps:
[0210] S41. At time period t, input the distribution network state into the policy neural network, output the charging and discharging states of each energy storage power station and the mean and standard deviation of the Gaussian distribution of the charging and discharging power, obtain a preliminary charging and discharging strategy through random sampling, and trim the charging and discharging strategy according to the constraint conditions of the energy storage power station in the distribution network model to determine the charging and discharging power of each energy storage power station;
[0211] S42. The start-stop state of the thermal power unit is determined by the day-ahead scheduling plan. Taking the minimum deviation between the intra-day scheduling output of the thermal power unit and the corresponding day-ahead scheduling plan output within time period t as the optimization objective, use the linear programming method for optimization to determine the intra-day scheduling output of the thermal power unit and the positive and negative reserve capacities allocated to the thermal power unit and the compressed air energy storage power station;
[0212] S43. Execute the intra-day scheduling strategy for time period t and collect the distribution network state information for the next time period , and repeat S41 until the intra-day scheduling is completed.
[0213] Furthermore, the S5 includes the following steps:
[0214] After completing the intra-day scheduling for one day, learn the state , action , reward and the next state for each time period within this day to optimize the scheduling strategy of the energy storage power station and participate in the intra-day scheduling of the distribution network for the next day.
[0215] In summary, the technical solution of the embodiment of the present application can be used for the scheduling of smart grids, realizing the efficient integration of renewable energy such as wind energy and solar energy, and improving the stability and reliability of the power grid through the intelligent scheduling of the compressed air energy storage system, and is particularly applicable to the power grid system with a high proportion of renewable energy grid connection;
[0216] It aims to solve the intra-day scheduling problem of the distribution network containing compressed air energy storage, cope with the intermittency and uncertainty of renewable energy (wind power and photovoltaic power), improve the reliability and stability of the distribution network, and at the same time reduce the operating cost of the distribution network.
[0217] Specifically, the following problems need to be solved:
[0218] First, the quality of the scheduling scheme of the traditional scheduling method based on simple rules is not high, and it cannot give full play to the role of complex constraint conditions and objective functions.
[0219] Second, the existing scheduling method based on model predictive control has a high computational complexity and depends on prediction data, and has poor adaptability to the uncertainty of renewable energy and load.
[0220] Furthermore, the beneficial effects of the technical solutions in the embodiments of the present application are as follows:
[0221] First, improve the reliability and stability of the distribution network: The method in the embodiments of the present application can effectively adjust the charge and discharge strategies of the compressed air energy storage system, making the distribution network more stable in the face of fluctuations in renewable energy sources such as wind power and photovoltaic power, and significantly reducing the grid operation risks brought by the uncertainty of renewable energy.
[0222] Second, flexibly adapt to complex constraint conditions: By adopting the proximal policy optimization algorithm, this method can adapt to multiple complex constraint conditions in grid scheduling, such as energy storage system capacity limitations, charge and discharge rate limitations, and start-stop constraints of generator sets, thus ensuring the feasibility and optimality of the scheduling plan.
[0223] Third, enhance the system's intelligent level: Through the introduction of the reinforcement learning model in the embodiments of the present application, the compressed air energy storage system can obtain the optimal scheduling strategy in continuous environmental interactions, with certain adaptive and learning abilities, and can automatically adjust the charge and discharge strategies during operation, further improving the intelligent level of scheduling.
[0224] In the second aspect, as shown in Figure 4 The embodiments of the present application provide an intraday scheduling device for a distribution network containing a compressed air energy storage system, and the device includes:
[0225] A pre-construction module, which is used to construct a distribution network model and a distribution network simulation environment containing a compressed air energy storage system;
[0226] An intelligent agent construction module, which is used to construct a reinforcement learning intelligent agent based on a neural network according to the characteristics of thermal power units and energy storage power stations in the distribution network model;
[0227] An intelligent agent training module, which is used to interact with the distribution network simulation environment using the reinforcement learning intelligent agent based on the proximal policy optimization algorithm, so that the reinforcement learning intelligent agent learns to obtain the optimal scheduling strategy and obtains the trained reinforcement learning intelligent agent;
[0228] An intraday scheduling module, which is used to perform intraday scheduling of the actual distribution network using the trained reinforcement learning intelligent agent, determine the charge and discharge power of the energy storage power station in each scheduling period, optimize the output and standby of each thermal power unit, obtain the scheduling strategy for each scheduling period, and complete the intraday scheduling for one day;
[0229] A subsequent scheduling module, which is used to update the scheduling strategy of the trained reinforcement learning intelligent agent after completing the intraday scheduling for one day, and control the intraday scheduling module to perform intraday scheduling for a new day.
[0230] In the embodiments of the present application, it aims to overcome the deficiencies of the prior art, learn the optimal strategy through the interaction between the agent and the environment, better cope with the intermittency and uncertainty of renewable energy, improve the reliability and stability of the distribution network, while reducing the operating cost, and provide new ideas and methods for the efficient operation of the distribution network.
[0231] It should be noted that the specific working content of each module in the intraday scheduling device of the distribution network with a compressed air energy storage system is as described in the corresponding content of each step in the intraday scheduling method of the distribution network with a compressed air energy storage system mentioned in the first aspect. Each module is used to execute the corresponding step and the sub-steps within that step.
[0232] It should be noted that the intraday scheduling device of the distribution network with a compressed air energy storage system mentioned in the second aspect is similar in terms of technical problems, technical means, and technical effects to the technical principle of the intraday scheduling method of the distribution network with a compressed air energy storage system mentioned in the first aspect, and will not be elaborated here.
[0233] In the description of the present application, it should be noted that the orientation or positional relationship indicated by terms such as "upper" and "lower" is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present application. Unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0234] It should be noted that in the present application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0235] The above description is only a specific implementation manner of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A day-ahead dispatching method for a distribution network with a compressed air energy storage system, characterized in that, The method includes the following steps: S1. Construct a distribution network model and a distribution network simulation environment including a compressed air energy storage system; S2. Based on the characteristics of thermal power units and energy storage power stations in the distribution network model, construct a reinforcement learning agent based on a neural network; S3. Based on the proximal policy optimization algorithm, use the reinforcement learning agent to interact with the distribution network simulation environment so that the reinforcement learning agent learns to obtain an optimal scheduling strategy and obtain a trained reinforcement learning agent; S4. Use the trained reinforcement learning agent to perform intraday scheduling of the actual distribution network, determine the charging and discharging power of the energy storage power station at each scheduling period, optimize the output and reserve of each thermal power unit, obtain the scheduling strategy for each scheduling period, and complete the intraday scheduling for one day; S5. After completing the intraday scheduling for one day, update the scheduling strategy of the trained reinforcement learning agent, repeat S4, and perform intraday scheduling for a new day; Among them, the distribution network model includes multiple thermal power units, wind power generation stations, photovoltaic power generation stations, multiple electrochemical energy storage power stations, a compressed air energy storage power station, and loads.
2. The intraday scheduling method for a distribution network with a compressed air energy storage system according to claim 1, characterized in that: The compression power constraint of the compressor in the compressed air energy storage power station is: ; Among them, is the power of the compressor at time period t, and as well as are the lower power limit and the upper power limit of the compressor respectively, is a binary variable representing the operating state of the compressor at time period t, being 1 indicates that the compressor is in the working state, being 0 indicates that the compressor is in the shutdown state; The expansion power constraint of the expander in the compressed air energy storage power station is: ; Among them, is the power of the expander at time period t, and and are the lower power limit and the upper power limit of the expander respectively, is a binary variable representing the operating state of the expander at time period t, being 1 indicates that the expander is in the working state, being 0 indicates that the expander is in the shutdown state; Limited by the internal structure of the compressed air energy storage system, the operating states of the compressor and the expander are constrained as follows: ; The mathematical model and constraint conditions of the gas storage chamber in the compressed air energy storage power station are: ; ; ; ; Among them, is the air pressure value of the gas storage chamber at time t, which is respectively restricted by the lower limit of the air pressure value and the upper limit of the air pressure value while and are respectively the air pressure rising rate and the air pressure falling rate of the gas storage chamber at time t, which are respectively proportional to the compression power and the expansion power at the same time period, and the proportionality coefficients are respectively and , is a scheduling period; The mathematical model and constraint conditions of the heat storage tank in the compressed air energy storage power station are: ; ; ; ; Among them, is the heat storage capacity of the heat storage tank during period t, which is respectively restricted by the lower limit of the heat storage capacity and the upper limit of the heat storage capacity while is the heat generation power of the heat exchanger during the air compression stage, which is proportional to the compression power during the same period, and the proportionality coefficient is . is the heat absorption power of the heat exchanger during the air expansion stage, which is proportional to the expansion power during the same period, and the proportionality coefficient is ; The reserve constraint in the compressed air energy storage power station is: ; ; Among them, and are the positive reserve capacity and the negative reserve capacity of the compressed air energy storage system in the compressed charging state respectively, and are the positive reserve capacity and the negative reserve capacity of the compressed air energy storage system in the expansion discharge state respectively; The mathematical model and constraints of the electrochemical energy storage power station are: ; ; ; Among them, is the energy stored in the i-th electrochemical energy storage power station at time t, while and are its lower and upper limits respectively, and are the charging power and discharging power of the i-th electrochemical energy storage power station at time t respectively, while and are the upper limits of power when the i-th electrochemical energy storage power station is in the charging or discharging state respectively, and are the charging and discharging efficiencies of the i-th electrochemical energy storage power station respectively. By using the binary variable to represent the working state of the i-th electrochemical energy storage power station at time t, being 1 indicates the charging state, being 0 indicates the discharging state; The power constraint of the thermal power unit is: ; ; ; Among them, represents the power of thermal power unit j at time period t, which is subject to the upper limit and the lower limit constraints in the operating state, while is a binary variable representing the operating state of thermal power unit j at time period t. Being 1 indicates that thermal power unit j is in the operating and starting state, being 0 indicates the shutdown state, which is determined by the day-ahead scheduling plan. and are the upper and lower ramping rate limits of thermal power unit j respectively; The reserve constraint of the thermal power unit is: ; ; Among them, and respectively represent the positive reserve capacity and negative reserve capacity of thermal power unit j at time t; The renewable energy output constraints are as follows: ; Among them, is the actual dispatching output of renewable energy in period t, while is the maximum output of renewable energy in period t; The power balance constraint of the distribution network model is: ; wherein, is the load at time period t, is the number of electrochemical energy storage power stations, is the number of thermal power units; The reserve constraint of the distribution network model is: ; The lower limit value of the reserve capacity of the distribution network model and the load and the output of renewable energy are linearly related, and their proportions are respectively and .
3. The day-ahead dispatching method for a distribution network with a compressed air energy storage system according to claim 1, wherein In step S2: The constructed reinforcement learning agent, based on the state of the distribution network in each time period t and the reward calculated through the reward function , outputs an action to adjust the charging and discharging powers of the compressed air energy storage power station and each electrochemical energy storage power station; Distribution network status at the current moment Including the renewable energy generation sequence and load sequence in the past 24 hours, the energy status of each electrochemical energy storage power station in the current period, the air pressure value in the gas storage chamber of the compressed air energy storage power station in the current period, the heat stored in the heat storage tank of the compressed air energy storage power station in the current period, the start-stop plan of each thermal power unit scheduled the day before, and the actual output in the previous period; Reward function for the reinforcement learning agent It is set as the negative of the deviation between the output of the thermal power unit, the charging and discharging power of the compressed air energy storage, and the charging and discharging power of the electrochemical energy storage power station at the current time period t and the corresponding day-ahead dispatch plan. The calculation formula is as follows: ; Among them, , and are the corresponding deviation coefficients, while , , , are the corresponding day-ahead scheduling plan values; The actions taken by the reinforcement learning agent for the current state of the distribution network include the charging and discharging powers of the compressed air energy storage power station and each electrochemical energy storage power station.
4. The intraday scheduling method for a distribution network with a compressed air energy storage system according to claim 3, characterized in that: The reinforcement learning agent includes a policy neural network and a value network; The policy neural network includes a gated recurrent unit and a fully connected layer. The processing process of the policy neural network for the state at each moment is as follows: Input the renewable energy generation sequence and load sequence of the past 24 hours into a gated recurrent unit to output temporal features; Input the energy state of each electrochemical energy storage power station at the current period, the air pressure value of the gas storage chamber of the compressed air energy storage power station, the heat stored in the heat storage tank, the start-stop plan of the thermal power unit for day-ahead scheduling, and the actual output of the previous period into a fully connected layer network to output state features; Input the temporal features and state features into another fully connected layer network together, so as to output the charging and discharging states of the compressed air energy storage power station and each electrochemical energy storage power station at the current moment, as well as the mean and standard deviation of the Gaussian distribution of the charging and discharging power.
5. The intraday scheduling method for a distribution network with a compressed air energy storage system according to claim 3, characterized in that: The value network of the reinforcement learning agent includes a fully connected layer that takes as input the state of the distribution network at time period t and outputs the value in this state .
6. The intraday dispatching method for a distribution network with a compressed air energy storage system according to claim 1, characterized in that, S3 includes the following steps: S31. Initialize model parameters: Initialize the parameters of the policy network and the value network, and set the learning rate , the discount factor , and the truncation parameter , as well as the smoothing discount coefficient for calculating the advantage function in generalized advantage estimation ; S32. Collect data: The reinforcement learning agent interacts with the distribution network simulation environment for multiple rounds to perform intraday scheduling; At each time period t, the distribution network state is input into the policy neural network, and the charging and discharging states of the compressed air energy storage power station and the electrochemical energy storage power station at the current time period, as well as the mean and standard deviation of the Gaussian distribution of the charging and discharging power, are output; After determining the charging and discharging power of each energy storage power station in the current period through random sampling, the reinforcement learning agent clips the charging and discharging power according to the constraint conditions of the energy storage power station to obtain an action , and the reinforcement learning agent adjusts the charging and discharging power of this period according to the action ; After determining the charging and discharging powers of each energy storage power station, a linear programming method is used to minimize the deviation between the intraday scheduling output of the thermal power unit and the day-ahead scheduling, and determine the output of each thermal power unit during this period and the positive and negative reserve capacities allocated to the thermal power units and compressed air energy storage power stations; Calculate the reward for this period through the reward function calculation expression and calculate the distribution network state for the next period through the state equation ; After completing the intraday scheduling for a day through multiple rounds of interaction, store the status of each time period , action , reward and the next state to form a complete trajectory; S33. Policy update: Smoothly estimate the advantage function for each time period t through generalized advantage estimation. The formula for the advantage function is: ; Among them, ; Calculate the loss function of the policy neural network. The formula for the advantage function is: ; Among them, and respectively represent the probabilities of taking action for state under the new and old policies; Calculate the loss function of the value network. The loss function formula is as follows: ; Execute the backpropagation update of the policy network and the value network, and update the network parameters according to the gradient descent method; S34. Multiple rounds of iteration: Repeat S32 and S33 until a specific number of iterations is satisfied.
7. The day-ahead dispatching method for a distribution network with a compressed air energy storage system according to claim 1, characterized in that, The said S4 includes the following steps: S41. At time period t, input the distribution network state into the policy neural network, output the charging and discharging states of each energy storage power station and the mean and standard deviation of the Gaussian distribution of the charging and discharging power, obtain a preliminary charging and discharging strategy through random sampling, and trim the charging and discharging strategy according to the constraint conditions of the energy storage power station in the distribution network model to determine the charging and discharging power of each energy storage power station. The start-stop state of the thermal power unit is determined by the day-ahead scheduling plan. With the goal of minimizing the deviation between the intraday scheduling output of the thermal power unit and the corresponding day-ahead scheduling plan output during time period t, use the linear programming method for optimization to determine the intraday scheduling output of the thermal power unit and the positive and negative reserve capacities allocated to the thermal power unit and the compressed air energy storage power station; Execute the intraday scheduling strategy for time period t and collect the distribution network status information for the next time period. , repeat S41 until the intraday scheduling is completed.
8. The intraday dispatching method for a distribution network with a compressed air energy storage system according to claim 1, wherein The said S5 includes the following steps: After completing the day-ahead dispatch, the status of each time period within the day , actions , rewards and the next status are used for learning to optimize the dispatch strategy of the energy storage power station and participate in the day-ahead dispatch of the distribution network on the next day.
9. A device for intraday dispatching of a distribution network with a compressed air energy storage system, characterized in that, The device includes: A pre-construction module, which is used to construct a distribution network model and a distribution network simulation environment including a compressed air energy storage system; An agent construction module, which is used to construct a reinforcement learning agent based on a neural network according to the characteristics of the thermal power units and energy storage power stations in the distribution network model; An agent training module, which is used to interact with the distribution network simulation environment using the reinforcement learning agent based on the proximal policy optimization algorithm, so that the reinforcement learning agent learns to obtain an optimal scheduling strategy and obtain a trained reinforcement learning agent; An intraday scheduling module, which is used to perform intraday scheduling of the actual distribution network using the trained reinforcement learning agent, determine the charging and discharging powers of the energy storage power stations at each scheduling time period, optimize the output and reserves of each thermal power unit, obtain the scheduling strategy for each scheduling time period, and complete the intraday scheduling for one day; A subsequent scheduling module, which is used to update the scheduling strategy of the trained reinforcement learning agent after completing the intraday scheduling for one day, and control the intraday scheduling module to perform intraday scheduling for a new day; Among them, the distribution network model includes multiple thermal power units, wind power generation stations, photovoltaic power generation stations, multiple electrochemical energy storage power stations, compressed air energy storage power stations, and loads.
Citation Information
Patent Citations
Novel power system dispatching optimization method based on reinforcement learning
CN118040669A
Dynamic co-optimization management for grid scale energy storage system (GSESS) market participation
US20160042369A1