A wind-solar-thermal storage system optimization configuration method and system based on time series difference method

By using a reinforcement learning model based on temporal difference method and SARSA algorithm, the operating status of thermal power units is dynamically adjusted, the configuration of wind, solar, thermal and energy storage system is optimized, the problem of low-load power curtailment of thermal power units is solved, operating costs are reduced and the utilization rate of wind and solar resources is improved.

CN114331025BActive Publication Date: 2025-10-24HUANENG CLEAN ENERGY RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111473491.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-10-24
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

In existing wind-solar-thermal-storage systems, thermal power units often waste electricity during low-load operation, resulting in high operating costs and low utilization of wind and solar resources.

Method used

A reinforcement learning model based on temporal difference method and SARSA algorithm is adopted to dynamically adjust the operating status of thermal power units. By optimizing the configuration method and system, the optimal strategy is determined to reduce the cumulative investment and operating costs and improve the utilization rate of wind and solar resources.

Benefits of technology

By dynamically adjusting the operating status of thermal power units and using the time-series difference algorithm, the cumulative investment and operating costs of the wind-solar-thermal-storage system are reduced under limited sampling conditions, while the utilization rate of wind and solar resources is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114331025B_ABST
    Figure CN114331025B_ABST
Patent Text Reader

Abstract

The application provides a wind-solar-thermal-storage system optimization configuration method and system based on a time sequence difference method. The method comprises the following steps: firstly, obtaining power grid demand power generation, wind turbine power generation, photovoltaic power generation, a preset constraint condition and an economic parameter at each moment in a historical period; secondly, determining total demand power generation of thermal power units and energy storage devices in the system based on the obtained data; then, dividing the state type of the system and training a reinforcement learning model established based on the wind-solar-thermal-storage integrated system based on a SARSA algorithm to obtain an optimal strategy of the system under different states; subsequently, calculating the cumulative operation cost of the system in a given period based on the optimal strategy; finally, modifying the preset constraint condition, selecting the preset constraint condition corresponding to the minimum cumulative investment and operation cost of the system under different constraints to optimize and configure the system. The technical scheme provided by the application improves the utilization rate of wind and light resources and saves operation cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of optimal configuration of systems, and particularly relates to a wind-solar-thermal-storage system optimal configuration method and system based on time series difference method. BACKGROUND

[0002] With the development of new energy, the wind-solar complementary power generation system is more and more widely used. However, the wind-solar complementary power generation system is greatly affected by climate and environment. The wind-solar-thermal-storage system established by using large-scale energy storage technology can ensure the continuity and reliability of load power consumption and reduce the waste of energy resources.

[0003] At present, the existing wind-solar-thermal-storage system defaults the continuous operation of thermal power generating units. When the thermal power generating units are not needed to output, the units are operated under the minimum load condition. Long-term low-load operation will cause the power to exceed the capacity of the energy storage device and the power is abandoned, which increases the operation cost and reduces the utilization rate of wind and light resources. SUMMARY

[0004] The present application provides a wind-solar-thermal-storage system optimal configuration method and system based on time series difference method to at least solve the technical problems of low utilization rate of wind and light resources and high operation cost in the related art.

[0005] The first aspect embodiment of the present application provides a wind-solar-thermal-storage system optimal configuration method based on time series difference method, which comprises the following steps:

[0006] obtaining the power grid demand power generation, the wind power generation of the wind power generating unit in the wind-solar-thermal-storage system, the photovoltaic power generation of the photovoltaic generating unit, a preset constraint condition and an economic parameter in each moment in a historical period;

[0007] determining the total demand power generation of the thermal power generating unit and the energy storage device in the wind-solar-thermal-storage system in each moment in the historical period according to the power grid demand power generation, the wind power generation of the wind power generating unit and the photovoltaic power generation of the photovoltaic generating unit in each moment in the historical period;

[0008] dividing the state of the wind-solar-thermal-storage system into different state types based on the operation state of the thermal power generating unit and the available power of the energy storage device, and randomly initializing the probability value of mutual transfer between states and the strategy corresponding to each state type;

[0009] establishing a reinforcement learning model based on the SARSA algorithm, taking the total demand power generation of the thermal power generating unit and the energy storage device in the wind-solar-thermal-storage system in each moment in the historical period as a sampling sequence, training the model, and obtaining an optimal strategy;

[0010] calculating the state of the wind-solar-thermal-storage system in each moment in a given period and the operation cost of the wind-solar-thermal-storage system corresponding to the state according to the optimal strategy, so as to calculate the cumulative investment and operation cost of the wind-solar-thermal-storage system in the given period;

[0011] recomputing the optimal strategy of each state and the cumulative investment operation cost of the system under the preset constraint condition within a given period, screening a minimum value from the cumulative investment operation cost of the system under different constraints, and optimizing the configuration of the wind-solar-thermal-storage system by using the preset constraint condition corresponding to the minimum value.

[0012] The preset constraint condition includes a capacity constraint of each power generation and energy storage device, a state constraint, and an initial state of the wind-solar-thermal-storage system.

[0013] The second aspect of the present application provides a wind-solar-thermal-storage system optimization configuration system based on a time sequence difference method, which comprises:

[0014] The acquisition module is configured to acquire the power grid demand power generation, the wind turbine power generation, the photovoltaic power generation, the preset constraint condition, and the economic parameter in the wind-solar-thermal-storage system at each time point in a historical period.

[0015] The determination module is configured to determine the total demand power generation of the thermal power unit and the energy storage device in the wind-solar-thermal-storage system at each time point in the historical period according to the power grid demand power generation, the wind turbine power generation, and the photovoltaic power generation in the wind-solar-thermal-storage system at each time point in the historical period.

[0016] The initialization module is configured to divide the state of the wind-solar-thermal-storage system into different state types based on the operating state of the thermal power unit and the available power of the energy storage device, and randomly initialize the probability value of mutual transfer between states and the strategy corresponding to each state type.

[0017] The optimal strategy module is configured to establish a reinforcement learning model based on the SARSA algorithm, take the total demand power generation of the thermal power unit and the energy storage device in the wind-solar-thermal-storage system at each time point in the historical period as a sampling sequence, train the model, and obtain the optimal strategy.

[0018] The calculation module is configured to calculate the state of the wind-solar-thermal-storage system at each time point within a given period and the operation cost of the system corresponding to the state according to the optimal strategy, so as to calculate the cumulative investment operation cost of the wind-solar-thermal-storage system within the given period.

[0019] The optimization configuration module is configured to modify the preset constraint condition, recompute the optimal strategy of each state and the cumulative investment operation cost of the system under the preset constraint condition within a given period, screen a minimum value from the cumulative investment operation cost of the system under different constraints, and optimize the configuration of the wind-solar-thermal-storage system by using the preset constraint condition corresponding to the minimum value.

[0020] The preset constraint condition includes a capacity constraint of each power generation and energy storage device, a state constraint, and an initial state of the wind-solar-thermal-storage system.

[0021] The third aspect of the present application provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the method of the first aspect of the present application when executing the computer program.

[0022] The fourth aspect of the present application provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect of the present application.

[0023] The embodiments of the present application provide at least the following beneficial effects:

[0024] To sum up, the present application provides a wind-solar-thermal-storage system optimization configuration method and system based on time series difference method, the method comprising: first, obtaining the power grid demand power generation, wind turbine power generation, photovoltaic power generation, preset constraint conditions and economic parameters at each time in the historical period; second, determining the total demand power generation of the thermal power unit and the energy storage device in the system based on the obtained data; third, dividing the state type of the system and training the reinforcement learning model established based on the wind-solar-thermal-storage integrated system based on the SARSA algorithm to obtain the optimal strategy of the system in different states; fourth, calculating the cumulative operating cost of the system in a given period based on the optimal strategy; and fifth, modifying the preset constraint condition, selecting the preset constraint condition corresponding to the minimum cumulative investment and operating cost of the system under different constraints to optimize and configure the system. The technical scheme provided by the present application can dynamically adjust the operating state of the thermal power unit, and utilize the time series difference algorithm to reduce the cumulative investment and operating cost of the integrated system in a given period as much as possible under the condition of limited sampling quantity, while improving the utilization rate of wind and light resources.

[0025] Additional aspects and advantages of the present application will be made apparent by the following description and the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0026] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description, taken in conjunction with the accompanying drawings, in which:

[0027] Figure 1 A flowchart of a wind-solar-thermal-storage system optimization configuration method based on time series difference method according to an embodiment of the present application is provided.

[0028] Figure 2 A state transition diagram according to an embodiment of the present application is provided.

[0029] Figure 3A specific flow chart of a wind-solar-thermal-storage system optimization configuration method based on a time series difference method according to an embodiment of the present application is provided.

[0030] Figure 4 A structure diagram of a wind-solar-thermal-storage system optimization configuration system based on a time series difference method according to an embodiment of the present application is provided.

[0031] Figure 5 A structure diagram of an optimal strategy module in a wind-solar-thermal-storage system optimization configuration system based on a time series difference method according to an embodiment of the present application is provided. DETAILED DESCRIPTION

[0032] Embodiments of the present application are described in detail below with reference to examples thereof illustrated in the attached drawings, in which the same or similar components are denoted by the same or similar reference numerals throughout. The embodiments described below by reference to the drawings are exemplary and are intended to explain the present application, and cannot be understood as limiting the present application.

[0033] In order for those skilled in the art to better understand the present application, the actual situation of new energy power generation is described in detail. The wind-solar complementary power generation system is greatly affected by climate and environment. Large-scale energy storage technology can ensure the continuity and reliability of load power consumption and reduce the waste of energy resources. Therefore, when designing a wind-solar-thermal-storage system, the optimal combination of load demand, wind power generation, photovoltaic power generation, thermal power generation and energy storage system capacity should be achieved to solve the problems of power supply reliability and power quality of the wind-solar complementary power generation system under relatively economic conditions.

[0034] The research paper on the optimization and scheduling method of the "wind, light, fire, storage, and storage" multi-energy complementary system considers the generation cost of conventional units under low load operation and climbing conditions on the basis of traditional coal costs and start-stop costs, and constructs a battery life loss cost model. The penalty cost calculation model of abandoned wind and light and the penalty cost calculation model of load shedding are introduced, thereby establishing the "wind, light, fire, storage, and storage" multi-energy complementary optimization scheduling model, and proposing to use the dynamic inertia weight particle swarm algorithm to solve the optimization objective of minimizing the total operation cost of the system.

[0035] However, most of the existing research results assume that the thermal power unit is in a continuous operation state. When the thermal power unit is not needed, the thermal power unit is operated under minimum load conditions, which may cause the thermal power unit to be in a long-term low-load operation state, and when the total excess power generation capacity exceeds the capacity of the energy storage device, the power is abandoned, which increases the operation cost and reduces the utilization rate of wind and light resources.

[0036] In order to solve the technical problems of high operation cost and low utilization rate of wind and light resources, the application aims to provide a wind-light-fire-storage system optimization configuration method, system, equipment and storage medium based on time sequence difference method, that is, the application optimizes the configuration of the wind-light-fire-storage system based on the time sequence difference method and adjusts the preset constraint condition, improves the utilization rate of wind and light resources, and saves the operation cost of the wind-light-fire-storage system.

[0037] The wind-light-fire-storage system optimization configuration method, system, equipment and storage medium based on time sequence difference method of the embodiments of the application are described below with reference to the accompanying drawings.

[0038] Embodiment 1

[0039] The application provides a wind-light-fire-storage system optimization configuration method based on time sequence difference method, Figure 1 The flowchart of the wind-light-fire-storage system optimization configuration method based on time sequence difference method provided by the embodiments of the present disclosure is shown in Figure 1 The method comprises the following steps:

[0040] Step 1: Obtain the power grid demand power generation, the power generation of wind turbine generators in the wind-light-fire-storage system, the power generation of photovoltaic generators, the preset constraint condition and the economic parameter at each time in the historical period;

[0041] It should be noted that the preset constraint condition comprises the capacity constraint, the state constraint of each power generation and energy storage device, and the initial state of the wind-light-fire-storage system.

[0042] Step 2: Determine the total demand power generation of the thermal power generator and the energy storage device in the wind-light-fire-storage system at each time in the historical period according to the power grid demand power generation, the power generation of wind turbine generators in the wind-light-fire-storage system and the power generation of photovoltaic generators in the historical period;

[0043] Step 3: Based on the operation state of the thermal power generator and the available power of the energy storage device, the state of the wind-light-fire-storage system is divided into different state types, and the probability value of mutual transfer between states and the strategy corresponding to each state type are randomly initialized;

[0044] In the embodiments of the present disclosure, the wind-light-fire-storage system is divided into different state types based on the operation state of the thermal power generator and the available power of the energy storage device, comprising:

[0045] The state of the wind-light-fire-storage system in which the thermal power generator is running and the available power of the energy storage device is greater than zero is divided into a first state;

[0046] The state of the wind-light-fire-storage system in which the thermal power generator is running and the available power of the energy storage device is equal to zero is divided into a second state;

[0047] The state in which the thermal power generating unit is shut down and the available power of the energy storage device is greater than zero in the wind-solar-thermal storage system is divided into a third state;

[0048] The state in which the thermal power generating unit is shut down and the available power of the energy storage device is equal to zero in the wind-solar-thermal storage system is divided into a fourth state.

[0049] Step 4: An optimal strategy is obtained by establishing a reinforcement learning model based on the SARS algorithm, taking the total demand power of the thermal power generating unit and the energy storage device in the wind-solar-thermal storage system at each time in the historical period as a sampling sequence, and training the model;

[0050] In the embodiment of the present disclosure, the reinforcement learning model is established based on the SARS algorithm, the total power demand at each time in the historical period is taken as a sampling sequence, and the optimal strategy under each state is obtained by training the model, including:

[0051] The initial state of the wind-solar-thermal storage system in the reinforcement learning model is initialized according to a preset constraint condition;

[0052] The initial state and the first sampling value in the sampling sequence are substituted into the pre-initialized action selection model to obtain an initial strategy corresponding to the initial state;

[0053] Based on the initial strategy, the action corresponding to the initial state and the next state corresponding to the action are determined;

[0054] The reward value of the state-action pair under the initial strategy is calculated based on the sampling value and the action corresponding to the initial state;

[0055] The next action corresponding to the state is determined based on the initial strategy of the next state;

[0056] The cumulative reward function of the state-action pair of the initial state, the reward value of the state-action pair under the initial strategy, and the cumulative reward function of the state-action pair of the next state are used to update the cumulative reward function of the state-action pair of the initial state and the strategy;

[0057] The next state and the next value of the sampling sequence are substituted into the reinforcement learning model, and all the above steps are repeated until all the values in the sampling sequence are traversed, and the training of the model is completed.

[0058] The strategy corresponding to each state in the trained model is the optimal strategy.

[0059] It should be noted that the action selection model is used to determine the action corresponding to the running state of the thermal power generating unit in the next time period from the running state of the thermal power generating unit in the current time period based on the state of the wind-solar-thermal storage system in the current time period and the total demand power in the next time period;

[0060] The total demand power generation includes: demand power generation being negative, demand power generation being positive and less than the current capacity of the energy storage device, demand power generation being greater than the current capacity of the energy storage device and less than the sum of the current capacity of the energy storage device and the maximum load of the thermal power unit, and demand power generation being greater than the sum of the current capacity of the energy storage device and the maximum load of the thermal power unit.

[0061] The operation state of the thermal power unit includes: shutdown and operation.

[0062] It should be noted that the reward value of the state-action pair is inversely proportional to the operation cost of the wind-solar-thermal-storage system;

[0063] The operation cost of the wind-solar-thermal-storage system mainly includes: coal cost of the thermal power unit, start-stop cost of the thermal power unit, maintenance cost of each device in the system, penalty cost of abandoned electricity, penalty cost of power shortage, and penalty cost of not meeting the normal use requirements of the device, etc.

[0064] It should be noted that the strategy is determined by the state transition probability;

[0065] The state transition probability is determined by the cumulative reward function of the state-action pair. If the i-th state has f selectable actions, there are f state-action pairs, and the cumulative reward function of the state-action pair can be obtained at initialization or calculated according to the sampling value;

[0066] The i-th state, the maximum one of the cumulative reward functions of the state-action pairs corresponding to the first action to the f-th action is taken as the optimal action corresponding to the i-th state in the state set, and the optimal action is the strategy in the state;

[0067] Wherein, f∈(1~δ), δ is the number of actions contained in the action set, i∈(1~N), N is the number of states contained in the state set.

[0068] For example, the calculation formula of the Q value Q t+1 in the t+1 iteration process in the cumulative reward function is as follows:

[0069] Q t+1 (s,a)=Q t (s,a)+α(r+γQ t (s',a')-Q t (s,a))

[0070] Q tis the Q value calculated during the t-th iteration, r is the reward value of the state-action pair selected in this calculation process, s is the current state, a is the current action, s' is the state after executing action a, a' is the action corresponding to the strategy of the s' state, α is the first preset parameter, γ is the second preset parameter, t∈(1~T), T is the iteration number threshold, and the sum of all iterations of all state-action cumulative reward functions is the number of samples in the sampling sequence.

[0071] Step 5: Calculate the state of the wind, solar, thermal and storage system at each moment in a given period and the operating cost of the wind, solar, thermal and storage system corresponding to the state according to the optimal strategy, thereby calculating the cumulative investment and operating cost of the wind, solar, thermal and storage system in the given period;

[0072] Step 6: Modify the preset constraints, recalculate the optimal strategy for each state and the cumulative investment and operating costs of the system under the preset constraints within a given period, select the minimum value from the cumulative investment and operating costs of the system under different constraints, and use the preset constraints corresponding to the minimum value to optimize the configuration of the wind, solar, thermal and energy storage system.

[0073] The specific method of this application is illustrated with examples based on the above configuration method:

[0074] In this embodiment, the startup state sequence of the thermal power unit is related to the equipment status and operating cost, and can be analyzed from the perspective of the unit's operating state transition. Every hour, the thermal power unit has two possible states: running and shut down. The energy storage device has two possible states: available power is 0 and available power is greater than 0. Therefore, the entire system has a total of four states, denoted as S0, S1, S2, and S3, and the corresponding state descriptions are:

[0075] S0: The thermal power unit is running and the available power of the energy storage device is greater than 0;

[0076] S1: The thermal power unit is running and the available power of the energy storage device is 0;

[0077] S2: The thermal power unit is shut down and the available power of the energy storage device is greater than 0;

[0078] S3: The thermal power unit is shut down and the available power of the energy storage device is 0;

[0079] The state transition diagram when the current state is S0 is as follows Figure 2The state is transferred to the next state according to the power demand at the next moment and the action of the thermal power unit, and a reward value r of the state transition is obtained, which is inversely proportional to the operation cost of the state transition. The action of the thermal power unit includes operation (A0) and shutdown (A1), and the power demand has four cases, namely, the demand is negative (Case 0), the demand is positive and less than the current capacity of the energy storage device (Case 1), the demand is greater than the current capacity of the energy storage device and less than the sum of the current capacity of the energy storage device and the maximum load of the thermal power unit (Case 2), and the demand is greater than the sum of the current capacity of the energy storage device and the maximum load of the thermal power unit (Case 3).

[0080] Since each state selects a certain action with a certain probability, each state-action pair is transferred to a certain state with a certain probability P, as shown by the arrows in Figure 2 , the current state is S0, and the power demand is Case 0, when the action A0 is performed, it will be transferred to the state S0 with a probability of P 000 , and to the state S1 with a probability of P 001 , therefore, when a certain state transition strategy maximizes the cumulative reward function, it is the optimal strategy, and the sequence of thermal power unit startup states obtained under the strategy minimizes the operation cost of the wind-solar-thermal-storage system. Since the reward value of state transition is different under different input parameters, different device operating states, different cost calculation methods and different device constraint conditions, the above two probabilities are unknown, at this time, a model-free reinforcement learning method can be used, such as the temporal difference learning method.

[0081] The specific flow chart of the wind-solar-thermal-storage system optimization configuration method based on the SARSA algorithm of the model-free temporal difference learning is shown in Figure 3 , and the specific steps are as follows:

[0082] F1: read in the power demand, the preset constraint conditions of each device, and the related economic parameters;

[0083] F2: initialize the current state s of the system, the current sampling step i, the cumulative reward function Q(s, a) of all state-action pairs, and the policy function Π(s) of all states;

[0084] F3: if the current sampling step i is less than or equal to the sampling sequence length, execute the single-step policy to step F4, otherwise go to step F9;

[0085] F4: determine the current action a according to the policy Π(s), and calculate the reward value r of this sampling and the operating state of each power generation and energy storage device in the integrated system, r is related to the operation cost, the smaller the cost, the greater the reward value;

[0086] F5: the next state s' is obtained according to the current state s and the current action a, and the next action a' is determined according to the policy Π(s');

[0087] F6: the formula Q t+1 (s,a)=Q t (s,a)+α(r+γQ t (s',a')-Q t (s,a)) is used to dynamically update the t+1th estimation value of the cumulative reward function Q of the state-action pair, wherein a is an update step, and γ is a reward discount;

[0088] F7: the policy Π(s) is updated to the action a'' that maximizes the Q value in the state s according to the updated Q(s,a);

[0089] F8: the step number i is increased by 1, and s' and a' are brought into step F3, and steps F3-F8 are repeatedly executed;

[0090] F9: after the full sampling is performed, the optimal policy Π' under the preset constraint condition can be obtained, and the state of the wind-solar-thermal storage system at each time in a given time period and the cumulative operation cost, power supply reliability index and the like of the wind-solar-thermal storage system corresponding to the state under the policy are saved;

[0091] F10: if it is necessary to adjust the preset constraint parameter to recalculate, returning to step F1, otherwise comparing the investment operation cost, power supply reliability index and the like obtained under different preset constraint parameters, and selecting the best configuration scheme of the wind-solar-thermal storage system.

[0092] In summary, the wind-solar-thermal storage system optimization configuration method based on the time series difference method provided by the present application firstly obtains the power grid demand power generation, the power generation of the wind turbine, the power generation of the photovoltaic turbine, the preset constraint condition and the economic parameter at each time in a historical period, secondly determines the total demand power generation of the thermal power unit and the energy storage device in the system based on the above-mentioned data, then divides the state type of the system and trains the reinforcement learning model established based on the wind-solar-thermal integrated system based on the SARSA algorithm, obtains the optimal strategy of the system under different states, then calculates the cumulative operation cost of the system in a given time period based on the optimal strategy, and finally modifies the preset constraint condition, selects the preset constraint condition corresponding to the minimum cumulative investment operation cost of the system under different constraints, and optimizes the configuration of the system. The technical scheme provided by the present application improves the utilization rate of wind and light resources and saves the operation cost.

[0093] Embodiment 2

[0094] Figure 4 The structure diagram of the wind-solar-thermal storage system optimization configuration system provided by the embodiment of the present application is as shown in Figure 4As shown, the system includes:

[0095] An acquisition module is used to obtain the power generation demand of the power grid at each moment in the historical period, the power generation of wind turbines in the wind-solar-thermal-storage system, the power generation of photovoltaic units, preset constraints and economic parameters;

[0096] A determination module is used to determine the total required power generation of thermal power units and energy storage equipment in the wind-solar-thermal-storage system at each moment in the historical period based on the power generation required by the power grid at each moment in the historical period, the power generation of wind turbines in the wind-solar-thermal-storage system, and the power generation of photovoltaic units;

[0097] The initialization module is used to divide the state of the wind, solar, thermal and energy storage system into different state types based on the operating status of the thermal power units and the available power of the energy storage equipment, and randomly initialize the probability values ​​of mutual transitions between states and the strategies corresponding to each state type;

[0098] The optimal strategy module is used to establish a reinforcement learning model based on the SARSA algorithm. The total power generation demand of thermal power units and energy storage equipment in the wind, solar, thermal and energy storage system at each moment in the historical period is used as a sampling sequence to train the model and obtain the optimal strategy.

[0099] A calculation module is used to calculate the state of the wind, solar, thermal and storage system at each moment in a given period and the operating cost of the system corresponding to the state according to the optimal strategy, so as to calculate the cumulative investment and operating cost of the wind, solar, thermal and storage system in the given period;

[0100] An optimization configuration module is used to modify preset constraints, recalculate the optimal strategy for each state and the cumulative investment and operating costs of the system under the preset constraints within a given period, select the minimum value from the cumulative investment and operating costs of the system under different constraints, and optimize the configuration of the wind, solar, thermal and energy storage system using the preset constraints corresponding to the minimum value;

[0101] Among them, the preset constraints include: capacity constraints, state constraints and initial state of each power generation and energy storage equipment and the wind, solar, thermal and storage system.

[0102] In the embodiment of the present disclosure, the wind-solar-thermal-storage system is divided into different status types based on the operating status of the thermal power unit and the available power of the energy storage device, including:

[0103] The state in which the thermal power generating unit in the wind-solar-thermal storage system is running and the available power of the energy storage device is greater than zero is classified as the first state;

[0104] The state in which the thermal power generating units in the wind-solar-thermal storage system are running and the available power of the energy storage device is zero is classified as the second state;

[0105] The state in which the thermal power generating unit is shut down and the available power of the energy storage device is greater than zero in the wind-solar-thermal storage system is divided into a third state;

[0106] The state in which the thermal power generating unit is shut down and the available power of the energy storage device is equal to zero in the wind-solar-thermal storage system is divided into a fourth state.

[0107] In the embodiments of the present disclosure, the optimal strategy module, as shown in the figure, comprises: Figure 5

[0108] An initialization unit is configured to initialize an initial state of the wind-solar-thermal storage system in the reinforcement learning model according to a preset constraint condition;

[0109] An initial strategy unit is configured to substitute the initial state and a first sampling value in the sampling sequence into the pre-initialized action selection model to obtain an initial strategy corresponding to the initial state;

[0110] A first determination unit is configured to determine an action corresponding to the initial state and a next state corresponding to the action based on the initial strategy;

[0111] A calculation unit is configured to calculate a reward value of a state-action pair under the initial strategy based on the sampling value and the action corresponding to the initial state;

[0112] A second determination unit is configured to determine a next action corresponding to the state based on the initial strategy of the next state;

[0113] An updating unit is configured to update a cumulative reward function of the state-action pair of the initial state, the reward value of the state-action pair under the initial strategy, and the cumulative reward function of the state-action pair of the next state based on the cumulative reward function of the state-action pair of the initial state.

[0114] A loop unit is configured to substitute the next state and a next value in the sampling sequence into the reinforcement learning model to repeat all the above steps until all the values in the sampling sequence are traversed, and the training of the model is completed.

[0115] An optimal strategy unit is configured to take the strategy corresponding to each state in the trained model as the optimal strategy.

[0116] It should be noted that the action selection model is configured to determine the action corresponding to the running state of the thermal power generating unit in the next time period based on the state of the wind-solar-thermal storage system in the current time period and the total demand power generation in the next time period.

[0117] ​The total demand power generation includes: demand power generation being negative, demand power generation being positive and less than the current capacity of the energy storage device, demand power generation being greater than the current capacity of the energy storage device and less than the sum of the current capacity of the energy storage device and the maximum load of the thermal power unit, and demand power generation being greater than the sum of the current capacity of the energy storage device and the maximum load of the thermal power unit.

[0118] The operation state of the thermal power unit includes: shutdown and operation.

[0119] It should be noted that the reward value of the state-action pair is inversely proportional to the operation cost of the wind-solar-thermal-storage system;

[0120] The operation cost of the wind-solar-thermal-storage system mainly includes the coal cost of the thermal power unit, the start-stop cost of the thermal power unit, the maintenance cost of each device in the system, the penalty cost of abandoned electricity, the penalty cost of power shortage, and the penalty cost of not meeting the normal use requirements of the device, etc.

[0121] It should be noted that the strategy is determined by the state transition probability;

[0122] The state transition probability is determined by the cumulative reward function of the state-action pair. If the i-th state has f selectable actions, there are f state-action pairs, and the cumulative reward function of the state-action pair can be obtained at initialization or calculated according to the sampling value;

[0123] The i-th state, the maximum one of the cumulative reward functions of the state-action pairs corresponding to the first action to the f-th action is taken as the optimal action corresponding to the i-th state in the state set, and the optimal action is the strategy in the state;

[0124] Wherein, f∈(1~δ), δ is the number of actions contained in the action set, i∈(1~N), N is the number of states contained in the state set.

[0125] For example, the calculation formula of the Q value Q t+1 in the t+1 iteration process in the cumulative reward function is as follows:

[0126] Q t+1 (s,a)=Q t (s,a)+α(r+γQ t (s',a')-Q t (s,a))

[0127] In the formula, Q tQ is a value calculated in the tth iteration process, r is a reward value of a state-action pair selected in the current calculation process, s is a current state, a is a current action, s' is a state after the action a is executed, a' is an action corresponding to the state s', a is a first preset parameter, g is a second preset parameter, t is in (1-T), T is a threshold of iteration number, and the sum of all iteration numbers of all state-action cumulative reward functions is a sample number of a sampling sequence.

[0128] To sum up, the wind, light, fire and storage system optimization configuration system based on the time sequence difference method is proposed in the application, the system comprises an acquisition module, a determination module, an initialization module, an optimal strategy module, a calculation module and an optimization configuration module. The application optimizes the configuration of the wind, light, fire and storage system based on the time sequence difference method and the adjusted preset constraint condition, improves the utilization rate of wind and light resources, and saves the operation cost of the wind, light, fire and storage system.

[0129] Embodiment 3

[0130] In order to realize the above-mentioned embodiments, the embodiments of the application further propose a computer device, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to realize the method described in Embodiment 1 of the application.

[0131] Embodiment 4

[0132] In order to realize the above-mentioned embodiments, the embodiments of the application further propose a non-transitory computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to realize the method described in Embodiment 1 of the application.

[0133] It should be noted that, in the description of the application, the terms "first", "second", etc. are only for the purpose of description, and cannot be understood as indicating or implying relative importance. In addition, in the description of the application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0134] In the description of the specification, the description referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the application. In the description of the specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples without contradiction.

[0135] Any processes or methods described in the flow charts or otherwise described herein can be understood as representing modules, segments, or portions of code that include one or more executable instructions for implementing the specified logical functions or steps, and the preferred embodiments of the application include additional or fewer steps, in other orders, with other functionality, in implementations of these preferred embodiments of the application. Thus, any of the steps, options, aspects, components, etc. discussed herein can be included or deleted in other embodiments of the application, and yet still be deemed to fall within the scope of the present application.

[0136] Although the embodiments of the present application have been shown and described above, it should be understood by those ordinary skilled in the art that the above embodiments are exemplary and cannot be construed as limiting the present application, and those ordinary skilled in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.

Claims

1. A method for optimal configuration of a wind-solar-thermal-storage system based on time-series difference method, characterized in that, The method comprises: acquiring power grid demand power generation, wind turbine generator power generation in a wind-solar-thermal-storage system, photovoltaic generator power generation, a preset constraint condition and an economic parameter at each time in a historical period; determining total demand power generation of the thermal turbine generator and the energy storage device in the wind-solar-thermal-storage system at each time in the historical period according to the power grid demand power generation, the wind turbine generator power generation in the wind-solar-thermal-storage system and the photovoltaic generator power generation; dividing the state of the wind-solar-thermal-storage system into different state types based on the operating state of the thermal turbine generator and the available power of the energy storage device, and randomly initializing the probability value of mutual transition between states and the strategy corresponding to each state type; establishing a reinforcement learning model based on the SARSA algorithm, taking the total demand power generation of the thermal turbine generator and the energy storage device in the wind-solar-thermal-storage system at each time in the historical period as a sampling sequence, training the model to obtain an optimal strategy; wherein the specific manner for determining the optimal strategy comprises: initializing the initial state of the wind-solar-thermal-storage system in the reinforcement learning model according to the preset constraint condition; substituting the initial state and the first sampling value in the sampling sequence into the pre-initialized action selection model to obtain an initial strategy corresponding to the initial state; determining an action corresponding to the initial state and a next state corresponding to the action based on the initial strategy; calculating a reward value of a state-action pair under the initial strategy based on the sampling value and the action corresponding to the initial state; determining a next action corresponding to the state based on the initial strategy of the next state; updating a cumulative reward function of the state-action pair of the initial state, the reward value of the state-action pair under the initial strategy and a cumulative reward function of the state-action pair of the next state based on the cumulative reward function of the state-action pair of the initial state; substituting the next state and a next sampling value in the sampling sequence into the reinforcement learning model to obtain an initial strategy corresponding to the next state, repeating the determination of the initial strategy corresponding to the next state until all values in the sampling sequence are traversed, completing the training of the model, and the strategy corresponding to each state in the trained model is the optimal strategy; calculating the state of the wind-solar-thermal-storage system at each time in a given period and the operating cost of the wind-solar-thermal-storage system corresponding to the state according to the optimal strategy, thereby calculating the cumulative investment and operating cost of the wind-solar-thermal-storage system in the given period; modifying the preset constraint condition, recalculating the optimal strategy of each state and the cumulative investment and operating cost of the system in the given period under the preset constraint condition, screening the minimum value from the cumulative investment and operating cost of the system under different constraints, and optimizing the configuration of the wind-solar-thermal-storage system by using the preset constraint condition corresponding to the minimum value; wherein the preset constraint condition comprises: capacity constraint, state constraint and initial state of the wind-solar-thermal-storage system of each power generation and energy storage device.

2. The method of claim 1, wherein, The dividing of the state of the wind-solar-thermal-storage system into different state types based on the operating state of the thermal turbine generator and the available power of the energy storage device comprises: dividing the state of the wind-solar-thermal-storage system in which the thermal turbine generator is running and the available power of the energy storage device is greater than zero into a first state; a state in which the thermal power generating unit in the wind-solar-thermal-storage system is running and the available power of the energy storage device is equal to zero is divided into a second state; a state in which the thermal power generating unit in the wind-solar-thermal-storage system is shut down and the available power of the energy storage device is greater than zero is divided into a third state; a state in which the thermal power generating unit in the wind-solar-thermal-storage system is shut down and the available power of the energy storage device is equal to zero is divided into a fourth state.

3. The method of claim 1, wherein, The action selection model is configured to determine, based on the state of the wind-solar-thermal-storage system at the current time and the operating state of the thermal power generating unit at the current time determined based on the total demand power generation at the next time, an action corresponding to a transition of the operating state of the thermal power generating unit at the current time to the operating state of the thermal power generating unit at the next time. The total demand power generation includes: negative demand power generation, positive demand power generation less than the current capacity of the energy storage device, demand power generation greater than the current capacity of the energy storage device and less than the sum of the current capacity of the energy storage device and the maximum load of the thermal power generating unit, and demand power generation greater than the sum of the current capacity of the energy storage device and the maximum load of the thermal power generating unit. The operating state of the thermal power generating unit includes: shut down and running.

4. The method of claim 1, wherein, The reward value of the state-action pair is inversely proportional to the operation cost of the wind-solar-thermal-storage system. The operation cost of the wind-solar-thermal-storage system includes the coal cost of the thermal power generating unit, the start-stop cost of the thermal power generating unit, the maintenance cost of each device in the system, the penalty cost of abandoned power, the penalty cost of power shortage, and the penalty cost of not meeting the normal use requirements of the device.

5. The method of claim 1, wherein, The strategy is determined by the state transition probability. The state transition probability is determined by the cumulative reward function of the state-action pair. If there are f available actions for the i-th state, there are f state-action pairs. The cumulative reward function of the state-action pair can be obtained at initialization or calculated according to the sampling value. In the i-th state, the maximum one in the cumulative reward functions of the state-action pairs corresponding to the first action to the f-th action is taken as the optimal action corresponding to the i-th state in the state set. The optimal action is the strategy in the state. wherein, , is the number of actions included in the action set, N is the number of states included in the state set.

6. The method of claim 1, wherein, The cumulative reward function is given by The formula for the calculation of the Q value in the next iteration is given by Where, For the The Q value calculated during the iteration, r is the reward value of the state-action pair selected in this calculation process, s is the current state, a is the current action, s' is the state after executing action a, and a' is the action corresponding to the strategy of the s' state. is the first preset parameter, is the second preset parameter, , is the iteration threshold, and the sum of all iterations of all state-action cumulative reward functions is the number of samples in the sampling sequence.

7. A system for optimal configuration of a wind-solar-thermal-storage system based on time series differentiation method, characterized in that, The system comprises: An acquisition module configured to acquire the power grid demand power generation, the power generation of the wind power generating unit, the power generation of the photovoltaic generating unit, a preset constraint condition, and an economic parameter in the wind-solar-thermal-storage system at each time in a historical period; A determination module configured to determine the total demand power generation of the thermal power generating unit and the energy storage device in the wind-solar-thermal-storage system at each time in the historical period based on the power grid demand power generation, the power generation of the wind power generating unit, and the power generation of the photovoltaic generating unit in the wind-solar-thermal-storage system at each time in the historical period; An initialization module configured to divide the states of the wind-solar-thermal-storage system into different state types based on the operating state of the thermal power generating unit and the available power of the energy storage device, and randomly initialize the probability values of mutual transitions between the states and the strategies corresponding to the state types. The optimal strategy module is configured to establish a reinforcement learning model based on a SARS algorithm, take total demand power of a thermal power unit and an energy storage device in a wind-solar-thermal-storage system at each time point in a historical period as a sampling sequence, train the model, and obtain an optimal strategy. The optimal strategy is determined in the following manner: an initial state of the wind-solar-thermal-storage system in the reinforcement learning model is initialized according to a preset constraint condition; the initial state and a first sampling value in the sampling sequence are substituted into a pre-initialized action selection model to obtain an initial strategy corresponding to the initial state; an action corresponding to the initial state and a next state corresponding to the action are determined based on the initial strategy; a reward value of a state-action pair in the initial strategy is calculated based on the sampling value and the action corresponding to the initial state; a next action corresponding to the state is determined based on the initial strategy of the next state; a cumulative reward function of the state-action pair of the initial state, the reward value of the state-action pair in the initial strategy, and a cumulative reward function of the state-action pair of the next state are used to update the cumulative reward function of the state-action pair of the initial state and the strategy; the next state and a next sampling value in the sampling sequence are substituted into the reinforcement learning model to obtain an initial strategy corresponding to the next state, and the initial strategy corresponding to the next state is repeatedly determined until all values in the sampling sequence are traversed, the training of the model is completed, and the strategy corresponding to each state in the trained model is the optimal strategy. The calculation module is configured to calculate a state of the wind-solar-thermal-storage system at each time point in a given period and an operation cost of the system corresponding to the state according to the optimal strategy, and thus calculate a cumulative investment and operation cost of the wind-solar-thermal-storage system in the given period. The optimal configuration module is configured to modify the preset constraint condition, recalculate the optimal strategy of each state and the cumulative investment and operation cost of the system in the given period under the preset constraint condition, select a minimum value from the cumulative investment and operation costs of the system under different constraints, and perform optimal configuration on the wind-solar-thermal-storage system by using the preset constraint condition corresponding to the minimum value. The preset constraint condition includes a capacity constraint of each power generation and energy storage device, a state constraint, and an initial state of the wind-solar-thermal-storage system.

8. A computer device, comprising: The computer program is executed by the processor to implement the method of any one of claims 1-6.

9. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Wind-photovoltaic-gas-storage combined dynamic economic dispatching optimization method based on Q learning

    CN111064229A

  • Dynamic power system economic dispatching method based on deep reinforcement learning

    CN112186743A