A frequency modulation-peak regulation collaborative optimization method and system of a molten salt heat storage system

By using hierarchical learning and reinforcement learning algorithms, the optimal slow and fast strategies for molten salt thermal energy storage systems are determined, solving the problems of computational complexity and training time consumption in power grid frequency regulation and peak shaving. This achieves efficient frequency regulation-peak shaving collaborative optimization, improving system stability and economy.

CN120454110BActive Publication Date: 2026-02-06XIAN THERMAL POWER RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510709730.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2026-02-06
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

In existing technologies, molten salt thermal energy storage systems face problems of computational complexity and training time consumption in power grid frequency regulation and peak shaving. In particular, a single time scale strategy cannot simultaneously handle second-level frequency regulation and hour-level peak shaving, resulting in an excessively large action space dimension, making it difficult to converge to the optimal solution.

Method used

A hierarchical learning approach is adopted, and the optimal slow strategy for long-term peak shaving and the optimal fast strategy for short-term frequency regulation of the molten salt thermal storage system are determined by reinforcement learning algorithm. The optimal strategy is then combined with frequency stability and the benchmark output plan of the generator set for collaborative optimization.

Benefits of technology

It improves the accuracy of long-term stable strategies, reduces computational load, increases frequency regulation response speed, reduces frequency deviation and cost, and extends energy storage life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120454110B_ABST
    Figure CN120454110B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of power grid frequency modulation and peak regulation, and particularly relates to a frequency modulation-peak regulation collaborative optimization method and system of a molten salt heat storage system, comprising: based on relevant state characteristic data and control management action characteristic data, determining an optimal slow strategy of the molten salt heat storage system in a long-time scale peak regulation process through a reinforcement learning algorithm, and determining a reference output plan of a generator set through the optimal slow strategy in the peak regulation process; determining a reference frequency of a power grid at each time through the reference output plan; according to the reference frequency of the power grid at each time, real-time state characteristic data and actual operation action characteristic data, determining an optimal fast strategy of the molten salt heat storage system in a short-time scale frequency modulation process through the reinforcement learning algorithm; and performing collaborative optimization of frequency modulation-peak regulation of the molten salt heat storage system through the optimal slow strategy and the optimal fast strategy. The present application improves frequency modulation response speed, reduces frequency deviation, saves cost and prolongs energy storage life.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid frequency modulation and peak regulation, and in particular to a frequency modulation-peak regulation collaborative optimization method and system of a molten salt thermal energy storage system. BACKGROUND

[0002] The molten salt thermal energy storage system (TES) has been widely used in the field of new energy power generation (especially solar thermal power, wind power, etc.). It mainly stores heat energy to balance the power grid load, achieve frequency modulation, peak regulation and energy management. In order to more effectively utilize the molten salt thermal energy storage system and improve its role in power dispatching, a frequency modulation-peak regulation collaborative optimization method is needed.

[0003] In order to ensure the stability and reliability of the power grid, solve the problems of power grid load and power generation fluctuation, unstable renewable energy generation and frequency fluctuation, etc., the stability of the system is maintained by frequency modulation and peak regulation. In the prior art, the frequency modulation and peak regulation of the power grid are usually carried out by a reinforcement learning algorithm of a single time scale strategy. However, the single time scale strategy cannot handle both second-level frequency modulation and hour-level peak regulation, which may lead to increased computational complexity and time consumption in the training process and difficulty in converging to the optimal solution when the action space dimension in the reinforcement learning is too large. This problem can be alleviated by simplifying the action space and hierarchical learning. SUMMARY

[0004] The present application provides a frequency modulation-peak regulation collaborative optimization method and system of a molten salt thermal energy storage system, which is used to solve the problems of computational complexity and training time consumption in the existing power grid frequency modulation and peak regulation.

[0005] The purpose of the present application can be achieved by the following technical solutions:

[0006] The first aspect of the present application provides a frequency modulation-peak regulation collaborative optimization method of a molten salt thermal energy storage system, comprising:

[0007] Obtain real-time state feature data in the molten salt thermal energy storage system and related state feature data affecting the efficiency of the molten salt thermal energy storage system; then obtain actual operation action feature data and control management action feature data in the molten salt thermal energy storage system;

[0008] Based on the related state feature data and the control management action feature data, determine the best slow strategy of the molten salt thermal energy storage system in the long time scale peak regulation process through a reinforcement learning algorithm, and determine the baseline output plan of the generator set through the best slow strategy in the peak regulation process; determine the baseline frequency of the power grid at each time through the baseline output plan;

[0009] According to the reference frequency of the power grid at each moment, the real-time state characteristic data and the actual operation action characteristic data, the optimal fast strategy of the molten salt heat storage system in the short-time scale frequency modulation process is determined through the reinforcement learning algorithm.

[0010] The frequency modulation-peak regulation coordination optimization of the molten salt heat storage system is performed through the optimal slow strategy and the optimal fast strategy.

[0011] Further, the optimal slow strategy of the molten salt heat storage system in the long-time scale peak regulation process is determined through the related state characteristic data and the control management action characteristic data, and the reinforcement learning algorithm, and the optimal slow strategy comprises:

[0012] A first state space of a long-time scale is constructed through the related state characteristic data, a first action space of a long-time scale is constructed through the control management action characteristic data, and the optimal slow strategy of the molten salt heat storage system in the long-time scale peak regulation process is determined through the first state space of a long-time scale and the first action space of a long-time scale and the reinforcement learning algorithm based on the strategy.

[0013] Further, the optimal slow strategy of the molten salt heat storage system in the long-time scale peak regulation process is determined through the first state space of a long-time scale and the first action space of a long-time scale and the reinforcement learning algorithm based on the strategy, and the optimal slow strategy comprises:

[0014] A target function of the slow strategy is determined according to the first state space of a long-time scale and the first action space of a long-time scale, and the target function of the slow strategy is specifically represented as:

[0015]

[0016] In the formula, γ is a discount factor, γ is a discount factor, γ represents a discount factor in the i-th time step, γ represents a discount factor in the i-th time step, r i represents a reward value in the i-th time step, r i represents a reward value in the i-th time step, N represents the total number of steps in the interaction process between the environment and the agent, μ represents an expected operation, that is, an expected value of decision-making according to the strategy π, V represents a long-term cumulative reward of all time steps;

[0017] The optimal slow strategy is obtained through the long-term cumulative reward of all time steps.

[0018] Further, the optimal fast strategy of the molten salt heat storage system in the short-time scale frequency modulation process is determined according to the reference frequency of the power grid at each moment, the real-time state characteristic data and the actual operation action characteristic data, and the reinforcement learning algorithm, and the optimal fast strategy comprises:

[0019] The preset time interval is 1 minute. to continuously obtain the real-time state feature data of the molten salt thermal storage system at each time point within the preset time period before the current time point to continuously obtain the real-time state feature data of the molten salt thermal storage system at each time point within the preset time period before the current time point

[0020] to obtain the maximum output power of the generator or the energy storage device;

[0021] by the reference frequency of the power grid at each time point, a frequency stability constraint is constructed, and a penalty value when the frequency deviates from the target frequency is obtained, and the penalty value when the frequency deviates from the target frequency is recorded as the penalty value of the first constraint; by the maximum output power of the generator or the energy storage device, a generator and energy storage capacity constraint is constructed, and a penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is obtained, and the penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is recorded as the penalty value of the second constraint;

[0022] a second state space of a short time scale is constructed by the real-time state feature data, and a second action space of a short time scale is constructed by the actual operation action feature data;

[0023] a target function of the fast strategy is constructed by the penalty values of the constraints, the second state space of the short time scale and the second action space of the short time scale; the optimal fast strategy of the molten salt thermal storage system in the short time scale frequency modulation process is determined by the target function of the fast strategy.

[0024] Further, the frequency stability constraint is constructed by the reference frequency of the power grid at each time point, and the penalty value when the frequency deviates from the target frequency includes:

[0025]

[0026] In the formula, represents the frequency of the power grid at each time point, represents the reference frequency of the power grid at each time point, represents a preset frequency error, represents a frequency violation penalty coefficient, represents a maximum value function, represents an absolute value symbol, represents the penalty value when the frequency deviates from the target frequency;

[0027] The frequency violation penalty coefficient is determined by hyperparameter tuning.

[0028] Further, the generator and energy storage capacity constraint is constructed by the maximum output power of the generator or the energy storage device, and the penalty value when the maximum output power limit of the generator or the energy storage device is exceeded includes:

[0029]

[0030] wherein, represents the generator or energy storage device output power at each time point, represents the maximum output power of the generator or energy storage device; represents the maximum function, represents the output power violation penalty coefficient, represents the penalty value when exceeding the maximum output power limit of the generator or energy storage device;

[0031] wherein, the penalty value when exceeding the maximum output power limit of the generator or energy storage device is determined by hyperparameter tuning.

[0032] Further, the target function of the fast strategy is constructed through the penalty values of the constraints, the second state space and the second action space of the short time scale; the optimal fast strategy of the molten salt thermal storage system in the short time scale frequency regulation process is determined through the target function of the fast strategy, comprising:

[0033]

[0034] represents the discount factor in the th update step, represents the reward value of the th update step, represents the total number of all update steps in the interaction process between the environment and the agent; represents the penalty value of the th constraint, represents the penalty strength of the th constraint, represents the expected operation, i.e. the expected value of decision-making according to the strategy represents the long-term cumulative reward of all update steps, represents the number of all constraints, represents the adjustment of the initial update time interval, i.e. the length of one update step is equal to the adjustment of the initial update time interval; wherein, the penalty strength of each constraint is determined by hyperparameter tuning;

[0035]

[0036] wherein, represents the value of the th reference time point of the th real-time state feature at the current time point, represents the average value of the values of all reference time points of the th real-time state feature data at the current time point, represents the number of reference time points, ​represents the number of all real-time state features, represents the absolute value symbol, represents a linear normalization function, represents a preset initial update time interval;

[0037] wherein a preset number of time points before the current time point are taken as reference time points of the current time point;

[0038] The best fast strategy is obtained through long-term cumulative rewards of all update steps.

[0039] The second aspect of the present application provides a frequency modulation-peak regulation collaborative optimization system of a molten salt thermal storage system, comprising:

[0040] The data acquisition module is used to acquire real-time state feature data and related state feature data affecting the efficiency of the molten salt thermal storage system in the molten salt thermal storage system, and then acquire actual operation action feature data and control management action feature data in the molten salt thermal storage system;

[0041] The long-time scale analysis module is used to determine the best slow strategy of the molten salt thermal storage system in a long-time scale peak regulation process based on the related state feature data and the control management action feature data through a reinforcement learning algorithm, determine the reference output plan of the generator set through the best slow strategy in the peak regulation process, and determine the reference frequency of the power grid at each time point through the reference output plan.

[0042] The short-time scale analysis module is used to determine the best fast strategy of the molten salt thermal storage system in a short-time scale frequency modulation process through a reinforcement learning algorithm according to the reference frequency of the power grid at each time point, the real-time state feature data and the actual operation action feature data.

[0043] The collaborative optimization module is used to perform collaborative optimization of frequency modulation-peak regulation of the molten salt thermal storage system through the best slow strategy and the best fast strategy.

[0044] The third aspect of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the frequency modulation-peak regulation collaborative optimization method of the molten salt thermal storage system when executing the computer program.

[0045] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the frequency modulation-peak regulation collaborative optimization method of the molten salt thermal storage system.

[0046] ​Compared with the prior art, the beneficial effects of the present application are: based on the relevant state feature data and the control management action feature data, the optimal slow strategy of the molten salt heat storage system in the long time scale peak regulation process is determined through the reinforcement learning algorithm, and the reference output plan of the generator set is determined through the optimal slow strategy in the peak regulation process; the reference frequency of the power grid at each time is determined through the reference output plan, which improves the accuracy of the long-term stable strategy; according to the reference frequency of the power grid at each time, the real-time state feature data and the actual operation action feature data, the optimal fast strategy of the molten salt heat storage system in the short time scale frequency regulation process is determined through the reinforcement learning algorithm, the accuracy of the transient strategy of the molten salt heat storage system is improved through the constraint of the long-term stable strategy, and the calculation amount is reduced based on the long-term stable strategy; the frequency regulation-peak regulation of the molten salt heat storage system is optimized through the optimal slow strategy and the optimal fast strategy, the frequency regulation response speed is improved, the frequency deviation is reduced, and the cost is saved and the service life of the energy storage is prolonged. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0048] Figure 1 A step flow diagram of a frequency regulation-peak regulation collaborative optimization method of a molten salt heat storage system is provided for the present application.

[0049] Figure 2 A module flow diagram of a frequency regulation-peak regulation collaborative optimization system of a molten salt heat storage system is provided for the present application. DETAILED DESCRIPTION

[0050] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0051] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0052] In view of the problems in the background art, a frequency regulation-peak regulation collaborative optimization method and system for a molten salt thermal storage system are designed, which has important practical significance.

[0053] As shown in Figure 1 The first aspect of the present application is to provide a frequency regulation-peak regulation collaborative optimization method for a molten salt thermal storage system, comprising the following steps:

[0054] Step S001: Collect various state feature data and various action feature data in the molten salt thermal storage system.

[0055] It should be noted that, in order to solve the problem of excessive dimension and the problem of calculation complexity and training time consumption caused by the fact that the single time scale strategy in the conventional technology cannot simultaneously process second-level frequency regulation and hour-level peak regulation, the present application classifies the dimensions for hierarchical learning, so as to simplify the state space and action space in the learning process.

[0056] Specifically, real-time state feature data in the molten salt thermal storage system and related state feature data affecting the efficiency of the molten salt thermal storage system are obtained; actual operation action feature data and control management action feature data in the molten salt thermal storage system are obtained.

[0057] The real-time state feature data includes power grid frequency, load demand, generator output (power generation capacity and power of the generator), and state of charge (SOC) of the energy storage system; the related state feature data includes power grid frequency, load demand, generator output, state of charge of the energy storage system, initial investment cost, operation and maintenance cost, energy conversion efficiency and cost, and weather state; the actual operation action feature data includes generator output power and charge-discharge power of the energy storage system; and the control management action feature data includes start-stop state of the unit, reference charge-discharge power and SOC value of the unit.

[0058] Step S002: based on the relevant state feature data and the control management action feature data, the optimal slow strategy of the molten salt heat storage system in the long time scale peak regulation process is determined through the reinforcement learning algorithm, and the reference output plan of the generator set is determined through the optimal slow strategy in the peak regulation process; the reference frequency of the power grid at each time is determined through the reference output plan.

[0059] It should be noted that in order to perform hierarchical learning on all dimensions, two strategy-based reinforcement learning algorithms are used to analyze the fast strategy and the slow strategy respectively. First, the start-stop state and the reference output plan of all generator sets are determined through economic optimization for the slow strategy, and the fast strategy is constrained by the reference output plan, so as to complete the frequency regulation-peak regulation collaborative optimization of the molten salt heat storage system in the transient change process.

[0060] Specifically, based on the relevant state feature data and the control management action feature data, the optimal slow strategy of the molten salt heat storage system in the long time scale peak regulation process is determined through the reinforcement learning algorithm;

[0061] Wherein, the first state space of long time scale is constructed through the relevant state feature data, and the first action space of long time scale is constructed through the control management action feature data; according to the first state space and the first action space of long time scale, the optimal slow strategy of the molten salt heat storage system in the long time scale peak regulation process is determined through the CPO algorithm (strategy reinforcement learning algorithm); wherein, the CPO (Conservative Policy Optimization) algorithm is a known technology, which will not be described in detail here.

[0062] It should be noted that in order to make a small range of adjustments to the grid frequency, load demand, generator output, and state of energy storage system, a long-term stable slow strategy can be developed to determine some reference values for analysis.

[0063] Specifically, the objective function of the slow strategy is determined according to the first state space and the first action space of long time scale, which is specifically represented as:

[0064]

[0065] In the formula, is a discount factor, represents the discount factor in the th time step, represents the reward value of the th time step, represents the total number of steps in the interaction process between the environment and the agent, represents the expected value of the expected operation, i.e. making decisions according to the strategy π, represents the long-term cumulative reward of all time steps;

[0066] The best slow strategy is obtained through the long-term cumulative reward of all time steps.

[0067] wherein the strategy at each step during the interaction between the environment and the agent is also adjusted, and the strategy adjustment process is specifically represented as:

[0068]

[0069] In the formula, represents the new strategy, represents the old strategy, represents the KL divergence (Kullback-Leibler divergence) between the new and old strategies, represents a preset control threshold; wherein in the embodiment, the preset control threshold is 0.01. is not specifically limited, and the implementer can determine it according to the specific situation. Wherein, the discount factor and the obtaining process of the reward value of each time step are known technologies, and will not be specifically described here. Wherein, in the embodiment, the time interval corresponding to one time step is 1 hour, and in the embodiment, the time interval corresponding to one time step is not specifically limited, and the implementer can determine it according to the specific situation.

[0070] wherein, in , when the KL divergence between the new and old strategies is less than or equal to the preset control threshold , the old strategy is replaced by the new strategy; when the KL divergence between the new and old strategies is greater than the preset control threshold , the old strategy is not adjusted. Wherein, the obtaining process of the KL divergence between the new and old strategies is a known technology, and will not be specifically described here.

[0071] The best slow strategy is obtained through the long-term cumulative reward of all time steps; the reference output plan of the generator set is determined through the best slow strategy; and the reference frequency of the power grid at each time is determined through the reference output plan.

[0072] Step S003: According to the reference frequency of the power grid at each time, the real-time state feature data and the actual operation action feature data, the best fast strategy of the molten salt heat storage system in the short-time scale frequency modulation process is determined through the reinforcement learning algorithm.

[0073] It should be noted that in the process of policy-based reinforcement learning (Reinforcement Learning, RL), in order to ensure that the agent can make decisions within a suitable range and achieve effective and stable learning effect, constraints are needed; the introduction of constraints helps to solve many problems, especially in practical applications, to ensure that the model not only learns the optimal strategy, but also ensures that the strategy meets certain safety, executability, efficiency and other requirements.

[0074] Specifically, according to the reference frequency of the power grid at each time, the real-time state feature data and the actual operation action feature data, the best fast strategy of the molten salt heat storage system in the short-time scale frequency modulation process is determined through the reinforcement learning algorithm; in this embodiment, the determination process of the fast strategy is determined through the DPG algorithm; the DPG (Deep Deterministic Policy Gradient) algorithm is a known technology, which will not be described in detail here.

[0075] Wherein, the preset time interval is used to continuously obtain the real-time state feature data of all time points in the molten salt heat storage system within a preset time length before the current time; in this embodiment, the preset time interval is 1 second, and the preset time length is 1 hour, wherein in this embodiment, the preset time interval and the preset time length are not specifically limited and can be determined according to specific conditions.

[0076] The maximum output power of the generator or energy storage device is obtained;

[0077] The frequency stability constraint is constructed by the reference frequency of the power grid at each time, and the penalty value when the frequency deviates from the target frequency is obtained; the frequency stability constraint is specifically represented as:

[0078]

[0079] In the formula, f represents the frequency of the power grid at each time, f ref represents the reference frequency of the power grid at each time, e f represents the preset frequency error, p f represents the frequency violation penalty coefficient, max represents the maximum value function, abs represents the absolute value symbol, p represents the penalty value when the frequency deviates from the target frequency;

[0080] wherein the frequency violation penalty coefficient is determined by hyperparameter tuning. The penalty value when the frequency deviates from the target frequency is denoted as the penalty value of the first constraint.

[0081] The generator and energy storage capacity constraints are constructed by the maximum output power of the generator or the energy storage device, and a penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is obtained; the generator and energy storage capacity constraints are specifically represented as:

[0082]

[0083] wherein, represents the output power of the generator or the energy storage device at each time point, represents the maximum output power of the generator or the energy storage device; represents the maximum function, represents the output power violation penalty coefficient, represents the penalty value when the maximum output power limit of the generator or the energy storage device is exceeded;

[0084] wherein the penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is determined by hyperparameter tuning. The penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is denoted as the penalty value of the second constraint.

[0085] The second state space of the short time scale is constructed by the real-time state feature data, and the second action space of the short time scale is constructed by the actual operation action feature data;

[0086] The objective function of the fast strategy is constructed by the penalty values of the constraints, the second state space and the second action space of the short time scale; the objective function of the fast strategy is specifically represented as:

[0087]

[0088] represents the discount factor in the th update step, represents the reward value in the th update step, represents the total number of all update steps in the interaction process between the environment and the agent; represents the penalty value of the th constraint, represents the penalty strength on the th constraint, represents the expected operation, i.e. the expected value of decision-making according to the strategy , and represents the long-term cumulative reward of all update steps, represents the number of all constraints, represents the adjustment initial update time interval, i.e. the length of one update step is equal to the adjustment initial update time interval; wherein the penalty strength of each constraint is determined by hyperparameter tuning.

[0089] It should be noted that, in the process of adjusting the generator output power and the charge-discharge power of the energy storage system, the smaller the fluctuation of all real-time state characteristic data of the second state space of the fast strategy, the shorter the interval in the power adjustment update process of the fast strategy, otherwise only the calculation amount will be increased; and the larger the fluctuation of all real-time state characteristic data of the second state space of the fast strategy, the longer the interval in the power adjustment update process of the fast strategy, otherwise the system will be unstable due to the delay of adjustment, therefore the update time in the process of adjusting the electric power of the fast strategy is adjusted, so as to more perfectly meet the adjustment update time interval.

[0090] Specifically, the preset initial update time interval is adjusted according to the fluctuation of all real-time state characteristic data, and an adjusted initial update time interval is obtained; the adjustment initial update time interval is specifically represented by a formula as follows:

[0091]

[0092] In the formula, represents the value of the i-th real-time state characteristic at the current time at the j-th reference time, represents the mean value of the values of the i-th real-time state characteristic data at all reference times at the current time, represents the number of reference times, represents the number of all real-time state characteristics, represents the absolute value symbol, represents a linear normalization function, represents the adjustment initial update time interval, represents the preset initial update time interval; wherein the preset initial update time interval is 0.1 seconds in the embodiment, and the preset initial update time interval is not specifically limited in the embodiment and can be determined according to specific conditions.

[0093] wherein the preset value times before the current time are taken as the reference times of the current time; wherein the preset value is 10 in the embodiment, and the preset value is not specifically limited in the embodiment and can be determined according to specific conditions. ​​​

[0094] The optimal fast strategy is obtained by long-term cumulative reward of all update steps.

[0095] Step S004: Coordinated optimization of frequency regulation and peak regulation of the molten salt thermal storage system is performed by the optimal slow strategy and the optimal fast strategy.

[0096] Coordinated optimization of frequency regulation and peak regulation of the molten salt thermal storage system is performed by the optimal slow strategy and the optimal fast strategy, so as to reduce the calculation complexity and shorten the training time consumption.

[0097] Through the above-mentioned coordinated optimization, the frequency regulation response speed of the molten salt thermal storage system is improved to 80 ms, the frequency deviation is reduced by 32%, the peak regulation economy is improved, the coal power start-stop cost is reduced by 18%, and the energy storage life is prolonged, that is, the SOC (State of Charge) fluctuation range is reduced to 45%.

[0098] As shown in Figure 2 The second aspect of the present application provides a frequency regulation-peak regulation coordinated optimization system of a molten salt thermal storage system, comprising the following steps:

[0099] The data acquisition module 101 is used to acquire real-time state characteristic data in the molten salt thermal storage system and related state characteristic data affecting the efficiency of the molten salt thermal storage system, and to acquire actual operation action characteristic data and control management action characteristic data in the molten salt thermal storage system.

[0100] The long-time scale analysis module 102 is used to determine the optimal slow strategy of the molten salt thermal storage system in a long-time scale peak regulation process based on the related state characteristic data and the control management action characteristic data through a reinforcement learning algorithm, to determine the reference power output plan of the generator set through the optimal slow strategy in the peak regulation process, and to determine the reference frequency of the power grid at each time through the reference power output plan.

[0101] The short-time scale analysis module 103 is used to determine the optimal fast strategy of the molten salt thermal storage system in a short-time scale frequency regulation process through a reinforcement learning algorithm according to the reference frequency of the power grid at each time, the real-time state characteristic data and the actual operation action characteristic data.

[0102] The coordinated optimization module 104 is used to perform coordinated optimization of frequency regulation and peak regulation of the molten salt thermal storage system through the optimal slow strategy and the optimal fast strategy.

[0103] The third aspect of the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements a frequency regulation-peak regulation coordinated optimization method of a molten salt thermal storage system when executing the computer program.

[0104] The fourth aspect of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, the computer program is executed by a processor to realize a frequency modulation-peak regulation collaborative optimization method of a molten salt heat storage system.

[0105] Those skilled in the art will appreciate that embodiments of the present application can be provided as methods, systems, or computer program products. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can be embodied in the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk memory, optical memory, etc.) having computer-usable program code embodied thereon.

[0106] The present application is described with reference to flowcharts and / or block diagrams of methods, systems, and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions, which are executed via the processor of the computer or other programmable data processing apparatus, generate a means for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the functions specified in the flowchart

[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction means, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the functions specified in the flowchart

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable data processing apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable data processing apparatus provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 one or more flows and / or blocks Figure 1 means for carrying out the functions specified in the flowchart

[0109] It should be pointed out finally that the above embodiments are only used to illustrate the technical solutions of the present application but not to limit it. Although the present application has been described in detail with reference to the above embodiments, it should be understood by those skilled in the art that the specific embodiments of the present application can be modified or equivalently replaced without departing from the spirit and scope of the present application, and any modification or equivalent replacement should be covered within the protection scope of the present application.

Claims

1. A method for frequency modulation and peak shaving collaborative optimization of a molten salt heat storage system, characterized in that, The application relates to a method for frequency and peak load collaborative optimization of a molten salt thermal storage system. Real-time state characteristic data and related state characteristic data influencing the efficiency of the molten salt thermal storage system are acquired; Actual operation action characteristic data and control management action characteristic data in the molten salt thermal storage system are further acquired; Based on the related state characteristic data and the control management action characteristic data, an optimal slow strategy of the molten salt thermal storage system in a long-time-scale peak load regulation process is determined through a reinforcement learning algorithm, a reference output plan of a generator set is determined through the optimal slow strategy in the peak load regulation process, and a reference frequency of a power grid at each moment is determined through the reference output plan. According to the reference frequency of the power grid at each time, real-time state feature data and actual operation action feature data, an optimal fast strategy of the molten salt heat storage system in a short-time scale frequency modulation process is determined through a reinforcement learning algorithm, including continuously acquiring, at preset time intervals real-time state feature data of the molten salt heat storage system at all times before a preset time length of a generator or an energy storage device; a frequency stability constraint is constructed through the reference frequency of the power grid at each time, and a penalty value when the frequency deviates from a target frequency is obtained, and the penalty value is recorded as a penalty value of a first constraint; a generator and energy storage capacity constraint is constructed through the maximum output power of the generator or the energy storage device, and a penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is obtained, and the penalty value is recorded as a penalty value of a second constraint; a second state space of a short-time scale is constructed through the real-time state feature data, and a second action space of a short-time scale is constructed through the actual operation action feature data; a target function of the fast strategy is constructed through the penalty values of the constraints, the second state space and the second action space of the short-time scale; and the optimal fast strategy of the molten salt heat storage system in the short-time scale frequency modulation process is determined through the target function of the fast strategy. The optimal slow strategy and an optimal fast strategy are used for collaborative optimization of frequency and peak load regulation of the molten salt thermal storage system.

2. The frequency modulation-peak regulation collaborative optimization method of a molten salt heat storage system according to claim 1, characterized in that, The optimal slow strategy of the molten salt thermal storage system in the long-time-scale peak load regulation process is determined through the related state characteristic data and the control management action characteristic data and a reinforcement learning algorithm. The optimal slow strategy of the molten salt thermal storage system in the long-time-scale peak load regulation process is determined through the first state space and the first action space of the long-time scale and a policy-based reinforcement learning algorithm.

3. The frequency modulation-peak regulation collaborative optimization method of a molten salt heat storage system according to claim 2, characterized in that, The optimal slow strategy of the molten salt thermal storage system in the long-time-scale peak load regulation process is determined through the first state space and the first action space of the long-time scale and a policy-based reinforcement learning algorithm. A target function of the slow strategy is determined according to the first state space and the first action space of the long-time scale, and the target function of the slow strategy is specifically represented as: wherein, is a discount factor, denotes the discount factor in the time step, denotes the reward value in the time step, denotes the total number of steps in all time steps during the interaction of the environment with the agent, denotes the expected operation, i.e. the expected value of making a decision according to the policy π, denotes the long-term cumulative reward in all time steps; The optimal slow strategy is obtained through long-term cumulative rewards of all time steps.

4. The frequency modulation-peak regulation collaborative optimization method of a molten salt heat storage system according to claim 1, characterized in that, The frequency stability constraint is constructed through the reference frequency of the power grid at each moment, and a penalty value when the frequency deviates from a target frequency is obtained. wherein denotes the frequency of the power grid at each time instant, denotes the reference frequency of the power grid at each time instant, denotes the preset frequency error, denotes the frequency violation penalty coefficient, denotes the max function, denotes the absolute value symbol, denotes the penalty value when the frequency deviates from the target frequency; The frequency violation penalty coefficient is determined through hyperparameter optimization.

5. The method of claim 1, wherein, The generator and energy storage capacity constraint is constructed through the maximum output power of the generator or the energy storage device, and a penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is obtained. wherein represents the generator or energy storage device output power at each time instant, represents the maximum output power of the generator or energy storage device; represents the max function, represents the output power violation penalty coefficient, represents the penalty value when the maximum output power limit of the generator or energy storage device is exceeded; The penalty value when the maximum output power limit of the generator or the energy storage device is exceeded is determined through hyperparameter optimization.

6. The frequency modulation-peak regulation collaborative optimization method of a fused salt heat storage system according to claim 1, characterized in that, The target function of the fast strategy is constructed through the penalty values of the constraints, the second state space and the second action space of the short-time scale; The optimal fast strategy of the molten salt thermal storage system in a short-time-scale frequency regulation process is determined through the target function of the fast strategy. Indicates the first Discount factor in each update step, Indicates the first The reward value for each update step. This represents the total number of update steps during the interaction between the environment and the agent. Indicates the first The penalty value for each constraint. Indicates the first The severity of the penalty for each constraint This indicates the desired operation, i.e., according to the strategy. Expected value for making decisions This represents the long-term cumulative reward for all update steps. This indicates the number of all constraints. This indicates that the initial update interval is adjusted, meaning the duration of one update step is equal to the adjusted initial update interval; the penalty strength for each constraint is determined through hyperparameter tuning. In the formula, represents the value of the first real-time state feature at the current time at the first reference time, represents the average value of the values of the first real-time state feature data at all reference times at the current time, represents the number of reference times, represents the number of all real-time state features, represents an absolute value symbol, represents a linear normalization function, represents a preset initial update time interval; wherein a preset value of time before the current time is taken as a reference time of the current time; the current time. The optimal fast strategy is obtained through long-term cumulative rewards of all update steps.

7. A frequency modulation-peak regulation collaborative optimization system of a molten salt heat storage system, characterized in that, The method comprises a data acquisition module, a long-time-scale analysis module, a short-time-scale analysis module and a collaborative optimization module, and the data acquisition module, the long-time-scale analysis module, the short-time-scale analysis module and the collaborative optimization module realize the frequency and peak load collaborative optimization method of the molten salt thermal storage system.

8. An electronic device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the frequency modulation-peak regulation collaborative optimization method of the molten salt heat storage system according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the frequency modulation-peak regulation collaborative optimization method of the molten salt heat storage system according to any one of claims 1-6.

Citation Information

Patent Citations

  • AGC unit dynamic optimization method based on deep reinforcement learning

    CN112186811A

  • Platform and method for participating in power grid frequency modulation based on MADDPG large-scale energy storage

    CN117353343A