Heat storage equipment system applied to photo-thermal energy storage and control method thereof

By collecting temperature and environmental data in real time and using reinforcement learning to optimize the flow rate and guiding structure of the central control unit, the problems of thermal stratification fluctuations and low energy efficiency in solar thermal energy storage equipment are solved, and dynamic optimization of thermal storage capacity and improvement of system energy efficiency are achieved.

CN121383459APending Publication Date: 2026-01-23SHANDONG ANYUE ENERGY SAVING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511412416.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing solar thermal energy storage equipment suffers from problems such as thermal stratification leading to unsteady fluctuations, low utilization rate of thermal storage capacity, easy equipment fatigue and slow control response, and low energy efficiency.

Method used

By collecting real-time temperature data inside the molten salt thermal storage tank and external environmental data, the central control unit uses reinforcement learning to calculate the optimal control action, drives the flow control valve and the flow guiding structure to adjust the molten salt flow, and combines the reward value to optimize the control strategy, thereby achieving dynamic optimization of thermal stratification in the molten salt thermal storage tank.

Benefits of technology

It effectively suppresses the unsteady fluctuations of the thermal stratification temperature gradient layer, improves the utilization rate of thermal storage capacity, reduces equipment fatigue, and enhances system energy efficiency and economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121383459A_ABST
    Figure CN121383459A_ABST
Patent Text Reader

Abstract

The invention discloses a heat storage equipment system applied to photo-thermal energy storage and a control method of the heat storage equipment system, and relates to the technical field of photo-thermal energy storage. Comprising the steps that S1, temperature data, detected by a temperature sensor array, of different heights in the fused salt heat storage tank and external environment data obtained by an external environment monitoring module are collected in real time, and the external environment data comprise illumination intensity, power grid loads and the real-time price of the electricity market; temperature data of different heights of a fused salt heat storage tank are collected in real time through a temperature sensor array, illumination intensity, power grid load and real-time price data of an electric power market are obtained in combination with an external environment monitoring module, and the data are input into a reinforcement learning central control unit after being preprocessed through filtering, denoising and the like; the unit outputs an optimal control action based on a trained model, a flow control valve is driven to accurately adjust the molten salt charging and discharging speed, a flow field in the tank is optimized through a flow guide structure, and unsteady fluctuation of a thermal stratification temperature gradient layer is effectively restrained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of photo-thermal energy storage, in particular to a heat storage equipment system applied to photo-thermal energy storage and a control method thereof. BACKGROUND

[0002] Photo-thermal energy storage is an advanced technology that converts solar energy into thermal energy and stores it, widely used in the field of solar thermal power generation. The system concentrates sunlight on an absorber through a concentrator, causing the heat transfer fluid, such as molten salt, to heat up, and then transports the high-temperature fluid to the heat storage equipment to store thermal energy. The heat storage equipment is usually composed of an insulated tank filled with heat storage materials, which can release heat in the absence of sunlight to continuously drive the turbine to generate electricity, effectively overcoming the intermittency of solar energy and improving the stability and efficiency of energy supply.

[0003] Although photo-thermal energy storage technology has made significant progress, the existing heat storage equipment system still has a subtle and specific operation problem. During the heat storage process, the inflow and outflow of the heat transfer fluid can cause thermal stratification in the tank, i.e., a temperature gradient layer is formed between the high-temperature fluid and the low-temperature fluid. However, due to the inertial effect of fluid flow and the complex interaction of thermal convection, this temperature gradient layer is prone to non-steady-state fluctuations, causing uneven heat distribution. This instability not only causes the actual utilization rate of heat storage capacity to decline, but also induces local thermal stress in frequent heat charging and discharging cycles, accelerating the fatigue of equipment materials. The current control method often relies on simple threshold adjustment and cannot dynamically optimize fluid parameters to suppress stratification fluctuations, making the system respond slowly when dealing with sudden changes in solar input and restricting the overall energy efficiency. SUMMARY

[0004] In view of the shortcomings of the prior art, the present application provides a heat storage equipment system applied to photo-thermal energy storage and a control method thereof, which solves the problems of non-steady-state fluctuations of thermal stratification, low utilization rate of heat storage capacity, fatigue of equipment, slow control response, and low energy efficiency compared with the prior art.

[0005] To achieve the above purpose, the present application is implemented by the following technical scheme: a heat storage equipment control method applied to photo-thermal energy storage, comprising:

[0006] S1, real-time collection of temperature data at different heights in the molten salt heat storage tank detected by a temperature sensor array and external environment data obtained by an external environment monitoring module, the external environment data including light intensity, grid load, and real-time electricity market price;

[0007] S2, the data acquisition and preprocessing module filters, denoises and normalizes the collected temperature data to eliminate measurement errors, and combines the preprocessed temperature data with external environmental data to form a current state observation value;

[0008] S3, inputting the state observation value into the reinforcement learning central control unit;

[0009] S4, the reinforcement learning central control unit calculates and outputs the optimal control action according to the current state observation value combined with the trained reinforcement learning model, the optimal control action including the opening degree instruction of the flow control valve and the parameter instruction of the flow guide structure;

[0010] S5, the actuator driving module receives the optimal control action and converts it into a driving signal to drive the flow control valve to adjust its opening degree, and to drive the flow guide structure to adjust its physical form, so as to dynamically adjust the filling speed, discharge speed and internal flow path of the molten salt;

[0011] S6, evaluating the influence of the optimal control action on the thermal stratification effect of the molten salt thermal storage tank, the system operation efficiency or the economic benefit, calculating the reward value based on at least one of the clarity of the thermal stratification interface, the energy utilization rate, the system response time or the operation cost;

[0012] S7, the reinforcement learning central control unit updates the control strategy in the reinforcement learning model based on the reward value to optimize its decision-making ability under different working conditions;

[0013] S8, repeating steps S1 to S7 to achieve continuous dynamic optimization of the thermal stratification effect of the molten salt thermal storage tank.

[0014] Further, the reinforcement learning model is a deep Q network DQN, an actor-critic Actor-Critic algorithm or a deep deterministic policy gradient DDPG algorithm.

[0015] Further, the flow guide structure includes an adjustable diffuser, a perforated plate or a flow guide vane.

[0016] Further, the formation of the current state observation value includes feature extraction of the preprocessed temperature data and integration with external environmental data to comprehensively reflect the internal thermodynamic state of the molten salt thermal storage tank and the external operating environment conditions.

[0017] Further, the evaluation of the influence of the optimal control action on the thermal stratification effect of the molten salt thermal storage tank includes calculating the temperature gradient to quantify the clarity of the thermal stratification interface according to the temperature data at different heights monitored by the temperature sensor array, calculating the energy utilization rate and the energy storage efficiency combined with the inlet and outlet temperatures and flow of the molten salt, and calculating the system response time according to the operation instruction and the actual response time.

[0018] Further, the actuator driving module receives the opening degree instruction output by the reinforcement learning central control unit, and drives the flow control valve to adjust the opening degree thereof, so as to accurately control the filling amount of the high-temperature molten salt or the discharging amount of the low-temperature molten salt; the actuator driving module also receives the adjustment parameter instruction, and drives the flow guide structure to adjust the flow guide angle, the aperture distribution or the position thereof, so as to change the flow field distribution and the mixing area of the molten salt in the heat storage tank.

[0019] Further, the control strategy in the reinforcement learning model is updated, including using the current state observation value, the executed action and the obtained reward value, adjusting the weight, the bias or the policy function parameter of the reinforcement learning model through the back propagation algorithm or the policy gradient algorithm, so as to enhance the ability of the reinforcement learning model to select a more optimal action in future decision-making.

[0020] Further, the steps S1 to S7 are cyclically executed, so as to continuously and dynamically optimize the heat stratification effect of the molten salt heat storage tank, and enable the system to adaptively maintain the heat stratification stability under the conditions of the change of the light intensity, the fluctuation of the power grid load or the disturbance of the system.

[0021] The application further provides a heat storage equipment system applied to the photo-thermal energy storage.

[0022] The molten salt heat storage tank is used for storing high-temperature molten salt and low-temperature molten salt, and is provided with a plurality of molten salt inlets and outlets.

[0023] The temperature sensor array is arranged at different heights in the molten salt heat storage tank respectively, and is connected with the data acquisition and preprocessing module, and is used for monitoring the temperature distribution and the heat stratification interface position information of the molten salt in the molten salt heat storage tank in real time.

[0024] The flow control valve is arranged on the inlet and outlet pipelines of the molten salt heat storage tank respectively, and is connected with the actuator driving module, and is used for accurately adjusting the filling flow and the discharging flow of the molten salt according to the instruction from the actuator driving module.

[0025] The flow guide structure is arranged in the molten salt heat storage tank, and is connected with the actuator driving module, and is used for guiding the molten salt flow in the molten salt heat storage tank, reducing the impact mixing of the molten salt and promoting the stable heat stratification.

[0026] The data acquisition and preprocessing module is connected with the temperature sensor array, and is used for receiving the real-time temperature data collected by the temperature sensor array, and performing filtering, denoising and normalization processing on the original temperature data, so as to eliminate the measurement error, and converting the processed temperature data into the state observation value recognizable by the reinforcement learning central control unit.

[0027] An external environment monitoring module is connected with the reinforcement learning central control unit, and is used for acquiring external environment information in real time, wherein the external environment information includes illumination intensity, power grid load demand and real-time price of the power market.

[0028] The reinforcement learning central control unit is connected with the data acquisition and preprocessing module and the external environment monitoring module respectively, and is used for receiving state observation values and external environment information as inputs, running a pre-trained reinforcement learning model internally, and outputting control action instructions according to the state observation values and the external environment information.

[0029] The actuator driving module is connected with the reinforcement learning central control unit, and is used for receiving the control action instructions output by the reinforcement learning central control unit, and converting the instructions into driving signals to drive the flow control valve to adjust the opening degree thereof and drive the flow guide structure to adjust the physical form thereof.

[0030] Further, the reinforcement learning central control unit trains and optimizes the control strategy thereof by using real-time temperature data, external environment data and reward values, so as to maximize the heat stratification effect, the system operation efficiency or the economic benefit, wherein the reward values include at least one of the following: the clarity of the heat stratification interface, the energy utilization rate, the system response time or the operation cost.

[0031] Compared with the prior art, the present application has the following advantages:

[0032] In the present application, the temperature sensor array is used to collect temperature data at different heights of the molten salt heat storage tank in real time, the illumination intensity, the power grid load and the real-time price of the power market are acquired by the external environment monitoring module, and the data are input into the reinforcement learning central control unit after being filtered and denoised. The unit outputs optimal control actions based on the trained model, drives the flow control valve to accurately adjust the charging and discharging speed of the molten salt, optimizes the flow field in the tank by the flow guide structure, and effectively suppresses the non-steady-state fluctuation of the heat stratification temperature gradient layer. At the same time, the control effect is evaluated and the reward value is calculated to update the model, so that the system can adapt to the working conditions such as illumination change and power grid load fluctuation, improve the utilization rate of the heat storage capacity, reduce local thermal stress to delay equipment fatigue, solve the problems of slow response and low energy efficiency of the existing threshold control, and realize continuous dynamic optimization of heat stratification and improvement of system energy efficiency and economic benefit. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 The figure is a method flowchart of the present application;

[0034] Figure 2 The figure is a system structure diagram of the present application. DETAILED DESCRIPTION

[0035] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0036] Please refer to Figure 1 The present application provides a heat storage device control method applied to photo-thermal energy storage, comprising:

[0037] S1, real-time collection of temperature data at different heights detected by a temperature sensor array in a molten salt heat storage tank, and external environment data obtained by an external environment monitoring module, the external environment data including light intensity, power grid load and real-time electricity market price;

[0038] S2, the data collection and preprocessing module filters, denoises and normalizes the collected temperature data for preprocessing to eliminate measurement errors, and combines the preprocessed temperature data with the external environment data to form a current state observation value;

[0039] S3, inputting the state observation value to a reinforcement learning central control unit;

[0040] S4, the reinforcement learning central control unit calculates and outputs an optimal control action according to the current state observation value in combination with a trained reinforcement learning model, the optimal control action including an opening degree instruction for adjusting a flow control valve and a parameter instruction for adjusting a flow guide structure;

[0041] S5, the actuator driving module receives the optimal control action and converts it into a driving signal to drive the flow control valve to adjust its opening degree, and to drive the flow guide structure to adjust its physical form, so as to dynamically adjust the filling speed, discharge speed and internal flow path of the molten salt;

[0042] S6, evaluating the influence of the optimal control action on the heat stratification effect of the molten salt heat storage tank, the system operation efficiency or the economic benefit, calculating a reward value based on at least one of the clarity of the heat stratification interface, the energy utilization rate, the system response time or the operation cost;

[0043] S7, the reinforcement learning central control unit updates the control strategy in the reinforcement learning model based on the reward value to optimize its decision-making ability under different working conditions;

[0044] S8, the steps S1 to S7 are repeatedly executed to realize continuous dynamic optimization of the heat stratification effect of the molten salt heat storage tank.

[0045] Specifically, in practical applications, an array of temperature sensors is first arranged at different heights inside the molten salt heat storage tank to ensure that the temperature of each region in the tank can be fully captured. The external environment monitoring module obtains the light intensity, power grid load, and real-time price of the power market in real time through a light detector, a power grid load monitoring terminal, and a power market data interface. The data acquisition and preprocessing module processes the temperature data using mean filtering and wavelet denoising methods, and then performs normalization to eliminate errors of different sensors. Subsequently, the temperature gradient and the position of the stratification interface in the temperature data are extracted as features, which are integrated with the external environment data to obtain the state observation value. The reinforcement learning central control unit has a model trained under a large number of working conditions. After receiving the state observation value, the model quickly calculates the flow control valve opening degree and the guide structure parameter command that adapt to the current working condition. The actuator driving module converts the command into an electrical signal to drive the flow control valve to accurately adjust the molten salt charging and discharging speed, and simultaneously controls the guide structure to change its physical form to optimize the molten salt internal flow path.

[0046] Then, the effect of the control action needs to be evaluated and the reward value is calculated. The reward value considers the thermal stratification interface clarity, energy utilization rate, system response time, and operating cost. Since the dimensions of each index are different, the single index is first normalized. For the thermal stratification interface clarity S and the energy utilization rate E, which are positive indicators, the larger the value, the better. The positive normalization formula is:

[0047]

[0048] where X' is the normalized value, X is the original value of the index, X min is the minimum value of the index in the historical working condition, and X max is the maximum value of the index in the historical working condition. For the system response time T and the operating cost C, which are negative indicators, the smaller the value, the better. The negative normalization formula is:

[0049]

[0050] The parameter definition is consistent with the positive normalization formula.

[0051] Then, the weights of each index are determined by the analytic hierarchy process (AHP). Experts in the field of solar thermal energy storage are invited to compare the importance of the four indexes pairwise, construct a judgment matrix, and calculate the weight coefficient of the thermal stratification interface clarity w1, the weight of the energy utilization rate w2, the weight of the system response time w3, and the weight of the operating cost w4, which satisfy w1+w2+w3+w4=1. The final calculation formula of the comprehensive reward value R is:

[0052] R=w1·S'+w2·E'+w3·T'+w4·C';

[0053] where S' is the normalized thermal stratification interface sharpness, E' is the normalized energy utilization, T' is the normalized system response time, and C' is the normalized operation cost.

[0054] The reinforcement learning central control unit adjusts the model parameter optimization control strategy according to the reward value R, and such a cycle operation can effectively suppress thermal stratification non-steady-state fluctuations, improve the utilization rate of heat storage capacity, reduce local thermal stress, and accelerate the response speed of the system to solar input mutations.

[0055] In this embodiment, the reinforcement learning model is a deep Q network DQN, an actor-critic algorithm, or a deep deterministic policy gradient DDPG algorithm.

[0056] Specifically, in actual implementation, if the system has a high accuracy requirement for control actions and the action type is discrete, a deep Q network DQN can be selected as the reinforcement learning model. The deep Q network DQN stores historical state-action-reward data through an experience replay mechanism, randomly samples data to train the network, and can improve the stability of the model, adapting to the scene of step adjustment of the flow control valve opening degree. When the system needs to simultaneously implement action decision and value evaluation and cope with continuously changing working conditions, an actor-critic algorithm is adopted. The actor network outputs a control action based on the current state, and the critic network evaluates the value of the action by calculating the time difference error. The two networks are updated and optimized in coordination, which is suitable for dynamic fine-tuning of the guide flow structure parameters. If the system working conditions are complex and the control action dimension is high, a deep deterministic policy gradient DDPG algorithm is more suitable. The deep deterministic policy gradient DDPG algorithm outputs continuous actions through a deterministic policy network, combines a target network and experience replay, and can stably learn in a high-dimensional continuous action space, meeting the demand for simultaneous adjustment of the molten salt charging and discharging speed and the guide flow structure parameters. After adopting these algorithms, the decision-making accuracy and response speed of the reinforcement learning model are significantly improved, and the model can better adapt to different operating conditions.

[0057] In this embodiment, the guide flow structure includes an adjustable diffuser, a porous plate, or a guide vane.

[0058] Specifically, the implementation of the guide flow structure needs to be combined with the size of the molten salt heat storage tank and the flow characteristics of the molten salt. When an adjustable diffuser is selected, it is installed at the molten salt inlet. The diffuser angle is adjusted through a driving mechanism to change the diffusion range of the molten salt entering the tank, avoiding direct impact of the molten salt on the low-temperature area in the tank. When a porous plate is used, the pore size and distribution of the porous plate are adjusted through an actuator according to the stratification of the molten salt in the tank to control the penetration speed of the molten salt in different areas and promote uniform flow of the molten salt. When a guide vane is used, the vane is arranged along the circumference and axis of the tank wall. The vane angle is adjusted to guide the molten salt to flow along a preset path, reducing the mixing of the molten salt. The adjustment of these guide flow structures can effectively optimize the flow field in the tank, reduce the impact and mixing of the molten salt, promote stable thermal stratification, and improve the heat storage efficiency.

[0059] In the present embodiment, forming the current state observation value includes feature extraction on the pre-processed temperature data and integration with external environment data to comprehensively reflect the internal thermodynamic state of the molten salt heat storage tank and the external operating environment conditions.

[0060] Specifically, in forming the current state observation value, the pre-processed temperature data is first subjected to feature extraction, the temperature gradient feature is obtained by calculating the temperature difference of adjacent height sensors, the thickness and position feature of the thermal stratification interface are determined in combination with the sensor distribution density, and the average temperature, temperature fluctuation amplitude and other features in the tank are extracted. Subsequently, the temperature features are unified in data format and aligned in dimension with the light intensity, power grid load and real-time price data obtained by the external environment monitoring module, and are integrated into multi-dimensional state observation values. This process can comprehensively reflect the internal thermodynamic state of the molten salt heat storage tank and the external operating environment conditions, providing a comprehensive decision basis for the reinforcement learning central control unit and improving the adaptability of the control strategy.

[0061] In the present embodiment, evaluating the influence of the optimal control action on the thermal stratification effect of the molten salt heat storage tank includes calculating the temperature gradient to quantify the clarity of the thermal stratification interface according to the temperature data at different heights monitored by the temperature sensor array, calculating the energy utilization rate and energy storage efficiency in combination with the inlet and outlet temperatures and flow rate of the molten salt, and calculating the system response time according to the operation instruction and actual response time.

[0062] Specifically, in evaluating the influence of the optimal control action, the temperature gradient is obtained by calculating the temperature change per unit height based on the temperature data at different heights collected by the temperature sensor array, the higher the stability of the temperature gradient, the higher the clarity of the thermal stratification interface, this data is obtained by the S index calculated by the reward value; the energy utilization rate and energy storage efficiency are obtained by calculating the ratio of input and output energy according to the energy conservation principle in combination with the inlet and outlet temperatures and flow rate of the molten salt, the energy utilization rate data is used to obtain the E index calculated by the reward value; the system response time is obtained by recording the time when the reinforcement learning central control unit outputs the operation instruction and the time when the system reaches a stable state after the actuator completes the action, the difference between the two is the system response time, this data is used to obtain the T index calculated by the reward value; the operating cost is obtained by statistical data of the association between the operating cost of the molten salt conveying energy consumption equipment and the electricity market price, this data is used to obtain the C index calculated by the reward value. Through these evaluation methods, the influence of the control action on the thermal stratification effect and system efficiency can be accurately judged, providing a reliable basis for reward value calculation and assisting model optimization.

[0063] In this embodiment, the actuator driving module receives the opening degree instruction output by the reinforcement learning central control unit, and drives the flow control valve to adjust its opening degree, so as to accurately control the charging amount of high-temperature molten salt or the discharging amount of low-temperature molten salt; the actuator driving module also receives the adjustment parameter instruction, and drives the flow guide structure to adjust its flow guide angle, aperture distribution or position, so as to change the flow field distribution and mixing area of the molten salt inside the heat storage tank.

[0064] Specifically, after receiving the opening degree instruction, the actuator driving module drives the valve core of the flow control valve to move through the motor, so as to accurately adjust the valve opening degree; when the charging amount of high-temperature molten salt needs to be increased, the opening degree is increased, and when the discharging amount of low-temperature molten salt needs to be reduced, the opening degree is reduced, so as to realize accurate control of the molten salt flow; after receiving the adjustment parameter instruction of the flow guide structure, for the diffuser, the adjustment mechanism is driven to change the diffusion angle, for the porous plate, the aperture size is adjusted through the hydraulic device, and for the flow guide vane, the vane deflection angle is adjusted by using the servo motor, so as to change the flow field distribution and mixing area of the molten salt inside the heat storage tank. This embodiment can realize fine control of the molten salt flow, effectively maintain the stability of thermal stratification, and improve the system operation efficiency.

[0065] In this embodiment, the control strategy in the reinforcement learning model is updated, including using the current state observation value, the executed action and the obtained reward value, adjusting the weight, bias or policy function parameter of the reinforcement learning model through the back propagation algorithm or the policy gradient algorithm, so as to enhance its ability to select better actions in future decisions.

[0066] Specifically, when updating the control strategy of the reinforcement learning model, the control action executed by the current state observation value and the reward value R form a training sample. If the back propagation algorithm is used, the reward value R is taken as the objective function, the error between the model prediction value and R is calculated, the gradient of each layer is calculated in reverse through the chain rule, the weight and bias of the model are adjusted, and the prediction error is reduced; if the policy gradient algorithm is used, the policy gradient is calculated according to the reward value R, the parameters of the policy function are adjusted through the gradient ascent method, and the probability of selecting the action with high reward value in subsequent decisions is increased. By continuously adjusting the model parameters, the reinforcement learning model can more accurately select the control action that adapts to the working condition in subsequent decisions, and improve the ability of the system to adapt to different working conditions.

[0067] In this embodiment, steps S1 to S7 are repeatedly executed, so as to continuously and dynamically optimize the thermal stratification effect of the molten salt heat storage tank, so that the system can adaptively maintain the stability of thermal stratification under the conditions of changes in light intensity, fluctuations in power grid load or system disturbances.

[0068] Specifically, during system operation, the process of data acquisition and preprocessing model decision action execution effect evaluation and model updating is continuously executed in a loop. When the light intensity suddenly increases, the system quickly captures the temperature change in the tank through the temperature sensor, the reinforcement learning central control unit adjusts the flow control valve opening degree and the flow guide structure parameters in time based on the real-time state observation value combined with the reward value calculation logic, to avoid thermal stratification fluctuation; when the power grid load fluctuates, the molten salt charging and discharging strategy is optimized combined with the real-time price of the electricity market, and the system efficiency and economic benefit are balanced through the reward value formula. This cyclic optimization method enables the system to adapt to different disturbance conditions and continuously maintain the stability of thermal stratification, ensuring long-term efficient operation of the system.

[0069] Referring to Figure 2 The application further provides a heat storage device system applied to photo-thermal energy storage, applied to the heat storage device control method applied to photo-thermal energy storage, comprising:

[0070] A molten salt heat storage tank is used to store high-temperature molten salt and low-temperature molten salt, and is provided with a plurality of molten salt inlets and outlets;

[0071] A temperature sensor array is arranged at different heights in the molten salt heat storage tank and is connected with the data acquisition and preprocessing module, and is used to monitor the temperature distribution and thermal stratification interface position information of the molten salt in the molten salt heat storage tank in real time;

[0072] A flow control valve is arranged on the inlet and outlet pipelines of the molten salt heat storage tank and is connected with the actuator driving module, and is used to accurately adjust the charging flow and discharging flow of the molten salt according to the instructions from the actuator driving module;

[0073] A flow guide structure is arranged in the molten salt heat storage tank and is connected with the actuator driving module, and is used to guide the flow of the molten salt in the molten salt heat storage tank, reduce the impact mixing of the molten salt and promote stable thermal stratification;

[0074] The data acquisition and preprocessing module is connected with the temperature sensor array, and is used to receive the real-time temperature data collected by the temperature sensor array, and perform filtering, denoising and normalization processing on the original temperature data to eliminate measurement errors, and convert the processed temperature data into state observation values recognizable by the reinforcement learning central control unit;

[0075] An external environment monitoring module is connected with the reinforcement learning central control unit, and is used to acquire external environment information in real time, wherein the external environment information includes light intensity, power grid load demand and real-time price of the electricity market;

[0076] The reinforcement learning central control unit is connected with the data acquisition and preprocessing module and the external environment monitoring module respectively, and is used for receiving state observation values and external environment information as inputs, running a pre-trained reinforcement learning model internally, and outputting control action instructions according to the state observation values and the external environment information;

[0077] The actuator driving module is connected with the reinforcement learning central control unit, and is used for receiving the control action instructions output by the reinforcement learning central control unit, and converting the instructions into driving signals to drive the flow control valve to adjust the opening degree thereof and to drive the flow guide structure to adjust the physical form thereof.

[0078] Specifically, the molten salt heat storage tank adopts a double-layer insulation structure, high-temperature molten salt and low-temperature molten salt are respectively stored in the upper and lower parts of the tank, and a plurality of molten salt inlets and outlets are arranged at different heights of the tank wall; an array of temperature sensors is arranged in the tank at the same height difference, and is connected with the data acquisition and preprocessing module through a data line to transmit temperature data in real time; flow control valves are installed on the molten salt inlet and outlet pipelines, a displacement sensor is arranged in the valve body to feed back the opening degree information, and a closed-loop control is formed with the actuator driving module; a flow guide structure is installed at the inlet and outlet and at key positions in the tank according to the tank type; the data acquisition and preprocessing module realizes data filtering and normalization through a hardware circuit, and transmits the processed data to the reinforcement learning central control unit; the external environment monitoring module accesses related monitoring equipment through wireless or wired mode, acquires environmental data in real time, and sends the data to the control unit; the reinforcement learning central control unit runs the model using an embedded processor, and outputs control instructions in combination with reward value calculation logic; the actuator driving module converts the instructions into driving signals to control the valve and the flow guide structure. The components work cooperatively to effectively solve the thermal stratification fluctuation problem and improve the system performance.

[0079] In the embodiment, the reinforcement learning central control unit trains and optimizes its control strategy by using real-time temperature data, external environment data and reward values to maximize the thermal stratification effect, system operation efficiency or economic benefit, and the reward values include at least one of the clarity of the thermal stratification interface, energy utilization rate, system response time or operation cost.

[0080] Specifically, the reinforcement learning central control unit pre-trains the model using simulated working condition data in the initial stage, receives temperature data and external environment data in real time after the system is running, and continuously updates the training sample library in combination with the reward value R after performing the control action. When the clarity of the thermal stratification interface is improved, the energy utilization rate is increased, the system response time is shortened or the operation cost is reduced, the reward value R is increased, and the control unit adjusts the model parameters and optimizes the control strategy by using a back propagation algorithm or a policy gradient algorithm accordingly. After continuous training, the model can more accurately match different working conditions, ensure the stability of thermal stratification, improve the system operation efficiency or reduce the operation cost, and realize multi-objective optimization.

[0081] In summary, the application collects temperature data at different heights of the molten salt heat storage tank in real time through a temperature sensor array, combines light intensity, power grid load and real-time electricity market price data obtained by an external environment monitoring module, and inputs the pre-processed data such as filtering and denoising into a reinforcement learning central control unit; the unit outputs optimal control actions based on the trained model to drive the flow control valve to accurately adjust the charging and discharging speed of the molten salt, and to optimize the flow field in the tank through the flow guide structure, effectively suppressing the non-steady-state fluctuations of the thermal stratification temperature gradient layer; at the same time, the control effect is evaluated and the reward value is calculated to update the model, so that the system can adapt to working conditions such as changes in light intensity and fluctuations in power grid load, improve the utilization rate of heat storage capacity, reduce local thermal stress to delay equipment fatigue, solve the problems of slow response and low energy efficiency of the existing threshold control, and realize continuous dynamic optimization of thermal stratification and improvement of system energy efficiency and economic benefits.

[0082] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device.

[0083] Although embodiments of the application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and alterations can be made thereto without departing from the principles and spirit of the application, and the scope of the application is defined by the appended claims and their equivalents.

Claims

1. A control method for a thermal storage device applied to solar thermal energy storage, characterized in that, include: S1. Real-time acquisition of temperature data at different heights inside the molten salt thermal storage tank, detected by a temperature sensor array, and external environmental data obtained by an external environmental monitoring module, including light intensity, grid load, and real-time electricity market price; S2. The data acquisition and preprocessing module filters, denoises, and normalizes the acquired temperature data to eliminate measurement errors. It then combines the preprocessed temperature data with external environmental data to form the current state observation value. S3. Input the state observation values ​​into the reinforcement learning central control unit; S4. The reinforcement learning central control unit calculates and outputs the optimal control action based on the current state observation and the trained reinforcement learning model. The optimal control action includes the instruction to adjust the opening of the flow control valve and the instruction to adjust the parameters of the flow guide structure. S5. The actuator drive module receives the optimal control action and converts it into a drive signal to drive the flow control valve to adjust its opening degree. At the same time, it drives the flow guide structure to adjust its physical shape in order to dynamically adjust the charging speed, discharge speed and internal flow path of the molten salt. S6. Evaluate the impact of the optimal control action on the thermal stratification effect, system operating efficiency, or economic benefits of the molten salt thermal storage tank, and calculate the reward value. The reward value is based on at least one of the following: the clarity of the thermal stratification interface, energy utilization rate, system response time, or operating cost. S7. The reinforcement learning central control unit updates the control strategy in the reinforcement learning model based on the reward value, and optimizes its decision-making ability under different operating conditions. S8. Repeat steps S1 to S7 to achieve continuous dynamic optimization of the thermal stratification effect of the molten salt thermal storage tank.

2. The control method for a thermal storage device applied to solar thermal energy storage according to claim 1, characterized in that, The reinforcement learning model is a Deep Q-Network (DQN), an Actor-Critic algorithm, or a Deep Deterministic Policy Gradient (DDPG) algorithm.

3. The control method for a thermal storage device applied to solar thermal energy storage according to claim 1, characterized in that, The flow guiding structure includes an adjustable diffuser, a perforated plate, or flow guide vanes.

4. The control method for a thermal storage device applied to solar thermal energy storage according to claim 1, characterized in that, The formation of current state observations includes feature extraction of preprocessed temperature data and integration with external environmental data to comprehensively reflect the internal thermodynamic state and external operating environment conditions of the molten salt thermal storage tank.

5. The control method for a thermal storage device applied to solar thermal energy storage according to claim 1, characterized in that, The evaluation of the impact of optimal control actions on the thermal stratification effect of the molten salt thermal storage tank includes calculating the temperature gradient based on temperature data at different heights monitored by the temperature sensor array to quantify the clarity of the thermal stratification interface, calculating the energy utilization rate and energy storage efficiency in combination with the inlet and outlet temperatures and flow rates of the molten salt, and calculating the system response time based on the operation instructions and actual response time.

6. The control method for a thermal storage device applied to solar thermal energy storage according to claim 1, characterized in that, The actuator drive module receives the opening command output by the reinforcement learning central control unit and drives the flow control valve to adjust its opening to precisely control the amount of high-temperature molten salt charged or the amount of low-temperature molten salt discharged. The actuator drive module also receives the adjustment parameter command and drives the flow guiding structure to adjust its flow guiding angle, orifice distribution or position to change the flow field distribution and mixing area of ​​molten salt inside the heat storage tank.

7. The control method for a thermal storage device applied to solar thermal energy storage according to claim 1, characterized in that, The updated control policy in the reinforcement learning model includes adjusting the weights, biases, or policy function parameters of the reinforcement learning model using the current state observations, the actions performed, and the reward values ​​obtained, through a backpropagation algorithm or a policy gradient algorithm, in order to enhance its ability to select better actions in future decisions.

8. A control method for a thermal storage device applied to solar thermal energy storage according to claim 1, characterized in that, The cyclic execution steps S1 to S7 enable continuous dynamic optimization of the thermal stratification effect of the molten salt thermal storage tank, allowing the system to adaptively maintain thermal stratification stability under conditions of changes in light intensity, fluctuations in grid load, or system disturbances.

9. A thermal energy storage system for solar thermal energy storage, comprising a control method for a thermal energy storage system for solar thermal energy storage as described in claims 1-8, characterized in that, include: Molten salt thermal storage tanks are used to store high-temperature molten salt and low-temperature molten salt, and are equipped with multiple molten salt inlets and outlets; A temperature sensor array is set at different heights inside the molten salt thermal storage tank and connected to the data acquisition and preprocessing module to monitor the temperature distribution and thermal stratification interface location information of the molten salt inside the molten salt thermal storage tank in real time. Flow control valves are installed on the inlet and outlet pipes of the molten salt thermal storage tank and connected to the actuator drive module. They are used to precisely adjust the filling and discharging flow of molten salt according to the instructions from the actuator drive module. A flow guiding structure is installed inside the molten salt thermal storage tank and connected to the actuator drive module. It is used to guide the flow of molten salt inside the molten salt thermal storage tank, reduce the impact mixing of molten salt, and promote stable thermal stratification. The data acquisition and preprocessing module is connected to the temperature sensor array. It is used to receive real-time temperature data acquired by the temperature sensor array, and to filter, denoise and normalize the raw temperature data to eliminate measurement errors. The processed temperature data is then converted into state observations that can be recognized by the reinforcement learning central control unit. The external environment monitoring module is connected to the reinforcement learning central control unit to acquire external environment information in real time, including light intensity, grid load demand and real-time electricity market prices. The reinforcement learning central control unit is connected to the data acquisition and preprocessing module and the external environment monitoring module, respectively. It is used to receive state observations and external environment information as inputs, run a pre-trained reinforcement learning model internally, and output control action commands based on state observations and external environment information. The actuator drive module is connected to the reinforcement learning central control unit. It receives control action commands output by the reinforcement learning central control unit and converts the commands into drive signals to drive the flow control valve to adjust its opening degree and drive the flow guide structure to adjust its physical shape.

10. A thermal energy storage system for solar thermal energy storage according to claim 9, characterized in that, The reinforcement learning central control unit trains and optimizes its control strategy using real-time temperature data, external environmental data, and reward values ​​to maximize thermal stratification effect, system operating efficiency, or economic benefits. The reward values ​​include at least one of the following: clarity of the thermal stratification interface, energy utilization rate, system response time, or operating cost.