A dehumidifier energy efficiency ratio dynamic optimization control method and system based on reinforcement learning

CN122708400APending Publication Date: 2026-09-08GUANGZHOU RACK TECH ELECTRO-MECHANICAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610683565.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-09-08

AI Technical Summary

Technical Problem

[0004]针对现有技术的不足,本发明提供了一种基于强化学习的除湿机能效比动态寻优控制方法及系统,解决了现有变频除湿机控制技术中,因制冷系统物理响应迟滞与算法寻优时间尺度不匹配,导致供需失衡、控制参数震荡难以稳定收敛的问题

Benefits of technology

(1)、该基于强化学习的除湿机能效比动态寻优控制方法,通过采集环境温湿度、压缩机排气温度等运行数据,并按时间顺序组成工况特征序列,使控制过程能够结合前后工况变化判断设备状态,通过环境湿度变化率计算湿负荷突变强度,并通过环境湿度变化率与压缩机排气温度变化率之间的响应差计算热力响应迟滞程度,使外部湿负荷变化和设备内部热力响应被分开判断。相比仅依靠湿度阈值或单一变化率判断工况的方式,能够更早发现室内湿负荷已经变化而制冷循环尚未完成响应的情况,减少因工况判断滞后造成的控制偏差。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122708400A_ABST
    Figure CN122708400A_ABST
Patent Text Reader

Abstract

This invention discloses a dynamic optimization control method and system for the energy efficiency ratio of dehumidifiers based on reinforcement learning, belonging to the field of dehumidifier control technology. The method collects dehumidifier operating data, preprocesses it to form a time-series operating condition feature sequence, inputs it into an operating condition identification model to obtain the operating condition type, and calculates the intensity of sudden changes in wet load and the degree of hysteresis in thermal response. Based on this, it adjusts the exploration utilization ratio and optimization step size; when the safety margin is sufficient, the step size increases with the intensity of the sudden change; when entering the protection warning interval, the step size is limited to not exceeding the safety upper limit. Operating parameters are output using a dual reward function. During execution, safety parameters are monitored in real time and the operating parameters are corrected. The corrected parameters are written into an experience pool. After satisfying the thermal steady-state criterion, the delayed reward is calculated and the algorithm is updated. This invention decomposes load-side changes and equipment-side responses, enabling optimization to take into account both wet load response and equipment protection. It updates the algorithm with real actions and steady-state feedback, matching experience learning with actual operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dehumidifier control technology, specifically to a dynamic optimization control method and system for dehumidifier energy efficiency ratio based on reinforcement learning. Background Technology

[0002] Variable frequency dehumidifiers are widely used in various scenarios due to their low energy consumption and good dehumidification effect. In actual operation, their energy efficiency ratio is affected by various factors such as ambient temperature and humidity, indoor humidity load, and equipment operating parameters, and can fluctuate by 30% to 50%. In terms of control strategy, existing variable frequency dehumidifiers mostly use fixed combinations of parameters such as compressor frequency, fan speed, and electronic expansion valve opening, or adjust based on conventional optimization methods such as PID and genetic algorithms. It is difficult to adapt to dynamic changes in operating conditions, resulting in insufficient dehumidification speed under high load and high energy consumption under low load.

[0003] The limitations of existing technologies include at least the following problems: When a variable frequency dehumidifier is running, it forms a closed-loop control process of condition perception, parameter optimization, and command execution. The refrigeration system has an inherent physical response delay to changes in thermal state. When the environmental humidity load changes abruptly, the external load demand has changed rapidly, while the internal refrigeration cycle of the equipment is still in a balanced transition state. When conventional optimization algorithms adjust control parameters based on real-time feedback, the physical state of the equipment has not yet reached the steady-state condition corresponding to the current operating condition. The asynchronous supply and demand response causes the control parameters to be repeatedly adjusted near the optimal range, making it difficult to achieve stable convergence. Frequent parameter fluctuations can also cause the compressor exhaust temperature and system high and low pressure to approach the protection threshold, triggering the safety mechanism and interrupting the optimization process. The matching degree between the control strategy and the equipment operating characteristics is insufficient, and the energy efficiency ratio is difficult to maintain at a relatively optimal level under dynamic operating conditions. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a dynamic optimization control method and system for the energy efficiency ratio of dehumidifiers based on reinforcement learning. This solves the problem in existing variable frequency dehumidifier control technologies where the physical response lag of the refrigeration system and the mismatch between the algorithm's optimization time scale lead to supply and demand imbalances and difficulty in achieving stable convergence of control parameter oscillations.

[0005] To achieve the above objectives, this invention provides the following technical solution: a dynamic optimization control method for the energy efficiency ratio of a dehumidifier based on reinforcement learning, comprising the following steps: collecting ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure, and the current operating frequency of the compressor during the operation of the dehumidifier; preprocessing these data to form a time-series of operating condition features; inputting the operating condition feature sequence into an operating condition identification model to obtain the current operating condition type; calculating the intensity of the wet load mutation based on the rate of change of ambient humidity; calculating the degree of thermal response hysteresis based on the response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature; and adjusting the exploration and utilization ratio and optimization step size of the reinforcement learning algorithm according to the intensity of the wet load mutation and the degree of thermal response hysteresis. When the safety margin is sufficient, the optimization step size increases with the intensity of the wet load mutation. When entering the protection warning zone, the optimization step size is limited to not exceeding the upper limit of the safety step size. Taking the operating condition characteristic sequence, the current operating condition type, the intensity of the wet load mutation, and the degree of thermodynamic response hysteresis as inputs, and taking the dual reward function composed of the energy efficiency ratio reward term and the dehumidification deviation suppression term as the objective, the optimal operating parameters are output through reinforcement learning algorithm. When executing the optimal operating parameters, the compressor exhaust temperature, system high pressure, and system low pressure are monitored in real time. When entering the protection warning zone, safety correction is performed, and the corrected actual execution parameters are used as operating instructions and written into the experience pool. After the thermodynamic steady-state criterion is met, the actual energy efficiency ratio and actual humidity data are collected, the delay reward is calculated, and the algorithm parameters are updated.

[0006] Further, the steps for assembling the operating condition feature sequence in chronological order after preprocessing are as follows: the collected ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure, and compressor current operating frequency are subjected to sliding filtering and normalization; the normalization results of multiple consecutive sampling times are arranged in chronological order to form the operating condition feature sequence.

[0007] Further, the steps for calculating the intensity of the wet load mutation and the degree of thermal response hysteresis are as follows: Analyze the trend of environmental humidity changes over multiple consecutive sampling times to obtain the rate of change of environmental humidity; determine the intensity of the wet load mutation based on the rate of change of environmental humidity; analyze the trend of compressor exhaust temperature changes over multiple consecutive sampling times to obtain the rate of change of compressor exhaust temperature; normalize the rate of change of environmental humidity and the rate of change of compressor exhaust temperature respectively to obtain the normalized rate of change of wet load and the normalized rate of change of equipment thermal response; calculate the degree of thermal response hysteresis based on the difference between the normalized rate of change of wet load and the normalized rate of change of equipment thermal response.

[0008] Furthermore, the steps for adjusting the exploration and utilization ratio and optimization step size of the reinforcement learning algorithm based on the intensity of the wet load mutation and the degree of thermal response hysteresis are as follows: determine the exploration ratio of the reinforcement learning algorithm based on the intensity of the wet load mutation; determine the utilization ratio of the reinforcement learning algorithm based on the degree of thermal response hysteresis; when the compressor discharge temperature, system high pressure, and system low pressure have not entered the protection warning range, the optimization step size increases with the increase of the intensity of the wet load mutation; when any one of the compressor discharge temperature, system high pressure, and system low pressure enters the protection warning range, the optimization step size is limited to within the preset safe step size upper limit.

[0009] Furthermore, the steps for constructing the dual reward function, which consists of an energy efficiency ratio reward term and a dehumidification deviation suppression term, are as follows: Collect the actual energy efficiency ratio and actual humidity data of the dehumidifier under the current operating parameters; determine the energy efficiency ratio reward term based on the deviation relationship between the actual energy efficiency ratio and the preset energy efficiency ratio benchmark value; determine the dehumidification deviation suppression term based on the deviation relationship between the actual humidity and the preset target humidity; and combine the energy efficiency ratio reward term and the dehumidification deviation suppression term into the output value of the dual reward function.

[0010] Furthermore, the steps for outputting optimal operating parameters using the reinforcement learning algorithm are as follows: The operating condition feature sequence, current operating condition type, intensity of wet load mutation, and degree of thermal response hysteresis are input into the decision network of the reinforcement learning algorithm; the decision network determines the adjustment range of the compressor operating frequency and the fan speed based on the current operating condition type and degree of thermal response hysteresis; when any one of the compressor exhaust temperature, system high pressure, or system low pressure enters the protection warning interval, the search range of candidate operating parameters is narrowed and the upper limit of the action change rate is reduced; the decision network outputs continuous action parameters based on the current state, and adds perturbations to the neighborhood of the continuous action parameters according to the exploration ratio to generate candidate operating parameters; the evaluation network uses a dual reward function as the optimization objective to evaluate the value of the candidate operating parameters; the candidate operating parameter with the highest value evaluation result is selected as the optimal operating parameter output.

[0011] Further, the steps for writing the corrected actual execution parameters into the experience pool, and then collecting actual energy efficiency ratio and actual humidity data to calculate the delay reward and update the algorithm parameters after the thermodynamic steady-state criterion is met are as follows: The corrected actual execution parameters are combined with the corresponding operating condition feature sequence, wet load mutation intensity, and thermodynamic response hysteresis degree, and then written into the experience pool in a state awaiting reward replenishment; within each sampling period after the actual execution parameters are executed, the compressor exhaust temperature change rate, system high pressure change rate, and system low pressure change rate are calculated; when the absolute values ​​of the compressor exhaust temperature change rate, system high pressure change rate, and system low pressure change rate are all less than their respective steady-state thresholds within a consecutive preset number of sampling periods, the thermodynamic steady-state criterion is determined to be met; after the thermodynamic steady-state criterion is met, actual energy efficiency ratio and actual humidity data are collected, the delay reward value is calculated, the delay reward value is added to the corresponding sample in the experience pool, and the network parameters of the reinforcement learning algorithm are updated according to the delay reward value.

[0012] Further, the steps for inputting the operating condition feature sequence into the operating condition identification model to obtain the current operating condition type are as follows: A hybrid structure of convolutional neural network and long short-term memory network is used to construct the operating condition identification model; static correlation features at each sampling time in the operating condition feature sequence are extracted using the convolutional neural network; temporal change features between multiple consecutive sampling times in the operating condition feature sequence are extracted using the long short-term memory network; the static correlation features and temporal change features are fused to output the current operating condition type, which includes high humidity and high load operating condition, low humidity and low load operating condition, and sudden change in wet load operating condition.

[0013] Furthermore, the safety correction steps when entering the protection warning zone are as follows: during the execution of optimal operating parameters, the compressor exhaust temperature, system high pressure, and system low pressure are collected in real time; when any one of the compressor exhaust temperature, system high pressure, and system low pressure enters the protection warning zone, the adjustment range of the compressor operating frequency and fan speed is limited; the operating parameters after limiting the adjustment range are output as the actual execution parameters after safety correction.

[0014] A dynamic optimization control system for the energy efficiency ratio of a dehumidifier based on reinforcement learning includes: a working condition acquisition unit, used to acquire ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure, and the current operating frequency of the compressor during the operation of the dehumidifier, and to assemble a working condition feature sequence in chronological order after preprocessing; a working condition identification unit, used to input the working condition feature sequence into a working condition identification model to obtain the current working condition type, and to calculate the intensity of the wet load mutation based on the rate of change of ambient humidity, and to calculate the degree of thermal response hysteresis based on the response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature; and an optimization decision unit, used to adjust the exploration and utilization ratio and optimization step size of the reinforcement learning algorithm according to the intensity of the wet load mutation and the degree of thermal response hysteresis, and to optimize when the safety margin is sufficient. The optimal step size increases with the intensity of sudden changes in wet load. When entering the protection warning zone, the optimal step size is limited to not exceeding the upper limit of the safe step size. The parameter calculation unit takes the operating condition characteristic sequence, the current operating condition type, the intensity of sudden changes in wet load, and the degree of thermodynamic response hysteresis as inputs, and takes the dual reward function composed of the energy efficiency ratio reward term and the dehumidification deviation suppression term as the objective, and outputs the optimal operating parameters through a reinforcement learning algorithm. The execution feedback unit monitors the compressor exhaust temperature, system high pressure, and system low pressure in real time when executing the optimal operating parameters. When entering the protection warning zone, it performs safety correction, uses the corrected actual execution parameters as operating instructions and writes them into the experience pool. After the thermodynamic steady-state criterion is met, it collects the actual energy efficiency ratio and actual humidity data, calculates the delay reward, and updates the algorithm parameters.

[0015] The present invention has the following beneficial effects: (1) This reinforcement learning-based dynamic optimization control method for dehumidifier energy efficiency ratio collects operating data such as ambient temperature and humidity, and compressor exhaust temperature, and assembles them into a time-series of operating condition characteristics. This allows the control process to judge the equipment status by combining changes in operating conditions before and after. It calculates the intensity of sudden changes in wet load by the rate of change of ambient humidity, and calculates the degree of thermal response hysteresis by the response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature. This separates the judgment of changes in external wet load and internal thermal response of the equipment. Compared with the method of judging operating conditions by relying solely on humidity thresholds or single rate of change, it can detect situations where the indoor wet load has changed but the refrigeration cycle has not yet completed its response earlier, reducing control deviations caused by lag in judging operating conditions.

[0016] (2) The dynamic optimization control method for dehumidifier energy efficiency ratio based on reinforcement learning links the optimization step size with the intensity of sudden change in wet load, the degree of hysteresis in thermal response, and the safety status of the equipment. This allows the reinforcement learning algorithm to adopt different optimization rhythms at different operating stages. When the wet load changes rapidly and the equipment has not entered the protection warning range, the optimization step size can be increased with the intensity of sudden change in wet load, so that the compressor frequency and fan speed can approach the appropriate operating range more quickly. When the compressor exhaust temperature, system high pressure, or system low pressure enters the protection warning range, the optimization step size is limited to the upper limit of the safe step size to avoid excessive action that triggers the equipment protection. This can take into account both the dehumidification demand response and the equipment operating boundary, so that the optimization process no longer blindly expands the search range, but is coordinated with the dehumidifier's own bearing capacity.

[0017] (3) The dehumidifier energy efficiency ratio dynamic optimization control method based on reinforcement learning monitors the compressor exhaust temperature, system high pressure and system low pressure in real time during the execution of the optimal operating parameters. When entering the protection warning range, the operating parameters are safely corrected and the corrected actual execution parameters are written into the experience pool as operating instructions. This can avoid the problem of the action record not corresponding to the operating result. After the thermodynamic steady state criterion is met, the actual energy efficiency ratio and actual humidity data are collected to calculate the delay reward, which reduces the interference of temperature and pressure fluctuations in the transition phase on the algorithm update. This makes the content learned by the algorithm closer to the correspondence between parameter adjustment, energy consumption change and dehumidification effect in the actual operation of the dehumidifier.

[0018] (4) The dehumidifier energy efficiency ratio dynamic optimization control system based on reinforcement learning provides continuous operating data through the operating condition acquisition unit, the operating condition identification unit judges the current operating condition and calculates the intensity of wet load change and the degree of thermal response hysteresis, the optimization decision unit adjusts the exploration utilization ratio and optimization step size accordingly, the parameter calculation unit outputs operating parameters with energy efficiency ratio reward item and dehumidification deviation suppression item as targets, and the execution feedback unit is responsible for safety correction, actual execution parameter input into the pool and delayed reward update. The units are connected one after another, so that the control command can be continuously corrected according to the equipment status change, thereby improving the energy efficiency ratio optimization effect of the dehumidifier under dynamic wet load.

[0019] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0020] Figure 1 This is a flowchart of a dynamic optimization control method for the energy efficiency ratio of a dehumidifier based on reinforcement learning, according to the present invention.

[0021] Figure 2 This is a block diagram of a dehumidifier energy efficiency ratio dynamic optimization control system based on reinforcement learning according to the present invention. Detailed Implementation

[0022] Please see Figure 1 This invention provides a technical solution: a dynamic optimization control method for the energy efficiency ratio of a dehumidifier based on reinforcement learning, comprising the following steps: collecting ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure, and the current operating frequency of the compressor during the operation of the dehumidifier; preprocessing the data to form a working condition feature sequence in chronological order; inputting the working condition feature sequence into a working condition identification model to obtain the current working condition type; calculating the intensity of the wet load mutation based on the rate of change of ambient humidity; calculating the degree of thermal response hysteresis based on the response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature; adjusting the exploration and utilization ratio and the optimization step size of the reinforcement learning algorithm according to the intensity of the wet load mutation and the degree of thermal response hysteresis; increasing the optimization step size with the intensity of the wet load mutation when the safety margin is sufficient; and limiting the optimization step size when entering the protection warning interval. If the safe step size limit is exceeded; the optimal operating parameters are output through a reinforcement learning algorithm, taking the operating condition feature sequence, current operating condition type, wet load mutation intensity, and thermal response hysteresis degree as inputs, and the dual reward function consisting of the energy efficiency ratio reward term and the dehumidification deviation suppression term as the objective; when executing the optimal operating parameters, the compressor exhaust temperature, system high pressure, and system low pressure are monitored in real time, and safety corrections are performed when entering the protection warning interval, and the corrected actual execution parameters are used as operating instructions and written into the experience pool; after the thermal steady-state criterion is met, the actual energy efficiency ratio and actual humidity data are collected, the delay reward is calculated, and the algorithm parameters are updated. The thermal steady-state criterion is that the absolute values ​​of the compressor exhaust temperature change rate, the absolute values ​​of the system high pressure change rate, and the absolute values ​​of the system low pressure change rate are all less than their respective steady-state thresholds within a continuous preset number of sampling periods.

[0023] Among them, the reinforcement learning algorithm adopts the improved DDPG network as the main architecture, and the experience pool adopts the priority experience replay mechanism to prioritize the storage of sample data related to wet load change conditions and safety corrections. The training process of the operating condition recognition model is as follows: historical operating data of the dehumidifier under three operating conditions: high humidity and high load, low humidity and low load, and sudden change in humidity load are collected. After preprocessing, the operating condition feature sequence for training is formed. The training set, validation set and test set are divided. The hybrid structure of convolutional neural network and long short-term memory network is trained. The network hyperparameters are adjusted through the validation set until the operating condition recognition accuracy of the model on the validation set reaches the preset threshold. Then the training stops. The operating condition recognition model is verified through the test set. The protection warning range is set based on the dehumidifier equipment parameters and industry safety standards; Sufficient safety margin means that the compressor discharge temperature, system high pressure, and system low pressure have not entered the corresponding protection warning range, and the difference between them and the corresponding protection threshold is greater than the preset safety margin threshold. In one embodiment, for example, the compressor exhaust temperature protection warning range can be set to 80℃~90℃ (equipment protection threshold is 95℃), the system high pressure protection warning range can be set to 1.8MPa~2.2MPa (equipment protection threshold is 2.5MPa), and the system low pressure drop warning range can be set to 0.2MPa~0.3MPa. The high pressure protection warning range corresponds to the upward trend of the system high pressure, and the low pressure drop warning range corresponds to the downward trend of the system low pressure. The equipment low pressure protection threshold is 0.15MPa. The preset statistical time window, preset sampling period, preset quantity sampling period, preset safety step size upper limit, and steady-state threshold are all set according to the dehumidifier model and actual operating scenario; In one implementation, for example, the preset statistical time window can be 5 to 10 minutes, the preset sampling period can be 10 to 30 seconds, the preset number of sampling periods can be 3 to 5, the preset safety step size upper limit can be 5 Hz / min (compressor frequency) or 50 r / min (fan speed), the steady-state threshold for the compressor exhaust temperature change rate can be 0.5℃ / min, the steady-state threshold for the system high pressure change rate can be 0.05 MPa / min, and the steady-state threshold for the system low pressure change rate can be 0.03 MPa / min.

[0024] Specifically, the steps for assembling the operating condition feature sequence in chronological order after preprocessing are as follows: The collected ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure, and compressor current operating frequency are subjected to sliding filtering and normalization processing, specifically as follows: A sliding filter is performed using a sliding window with a length of 3 to 5 sampling periods. The mean of the collected single-class data is calculated for each window, and the original data in the window is replaced with the mean of the window. This removes abnormal data that exceeds the allowable error range of the sensor or deviates significantly from the adjacent sampling trend. In one implementation, the sampling period is 20s, and the sliding window length is set to 3 sampling periods, that is, the average value of 3 sets of data within 60s is taken, and abnormal data such as instantaneous fluctuations of ambient humidity exceeding a preset humidity fluctuation threshold and instantaneous fluctuations of compressor exhaust temperature exceeding a preset temperature fluctuation threshold are removed. The normalization process uses the min-max normalization method to uniformly map the filtered data to the numerical range of [0,1]. The normalized results from multiple consecutive sampling times are arranged in chronological order to form a working condition characteristic sequence, which is as follows: Normalized data from 10 to 20 consecutive sampling periods are selected. The normalized data at each sampling time includes values ​​of six dimensions: ambient temperature, ambient humidity, compressor exhaust temperature, system high pressure, system low pressure, and compressor current operating frequency. These data are arranged in chronological order of sampling time to form a working condition feature sequence with a dimension of (number of sampling periods × 6). In one implementation, the sampling period can be 20s, and normalized data within 15 consecutive sampling periods, i.e. 300s, are selected to form a working condition feature sequence with a dimension of 15×6.

[0025] In this implementation plan, by filtering, normalizing, and arranging the operating data from multiple consecutive sampling times in chronological order, the dehumidifier's operating condition is no longer determined by instantaneous data at a single moment, but rather by combining the changing trends of ambient humidity, exhaust temperature, system pressure, and compressor frequency over a period of time.

[0026] Specifically, the steps for calculating the intensity of the wet load mutation based on the rate of change of ambient humidity, and for calculating the degree of thermal response hysteresis based on the normalized response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature, are as follows: By analyzing the trend of environmental humidity changes over multiple consecutive sampling times, the rate of change of environmental humidity is obtained, which is as follows: Select environmental humidity data for 5 to 8 consecutive sampling periods, and use linear fitting or difference method to calculate the rate of change of environmental humidity, i.e., the rate of change of environmental humidity, in units of %RH / min. The calculation method of difference method is to divide the difference of environmental humidity between two adjacent sampling times by the sampling period, and then take the average of multiple consecutive differences as the current rate of change of environmental humidity. In one implementation, for example, the sampling period can be 20s (i.e. 1 / 3min), and the ambient humidity at 6 consecutive sampling times is 60%RH, 62%RH, 65%RH, 67%RH, 69%RH, and 72%RH, respectively. The adjacent differences are 2%RH, 3%RH, 2%RH, 2%RH, and 3%RH, respectively. The change rate corresponding to each difference is 6%RH / min, 9%RH / min, 6%RH / min, 6%RH / min, and 9%RH / min. The average of these 5 change rates, 7.2%RH / min, is taken as the current ambient humidity change rate. The intensity of sudden change in wet load is determined based on the rate of change in ambient humidity. The greater the rate of change in ambient humidity, the greater the intensity of sudden change in wet load. Specifically: The rate of change in ambient humidity is divided into multiple levels, corresponding to different intensities of sudden changes in moisture load. A linear mapping method is used to convert the rate of change in ambient humidity into a moisture load sudden change intensity value between 0 and 1. The mapping formula is as follows: ; in, The intensity of sudden change in wet load, The current rate of change in ambient humidity. This is the minimum threshold for the rate of change of ambient humidity (i.e., the rate of change without abrupt changes, usually set to 0.5%RH / min). This is the maximum threshold for the rate of change of ambient humidity (i.e., the rate of change during extreme abrupt changes, usually set to 10%RH / min). In one implementation, for example, the current rate of change of ambient humidity can be 7.2%RH / min. Substituting this into the formula, the intensity of the sudden change in wet load is calculated to be approximately 0.71, which is considered a high intensity of sudden change. The compressor exhaust temperature variation trend was analyzed at multiple consecutive sampling times to obtain the compressor exhaust temperature variation rate, which is as follows: Using the same method as for calculating the rate of change of ambient humidity, compressor exhaust temperature data for 5 to 8 consecutive sampling periods were selected. The temperature difference between adjacent sampling times was calculated by the difference method, and the difference was divided by the sampling period to obtain the rate of change of temperature for a single time period. The average of the rate of change for multiple consecutive time periods was then taken as the current rate of change of compressor exhaust temperature, with the unit being ℃ / min. In one implementation, for example, the sampling period can be 20s, and the compressor exhaust temperature at 6 consecutive sampling times is 65℃, 66℃, 68℃, 69℃, 70℃, and 72℃, respectively, with adjacent differences of 1℃, 2℃, 1℃, 1℃, and 2℃, respectively, and the corresponding change rates are 3℃ / min, 6℃ / min, 3℃ / min, 3℃ / min, and 6℃ / min, respectively. The average value of 4.2℃ / min is taken as the current compressor exhaust temperature change rate. The rates of change in ambient humidity and compressor exhaust temperature were normalized to obtain the normalized rates of change in wet load and equipment thermal response, respectively. Using the same min-max normalization method as the aforementioned preprocessing steps, the rate of change of ambient humidity is normalized to the [0,1] interval to obtain the normalized rate of change of wet load. Similarly, the rate of change of compressor exhaust temperature is normalized to the [0,1] interval to obtain the normalized rate of change of equipment thermal response. The degree of thermal response hysteresis is calculated based on the difference between the normalized rate of change of wet load and the normalized rate of change of equipment thermal response. The larger the difference, the greater the degree of thermal response hysteresis. Specifically: The thermal response hysteresis is calculated as follows: when the normalized rate of change of wet load is greater than the normalized rate of change of equipment thermal response, the difference between the two is taken as the degree of thermal response hysteresis. When the normalized rate of change of wet load is less than or equal to the normalized rate of change of equipment thermal response, the hysteresis degree of thermal response is set to zero. This means that the hysteresis judgment mainly corresponds to the situation where the equipment thermal response lags behind the change of wet load. The value of the hysteresis degree ranges from 0 to 1. The closer the value is to 1, the greater the difference between the thermal response on the equipment side and the change on the load side, and the more severe the hysteresis. The calculation formula is as follows: ; Where H represents the degree of thermal hysteresis. The normalized rate of change of wet load, The normalized rate of change of the equipment's thermal response; In one implementation, =0.71, =0.51, and the calculated thermal response hysteresis is 0.20, which is a moderate hysteresis. like =0.4, =0.6, then the degree of hysteresis in the thermal response is taken as 0, that is, it is determined that there is no hysteresis.

[0027] In this implementation scheme, the external humidity load change and the internal thermal response state of the equipment are characterized separately by calculating the intensity of the humidity load mutation and the degree of thermal response hysteresis. The intensity of the humidity load mutation reflects the rate of change in indoor humidity demand, while the degree of thermal response hysteresis reflects whether equipment-side parameters such as compressor exhaust temperature follow the load change in a timely manner.

[0028] Specifically, the steps for adjusting the exploration and utilization ratio and optimization step size of the reinforcement learning algorithm based on the intensity of the wet load mutation and the degree of thermal response hysteresis are as follows: The exploration ratio of the reinforcement learning algorithm is determined based on the intensity of the wet load mutation; the greater the intensity of the wet load mutation, the greater the exploration ratio. Specifically: The exploration ratio ranges from 0.1 to 0.8. A linear mapping method is used to convert the wet load mutation intensity (0-1) into the corresponding exploration ratio. The mapping formula is as follows: ; in, To explore proportions, The intensity of sudden change in wet load; In one implementation, for example, the wet load mutation intensity can be 0.71. Substituting this into the formula, the exploration ratio is approximately 0.60. The algorithm increases the amplitude of continuous action disturbance or increases the number of candidate actions generated based on this exploration ratio to improve the search strength for new operating parameters. The utilization ratio of the reinforcement learning algorithm is determined based on the degree of thermal response hysteresis. The greater the thermal response hysteresis, the smaller the utilization ratio. When the utilization ratio is reduced based on the degree of thermal response hysteresis, the reduced portion is added equally to the exploration ratio, or the exploration ratio and utilization ratio are normalized so that the sum of the exploration ratio and the utilization ratio is kept at 1. Specifically: The sum of the proportion and the exploration proportion is 1, that is, the proportion is used. At the same time, the degree of thermal response hysteresis is combined for correction. When the degree of thermal response hysteresis H≥0.5 (severe hysteresis), the reduced utilization ratio is simultaneously added to the exploration ratio, or the exploration ratio and utilization ratio are normalized so that the sum of the two is kept at 1, so as to avoid the failure of existing optimal parameters due to equipment hysteresis. In one implementation, for example, the exploration ratio is 0.60 and the original utilization ratio is 0.40. If the thermal response hysteresis is 0.6, the utilization ratio is reduced by 0.1 to 0.30, while the exploration ratio is increased by 0.1 to 0.70, so that the sum of the two is still 1. The hysteresis condition is adapted by increasing the amplitude of continuous action disturbance or the number of candidate actions. When the compressor discharge temperature, system high pressure, and system low pressure have not entered the protection warning range, the optimization step size increases with the intensity of the wet load change, specifically as follows: The optimization step size is divided into compressor frequency optimization step size and fan speed optimization step size. Both increase linearly with the intensity of the wet load mutation. The value range of the compressor frequency optimization step size is 1~5Hz / min, and the value range of the fan speed optimization step size is 10~50r / min. The mapping formulas are as follows: ( (Optimization step size for compressor frequency). ( (Optimization step size for fan speed). In one implementation, for example, the wet load mutation intensity can be 0.71, and the calculated compressor frequency optimization step size is 3.84 Hz / min (rounded to 4 Hz / min), and the fan speed optimization step size is 38.4 r / min (rounded to 40 r / min), to ensure that the parameter adjustment speed can match the load change under the sudden change condition; When any one of the compressor discharge temperature, system high pressure, or system low pressure enters the protection warning range, the optimization step size is limited to within the preset safe step size upper limit, specifically: The preset safety step size upper limit is set according to the equipment safety requirements; In one implementation, for example, the upper limit of the safe step size for the compressor frequency can be 2-3 Hz / min, and the upper limit of the safe step size for the fan speed can be 20-30 r / min. When the compressor exhaust temperature reaches 85℃ (entering the protection warning range), no matter how large the intensity of the current wet load change, the compressor frequency optimization step size will not exceed the preset safe step size upper limit, and the fan speed optimization step size will not exceed the preset safe step size upper limit, so as to avoid the equipment triggering protection shutdown due to excessive parameter adjustment.

[0029] In this implementation plan, the intensity of sudden changes in wet load, the degree of thermal response hysteresis, and the protection warning status jointly participate in the optimization rhythm adjustment, so that the exploration ratio, utilization ratio, and optimization step size are adjusted according to the operating conditions. When the wet load changes significantly and the equipment safety margin is sufficient, the exploration ratio and optimization step size are appropriately increased, which is conducive to quickly finding the compressor frequency and fan speed suitable for the current load. When the equipment is close to the protection warning range, the optimization step size is limited to prevent excessive control action from causing the exhaust temperature or system pressure to continue to rise.

[0030] Specifically, the steps for constructing the dual reward function, which consists of an energy efficiency ratio reward term and a dehumidification deviation suppression term, are as follows: The actual energy efficiency ratio and actual humidity data of the dehumidifier under the current operating parameters are collected. The actual energy efficiency ratio is the ratio of dehumidification capacity to input electrical energy within a preset statistical time window, specifically: The preset statistical time window is set to 5-10 minutes to reduce the impact of transient data on the calculation results. The dehumidification capacity can be directly collected by a condensate water sensor or calculated by the difference in humidity between the inlet and outlet air and the air volume. The input electrical energy is collected in real time and integrated by the power acquisition module. In this method, the energy efficiency ratio is the dehumidification capacity corresponding to the unit input electrical energy of the dehumidifier, and its calculation formula is as follows: ; in, This is the actual energy efficiency ratio. W represents the dehumidification capacity within the preset statistical time window (unit: kg), and W represents the input electrical energy within the preset statistical time window (unit: kWh). In one implementation, for example, the preset statistical time window can be 5 minutes, the collected dehumidification amount can be 0.3 kg, the input electrical energy can be 0.1 kWh, and the calculated actual energy efficiency ratio is 3.0; The energy efficiency ratio (EER) bonus is determined based on the deviation between the actual EER and the preset EER benchmark. When the actual EER is higher than the preset EER benchmark, the bonus is positive; when the actual EER is lower than the preset EER benchmark, the bonus is reduced. Specifically: The preset energy efficiency ratio benchmark value is set based on the dehumidifier's rated parameters and actual operating experience, typically between 2.5 and 3.5. The energy efficiency ratio bonus uses a linear bonus method, calculated using the following formula: ; in, For energy efficiency ratio bonus items, The energy efficiency ratio bonus weight (with a value range of 0.4 to 0.6) This is the preset energy efficiency ratio benchmark value; In one implementation, for example The actual energy efficiency ratio is 3.2, and the calculated energy efficiency ratio bonus is 0.53, which is a positive bonus. If the actual energy efficiency ratio is 2.8, the calculated energy efficiency ratio bonus is 0.47, and the bonus is reduced. The dehumidification deviation suppression term is determined based on the deviation between the actual humidity and the preset target humidity. The greater the deviation between the actual humidity and the preset target humidity, the smaller the dehumidification deviation suppression term (i.e., the greater the suppression effect). Specifically: The preset target humidity is set according to user needs, typically between 40%RH and 60%RH. The dehumidification deviation suppression term uses a linear evaluation method, and the calculation formula is as follows: ; in, This is a dehumidification deviation suppression term. The dehumidification deviation suppression weight (with a value range of 0.4 to 0.6) This represents the actual humidity. To preset the target humidity, , These represent the maximum and minimum values ​​of ambient humidity (30%RH and 90%RH), respectively. In one implementation, for example The actual humidity is 55%RH, the deviation is 5%RH, and the calculated dehumidification deviation suppression term is 0.46. The larger the deviation, the smaller the suppression term (i.e., the greater the suppression effect). The energy efficiency ratio bonus and the dehumidification deviation suppression term are combined into the output value of the dual bonus function, which is as follows: The output of the dual reward function is the sum of the energy efficiency ratio reward term and the dehumidification deviation suppression term, allowing the algorithm to simultaneously consider energy efficiency optimization and dehumidification effect during the optimization process. The calculation formula is as follows: ; Where R is the output value of the dual reward function. After being limited, the output value of the dual reward function is mapped to the range of 0 to 1. The larger the output value, the higher the comprehensive evaluation value of the current operating parameters. In one implementation, for example, the energy efficiency ratio bonus is 0.53 and the dehumidification deviation suppression is 0.46, and the calculated output value of the dual reward function is 0.99, which remains 0.99 after limiting.

[0031] In this implementation scheme, by incorporating the energy efficiency ratio reward term and the dehumidification deviation suppression term into the reward function, the reinforcement learning algorithm considers both energy-saving and dehumidification effects during the optimization process. This method incorporates the relationship between dehumidification amount and input electrical energy, as well as the deviation between actual humidity and target humidity, into the evaluation to constrain the operating parameters to achieve a balance between energy consumption and dehumidification effect.

[0032] Specifically, the steps for outputting the optimal operating parameters using a reinforcement learning algorithm are as follows: The operating condition feature sequence, current operating condition type, intensity of wet load abrupt change, and degree of thermal response hysteresis are input into the decision network of the reinforcement learning algorithm, specifically as follows: The decision network adopts a fully connected neural network structure. The number of nodes in the input layer is consistent with the total dimension of the working condition feature sequence, the current working condition type (using one-hot encoding, 3 working conditions correspond to 3 nodes), the intensity of wet load change (1 node), and the degree of thermal response hysteresis (1 node). In one implementation, for example, the dimension of the operating condition feature sequence can be 15×6=90, plus 3 operating condition type nodes, 1 wet load change intensity node, and 1 thermal response hysteresis degree node, the total number of nodes in the input layer is 95, the hidden layer is set to 2 to 3 layers, the number of nodes in each layer is 64 to 128, and the number of nodes in the output layer is 2 (corresponding to compressor frequency and fan speed respectively). After the above input data is normalized, it is input into the input layer of the decision network, and then processed by the activation function (using the ReLU function) of the hidden layer in sequence before being passed to the output layer. The decision network determines the adjustment range of the compressor operating frequency and the fan speed based on the current operating condition type and the degree of thermal response hysteresis. The greater the thermal response hysteresis, the wider the search range of candidate operating parameters is, and the lower the upper limit of the rate of change for a single action is. Specifically: Different operating conditions correspond to different basic adjustment ranges. Under high humidity and high load conditions, the compressor frequency adjustment range is 40 to 60 Hz, and the fan speed adjustment range is 800 to 1200 r / min. Under low humidity and low load conditions, the compressor frequency adjustment range is 20 to 40 Hz, and the fan speed adjustment range is 400 to 800 r / min; Under conditions of sudden change in wet load, the basic adjustment range is consistent with that under conditions of high humidity and high load. The search range is expanded according to the degree of thermal response hysteresis. For every 0.1 increase in hysteresis, the compressor frequency adjustment range is expanded by 2-3 Hz (both upper and lower limits are expanded), the fan speed adjustment range is expanded by 50-100 r / min (both upper and lower limits are expanded), and the upper limit of the single action change rate is reduced by 0.5 Hz / min (compressor frequency) and 5 r / min (fan speed). In one implementation, for example, if the current operating condition is a sudden change in wet load and the thermal response hysteresis is 0.7 (severe hysteresis), then the compressor frequency adjustment range is expanded to 33-67Hz, the fan speed adjustment range is expanded to 650-1350r / min, and the upper limit of the single action change rate is reduced to 3.5Hz / min (compressor frequency) and 35r / min (fan speed). When any one of the compressor discharge temperature, system high pressure, or system low pressure enters the protection warning range, the search range of candidate operating parameters is narrowed and the upper limit of the rate of change of action is reduced. Specifically: When any parameter enters the protection warning range, the compressor frequency adjustment range is narrowed by 5 to 10 Hz, the fan speed adjustment range is narrowed by 100 to 200 r / min, and the upper limit of the single action change rate is reduced by 1 to 2 Hz / min (compressor frequency) and 10 to 20 r / min (fan speed). In one implementation, for example, if the compressor exhaust temperature enters the protection warning range (85°C), the expanded compressor frequency adjustment range is reduced to 38-62Hz, the fan speed adjustment range is reduced to 750-1250r / min, and the upper limit of the single action change rate is further reduced to 2.5Hz / min (compressor frequency) and 25r / min (fan speed). The decision network outputs continuous action parameters based on the current state, and generates candidate operating parameters by adding perturbations to the neighborhood of the continuous action parameters according to the exploration ratio. The evaluation network evaluates the value of the candidate operating parameters using a dual reward function as the optimization objective. Specifically: The decision network outputs continuous action parameters based on the current state, and adds disturbance noise or generates several candidate actions near the output action according to the exploration ratio. It evaluates the network input candidate operating parameters, current operating condition feature sequence, wet load mutation intensity and thermal response hysteresis. Based on the value function trained by the dual reward function, it outputs the value assessment value of the candidate operating parameter. The larger the value assessment value, the higher the expected reward corresponding to the candidate parameter. In one implementation, for example, the decision network generates 5 sets of candidate parameters, and the evaluation network calculates 5 value assessment values ​​according to the dual reward function, which are 0.85, 0.92, 0.88, 0.95 and 0.90 respectively; The candidate operating parameter with the highest value assessment result is selected as the optimal operating parameter output, specifically as follows: The value assessment values ​​output by the evaluation network are compared, and the candidate operating parameter with the highest value assessment value is taken as the optimal operating parameter under the current working condition and output to the actuator. In one implementation, for example, among the value assessment values ​​of 5 sets of candidate parameters, 0.95 is the highest value, and the corresponding candidate parameters (compressor frequency 55Hz, fan speed 1000r / min) are the optimal operating parameters, which are output to the drive modules of the compressor and fan.

[0033] In this implementation scheme, candidate operating parameters are generated by a decision network and evaluated by an evaluation network. This makes the output of compressor frequency and fan speed no longer a fixed combination of parameters or a simple threshold control result, but can be dynamically determined according to the current operating condition type, the intensity of wet load change, and the degree of thermal response hysteresis. For operating conditions with wet load change or equipment response hysteresis, the algorithm can expand the candidate search range to find more suitable control parameters, while avoiding parameter change by using the upper limit of the rate of change of action.

[0034] Specifically, the steps for writing the corrected actual execution parameters into the experience pool, collecting actual energy efficiency ratio and actual humidity data to calculate the delay reward and update the algorithm parameters after the thermodynamic steady-state criterion is met are as follows: The actual execution parameters after safety correction are combined with the corresponding operating condition characteristic sequence, wet load mutation intensity, and thermal response hysteresis degree, and then written into the experience pool as a state awaiting supplementary rewards. Specifically: The experience pool adopts a priority experience replay pool, and the stored data format is (operating condition characteristic sequence, wet load mutation intensity, thermal response hysteresis degree, actual execution parameters, delayed reward, operating condition characteristic sequence at the next moment). After the actual execution parameters are output, the sample is first written into the experience pool in the state of waiting for supplementary reward. Once the thermodynamic steady-state criterion is met, the calculated delayed reward will be added to the corresponding sample. The actual execution parameters are the compressor frequency and fan speed after safety correction. When writing, the priority is set according to the intensity of the wet load change and whether safety correction has been performed. The greater the intensity of the wet load change and the higher the priority of the sample that has been safety corrected. In one implementation, for example, the actual execution parameters after safety correction can be a compressor frequency of 50Hz and a fan speed of 950r / min. This parameter is combined with the corresponding operating condition characteristic sequence, wet load mutation intensity of 0.71, and thermal response hysteresis degree of 0.20, and the priority is set to 0.8 (priority range 0 to 1) so that the supplementary reward status can be written into the experience pool. Within each sampling period after the actual execution of the parameters, the compressor exhaust temperature change rate, the system high pressure change rate, and the system low pressure change rate are calculated, specifically as follows: The compressor exhaust temperature, system high pressure, and system low pressure data are collected once in each sampling period (10-30s). The rate of change of each parameter is calculated using the differential method, which is to subtract the parameter value at the previous sampling time from the parameter value at the current sampling time, and then divide by the sampling period to obtain the rate of change of the parameter within the sampling period. The rate of change of the three parameters is calculated and recorded respectively. In one implementation, for example, the sampling period can be 20s. If the compressor exhaust temperature was 70℃ at the previous sampling time and 71℃ at the current sampling time, then the compressor exhaust temperature change rate is (71-70) / (20 / 60) = 3℃ / min. Similarly, the system high pressure change rate is calculated to be 0.04MPa / min, and the system low pressure change rate is 0.02MPa / min. When the absolute values ​​of the compressor exhaust temperature change rate, the absolute values ​​of the system high-pressure change rate, and the absolute values ​​of the system low-pressure change rate are all less than their respective steady-state thresholds within a consecutive preset number of sampling periods, the thermodynamic steady-state criterion is determined to be satisfied. Specifically: The preset sampling period is 3 to 5 times, and the steady-state threshold is set according to the characteristics of the device. In one implementation, for example, the preset sampling period can be 3, and the steady-state thresholds can be 0.5℃ / min, 0.05MPa / min, and 0.03MPa / min, respectively. Within 3 consecutive sampling periods, the absolute values ​​of the rate of change of the three parameters are 0.4℃ / min, 0.04MPa / min, and 0.02MPa / min, respectively, all of which are less than the corresponding steady-state thresholds, and the thermodynamic steady-state criterion is satisfied. After satisfying the thermodynamic steady-state criterion, actual energy efficiency ratio and actual humidity data are collected, the delayed reward value is calculated, the delayed reward value is added to the corresponding sample in the experience pool, and the network parameters of the reinforcement learning algorithm are updated according to the delayed reward value. Specifically: After satisfying the thermodynamic steady-state criterion, actual energy efficiency ratio and actual humidity data are collected according to the method of claim 5, and the output value of the dual reward function is calculated. This output value is the delayed reward value. The delayed reward value is added to the corresponding sample in the experience pool. Then, a certain number of samples (usually 32 to 64) are drawn from the experience pool according to the sample priority to update the parameters of the decision network and the evaluation network. The gradient descent method is used to minimize the prediction error of the evaluation network and update the parameters of the decision network at the same time, so that the decision network can output better operating parameters. In one implementation, for example, after satisfying the thermodynamic steady-state criterion, the actual energy efficiency ratio collected is 3.3 and the actual humidity is 50%RH. The calculated delayed reward value is 1.0. This reward value is added to the corresponding sample in the experience pool, 32 samples are extracted to update the network parameters, and the algorithm iteration is completed.

[0035] In this implementation scheme, by writing the actual execution parameters after safety correction into the experience pool and calculating the delayed reward after the thermodynamic steady-state criterion is met, the impact of unexecuted parameters and fluctuation data during the transition phase on algorithm updates can be reduced. Since the dehumidification mechanism has a thermodynamic response process in the cooling cycle, the energy efficiency ratio and humidity changes after parameter adjustment need to go through a certain transition process before stabilizing. Therefore, delaying the reward calculation until the thermodynamic steady state allows the reward value to more realistically reflect the actual effect of the current operating parameters. The priority experience playback mechanism can improve the utilization rate of wet load mutations and safety correction samples.

[0036] Specifically, the steps for inputting the operating condition feature sequence into the operating condition identification model to obtain the current operating condition type are as follows: A hybrid structure of convolutional neural networks and long short-term memory networks is used to construct a working condition identification model, specifically as follows: The working condition recognition model consists of convolutional layers, pooling layers, LSTM layers, and fully connected layers. There are 2 to 3 convolutional layers used to extract static correlation features. The kernel size of each convolutional layer is 3×3, the stride is 1, and the padding method is same. The pooling layer uses max pooling with a kernel size of 2×2 and a stride of 2. The LSTM layer consists of 1-2 layers with 64-128 hidden units, used to extract temporal variation features; the fully connected layer consists of 1-2 layers, with the output layer using the softmax activation function to output probability values ​​for three operating conditions (high humidity and high load, low humidity and low load, and sudden change in wet load); the model uses the cross-entropy loss function, the Adam optimizer, and the learning rate is set to 0.001-0.005; In one implementation, the specific parameters can be adjusted according to the dehumidifier model, compressor specifications, refrigerant type, and application scenario; The static correlation features at each sampling time point in the working condition feature sequence are extracted using a convolutional neural network, specifically as follows: The operating condition feature sequence is converted into a two-dimensional matrix format (number of sampling periods × feature dimension), input into a convolutional layer, and the convolutional layer performs convolution operation on the matrix through convolution kernels to extract the correlation between the six-dimensional features at each sampling time (such as the correlation between ambient humidity and compressor exhaust temperature, and the correlation between system high and low pressure and compressor frequency). After dimensionality reduction processing by pooling layer, key static correlation features are retained and data redundancy is reduced. In one implementation, for example, the working condition feature sequence can be a 15×6 matrix, which, after being processed by two convolutional layers and one pooling layer, outputs a static correlation feature matrix with a dimension of 4×3. The temporal variation features between multiple consecutive sampling times in the working condition feature sequence are extracted using a Long Short-Term Memory (LSTM) network. Specifically: The static correlation features extracted by the convolutional neural network are input into the LSTM layer. The LSTM layer captures the feature change trend between consecutive sampling times through the synergistic effect of the forget gate, input gate, and output gate, such as the continuous rise / fall trend of ambient humidity and the hysteresis response trend of compressor exhaust temperature, which can effectively explore the dynamic change law of the working condition. In one implementation, for example, a 4×3 static correlation feature matrix is ​​input into an LSTM layer (with 64 hidden units), and the output is a temporal variation feature vector with a dimension of 64. By integrating static correlation features and temporal variation features, the current operating condition type is output. The current operating condition type includes high humidity and high load, low humidity and low load, and sudden change in humidity load, which are specifically as follows: The static correlation feature matrix is ​​flattened and concatenated with the temporal variation feature vector to obtain the fused feature vector. The fused feature vector is input into the fully connected layer, and after being processed by the activation function, the probability values ​​of three working conditions are output. The working condition with the highest probability value is the current working condition. In one implementation, for example, the fused feature vector dimension is 4×3+64=76. After inputting into the fully connected layer, the probability of high humidity and high load is 0.05, the probability of low humidity and low load is 0.15, and the probability of wet load change is 0.80. Therefore, the current operating condition is determined to be a wet load change condition. In one embodiment, the criteria for determining high humidity and high load conditions can be ambient humidity ≥70%RH and compressor frequency ≥50Hz, the criteria for determining low humidity and low load conditions can be ambient humidity ≤50%RH and compressor frequency ≤30Hz, and the criteria for determining wet load change conditions can be wet load change intensity ≥0.5.

[0037] In this implementation scheme, a convolutional neural network is used to extract the correlation between ambient humidity, exhaust temperature, system pressure and compressor frequency within the same sampling time, and a long short-term memory network is used to extract the changing trend between consecutive sampling times.

[0038] Specifically, the steps for safety correction when entering a protection warning zone are as follows: During the execution of optimal operating parameters, the compressor discharge temperature, system high pressure, and system low pressure are collected in real time, specifically: The corresponding sensors are used to collect three parameters in real time. The collection frequency is consistent with the sampling period (10-30s / time). The collected data is transmitted to the control module in real time. The control module monitors the data in real time and determines whether the protection warning zone has been entered. The measurement accuracy of the sensors meets the equipment control requirements. The compressor exhaust temperature sensor has an accuracy of ±0.5℃, and the system high and low pressure sensors have an accuracy of ±0.01MPa. In one embodiment, for example, the sampling period can be 20s, and the real-time collected compressor exhaust temperature is 86℃, the system high pressure is 2.0MPa, and the system low pressure is 0.25MPa. When any one of the compressor discharge temperature, system high pressure, or system low pressure enters the protection warning range, the adjustment range of the compressor operating frequency and fan speed is limited, specifically: preset adjustment range limit value; In one implementation, for example, the single adjustment range of the compressor frequency can not exceed 1-2 Hz, and the single adjustment range of the fan speed can not exceed 10-20 r / min. At the same time, targeted restrictions are made according to the type of parameter entering the warning range. If the compressor exhaust temperature enters the warning range, the increase in compressor frequency is restricted first. If the system high pressure enters the warning range, the compressor frequency adjustment range is reduced and the fan speed adjustment range is appropriately increased. If the system low pressure enters the warning range, the compressor frequency increase or decrease is restricted first, and the fan speed is maintained or appropriately increased to improve the heat exchange state on the evaporator side. In one implementation, for example, when the compressor discharge temperature reaches 86°C and enters the warning range, the optimal operating parameters are a compressor frequency of 55Hz and a fan speed of 1000r / min. At this time, the single adjustment range of the compressor frequency is limited to no more than 1Hz, and the single adjustment range of the fan speed is limited to no more than 10r / min. The operating parameters after limiting the adjustment range are output as the actual execution parameters after safety correction, specifically: Based on the adjustment range limit, the optimal operating parameters are corrected. The corrected parameters must ensure that the relevant safety parameters gradually move away from the protection warning range, while getting as close as possible to the optimal operating parameters. In one implementation, for example, the optimal compressor frequency is 55Hz, and the single adjustment range is limited to no more than 1Hz, which is corrected to 54Hz. The optimal fan speed is 1000r / min, and the single adjustment range is limited to no more than 10r / min, which is corrected to 1010r / min (appropriately increasing the fan speed helps to reduce the exhaust temperature). 54Hz and 1010r / min are used as the actual execution parameters after safety correction and output to the actuator.

[0039] In this implementation scheme, the compressor exhaust temperature, system high pressure and system low pressure are monitored in real time during the execution of optimal operating parameters, and the compressor frequency and fan speed are limited and corrected when entering the protection warning range, so that the parameters output by the reinforcement learning algorithm are constrained by the equipment safety boundary before execution.

[0040] Please see Figure 2 This invention provides a technical solution: a dynamic optimization control system for the energy efficiency ratio of a dehumidifier based on reinforcement learning, comprising: a working condition acquisition unit, used to acquire ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure, and the current operating frequency of the compressor during the operation of the dehumidifier, and to assemble a working condition feature sequence in chronological order after preprocessing; a working condition identification unit, used to input the working condition feature sequence into a working condition identification model to obtain the current working condition type, and to calculate the intensity of the wet load mutation based on the rate of change of ambient humidity, and to calculate the degree of thermal response hysteresis based on the response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature; and an optimization decision unit, used to adjust the exploration and utilization ratio and the optimization step size of the reinforcement learning algorithm according to the intensity of the wet load mutation and the degree of thermal response hysteresis. When the safety margin is sufficient, the optimization step size increases with the intensity of the wet load mutation, and optimization is restricted when the dehumidifier enters the protection warning range. The step size does not exceed the upper limit of the safe step size; the parameter calculation unit is used to take the operating condition feature sequence, the current operating condition type, the intensity of the wet load change and the degree of thermal response hysteresis as input, and the dual reward function composed of the energy efficiency ratio reward term and the dehumidification deviation suppression term as the objective, and output the optimal operating parameters through the reinforcement learning algorithm; the execution feedback unit is used to monitor the compressor exhaust temperature, system high pressure and system low pressure in real time when executing the optimal operating parameters, and to perform safety correction when entering the protection warning interval. The corrected actual execution parameters are used as the operating instructions and written into the experience pool. After the thermal steady-state criterion is met, the actual energy efficiency ratio and actual humidity data are collected, the delay reward is calculated and the algorithm parameters are updated. The thermal steady-state criterion is that the absolute values ​​of the compressor exhaust temperature change rate, the absolute values ​​of the system high pressure change rate, and the absolute values ​​of the system low pressure change rate are all less than their respective steady-state thresholds within a continuous preset number of sampling periods.

[0041] The operating condition acquisition unit includes multiple sensors and a preprocessing module. The sensors include an ambient temperature and humidity sensor, a compressor exhaust temperature sensor, a system high-pressure sensor, and a system low-pressure sensor. The compressor frequency detection module is integrated into the compressor drive module and is used to obtain the current operating frequency of the compressor. The sensors are installed at the corresponding positions on the dehumidifier. The ambient temperature and humidity sensor is installed at the air inlet of the dehumidifier, the compressor exhaust temperature sensor is installed on the compressor exhaust pipe, the system high-pressure sensor is installed on the condenser inlet pipe, and the system low-pressure sensor is installed on the evaporator outlet pipe. The current operating frequency of the compressor is obtained by feedback from the compressor drive module. The preprocessing module is electrically connected to each sensor to receive the data collected by the sensor, perform sliding filtering and normalization processing, and then assemble the working condition feature sequence in time order. The operating condition recognition unit includes a model storage module and an operating condition analysis module. The model storage module is used to store the trained CNN-LSTM hybrid structure operating condition recognition model, and the operating condition analysis module is used to call the model to process the operating condition feature sequence, and at the same time calculate the rate of change of ambient humidity and the rate of change of compressor exhaust temperature, thereby obtaining the intensity of wet load change and the degree of thermal response hysteresis. The optimization decision-making unit and parameter calculation unit are integrated in the same control module. They are implemented using industrial controllers, embedded processors, edge computing modules, FPGAs, or control chips with neural network inference capabilities. The built-in reinforcement learning algorithm program includes decision network, evaluation network, and experience pool storage module, which can quickly complete parameter adjustment, optimization calculation, and optimal parameter output. The execution feedback unit includes an execution module and a feedback module. The execution module is electrically connected to the drive module of the compressor and fan, and is used to receive the optimal operating parameters or the actual execution parameters after safety correction, and drive the compressor and fan to run. The feedback module is electrically connected to each sensor and execution module to monitor operating parameters in real time, determine whether the protection warning range has been entered, perform safety corrections, and collect steady-state energy efficiency ratio and humidity data. It also calculates the delay reward and feeds it back to the optimization decision unit and parameter calculation unit to complete the algorithm parameter update.

[0042] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0043] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A dynamic optimization control method for the energy efficiency ratio of a dehumidifier based on reinforcement learning, characterized in that, Includes the following steps: The ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure, and current compressor operating frequency are collected during the operation of the dehumidifier. After preprocessing, they are arranged into a time sequence of operating condition characteristics. Input the operating condition feature sequence into the operating condition identification model to obtain the current operating condition type, calculate the intensity of wet load change based on the rate of change of ambient humidity, and calculate the degree of thermal response hysteresis based on the response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature. The exploration and utilization ratio and optimization step size of the reinforcement learning algorithm are adjusted according to the intensity of wet load change and the degree of thermal response hysteresis. When the safety margin is sufficient, the optimization step size increases with the intensity of wet load change. When entering the protection warning range, the optimization step size is limited to not exceeding the upper limit of the safety step size. Taking the operating condition feature sequence, current operating condition type, wet load abrupt change intensity and thermal response hysteresis as input, and the dual reward function composed of energy efficiency ratio reward term and dehumidification deviation suppression term as target, the optimal operating parameters are output through reinforcement learning algorithm; When executing optimal operating parameters, the compressor discharge temperature, system high pressure and system low pressure are monitored in real time. When entering the protection warning range, safety correction is performed, and the corrected actual execution parameters are used as operating instructions and written into the experience pool. Once the thermodynamic steady-state criterion is met, the actual energy efficiency ratio and actual humidity data are collected, the delay reward is calculated, and the algorithm parameters are updated.

2. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps for assembling the operating condition characteristic sequence in chronological order after preprocessing are as follows: The collected ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure and compressor current operating frequency are subjected to sliding filtering and normalization. The normalized results from multiple consecutive sampling times are arranged in chronological order to form a working condition feature sequence.

3. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps for calculating the intensity of the wet load mutation and the degree of hysteresis in the thermal response are as follows: The environmental humidity change trend was analyzed at multiple consecutive sampling times to obtain the environmental humidity change rate; The intensity of sudden change in wet load is determined based on the rate of change in ambient humidity. The compressor exhaust temperature change trend was analyzed at multiple consecutive sampling times to obtain the compressor exhaust temperature change rate. The change rates of ambient humidity and compressor exhaust temperature were normalized to obtain the normalized change rate of wet load and the normalized change rate of equipment thermal response. The degree of thermal response hysteresis is calculated based on the difference between the normalized rate of change of wet load and the normalized rate of change of equipment thermal response.

4. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps for adjusting the exploration and utilization ratio and optimization step size of the reinforcement learning algorithm based on the intensity of the wet load mutation and the degree of thermal response hysteresis are as follows: The exploration ratio of the reinforcement learning algorithm is determined based on the intensity of the sudden change in wet load. The utilization ratio of reinforcement learning algorithms is determined based on the degree of thermal response hysteresis; When the compressor discharge temperature, system high pressure, and system low pressure have not entered the protection warning range, the optimization step size increases with the increase of the intensity of the wet load change. When any one of the compressor discharge temperature, system high pressure, or system low pressure enters the protection warning range, the optimization step size is limited to the preset safe step size upper limit.

5. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps for constructing the dual reward function, which consists of an energy efficiency ratio reward term and a dehumidification deviation suppression term, are as follows: Collect the actual energy efficiency ratio and actual humidity data of the dehumidifier under the current operating parameters; The energy efficiency ratio bonus is determined based on the deviation between the actual energy efficiency ratio and the preset energy efficiency ratio benchmark value; Determine the dehumidification deviation suppression term based on the deviation relationship between the actual humidity and the preset target humidity; The energy efficiency ratio bonus and the dehumidification deviation suppression term are combined into the output value of a dual reward function.

6. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps to output the optimal operating parameters using a reinforcement learning algorithm are as follows: The operating condition feature sequence, current operating condition type, intensity of wet load mutation, and degree of thermal response hysteresis are input into the decision network of the reinforcement learning algorithm; The decision network determines the adjustment range of the compressor operating frequency and the fan speed based on the current operating condition type and the degree of thermal response hysteresis. When any one of the compressor discharge temperature, system high pressure, or system low pressure enters the protection warning range, the search range of candidate operating parameters is narrowed and the upper limit of the rate of change of action is reduced. The decision network generates candidate operating parameters within the adjustment range, and the evaluation network evaluates the value of the candidate operating parameters based on historical experience samples and the value function corresponding to the dual reward function. The candidate operating parameter with the highest value assessment result is selected as the optimal operating parameter output.

7. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps for writing the corrected actual execution parameters into the experience pool, and then collecting actual energy efficiency ratio and actual humidity data to calculate the delay reward and update the algorithm parameters after the thermodynamic steady-state criterion is met are as follows: The actual execution parameters after safety correction are combined with the corresponding operating condition characteristic sequence, wet load mutation intensity, and thermal response hysteresis degree and written into the experience pool as a state to be supplemented with rewards. In each sampling period after the actual execution parameters are executed, the compressor exhaust temperature change rate, the system high pressure change rate, and the system low pressure change rate are calculated. When the absolute values ​​of the compressor exhaust temperature change rate, the absolute values ​​of the system high pressure change rate, and the absolute values ​​of the system low pressure change rate are all less than their respective steady-state thresholds within a consecutive preset number of sampling periods, the thermodynamic steady-state criterion is determined to be satisfied. After satisfying the thermodynamic steady-state criterion, actual energy efficiency ratio and actual humidity data are collected, the delayed reward value is calculated, the delayed reward value is added to the corresponding sample in the experience pool, and the network parameters of the reinforcement learning algorithm are updated according to the delayed reward value.

8. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps to input the working condition feature sequence into the working condition identification model to obtain the current working condition type are as follows: A working condition identification model is constructed using a hybrid structure of convolutional neural network and long short-term memory network. Static correlation features at each sampling time in the working condition feature sequence are extracted using a convolutional neural network; Temporal variation features between multiple consecutive sampling times in the working condition feature sequence are extracted using a long short-term memory network. By integrating static correlation features and temporal change features, the current operating condition type is output, which includes high humidity and high load conditions, low humidity and low load conditions, and sudden change in humidity load conditions.

9. The method for dynamic optimization control of dehumidifier energy efficiency ratio based on reinforcement learning according to claim 1, characterized in that, The steps for safety correction when entering a protection warning zone are as follows: During the execution of optimal operating parameters, the compressor discharge temperature, system high pressure, and system low pressure are collected in real time. When any one of the compressor discharge temperature, system high pressure, or system low pressure enters the protection warning range, the adjustment range of the compressor operating frequency and fan speed is limited. The operating parameters after limiting the adjustment range will be output as the actual execution parameters after safety correction.

10. A dehumidifier energy efficiency ratio dynamic optimization control system based on reinforcement learning, employing the dehumidifier energy efficiency ratio dynamic optimization control method based on reinforcement learning as described in any one of claims 1-9, characterized in that, include: The operating condition acquisition unit is used to collect ambient temperature and humidity, compressor exhaust temperature, system high pressure, system low pressure and compressor current operating frequency during the operation of the dehumidifier. After preprocessing, the data are arranged into an operating condition feature sequence in chronological order. The operating condition identification unit is used to input the operating condition feature sequence into the operating condition identification model to obtain the current operating condition type, calculate the intensity of wet load change based on the rate of change of ambient humidity, and calculate the degree of thermal response hysteresis based on the response difference between the rate of change of ambient humidity and the rate of change of compressor exhaust temperature. The optimization decision unit is used to adjust the exploration and utilization ratio and optimization step size of the reinforcement learning algorithm according to the intensity of wet load change and the degree of thermal response hysteresis. When the safety margin is sufficient, the optimization step size increases with the intensity of wet load change. When entering the protection warning range, the optimization step size is limited to not exceed the upper limit of the safety step size. The parameter calculation unit is used to take the operating condition feature sequence, current operating condition type, wet load mutation intensity and thermal response hysteresis as inputs, and the dual reward function composed of energy efficiency ratio reward term and dehumidification deviation suppression term as the target, and output the optimal operating parameters through reinforcement learning algorithm. The execution feedback unit is used to monitor the compressor exhaust temperature, system high pressure and system low pressure in real time when executing the optimal operating parameters. When entering the protection warning range, it performs safety correction, uses the corrected actual execution parameters as the operating instructions and writes them into the experience pool. After the thermodynamic steady-state criterion is met, it collects the actual energy efficiency ratio and actual humidity data, calculates the delay reward and updates the algorithm parameters.