Intelligent Control Method for Evaporative Cooling Air Conditioning System Based on Deep Reinforcement Learning
Through the intelligent control method based on deep reinforcement learning, the multi-objective control problem of evaporative refrigerated air conditioning system in complex environments is solved, adaptive adjustment and precise control are achieved, and the system's self-optimization ability and user experience are improved.
Patent Information
- Application Number
- CN202510657114.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-21
AI Technical Summary
When traditional evaporative air conditioning systems face complex coupling and dynamic disturbances, they are difficult to take into account multiple target requirements such as optimal energy consumption, humidity control and load fluctuations, and it is difficult for users to obtain the precise air conditioning system status through human senses or simple parameters, resulting in inconvenience in maintenance and use.
An intelligent control method based on deep reinforcement learning is adopted, by obtaining the spatial elements of the refrigeration area, dividing functional blocks, deploying sensor units, performing data preprocessing and feature extraction, an intelligent control policy network is built, and control instructions are output to adjust valve opening and fan frequency, and adaptive adjustment is achieved.
It realizes adaptive adjustment of the evaporative refrigerated air conditioning system, balances the conflicts between various regulatory targets, provides precise control instructions, and improves the system's self-optimization ability and user experience.
Smart Images

Figure CN120194407B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of refrigeration equipment regulation, and more specifically, relates to an intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning. Background Art
[0002] In an evaporative cooling air conditioning system, if the supply air temperature is simply controlled within a fixed range, various coupling complexities and dynamic disturbances will be faced during actual operation, making it difficult for traditional PID control to balance multiple objective requirements such as optimal energy consumption, humidity control, equipment wear, and load fluctuations.
[0003] Currently, when traditional air conditioning systems mainly control cooling, heating, ventilation, dehumidification, etc. modes through a remote control, only beeping sounds and simple symbols are used to indicate mode switching, and it is difficult to accurately obtain parameters such as the current operating voltage of the unit, compressor operation, indoor and outdoor motor operation, indoor and outdoor temperatures, supply air temperature, and power consumption. It can only be collected by human senses or with the help of other devices, resulting in users being difficult to effectively and scientifically maintain and use the air conditioning system. Such problems are particularly significant in the maintenance and use of evaporative cooling air conditioning systems.
[0004] First, in an evaporative cooling air conditioning system, in addition to the supply air temperature, it is also necessary to maintain operating condition variables such as indoor humidity, cooling water flow rate, and pressure difference before and after the tower. These variables are highly coupled, and the operation of each actuator (fan, water pump, valve) does not have a linear superposition effect on the entire evaporative cooling air conditioning system. For example, increasing the fan speed is beneficial to the decrease of the supply air temperature, but it is more likely to cause humidity fluctuations and energy consumption increase. Second, the evaporative cooling air conditioning system is more susceptible to environmental changes such as outdoor temperature, relative humidity, indoor human flow, and window opening. Traditional PID control technology requires frequent manual readjustment of gain parameters and is difficult to achieve self - adaptation. Finally, the heat transfer efficiency of the evaporative cooling tower, water film thickness, and fan performance curve all have strong nonlinearity and uncertainty, and it is difficult to calibrate in practice. Summary of the Invention
[0005] To solve the deficiencies in the prior art, the purpose of the present invention is to address the above - mentioned defects and further propose an intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning.
[0006] The present invention adopts the following technical solutions.
[0007] The first aspect of the present invention discloses an intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning, and the method includes:
[0008] Obtain the spatial elements of the refrigeration area and convert the spatial elements into a digital model to generate a functional block division strategy for the refrigeration area based on the digital model;
[0009] Based on the above functional block division strategy, obtain the sensor data on each functional block fusion node, and preprocess the sensor data to obtain a multi-dimensional feature data set;
[0010] Fuse the multi-dimensional feature data set into performance indicators to output the performance evaluation vectors of each functional block;
[0011] Based on the performance evaluation vectors, construct an experience replay pool to fuse the global state vectors, and perform offline training on the DDPG algorithm to obtain an intelligent control strategy network;
[0012] Input the real-time performance indicators of the refrigeration area into the intelligent control strategy network for processing, and output the control instructions for the evaporative cooling air conditioning system;
[0013] In response to the control instructions, control the valve opening and fan frequency of each functional block;
[0014] Among them, the spatial elements include building floor plans, cooling load distributions, and pipe network structures, and the digital model includes a digitalized spatial coordinate set, load mapping, and pipe network topology diagrams; the sensor data includes temperature, humidity, flow rate, and pressure difference.
[0015] Further, the method of obtaining the spatial elements of the refrigeration area, converting the spatial elements into a digital model, and generating the functional block division strategy for the refrigeration area based on the digital model includes:
[0016] Based on the digitalized spatial coordinate set, load mapping, and pipe network topology diagrams, divide the refrigeration area into multiple functional blocks;
[0017] Select key nodes from each functional block, and set sensor units at the key nodes to generate the functional block division strategy; the sensor units include temperature and humidity sensors, flow meters, and pressure difference gauges.
[0018] Further, the method of obtaining the sensor data on each functional block fusion node based on the functional block division strategy, and preprocessing the sensor data to obtain a multi-dimensional feature data set includes:
[0019] Receive the sensor data from each sensor unit, and construct an initial sequence matrix based on the sensor data according to the sampling period and time series;
[0020] Perform moving average filtering on each column of sensor data in the initial sequence matrix to obtain a smoothed sequence matrix;
[0021] Fill in the missing values in the smoothed sequence matrix, and normalize the filled smoothed sequence matrix through linear mapping to obtain the multi-dimensional feature dataset.
[0022] Further, fusing the multi-dimensional feature dataset into performance metrics to output the performance evaluation vectors of each functional block includes:
[0023] Align the timestamps of the time-series data of each functional block in the multi-dimensional feature dataset through the Network Time Protocol, and linearly interpolate and fill in the sensor data with a time error exceeding the first threshold;
[0024] Judge the health status of the sensor device through the heartbeat detection frequency and residual statistics, and assign health weights to each sensor data according to the degree of health, so as to fill in the sensor data with a health weight lower than the second threshold to obtain the health-weighted time-series data;
[0025] Perform moving window median filtering on the health-weighted time-series data, and perform weighted averaging on the sensor data at the same moment according to the health weight to obtain the fused time-series curve of each functional block;
[0026] Statistically calculate the performance metrics of the fused time-series curve according to the set moving window; the performance metrics include cooling efficiency, temperature and humidity deviation, and energy load;
[0027] Mark the performance metrics with a moving average abnormal jump exceeding the third threshold as abnormal metrics, and package the abnormal metrics and performance metrics to obtain the performance evaluation vector.
[0028] Further, constructing an experience replay pool based on the performance evaluation vector to fuse the global state vector to perform offline training on the DDPG algorithm to obtain an intelligent control strategy network, including:
[0029] Concatenate the performance evaluation vectors of each functional block in time series to form the global state vector of the refrigeration area;
[0030] When the experience replay pool reaches the capacity limit, sort each global state vector according to the TD error size of the global state vector, and retain the global state vectors with a TD error exceeding the fourth threshold to obtain the global state samples.
[0031] Further, constructing an experience replay pool based on the performance evaluation vector to fuse the global state vector to perform offline training on the DDPG algorithm to obtain an intelligent control strategy network, further including:
[0032] Determine whether the temperature difference or energy consumption of each functional block exceeds a preset high warning threshold; if so, mark the functional block with a first label; if not, mark the functional block with a second label;
[0033] When multiple adjacent functional blocks are marked with the first label, add a coupling weight to the multiple adjacent functional blocks, and splice the functional blocks carrying the coupling weight with other functional blocks to obtain a fused state vector;
[0034] Use the fused state vector for offline training with the DDPG algorithm based on the Actor-Critic framework to obtain the intelligent control strategy network.
[0035] Further, inputting the real-time performance index of the refrigeration area into the intelligent control strategy network for processing and outputting the control instruction of the evaporative cooling air conditioning system includes:
[0036] Obtain the real-time performance indexes of each functional block in the refrigeration area, and screen the real-time performance indexes according to the health labels of the real-time performance indexes, so as to eliminate and fill the sensor data with a health degree lower than the fifth threshold, and obtain the four-element performance values of each functional block;
[0037] Splice the four-element performance values of all functional blocks, and carry the global state vector and health weight of each functional block during splicing to output a real-time global state vector;
[0038] Call the intelligent control strategy network to process the real-time global state vector to generate the recommended valve opening and fan frequency of each functional block of the evaporative cooling air conditioning system, and output the control instruction executed according to the recommended valve opening and fan frequency.
[0039] Further, the method further includes:
[0040] When a sensor device of any functional block fails, determine that the functional block is a faulty block, switch the faulty block to a preset PID redundant loop, and send an alarm message to the operation and maintenance terminal at the same time;
[0041] Generate a global action vector based on the recommended valve opening and fan frequency, and send the global action vector to the control center, and the control center issues the control instruction to the PLC through the fieldbus.
[0042] A second aspect of the present invention discloses a terminal, including a processor and a storage medium; characterized in that:
[0043] The storage medium is used to store instructions;
[0044] The processor is configured to operate according to the instructions to execute the steps of the method described in the first aspect.
[0045] A third aspect of the present invention discloses a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the steps of the method described in the first aspect.
[0046] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention has the following advantages:
[0047] (1) In view of the temperature and humidity gradient and pipe network diversity in the large space of the computer room, the present invention deploys data acquisition and preprocessing units in zones, realizing near-real-time perception and preliminary fusion nearby, providing a data basis for subsequent spatial cooling analysis and regulation.
[0048] (2) The present invention fuses multi-dimensional raw data into comparable performance indicators, which can not only reflect the cooling efficiency of each zone, but also measure energy consumption and comfort, providing a refined view for intelligent decision-making. At the same time, the DDPG algorithm of the Actor-Critic architecture is introduced to optimize and verify the strategy. Through training, an intelligent control strategy network that can be directly deployed is obtained, which can learn the optimal operation strategy in a high-dimensional state space, balance the conflicts between various regulation objectives, and realize the adaptive regulation of the evaporative cooling air conditioning system. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a schematic flowchart of an intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention. DETAILED DESCRIPTION
[0050] The following further describes the present application with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present application.
[0051] As Figure 1 shown, in one embodiment, an intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning includes the following steps:
[0052] Step S110, obtaining the spatial elements of the refrigeration area and converting the spatial elements into a digital model to generate a functional block division strategy for the refrigeration area based on the digital model.
[0053] Among them, the spatial elements include the building floor plan, cooling load distribution, and pipe network structure, and the digital model includes a digitalized spatial coordinate set, load mapping, and pipe network topology diagram.
[0054] In some embodiments, for the intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention, step S110 specifically includes the following steps:
[0055] Step S111: Based on the digital spatial coordinate set, load mapping, and pipe network topology map, divide the refrigeration area into multiple functional blocks.
[0056] Step S112: Select key nodes from each functional block and set sensor units at the key nodes to generate a functional block division strategy; the sensor units include temperature and humidity sensors, flow meters, and differential pressure gauges.
[0057] In a specific embodiment, the intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention includes steps 1 to 5:
[0058] Step 1: System block division and fusion node deployment.
[0059] Specifically, for the temperature and humidity gradients and pipe network diversity in the large space of the computer room, deploy data acquisition and preprocessing units in zones to achieve near-real-time perception and preliminary fusion.
[0060] It includes the following steps:
[0061] Step 1.1: Basic data acquisition and digital modeling.
[0062] Specifically, first, obtain the building floor plan: including the two-dimensional plane layout of each air supply outlet, air return outlet, and wet curtain in the computer room; cold load distribution: the instantaneous cold load index corresponding to each air supply outlet (such as the cooling capacity collected in real time by the energy consumption management system); pipe network structure: the main and branch pipeline directions and valve node positions between the evaporative cooling tower and the air supply outlet. Convert paper or CAD drawings and operation data into a computable digital model, laying the foundation of "where, how much load, and how the pipelines are connected" and providing the three elements of space - heat - connectivity for subsequent functional block division.
[0063] Step 1.2: Functional block division.
[0064] Specifically, classify the air supply outlets and air return circuits with small heat load differences into the same block to avoid uneven heating and cooling in the large area, ensure that the air supply circuit and the cooling water circuit in the same block are connected or nearly connected in the physical pipeline network, reduce the response delay caused by long-distance transmission of cross-regional pipelines, and try to keep each block coherent in the room layout and reduce the situation where one block spans multiple rooms. Finally, output several functional block division schemes. The present invention solves the problems of excessive heating and cooling gradients and chaotic cross-control of the pipe network in a single area, providing clear spatial boundaries for subsequent parallel processing in zones and local optimization.
[0065] Step 1.3: Precise positioning of edge fusion nodes.
[0066] Specifically, in each functional block, the following "key points" are selected as the primary deployment locations:
[0067] At the wet curtain inlet: This is the key position where the evaporative cold water directly contacts the air, and the temperature and humidity of the inlet and outlet water need to be monitored.
[0068] At the pump outlet branch point: The intersection of the main water supply pipe and the branch pipes, and the water volume distribution of each branch is monitored through a flow meter.
[0069] At the air supply pipe bifurcation point: The node before the main air supply pipe bifurcates into each air supply outlet, and the overall air supply characteristics are monitored with temperature and humidity sensors.
[0070] At the valve control node: The positions of each regulating valve, and the flow resistance changes caused by the valve opening are detected with a differential pressure sensor.
[0071] After that, a fusion node (including temperature and humidity sensors, flow meters, differential pressure gauges, and edge computing units) is preferentially deployed at each key point. If the area or load of a block is large, additional nodes are added at equal distances or according to the load gradient distribution in the middle of the area or the remaining hot spots to ensure that there is no monitoring blind spot in the entire block. This process ensures that the key parts of heat exchange and water circuits in each block are monitored in real time and with high precision. The edge nodes complete preliminary data cleaning and fusion locally, reducing the burden on the central system and accelerating the response speed.
[0072] Step S120: Based on the functional block division strategy, obtain the sensor data on the fusion nodes of each functional block, and preprocess the sensor data to obtain a multi-dimensional feature data set.
[0073] Among them, the sensor data includes temperature, humidity, flow rate, and differential pressure.
[0074] In some embodiments, for the intelligent control method of the evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention, step S120 specifically includes the following steps:
[0075] Step S121: Receive the sensor data from each sensor unit, and construct an initial sequence matrix based on the sensor data according to the sampling period and time series.
[0076] Step S122: Perform a moving average filtering process on the sensor data in each column of the initial sequence matrix to obtain a smoothed sequence matrix.
[0077] Step S123: Fill in the missing values in the smoothed sequence matrix, and perform normalization processing on the filled smoothed sequence matrix through linear mapping to obtain a multi-dimensional feature data set.
[0078] In a specific embodiment, for the intelligent control method of the evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention, step 2 is multi-dimensional feature acquisition and preprocessing. Step 2 is used to clean and unify the data formats of various sensors, solve noise and missing value interferences, and provide high-quality input for subsequent fusion analysis, including the following steps:
[0079] Step 2.1, raw signal acquisition.
[0080] Specifically, first obtain the sensor data on the fusion nodes of each functional block, including the supply air temperature sensor data , relative humidity sensor data , circulating water flowmeter output data and wet curtain differential pressure sensor data , and construct a raw time series matrix according to the time series of these sensor data :
[0081]
[0082] Among them, is the total time, and t is the moment.
[0083] The above raw time series matrix is formed by synchronously reading the four types of sensor data at a fixed sampling period (such as 1 s) and caching them locally. This process unifies the analog / digital signals of heterogeneous sensors into a structured time series, laying a foundation for subsequent cleaning and analysis.
[0084] Step 2.2, noise filtering.
[0085] Specifically, apply a moving average filter to each column of data in matrix R to independently process temperature, humidity, flow rate, and differential pressure, and select a window length w (such as 5 - 15 points). Its expression is:
[0086]
[0087] Among them, is the smoothed result and serves as an element in the smoothed matrix S. t is the moment, and k represents the unit time corresponding to the window length.
[0088] This process eliminates the spikes and jitters caused by sensor noise and pipeline pulsation, improves the smoothness of the sensor signal, and avoids subsequent misjudgments.
[0089] Step 2.3, missing value filling.
[0090] Specifically, when there is a single-point missing value in the smoothed matrix S, linear interpolation before and after is used for filling. If more than three consecutive points are missing, the historical average value of the same period of this functional block or the data of adjacent synchronous blocks is used for backfilling to ensure continuity, and finally a complete matrix I without missing values is output. This process fills the data holes caused by network jitter or temporary sensor failure, ensures the integrity of the feature sequence, and prevents the occurrence of NAN during training.
[0091] Step 2.4, feature normalization.
[0092] Specifically, each channel in the complete matrix I is linearly mapped to the interval [0,1], and its mapping formula is:
[0093]
[0094] where represents the value of the j-th feature (select one from temperature, humidity, flow rate, and pressure difference) at time t, represents the mapping result corresponding to the value of the j-th feature (select one from temperature, humidity, flow rate, and pressure difference) at time t, and respectively represent the lower and upper limits of the range of the sensor corresponding to this feature.
[0095] This process unifies physical quantities with different dimensions, eliminates the numerical scale difference, and can ensure stable convergence during the training process of the deep reinforcement learning model.
[0096] Step S130, fuse the multi-dimensional feature dataset into performance indicators to output the performance evaluation vector of each functional block.
[0097] In some embodiments, for the intelligent control method of the evaporative cooling air conditioner system based on deep reinforcement learning provided by the present invention, step S130 specifically includes the following steps:
[0098] Step S131, align the timestamps of the time series data of each functional block in the multi-dimensional feature dataset through the Network Time Protocol, and linearly interpolate and fill the sensor data with a time error exceeding the first threshold.
[0099] Step S132, judge the health status of the sensor devices through the heartbeat detection frequency and residual statistics, and assign health weights to the sensor data according to the degree of health, so as to fill the sensor data with a health weight lower than the second threshold to obtain health-weighted time series data.
[0100] Step S133, perform moving window median filtering on the health-weighted time series data, and perform weighted average on the sensor data at the same moment according to the health weights to obtain the fused time series curve of each functional block.
[0101] Step S134, statistically analyze the performance metrics of the fusion time-series curve according to the set sliding window; the performance metrics include cooling efficiency, temperature and humidity deviation, and energy load.
[0102] Step S135, mark the performance metrics with a sliding average abnormal jump exceeding the third threshold as abnormal metrics, and package the abnormal metrics and performance metrics to obtain a performance evaluation vector.
[0103] In a specific embodiment, for the intelligent control method of an evaporative cooling air-conditioning system based on deep reinforcement learning provided by the present invention, step 3, distributed fusion analysis generates regional performance metrics. Step 3 is used to fuse multi-dimensional raw data into comparable performance metrics, which can not only reflect the cooling efficiency of each area, but also measure energy consumption and comfort, providing a refined view for intelligent decision-making, including the following steps:
[0104] Step 3.1, time series alignment and robust weighting.
[0105] Specifically, first, align the timestamps of each node with the Network Time Protocol (NTP), and linearly interpolate and fill in the data with an error exceeding half of the sampling period. For sensor drift or disconnection, determine the health status through the heartbeat detection frequency and residual statistics within the previous and next two minutes, and dynamically assign weights according to the health degree. For the node data with extremely low weights or long-term dropouts, directly set the weight to zero and fill it with the data of the nearest neighbor node. Finally, output the health-weighted time series data. This process can ensure the time consistency of the fused basic data, dynamically reduce the influence of unreliable nodes, and improve the subsequent fusion quality.
[0106] Step 3.2, abnormal rejection and weighted fusion.
[0107] Specifically, use median filtering in a sliding window (window length 10s) to eliminate extreme pulses in the time series of the health-weighted time series data output in step 3.1, and calculate the weighted average of the readings of each node at the same moment according to the health weights to obtain a fused value sequence representing the entire block. Finally, integrate and output the fused time series curve of each physical quantity (temperature difference, humidity, flow rate, pressure difference). This process uses median filtering to eliminate occasional spikes, and weighted averaging ensures that the fused readings are both stable and representative of the overall situation.
[0108] Step 3.3, key index period statistics.
[0109] Specifically, a 1-minute sliding window is applied to the fused time-series curve, and indicators within each window are calculated at a 10s step size. A rolling queue of the most recent 6 windows is maintained. If the sliding average has an abnormal jump exceeding 10%, it is marked as "indicator abnormal" and enters the priority alarm flow. Three indicators within the sliding statistical window are obtained: cooling efficiency (average temperature difference), humidity deviation (average humidity deviation), and energy consumption load (flow × average pressure difference). This process provides a basis for multi-period comparison in comprehensive evaluation by regularly assessing the short-term performance of the block and capturing sudden anomalies.
[0110] Step 3.4, abnormal priority and comprehensive vector packaging.
[0111] Specifically, based on the three indicators and the abnormal marks obtained in Step 3.3, if the "cooling efficiency" or "humidity deviation" triggers an abnormal mark, the comprehensive score directly takes this indicator and enters priority adjustment. Otherwise, after normalizing the three indicators according to the preset comfort weight (0.6) and energy-saving weight (0.4), the weighted sum is calculated, and the three original statistical values and the comprehensive score are output together as a performance vector. This process achieves a balance between comfort and energy saving under most normal operating conditions. Once serious discomfort or energy consumption anomalies occur, the specific problem item is preferentially fed back to guide the intelligent agent to focus on correction.
[0112] Step S140, construct an experience replay pool based on the performance evaluation vector to fuse the global state vector for offline training of the DDPG algorithm, and obtain the intelligent control strategy network.
[0113] In some embodiments, for the intelligent control method of the evaporative cooling air-conditioning system based on deep reinforcement learning provided by the present invention, Step S140 specifically includes the following steps:
[0114] Step S141, splice the performance evaluation vectors of each functional block in time series to form the global state vector of the refrigeration area.
[0115] Step S142, when the experience replay pool reaches the capacity limit, sort the global state vectors according to the TD error size of the global state vectors to retain the global state vectors whose TD error exceeds the fourth threshold, and obtain the global state samples.
[0116] In some embodiments, for the intelligent control method of the evaporative cooling air-conditioning system based on deep reinforcement learning provided by the present invention, Step S140 specifically further includes the following steps:
[0117] Step S143, determine whether the temperature difference or energy consumption of each functional block exceeds the preset high warning threshold; if so, mark the functional block with the first label; if not, mark the functional block with the second label.
[0118] Step S144, when multiple adjacent functional blocks are marked with the first label, add coupling weights to the multiple adjacent functional blocks, and splice the functional blocks carrying the coupling weights with other functional blocks to obtain a fused state vector.
[0119] Step S145, use the fused state vector for offline training with the DDPG algorithm based on the Actor-Critic framework to obtain an intelligent control policy network.
[0120] In a specific embodiment, for the intelligent control method of an evaporative cooling air-conditioning system based on deep reinforcement learning provided by the present invention, step 4, offline deep reinforcement learning policy training. Step 4 takes the performance evaluation vectors of each functional block output in step 3, the corresponding historical control actions, and reward signals as inputs, performs offline training through constructing an experience replay pool, fusing the global state, and using the DDPG algorithm with the Actor–Critic architecture, and tunes and verifies the policy. Finally, an intelligent control policy network that can be directly deployed is output. It includes the following steps:
[0121] Step 4.1, construct an experience replay pool.
[0122] Specifically, first, determine the performance evaluation vectors (quadruple: temperature difference, humidity deviation, energy consumption load, comprehensive score) of all blocks in each control period (for example, every 1 minute); the corresponding action vectors (cooling water valve opening and fan frequency of each block) in the same period, and the reward value (the global reward scalar synthesized according to the regional priority after calculating the "comfort - energy consumption" return of each block with a reward function). Secondly, for the recorded offline simulation or on-site log, sequentially read the performance vectors of each block in time series and splice them into a global state, and at the same time read the corresponding actions and rewards, and then read the state of the next period to generate a complete sample.
[0123] To prevent the policy from only learning normal working conditions, establish a "high-return sample pool" and a "low-return sample pool", and save samples with returns greater than the upper percentile and lower than the lower percentile with a higher probability respectively, and ordinary samples enter the pool with a lower probability to ensure that the replay pool takes into account both extreme and normal experiences. When the replay pool reaches the capacity limit, sort the samples according to the TD error (i.e., the Critic estimation error) size, preferentially retain samples with large errors, and eliminate the samples with the lowest priority.
[0124] Regularly save the snapshot of the replay pool to the database for easy recovery when training restarts, support multi-version management, and record the sample range used in each batch of training for traceability. Finally, output an experience replay pool DDD in the form of a double-ended circular queue, storing the quadruple of "global state - action - reward - next global state", with a configurable capacity (such as 100,000 entries).
[0125] This process breaks the temporal correlation, improves the training stability and efficiency, ensures that the policy has sufficient learning samples under various working conditions (extreme and normal) through balancing and prioritization strategies, and the replay pool persistence provides a data basis for experiment reproducibility and cross-team collaboration.
[0126] Step 4.2, Global state fusion and regional coupling.
[0127] Specifically, first determine the original global state vector in the replay pool (the concatenation of the performance quadruples of each block) and the predefined "high alert threshold" and "coupling regulation threshold" - such as a temperature difference exceeding 3°C or an energy consumption index exceeding 80% of the historical peak. Perform threshold judgment on the temperature difference and energy consumption indicators of each block: if the temperature difference or energy consumption of a certain block exceeds the "high alert threshold", the label of this block is "urgently needed to be adjusted"; otherwise, the label is "normal". If multiple (≥2) adjacent blocks (connected positions in the pipe network topology) are simultaneously marked as "urgently needed to be adjusted", then trigger the "coupling regulation" mode; add an additional one-dimensional "coupling weight" to these blocks, and the value can be allocated according to the proportion of the number of blocks exceeding the threshold (for example, if two blocks exceed the threshold simultaneously, the weight is 0.8; if three, it is 0.9). In addition, it is also necessary to concatenate the four-dimensional performance vectors of all blocks in a fixed order together with their respective "urgently needed to be adjusted" binary labels and "coupling weights". If the overall is a "coupling regulation" scenario, then append a "coupling trigger" flag bit at the end of the vector. Finally, output the fused global state vector with coupling marks and priority attention coefficients for input to the Actor network.
[0128] This process enables the coexistence of block independence and coupling scenarios, enabling the policy to "locally independently adjust" in general situations and also "cooperatively optimize" when the adjacent area linkage is unbalanced. The label and weight mechanism allows the network to automatically identify and focus on the blocks that urgently need to be adjusted and their mutual influences, effectively avoiding isolated decision-making in each area.
[0129] Step 4.3, DDPG algorithm training and cross-region interaction modeling.
[0130] Specifically, based on the fused global state vector (including coupling marks), the corresponding global action vector, and the global rewards and next global states in the replay pool obtained in the previous steps, select the Actor-Critic architecture. Among them, the Actor takes the fused state as input and outputs continuous global actions (the valve openings and fan frequencies of all blocks), and has a coupling perception layer for processing adjacent block interaction features; the Critic takes the state-action pair as input and outputs the estimated reward value of this combination under the off-policy for evaluating the quality of the action.
[0131] In this embodiment, during the critic update process, a small batch (e.g., 64) of samples is randomly sampled to calculate the target reward (based on the next state and the estimated reward of the target network). The mean squared loss between the current critic output and the target reward is minimized. During the actor update process, the gradient of the critic's output with respect to the current policy action is used to adjust the actor parameters to maximize the expected reward estimated by the critic. A "neighborhood convolution" or "self-attention" module is added after the actor network input layer to perform local interactive learning on all concatenated block state vectors, enabling the network to automatically extract coupling features between adjacent blocks. During the initial training phase, time-correlated noise (e.g., the Ornstein-Uhlenbeck process) is added to the actor output to encourage diverse exploration. Soft updates (τ = 0.005) are also used to periodically synchronize with the target network to reduce the risk of training divergence.
[0132] In this embodiment, DDPG can process high-dimensional continuous action spaces and directly output multi-block coordinated adjustment instructions. The actor-critic separation and coupled perception layer jointly ensure that the strategy has strong coordination and generalization capabilities in cross-region coupling scenarios. Details such as exploration and target network soft updates improve the stability of the training process and the robustness of the final strategy.
[0133] Step 4.4: Offline verification, tuning, and model solidification.
[0134] Specifically, a multi-scenario simulation environment was used: Scenario A: A sudden increase in heat load in two adjacent blocks (coupling trigger); Scenario B: A single block with high load and the remaining blocks with low load (independent scenario); and Scenario C: A uniform medium load across the entire block (conventional scenario). At least 2,000 simulation steps were run in each scenario, periodically recording changes in temperature and humidity differences, energy load, and coupling across each block. The coupled scenario measured whether the strategy synchronized water flow and fan frequency increases in two adjacent blocks experiencing sudden heat load increases without accidentally impacting other blocks. The independent scenario verified whether the strategy only adjusted the high-load block, maintaining or fine-tuning the actions of other blocks. The conventional scenario tested whether the overall energy consumption reduction rate and comfort deviation met the predetermined targets (e.g., comfort deviation <1°C, energy consumption reduction >10%). If performance metrics were not met in any scenario, the deviation between the critic's value estimate and the actual reward was analyzed, and penalty weights were appropriately adjusted or the specific scenario samples were resampled and retrained. In addition, after all scenarios are passed, the final model weight file is output and a detailed verification report including scenario comparison curves, convergence trends and sample coverage is generated.
[0135] During this process, offline verification ensures that the policy can respond correctly in both coupled and independent complementary scenarios, taking into account both local and global optimality. The tuning mechanism enables the model to continuously improve until it meets the multi-dimensional deployable criteria. Model solidification and reporting provide a reliable basis for on-site one-key deployment and later maintenance.
[0136] Step S150: Input the real-time performance indicators of the refrigeration area into the intelligent control policy network for processing, and output the control instructions for the evaporative cooling air conditioning system.
[0137] In some embodiments, for the intelligent control method of the evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention, step S150 specifically includes the following steps:
[0138] Step S151: Obtain the real-time performance indicators of each functional block in the refrigeration area, and screen the real-time performance indicators according to the health labels of the real-time performance indicators, so as to eliminate and fill in the sensor data with a health level lower than the fifth threshold, and obtain the four-element performance values of each functional block.
[0139] Step S152: Concatenate the four-element performance values of all functional blocks, and carry the global state vector and health weight of each functional block during concatenation to output the real-time global state vector.
[0140] Step S153: Invoke the intelligent control policy network to process the real-time global state vector, generate the recommended valve opening and fan frequency of each functional block of the evaporative cooling air conditioning system, and output the control instructions to be executed according to the recommended valve opening and fan frequency.
[0141] Step S160: Respond to the control instructions to control the valve opening and fan frequency of each functional block.
[0142] In some embodiments, for the intelligent control method of the evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention, the method further includes the following steps:
[0143] Step S210: When a sensor device of any functional block fails, determine that the functional block is a faulty block, switch the faulty block to a preset PID redundant loop, and simultaneously send an alarm message to the operation and maintenance terminal.
[0144] Step S220: Generate a global action vector based on the recommended valve opening and fan frequency, and send the global action vector to the control center. The control center issues the control instructions to the PLC through the fieldbus.
[0145] In a specific embodiment, for the intelligent control method of the evaporative cooling air conditioning system based on deep reinforcement learning provided by the present invention, step 5 is online predictive control and fault emergency switching. In step 5, the optimal control actions predicted by the policy network are sent down to respond to the fluctuations of the cooling and heating loads in advance. When it is detected that the water pump stops rotating, the sensor fails, or the performance index exceeds the limit, the preset PID redundant loop is automatically triggered and an alarm is given to ensure the basic cooling demand, realizing predictive closed-loop control and having the ability of fault tolerance and degradation, and ensuring the stable operation of the system under sudden conditions. It includes the following steps:
[0146] Step 5.1, obtaining the real-time performance vector.
[0147] Specifically, based on the fused performance evaluation vector (quadruple: average temperature difference, humidity deviation, energy consumption load, comprehensive score) of each functional block at the current moment and the latest sensor health and weight information (used to verify the data reliability), each edge fusion node immediately uploads the quadruple performance evaluation of this block to the control center. The control center checks the health tags of each node, eliminates the data with a health lower than the threshold (such as 0.7), and fills it with the adjacent area or historical value. Finally, the quadruple performance values of all blocks are spliced in a predetermined order and an urgent adjustment or coupling trigger flag is added to form a global state vector. This process can ensure that the state data used for online control is the latest and reliable, providing a clear and comparable full-system observation for the subsequent policy network and supporting cross-block collaborative optimization.
[0148] Step 5.2, generating predictive control actions.
[0149] Specifically, based on the obtained global state vector and the pre-trained Actor network, the control center inputs the global state vector into the pre-trained Actor network to output an initial action suggestion, that is, recommend a global action vector, including the valve opening and fan frequency of each block. If any two adjacent blocks are simultaneously marked with "coupling trigger", then apply a coupling adjustment strategy to their corresponding actions, that is, slightly synchronously increase the fan frequencies of the two blocks or coordinate and balance the valve openings to jointly relieve the high load; otherwise, independently optimize the actions of each block to ensure local rapid response. In addition, range constraints (such as the valve opening is in [0,1], and the fan frequency is in [10,60] Hz) and slope limits (to avoid sudden changes) are also required for each action in the recommended global action vector. This process realizes online predictive control through the policy network, responds to load fluctuations in advance, and the condition judgment and coupling adjustment also ensure that when the coordination between adjacent blocks is unbalanced, they can exert force simultaneously, avoiding the performance deterioration of other blocks caused by the actions of a single block.
[0150] Step 5.3, fault detection and emergency switching.
[0151] Specifically, based on the real-time sensor raw data, the fusion performance vector, and the device operation status alarm signals (such as zero water pump flow or abnormal wet curtain pressure difference), if it is detected that the water flow in a certain block drops to zero briefly or the pressure difference exceeds the maximum safety threshold (such as 5 kPa), it is determined that there is a failure in the water circuit or sensor of that block. If the wet curtain inlet humidity sensor reads out of range continuously for 3 times, it is determined that the measurement point fails. Immediately switch the faulty block to the pre-tuned "PID redundant loop", which will use fixed PID parameters to maintain the minimum supply air temperature and the lowest energy consumption. Other normal blocks continue to adopt the actions output by the deep policy network. At the same time, trigger an operation and maintenance alarm, and push the fault information to the operation and maintenance personnel through the SCADA system. This process can instantly restore to the controllable traditional PID control when a key device or measurement point fails, ensure the basic cooling demand, isolate the non-faulty blocks from the faulty blocks, and prevent the fault from spreading and affecting the overall system.
[0152] Step 5.4, control command execution and closed-loop feedback.
[0153] Specifically, based on the finally selected global action vector (including predictive and emergency commands), the control center issues valve and fan commands to the PLC through the fieldbus (such as Modbus / TCP). After the PLC performs rate limiting and anti-jitter processing on the commands, it outputs them to the frequency converter and servo valve. Immediately after execution, the actual execution value is collected, compared with the command value, and an error report is generated and returned to the edge fusion node. This process realizes the implementation of actions and visual monitoring, ensures that the commands are executed correctly, detects the command and execution deviation through closed-loop feedback, and provides real error data for the preprocessing and training of the next cycle.
[0154] Step 5.5, online log recording and policy evaluation.
[0155] Specifically, write the global state vector, actions, rewards, and execution feedback of each cycle into the time series database. Automatically count the energy consumption, comfort deviation, and number of faults per hour / day to generate KPI Charts. If it is found that the reward of the policy drops severely continuously under a certain working condition, immediately trigger the "online policy switching" or "offline training again" process. This process can ensure the traceability and transparency of the online operation of the policy, timely detect policy failure or system changes through real-time evaluation, and support online fine-tuning or trigger the next round of offline training.
[0156] The present disclosure can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0157] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0158] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0159] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0160] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0161] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0162] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0163] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. An intelligent control method for an evaporative cooling air conditioning system based on deep reinforcement learning, characterized in that, The method includes: Obtaining the spatial elements of the refrigeration area, and converting the spatial elements into a digital model to generate a functional block division strategy for the refrigeration area based on the digital model; Based on the functional block division strategy, obtaining the sensor data on the fusion nodes of each functional block, and preprocessing the sensor data to obtain a multi-dimensional feature data set; Fusing the multi-dimensional feature data set into performance indicators to output the performance evaluation vectors of each functional block; Constructing an experience replay pool to fuse the global state vector based on the performance evaluation vector to perform offline training on the DDPG algorithm to obtain an intelligent control strategy network; Inputting the real-time performance indicators of the refrigeration area into the intelligent control strategy network for processing, and outputting the control instructions of the evaporative cooling air conditioning system; Controlling the valve opening and fan frequency of each functional block in response to the control instructions; Wherein, the spatial elements include the building floor plan, the cooling load distribution, and the pipe network structure, and the digital model includes a digitalized spatial coordinate set, a load mapping, and a pipe network topology map; the sensor data includes temperature, humidity, flow rate, and pressure difference.
2. The intelligent control method of the evaporative cooling air-conditioning system based on deep reinforcement learning according to claim 1, characterized in that, The obtaining the spatial elements of the refrigeration area, and converting the spatial elements into a digital model to generate a functional block division strategy for the refrigeration area includes: Dividing the refrigeration area into multiple functional blocks based on the digitalized spatial coordinate set, the load mapping, and the pipe network topology map; Selecting key nodes from each functional block, and setting sensor units at the key nodes to generate the functional block division strategy; the sensor units include temperature and humidity sensors, flow meters, and differential pressure gauges.
3. The intelligent control method of the evaporative cooling air-conditioning system based on deep reinforcement learning according to claim 2, characterized in that, The based on the functional block division strategy, obtaining the sensor data on the fusion nodes of each functional block, and preprocessing the sensor data to obtain a multi-dimensional feature data set includes: Receiving the sensor data from each sensor unit, and constructing an initial sequence matrix based on the sensor data according to the sampling period and time series; Performing a moving average filtering process on the sensor data of each column in the initial sequence matrix to obtain a smoothed sequence matrix; Performing missing value filling on the smoothed sequence matrix, and performing normalization processing on the filled smoothed sequence matrix through linear mapping to obtain the multi-dimensional feature data set.
4. The intelligent control method of the evaporative cooling air conditioning system based on deep reinforcement learning according to claim 3, characterized in that, The fusing the multi-dimensional feature data set into performance indicators to output the performance evaluation vectors of each functional block includes: Aligning the time stamps of the time series data of each functional block in the multi-dimensional feature data set through the Network Time Protocol, and linearly interpolating and filling the sensor data with a time error exceeding the first threshold; Judging the health status of the sensor devices through the heartbeat detection frequency and residual statistics, and assigning health weights to each sensor data according to the health degree, so as to fill the sensor data with a health weight lower than the second threshold to obtain health-weighted time series data; Performing a moving window median filtering on the health-weighted time series data, and performing a weighted average on the sensor data at the same moment according to the health weights to obtain the fusion time series curve of each functional block; Statistically analyze the performance metrics of the fused time-series curve according to a set sliding window; the performance metrics include cooling efficiency, temperature and humidity deviation, and energy load; Mark the performance metrics with a sliding average abnormal jump exceeding the third threshold as abnormal metrics, and package the abnormal metrics and performance metrics to obtain the performance evaluation vector.
5. The intelligent control method of the evaporative cooling air-conditioning system based on deep reinforcement learning according to claim 4, characterized in that, Constructing an experience replay pool to fuse the global state vector based on the performance evaluation vector to perform offline training on the DDPG algorithm to obtain an intelligent control strategy network, including: Concatenate the performance evaluation vectors of each functional block in time series to form the global state vector of the refrigeration area; When the experience replay pool reaches the capacity limit, sort each global state vector according to the TD error size of the global state vector to retain the global state vectors with TD errors exceeding the fourth threshold to obtain global state samples.
6. The intelligent control method of the evaporative cooling air-conditioning system based on deep reinforcement learning according to claim 5, characterized in that, Constructing an experience replay pool to fuse the global state vector based on the performance evaluation vector to perform offline training on the DDPG algorithm to obtain an intelligent control strategy network, further including: Judge whether the temperature difference or energy consumption of each functional block exceeds a preset high warning threshold; if so, mark the functional block with the first label; if not, mark the functional block with the second label; When multiple adjacent functional blocks are marked with the first label, add a coupling weight to the multiple adjacent functional blocks, and splice the functional blocks carrying the coupling weight with other functional blocks to obtain a fused state vector; Use the fused state vector as the DDPG algorithm with the Actor-Critic framework for offline training to obtain the intelligent control strategy network.
7. The intelligent control method of the evaporative cooling air conditioning system based on deep reinforcement learning according to claim 6, characterized in that, Inputting the real-time performance metrics of the refrigeration area into the intelligent control strategy network for processing and outputting control instructions for the evaporative cooling air-conditioning system, including: Obtain the real-time performance metrics of each functional block in the refrigeration area, and screen the real-time performance metrics according to the health labels of the real-time performance metrics to eliminate and fill the sensor data with a health level lower than the fifth threshold to obtain the four-element performance values of each functional block; Concatenate the four-element performance values of all functional blocks, and carry the global state vector and health weight of each functional block during concatenation to output a real-time global state vector; Call the intelligent control strategy network to process the real-time global state vector to generate the recommended valve opening and fan frequency of each functional block of the evaporative cooling air-conditioning system, and output control instructions executed according to the recommended valve opening and fan frequency.
8. The intelligent control method of the evaporative cooling air-conditioning system based on deep reinforcement learning according to claim 7, wherein The method further includes: When a sensor device of any functional block fails, determine that the functional block is a faulty block, switch the faulty block to a preset PID redundant loop, and send an alarm message to the operation and maintenance terminal at the same time; Generate a global action vector based on the recommended valve opening and fan frequency, and send the global action vector to the control center, and the control center issues the control instructions to the PLC through the field bus.
9. A terminal, including a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Subway station air conditioning system energy-saving control method based on deep reinforcement learning
CN113283156A
Central air conditioner non-inductive automatic decision-making method based on deep reinforcement learning
CN116753600A