Intelligent regulation and control method for cooling tower top fan

By collecting multi-dimensional data from cooling towers to construct multi-scale feature vectors and using a reinforcement learning model to dynamically adjust reward weights, the system addresses the issues of insufficient state perception and dynamic reward mechanism in complex scenarios of the intelligent control system for cooling tower fans, thereby achieving high efficiency, energy saving, and stable operation.

CN121474929APending Publication Date: 2026-02-06GUANGZHOU SINGLE BEAM ALL STEEL COOLING TOWER EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202512004740.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing intelligent control systems for cooling tower fans suffer from insufficient state perception and modeling granularity, weak robustness, and inability to adapt to environmental changes in real time when facing complex multi-objective scenarios. Furthermore, the control reward mechanism lacks dynamism, resulting in an inability to balance energy saving and stability, and frequent imbalances.

Method used

Multi-dimensional state data of the cooling tower's operating environment are collected to construct multi-scale feature vectors. The reward weights are dynamically adjusted through a reinforcement learning model to generate a fan speed control strategy. The model parameters are optimized through closed-loop execution and feedback to achieve real-time adaptation to the environment and multi-objective optimization.

Benefits of technology

It improves the accuracy of environmental modeling and the adaptability of control strategies for cooling tower systems, reduces energy consumption, enhances user comfort and overall system performance, and reduces energy waste and human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121474929A_ABST
    Figure CN121474929A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent regulation and control method for a cooling tower top fan, and the method comprises the steps: collecting the time, space and equipment dimension state information of a cooling tower in real time, constructing a multi-scale feature vector, combining the operation stage recognition with the external energy-saving target priority, dynamically adjusting the multi-target weight distribution, and generating a composite reward function, thereby achieving the intelligent regulation and control of the cooling tower top fan. Inputting a reinforcement learning model to carry out fan rotating speed optimization control; and the system performs closed-loop collection and execution feedback and performs online fine adjustment on the model, adapts to operation stage switching, realizes adaptive efficient adjustment of the fan under multiple targets, improves the operation energy-saving level of the cooling tower, the system stability and the user experience, and has both flexibility and environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for cooling tower fans, and in particular to an intelligent control method for cooling tower top fans. Background Technology

[0002] Currently, intelligent regulation and multi-objective optimization control technology for cooling tower fans is an important development direction in the field of industrial refrigeration and energy management. With the continuous advancement of artificial intelligence, the Internet of Things, and automation technologies, more and more cooling tower systems are adopting intelligent control strategies to achieve comprehensive improvements in energy saving, stable operation, and user comfort by automatically adjusting key components such as fans and pumps. Reinforcement learning, as an increasingly widely used intelligent optimization tool in recent years, has been widely introduced into the field of industrial process control, especially in systems with complex, multi-objective dynamic constraints. Current mainstream solutions in the industry mostly rely on static rules, traditional PID control, or single-objective optimization strategies. Although some solutions introduce technologies such as deep reinforcement learning and fuzzy control to achieve a certain degree of adaptive adjustment, in complex scenarios where energy saving, equipment stability, and user comfort coexist, the commonly used reward function weights, fixed operating condition thresholds, and target priorities set manually or based on experience are difficult to balance multiple objectives in real time when operating conditions change. This results in control strategies often being biased towards a single objective and failing to continuously achieve optimal overall performance. The intelligent control system for cooling towers also faces two major technical challenges: First, the granularity and robustness of state perception and modeling are insufficient. Existing reinforcement learning or multi-objective optimization frameworks mostly adopt a single-scale environmental state definition, which cannot take into account the multi-dimensional and multi-scale dynamic characteristics of time (such as load fluctuation trends), space (such as temperature distribution gradients), equipment operating conditions, and control states. This makes the state input less sensitive to complex environmental changes, weakens the adaptability and generalization ability of the control strategy, and lacks the ability to intelligently identify and switch between operating stages such as startup, stabilization, and load fluctuations, making it impossible to adjust the target weights and strategy generation mechanism in a targeted manner. Second, the dynamics of the control reward mechanism are insufficient. Most methods rely on static reward function design, and the weights of sub-objectives such as energy saving, stability, and comfort are mostly manually set or only have simple dynamic adjustability. They cannot be intelligently and stably dynamically adjusted according to real-time operating conditions, changes in external objectives (such as energy consumption assessment periods), and system operating status, resulting in weak multi-objective balance and frequent imbalances such as "energy saving priority but temperature control fluctuations" and "rapid response but increased energy consumption". Summary of the Invention

[0003] In order to solve the above-mentioned technical problems, the present invention provides a method for intelligent control of cooling tower top fans.

[0004] The technical solution of this invention is implemented as follows: A method for intelligent control of a cooling tower top fan, comprising: S1: Collect time, space and equipment status data in the operating environment of the cooling tower. The time dimension includes real-time temperature and humidity and load fluctuation information, the space dimension includes temperature sensor data in different areas of the tower, and the equipment dimension includes fan speed, water pump start and stop status and energy consumption data. S2: Normalize the collected multi-source state data and construct a multi-scale state feature vector to characterize the dynamic operating conditions of the cooling tower at different operating stages. S3: Based on the current operation stage identification result, match the current stage label from the preset operation stage label library. The operation stages include the startup stage, the stable operation stage, and the load fluctuation stage. S4: Based on the matched operational phase labels and the priority of external energy-saving goals, dynamically adjust the weight coefficients of the three sub-goals of energy saving, stability and user comfort in the total reward function to generate a dynamic reward weight allocation strategy; S5: Input the multi-scale state feature vector and dynamic reward weight allocation strategy into the reinforcement learning model, generate wind turbine speed control action suggestions through the policy network, and calculate the corresponding action value function; S6: Based on the fan speed control action suggestions output by the reinforcement learning model, send the speed adjustment command to the cooling tower fan control system and execute the fan speed adjustment operation; S7: Collect system feedback data after the fan is executed, including actual speed, tower temperature change rate and energy consumption change, to evaluate the actual effect of the control action; S8: Update the experience replay buffer of the reinforcement learning model based on system feedback data, and fine-tune the model parameters online through the target network to improve the model's adaptability under the current working conditions; S9: Determine if the current running phase has switched. If it has switched, regenerate the dynamic reward weight allocation strategy based on the new running phase label and trigger the policy update process of the reinforcement learning model. S10: Based on historical control actions and system feedback data, construct multi-objective optimization performance evaluation indicators for subsequent model training strategy iteration optimization and parameter tuning.

[0005] The intelligent control method for cooling tower top fan provided in this application has the following beneficial effects: (1) This invention uses real-time acquisition and normalization of multi-dimensional data of time, space and equipment, and multi-scale feature extraction. The model can accurately reflect the dynamic changes of the cooling tower operating environment. The special processing flow, such as sliding window statistics, Kalman filtering, spatial principal component analysis and equipment status consistency test, greatly improves the accuracy of status input and the sensitivity of the model to environmental fluctuations. Compared with traditional single or low-dimensional monitoring methods, the modeling accuracy of environmental status is greatly improved, laying a solid data foundation for subsequent strategy optimization and online decision-making. (2) The mechanisms designed in this invention, such as operation phase label recognition, energy-saving priority input, logical mapping allocation, fuzzy logic adaptive adjustment, and historical strategy fusion, realize the continuous dynamic adjustment of reward weight under different operating conditions and external target driving. The weight allocation always matches the operating phase of the cooling tower with the current energy-saving / comfort / stability requirements in real time, which solves the problem that the existing static weight mechanism has a single decision and cannot take into account multiple objectives. Compared with the traditional static weight method, this invention can significantly improve the comprehensive performance index of the system and significantly reduce the risk of energy efficiency fluctuation and user experience decline caused by the bias of energy-saving and stability weights. (3) Under the input of multi-scale state features and dynamic reward allocation, the deep neural network strategy layer of this invention can generate the optimal wind turbine control action for real-time operating conditions, and evaluate and adjust the output scheme through value network / exploration mechanism. The innovation also includes experience playback and online fine-tuning of the target network, and dynamic iteration of strategy triggered by operating condition switching, so that the model can always maintain the rationality of decision-making in scenarios such as load fluctuation. Compared with the existing control methods that are manually set or based on single-objective dynamic adjustment, the actual energy consumption of the wind turbine is significantly reduced, the temperature change rate control accuracy is significantly improved, and the user comfort is kept at a high level. (4) The present invention automatically encapsulates and sends control action suggestions to the frequency converter, and realizes closed-loop execution through industrial communication protocol. With multi-dimensional feedback monitoring and data synchronization, it further improves the whole process closed loop from intelligent decision-making to actual adjustment, improves the system response speed, and reduces energy waste and manual intervention requirements, so that intelligent control can be stably and reliably applied in large-scale cooling tower groups. Attached Figure Description

[0006] Figure 1 This is a flowchart of an intelligent control method for a cooling tower top fan according to the present invention; Figure 2 This is a sub-flowchart of a method for intelligent control of a cooling tower top fan according to the present invention; Figure 3 This is another sub-flowchart of the intelligent control method for the top fan of a cooling tower according to the present invention. Detailed Implementation

[0007] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0008] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0009] like Figure 1 As shown, this application provides an intelligent control method for a cooling tower top fan, specifically including: S1: Collect time, space and equipment status data in the operating environment of the cooling tower. The time dimension includes real-time temperature and humidity and load fluctuation information, the space dimension includes temperature sensor data in different areas of the tower, and the equipment dimension includes fan speed, water pump start and stop status and energy consumption data. S2: Normalize the collected multi-source state data and construct a multi-scale state feature vector to characterize the dynamic operating conditions of the cooling tower at different operating stages. S3: Based on the current operation stage identification result, match the current stage label from the preset operation stage label library. The operation stages include the startup stage, the stable operation stage, and the load fluctuation stage. S4: Based on the matched operational phase labels and the priority of external energy-saving goals, dynamically adjust the weight coefficients of the three sub-goals of energy saving, stability and user comfort in the total reward function to generate a dynamic reward weight allocation strategy; S5: Input the multi-scale state feature vector and dynamic reward weight allocation strategy into the reinforcement learning model, generate wind turbine speed control action suggestions through the policy network, and calculate the corresponding action value function; S6: Based on the fan speed control action suggestions output by the reinforcement learning model, send the speed adjustment command to the cooling tower fan control system and execute the fan speed adjustment operation; S7: Collect system feedback data after the fan is executed, including actual speed, tower temperature change rate and energy consumption change, to evaluate the actual effect of the control action; S8: Update the experience replay buffer of the reinforcement learning model based on system feedback data, and fine-tune the model parameters online through the target network to improve the model's adaptability under the current working conditions; S9: Determine if the current running phase has switched. If it has switched, regenerate the dynamic reward weight allocation strategy based on the new running phase label and trigger the policy update process of the reinforcement learning model. S10: Based on historical control actions and system feedback data, construct multi-objective optimization performance evaluation indicators for subsequent model training strategy iteration optimization and parameter tuning.

[0010] Step S1: Collect time-dimensional, spatial-dimensional, and equipment-dimensional status data of the cooling tower's operating environment. The time dimension includes real-time temperature and humidity and load fluctuation information; the spatial dimension includes temperature sensor data from different areas of the tower; and the equipment dimension includes fan speed, water pump start / stop status, and energy consumption data. Specifically, this includes: S1.1: Collect real-time temperature and humidity data in the cooling tower's operating environment to obtain the current environmental thermodynamic state parameters, which serve as the basic input for time-dimensional features; Real-time temperature and humidity signals of the cooling tower's operating environment are obtained using a high-precision digital temperature and humidity sensor array (parameter: temperature resolution). ℃, humidity resolution (%RH), to achieve continuous data acquisition; Furthermore, through the analog-to-digital conversion module (parameter: sampling precision) Bits, sampling rate (Hz), converting the analog or semi-digital signal output by the sensor into a digital data stream in a unified format, and obtaining a structured raw temperature and humidity data matrix; Furthermore, a Kalman filter algorithm (parameters: dynamic update of process noise covariance matrix Q and measurement noise covariance matrix R) is used to perform real-time filtering on the raw temperature and humidity data, and generate low-noise instantaneous temperature and humidity estimates. Furthermore, the mean-standard deviation extraction method using a sliding time window (window length) is employed. Seconds, step size The filtered data (in seconds) is used to calculate time-series characteristics to obtain the temperature and humidity change rate index. The temperature change rate is calculated using the following formula:

[0011] in, This is the temperature value. This is the current timestamp. Indicates the previous timestamp; Furthermore, through a data integrity verification algorithm (parameter: timestamp continuity threshold) Seconds, data loss tolerance %) Anomalies are removed and missing values ​​are filled in for the rate of change of temperature and humidity and real-time humidity values, forming a complete time-dimensional feature base input; Through the above-mentioned multi-level acquisition, conversion, filtering and feature calculation processing, the original sensor output signal of the previous step is transformed into a stable temperature and humidity feature vector that characterizes the current thermodynamic state of the cooling tower, so as to realize the accurate modeling of the working condition of the reinforcement learning model in the time dimension. For example, in a facility with a rated water treatment capacity of The industrial cooling tower, with a capacity of tons per hour, is equipped with eight digital temperature and humidity sensors, distributed at the top, middle, and air inlet of the tower. The temperature resolution of the sensors is [missing information]. ℃, humidity resolution is %RH; The sampling accuracy of the analog-to-digital converter module is Bits, sampling rate Hz. During Kalman filtering, the initial value of the process noise covariance matrix Q is set to... The initial value of the measurement noise covariance matrix R is set to And adjust the sliding time window length according to the data drift dynamics. Seconds, step size The calculated temperature change rate range is seconds. ℃ / min to ℃ / min, humidity change rate range is %RH / min to %RH / min. After integrity verification, the effective sampling rate of the time dimension feature vector remains at %RH / min. More than 90% of these feature vectors, when input into the reinforcement learning model training module, reduce the response latency of the model's energy-saving strategy to environmental changes. %above; S1.2: Collect load fluctuation information in the cooling tower operating environment, obtain real-time power change data based on the power monitoring module, and construct dynamic load characteristics in the time dimension; When collecting load fluctuation information in the cooling tower operating environment, the input condition is the real-time power signal output by the power monitoring module arranged on the power supply circuit of the fan and water pump. This signal is converted into digital data through the PLC acquisition interface. High-precision power monitoring module adopted (parameter: sampling rate ≥ Hz, power measurement error ≤ %), to achieve synchronous sampling of voltage and current and calculation of instantaneous power; Furthermore, the sliding window averaging method (parameter: window duration) is used. Seconds, step size The instantaneous power data is denoised (in seconds) to obtain a smoothed power curve for analyzing short-period fluctuation characteristics; Furthermore, the power differential calculation method is used to quantify the instantaneous load change and generate the original feature vector of load fluctuation; Furthermore, the Fast Fourier Transform (FFT) algorithm is applied (parameter: number of sampling points). The Hanning window function is used to perform spectral analysis on the load differential sequence, extract the main fluctuation frequency components and their amplitude indices, and is used to characterize the periodic characteristics of load fluctuations. Furthermore, normalization processing (method: Z-score standardization) is used to unify the load fluctuation amplitude and frequency characteristics to a zero mean and unit variance scale, generating a dynamic load feature vector in the time dimension; Through the above algorithm chain, the instantaneous power data of the previous step is transformed into dynamic load features with amplitude and frequency indicators in the time dimension, so as to realize the quantitative expression of load fluctuation in the input space of the reinforcement learning state. For example, in a certain cooling tower system, the sampling rate of the power monitoring module is set to... Hz, instantaneous power measurement error is The raw instantaneous power sequence acquired during runtime is smoothed by a 5-second window with a 1-second step size to obtain the mean of the smoothed power sequence. kW. The power differential calculation yielded the result. Maximum load change within seconds kW, FFT spectrum analysis shows that the main fluctuation frequency is Hz, amplitude kW. After Z-score standardization, the load fluctuation amplitude index is normalized to... The frequency index is normalized to This feature vector is used in the state input of the reinforcement learning model to identify load fluctuation phases, thereby improving the adaptability of the policy under load change conditions; S1.3: Collect temperature sensor data from different areas of the cooling tower body, obtain temperature field distribution information based on a distributed sensor network, and extract thermal field gradient features in the spatial dimension. S1.4: Collect the operating status data of the cooling tower fan, obtain the current speed information of the fan based on the speed encoder, and obtain key execution feedback features in the equipment dimension; S1.5: Collect start-up and shutdown status and energy consumption data of cooling tower water pumps, and obtain water pump operation cycle and energy consumption indicators based on PLC controller and power metering module to supplement the energy consumption characteristics in the equipment dimension.

[0012] Step S2: Normalize the collected multi-source state data and construct a multi-scale state feature vector to characterize the dynamic operating conditions of the cooling tower at different operating stages. Specifically, this includes: S2.1: Perform missing value imputation and outlier removal on the raw state data collected from the cooling tower's operating environment in terms of time, space, and equipment dimensions to obtain a complete and reliable initial state dataset. The time dimension data includes real-time temperature, humidity, and load fluctuation information; the space dimension data includes temperature sensor data from different areas of the tower; and the equipment dimension data includes fan speed, water pump start / stop status, and energy consumption data. For the original status data collected from the cooling tower operating environment in terms of time, space and equipment dimensions, a data integrity detection method (parameters: timestamp continuity threshold 0.5 seconds, data loss rate threshold 2%) is used to detect the missing data streams and collection anomalies in each dimension. Furthermore, the missing temperature and humidity values ​​in the time dimension data are filled in using a linear interpolation algorithm (parameter: interpolation window size is 3 points), and a basic time dimension dataset with continuity is obtained. Furthermore, based on the sample difference test algorithm (parameter: 3σ principle, σ is the sample standard deviation), outlier removal processing is performed on temperature sensor data in different regions of the spatial dimension. Records that deviate from the mean of the current region by more than three times the standard deviation are deleted from the dataset, thereby improving the credibility of the spatial thermal field data. Furthermore, a sliding window smoothing algorithm (parameter: window length is 5 seconds) is used to smooth the fan speed, water pump start-up and shutdown status and energy consumption data of the equipment dimension, so as to reduce the interference of transient spikes on subsequent feature calculations. Furthermore, the data collected from the device dimension is compared across variables using a consistency verification method (parameter: device status logic rule set) to generate a device status consistency identifier, which is used to confirm the reliability and logical rationality of the collected data. Through the above-mentioned multiple algorithm processing methods, the data stream of the previous step is transformed into a complete and reliable initial state dataset, thereby improving the stability of the data structure and the accuracy of subsequent normalization processing. For example, in a cooling tower operation monitoring system, the temperature and humidity sensors in the time dimension sample at a frequency of 1Hz. During the data collection on a certain day, three points of missing temperature and humidity data occurred, with each missing point lasting no more than 2 seconds. A linear interpolation algorithm with a window size of 3 points was used to fill the missing points, resulting in a continuous curve. In the spatial dimension, five temperature sensors were deployed. Among them, sensor number 2 showed five outlier values ​​during data collection, which deviated from the sensor's current data mean by more than three standard deviations. After removal using the 3σ principle, 410 normal samples were retained. In the equipment dimension, the fan speed was sampled once per second and smoothed using a 5-second sliding window to reduce the instantaneous spikes of the original waveform to no more than 5% of the rated value. The pump start / stop status and energy consumption data were checked using logical rules, and one abnormal record that did not meet the start / stop status switching conditions was detected and removed. Finally, an initial state dataset with complete and consistent time, spatial, and equipment dimensions was formed. S2.2: The cleaned multi-source state data is normalized based on the Z-score normalization method to eliminate the impact of the difference in the data units of different sensors on subsequent feature fusion and obtain a normalized state data matrix under a unified scale. S2.3: The sliding time window algorithm is used to extract time-series features from the normalized time dimension data, and the temperature and humidity change rate, load fluctuation mean and standard deviation within the window are calculated to obtain a time-series feature vector characterizing the cooling tower operation trend. S2.4: Based on the spatial correlation analysis method, the temperature distribution features of multiple regions are extracted from the normalized spatial dimension data. The principal component analysis (PCA) algorithm is used to reduce the dimensionality and obtain the low-dimensional spatial state feature vector in the spatial dimension, which is used to characterize the thermal field distribution characteristics inside the cooling tower. S2.5: Perform equipment operation status encoding conversion processing on the normalized equipment dimension data, map the fan speed, water pump start-stop status and energy consumption data into continuous and discrete feature variables respectively, and combine the time series feature vector and spatial state feature vector to generate multi-scale state feature vector, which serves as the state input space of the reinforcement learning model.

[0013] Step S3: Based on the current operation phase identification result, match the current phase label from a preset operation phase label library. The operation phases include the startup phase, stable operation phase, and load fluctuation phase. Figure 2 As shown, it specifically includes: S3.1: Perform sliding window statistical processing on the time dimension features in the multi-scale state feature vector to calculate the temperature and humidity change rate and load fluctuation amplitude within the current time period in order to extract the operating trend features; S3.2: Based on the time series data of the operating status of the fan and water pump, a state transition detection algorithm is used to identify the change points of the equipment operating mode in order to determine whether the cooling tower is in the start-up, stable or load fluctuation stage; S3.3: Perform spatial gradient analysis on the temperature sensor data of multiple regions of the tower in the spatial dimension, calculate the temperature difference between adjacent regions and its changing trend, so as to help identify the spatial thermal distribution characteristics of the current operation stage; S3.4: The time-dimensional trend features, equipment operating status change point information and spatial temperature gradient features are fused to generate a comprehensive operating status feature vector, which is used as the input of the operating phase identification model; The input time-dimensional trend features, equipment operating status change points and spatial temperature gradient features are loaded into the feature fusion buffer respectively. The feature standardization matching algorithm (parameters: feature mean μ, standard deviation σ) is used to achieve comparability processing of features of each dimension under a unified statistical scale. Furthermore, a weighted feature concatenation algorithm (parameter: time feature weights) is used. Equipment feature weights Spatial feature weights This enables vector-level fusion of multi-dimensional features and yields a preliminary comprehensive feature vector. ; Furthermore, a mutual information feature correlation analysis method (parameter: mutual information threshold τ) is employed to measure the correlation of features in each dimension within the comprehensive vector, and to generate a feature importance score matrix. ; Furthermore, by employing a principal component analysis (PCA) dimensionality reduction algorithm (parameter: retaining feature variance contribution rate ≥ 95%), the dimensionality of the comprehensive feature vector is compressed, generating a low-dimensional comprehensive operational state feature vector. ; Through feature normalization fusion processing, the low-dimensional comprehensive feature vector from the previous step is transformed... This is transformed into the input feature set of the recognition model during the runtime phase, thereby improving the recognition accuracy and stability of the model under different input dimensions; For example, in the implementation of a certain cooling tower system, the time-dimensional trend characteristics include the rate of temperature change. for ℃ / min, the points of change in equipment operating status are obtained by the speed change detection algorithm, and the frequency of change is . The spatial temperature gradient characteristics, including the temperature difference between the top and bottom of the tower, are measured per hour. ℃. Standardization process is set as follows: The mean of each feature, Output the standardized value for each standard deviation. Weighting =0.4, =0.35, =0.25, weighted concatenation yields the fused feature vector F. The mutual information threshold τ is set to 0.05. After deleting low-correlation features, PCA is performed with 3 principal components, achieving a cumulative variance contribution rate of 96%. The resulting low-dimensional feature vector F' is then input into the running recognition model. The model's output labels match those in the actual running phase, improving the recognition accuracy to 98%. S3.5: Input the comprehensive operating status feature vector into the pre-trained operating phase classifier, and output the operating phase labels of the startup phase, stable operating phase and load fluctuation phase, so as to trigger and execute the dynamic reward weight allocation strategy. Based on the data input conditions of the comprehensive operational status feature vector, a convolutional neural network is used (parameters: the kernel sizes of the three convolutional layers are respectively...). , , The activation function is Extract spatiotemporal pattern features to achieve joint feature encoding of time trends, spatial gradients, and equipment operation modes; Furthermore, through a Long Short-Term Memory network (parameter: number of hidden units) Time step Temporal dependency modeling is performed on the feature sequence after convolutional encoding, and a high-dimensional state feature tensor containing information on stage changes is obtained; Furthermore, a Softmax classifier is employed (parameter: number of classes). Probability distribution calculations are performed on the high-dimensional state feature tensor to obtain the classification probability vectors of the cooling tower in the startup phase, stable operation phase, and load fluctuation phase. ; Furthermore, the optimal label is determined from the classification probability vector using the maximum likelihood selection algorithm, calculated as follows:

[0014] in, The output is the runtime stage label. This is the classification probability vector; The labels of the running phase are passed to the dynamic reward weight allocation strategy module through the label output interface, so as to trigger the adjustment of the strategy weight of the subsequent reinforcement learning model. For example, a cooling tower system operates in an industrial cooling scenario, and the dimension of the comprehensive operating status feature vector is: The parameters of the three convolutional kernels in the convolutional neural network are set as follows: , , The padding mode is set to "same," and batch normalization is used to stabilize the training process. The convolutional encoding results are fed into the Long Short-Term Memory network, with a hidden unit count of [number missing]. Time step This is to capture the temporal feature signals of the phase transition. In the classification probability vector output by the Softmax classifier, the probability of the initiation phase is... The probability of stable operation is The probability of a load fluctuation phase is The maximum likelihood selection algorithm is used to determine the running stage label. The tag is passed to the S4 dynamic reward weight allocation strategy module to achieve a balanced increase in energy-saving weight and comfort weight during the stable operation phase, thereby improving the overall optimization effect.

[0015] Step S4: Based on the matched operational phase labels and the priority of external energy-saving goals, dynamically adjust the weight coefficients of the three sub-goals—energy saving, stability, and user comfort—in the total reward function to generate a dynamic reward weight allocation strategy. For example... Figure 3 As shown, it specifically includes: S4.1: Based on the operation phase identification results, match the current operation phase label from the preset operation phase label library. The operation phase labels include startup phase, stable operation phase and load fluctuation phase, as the basic input conditions for dynamic weight adjustment. S4.2: Obtain the energy-saving target priority parameters from external input. These parameters are provided by the industrial scheduling system or energy management system and are used to characterize the relative importance of the three optimization targets of energy saving, stability, and user comfort within the current time period. S4.3: Perform logical mapping processing on the operation phase labels and energy-saving target priority parameters to determine the initial weight allocation ratio of the three sub-targets of energy saving, stability and user comfort in the current operating environment; For the matched operation phase labels and external energy-saving target priority parameters, a multi-objective logical mapping algorithm (input parameters: operation phase label set {start-up phase, stable operation phase, load fluctuation phase}, energy-saving target priority triplet) is used to realize the linear-nonlinear coupling calculation of operation phase and target priority. Furthermore, by using a multi-dimensional weight initialization function (parameters: stage category encoding, energy-saving priority numerical vector), the stage category and priority numerical values ​​are combined and mapped to obtain a preliminary multi-objective optimization weight vector. ,in Indicates energy saving weight, Indicates stability weights, Indicates the comfort level weight; Furthermore, the normalized constraint calculation formula is adopted. This ensures that the total weight coefficient of the multi-objective system satisfies the unitization constraint in the numerical space, while maintaining the relative proportional relationship between different objectives. Furthermore, by using a phase-specific adjustment coefficient matrix (parameters: operation phase category code, external target priority vector), the initial weight vector is fine-tuned within the phase, so that the stability weight is increased in the startup phase and the energy-saving weight is appropriately reduced in the load fluctuation phase to meet comfort requirements. Through the above logical mapping and normalization process, the operation phase label and energy-saving target priority parameter are transformed into the initial three-target weight allocation ratio under the current operating environment, so as to realize the dynamic adaptation of weight configuration and operating conditions. For example, during the stable operation phase, an external energy-saving target priority parameter vector is received. The operational phase category code is as follows (Indicating stability), an initial weight vector is generated using a multi-objective logical mapping algorithm. Then a normalization test was performed. The summation constraint is confirmed to be satisfied. Under the influence of the stage-specific adjustment coefficient matrix (the parameter matrix increases the stability weights by 0.05), the adjusted weights are obtained. As the input condition for the subsequent S4.4 fuzzy logic algorithm, in the actual test comparison, this weight ratio improved the stability index by 3% while maintaining the energy-saving effect, and the user comfort did not decrease significantly, which verified the effectiveness of dynamic adaptation. S4.4: Based on the fuzzy logic control algorithm, the initial weight allocation ratio is nonlinearly adjusted to adapt to the uncertainty changes during the transition process of the operation phase, and a dynamic weight adjustment factor is generated; S4.5: The dynamic weight adjustment factor and the historical weight allocation strategy are weighted and fused to generate the final dynamic reward weight allocation strategy, which serves as the input parameter for training and control decisions of the reinforcement learning model. The input conditions include the dynamic weight adjustment factor obtained through the fuzzy logic control algorithm, and the weight allocation strategy sequence in the historical records; A weighted fusion algorithm (parameters: adjustment factor weights, normalized coefficients of historical strategy weights) is used to realize the linear combination calculation of dynamic weight adjustment factors and historical weight strategies. Furthermore, by adjusting the contribution of historical strategies during the fusion process through a time decay factor, the weighting coefficients are corrected according to a time-decreasing function, and the time-weighted fusion result is obtained. Furthermore, the sliding window mean smoothing method (parameter: window length is determined by the duration of the running phase) is used to smooth the fusion weight curve and generate a weight allocation vector without abrupt changes. Furthermore, a normalization constraint is adopted (parameter: the sum of the weights of the three types of sub-objectives is constrained to be...). The smoothed weight vector is proportionally normalized to ensure that the sum of the weights for energy saving, stability, and user comfort remains within a fixed constraint range. Through the above weighted fusion and normalization process, the adjustment factor and historical strategy of the previous step are transformed into the final reward weight allocation strategy that satisfies dynamic adaptation and multi-objective balance, so as to achieve the optimal policy adaptation of the reinforcement learning model under dynamic conditions. For example, when the cooling tower is in a load fluctuation phase, the energy-saving target priority issued by the external energy management system is 0.6, stability is 0.3, and user comfort is 0.1. The dynamic weight adjustment factor output by the fuzzy logic control module is [0.65, 0.25, 0.10], the historical weight allocation strategy is [0.5, 0.4, 0.1], and the weight normalization coefficient is set to... The time decay factor is set to After weighted fusion calculation, the fusion result vector is [0.575, 0.29, 0.135]. Then, normalization is performed, and the normalized result is [0.575, 0.29, 0.135] (the sum in this example is already [0.575, 0.29, 0.135]). The final dynamic reward weight allocation strategy is input into the reinforcement learning model to achieve an optimal balance between energy conservation and real-time policy generation.

[0016] Step S5: Input the multi-scale state feature vector and the dynamic reward weight allocation strategy into the reinforcement learning model, generate wind turbine speed control action suggestions through the policy network, and calculate the corresponding action value function. Specifically, this includes: S5.1: Using multi-scale state feature vectors as the state input of the reinforcement learning model, and constructing a composite reward function with weight factors based on the dynamic reward weight allocation strategy configured according to the current operating stage and energy-saving target priority, as the basis for model training and policy update; S5.2: Perform forward propagation computation on the policy network in the reinforcement learning model, and generate the probability distribution of the fan speed control action based on the input state feature vector and dynamic reward weights, so as to output the optimal control action suggestion; The multi-scale state feature vector and dynamic reward weight allocation strategy are used as inputs to the policy network. The forward propagation algorithm of deep neural network (the structural parameters include the dimension of the input layer corresponding to the dimension of the state feature vector, the activation function of the hidden layer is ReLU, and the activation function of the output layer is Softmax) is used to realize the mapping from state to action probability distribution. Furthermore, the weighted input of the hidden layer is calculated by matrix multiplication of the weight matrix and the input vector. The formula is as follows:

[0017] in This is the first layer weight matrix. For the state input vector, It is the bias vector; Furthermore, the hidden layer state vector is obtained by performing element-wise transformation on the weighted input of the hidden layer using a nonlinear activation function. The formula is as follows:

[0018] Furthermore, through the hidden layer state vector With the output layer weight matrix Perform matrix multiplication and add output layer bias. The unnormalized action values ​​of the output layer are obtained. The formula is as follows:

[0019] Furthermore, the unnormalized action values ​​of the output layer are normalized using the Softmax function to obtain the action probability distribution. ; By weighting the dynamic reward weight and the output action probability, the action with the highest comprehensive score is obtained as the optimal control action suggestion, thereby realizing the dynamic adaptability of the wind turbine speed control strategy. For example, during the stable operation phase of the cooling tower, the input state feature vector has a dimension of 32, and the dynamic reward weight allocation is energy saving: stability: comfort = 0.5:0.3:0.2. The first layer weight matrix of the policy network has a size of 64×32, the bias is a 64-dimensional zero vector, and the activation function is ReLU. After matrix multiplication and activation function processing, the hidden layer vector has a dimension of 64, and the second layer weight matrix has a size of 5×64, corresponding to 5 discrete speed control candidate actions. After Softmax processing, the action probability distribution is [0.05, 0.1, 0.7, 0.1, 0.05]. After adjusting the reward weight, the third action (speed of 900 rpm) has the highest comprehensive score and is recommended as the optimal control action. In actual execution, this action achieved an 8% reduction in energy consumption, temperature fluctuation control within ±0.5℃, and a user comfort score of 95 / 100. S5.3: Calculate the action value function of the generated wind turbine speed control action based on Q network or value network to obtain the expected reward value of each candidate action in the current state, so as to evaluate the comprehensive performance of different actions under multi-objective optimization; S5.4: Based on the evaluation results of the action value function, and in conjunction with the exploration-exploitation strategy (such as the ε-greedy strategy), select the final wind turbine speed control action to balance the relationship between strategy exploration and the current optimal action utilization; S5.5: Encapsulate the selected fan speed control action suggestions into a control command format, and attach the action value assessment results and strategy network output confidence information for subsequent control execution and feedback loop use.

[0020] Step S6: Based on the fan speed control action suggestion output by the reinforcement learning model, a speed adjustment command is sent to the cooling tower fan control system to execute the fan speed adjustment operation. Specifically, this includes: S6.1: Analyze the fan speed control action suggestions output by the reinforcement learning model, extract the target speed value and control execution cycle parameters, and generate a standard control instruction format that is compatible with the cooling tower fan control system; S6.2: Based on the parsed standard control command format, a speed adjustment command is sent to the cooling tower fan frequency converter via the Modbus / TCP protocol. The adjustment command includes a frequency setpoint, acceleration slope, and operating mode switching signal. The parsed standard control command format is encapsulated using the Modbus / TCP protocol (parameter configuration includes TCP port number 502, slave address, and function code 06) to convert the target speed value, acceleration slope, and operating mode switching signal into a data frame structure that conforms to the frequency converter communication protocol. Furthermore, through the instruction field mapping algorithm (parameter mapping table definition range: speed 0~1500rpm; acceleration slope 0~20rpm / s; operating mode 0~2), the physical quantities in the control instructions are converted into register addresses and numerical codes that the frequency converter can recognize, and the register write data blocks are obtained; Furthermore, through the CRC16 checksum algorithm (polynomial) initial value This allows for data integrity verification of the encapsulated Modbus data frame and the generation of a checksum field. Furthermore, processing is performed via TCP socket transmission (setting a transmission timeout). ms, number of retries This enables the verification of data frames to be sent to the network interface of the cooling tower fan frequency converter, and the receipt of execution confirmation response data returned by the controller. Through the Modbus / TCP encapsulation and transmission processing method described above, the speed control command parsed in the previous step is converted into network transmission data that conforms to industrial communication standards and has data integrity protection, so as to realize the remote and controllable issuance of the fan speed adjustment command; For example, in a cooling tower fan control scenario, the reinforcement learning model outputs the target rotational speed value. rpm, acceleration slope The speed is rpm / s, and the operating mode switching signal is mode 1. The instruction format parsing result includes the speed field 900, the slope field 10, and the mode field 01. An instruction field mapping algorithm is used to map the speed to a register address. Value encoding ; Slope mapping to register address Value encoding ; Mode mapping to register address Value encoding The Modbus function code 06 is used to write a single-register data frame. The network layer uses TCP port 502, and the generated unchecked data is 12 bytes long. The checksum field is calculated using the CRC16 checksum algorithm. The data frame was sent via TCP socket and appended to the end of the data frame. Under the 1-second timeout detection and 3-retransmission strategy, the frequency converter successfully returned an execution confirmation response. The response analysis results showed that the target speed had been loaded into the controller parameter area and was in the adjustment process. The actual monitored fan speed steadily increased from the original speed of 700 rpm to 900 rpm within 5 seconds, which is consistent with the control intent output by the reinforcement learning model. S6.3: After receiving the speed adjustment command, the wind turbine frequency converter controller performs closed-loop PID regulation control on the wind turbine drive motor to accurately adjust the wind turbine speed to the target speed value suggested by the reinforcement learning model; S6.4: During the wind turbine speed regulation process, real-time speed, current and voltage signals fed back by the frequency converter are collected to generate wind turbine operating status monitoring data for subsequent performance evaluation and online model fine-tuning; S6.5: The executed wind turbine operation status monitoring data is transmitted to the edge computing node via industrial Ethernet and marked as the control action execution record under the current operation stage, so as to be used for updating the experience playback buffer of the reinforcement learning model.

[0021] Step S7: Collect system feedback data after the fan operation, including actual rotational speed, tower temperature change rate, and energy consumption change, to evaluate the actual effect of the control action. Specifically, this includes: S7.1: Collect the actual speed data of the cooling tower fan after the speed adjustment command is executed, and use the speed sensor and PLC controller communication interface to obtain the real-time speed feedback value, so as to obtain the accurate measurement results of the actual operating status of the fan; S7.2: Based on real-time temperature data collected by distributed temperature sensors inside the tower, calculate the rate of temperature change per unit time, and use a sliding window average filtering algorithm to smooth the temperature data to improve the stability of temperature change trend identification. S7.3: Collect the overall energy consumption change of the cooling tower, obtain the power difference of key equipment such as fans and water pumps before and after the control action based on the electricity meter and data acquisition module, and obtain the energy consumption change per unit time through integral calculation to quantify the energy saving effect; S7.4: The collected actual rotation speed, temperature change rate and energy consumption change data are timestamped and synchronized using linear interpolation to ensure the consistency and comparability of the feedback data in the time dimension. S7.5: Generate a multi-dimensional feedback feature vector based on the synchronized system feedback data, including rotational speed error, temperature response rate, and energy consumption change amplitude, as input for updating the experience replay buffer of the reinforcement learning model to support online learning and policy optimization of the model.

[0022] Step S8: Update the experience replay buffer of the reinforcement learning model based on system feedback data, and fine-tune the model parameters online through the target network to improve the model's adaptability under the current operating conditions. Specifically, this includes: S8.1: Preprocess the collected system feedback data, which includes actual rotational speed, tower temperature change rate and energy consumption change. Remove abnormal data points through a filtering algorithm to obtain clean feedback sample data. S8.2: Bind the preprocessed feedback sample data with the corresponding state-action pairs to generate the experience tuples (state, action, reward, next_state) required for training the reinforcement learning model, so as to build the updated dataset of the experience replay buffer; S8.3: Based on the generated experience tuples, perform the update operation of the experience replay buffer, and use the first-in-first-out (FIFO) strategy to replace the old data in order to maintain the freshness and representativeness of the data in the experience pool, thereby improving the timeliness and working condition matching of the model training samples. S8.4: Based on the updated experience replay buffer, the double-delay deep deterministic policy gradient algorithm (TD3) is used to fine-tune the policy network parameters of the reinforcement learning model online in order to optimize the control policy output accuracy of the model in the current running stage. S8.5: The target network performs a soft update operation on the fine-tuned policy network parameters, and uses an exponential smooth update method to synchronously adjust the target network parameters to improve the stability and convergence speed of model training and enhance the model's generalization ability in dynamic environments.

[0023] Step S9: Determine whether a switch has occurred in the current running phase. If a switch has occurred, regenerate the dynamic reward weight allocation strategy based on the new running phase label and trigger the policy update process of the reinforcement learning model. Specifically, this includes: S9.1: Real-time monitoring of the multi-scale state feature vector of the cooling tower fan system to identify the continuity and stability of the current operating phase. The multi-scale state feature vector includes state information in the time dimension, spatial dimension, and equipment dimension. S9.2: Based on a preset operating phase tag library, a clustering analysis algorithm is used to classify the current multi-scale state feature vector and determine whether it meets the operating phase switching conditions. The operating phases include the startup phase, the stable operating phase, and the load fluctuation phase. S9.3: If a change in the running phase is detected, a new running phase label is matched from the running phase label library and used as the input condition for the dynamic reward weight allocation strategy generation module. S9.4: Based on the matched operation phase labels and the energy-saving target priority parameters input from the outside, a multi-objective optimization algorithm is used to dynamically adjust the weight coefficients of the three sub-objectives of energy saving, stability and user comfort in the total reward function, and generate a dynamic reward weight allocation strategy. S9.5: Input the generated dynamic reward weight allocation strategy and the latest multi-scale state feature vector into the reinforcement learning model to trigger the model's policy update process, so as to adapt to the multi-objective control requirements under the current operation stage.

[0024] Step S10: Based on historical control actions and system feedback data, construct a multi-objective optimization performance evaluation index for subsequent iterative optimization of model training strategies and parameter tuning. Specifically, this includes: S10.1: Perform time alignment processing on the historical control action sequence and the corresponding system feedback data to obtain the causal mapping relationship between control actions and system responses; S10.2: Based on the aligned control action-feedback data pairs, the sliding window algorithm is used to extract the dynamic response features within the control cycle, including the temperature change rate, energy consumption fluctuation coefficient and fan response delay time. S10.3: Based on the extracted dynamic response features, construct a multi-objective performance evaluation dimension, including energy efficiency index, system stability index and control response speed index, which correspond to the three sub-objectives of the reinforcement learning model respectively. S10.4: The normalized weighted summation method is used to comprehensively score the multi-objective performance evaluation dimensions and generate a comprehensive performance evaluation index to quantify the overall control effect of the model in the current operation stage. S10.5: Based on comprehensive performance evaluation metrics and corresponding operational stage labels, establish a performance evaluation dataset and associate it with the training logs of the reinforcement learning model for subsequent policy optimization and hyperparameter tuning.

[0025] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0026] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and rules of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for intelligent control of a cooling tower top fan, characterized in that, Includes the following steps: S1: Collect time, space and equipment status data in the operating environment of the cooling tower. The time dimension includes real-time temperature and humidity and load fluctuation information, the space dimension includes temperature sensor data in different areas of the tower, and the equipment dimension includes fan speed, water pump start and stop status and energy consumption data. S2: Normalize the collected multi-source state data and construct multi-scale state feature vectors; S3: Based on the identification results of the current running stage, match the current stage label from the preset running stage label library; S4: Based on the matched operational phase labels and the priority of external energy-saving goals, dynamically adjust the weight coefficients of the three sub-goals of energy saving, stability and user comfort in the total reward function to generate a dynamic reward weight allocation strategy; S5: Input the multi-scale state feature vector and the dynamic reward weight allocation strategy into the reinforcement learning model to generate wind turbine speed control action suggestions and calculate the corresponding action value function; S6: Based on the aforementioned fan speed control action suggestion, send a speed adjustment command to the cooling tower fan control system to execute the fan speed adjustment operation; S7: Collect system feedback data after the fan is executed, including actual speed, tower temperature change rate and energy consumption change, and evaluate the actual effect of the control action.

2. The intelligent control method for the cooling tower top fan according to claim 1, characterized in that, Following step S7, the following is also included: S8: Update the experience replay buffer of the reinforcement learning model based on system feedback data, and fine-tune the model parameters online through the target network; S9: Determine if the current running phase has switched. If it has switched, regenerate the dynamic reward weight allocation strategy based on the new running phase label and trigger the policy update process of the reinforcement learning model. S10: Construct multi-objective optimization performance evaluation indicators based on historical control actions and system feedback data.

3. The intelligent control method for the cooling tower top fan according to claim 1, characterized in that, Step S1 specifically includes: Real-time temperature and humidity data of the cooling tower operating environment are collected to obtain the current environmental thermodynamic state parameters; Load fluctuation information in the cooling tower operating environment is collected, and real-time power change data is obtained based on the power monitoring module to construct dynamic load characteristics in the time dimension; Temperature sensor data from different areas of the cooling tower are collected, and temperature field distribution information is obtained based on a distributed sensor network to extract thermal field gradient features in the spatial dimension. The operating status data of the cooling tower fan is collected, and the current speed information of the fan is obtained based on the speed encoder. Key execution feedback features in the equipment dimension are obtained. The start-stop status and energy consumption data of the cooling tower water pump are collected. The pump operation cycle and energy consumption indicators are obtained based on the PLC controller and the power metering module to supplement the energy consumption characteristics of the equipment.

4. The intelligent control method for the cooling tower top fan according to claim 3, characterized in that, The real-time temperature and humidity data acquisition uses a digital temperature and humidity sensor with a temperature resolution of 0.1℃ and a humidity resolution of 0.1%RH. The sampling accuracy is 16 bits and the sampling rate is 1Hz after analog-to-digital conversion. The temperature and humidity change rate is calculated using Kalman filtering and sliding time window mean-standard deviation, and data integrity is verified to finally form a time dimension feature vector.

5. The intelligent control method for the cooling tower top fan according to claim 1, characterized in that, Step S2 specifically includes: Missing values ​​were filled and outliers were removed from the original state data collected from the cooling tower's operating environment in terms of time, space, and equipment dimensions to obtain a complete and reliable initial state dataset. The cleaned multi-source state data is normalized using a standardization method to obtain a normalized state data matrix at a unified scale. A sliding time window algorithm is used to extract time-series features from the normalized time dimension data, calculate the temperature and humidity change rate, load fluctuation mean and standard deviation within the window, and obtain a time-series feature vector characterizing the cooling tower's operating trend. Based on spatial correlation analysis, multi-regional temperature distribution features are extracted from normalized spatial dimension data. Principal component analysis is used for dimensionality reduction to obtain low-dimensional spatial state feature vectors in the spatial dimension. The normalized equipment dimension data is processed by equipment operation status encoding conversion, which maps the fan speed, water pump start-stop status and energy consumption data into continuous and discrete feature variables, respectively. The time series feature vector and the spatial state feature vector are then combined to generate a multi-scale state feature vector.

6. The intelligent control method for a cooling tower top fan according to claim 1, characterized in that, Step S3 specifically includes: Sliding window statistical processing is performed on the time dimension features in the multi-scale state feature vector to calculate the temperature and humidity change rate and load fluctuation amplitude in the current time period and extract the time dimension operation trend features. Based on the time series data of the operating status of the fan and water pump, a state transition detection algorithm is used to identify the equipment operating status change points of the equipment operating mode and determine whether the cooling tower is in the start-up, stable or load fluctuation stage. Spatial gradient analysis is performed on the temperature sensor data of multiple regions of the tower in the spatial dimension to calculate the temperature difference between adjacent regions and its changing trend, thereby obtaining the spatial temperature gradient characteristics. The time-dimensional operational trend features, the equipment operational status change point information, and the spatial temperature gradient features are fused to generate a comprehensive operational status feature vector; The comprehensive operational status feature vector is input into a pre-trained operational phase classifier, which outputs operational phase labels.

7. The intelligent control method for a cooling tower top fan according to claim 6, characterized in that, The operation phase labels include the startup phase, the stable operation phase, and the load fluctuation phase.

8. The intelligent control method for a cooling tower top fan according to claim 1, characterized in that, Step S4 specifically includes: Based on the operation phase identification results, the current operation phase label is matched from the preset operation phase label library. The operation phase labels include startup phase, stable operation phase and load fluctuation phase. Obtain energy-saving target priority parameters from external input, which are provided by an industrial scheduling system or an energy management system; Logical mapping is performed between the operation phase labels and the energy-saving target priority parameters to determine the initial weight allocation ratio of the three sub-targets of energy saving, stability and user comfort in the current operating environment; The initial weight allocation ratio is nonlinearly adjusted based on a fuzzy logic control algorithm to generate a dynamic weight adjustment factor. The dynamic weight adjustment factor is weighted and fused with the historical weight allocation strategy to generate the final dynamic reward weight allocation strategy.

9. The intelligent control method for a cooling tower top fan according to claim 1, characterized in that, Step S5 specifically includes: Using multi-scale state feature vectors as the state input of the reinforcement learning model, and based on the dynamic reward weight allocation strategy configured with the current operating stage and energy-saving target priority, a composite reward function with weight factors is constructed. Forward propagation computation is performed on the policy network in the reinforcement learning model to generate the probability distribution of the wind turbine speed control action based on the input state feature vector and dynamic reward weight; The action value function is calculated for the fan speed control action to obtain the expected return value of each candidate action in the current state; Based on the evaluation results of the action value function, the final wind turbine speed control action is selected by combining the exploration-utilization strategy; The selected wind turbine speed control action suggestions are encapsulated into a control command format, along with action value assessment results and policy network output confidence information.

10. The intelligent control method for a cooling tower top fan according to claim 9, characterized in that, In step S5, the policy network adopts a multi-scale feature vector dimension corresponding to the input layer, a ReLU activation function for the hidden layer, and a Softmax activation function for the output layer. The wind turbine speed control action is output through a weighted probability distribution. The action value function is calculated by a Q network or a value network. Finally, the exploration-utilization policy is combined to generate the final adjustment suggestion and issue control commands in a standard command format.