Intelligent water temperature control system and method based on machine learning

Through an intelligent water temperature control system based on machine learning, using time-varying dynamics model and multi-branch LSTM network, the problem of difficulty in achieving precise adjustment in traditional water temperature control methods is solved, and efficient and stable water temperature control and energy optimization are achieved.

CN119620805BActive Publication Date: 2025-06-06SHENZHEN DAYIN MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510149593.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-06
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Traditional water temperature control methods are difficult to achieve accurate and rapid temperature regulation, resulting in unstable energy waste and control performance, and it is difficult to effectively utilize massive real-time data to optimize control strategies.

Method used

Using an intelligent water temperature control system based on machine learning, we use a data on water temperature, flow, pressure and environmental parameter, and build a time-varying dynamic model. The initial water temperature prediction model is trained using a multi-branch long and short-term memory network to detect and predict the system operation abnormality, solve the optimal control strategy, and update the model parameters dynamically.

Benefits of technology

It realizes precise control of intelligent water temperature, improves the performance of the control system, reduces energy consumption, and enhances the adaptability and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119620805B_ABST
    Figure CN119620805B_ABST
Patent Text Reader

Abstract

The present invention relates to an intelligent water temperature control system and method based on machine learning. The method: collects water temperature, flow, pressure and environmental parameter data of a constant water temperature control system to generate a standardized time series data set; constructs a time-varying dynamics model based on the standardized time series data set, and uses a variable forgetting factor generalized subspace tracker to identify parameters to obtain time-varying pseudo-modal temperature parameters; trains a multi-branch long short-term memory network based on the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model; performs system operation anomaly detection and prediction based on the initial water temperature prediction model to obtain anomaly detection and prediction results; solves a control strategy based on the anomaly detection and prediction results to obtain an optimal water temperature control strategy; performs water temperature control based on the optimal water temperature control strategy, and updates the initial water temperature prediction model to obtain a target water temperature prediction model. The implementation of the present invention realizes precise control of intelligent water temperature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to an intelligent water temperature control system and method based on machine learning. Background Art

[0002] The precise water temperature control of the constant water temperature control system not only affects product quality and production efficiency, but is also directly related to energy consumption and user comfort. The traditional water temperature control method mainly relies on the PID controller with fixed parameters. This method is often difficult to achieve accurate and fast temperature regulation when facing complex and changing environmental and load conditions, resulting in energy waste and unstable control performance.

[0003] With the rapid development of Internet of Things and sensor technology, water temperature control systems can obtain more dimensional and higher precision real-time data. However, how to effectively use these massive data to optimize control strategies has become an urgent problem to be solved. Traditional data analysis methods are difficult to fully explore the deep information contained in the data and cannot adapt to the dynamic changes and nonlinear characteristics of the system, which limits the performance improvement space of the control system. Summary of the invention

[0004] The main purpose of the present invention is to provide an intelligent water temperature control system and method based on machine learning to achieve precise control of intelligent water temperature.

[0005] To achieve the above object, the present invention provides an intelligent water temperature control method based on machine learning, comprising the following steps:

[0006] Collect water temperature, flow, pressure and environmental parameter data of the constant water temperature control system to generate a standardized time series data set;

[0007] A time-varying dynamics model is constructed based on the standardized time series data set, and a variable forgetting factor generalized subspace tracker is used to perform parameter identification to obtain time-varying pseudo-modal temperature parameters;

[0008] Training a multi-branch long short-term memory network according to the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model;

[0009] Performing system operation abnormality detection and prediction based on the initial water temperature prediction model to obtain abnormality detection and prediction results;

[0010] Solving the control strategy based on the abnormality detection and prediction results to obtain the optimal water temperature control strategy;

[0011] Water temperature control is performed based on the optimal water temperature control strategy, and the initial water temperature prediction model is updated to obtain a target water temperature prediction model.

[0012] The present invention also provides an intelligent water temperature control system based on machine learning, comprising:

[0013] The acquisition module is used to collect the water temperature, flow, pressure and environmental parameter data of the constant water temperature control system and generate a standardized time series data set;

[0014] A construction module is used to construct a time-varying dynamic model based on the standardized time series data set, and use a variable forgetting factor generalized subspace tracker to perform parameter identification to obtain time-varying pseudo-modal temperature parameters;

[0015] A training module, used for training a multi-branch long short-term memory network according to the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model;

[0016] A detection module, used to perform system operation anomaly detection and prediction based on the initial water temperature prediction model to obtain anomaly detection and prediction results;

[0017] A solution module, used to solve the control strategy based on the abnormality detection and prediction results to obtain the optimal water temperature control strategy;

[0018] A control module is used to perform water temperature control based on the optimal water temperature control strategy and update the initial water temperature prediction model to obtain a target water temperature prediction model.

[0019] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.

[0020] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above-mentioned methods are implemented.

[0021] In summary, the technical solution provided by the present invention adopts wavelet transform and local anomaly factor algorithm to effectively remove noise and outliers and improve data quality. Multivariate interpolation processing technology completes missing data to ensure the continuity and integrity of data. Normalization and time window division processing make the data more suitable for the training of machine learning models. Based on time-varying dynamics model and variable forgetting factor generalized subspace tracker, real-time identification of system parameters is realized. Through Markov superposition and recursive parameter identification, the nonlinear and time-varying characteristics of the system are captured. Dynamic adjustment of variable forgetting factor improves the adaptability of the model to system changes. Multi-branch LSTM network structure processes water temperature, flow, pressure and environmental parameters respectively to improve the pertinence of feature extraction. Introduce attention mechanism and residual connection to enhance the ability of the model to capture long-term dependencies. Differential overfitting mitigation and regularization technology are used to improve the generalization ability of the model. Combined with Mahalanobis distance and local anomaly factor algorithm, the sensitivity and accuracy of anomaly detection are improved. The threshold is dynamically adjusted using time-varying pseudo-modal temperature parameters to adapt to the dynamic changes of the system. Based on Markov decision process and deep reinforcement learning, adaptive optimization of control strategy is achieved. Trust region strategy optimization algorithm is used to improve the stability and convergence speed of strategy update. The robustness of control strategy is enhanced through experience replay and exploratory sampling. Sliding time window and online feature reconstruction are used to achieve dynamic update of model parameters. Local update of LSTM layer parameters improves the model's adaptability to the latest data while maintaining the memory of historical information, thereby achieving precise control of intelligent water temperature. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the steps of an intelligent water temperature control method based on machine learning in one embodiment of the present invention;

[0023] Figure 2 It is a structural block diagram of an intelligent water temperature control system based on machine learning in one embodiment of the present invention;

[0024] Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0025] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0027] Reference Figure 1 , this embodiment provides an intelligent water temperature control method based on machine learning, comprising the following steps:

[0028] S1, collects water temperature, flow, pressure and environmental parameter data of the constant water temperature control system to generate a standardized time series data set;

[0029] Specifically, multiple sensors are installed on the water inlet, outlet, heating unit and water storage device of the constant water temperature control system. These sensors are used to collect key parameters such as water temperature, flow rate, pressure and ambient temperature during the operation of the system. Through these sensors, the original sensor data is obtained, which contains multi-dimensional information of the water temperature control system at different time points. The original data is processed by wavelet transform. Wavelet transform is an effective signal processing method that can decompose data into frequency components of different scales, realize denoising, and obtain denoised data. Outlier detection is performed on the denoised data. The density-based Local Outlier Factor (LOF) algorithm is adopted. This algorithm determines whether it is an outlier by calculating the local density of each data point. If a data point has a lower density than other points in its neighborhood, it is marked as an outlier. After being processed by the LOF algorithm, an outlier marking result is obtained, which identifies which data points are statistically abnormal. Based on the outlier marking result, the denoised data is processed by outlier removal to obtain data after outliers are removed. Multivariate interpolation is performed on the missing values ​​in the data after outliers are removed, and the missing values ​​are inferred by using other relevant data to obtain the completed data. The completed data is normalized to convert data of different dimensions to the same scale range, usually between 0 and 1, so as to avoid the weight imbalance problem caused by different data scales during model training. The normalized data is divided into time windows, and the continuous time series data is divided into multiple subsequences according to the set window length. Key features are extracted in each time window to generate a standardized time series data set.

[0030] S2, a time-varying dynamic model is constructed based on the standardized time series data set, and a variable forgetting factor generalized subspace tracker is used for parameter identification to obtain the time-varying pseudo-modal temperature parameters;

[0031] Specifically, the standardized time series data set is decomposed into subsystems. The entire water temperature control system is divided into multiple subsystems with different physical characteristics and functions, and each subsystem corresponds to an independent data set. Through this decomposition method, the thermodynamic characteristics of different subsystems are captured respectively, and a thermodynamic model is constructed on this basis. The model is specifically expressed as a time-varying temperature-flow coupled dynamic equation, which describes the dynamic relationship between water temperature and flow. The coupled dynamic equation is Markov superposed to obtain a component matrix for parameter identification. Through Markov superposition, the complex dynamic system is decomposed into multiple Markov states. Based on the obtained component matrix, the mean square error of the system is calculated, and the initial mean square error value reflects the stability and error distribution of the system under the current conditions. According to the calculated initial mean square error value, the initial value of the variable forgetting factor is set. The variable forgetting factor is a parameter used to adjust the model weight. By applying the initial variable forgetting factor to the standardized time series data set, the generalized subspace tracking algorithm is used to perform the initial subspace estimation. The subspace tracking algorithm is used to update the model parameters in real time in a changing environment, so that the model can reflect the dynamic changes of the system in a timely manner. Based on the initial subspace estimation results and the initial variable forgetting factor, recursive parameter identification is performed to obtain the first round of identification results, which represent the characteristics of the current system under time-varying conditions. According to the first round of identification results, the mean square error of the system is recalculated to obtain an updated mean square error value. Based on the updated mean square error value, the variable forgetting factor is dynamically adjusted to obtain an optimized variable forgetting factor. By dynamically adjusting the forgetting factor, the model can adapt to system changes more sensitively and improve the accuracy of parameter identification. The optimized variable forgetting factor is used to iteratively optimize the first round of identification results. After multiple optimization cycles, the final time-varying pseudo-modal temperature parameters are obtained.

[0032] The initial mean square error value is exponentially smoothed to weaken the random fluctuations in the data, retain its overall trend, and obtain a more stable mean square error sequence. Based on the smoothed mean square error sequence, the initial value of the variable forgetting factor is calculated. The initial variable forgetting factor is one of the key parameters of the generalized subspace tracking algorithm and directly affects the accuracy of the subsequent subspace estimation. The standardized time series data set is subjected to singular value decomposition to decompose the data set into a left singular matrix, a singular value matrix, and a right singular matrix. The left singular matrix describes the row-wise characteristics of the data set, while the singular value matrix provides the eigenvalue information of the data set. By combining the left singular matrix and the singular value matrix, an initial observation matrix is ​​constructed, which is used to capture the main dynamic features in the time series data set. The initial observation matrix is ​​subjected to QR decomposition to decompose a matrix into an orthogonal matrix and an upper triangular matrix. The column vectors of the orthogonal matrix are orthogonal to each other and have important mathematical properties, while the upper triangular matrix contains the main linear combination coefficients. Based on these two matrices, the generalized subspace projection matrix is ​​calculated, which is used to map high-dimensional data into a lower-dimensional subspace while retaining the main features of the data. Recursive least squares estimation is performed on the standardized time series data set based on the initial variable forgetting factor and the initial projection matrix. Recursive least squares estimation is a step-by-step optimization algorithm that can update model parameters in real time when new data arrives to improve the accuracy of the estimation. While performing the initial state estimation, the initial subspace model is constructed based on the obtained estimation results and the initial projection matrix. The initial subspace model represents the main dynamic characteristics of the data. The initial subspace estimate is obtained.

[0033] S3, training a multi-branch long short-term memory network based on the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model;

[0034] Specifically, the standardized time series data set is divided into time steps, and the data set is converted into an input sequence with a fixed time window. The input sequence is feature expanded based on the time-varying pseudo-modal temperature parameter to increase the input information of the model and obtain a richer expanded feature sequence, including multiple data types such as water temperature, flow, pressure and environmental parameters. According to the data type of the expanded feature sequence, the expanded feature sequence is divided into four sub-networks, corresponding to the water temperature branch LSTM sub-network, the flow branch LSTM sub-network, the pressure branch LSTM sub-network and the environmental parameter branch LSTM sub-network. Each branch contains an independent LSTM sub-network to perform special processing for different types of data. In the construction of the water temperature branch LSTM sub-network, the first input layer, the first LSTM layer and the first output layer are designed. The first LSTM layer contains 64 neurons and uses the tanh activation function to ensure that the neuron output is between -1 and 1. In order to prevent overfitting, Dropout regularization is introduced between the first input layer and the first LSTM layer. The regularization technology improves the generalization ability of the model by randomly discarding the output of some neurons. In the construction of the traffic branch LSTM subnetwork, the second input layer, the second LSTM layer and the second output layer are designed. The second LSTM layer contains 32 neurons and uses the ReLU activation function. The selection of the activation function makes the neuron output more sparse, thereby enhancing the nonlinear expression ability of the model. A batch normalization layer is added to the second LSTM layer to speed up the convergence of the model and reduce the impact of internal covariate shift. In the construction of the pressure branch LSTM subnetwork, the third input layer, two stacked third LSTM layers and the third output layer are designed. Each third LSTM layer contains 48 neurons and uses the Leaky ReLU activation function. This activation function can maintain a certain output when the input is less than zero, thereby avoiding the "death" of neurons. In order to enhance the deep learning ability of the network, residual connections are introduced between the two stacked LSTM layers, so that the network can still effectively transmit information while the depth increases, preventing the occurrence of the gradient vanishing problem. In the environmental parameter branch LSTM subnetwork, the fourth input layer, the fourth LSTM layer, the attention layer and the fourth output layer are constructed. The fourth LSTM layer contains 40 neurons and uses the ELU activation function. The choice of ELU activation function helps to accelerate the learning process of the neural network and reduce the impact of offset. The attention layer uses a self-attention mechanism to enhance the model's feature representation ability by calculating the correlation between different parts of the input sequence, ensuring that the model can focus on important temporal features. After splicing the outputs of these four branches, a fused feature vector is obtained. The fused feature vector is fully connected to obtain a comprehensive feature representation.This feature representation is used to construct a bidirectional LSTM layer, which contains 128 neurons and uses the tanh activation function. The bidirectional LSTM layer processes the input sequence in the forward and reverse directions, so that the model can capture the dependencies between the sequences at the same time. In order to enhance the ability to capture time series features, a jump connection is added between the forward and reverse LSTMs to ensure the effective transmission of information. In order to improve the attention mechanism of the model, a multi-head self-attention mechanism is applied to the time series dependency features. Eight attention heads are set, and the dimension of each head is 16. The multi-head mechanism can focus on different features from multiple angles at the same time to obtain attention weighted features. The weighted features are input into the fully connected layer to obtain the predicted output. During the training of the model, the mean square error is used as the loss function, and the Adam optimizer is selected for end-to-end training. The initial value of the learning rate is set to 0.001, and the cosine annealing strategy is introduced to dynamically adjust the learning rate to ensure that the model can gradually converge during the training process to obtain the initial water temperature prediction model.

[0035] The error between the model's predicted output and the actual water temperature value is calculated. The mean square error is calculated to quantify the gap between the model's predicted value and the actual value. The mean square error loss value can provide a clear optimization direction for the model training process. Based on the mean square error loss value, the gradient of the network parameters is calculated to obtain the initial gradient value, which guides the model to gradually approach the optimal solution. The network parameters are updated using the Adam optimizer. The Adam optimizer is an optimization algorithm that combines momentum and adaptive learning rate. It can accelerate the convergence speed of the model and improve the robustness of the model. When using the Adam optimizer for parameter update, the momentum parameter is set to 0.9. This parameter controls the degree of gradient accumulation, so that the update direction depends not only on the current gradient, but also considers the past gradient history. The exponential decay rate of the second-order moment estimate is set to 0.999 to ensure better adaptation to different gradient changes during the gradient estimation process. In order to improve the numerical stability of the calculation, a very small constant (set to 1e-8) is introduced in the Adam optimizer to prevent the denominator from being zero. With these settings, the network parameters after the first round of updates are obtained. As the training progresses, in order to prevent the model from falling into the local optimal solution or overfitting problems under a fixed learning rate, the cosine annealing strategy is applied to the training rounds to dynamically adjust the learning rate. This strategy can adjust the learning rate at different stages of training, so that the model updates the parameters with a larger step in the early stage, and gradually reduces the step in the later stage to fine-tune the parameters. By calculating the current learning rate, the dynamically adjusted learning rate is obtained, and the forward propagation calculation is performed based on this to generate a new iterative prediction output. At this time, the difference between the iterative prediction output and the actual water temperature value is calculated again to obtain the differential error. In order to improve the generalization ability and robustness of the model, the differential error is subjected to differential overfitting mitigation. By introducing the regularization term, the overly complex feature learning of the model is suppressed to prevent the model from overfitting the training data. Through the regularized loss function, the model can maintain a balance between a good fit to the training data and the ability to predict new data. Based on the regularized loss function, the back propagation algorithm is used to calculate the new gradient, and the network parameters are updated in combination with the momentum and adaptive learning rate mechanism of the Adam optimizer. This process is repeated, and as the training progresses, the model gradually adjusts and optimizes its internal parameters, eventually obtaining a fully trained initial water temperature prediction model.

[0036] S4, performing system operation anomaly detection and prediction based on the initial water temperature prediction model to obtain anomaly detection and prediction results;

[0037] Specifically, the target operation data of the constant water temperature control system is preprocessed to eliminate the noise and dimensional differences in the data and ensure the quality and consistency of the data. Through standardization, the target time series data of different dimensions are converted to the same scale to obtain standardized target time series data. The processed time series data is input into the initial water temperature prediction model. In this model, the water temperature, flow, pressure and environmental parameter data are respectively input into the corresponding LSTM branch network to generate their own characteristic subsequences. The water temperature data is input into the water temperature branch LSTM subnetwork to generate a water temperature characteristic subsequence; the flow data is input into the flow branch LSTM subnetwork to generate a flow characteristic subsequence; the pressure data is input into the pressure branch LSTM subnetwork to generate a pressure characteristic subsequence; the environmental parameter data is input into the environmental parameter branch LSTM subnetwork to generate an environmental parameter characteristic subsequence. The water temperature characteristic subsequence is input into the first input layer of the water temperature branch LSTM subnetwork. The subnetwork processes the water temperature data through a 64-neuron LSTM layer. The memory unit of the LSTM layer can capture the time dependency in the data, and the output is processed nonlinearly through the tanh activation function, so that the generated water temperature feature vector can reflect the complex time series characteristics. Similarly, the flow feature subsequence is input into the second input layer of the flow branch LSTM subnetwork. After being processed by the 32-neuron LSTM layer, the ReLU activation function is used to activate the output, so that the network can effectively learn nonlinear features. In order to ensure the stability of the model and accelerate the convergence speed, batch normalization is added after the ReLU activation function to improve the quality of the flow feature vector. The pressure feature subsequence is input into the third input layer of the pressure branch LSTM subnetwork. The network consists of two stacked LSTM layers, each containing 48 neurons. Through the deep stacking structure, the network can better learn higher-level features. The leaky ReLU activation function is used to process the inter-layer output to avoid the "neuron death" problem, and the information is transmitted between layers through the residual connection to obtain the pressure feature vector. The environmental parameter feature subsequence is input into the fourth input layer of the environmental parameter branch LSTM subnetwork, and processed by the 40-neuron LSTM layer. The ELU activation function is used to activate the output to improve the performance of the model when processing environmental data. The environmental parameter branch introduces a self-attention mechanism, which processes data at different time steps in a weighted manner to ensure that the model can focus on the time points that are most useful for prediction and generate environmental parameter feature vectors. The water temperature feature vector, flow feature vector, pressure feature vector and environmental parameter feature vector are concatenated and processed through a fully connected layer to obtain the target feature representation. The target feature representation is input into the bidirectional LSTM layer. The bidirectional LSTM layer contains 128 neurons, which process the input sequence forward and backward respectively, capture the bidirectional time dependency in the sequence, and retain important feature information between layers through skip connections to generate target time series features.In order to improve the accuracy of the model, an 8-head self-attention mechanism is applied to the target time series features. The dimension of each attention head is set to 16. The multi-head mechanism can weight the data from multiple angles to ensure that the model captures important features more comprehensively and accurately. The attention weighted features are processed by the fully connected layer to obtain the final water temperature prediction value. After completing the water temperature prediction, the prediction performance of the model is evaluated by calculating the error between the predicted water temperature value and the actual water temperature value. In the anomaly detection process, the prediction error is analyzed using the Mahalanobis distance. The Mahalanobis distance can quantify the degree of abnormality of the error in multidimensional space and help identify potential anomalies. At the same time, combined with the local anomaly factor algorithm, the anomaly points are confirmed by comparing the local density of each data point. This process can effectively identify isolated abnormal data. In order to better adapt to the dynamic changes of the system, the threshold of anomaly detection is dynamically adjusted in combination with the time-varying pseudo-modal temperature parameter to make the detection results more accurate. Generate anomaly detection and prediction results of system operation.

[0038] S5, solving the control strategy based on the abnormality detection and prediction results to obtain the optimal water temperature control strategy;

[0039] Specifically, the state space modeling of abnormal detection and prediction results is carried out. Water temperature, flow, pressure and environmental parameters are taken as state variables, and heating power and flow regulation are taken as control variables. By introducing these variables into the state space, a Markov decision process model is constructed. The core of the Markov decision process model lies in its state transfer mechanism and reward mechanism, which describe the state change of the system after taking a certain control action and the immediate reward that can be obtained. Based on the constructed Markov decision process model, two important neural networks are established: the value function network and the policy network. The value function network is used to estimate the expected cumulative reward in a certain state, while the policy network is used to generate the control action to be taken in the current state. In terms of design, the value function network adopts a three-layer feedforward neural network structure, which is suitable for processing nonlinear relationships in continuous state space. The policy network adopts a Gaussian policy model, which can better handle the problem of continuous action space by generating control actions that conform to Gaussian distribution. The parameters of the initial neural network structure are initialized. The weights of the neural network are initialized using the Xavier method, and the stability of the network is maintained by controlling the variance of the weights to prevent the problem of gradient disappearance or explosion. At the same time, the bias of the network is initialized to zero, which helps to simplify the learning process of the network. Based on the initialized value function network and policy network, the current state is sampled, and the policy network is used to generate control actions. The control action is input into the environment model for simulation and execution to obtain the next state of the system and the immediate reward. The immediate reward is usually based on the control goal setting, such as maintaining the water temperature within the target range or minimizing energy consumption. Using this information, the target value of the value function is calculated using the temporal difference learning method. The temporal difference learning method can gradually approach the true value function when the state and reward are not completely known. Through the back propagation algorithm, the parameters of the value function network are updated to more accurately reflect the expected rewards under different states. Based on the updated value function estimate, the policy gradient method is used to calculate the gradient of the policy network, and the trust region policy optimization algorithm is used to update the parameters of the policy network. The trust region policy optimization algorithm is an optimization method that can ensure the stability of the policy update process. By limiting the step size of each update, it ensures that the policy will not deviate too far during the optimization process. When exploring new strategies, try to avoid destroying existing good strategies, gradually improve the control strategy, and obtain an improved control strategy. In order to improve the quality of the control strategy, exploratory sampling is performed on the improved control strategy. Control actions are executed in multiple time steps, the changes in system states are observed, and the state transition information is stored in the experience replay buffer. The experience replay buffer can save transfer samples under multiple different states. These samples will be randomly extracted in subsequent training to avoid the influence of correlation between samples on the training process. By repeatedly extracting batch samples, the network parameters are continuously optimized until the performance of the policy network converges or reaches the preset number of iterations, forming the optimal water temperature control strategy.

[0040] S6, executing water temperature control based on the optimal water temperature control strategy, and updating the initial water temperature prediction model to obtain a target water temperature prediction model.

[0041] Specifically, the optimal water temperature control strategy is discretized, and the continuous control actions are mapped to a limited set of control instructions to obtain a discrete control instruction sequence. Based on the discrete control instruction sequence, the heating power and flow of the constant water temperature control system are adjusted in real time to achieve precise control of the water temperature. In the process of executing the control instructions, the feedback data of water temperature, flow, pressure and environmental parameters are collected in real time to obtain a real-time feedback data set. The real-time feedback data set is processed by sliding time window, and a sliding window data subset is formed by extracting the data samples of the latest N time steps to capture the most recent dynamic changes and ensure that the model update can reflect the current system status in a timely manner. Based on the sliding window data subset, the initial water temperature prediction model is reconstructed online. The process of online feature reconstruction involves dynamic weight adjustment of water temperature, flow, pressure and environmental parameter features. The model re-evaluates and adjusts the influence of different features according to the current data characteristics to ensure that the features of the model input can more accurately reflect the current system status. Through dynamic weight adjustment, the model can respond to various changes more flexibly and improve its prediction accuracy. The parameters of the initial water temperature prediction model are locally updated. The update process includes adjusting the weight matrix and bias vector of the LSTM layer. By locally updating these parameters, the model can better adapt to new data and gradually improve its prediction performance. In this process, each local update will gradually make the model approach the target water temperature prediction model. Through multiple iterations and online updates, the target water temperature prediction model is obtained.

[0042] In one example, water temperature, flow, pressure, and environmental parameter data of a constant water temperature control system are collected to generate a standardized time series data set, including:

[0043] Install multiple sensors on the water inlet, water outlet, heating unit and water storage device of the constant water temperature control system;

[0044] Collect water temperature, flow, pressure and ambient temperature data based on multiple sensors to obtain raw sensor data;

[0045] Perform wavelet transform on the original sensor data to obtain denoised data, and perform outlier detection on the denoised data based on the density-based local anomaly factor algorithm to obtain outlier marking results;

[0046] Based on the outlier marking results, outliers in the denoised data are removed to obtain data after outlier removal, and multivariate interpolation processing is performed on the missing values ​​in the data after outlier removal to obtain completed data;

[0047] The completed data is normalized to obtain normalized data, and the normalized data is divided into time windows and features are extracted to obtain a standardized time series data set.

[0048] In this example, multiple sensors are installed on the water inlet, outlet, heating unit and water storage device of the constant water temperature control system to collect the system's operating data, including multiple parameters such as water temperature, flow, pressure and ambient temperature, and obtain the original sensor data. The original sensor data is processed by wavelet transform to decompose the original data into different frequency components, effectively removing high-frequency noise. The local anomaly factor algorithm of density is used to detect outliers on the denoised data. The local anomaly factor algorithm determines whether each data point is an outlier by calculating the density of each data point relative to its neighborhood data points. If the density of a data point is significantly lower than other points in its neighborhood, the point is marked as an outlier. The formula is as follows:

[0049] ;

[0050] in, Represents data points The local anomaly factor, It is a data point of Neighbor set, It is a data point The local reachability density of reflects the density relationship between the point and its neighboring points. Significantly greater than 1, it indicates It may be an outlier. Based on the outlier labeling results, the denoised data is processed to remove the outliers from the data set. Multivariate interpolation is performed on the data after outliers are removed. Multivariate interpolation is a method of using existing data points to infer the value of missing data points. By analyzing the distribution of adjacent data points, reasonable missing value completion results are inferred. The completed data is normalized to convert data of different dimensions to the same scale, usually mapping the data to the interval of [0, 1] or [-1, 1]. Normalization can eliminate the dimensional differences between different features, so that each feature contributes more evenly to the results during model training, and avoids some features dominating the learning process of the entire model due to their large values. Time window division and feature extraction are performed on the normalized data. Time window division refers to dividing continuous time series data into multiple subsequences according to a fixed time step, and each subsequence corresponds to data within a period of time. In this way, the time dependency and change trend in the data can be captured, so that the model can better understand the dynamic characteristics of the data. Feature extraction is to extract key statistical features in each time window, such as mean, standard deviation, maximum value, minimum value, etc., so that the model can learn the potential patterns in the data and obtain a standardized time series data set.

[0051] In one example, a time-varying dynamics model is constructed based on a standardized time series data set, and a variable forgetting factor generalized subspace tracker is used for parameter identification to obtain time-varying pseudo-modal temperature parameters, including:

[0052] The standardized time series data set is decomposed into subsystems to obtain multiple subsystem data sets, and a thermodynamic model is constructed based on the multiple subsystem data sets to obtain a time-varying temperature-flow coupling kinetic equation;

[0053] The time-varying temperature-flow coupling dynamic equation is subjected to Markov superposition to obtain a component matrix, and the system mean square error is calculated based on the component matrix to obtain an initial mean square error value;

[0054] The initial value of the variable forgetting factor is set according to the initial mean square error value to obtain the initial variable forgetting factor, and the generalized subspace tracking is performed on the standardized time series data set to obtain the initial subspace estimation;

[0055] Recursive parameter identification is performed based on the initial subspace estimation and the initial variable forgetting factor to obtain the first round of identification results, and the system mean square error is updated according to the first round of identification results to obtain an updated mean square error value;

[0056] The variable forgetting factor is dynamically adjusted based on the updated mean square error value to obtain the optimized variable forgetting factor. The first round of identification results are iteratively optimized based on the optimized variable forgetting factor to obtain the time-varying pseudo-modal temperature parameters.

[0057] In this example, the standardized time series data set is subsystem decomposed to split the data of the entire system into multiple subsystem data sets, each of which corresponds to a different part or functional module of the system. Through subsystem decomposition, the characteristics and dynamic behaviors of each subsystem are captured separately. For example, in a constant water temperature control system, the water inlet system, heating system, circulation system, etc. are regarded as independent subsystems, and the time series data of their respective water temperature, flow, pressure and other environmental parameters are separated. Based on the subsystem data sets, the respective thermodynamic models are constructed. The thermodynamic model describes the energy conversion and transfer process within each subsystem, especially the coupling relationship between water temperature and flow. The coupling relationship is usually nonlinear and changes with time. By combining the thermodynamic models of the subsystems, the overall time-varying temperature-flow coupling dynamic equation is obtained. This equation is used to describe the state of the system at any time. Its form is usually a partial differential equation or an ordinary differential equation, reflecting the complex interaction between water temperature and flow. The time-varying temperature-flow coupling dynamic equation is subjected to Markov superposition processing, and the state variables in the dynamic equation are converted into Markov states to describe the time evolution process of the system. Through Markov superposition, complex dynamic systems are decomposed into transition processes between multiple states, which are represented in the form of component matrices. The component matrix contains the behavior information of the system in each state. Based on the component matrix, the mean square error of the system is calculated. The mean square error value can be calculated by the following formula:

[0058] ;

[0059] in, Indicates The actual observed values, represents the corresponding predicted value, Represents the number of samples. The smaller the mean square error value, the more accurate the prediction model of the system is and the higher the stability is. Based on the initial mean square error value, set the initial value of the variable forgetting factor. The variable forgetting factor is a key parameter in the recursive algorithm, which is used to control the weight of the influence of new and old data on the model. When the initial mean square error value is large, it may be necessary to set a smaller variable forgetting factor so that the model can adapt to new data faster; conversely, if the mean square error value is small, a larger variable forgetting factor can be set to maintain the stability of the model. After determining the variable forgetting factor, generalized subspace tracking is performed on the standardized time series data set. Generalized subspace tracking is an online algorithm that can update the model in real time as data gradually arrives and track the changes in the subspace. Through generalized subspace tracking, the initial subspace estimate is obtained, which reflects the main dynamic characteristics of the data set. Recursive parameter identification is performed based on the initial subspace estimate and the initial variable forgetting factor. Recursive parameter identification is an iterative optimization process that uses new data to gradually adjust model parameters to improve the model's fit to the system dynamics. In the first round of recursive parameter identification, the model generates preliminary identification results, which represent the current state parameters of the system. Based on this identification result, the mean square error value of the system is updated. As the mean square error value is updated, the variable forgetting factor is dynamically adjusted so that the model can better adapt to new data changes in the subsequent identification process. The optimized variable forgetting factor is used to iteratively optimize the first round of identification results. This iterative optimization will gradually approach the real dynamics of the system and obtain the final time-varying pseudo-modal temperature parameters.

[0060] In one example, an initial value of a variable forgetting factor is set according to an initial mean square error value to obtain an initial variable forgetting factor, and generalized subspace tracking is performed on a standardized time series data set to obtain an initial subspace estimate, including:

[0061] Performing exponential smoothing on the initial mean square error value to obtain a smoothed mean square error sequence, and calculating the initial value of the variable forgetting factor based on the smoothed mean square error sequence to obtain an initial variable forgetting factor;

[0062] Perform singular value decomposition on the standardized time series data set to obtain a left singular matrix, a singular value matrix, and a right singular matrix, and construct an observation matrix based on the left singular matrix and the singular value matrix to obtain an initial observation matrix;

[0063] Perform QR decomposition on the initial observation matrix to obtain an orthogonal matrix and an upper triangular matrix, and calculate the generalized subspace projection matrix based on the orthogonal matrix and the upper triangular matrix to obtain the initial projection matrix;

[0064] Recursive least squares estimation is performed on the standardized time series data set based on the initial variable forgetting factor and the initial projection matrix to obtain an initial state estimate, and an initial subspace model is constructed based on the initial state estimate and the initial projection matrix to obtain an initial subspace estimate.

[0065] In this example, the initial mean square error value is exponentially smoothed to reduce fluctuations in the data and highlight the overall trend. For the initial mean square error value, exponential smoothing can better eliminate the impact of accidental fluctuations by decreasing the weight, so that the smoothed mean square error sequence can better represent the actual change trend of the system. The calculation formula for exponential smoothing is as follows:

[0066] ;

[0067] in, Indicates at time The smoothed mean square error value at time, It's in time The initial mean square error value at time, is the smoothing coefficient, and its value range is , which controls the weight of new and old data. Through the above formula, a smoothed mean square error sequence is generated. Based on the smoothed mean square error sequence, the initial value of the variable forgetting factor is calculated. The variable forgetting factor is a key parameter in the recursive algorithm, which is used to balance the impact of new and old data at each update. The smoothed mean square error value can reflect the degree of change of the system. If the mean square error value is large, it means that the system changes more drastically, so a smaller variable forgetting factor needs to be set so that the model can respond quickly to new changes; conversely, when the mean square error value is small, it means that the system is relatively stable, and a larger variable forgetting factor can be set to maintain the stability of the model and the weight of historical data. The initial variable forgetting factor is obtained by analyzing the mean square error value. Perform singular value decomposition on the standardized time series data set and decompose the data matrix into three parts: left singular matrix, singular value matrix and right singular matrix. Assume that the data matrix is , its singular value decomposition is expressed as:

[0068] ;

[0069] in, is a left singular matrix, containing the main directional characteristics of the data; is a singular value matrix containing diagonal elements that represent the magnitude of the eigenvalues ​​in each direction; It is a right singular matrix, which reflects the projection of the original data in the direction of each singular value. Through singular value decomposition, high-dimensional data can be effectively reduced in dimension and the most representative features in the data can be extracted. and the singular value matrix , construct the observation matrix, which retains the most important information in the data set and forms the initial observation matrix. Perform QR decomposition on the initial observation matrix. QR decomposition is an algorithm that decomposes a matrix into an orthogonal matrix and an upper triangular matrix. Suppose the initial observation matrix is , its QR decomposition is expressed as:

[0070] ;

[0071] in, It is an orthogonal matrix, which has the property that column vectors are orthogonal to each other; is an upper triangular matrix containing the main linear combination coefficients. Through QR decomposition, the complex structure of the initial observation matrix is ​​simplified into an easy-to-handle form. Based on the orthogonal matrix and the upper triangular matrix , calculate the generalized subspace projection matrix. The generalized subspace projection matrix can project the original data into a low-dimensional subspace, retain the main features of the data and remove noise, and obtain the initial projection matrix. Based on the initial variable forgetting factor and the initial projection matrix, recursive least squares estimation is performed on the standardized time series data set. Recursive least squares estimation is an online learning algorithm that can update model parameters in real time as data gradually arrives. Through the recursive least squares estimation algorithm, each time new data arrives, the variable forgetting factor is used to adjust the model's weight for the new daily data, while maintaining the stability of the model and quickly responding to changes in new data. The goal of the recursive least squares estimation algorithm is to minimize the following objective function:

[0072] ;

[0073] in, represents the loss function, It is the forgetfulness factor. It is The observations at each time step, is the eigenvector, is the parameter vector to be estimated. By optimizing the objective function, the initial state estimate is obtained. Based on the initial state estimate and the initial projection matrix, the initial subspace model is constructed. This model is the optimal representation of the system under the current conditions and can reflect the main dynamic characteristics in the data. By constructing the subspace model, the state changes of the system can be effectively tracked and the initial subspace estimate can be obtained.

[0074] In one example, a multi-branch long short-term memory network is trained based on time-varying pseudo-modal temperature parameters and a standardized time series data set to obtain an initial water temperature prediction model, including:

[0075] The standardized time series data set is divided into time steps to obtain an input sequence with a fixed time window, and the input sequence is feature expanded based on the time-varying pseudo-modal temperature parameter to obtain an expanded feature sequence;

[0076] According to the data type of the expanded feature sequence, it is divided into a water temperature branch LSTM sub-network, a flow branch LSTM sub-network, a pressure branch LSTM sub-network, and an environmental parameter branch LSTM sub-network. Each branch contains an independent LSTM sub-network, resulting in four parallel LSTM branches.

[0077] The water temperature branch LSTM subnetwork is constructed, including the first input layer, the first LSTM layer and the first output layer. The first LSTM layer contains 64 neurons and uses the tanh activation function. Dropout regularization is used between the first input layer and the first LSTM layer.

[0078] The traffic branch LSTM subnetwork is constructed, including the second input layer, the second LSTM layer, and the second output layer. The second LSTM layer contains 32 neurons and uses the ReLU activation function. A batch normalization layer is added after the second LSTM layer.

[0079] The pressure branch LSTM sub-network is constructed, including the third input layer, two stacked third LSTM layers and the third output layer. Each third LSTM layer contains 48 neurons, using the Leaky ReLU activation function, and adding residual connections between the two stacked third LSTM layers.

[0080] The environmental parameter branch LSTM sub-network is constructed, including the fourth input layer, the fourth LSTM layer, the attention layer and the fourth output layer. The fourth LSTM layer contains 40 neurons, uses the ELU activation function, and the attention layer adopts the self-attention mechanism;

[0081] The outputs of the water temperature branch LSTM subnetwork, the flow branch LSTM subnetwork, the pressure branch LSTM subnetwork, and the environmental parameter branch LSTM subnetwork are spliced ​​to obtain a fused feature vector, and the fused feature vector is fully connected to obtain a comprehensive feature representation;

[0082] A bidirectional LSTM layer is constructed based on the comprehensive feature representation, which contains 128 neurons, uses the tanh activation function, and adds skip connections between the forward and reverse LSTMs to obtain temporal dependency features;

[0083] A multi-head self-attention mechanism is applied to the time-dependent features, including 8 attention heads, each with a dimension of 16, to obtain the attention-weighted features, and the attention-weighted features are input into the fully connected layer to obtain the predicted output;

[0084] The mean square error was used as the loss function, and the Adam optimizer was used for end-to-end training. The initial value of the learning rate was set to 0.001, and the cosine annealing strategy was used to dynamically adjust the learning rate to obtain the initial water temperature prediction model.

[0085] In this example, the standardized time series data set is divided into time steps, and the continuous time series data is divided into input sequences with fixed time windows. The input sequence is feature expanded based on the time-varying pseudo-modal temperature parameters. The time-varying pseudo-modal temperature parameters reflect the dynamic temperature changes of the system in different time periods. By combining these parameters with the input sequence, the feature dimension of the input is expanded, so that the model can better capture the complex temperature dynamic behavior. According to the data type of the expanded feature sequence, it is divided into four independent LSTM subnetworks: water temperature branch LSTM subnetwork, flow branch LSTM subnetwork, pressure branch LSTM subnetwork and environmental parameter branch LSTM subnetwork. Each subnetwork is responsible for processing different types of data and extracting related features respectively to form four parallel LSTM branches. For the water temperature branch LSTM subnetwork, its structural design includes a first input layer, a first LSTM layer and a first output layer. The first LSTM layer contains 64 neurons, and the input data is processed using the tanh activation function. The tanh function can effectively process the nonlinear relationship in the time series data. The output value fluctuates between -1 and 1, which helps to stabilize the training process. In order to prevent the model from overfitting, Dropout regularization is added between the first input layer and the first LSTM layer to randomly discard some neuron outputs and enhance the generalization ability of the model. The design of the flow branch LSTM subnetwork includes the second input layer, the second LSTM layer and the second output layer. The second LSTM layer contains 32 neurons and uses the ReLU activation function, which has the characteristics of nonlinearity and sparsity, enabling the model to better learn the characteristics of complex data. In order to improve the stability and training speed of the model, a batch normalization layer is added after the second LSTM layer. The batch normalization layer standardizes each batch of data so that each layer of the model can be trained on a more stable distribution, accelerates convergence and prevents gradient disappearance. The pressure branch LSTM subnetwork includes the third input layer, two stacked third LSTM layers and the third output layer. Each third LSTM layer contains 48 neurons and uses the Leaky ReLU activation function. Leaky ReLU still retains a small gradient when there is a negative input, thereby preventing neurons from "dying" during training. A residual connection is added between the two stacked third LSTM layers. The residual connection directly adds the input to the output, allowing the deep network to better transmit information and avoid the gradient vanishing problem in the deep network. The environmental parameter branch LSTM subnetwork includes the fourth input layer, the fourth LSTM layer, the attention layer, and the fourth output layer. The fourth LSTM layer contains 40 neurons and uses the ELU activation function. The ELU activation function can provide a greater response to negative inputs, which helps to speed up the training speed and convergence performance of the model. An attention layer with a self-attention mechanism is added to the fourth LSTM layer.The self-attention mechanism calculates the correlation between different parts of the input sequence, so that the model can automatically focus on the most important features for the output results, thereby improving the prediction accuracy of the model. The outputs of the water temperature branch LSTM subnetwork, the flow branch LSTM subnetwork, the pressure branch LSTM subnetwork, and the environmental parameter branch LSTM subnetwork are concatenated to form a fused feature vector. The fused feature vector integrates the feature information extracted by each subnetwork and provides a comprehensive representation of the system state. The fused feature vector is fully connected to obtain a comprehensive feature representation. A bidirectional LSTM layer is constructed based on the comprehensive feature representation. The bidirectional LSTM layer contains 128 neurons and uses the tanh activation function. The bidirectional LSTM layer enables the model to capture more complex temporal patterns in the sequence by simultaneously considering the forward and backward dependencies of the input sequence. Skip connections are added between the forward and backward LSTMs to ensure the effective transfer of information between time steps and obtain temporal dependency features. In order to enhance the model's ability to capture key features, a multi-head self-attention mechanism is applied to the temporal dependency features. The multi-head self-attention mechanism contains 8 attention heads, each with a dimension of 16. This mechanism weights the time-dependent features from multiple perspectives by processing them in parallel between different heads, making the features of the final output more expressive. The attention-weighted features are input into the fully connected layer to obtain the final prediction output. During the training of the model, the mean square error (MSE) is used as the loss function. The mean square error is a common indicator for evaluating the prediction effect of the regression model. Its formula is:

[0086] ;

[0087] in, For the The true value of the samples, is the predicted value of the model, is the number of samples. The smaller the MSE, the closer the prediction result of the model is to the true value. The Adam optimizer is used for end-to-end training. The Adam optimizer is an optimization algorithm with adaptive learning rate. It combines the advantages of momentum and RMSProp to speed up training and improve the convergence of the model. The initial value of the learning rate is set to 0.001 to ensure that the model has a large adjustment range in the early stage to converge quickly. In order to improve the training effect, the cosine annealing strategy is used to dynamically adjust the learning rate. Cosine annealing gradually reduces the learning rate at different stages of training, so that the model can make more refined parameter adjustments in the later stage of training, and finally obtain a stable initial water temperature prediction model.

[0088] In one example, the mean square error is used as the loss function, the Adam optimizer is used for end-to-end training, the initial value of the learning rate is set to 0.001, and the cosine annealing strategy is used to dynamically adjust the learning rate to obtain the initial water temperature prediction model, including:

[0089] The error between the predicted output and the actual water temperature value is calculated to obtain the mean square error loss value, and the gradient of the network parameters is calculated based on the mean square error loss value to obtain the initial gradient value;

[0090] Based on the initial gradient value, the Adam optimizer is used to update the parameters, where the momentum parameter is set to 0.9, the exponential decay rate of the second-order moment estimate is set to 0.999, and the numerical stability constant is set to 1e-8 to obtain the network parameters after the first round of updates;

[0091] Apply the cosine annealing strategy to the training round, calculate the current learning rate, obtain the dynamically adjusted learning rate, perform forward propagation calculation based on the dynamically adjusted learning rate, obtain the iterative prediction output, and calculate the difference between the iterative prediction output and the actual value to obtain the differential error;

[0092] Perform differential overfitting on the differential error to obtain the regularized loss function;

[0093] Based on the regularized loss function, the back propagation algorithm is used to calculate the gradient, and the momentum and adaptive learning rate mechanism of the Adam optimizer are combined to update the network parameters to obtain the initial water temperature prediction model.

[0094] In this example, the error between the predicted output and the actual water temperature is calculated, and the mean square error is used as the loss function. The mean square error can measure the difference between the predicted value and the true value, and its calculation formula is:

[0095] ;

[0096] in, represents the number of samples, For the The actual water temperature value of the sample, For the The predicted output of samples. By calculating the MSE, the mean square error loss value is obtained. The smaller the value, the more accurate the model's prediction. The gradient of the network parameters is calculated based on the mean square error loss value. The gradient reflects the rate of change of the loss function with respect to the model parameters. In deep learning, the backpropagation algorithm is often used to calculate the gradient. This algorithm transfers the error layer by layer through the chain rule so that the weights and biases of each layer can obtain the corresponding gradient information and obtain the initial gradient value. Based on the initial gradient value, the Adam optimizer is used to update the network parameters. The Adam optimizer is an optimization algorithm that combines momentum and adaptive learning rate. It dynamically adjusts the learning rate by calculating the moving average of the first-order moment and the second-order moment, thereby accelerating the convergence speed and avoiding falling into the local optimal solution. The update formula of the Adam optimizer is as follows:

[0097] ;

[0098] ;

[0099] ;

[0100] ;

[0101] ;

[0102] in, and are estimates of the first and second moments of the gradient, respectively. is the current gradient, and is the exponential decay rate of the first-order moment and second-order moment estimates, usually set to 0.9 and 0.999, and are the bias-corrected first- and second-order moment estimates, is the learning rate, is a numerical stability constant, usually set to . Through the calculation of these formulas, the network parameters after the first round of updates are obtained. In order to improve the optimization effect of the model, the cosine annealing strategy is applied to the training rounds. The cosine annealing strategy gradually reduces the learning rate by simulating the attenuation characteristics of the cosine function, so that the model can make more precise parameter adjustments in the later stage of training. The learning rate calculation formula of cosine annealing is as follows:

[0103] ;

[0104] in, It is The learning rate during the training round, and are the minimum and maximum learning rates, respectively, is the current training epoch, is the preset maximum number of training rounds. The dynamically adjusted learning rate is calculated through the cosine annealing strategy. Based on the dynamically adjusted learning rate, the forward propagation calculation is performed to obtain the iterative prediction output. The difference between the iterative prediction output and the actual water temperature value is calculated again to obtain the differential error, which reflects the improvement of the model performance in each round of training. In order to prevent the model from overfitting during the training process, the differential error is subjected to differential overfitting mitigation processing. Overfitting means that the model fits the training data too well and cannot be generalized to new data. Differential overfitting mitigation is usually achieved by adding regularization terms. The regularization term can limit the complexity of the model so that the model can maintain the ability to predict unseen data while fitting the training data. The regularized loss function usually consists of two parts: an error term and a regularization term, and its form is as follows:

[0105] ;

[0106] in, is the regularization coefficient, which is used to balance the influence of the error term and the regularization term. Based on the regularized loss function, the back propagation algorithm is used to calculate the new gradient, and the momentum and adaptive learning rate mechanism of the Adam optimizer are combined to update the network parameters again. After multiple iterations and optimizations, a converged initial water temperature prediction model is finally obtained.

[0107] In one example, system operation anomaly detection and prediction are performed based on the initial water temperature prediction model, and anomaly detection and prediction results are obtained, including:

[0108] The target operation data of the constant water temperature control system is preprocessed to obtain standardized target time series data, and the standardized target time series data is input into the initial water temperature prediction model. The water temperature, flow, pressure and environmental parameter data are respectively input into the corresponding LSTM branches to obtain the water temperature feature subsequence, flow feature subsequence, pressure feature subsequence and environmental parameter feature subsequence;

[0109] The water temperature feature subsequence is input into the first input layer of the water temperature branch LSTM subnetwork, and processed by the first LSTM layer of 64 neurons and the tanh activation function to obtain the water temperature feature vector;

[0110] The traffic feature subsequence is input into the second input layer of the traffic branch LSTM subnetwork, and is processed through the second LSTM layer of 32 neurons, the ReLU activation function, and batch normalization to obtain the traffic feature vector.

[0111] The pressure feature subsequence is input into the third input layer of the pressure branch LSTM subnetwork, and processed by the stacked third LSTM layer with two layers of 48 neurons each and the Leaky ReLU activation function to obtain the pressure feature vector;

[0112] The environmental parameter feature subsequence is input into the fourth input layer of the environmental parameter branch LSTM subnetwork, and processed by the fourth LSTM layer of 40 neurons, the ELU activation function and the self-attention mechanism to obtain the environmental parameter feature vector;

[0113] The water temperature feature vector, flow feature vector, pressure feature vector and environmental parameter feature vector are concatenated and processed through a fully connected layer to obtain the target feature representation;

[0114] The target feature representation is input into the bidirectional LSTM layer, processed by the forward and reverse LSTM of 128 neurons, and skip connections are applied to obtain the target time series feature;

[0115] An 8-head self-attention mechanism is applied to the target time series features, with each head dimension of 16, to obtain the attention-weighted features, which are then processed through a fully connected layer to obtain the water temperature prediction value;

[0116] The prediction error is calculated based on the predicted water temperature value and the actual water temperature value, and the Mahalanobis distance and local anomaly factor algorithm are used to detect anomalies in the prediction error. At the same time, the threshold is dynamically adjusted in combination with the time-varying pseudo-modal temperature parameters to obtain anomaly detection and prediction results.

[0117] In this example, the target operation data of the constant water temperature control system is preprocessed, and the data is converted into a uniform scale range to obtain standardized target time series data. The standardized target time series data is input into the initial water temperature prediction model. The model consists of multiple LSTM (Long Short-Term Memory Network) branches, each of which is responsible for processing different types of data. The water temperature data is input into the water temperature branch LSTM subnetwork, the flow data is input into the flow branch LSTM subnetwork, the pressure data is input into the pressure branch LSTM subnetwork, and the environmental parameter data is input into the environmental parameter branch LSTM subnetwork. Each branch network will extract the corresponding features to generate a water temperature feature subsequence, a flow feature subsequence, a pressure feature subsequence, and an environmental parameter feature subsequence. In the water temperature branch LSTM subnetwork, the water temperature feature subsequence is input into the first input layer. The subsequence passes through an LSTM layer containing 64 neurons. The LSTM layer can effectively capture the time dependency in the input data through its built-in memory unit and gating mechanism. The tanh activation function is used to perform nonlinear processing on the output of the LSTM layer so that the output value is limited between -1 and 1, which helps to maintain the stability of the network and the gradient flow during training. After this series of processing, a water temperature feature vector is generated, which summarizes the main dynamic features in the water temperature data. For the flow branch LSTM sub-network, the flow feature sub-sequence is input to the second input layer. The sequence is processed by an LSTM layer containing 32 neurons. The ReLU activation function is used to activate the output of the LSTM layer. The main advantages of the ReLU activation function are its nonlinearity and sparsity, which enables the model to learn complex features more effectively. A batch normalization layer is added to the LSTM layer. The function of this layer is to standardize the output of the LSTM layer and reduce the impact of internal covariate shift, thereby speeding up the training speed of the model and improving stability. After processing, a flow feature vector is generated. For the pressure branch LSTM sub-network, the pressure feature sub-sequence is input to the third input layer. The sub-sequence is processed by two stacked LSTM layers, each of which contains 48 neurons. The Leaky ReLU activation function is used to activate the output of the LSTM layer. The activation function allows a small gradient to be maintained in the case of negative input, thereby avoiding the "death" phenomenon of neurons. In order to enhance the deep learning ability of the network, residual connections are added between the two LSTM layers. This connection method can effectively alleviate the gradient vanishing problem in deep networks and ensure that the network can learn useful features in the data at a deeper level. Generate pressure feature vectors. For the environmental parameter branch LSTM sub-network, the environmental parameter feature sub-sequence is input to the fourth input layer. The sequence is processed by an LSTM layer containing 40 neurons, and the ELU activation function is used to activate the output of the LSTM layer. The ELU activation function provides an exponential function when the input is negative, which can accelerate the convergence of the network and reduce the impact of the offset.The attention layer of the self-attention mechanism is added. The self-attention mechanism enables the model to automatically focus on the most important parts for prediction by calculating the correlation between different parts of the input sequence, improves the prediction accuracy of the model, and generates the environmental parameter feature vector. The water temperature feature vector, flow feature vector, pressure feature vector and environmental parameter feature vector are concatenated to form a fused feature vector. The fused feature vector is processed by the fully connected layer to refine and fuse the feature information to obtain the target feature representation. The target feature representation is input into a bidirectional LSTM layer. The bidirectional LSTM layer contains 128 neurons and uses the tanh activation function. The bidirectional LSTM layer enables the model to capture the dynamic patterns in the time series more comprehensively by considering the forward and backward dependencies of the input sequence at the same time. Skip connections are added between the forward and backward LSTMs to ensure the effective transmission of information between time steps to obtain the target time series features. In order to enhance the performance of the model, a multi-head self-attention mechanism is applied to the target time series features. The multi-head self-attention mechanism contains 8 attention heads, each with a dimension of 16. The self-attention mechanism can perform weighted processing on the target time series features from multiple angles through parallel processing between different heads, making the final output features more expressive. The attention weighted features are input into the fully connected layer, and after processing, the final water temperature prediction value is obtained. The prediction error is calculated by comparing the predicted water temperature value with the actual water temperature value. The calculation of the prediction error is an important indicator to measure the performance of the model. The mean square error (MSE) is usually used to measure the gap between the predicted value and the actual value. The MSE formula is:.

[0118] ;

[0119] in, Indicates the actual water temperature value. represents the model prediction value, is the number of samples. The smaller the MSE, the closer the prediction result of the model is to the true value. The Mahalanobis distance and local anomaly factor algorithm are used to detect anomalies in the prediction error. The Mahalanobis distance can quantify the similarity between multidimensional data and is used to identify outliers. The formula is as follows:

[0120] ;

[0121] in, is a data point, is the data mean, is the covariance matrix. The local anomaly factor algorithm determines whether a data point is an outlier by comparing its local density. Combined with the time-varying pseudo-modal temperature parameter, the detection threshold is dynamically adjusted to accurately identify anomalies. The anomaly detection and prediction results are obtained.

[0122] In one example, a control strategy is solved based on anomaly detection and prediction results to obtain an optimal water temperature control strategy, including:

[0123] The state space modeling of the abnormal detection and prediction results is carried out, and the water temperature, flow rate, pressure and environmental parameters are used as state variables, and the heating power and flow rate regulation are used as control variables to obtain the Markov decision process model;

[0124] The value function network and policy network are constructed based on the Markov decision process model, where the value function network adopts a three-layer feedforward neural network structure and the policy network adopts a Gaussian policy model to obtain the initial neural network structure;

[0125] Initialize the parameters of the initial neural network structure, use the Xavier method to initialize the weights, initialize the bias to zero, and obtain the initialized value function network and policy network;

[0126] Based on the initialized value function network and policy network, the current state is sampled, the policy network is used to generate control actions, and the execution is simulated through the environment model to obtain the next state and immediate reward;

[0127] According to the next state and immediate reward, the target value of the value function is calculated using the temporal difference learning method, and the network parameters of the value function are updated through back propagation to obtain the updated value function estimate;

[0128] Based on the updated value function estimate, the policy gradient method is used to calculate the gradient of the policy network, and the policy network parameters are updated through the trust region policy optimization algorithm to obtain the improved control strategy;

[0129] Exploratory sampling is performed on the improved control strategy. Actions are executed in multiple time steps, state transition samples are collected, and an experience replay buffer is obtained. Based on the experience replay buffer, batch samples are randomly selected until convergence or the preset number of iterations is reached to obtain the optimal water temperature control strategy.

[0130] In this example, state space modeling is performed for anomaly detection and prediction results. Water temperature, flow, pressure, and environmental parameters are used as state variables of the system to describe the operating state of the system at any given time point. Heating power and flow regulation are used as control variables, which are parameters that can be directly adjusted by the system to affect the change of state variables. A Markov decision process (MDP) model is constructed by introducing state variables and control variables into the state space. The core of the MDP model lies in its state transition and reward mechanism, which describe the state change of the system after taking a certain control action and the corresponding reward value. Based on the MDP model, two key neural networks are constructed: the value function network and the policy network. The value function network is used to estimate the expected cumulative reward in a certain state, while the policy network is used to generate the control action to be taken in the current state. The value function network adopts a three-layer feedforward neural network structure, which is suitable for processing nonlinear relationships in continuous state space. The feedforward neural network approximates complex functional relationships and accurately predicts the value of a certain state through the combination of multi-layer linear transformation and nonlinear activation function. The policy network adopts the Gaussian policy model, which can better handle the problem of continuous action space by generating control actions that conform to the Gaussian distribution. The Gaussian policy model usually describes the distribution of control actions by means of mean and variance. This method can retain a certain degree of exploratory nature when generating actions, making the policy network more flexible when searching for the optimal control strategy. Initialize the parameters of the initial neural network structure. In order to ensure the training efficiency and stability of the network, the Xavier method is used to initialize the weights of the network. The Xavier initialization method controls the variance of the weights so that the variance of the input and output of each layer remains consistent, thereby avoiding the problem of gradient disappearance or explosion during training. The bias of the network is initialized to zero, which simplifies the training process of the network and helps the network quickly enter an effective learning state. Through these initialization steps, the initialized value function network and policy network are obtained. Based on the initialized value function network and policy network, the current state is sampled and representative features are extracted from the current state. The policy network is used to generate control actions, which will affect the state transition of the system. By inputting the control action into the environment model and simulating its execution process, the next state and immediate reward are obtained. The immediate reward is usually defined based on the system control goal, such as keeping the water temperature within the target range or minimizing energy consumption. Based on the next state and immediate reward, the target value of the value function is calculated using the Temporal Difference Learning (TD Learning) method. Temporal Difference Learning is a reinforcement learning method that can gradually approximate the true value function without fully understanding the environmental model. TD learning corrects the parameters of the value function network by comparing the current estimate with the estimate of the next state. The specific TD target value calculation formula is as follows:

[0131] ;

[0132] in, Indicates the current status The value of It’s an instant reward. is the discount factor, controlling the influence of future rewards, is the learning rate, which determines the stride of each update. Through the back-propagation algorithm, the error is propagated backward to each layer of the network, the parameters of the value function network are updated, and the updated value function estimate is obtained. Based on the updated value function estimate, the policy gradient method is used to calculate the gradient of the policy network. The policy gradient method adjusts the parameters of the policy network by calculating the contribution of the control action output by the policy network to the expected reward, so as to gradually improve the overall decision quality. To ensure the stability of the policy update, the trust region policy optimization algorithm is used. The trust region policy optimization algorithm ensures that the policy does not change drastically by limiting the amplitude of each policy update, so as to avoid the network from falling into the local optimal solution or causing performance degradation during the training process. The updated policy network represents an improved control policy. In order to verify and optimize the control strategy, exploratory sampling is performed. By executing control actions in multiple time steps, observing the state changes of the system, and collecting state transition samples. These samples are stored in the experience replay buffer, which can save a large amount of historical data and provide a rich source of samples for subsequent training. By randomly extracting batch samples from the experience replay buffer, multiple rounds of training and optimization are performed until the performance of the policy network converges or reaches the preset number of iterations, and finally the optimal water temperature control strategy is obtained.

[0133] In one example, water temperature control is performed based on the optimal water temperature control strategy, and the initial water temperature prediction model is updated to obtain a target water temperature prediction model, including:

[0134] Discretize the optimal water temperature control strategy, map the continuous control actions to a limited set of control instructions, and obtain a discrete control instruction sequence;

[0135] Based on the discrete control instruction sequence, the heating power and flow rate of the constant water temperature control system are adjusted, and the real-time water temperature, flow rate, pressure and environmental parameter data during the execution process are collected to obtain a real-time feedback data set;

[0136] Perform sliding time window processing on the real-time feedback data set, extract data samples of the most recent N time steps, and obtain the sliding window data subset;

[0137] Based on the sliding window data subset, the initial water temperature prediction model is reconstructed online, including dynamic weight adjustment of water temperature, flow, pressure and environmental parameter characteristics to obtain updated input features;

[0138] According to the updated input features, the parameters of the initial water temperature prediction model are locally updated, including updating the weight matrix and bias vector of the LSTM layer, to obtain the target water temperature prediction model.

[0139] In this example, the continuous control action is discretized. Continuous control actions are usually represented as a series of adjustment parameters that can take any value, such as heating power and water flow. The continuous control action is mapped to a finite set of control instructions to obtain a discretized control instruction sequence. The heating power and flow of the constant water temperature control system are adjusted based on the discretized control instruction sequence. The control system gradually adjusts the power output and water flow of the heating device according to the discretized instructions to maintain the stability of the target water temperature. While executing these control instructions, the system collects data on water temperature, flow, pressure and environmental parameters in real time to reflect the changes in various state variables of the system during execution. By monitoring the state variables, a real-time feedback data set is formed. The real-time feedback data is processed by sliding time window. Sliding time window processing is a commonly used time series data processing method. By sliding a fixed-length window on the time series, the data samples of the most recent N time steps are extracted. Each sliding window contains continuous data within a period of time, which can capture the dynamic change characteristics of the system during that period of time. The initial water temperature prediction model is reconstructed online based on the sliding window data subset. The input features of the model are dynamically adjusted according to the latest system status. This includes redistributing the weights of features such as water temperature, flow, pressure, and environmental parameters to ensure that the model can focus on the most important features at the moment. For example, when the ambient temperature changes significantly, the model may need to increase the weight of environmental parameters in the input features to better adapt to the water temperature prediction needs under the new conditions. The updated input features are obtained through online feature reconstruction. According to the updated input features, the parameters of the initial water temperature prediction model are locally updated, including updating the weight matrix and bias vector of the LSTM layer. The LSTM layer is the core part of the long short-term memory network, which can capture long-term and short-term dependencies in time series data. The weight matrix controls how the input data affects the output of each neuron, while the bias vector is a fixed value added by each neuron when calculating the output. By locally updating the weight matrix and bias vector, the model can better adapt to the latest input features, improve the prediction accuracy and the response speed of the system. Through a series of local updates and online learning, the initial water temperature prediction model gradually evolves into the target water temperature prediction model. The target water temperature prediction model can not only adapt to the changes in the system state in real time, but also provide more accurate water temperature prediction results, thereby optimizing the execution effect of the control strategy.

[0140] Reference Figure 2 , this embodiment provides an intelligent water temperature control system based on machine learning, including:

[0141] Acquisition module 1, used to collect water temperature, flow, pressure and environmental parameter data of the constant water temperature control system and generate a standardized time series data set;

[0142] Building module 2, for building a time-varying dynamic model based on a standardized time series data set, and using a variable forgetting factor generalized subspace tracker for parameter identification to obtain time-varying pseudo-modal temperature parameters;

[0143] Training module 3, used to train a multi-branch long short-term memory network according to the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model;

[0144] Detection module 4, used to detect and predict system operation anomalies based on the initial water temperature prediction model, and obtain anomaly detection and prediction results;

[0145] A solution module 5 is used to solve the control strategy based on the abnormality detection and prediction results to obtain the optimal water temperature control strategy;

[0146] The control module 6 is used to perform water temperature control based on the optimal water temperature control strategy, and to update the initial water temperature prediction model to obtain a target water temperature prediction model.

[0147] In this embodiment, for the specific implementation of each unit in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.

[0148] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a display screen, an input system, a network interface and a database connected through a system bus. Among them, the processor designed by the computer is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.

[0149] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0150] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented. It can be understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0151] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided by the present invention and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM.

[0152] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, system, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, system, article or method. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, system, article or method including the element.

[0153] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An intelligent water temperature control method based on machine learning, characterized in that: The following steps are involved: Collect water temperature, flow, pressure and environmental parameter data of the constant water temperature control system to generate a standardized time series data set; The method comprises the following steps: constructing a time-varying dynamic model based on the standardized time series data set, and performing parameter identification using a variable forgetting factor generalized subspace tracker to obtain time-varying pseudo-modal temperature parameters; specifically comprising: performing subsystem decomposition on the standardized time series data set to obtain multiple subsystem data sets, and constructing a thermodynamic model based on the multiple subsystem data sets to obtain a time-varying temperature-flow coupling dynamic equation; performing Markov superposition on the time-varying temperature-flow coupling dynamic equation to obtain a component matrix, and calculating the system mean square error based on the component matrix to obtain an initial mean square error value; setting the system mean square error value according to the initial mean square error value; The initial value of the variable forgetting factor is set to obtain the initial variable forgetting factor, and the generalized subspace tracking is performed on the standardized time series data set to obtain the initial subspace estimate; recursive parameter identification is performed based on the initial subspace estimate and the initial variable forgetting factor to obtain the first round of identification results, and the system mean square error is updated according to the first round of identification results to obtain an updated mean square error value; the variable forgetting factor is dynamically adjusted based on the updated mean square error value to obtain the optimized variable forgetting factor, and the first round of identification results are iteratively optimized based on the optimized variable forgetting factor to obtain time-varying pseudo-modal temperature parameters; Training a multi-branch long short-term memory network according to the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model; Performing system operation abnormality detection and prediction based on the initial water temperature prediction model to obtain abnormality detection and prediction results; Solving the control strategy based on the abnormality detection and prediction results to obtain the optimal water temperature control strategy; Water temperature control is performed based on the optimal water temperature control strategy, and the initial water temperature prediction model is updated to obtain a target water temperature prediction model.

2. The intelligent water temperature control method based on machine learning according to claim 1 is characterized in that: The water temperature, flow rate, pressure and environmental parameter data of the constant water temperature control system are collected to generate a standardized time series data set, including: Installing multiple sensors on the water inlet, water outlet, heating unit and water storage device of the constant water temperature control system; Collect water temperature, flow rate, pressure and ambient temperature data based on the multiple sensors to obtain raw sensor data; Performing wavelet transform processing on the raw sensor data to obtain denoised data, and performing outlier detection on the denoised data based on a density-based local anomaly factor algorithm to obtain an outlier marking result; Based on the outlier marking result, outliers in the denoised data are removed to obtain outlier-removed data, and multivariate interpolation processing is performed on missing values ​​in the outlier-removed data to obtain completed data; The completed data is normalized to obtain normalized data, and the normalized data is divided into time windows and features are extracted to obtain a standardized time series data set.

3. The intelligent water temperature control method based on machine learning according to claim 1 is characterized in that: The step of setting the initial value of the variable forgetting factor according to the initial mean square error value to obtain the initial variable forgetting factor, and performing generalized subspace tracking on the standardized time series data set to obtain an initial subspace estimate includes: Performing exponential smoothing on the initial mean square error value to obtain a smoothed mean square error sequence, and calculating an initial value of a variable forgetting factor based on the smoothed mean square error sequence to obtain an initial variable forgetting factor; Performing singular value decomposition on the standardized time series data set to obtain a left singular matrix, a singular value matrix, and a right singular matrix, and constructing an observation matrix based on the left singular matrix and the singular value matrix to obtain an initial observation matrix; Performing QR decomposition on the initial observation matrix to obtain an orthogonal matrix and an upper triangular matrix, and calculating a generalized subspace projection matrix based on the orthogonal matrix and the upper triangular matrix to obtain an initial projection matrix; Recursive least squares estimation is performed on the standardized time series data set based on the initial variable forgetting factor and the initial projection matrix to obtain an initial state estimate, and an initial subspace model is constructed based on the initial state estimate and the initial projection matrix to obtain an initial subspace estimate.

4. The intelligent water temperature control method based on machine learning according to claim 1 is characterized in that: The method of training a multi-branch long short-term memory network according to the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model includes: Dividing the standardized time series data set into time steps to obtain an input sequence with a fixed time window, and performing feature expansion on the input sequence based on the time-varying pseudo-modal temperature parameter to obtain an expanded feature sequence; According to the data type of the expanded feature sequence, it is divided into a water temperature branch LSTM subnetwork, a flow branch LSTM subnetwork, a pressure branch LSTM subnetwork and an environmental parameter branch LSTM subnetwork, each branch contains an independent LSTM subnetwork, and four parallel LSTM branches are obtained; The water temperature branch LSTM subnetwork is constructed, including a first input layer, a first LSTM layer and a first output layer, wherein the first LSTM layer contains 64 neurons, uses a tanh activation function, and Dropout regularization is used between the first input layer and the first LSTM layer; The traffic branch LSTM subnetwork is constructed, including a second input layer, a second LSTM layer, and a second output layer, wherein the second LSTM layer contains 32 neurons, uses a ReLU activation function, and a batch normalization layer is added after the second LSTM layer; The pressure branch LSTM subnetwork is constructed, including a third input layer, two stacked third LSTM layers and a third output layer, each third LSTM layer contains 48 neurons, a Leaky ReLU activation function is used, and a residual connection is added between the two stacked third LSTM layers; The environmental parameter branch LSTM subnetwork is constructed, including a fourth input layer, a fourth LSTM layer, an attention layer and a fourth output layer, wherein the fourth LSTM layer contains 40 neurons, uses an ELU activation function, and the attention layer adopts a self-attention mechanism; The outputs of the water temperature branch LSTM subnetwork, the flow branch LSTM subnetwork, the pressure branch LSTM subnetwork, and the environmental parameter branch LSTM subnetwork are spliced ​​to obtain a fused feature vector, and the fused feature vector is fully connected to obtain a comprehensive feature representation; Based on the comprehensive feature representation, a bidirectional LSTM layer is constructed, which includes 128 neurons, uses a tanh activation function, and adds a jump connection between the forward and reverse LSTMs to obtain a temporal dependency feature; Applying a multi-head self-attention mechanism to the temporal dependency feature, including 8 attention heads, each with a dimension of 16, to obtain an attention-weighted feature, and inputting the attention-weighted feature into a fully connected layer to obtain a prediction output; The mean square error was used as the loss function, and the Adam optimizer was used for end-to-end training. The initial value of the learning rate was set to 0.001, and the cosine annealing strategy was used to dynamically adjust the learning rate to obtain the initial water temperature prediction model.

5. The intelligent water temperature control method based on machine learning according to claim 4 is characterized in that: The mean square error is used as the loss function, the Adam optimizer is used for end-to-end training, the initial value of the learning rate is set to 0.001, and the cosine annealing strategy is used to dynamically adjust the learning rate to obtain the initial water temperature prediction model, including: Performing error calculation on the predicted output and the actual water temperature value to obtain a mean square error loss value, and calculating the gradient of the network parameter based on the mean square error loss value to obtain an initial gradient value; Based on the initial gradient value, the Adam optimizer is used to update the parameters, where the momentum parameter is set to 0.9, the exponential decay rate of the second-order moment estimate is set to 0.999, and the numerical stability constant is set to 1e-8, to obtain the network parameters after the first round of updates; Applying a cosine annealing strategy to the training round, calculating the current learning rate to obtain a dynamically adjusted learning rate, performing a forward propagation calculation based on the dynamically adjusted learning rate to obtain an iterative prediction output, and calculating the difference between the iterative prediction output and the actual value to obtain a differential error; Perform differential overfitting relief on the differential error to obtain a regularized loss function; Based on the regularized loss function, the back propagation algorithm is used to calculate the gradient, and the momentum and adaptive learning rate mechanism of the Adam optimizer are combined to update the network parameters to obtain the initial water temperature prediction model.

6. The intelligent water temperature control method based on machine learning according to claim 5 is characterized in that: The system operation abnormality detection and prediction is performed based on the initial water temperature prediction model to obtain abnormality detection and prediction results, including: Preprocess the target operation data of the constant water temperature control system to obtain standardized target time series data, input the standardized target time series data into the initial water temperature prediction model, input the water temperature, flow, pressure and environmental parameter data into the corresponding LSTM branches respectively, and obtain the water temperature feature subsequence, flow feature subsequence, pressure feature subsequence and environmental parameter feature subsequence; Input the water temperature feature subsequence into the first input layer of the water temperature branch LSTM subnetwork, and process it through the first LSTM layer of 64 neurons and the tanh activation function to obtain a water temperature feature vector; Input the traffic feature subsequence into the second input layer of the traffic branch LSTM subnetwork, and obtain a traffic feature vector after passing through a second LSTM layer of 32 neurons, a ReLU activation function, and batch normalization processing; Input the pressure feature subsequence into the third input layer of the pressure branch LSTM subnetwork, and process it through two stacked third LSTM layers with 48 neurons each and a Leaky ReLU activation function to obtain a pressure feature vector; Inputting the environmental parameter feature subsequence into the fourth input layer of the environmental parameter branch LSTM subnetwork, and processing it through the fourth LSTM layer of 40 neurons, the ELU activation function and the self-attention mechanism to obtain an environmental parameter feature vector; The water temperature feature vector, the flow feature vector, the pressure feature vector and the environmental parameter feature vector are concatenated and processed through a fully connected layer to obtain a target feature representation; The target feature representation is input into a bidirectional LSTM layer, processed by a forward and reverse LSTM of 128 neurons, and a skip connection is applied to obtain a target time series feature; Applying an 8-head self-attention mechanism to the target time series features, each with a head dimension of 16, obtains attention-weighted features, and processes them through a fully connected layer to obtain a water temperature prediction value; The prediction error is calculated based on the water temperature prediction value and the actual water temperature value, and the prediction error is detected for anomalies using the Mahalanobis distance and local anomaly factor algorithm. At the same time, the threshold is dynamically adjusted in combination with the time-varying pseudo-modal temperature parameter to obtain anomaly detection and prediction results.

7. The intelligent water temperature control method based on machine learning according to claim 1 is characterized in that: The control strategy is solved based on the abnormality detection and prediction results to obtain the optimal water temperature control strategy, including: Performing state space modeling on the abnormality detection and prediction results, taking water temperature, flow rate, pressure and environmental parameters as state variables, and heating power and flow rate regulation as control variables, to obtain a Markov decision process model; Based on the Markov decision process model, a value function network and a policy network are constructed, wherein the value function network adopts a three-layer feedforward neural network structure, and the policy network adopts a Gaussian policy model to obtain an initial neural network structure; Initializing the parameters of the initial neural network structure, using the Xavier method to initialize the weights, initializing the bias to zero, and obtaining an initialized value function network and a strategy network; Based on the initialized value function network and policy network, the current state is sampled, a control action is generated using the policy network, and the control action is executed through environmental model simulation to obtain the next state and immediate reward; According to the next state and the immediate reward, a target value of the value function is calculated using a temporal difference learning method, and a network parameter of the value function is updated by back propagation to obtain an updated value function estimate; Based on the updated value function estimate, the policy network gradient is calculated using the policy gradient method, and the policy network parameters are updated using the trust region policy optimization algorithm to obtain an improved control strategy; The improved control strategy is subjected to exploratory sampling, actions are executed in multiple time steps, state transition samples are collected, and an experience replay buffer is obtained. Based on the experience replay buffer, batch samples are randomly extracted until convergence or a preset number of iterations is reached to obtain the optimal water temperature control strategy.

8. The intelligent water temperature control method based on machine learning according to claim 1 is characterized in that: The performing of water temperature control based on the optimal water temperature control strategy and updating the initial water temperature prediction model to obtain a target water temperature prediction model includes: Discretizing the optimal water temperature control strategy, mapping continuous control actions to a finite set of control instructions, and obtaining a discrete control instruction sequence; Based on the discrete control instruction sequence, the heating power and flow rate of the constant water temperature control system are adjusted, and real-time water temperature, flow rate, pressure and environmental parameter data during the execution process are collected to obtain a real-time feedback data set; Performing sliding time window processing on the real-time feedback data set, extracting data samples of the most recent N time steps, and obtaining a sliding window data subset; Based on the sliding window data subset, online feature reconstruction is performed on the initial water temperature prediction model, including dynamic weight adjustment of water temperature, flow, pressure and environmental parameter features to obtain updated input features; According to the updated input features, the parameters of the initial water temperature prediction model are locally updated, including updating the weight matrix and bias vector of the LSTM layer, to obtain a target water temperature prediction model.

9. An intelligent water temperature control system based on machine learning, characterized in that: Used to execute the intelligent water temperature control method based on machine learning according to claim 1, the intelligent water temperature control system based on machine learning comprises: The acquisition module is used to collect the water temperature, flow, pressure and environmental parameter data of the constant water temperature control system and generate a standardized time series data set; A construction module is used to construct a time-varying dynamic model based on the standardized time series data set, and use a variable forgetting factor generalized subspace tracker to perform parameter identification to obtain time-varying pseudo-modal temperature parameters; A training module, used for training a multi-branch long short-term memory network according to the time-varying pseudo-modal temperature parameters and the standardized time series data set to obtain an initial water temperature prediction model; A detection module, used to perform system operation anomaly detection and prediction based on the initial water temperature prediction model to obtain anomaly detection and prediction results; A solution module, used to solve the control strategy based on the abnormality detection and prediction results to obtain the optimal water temperature control strategy; A control module is used to perform water temperature control based on the optimal water temperature control strategy and update the initial water temperature prediction model to obtain a target water temperature prediction model.

Citation Information

Patent Citations

  • Generalized subspace tracing-based time-varying structure mode parameter identification method

    CN107729592A

  • Power system dynamic state estimation method and system based on long short-term memory network

    CN117318021A