Water quality prediction method based on Transform-LSTM fusion model
Through the Transformer-LSTM fusion model, the problems of medium- and long sequence modeling, multivariate nonlinear relationship analysis and spatial and temporal feature fusion are solved, and efficient and stable water quality prediction effect is achieved.
Patent Information
- Application Number
- CN202510513910.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-01
AI Technical Summary
The existing water quality prediction methods have shortcomings in long-sequence modeling, multivariate nonlinear relationship analysis, spatiotemporal and spatial characteristics fusion and generalization capabilities, and are difficult to meet the needs of real-time prediction.
Transformer-LSTM fusion model is adopted to capture long-distance timing dependence through self-attention mechanism and position coding, and strengthen local dynamic learning with the gating mechanism, integrate spatiotemporal features, optimize gradient propagation and regularization technologies to improve model accuracy and stability.
It significantly improves the modeling accuracy of multivariate nonlinear relationships, enhances the interactive modeling ability of spatiotemporal features, and improves the prediction accuracy and stability of the model in new data and complex environments.
Smart Images

Figure CN120408555A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of water quality time series data prediction and artificial intelligence, and particularly relates to a water quality prediction method based on a Transformer-LSTM fusion model. Background Art
[0002] At present, there are mainly three types of water quality prediction methods for various water bodies: (1) Traditional statistical models (such as ARIMA) are based on linear assumptions, with high computational efficiency, but they have insufficient ability to capture long-distance time dependence relationships, are difficult to model the complex non-linear associations between water quality parameters, and have strict requirements for data stationarity, while actual water quality data often shows non-stationarity due to factors such as seasons and meteorology. (2) Process-driven models (such as MIKE, EFDC) simulate ecological processes through physical equations, but face two major bottlenecks: the generalization ability is limited due to parameter calibration relying on expert experience, and the computational complexity is high and it is difficult to meet the real-time prediction requirements. (3) Neural network models (such as LSTM, CNN) have the ability of non-linear modeling, however, a single model has the problem of gradient disappearance and is difficult to capture long sequence dependencies, and a hybrid model is prone to parameter explosion due to multi-dimensional input. Existing methods often separate the spatio-temporal features for processing, ignoring the spatio-temporal evolution coupling law of lake water quality, and the model is prone to overfitting local noise, and the prediction performance significantly decreases under new monitoring sites or extreme events. In summary, existing methods have obvious deficiencies in long sequence modeling, multi-variable non-linear relationship analysis, spatio-temporal feature fusion and generalization ability, and there is an urgent need for innovative methods that take into account both efficiency and robustness. Summary of the Invention
[0003] In view of the technical bottlenecks of existing water quality prediction methods in aspects such as long-sequence modeling, multi-variable non-linear relationship analysis, spatio-temporal feature fusion, and generalization ability, a water quality prediction method based on a Transformer-LSTM fusion model is proposed, and breakthroughs are achieved through the following technical solutions: First, long-sequence dependence and global feature capture. The self-attention mechanism of Transformer is adopted, combined with sine-cosine position encoding, to explicitly model the long-distance temporal dependence relationship of water quality data, breaking through the limitation of the vanishing gradient of LSTM. Second, local dynamic pattern optimization. The LSTM layer receives the high-order features encoded by Transformer, and with the help of gating mechanisms such as input gate, forget gate, and output gate, strengthens the fine-grained learning of local temporal dynamics and improves the modeling accuracy of multi-variable non-linear relationships. Third, deep spatio-temporal feature fusion. Spatio-temporal position encoding is introduced into Transformer to combine the spatial distribution of water quality parameters (such as the location of monitoring stations) with time series information to achieve spatio-temporal interaction modeling and solve the limitation of spatio-temporal separation in traditional methods. Fourth, enhanced generalization ability. Gradient propagation is optimized through residual connection and layer normalization, combined with Dropout regularization technology, to suppress the overfitting of the model to the noise of training data and improve the accuracy and stability of prediction.
[0004] The technical solution adopted by the present invention is: A water quality prediction method based on a Transformer-LSTM fusion model, comprising the following steps:
[0005] (1) Obtain water quality data from water quality monitoring stations of a certain lake, reservoir, river, etc.
[0006] (2) Preprocess the obtained water quality data.
[0007] (3) Conduct a correlation analysis on the preprocessed water quality data and screen out water quality characteristic data.
[0008] (4) Divide the screened water quality data into a training set, a validation set, and a test set.
[0009] (5) Input the divided water quality data into the Transformer-LSTM fusion model.
[0010] (6) The Transformer encoder embeds the temporal information of water quality data through position encoding, uses the multi-head attention mechanism to extract the global dependence relationship between features, and combines residual connection and layer normalization to optimize gradient propagation;
[0011] (7) LSTM receives the high-order features encoded by Transformer and captures local temporal dynamic patterns;
[0012] (8) The regression output layer uses a linear activation function to map the extracted water quality time series features to specific prediction results and calculates a series of evaluation indicators to objectively measure the performance of the model.
[0013] Step (1) is as follows: water quality monitoring data for a certain period of time is obtained from a water quality monitoring station of a lake, reservoir or river, including water quality indicators such as water temperature, pH, dissolved oxygen, and total phosphorus.
[0014] Step (2) is as follows: The acquired water quality data may be missing or contain outliers in consecutive time periods due to monitoring equipment failure or extreme weather conditions. The missing values are filled using cubic spline interpolation, and outliers are removed using the 3σ principle to improve the quality of the dataset.
[0015] Step (3) is as follows: Use the Pearson method to perform correlation analysis on the pre-processed water quality data to measure the correlation between water quality parameters, so as to select features with high correlation with the prediction target as input variables. The Pearson correlation coefficient calculation formula is as follows:
[0016]
[0017] Among them, x i and y i are the i-th observation values of the two variables, and are the means of the two variables, and n is the number of samples.
[0018] Step (4) is as follows: the selected water quality characteristic data are divided into a training set, a validation set, and a test set in a ratio of 8:1:1. The training set is used to train the network to learn the patterns and characteristics in the data; the validation set is used to evaluate the performance of the network during training and adjust hyperparameters; and the test set is used to evaluate the network's final generalization ability and prediction performance.
[0019] Step (5) is as follows: the divided water quality data is input into the Transformer-LSTM fusion model. The input layer receives the water quality time series data and converts it into a word embedding vector (water quality feature embedding vector) X∈R T×D , where T is the time step, D is the feature dimension, and R is a real matrix of dimension T × D. The Transformer encoder consists of N encoding layers, each of which consists of a multi-head attention module and a feedforward neural network. The LSTM layer includes M memory cells and a hidden layer dimension of H. The output layer is mapped to the prediction target dimension O through a fully connected layer.
[0020] Step (6) is as follows: Transformer embeds the time series information of water quality data through position encoding, and the formula is as follows:
[0021]
[0022] Among them, PE is the positional encoding, pos is the sequence position, and i is the feature dimension index.
[0023] Immediately afterwards, the multi-head attention mechanism is used to extract the global dependencies between features, and the formula is as follows:
[0024] MultiHead(Q, K, V) = concat(head1,..., head h )W o
[0025] head i = Attention(QW i Q , KW i K , VW i V )
[0026]
[0027] Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and W o is the matrix for linearly transforming the concatenated multi-head outputs in multi-head attention to integrate features, represents the matrix for linearly transforming the query matrix Q, mapping Q to D × d k dimensions, represents the matrix for linearly transforming the key matrix K, mapping K to D × d k dimensions, is the matrix for linearly transforming the value matrix V, mapping V to D × d v dimensions, h is the number of attention heads, d k is the key dimension, and d v is the value dimension.
[0028] After that, residual connection and layer normalization are combined to optimize gradient propagation, and the formula is as follows:
[0029] z = LayerNorm(x + SubLayer(x))
[0030]
[0031] Among them, μ is the mean, σ 2 is the variance, γ and β are learnable parameters, x is the original data input to the current module, LayerNorm(x) is the operation of performing layer normalization on x, and ε is a very small number to prevent division by zero in the denominator.
[0032] Step (7) is specifically as follows: The LSTM layer receives the high-order features encoded by the Transformer, captures the local temporal dynamic patterns, and the formula is as follows:
[0033] The update formula of the LSTM cell is:
[0034]
[0035] h t = o t ·tanh(c t )
[0036] where, i t is the input gate, f t is the forget gate, o t is the output gate, c t is the cell state at time t, is the candidate memory unit, h t-1 is the hidden state at time t-1, h t is the hidden state at time t, is the sigmoid function, tanh is the hyperbolic tangent function, x t is the input at time t, W i is the weight matrix of the input x t to the input gate, W f is the weight matrix of the input x t to the forget gate, W o is the weight matrix of the input x t to the output gate, W c is the weight matrix of the input x t to the candidate cell state, U i is the weight matrix of the previous hidden state h t-1 to the input gate, U f is the weight matrix of the previous hidden state h t-1 to the forget gate, U o is the weight matrix of the previous hidden state h t-1 to the output gate, U c wis the weight matrix of the previous hidden state h t-1 to the candidate cell state, b i is the bias term of the input gate, b f is the bias term of the forget gate, b o is the bias term of the output gate, b c is the bias term of the candidate cell state.
[0037] Step (8) is specifically as follows: The regression output layer maps the extracted temporal features to the specific continuous value prediction results, and adopts the mean square error (MSE) loss function. The formula is as follows:
[0038]
[0039] Among them, n is the number of samples, and y i is the i-th observation of the variable, that is, y i is the true value, is the predicted value.
[0040] Finally, the root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R 2 ) are used as evaluation indicators for the network prediction performance, and the formulas are as follows:
[0041]
[0042] Among them, is the mean of the true values.
[0043] The beneficial effects of the present invention are:
[0044] (1) Optimization of long-sequence dependence and multi-variable relationships: The self-attention mechanism of Transformer is combined with positional encoding to break through the capture limit of traditional models for long-distance time dependence. At the same time, the gating mechanism of LSTM is used to strengthen local dynamic feature learning, significantly improving the modeling accuracy of multi-variable non-linear relationships.
[0045] (2) Deep fusion of spatio-temporal features: The spatio-temporal attention module in Transformer integrates multi-source water quality time series data, and realizes the interactive modeling of time series and spatial distribution through joint spatio-temporal positional encoding, enhancing the ability to capture the spatio-temporal evolution law of water quality.
[0046] (3) Improvement of generalization ability and prediction accuracy: By improving the model structure design, the adaptability of the model to different scenarios and data distributions is enhanced. Specifically, by integrating multi-dimensional information, optimizing the feature extraction method, and using regularization techniques to prevent overfitting, the model can still maintain stable prediction performance in new data or complex environments, improving the accuracy and reliability of prediction results. Description of the Drawings
[0047] Figure 1 is a schematic flow chart of a water quality prediction method based on a Transformer-LSTM fusion model of the present invention;
[0048] Figure 2 is a schematic diagram of the Transformer model structure;
[0049] Figure 3 is a schematic diagram of the LSTM model structure;
[0050] Figure 4This is a schematic diagram of the Transformer-LSTM model structure;
[0051] Figure 5 This is a schematic diagram of the actual effect of the present invention on the prediction of dissolved oxygen concentration;
[0052] Figure 6 It is a density scatter diagram showing the prediction effect of the present invention on dissolved oxygen concentration. DETAILED DESCRIPTION
[0053] To make the purpose, technical solutions and advantages of the present invention clearer, Figure 1-6 The embodiments of the present invention are described in further detail.
[0054] Example 1: To build a neural network model to accurately predict water quality change trends, see Figure 1 , an embodiment of the present invention proposes a water quality prediction method based on a Transformer-LSTM fusion model, comprising the following steps:
[0055] (1) Obtain water quality data from a water quality monitoring station at a lake, reservoir, river, etc.; specifically, the water quality data obtained in this embodiment is derived from the China National Environmental Monitoring Center. Water quality monitoring data from the Caohai Center and Broken Bridge Station of Dianchi Lake are selected, including 11 indicators: water temperature, pH, dissolved oxygen, total phosphorus, total nitrogen, ammonia nitrogen, chlorophyll, algae density, conductivity, turbidity, and permanganate index. The data set time range is from 2021 to 2024, and the water quality data time step is 4 hours (monitoring once every 4 hours).
[0056] (2) Preprocess the acquired water quality data; specifically, if the water quality data is missing or has abnormal values in a continuous time period due to monitoring equipment failure or extreme weather conditions, the missing values are filled using cubic spline interpolation and the abnormal values are eliminated using the 3σ principle. The formula is as follows:
[0057] S i (x) = a i +b i (x'-x k )+c i (x'-x k ) 2 +d i (x'-x k ) 3
[0058] Among them, a i 、b i 、c i d i is the coefficient to be determined, x kis the horizontal coordinate of the known data point, and x' is any horizontal coordinate value of the interpolation result to be calculated within the interpolation interval.
[0059] After filling the missing values, the 3σ principle is used to eliminate outliers in the water quality data. The 3σ principle is based on the normal distribution and is used to determine whether the data is an outlier. The relevant formula and calculation steps are as follows:
[0060] 1. Calculate the mean. The mean reflects the central tendency of the data. The formula is as follows:
[0061]
[0062] 2. Calculate the standard deviation. The standard deviation measures the degree of dispersion of the data. The formula is as follows:
[0063]
[0064] 3. According to the 3σ principle, approximately 99.73% of the data will fall within the range of μ±3σ. If a data exceeds this range, it can be judged as an outlier.
[0065] Where μ is the mean, σ is the standard deviation, and n is the number of samples.
[0066] (3) Performing correlation analysis on the pre-processed water quality data to screen out water quality characteristic data; specifically, this embodiment uses dissolved oxygen as the prediction target, adopts the Pearson correlation analysis method to perform correlation analysis on the water quality time series data, selects water quality parameters with a correlation threshold of dissolved oxygen |r| ≥ 0.5 as input features, and finally selects water temperature, pH, total phosphorus, ammonia nitrogen, and total nitrogen as input features. The formula of the Pearson correlation analysis method is as follows:
[0067]
[0068] Among them, x i and y i are the i-th observation values of the two variables, and are the means of the two variables, and n is the number of samples.
[0069] (4) Dividing the screened water quality data into a training set, a validation set, and a test set; specifically, in this embodiment, the screened water quality characteristic data is divided into a training set, a validation set, and a test set in a ratio of 8:1:1. The training set is used to train the network so that it learns the patterns and features in the data; the validation set is used to evaluate the performance of the network during training and adjust hyperparameters; and the test set is used to evaluate the network's final generalization ability and prediction performance.
[0070] (5) Input the divided water quality data into the Transformer-LSTM fusion model, such asFigure 4 As shown below: In parameter settings, set the time window T = 12 (historical data for 48 hours) to predict the dissolved oxygen concentration for the next 4 hours. The shape of the input sequence is "number of samples, 12, number of features", and the output shape is "number of samples, 1".
[0071] (6) The Transformer embeds the temporal information of water quality data through positional encoding, uses the multi-head attention mechanism to extract the global dependencies between features, and combines residual connections and layer normalization to optimize gradient propagation, as Figure 2 shown below: The input embedding layer converts the water quality time series data (shape: "number of samples, time window, number of features") into high-dimensional embedding vectors; a positional encoding matrix is generated through the sine-cosine function and added element-wise to the input embedding vector to endow the model with temporal position perception ability. The formula is as follows:
[0072]
[0073] where pos ∈ [0, 11]: the time step position (corresponding to historical data for 48h, with each 4h as a time step), D is the feature dimension, i.e., the number of input features, and i is the index of the feature dimension.
[0074] The Transformer multi-head attention module consists of h independent attention heads. Each head generates query (Q), key (K), and value (V) matrices through linear transformation to calculate the global dependencies between features. Among them, the number of heads h is set to 4, the attention key d k is set to 64, and the attention value d v dimension is set to 64.
[0075] Calculation process: MultiHead(Q, K, V) = concat(head1,..., head h )W o
[0076] head i = Attention(QW i Q , KW i K , VW i V )
[0077]
[0078] Output dimension: 64 × 4 = 256 dimensions, linearly projected back to 64 dimensions.
[0079] Residual connection and layer normalization: Residual connections are introduced after multi-head attention and the feed-forward neural network, and layer normalization is combined to optimize gradient propagation. The formula is as follows:
[0080] z = LayerNorm(x + SubLayer(x))
[0081]
[0082] where μ is the mean, σ 2 is the variance, γ and β are learnable parameters, and the activation function is GELU.
[0083] (7) The LSTM receives the high - order features encoded by the Transformer, learns the local dynamic patterns of water quality data through the gating mechanism, and makes up for the deficiency of the Transformer's low sensitivity to short - term fluctuations. As Figure 3 shown below: The LSTM layer receives high - order features with a dimension of 64. In the LSTM network, the hidden layer dimension is set to 32, the number of memory units is set to 2, the learning rate is set to 0.001, the number of training epochs is set to 200, the dropout rate is set to 0.02, and the LSTM formula is as follows:
[0084] i t = σ(W i x t + U i h t-1 + b i )
[0085] f t = σ(W f x t + U f h t-1 + b f )
[0086] o t = σ(W o x t + U0h t-1 + b o )
[0087]
[0088] h t = o t ·tanh(c t )
[0089] where i t is the input gate, f t is the forget gate, o t is the output gate, c t is the cell state at time t, is the candidate memory unit, h t-1 is the hidden state at time t - 1, h tis the hidden state at time t, is the sigmoid function, tanh is the hyperbolic tangent function, and x t is the input at time t, and W i input x t to the weight matrix of the input gate, W f is the input x t to the weight matrix of the forget gate, W o is the input x t to the weight matrix of the output gate, W c is the input x t to the weight matrix of the candidate cell state, U i is the hidden state h at the previous time t-1 to the weight matrix of the input gate, U f is the hidden state h at the previous time t-1 to the weight matrix of the forget gate, U o is the hidden state h at the previous time t-1 to the weight matrix of the output gate, U c is the hidden state h at the previous time t-1 to the weight matrix of the candidate cell state, b i is the bias term of the input gate, b f is the bias term of the forget gate, b o is the bias term of the output gate, b c is the bias term of the candidate cell state.
[0090] (8) The regression output layer uses a linear activation function to map the extracted water quality time series features to specific prediction results and calculates a series of evaluation metrics to objectively measure the performance of the model. Specifically as follows:
[0091] The root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R 2 ) are used as evaluation metrics for the prediction performance of the fusion model, where the smaller the RMSE and MAE, the better, and R 2 is between 0 and 1, and the closer it is to 1, the better the prediction effect. The formula is as follows:
[0092]
[0093] where, is the mean of the true values.
[0094] In the actual prediction of dissolved oxygen, the measured results of the Transformer-LSTM fusion model and other models were compared.
[0095] Table 1 Performance of different models in predicting dissolved oxygen at 2 water quality monitoring stations
[0096]
[0097] Table 1 shows the MAE, RMSE, and R of the prediction effects of the LSTM, Transformer, and Transformer-LSTM fusion models on the dissolved oxygen prediction at the Caohai Center and Duanqiao water quality monitoring stations. 2 The performance of the three evaluation indicators. At the Caohai Center station, the MAE of the Transformer-LSTM fusion model is 0.614 (lower than 0.801 of LSTM and 0.771 of Transformer), the RMSE is 0.823 (better than 1.119 of LSTM and 0.943 of Transformer), and the R 2 is 0.931 (higher than 0.894 of LSTM and 0.91 of Transformer); at the Duanqiao station, its MAE is 0.482 (lower than 0.581 of LSTM and 0.555 of Transformer), the RMSE is 0.632 (better than 0.761 of LSTM and 0.779 of Transformer), and the R 2 is 0.912 (higher than 0.845 of LSTM and 0.856 of Transformer). Overall, in the comparison of the two stations, the Transformer-LSTM fusion model has the smallest error in the MAE and RMSE indicators, and the R 2 indicator has the best fitting effect. Compared with the LSTM and Transformer models, it shows better prediction accuracy, stronger error control ability, and more excellent data fitting ability. As Figure 5 shown, it presents the visualization diagram of the prediction results of the test set when the lake water quality prediction method based on the Transformer-LSTM fusion model predicts the dissolved oxygen concentration. The horizontal axis represents the water quality time series data (corresponding to the time series observation points), and the vertical axis is the value of the dissolved oxygen. The red curve in the figure represents the true value of the dissolved oxygen, and the black curve is the predicted value of the model. Overall, the black predicted value curve can closely follow the change trend of the red true value curve, and their fluctuation patterns are highly consistent at most time points, reflecting the effective capture ability of the model for the change law of dissolved oxygen. Although there are slight deviations between the predicted value and the true value at some time points (such as around 600 on the horizontal axis), it can still show the large fluctuation trend of the true value, fully indicating that the Transformer-LSTM fusion model has good effectiveness and reliability in the prediction task and can provide a reliable reference basis for water quality prediction.
[0098] Figure 6The test set density scatter plot showing the dissolved oxygen prediction effect of the Transformer-LSTM fusion model, where the horizontal axis represents the true value of dissolved oxygen and the vertical axis represents the predicted value of the model, covering 800 data points. The gray straight line in the figure is the fitting line of the predicted value and the true value, corresponding to the equation "y = 0.954x + 0.231", indicating a significant linear positive correlation between the predicted value and the true value. When the true value changes, the predicted value can respond reasonably according to this rule. The coefficient of determination "R 2 = 0.931", indicating that the predicted value of the model is highly consistent with the true value, with excellent fitting effect, highlighting the model's ability to deeply capture the variation law of dissolved oxygen. The density of the scatter points is reflected by the color scale (L to H) on the right side. The area near the fitting line is mainly in warm colors (yellow, orange), indicating that most data points are concentrated here, that is, a large number of predicted values are closely distributed around the true value, with only a small number deviating. This distribution characteristic not only demonstrates the prediction accuracy of the model in most scenarios but also reflects its stable ability to capture the internal relationship between the true value and the predicted value of dissolved oxygen, fully verifying the reliability, accuracy, and stability of the Transformer-LSTM fusion model in water quality prediction tasks.
[0099] Although the embodiments or examples of the present application have been described in conjunction with the accompanying drawings, it should be clear that the above methods are only exemplary embodiments or examples, and the protection scope of the present invention is not limited by these embodiments or examples, but only by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the execution order of each step can be different from that described in the present application. Furthermore, the elements in the embodiments or examples can be combined in various ways. With the development of technology, many elements described here can be replaced by equivalent elements that appear after the present application, which is crucial.
Claims
1. A water quality prediction method based on a Transformer-LSTM fusion model, characterized in that: The following steps are involved: (1) Obtain water quality data from water quality monitoring stations of water bodies; (2) Preprocessing the acquired water quality data; (3) Conduct correlation analysis on the pre-processed water quality data to screen out water quality characteristic data; (4) Divide the screened water quality data into training set, validation set and test set; (5) Input the divided water quality data into the Transformer-LSTM fusion model; (6) The Transformer encoder embeds the temporal information of water quality data through position encoding, uses the multi-head attention mechanism to extract the global dependency between features, and combines residual connections with layer normalization to optimize gradient propagation; (7) The LSTM layer receives the high-order features encoded by the Transformer and captures the local temporal dynamic patterns; (8) The regression output layer uses a linear activation function to map the extracted water quality temporal features to specific prediction results and calculates a series of evaluation indicators to objectively measure the performance of the model.
2. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that, Step (1) is as follows: obtain water quality data for a continuous period of time at a water quality monitoring station in a lake, including water temperature, pH, dissolved oxygen, and total phosphorus.
3. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that Step (2) is as follows: use the cubic spline interpolation method to fill the missing values in the water quality data, and use the 3σ principle to eliminate outliers in the water quality data.
4. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that Step (3) is as follows: the pre-processed water quality data is subjected to correlation analysis using the Pearson method to measure the correlation between water quality parameters, so as to select features with a high correlation with the prediction target as input variables. The Pearson correlation coefficient calculation formula is as follows: where x i and y i are the i-th observations of two variables respectively, and are the means of the two variables respectively, and n is the sample size.
5. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that, Step (4) is as follows: the screened water quality characteristic data are divided into a training set, a validation set, and a test set in a ratio of 8:1:1, wherein the training set is used to train the network to learn the patterns and features in the data; the validation set is used to evaluate the performance of the network during the training process and adjust the hyperparameters; and the test set is used to evaluate the network's final generalization ability and prediction performance.
6. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that, Step (5) is specifically as follows: Input the divided water quality data into the Transformer-LSTM fusion model. The input layer receives the water quality time series data and converts it into a word embedding vector X ∈ R T×D , where T is the time step, D is the feature dimension, R is a real number matrix of dimension T×D. The Transformer encoder contains N encoding layers, each encoding layer consists of a multi-head attention module and a feed-forward neural network. The LSTM layer includes M memory units, the hidden layer dimension is H, and the output layer is mapped to the prediction target dimension O through a fully connected layer.
7. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that Step (6) is as follows: Transformer embeds the time series information of water quality data through position encoding, and the calculation formula is: Among them, PE is the position code, pos is the sequence position, and i is the feature dimension index; Then, the multi-head attention mechanism is used to extract the global dependency between features. The formula is as follows: MultiHead(Q,K,V)=concat(head1,...,head h )W o head i = Attention(QW i Q ,KW i K ,VW i V ) Among them, Q is the query matrix, K is the key matrix, V is the value matrix, and W o is the matrix for linearly transforming the concatenated multi-head outputs in multi-head attention to integrate features, represents the matrix for linearly transforming the query matrix Q, mapping Q to D×d k dimension, represents the matrix for linearly transforming the key matrix K, mapping K to D×d k dimension, is the matrix for linearly transforming the value matrix V, mapping V to D×d v dimension, h is the number of attention heads, d k is the key dimension, d v is the value dimension; Then, the residual connection and layer normalization are combined to optimize the gradient propagation. The formula is as follows: z = LayerNorm(x + SubLayer(x)) where μ is the mean, σ 2 is the variance, γ and β are learnable parameters, x is the original data input to the current module, LayerNorm(x) is the layer normalization operation on x, and ε is a number to prevent division by zero in the denominator.
8. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that Step (7) is as follows: The LSTM layer receives the high-order features encoded by the Transformer and captures the local temporal dynamic pattern. The formula is as follows: The LSTM unit update formula is: h t = o t ·tanh(c t ) Among them, i t is the input gate, f t is the forget gate, o t is the output gate, c t is the cell state at time t, is the candidate memory unit, h t-1 is the hidden state at time t-1, h t is the hidden state at time t, is the sigmoid function, tanh is the hyperbolic tangent function, x t is the input at time t, W i input x t to the weight matrix of the input gate, W f is the input x t to the weight matrix of the forget gate, W o is the input x t to the weight matrix of the output gate, W c is the input x t to the weight matrix of the candidate cell state, U i is the weight matrix of the previous hidden state h t-1 to the input gate, U f is the weight matrix of the previous hidden state h t-1 to the forget gate, U o is the weight matrix of the previous hidden state h t-1 to the output gate, U c is the weight matrix of the previous hidden state h t-1 to the candidate cell state, b i is the bias term of the input gate, b f is the bias term of the forget gate, b o is the bias term of the output gate, b c is the bias term of the candidate cell state.
9. A water quality prediction method based on a Transformer-LSTM fusion model according to claim 1, characterized in that, Step (8) is as follows: The regression output layer maps the extracted time series features to specific continuous value prediction results, using the mean square error (MSE) loss function, which is as follows: where n is the number of samples, and y i is the i-th observation of the variable, and is the predicted value; Finally, the root mean square error RMSE, mean absolute error MAE, and coefficient of determination R 2 are used as evaluation metrics for the network prediction performance, and the formulas are as follows: Among them, is the mean of the true values.
Citation Information
Patent Citations
Space-time multi-source offshore water quality time sequence prediction method of LSTM coupling mechanism model
CN116187210A
TCN-LSTM water quality prediction method based on variational mode decomposition and attention mechanism
CN119624217A
Cited By
Space intelligent watershed hydrology and water quality multivariable prediction method based on quaternion time-frequency multi-scale
CN120822194A
Transform-BiLSTM water quality prediction method based on shop interpretability analysis
CN121188745A
Physical law constrained marine ecological variable collaborative prediction method and prediction system thereof
CN121234728A
Expressway interleaving area lane change accident risk early warning method and system based on vehicle trajectory data
CN121305922A
Water quality interval prediction method based on integrated deep learning model
CN121327608A