Water quality prediction model and prediction method based on spatial-temporal feature fusion and LSTM memory network
By using the model of spatiotemporal and spatial feature fusion and LSTM memory network in water quality prediction, the limitations of the prior art in processing complex time series data and multimodal data fusion are solved, and high-accurate water quality prediction is achieved.
Patent Information
- Application Number
- CN202510226378.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
Existing water quality prediction technologies have limitations in processing complex time series data, multimodal data fusion, and capturing nonlinear relationships, making it difficult to achieve highly accurate water quality prediction.
The water quality prediction model based on spatiotemporal feature fusion and LSTM memory network is adopted. The characteristics of time series data and spatial data are extracted respectively through the time convolution network and the classical convolution network, and they are integrated through the feature fusion module, and the timing modeling is used for LSTM to finally generate the water quality prediction results.
This model can effectively process the timing dependence and nonlinear relationships in water quality monitoring data, flexibly integrate multi-source information, improve the accuracy of water quality prediction, and provide strong support for river basin pollution traceability and governance solutions optimization.
Smart Images

Figure CN120067993A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental science and technology, and particularly to a water quality prediction model and prediction method based on spatio-temporal feature fusion and LSTM memory network. Background Art
[0002] In the field of environmental science and technology, especially in water quality monitoring and prediction, the existing technical methods mainly include traditional statistical methods, machine learning algorithms, and preliminary deep learning applications. Traditional methods such as linear regression and time series analysis (such as ARIMA model) have played a certain role in water quality prediction, but they have limitations in dealing with non-linear relationships and complex time series data. With the development of machine learning technology, algorithms such as support vector machine (SVM) and random forest (RF) have been introduced into water quality prediction. These algorithms can handle non-linear relationships and have good robustness, but they are less efficient in dealing with large-scale data sets and are not as effective as specialized time series models in dealing with time series data.
[0003] In recent years, deep learning technology has been widely used in the field of water quality prediction. Recurrent neural networks (RNNs), especially long short-term memory networks (LSTMs) and gated recurrent units (GRUs), have been used for water quality prediction because they can effectively handle long-term dependencies in time series data. However, these models may not fully utilize spatial features. On the other hand, convolutional neural networks (CNNs), although mainly used for image recognition and processing, have also been applied to the processing of time series data, such as capturing local features through one-dimensional convolution. However, using CNNs alone may not be able to fully capture the long-term dependencies of time series data.
[0004] In addition, most of the existing water quality prediction models only consider a single type of data, such as water quality monitoring data itself. However, the change of water quality is affected by multiple factors, including meteorological conditions, hydrological factors, etc. The existing technologies usually lack effective methods to fuse these multi-modal data.
[0005] In summary, although the existing water quality prediction technologies can provide prediction results to a certain extent, they have certain limitations in dealing with complex time series data, multi-modal data fusion, and capturing non-linear relationships. These technologies may encounter challenges in dealing with large-scale data sets and complex environmental problems, especially in situations where real-time response and optimized resource allocation are required. Therefore, a new method is needed to overcome these limitations and improve the accuracy of water quality prediction. Summary of the Invention
[0006] The object of the present invention is to provide a water quality prediction model and prediction method based on spatio-temporal feature fusion and LSTM memory network. This model can effectively handle the temporal dependence and non-linear relationships in water quality monitoring data, and at the same time flexibly integrate multi-source information to achieve accurate prediction of changes in water quality (such as indicators like pH, ammonia nitrogen concentration, total nitrogen concentration, total phosphorus concentration, dissolved oxygen, etc.), providing strong support for basin pollution source tracing and optimization of treatment plans. Thus, the foregoing problems existing in the prior art are solved.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A deep learning water quality prediction model based on spatio-temporal feature fusion and LSTM, comprising:
[0009] One or more temporal convolutional networks for extracting features from the input time series data. The temporal convolutional network consists of multiple TemporalBlock modules. Each TemporalBlock module contains two 1D convolutional layers with weight normalization, ReLU activation function, Dropout layer, residual connection, and adaptive average pooling layer, and the weights of the convolutional layers are initialized using the Xavier initialization method;
[0010] One or more classical convolutional networks for extracting features from the input spatial data. The classical convolutional network contains one or more convolutional layers, one pooling layer, and one fully connected layer, with parameters initialized using the Xavier initialization method, having a feature fusion mechanism. Each convolutional layer contains a convolution operation, ReLU activation function, Dropout layer, and pooling layer;
[0011] At least one additional temporal convolutional network for processing auxiliary time series data related to the main time series data;
[0012] A long short-term memory network for performing temporal modeling on the output from the temporal convolutional network. The long short-term memory network contains a long short-term memory network layer and a fully connected layer, with the weights of the long short-term memory network and the fully connected layer initialized using the Xavier initialization method, having a hidden state management mechanism;
[0013] A feature fusion module for concatenating the features extracted by multiple temporal convolutional network modules and inputting them into the long short-term memory network for temporal modeling. It is also used to fuse the temporal features extracted by the temporal convolutional network modules, the spatial features extracted by the classical convolutional network, and the temporal modeling results of the long short-term memory network to generate the final prediction output;
[0014] An output module for flexibly adjusting the output to single-step prediction and multi-step prediction, and used to generate the final prediction result.
[0015] Preferably, the model further includes a multi-task learning mechanism that optimizes multiple related tasks simultaneously by sharing some network parameters to improve the generalization ability of the model; and an adaptive learning rate adjustment mechanism that dynamically adjusts the learning rate according to the loss change during training to accelerate model convergence and improve training stability, and supports running in a multi-GPU environment to accelerate model training.
[0016] Preferably, the specific structure of the TemporalBlock module includes:
[0017] The first convolutional layer: the number of input channels is input_channel, the number of output channels is output_channel, the kernel size is kernel_size, the dilation rate is dilation, and the padding size is (kernel_size - 1) * dilation;
[0018] The second convolutional layer: the number of input channels is the same as the number of output channels, and other parameters are the same as those of the first convolutional layer;
[0019] The downsampling layer: when the number of input channels does not match the number of output channels, use a 1x1 convolution to adjust the number of channels;
[0020] The ReLU activation function: applied after each layer of convolution to introduce non-linearity;
[0021] The Dropout layer: applied after each layer of convolution to randomly discard some neurons to prevent overfitting.
[0022] Preferably, the basic building unit configuration of the SpatialBlock of the classical convolutional network is as follows:
[0023] The number of input channels is configured according to the feature dimension of the input data;
[0024] The number of output channels can be adjusted according to the model requirements;
[0025] The kernel size is 3×3 or 5×5;
[0026] The stride is usually set to 1 in the convolutional layer and can be set to 2 in the pooling layer;
[0027] The padding adopts the "same" padding method;
[0028] The pooling method adopts the max pooling operation;
[0029] The Dropout probability is set to 0.2 - 0.5 during training.
[0030] Preferably, the basic building unit configuration of the TemporalBlock of the temporal convolutional network is as follows:
[0031] The number of input channels is configured according to the number of features of the input data;
[0032] The number of output channels is adjusted according to the requirements of the model;
[0033] The size of the convolutional kernel is 2 or 3;
[0034] The stride is set to 1;
[0035] The dilation rate increases gradually;
[0036] Appropriate padding is adopted to keep the output sequence length unchanged;
[0037] The Dropout probability is used to reduce overfitting.
[0038] Preferably, the parameter configuration of the long short-term memory network is as follows:
[0039] The input size is configured according to the number of input features;
[0040] The hidden layer size determines the size of the state vector inside the long short-term memory network;
[0041] The output size is adjusted according to the requirements of the model;
[0042] The number of layers is selected according to the model complexity and the amount of data.
[0043] A water quality prediction method for a deep learning model based on the same concept includes the following steps:
[0044] Data preprocessing: Clean multi-source input data such as the cross-section's own monitoring data, upstream monitoring data of the cross-section, hydrological data, and meteorological data. Use linear interpolation or time series interpolation to fill in missing values, monitor and process outliers through statistical methods, and remove noise using the moving average method; Structure the data, convert time series data into multi-dimensional 1 arrays, convert spatial data into matrix form, and use the Z-score normalization method to convert data with different dimensions into the same scale;
[0045] Feature extraction: Use multiple temporal convolutional networks or classical convolutional network encoders to extract features from different types of preprocessed input data. Among them, the classical convolutional network encoder inputs the formatted spatial data, extracts local spatial features through convolutional operations, introduces non-linearity using the ReLU activation function, reduces the dimension of the feature map through the max pooling layer, and stacks multiple layers to gradually extract spatial features from local to global; The temporal convolutional network encoder inputs the formatted time series data, uses causal convolution to ensure the causal relationship of the time series, expands the receptive field through convolutional layers with different dilation rates, uses residual connections to ensure gradient flow, and introduces a Dropout layer during training to reduce the risk of overfitting;
[0046] Data fusion: Concatenate the outputs from different temporal convolutional networks or classical convolutional network encoders to ensure that the feature dimensions of the outputs from different encoders are consistent, forming a unified feature representation;
[0047] Sequence prediction: Use a long short-term memory network to perform sequence prediction on the fused features. Select single-step prediction or multi-step prediction according to the prediction target. The long short-term memory network processes the long-term dependencies of time series data through a gating mechanism and captures the dynamic changes in the time series through the transmission of hidden states;
[0048] Model training and optimization: Use the mean squared error to evaluate the gap between the predicted results and the actual results, and optimize the parameters through methods such as backpropagation and gradient descent. During the training process, repeat the processes of forward propagation, loss calculation, backpropagation, and parameter update until the model performance reaches the expectation or meets the stopping conditions. Evaluate the model performance on the test set, and adjust the hyperparameters or adopt regularization methods to further optimize the model.
[0049] The beneficial effects of the present invention are as follows: The beneficial effects of the deep learning model and method for water quality prediction based on spatio-temporal feature fusion and LSTM include the following points:
[0050] 1. Multi-source data fusion and feature extraction: Use a temporal convolutional network (TCN) to extract features from time series data and a classical convolutional network (CNN) to extract features from spatial data respectively. It can also process auxiliary time series data, and has a feature fusion module, which can effectively fuse multi-source data features, comprehensively capture the features of water quality data in time and space, and improve the prediction accuracy.
[0051] 2. Prevent overfitting: Dropout layers are set in the TemporalBlock module of TCN, the SpatialBlock basic building unit of CNN, and LSTM. By randomly discarding some neurons, the risk of model overfitting is reduced, and the generalization ability of the model is enhanced.
[0052] 3. Efficient network structure: The TCN module adopts structures such as a 1D convolutional layer with weight normalization, a ReLU activation function, and a residual connection. The CNN has a feature fusion mechanism, and the LSTM has a hidden state management mechanism. These structural designs enable the model to learn and extract features more efficiently when processing data, improving the performance of the model.
[0053] 4. Flexible prediction method: The output module can be flexibly adjusted to single-step prediction and multi-step prediction, which can meet the water quality prediction needs in different scenarios and increase the practicality of the model.
[0054] 5. Multi-task Learning and Adaptive Learning Rate: The model incorporates a multi-task learning mechanism that optimizes multiple related tasks by sharing some network parameters, enhancing generalization ability. The adaptive learning rate adjustment mechanism can dynamically adjust the learning rate according to the training loss, accelerating model convergence and improving training stability.
[0055] 6. Multi-GPU Support: It supports running in a multi-GPU environment, which can accelerate the model training process, shorten the training time, and improve development efficiency.
[0056] 7. Comprehensive Data Preprocessing: In the water quality prediction method, comprehensive preprocessing is performed on multi-source input data, including cleaning, filling missing values, handling outliers, removing noise, data structuring, and normalization, etc. This provides high-quality data for subsequent feature extraction and model training, facilitating the improvement of model performance.
[0057] 8. Effective Model Training and Optimization: Appropriate loss functions such as Mean Squared Error (MSE) are used to evaluate prediction results. Parameter optimization is carried out through methods such as backpropagation and gradient descent. The performance is evaluated on the test set, and hyperparameters are adjusted or regularization and other methods are adopted to further optimize the model, continuously improving the prediction accuracy and stability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic diagram of the model structure of the present invention;
[0059] Figure 2 is a flowchart of the prediction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0061] Refer to Figure 1 and Figure 2 A deep learning model for water quality prediction based on spatio-temporal feature fusion and LSTM shown, includes:
[0062] One or more temporal convolutional networks for extracting features from the input time series data. The temporal convolutional network consists of multiple TemporalBlock modules. Each TemporalBlock module contains two 1D convolutional layers with weight normalization, ReLU activation function, Dropout layer, residual connection, and adaptive average pooling layer, and the Xavier initialization method is used to initialize the weights of the convolutional layers.
[0063] In this embodiment, it should be noted that:
[0064] Two 1D convolutional layers with weight normalization, used to extract local features in time series data;
[0065] ReLU activation function, used to introduce non-linearity;
[0066] Dropout layer, used to prevent overfitting;
[0067] Residual connection, used to directly add the input to the output to alleviate the vanishing gradient problem;
[0068] Adaptive average pooling layer, used to adjust the output to a fixed length;
[0069] Method for initializing weights, using the Xavier initialization method to initialize the weights of the convolutional layers.
[0070] One or more classical convolutional networks, used to extract features from the input spatial data. The classical convolutional network contains one or more convolutional layers, one pooling layer, and one fully connected layer. The parameters are initialized using the Xavier initialization method, and it has a feature fusion mechanism. Each convolutional layer includes a convolution operation, a ReLU activation function, a Dropout layer, and a pooling layer.
[0071] In this embodiment, it should be noted that: one or more convolutional layers: used to extract features from the input spatial data. Each convolutional layer includes a convolution operation, a ReLU activation function, and a Dropout layer;
[0072] One ReLU activation function: used to introduce non-linearity and enhance the expressive power of the model;
[0073] One Dropout layer: used to randomly discard some neurons during training to prevent the model from overfitting;
[0074] One pooling layer: used to reduce the spatial dimension of the feature map, reduce the computational amount, and enhance the robustness of the model;
[0075] One fully connected layer: used to map the features extracted by the convolutional layer to the target output space;
[0076] Method for initializing weights: using Xavier initialization or other statistical initialization methods to initialize the parameters of the CNN;
[0077] Feature fusion mechanism: fusing the spatial features extracted by the CNN with the temporal features extracted by the TCN to enhance the comprehensive modeling ability of the model.
[0078] At least one additional temporal convolutional network, used to process auxiliary time series data related to the main time series data;
[0079] A long short-term memory network for temporal modeling of the output from a temporal convolutional network. The long short-term memory network includes a long short-term memory (LSTM) layer and a fully connected layer. The weights of the long short-term memory network and the fully connected layer are initialized using the Xavier initialization method, and it has a hidden state management mechanism.
[0080] In this embodiment, it should be noted that:
[0081] Long short-term memory (LSTM): The input size is input_size, the hidden layer size is hidden_size, the number of layers is num_layers, and dropout is supported;
[0082] Fully connected layer: Used to map the output of the LSTM to the target output space, and the output size is output_size;
[0083] Method for initializing weights: Use the Xavier initialization method to initialize the weights of the LSTM and the fully connected layer;
[0084] Hidden state management mechanism: By initializing the hidden state and cell state, and effectively managing these states between different batches to ensure the continuity of model training.
[0085] Feature fusion module, used to concatenate the features extracted by multiple temporal convolutional network (TCN) modules and input them into the long short-term memory network for temporal modeling. It is also used to fuse the temporal features extracted by the temporal convolutional network (TCN) module, the spatial features extracted by the classical convolutional network (CNN), and the temporal modeling results of the long short-term memory network (LSTM) to generate the final prediction output;
[0086] Output module, flexibly adjust the output to single-step prediction and multi-step prediction, used to generate the final prediction result.
[0087] Preferably, the model further includes a multi-task learning mechanism, which optimizes multiple related tasks simultaneously by sharing some network parameters to improve the generalization ability of the model; and an adaptive learning rate adjustment mechanism, which dynamically adjusts the learning rate according to the loss change during training to accelerate model convergence and improve training stability, and supports running in a multi-GPU environment to accelerate model training.
[0088] Preferably, the specific structure of the TemporalBlock module includes:
[0089] The first convolutional layer: The number of input channels is input_channel, the number of output channels is output_channel, the convolutional kernel size is kernel_size, the dilation rate is dilation, and the padding size is (kernel_size - 1) * dilation;
[0090] The second convolutional layer: The number of input channels is the same as the number of output channels, and other parameters are the same as those of the first convolutional layer;
[0091] The downsampling layer: When the number of input channels does not match the number of output channels, a 1x1 convolution is used to adjust the number of channels;
[0092] The ReLU activation function: Applied after each convolutional layer to introduce non-linearity;
[0093] The Dropout layer: Applied after each convolutional layer to randomly discard some neurons to prevent overfitting.
[0094] Preferably, the SpatialBlock basic building unit of the classical convolutional neural network (CNN) is configured as follows:
[0095] The number of input channels is configured according to the feature dimension of the input data;
[0096] The number of output channels can be adjusted according to the model requirements;
[0097] The convolutional kernel size is 3×3 or 5×5;
[0098] The stride is usually set to 1 in the convolutional layer and can be set to 2 in the pooling layer;
[0099] The padding adopts the "same" padding method;
[0100] The pooling method adopts the max pooling operation;
[0101] The Dropout probability is set to 0.2 - 0.5 during the training process.
[0102] In this embodiment, (1) CNN (Convolutional Neural Network):
[0103] Definition: CNN is a deep learning architecture based on convolutional operations that efficiently extracts spatial features through local receptive fields and weight sharing mechanisms. Its core components include convolutional layers, pooling layers, and activation functions, which can automatically learn the spatial hierarchical features in the input data.
[0104] Function: In the present invention, as an important part of the hybrid architecture, CNN focuses on extracting spatial features in water quality monitoring data, such as the spatial correlation between monitoring stations, the spatial distribution pattern of pollutants, etc. Through multi-layer convolution operations, CNN can capture spatial dependencies from local to global, providing rich spatial context information for subsequent TCN time feature extraction.
[0105] Implementation details:
[0106] SpatialBlock: This is the basic building block of the CNN network part of the model, used to extract spatial features in water quality monitoring data, such as the spatial relationship between key pollution emission sources and water quality monitoring sections. Each SpatialBlock consists of a convolutional layer, a ReLU activation function, and a pooling layer, gradually capturing spatial dependencies from local to global through multi-layer stacking. The following is an introduction to the network model structure of a single SpatialBlock:
[0107] Number of input channels: Configured according to the feature dimensions of the input data, such as pollution emission intensity, river distance, monitoring point location, etc.
[0108] Number of output channels: Can be adjusted according to model requirements, usually increasing layer by layer to extract more complex spatial features.
[0109] Convolution kernel size: Medium-sized convolution kernels (such as 3×3 or 5×5) are used to balance local feature extraction and computational efficiency.
[0110] Stride: Usually set to 1 to preserve the integrity of spatial information; can be set to 2 in the pooling layer to reduce the feature map size and enhance model robustness.
[0111] Padding: Use the "same" padding method to ensure that the spatial size of the output feature map is the same as the input, avoiding information loss.
[0112] Pooling method: Use the MaxPooling operation to highlight significant spatial features and reduce the risk of overfitting.
[0113] Dropout probability: Introduce a Dropout layer during training (such as a probability set to 0.2 - 0.5) to enhance the generalization ability of the model.
[0114] Preferably, the basic building block TemporalBlock of the temporal convolutional network (TCN) is configured as follows:
[0115] The number of input channels is configured according to the number of features of the input data;
[0116] The number of output channels is adjusted according to the requirements of the model;
[0117] The convolutional kernel size is 2 or 3;
[0118] The stride is set to 1;
[0119] The dilation rate increases gradually;
[0120] Appropriate padding is used to keep the output sequence length unchanged;
[0121] The Dropout probability is used to reduce overfitting.
[0122] In this embodiment, Temporal Convolutional Network (TCN):
[0123] Definition: TCN is an architecture based on a one-dimensional convolutional neural network, specifically designed for processing time series data. It maintains information on long-term dependencies by using causal convolutions and residual connections, and increases the receptive field without increasing the computational cost through dilated convolutions.
[0124] Function: In the present invention, TCN is used to extract the time series features in the input data, especially the spatio-temporal characteristics in water quality monitoring data and other data.
[0125] Implementation details:
[0126] TemporalBlock: This is the basic building block of the TCN network part of the model. It extracts temporal features by stacking convolutional layers and ensures the flow of gradients through residual connections. Each TemporalBlock consists of a convolutional layer, a ReLU activation function, and a Dropout layer, and expands the receptive field by using dilation rates of different scales. The following is an introduction to the network model structure of a single TemporalBlock:
[0127] Number of input channels: Configured according to the number of features of the input data.
[0128] Number of output channels: Can be adjusted according to the requirements of the model.
[0129] Convolutional kernel size: Usually, a smaller kernel size (such as 2 or 3) can capture local features.
[0130] Stride: Usually set to 1 to maintain the integrity of the time series.
[0131] Dilation rate: The gradually increasing dilation rate allows the model to capture dependencies at different time scales.
[0132] Padding: Appropriate padding is used to keep the output sequence length unchanged.
[0133] Dropout probability: Used to reduce overfitting.
[0134] Preferably, the parameter configuration of the LSTM is as follows:
[0135] The input size is configured according to the number of input features;
[0136] The hidden layer size determines the size of the state vector inside the LSTM;
[0137] The output size is adjusted according to the requirements of the model;
[0138] The number of layers is selected according to the model complexity and the amount of data;
[0139] The Dropout probability is used to reduce overfitting.
[0140] In this embodiment: Long Short-Term Memory Network (LSTM):
[0141] Definition: LSTM is a special type of Recurrent Neural Network (RNN) that effectively solves the problems of vanishing gradients and exploding gradients through a gating mechanism and can handle long sequence data.
[0142] Function: In the present invention, LSTM is used to handle the long-term dependencies of time series data, especially in the stage after the TCN extracts features.
[0143] Implementation details:
[0144] Input size: Configured according to the number of input features.
[0145] Hidden layer size: Determines the size of the state vector inside the LSTM.
[0146] Output size: Can be adjusted according to the requirements of the model.
[0147] Number of layers: Can be selected according to the model complexity and the amount of data.
[0148] Dropout probability: Used to reduce overfitting.
[0149] This model supports running in a multi-GPU environment to accelerate model training.
[0150] The present invention also includes: Multi-source data processing:
[0151] Multiple inputs: The model of the present invention can receive multiple inputs, including cross-section self-monitoring data, upstream cross-section monitoring data, hydrological data, meteorological data, etc.
[0152] Multiple TCN and CNN encoders: For each type of data, a dedicated TCN or CNN encoder is designed to extract specific data features.
[0153] Data fusion: Different types of input data are processed by their respective TCN or CNN encoders, and the results are concatenated as the input to the LSTM, thus achieving the fusion of multi-source data.
[0154] The working process of this model:
[0155] Data preprocessing:
[0156] 1) Data cleaning
[0157] ① Missing value handling: For missing values in the data, linear interpolation or time series interpolation is used for filling to ensure the continuity of the data.
[0158] ② Outlier handling: Outliers are monitored and processed through statistical methods (such as Z-score) to ensure the reliability of the data.
[0159] ③ Data denoising: The moving average method is used to remove noise in the data.
[0160] 2) Data formatting
[0161] ① Data structuring: The data is converted into a format suitable for the model. For example, time series data is converted into a multi-dimensional array, and spatial data is converted into a matrix form.
[0162] ② Data normalization: The Min-Max or Z-score normalization method is used to convert data with different dimensions into the same scale.
[0163] Feature extraction:
[0164] Multiple TCN or CNN encoders are used to extract features from different types of input data.
[0165] 1) CNN encoder
[0166] ① Input data: The formatted spatial data (such as the spatial relationship between monitoring stations) is input into the CNN encoder.
[0167] ② Convolution operation: Local spatial features are extracted through the convolutional layer, and the ReLU activation function is used to introduce non-linearity.
[0168] ③ Pooling operation: The dimension of the feature map is reduced through the max-pooling layer, and significant spatial features are retained.
[0169] ④ Multi-layer stacking: Through the stacking of multiple convolutional layers and pooling layers, spatial features from local to global are gradually extracted.
[0170] 2) TCN encoder
[0171] ① Input data: Input the formatted time series data (such as water quality monitoring data, meteorological data) into the TCN encoder.
[0172] ② Causal convolution: Use causal convolution to ensure the causality of the time series and avoid leakage of future information.
[0173] ③ Dilated convolution: Through convolutional layers with different dilation rates, expand the receptive field and capture dependencies at different time scales.
[0174] ④ Residual connection: Ensure the flow of gradients through residual connections and avoid the problem of gradient vanishing.
[0175] ⑤ Dropout layer: Introduce a Dropout layer during training to reduce the risk of overfitting.
[0176] Data fusion:
[0177] Concatenate the outputs from different TCN or CNN encoders together to form a unified feature representation.
[0178] 1) Feature concatenation and fusion
[0179] Concatenate the output features from different encoders (such as TCN encoder and CNN encoder) to form a unified feature representation. For example, concatenate the encoded features of historical data, upstream data, hydrological data, and meteorological data in the feature dimension.
[0180] Feature dimension alignment: Ensure that the feature dimensions of the outputs from different encoders are consistent for subsequent processing. For example, by adjusting the convolutional kernel size, stride, or padding, align the output features of different encoders in the time step and channel dimensions.
[0181] Sequence prediction:
[0182] Use LSTM for sequence prediction on the fused features. Depending on the prediction target, single-step prediction or multi-step prediction can be selected.
[0183] 1) LSTM modeling
[0184] ① Input features: Input the fused feature vector into the LSTM network.
[0185] ② Temporal modeling: Process the long-term dependencies of time series data through the gating mechanism (input gate, forget gate, output gate) of LSTM.
[0186] ③ Hidden state management: Capture the dynamic changes in the time series through the transmission of hidden states.
[0187] 2) Prediction output
[0188] ①Multi-step prediction: The prediction results for all time steps within a future period are generated at once by the LSTM decoder.
[0189] ②Single-step prediction: After generating the prediction results for all time steps within a future period by the LSTM decoder, only the output of the last time step is taken as the prediction value.
[0190] ③The model flexibly supports single-step and multi-step predictions by adjusting the output parameters of the LSTM.
[0191] Model training and optimization:
[0192] The mean squared error (MSE) or other suitable loss functions are used to evaluate the gap between the prediction results and the actual results. Parameter optimization is carried out by methods such as backpropagation and gradient descent.
[0193] 1) Define the loss function
[0194] The mean squared error (MSE) is selected as the loss function to measure the gap between the predicted value and the true value.
[0195] 2) Forward propagation
[0196] The input data is calculated through the model to generate the predicted output.
[0197] 3) Calculate the loss
[0198] According to the loss function, calculate the gap between the predicted value and the true value to obtain the loss value of the current model.
[0199] 4) Backward propagation
[0200] Calculate the gradient of the loss with respect to the model parameters by the chain rule, preparing for subsequent updates.
[0201] 5) Parameter update
[0202] Use the optimization algorithm of gradient descent to update the model parameters according to the calculated gradient, reducing the loss value.
[0203] 6) Iterative training
[0204] Repeat the processes of forward propagation, calculating the loss, backward propagation, and parameter update until the model performance reaches the expectation or meets the stopping condition.
[0205] 7) Model evaluation and optimization
[0206] Evaluate the model performance on the test set, and further optimize the model by adjusting hyperparameters or adopting regularization and other methods.
[0207] Through the above steps, the training and optimization of the model are completed. The core of the whole process is to continuously adjust the model parameters to make the prediction results closer to the true values.
[0208] A water quality prediction method for a deep learning model based on the same concept, comprising the following steps:
[0209] Data preprocessing: Clean multi-source input data such as the cross-section's own monitoring data, upstream monitoring data of the cross-section, hydrological data, and meteorological data. Use linear interpolation or time series interpolation to fill in missing values, monitor and process outliers through statistical methods, and remove noise using the moving average method; Structure the data, convert time series data into multi-dimensional arrays, and convert spatial data into matrix form. Then use the Min-Max or Z-score normalization method to convert data with different dimensions into the same scale;
[0210] Feature extraction: Use multiple temporal convolutional networks or classical convolutional network encoders to extract features from different types of preprocessed input data. Among them, the classical convolutional network encoder inputs the formatted spatial data, extracts local spatial features through convolutional operations, introduces non-linearity using the ReLU activation function, reduces the dimension of the feature map through the max pooling layer, and stacks multiple layers to gradually extract spatial features from local to global; The temporal convolutional network encoder inputs the formatted time series data, uses causal convolution to ensure the causal relationship of the time series, expands the receptive field through convolutional layers with different dilation rates, and uses residual connections to ensure the flow of gradients. Introduce a Dropout layer during training to reduce the risk of overfitting;
[0211] Data fusion: Concatenate the outputs from different temporal convolutional networks or classical convolutional network encoders to ensure that the feature dimensions of the outputs from different encoders are consistent, forming a unified feature representation;
[0212] Sequence prediction: Use a long short-term memory network to perform sequence prediction on the fused features. Select single-step prediction or multi-step prediction according to the prediction target. The long short-term memory network processes the long-term dependence of time series data through a gating mechanism and captures the dynamic changes in the time series through the transfer of hidden states;
[0213] Model training and optimization: Use the mean squared error (MSE) or other suitable loss functions to evaluate the gap between the prediction results and the actual results, and perform parameter optimization through methods such as backpropagation and gradient descent. During training, repeat the processes of forward propagation, calculating the loss, backpropagation, and parameter update until the model performance reaches the expectation or meets the stopping conditions. Evaluate the model performance on the test set, adjust the hyperparameters, or use regularization methods to further optimize the model.
[0214] Example:
[0215] In a water quality prediction project for a river basin, water quality monitoring data from multiple monitoring sections of the river basin were collected, covering historical data of indicators such as pH value, ammonia nitrogen concentration, total nitrogen concentration, total phosphorus concentration, dissolved oxygen content, etc. At the same time, upstream monitoring data of the section, as well as meteorological data (such as temperature, precipitation, wind speed, etc.) from surrounding meteorological stations and hydrological data (such as water level, flow, etc.) from the hydrological department were obtained.
[0216] First, data preprocessing was performed. For some water quality data with more missing values, time series interpolation method was used to fill them. The Z-score method was used to identify and process outliers. The moving average method was used to remove data noise. The time series water quality monitoring data was organized into a multidimensional array, the spatial data was converted into a matrix form, and the Min-Max normalization method was used to unify the data scale.
[0217] In the feature extraction stage, the formatted spatial data is input into the CNN encoder. SpatialBlock is used as the basic unit, and the number of input channels is configured according to the input data characteristics. As the network layers increase, the number of output channels is gradually adjusted to extract more complex features. A 3×3 convolution kernel is used, the step size is set to 1 in the convolution layer, and the pooling layer is set to 2. The "same" padding method is used, and the maximum pooling operation highlights significant features. The Dropout probability is set to 0.3 during training.
[0218] Multiple SpatialBlocks are stacked to gradually extract spatial features from local to global. The formatted time series data is input into the TCN encoder. TemporalBlock is used as the basic unit. The number of input channels is configured according to the data characteristics, the number of output channels is adjusted, a smaller convolution kernel (such as 3) is used, the step size is set to 1, the expansion rate is gradually increased, and reasonable padding is used to keep the sequence length unchanged. The Dropout probability is set to 0.2, and the time series features are extracted through causal convolution, dilated convolution and other operations.
[0219] Next, the outputs of the CNN and TCN encoders are fused to ensure that the feature dimensions are aligned and then spliced to form a unified feature representation. The fused features are input into the LSTM for sequence prediction. According to actual needs, if you want to predict the water quality changes in the next week, select the multi-step prediction mode; if you only focus on the water quality of the next day, select the single-step prediction mode.
[0220] In the model training and optimization phase, the mean square error (MSE) was used as the loss function, and the back propagation and gradient descent algorithms were used to optimize the parameters. The historical data was divided into training set, validation set, and test set in proportion, and multiple iterations of training were performed on the training set. The hyperparameters were adjusted according to the validation set, and finally the model performance was evaluated on the test set. After multiple rounds of optimization, the model showed high accuracy in the water quality prediction of the basin, and can provide a reliable decision-making basis for water quality management and pollution control in the basin.
[0221] In summary, the beneficial effects of this application are as follows:
[0222] 1. Multi-source data fusion and feature extraction: The time series data is subjected to feature extraction through a Temporal Convolutional Network (TCN), and the spatial data is subjected to feature extraction through a Classical Convolutional Network (CNN). It can also process auxiliary time series data, and there is a feature fusion module, which can effectively fuse the features of multi-source data, comprehensively capture the features of water quality data in time and space, and improve the prediction accuracy.
[0223] 2. Prevention of overfitting: Dropout layers are set in the TemporalBlock module of the TCN, the basic building unit of the SpatialBlock of the CNN, and the LSTM. By randomly discarding some neurons, the risk of model overfitting is reduced, and the generalization ability of the model is enhanced.
[0224] 3. Efficient network structure: The TCN module adopts structures such as a 1D convolutional layer with weight normalization, a ReLU activation function, and a residual connection. The CNN has a feature fusion mechanism, and the LSTM has a hidden state management mechanism. These structural designs enable the model to learn and extract features more efficiently when processing data, improving the performance of the model.
[0225] 4. Flexible prediction method: The output module can be flexibly adjusted to single-step prediction and multi-step prediction, which can meet the water quality prediction requirements in different scenarios and increase the practicality of the model.
[0226] 5. Multi-task learning and adaptive learning rate: The model includes a multi-task learning mechanism, which optimizes multiple related tasks by sharing some network parameters to improve the generalization ability; the adaptive learning rate adjustment mechanism can dynamically adjust the learning rate according to the training loss, accelerating the model convergence and improving the training stability.
Claims
1. A deep learning model for water quality prediction based on spatiotemporal feature fusion and LSTM, characterized in that: include: One or more temporal convolutional networks, used for extracting features from input time series data, wherein the temporal convolutional network comprises a plurality of TemporalBlock modules, each of which comprises two 1D convolutional layers with weight normalization, a ReLU activation function, a Dropout layer, a residual connection and an adaptive average pooling layer, and the convolutional layer weights are initialized using the Xavier initialization method; One or more classic convolutional networks, used for extracting features from input spatial data, wherein the classic convolutional network comprises one or more convolutional layers, a pooling layer, and a fully connected layer, and uses the Xavier initialization method to initialize parameters, and has a feature fusion mechanism, wherein each convolutional layer comprises a convolution operation, a ReLU activation function, a Dropout layer, and a pooling layer; at least one additional temporal convolutional network for processing auxiliary time series data associated with the primary time series data; A long short-term memory network, used for temporal modeling of the output from the temporal convolutional network, wherein the long short-term memory network comprises a long short-term memory network layer and a fully connected layer, and the weights of the long short-term memory network and the fully connected layer are initialized using the Xavier initialization method, and has a hidden state management mechanism; The feature fusion module is used to splice the features extracted by multiple temporal convolutional network modules and input them into the long short-term memory network for time series modeling. It is also used to fuse the time features extracted by the temporal convolutional network module, the spatial features extracted by the classical convolutional network, and the time series modeling results of the long short-term memory network to generate the final prediction output; The output module flexibly adjusts the output into single-step prediction and multi-step prediction to generate the final prediction results.
2. The water quality prediction deep learning model based on spatiotemporal feature fusion and LSTM according to claim 1 is characterized in that: The model also includes a multi-task learning mechanism that improves the generalization ability of the model by sharing some network parameters and optimizing multiple related tasks at the same time; and an adaptive learning rate adjustment mechanism that dynamically adjusts the learning rate according to the change in loss during training to accelerate model convergence and improve training stability, and supports running in a multi-GPU environment to accelerate model training.
3. The water quality prediction deep learning model based on spatiotemporal feature fusion and LSTM according to claim 2 is characterized in that: The specific structure of the TemporalBlock module includes: First convolutional layer: the number of input channels is input_channel, the number of output channels is output_channel, the convolution kernel size is kernel_size, the expansion rate is dilation, and the padding size is (kernel_size-1)*dilation; Second convolutional layer: The number of input channels is the same as the number of output channels, and other parameters are consistent with the first convolutional layer; Downsampling layer: When the number of input channels does not match the number of output channels, 1x1 convolution is used to adjust the number of channels; The ReLU activation function: applied after each convolution layer to introduce nonlinearity; The Dropout layer: Applied after each convolution layer, randomly discarding some neurons to prevent overfitting.
4. The water quality prediction deep learning model based on spatiotemporal feature fusion and LSTM according to claim 2 is characterized in that: The SpatialBlock basic building block of the classic convolutional network is configured as follows: The number of input channels is configured according to the feature dimension of the input data; The number of output channels can be adjusted according to model requirements; The convolution kernel size is 3×3 or 5×5; The stride is usually set to 1 in the convolution layer and can be set to 2 in the pooling layer; The filling method is "same"; The pooling method adopts the maximum pooling operation; The dropout probability is set to 0.2-0.5 during training.
5. The water quality prediction deep learning model based on spatiotemporal feature fusion and LSTM according to claim 2 is characterized in that: The TemporalBlock basic building block of the temporal convolutional network is configured as follows: The number of input channels is configured according to the number of features of the input data; The number of output channels is adjusted according to the needs of the model; The convolution kernel size is 2 or 3; The step size is set to 1; The expansion rate gradually increases; Use appropriate padding to keep the output sequence length unchanged; Dropout probability is used to reduce overfitting.
6. The water quality prediction deep learning model based on spatiotemporal feature fusion and LSTM according to claim 2 is characterized in that: The parameter configuration of the long short-term memory network is as follows: The input size is configured based on the number of input features; The size of the hidden layer determines the size of the state vector inside the LSTM network; The output size is adjusted according to the needs of the model; The number of layers is selected based on model complexity and data size.
7. A water quality prediction method based on the deep learning model according to any one of claims 1 to 6, characterized in that: The following steps are involved: Data preprocessing: Clean the multi-source input data such as the section’s own monitoring data, upstream monitoring data, hydrological data, and meteorological data, use linear interpolation or time series interpolation to fill missing values, monitor and process outliers through statistical methods, and use the moving average method to remove noise; structure the data, convert the time series data into a multidimensional array, convert the spatial data into a matrix form, and use the Z-score normalization method to convert data of different dimensions into the same scale; Feature extraction: Use multiple temporal convolutional networks or classic convolutional network encoders to extract features from different types of preprocessed input data. The classic convolutional network encoder inputs formatted spatial data, extracts local spatial features through convolution operations, introduces nonlinearity using the ReLU activation function, reduces the feature map dimension through the maximum pooling layer, and gradually extracts spatial features from local to global through multi-layer stacking. The temporal convolutional network encoder inputs formatted time series data, uses causal convolution to ensure the causal relationship of the time series, expands the receptive field through convolutional layers with different expansion rates, ensures gradient flow with the help of residual connections, and introduces the Dropout layer during training to reduce the risk of overfitting. Data fusion: splicing the outputs from different temporal convolutional networks or classic convolutional network encoders to ensure that the feature dimensions of the outputs of different encoders are consistent and form a unified feature representation; Sequence prediction: Use the long short-term memory network to perform sequence prediction on the fused features. Select single-step prediction or multi-step prediction according to the prediction target. The long short-term memory network processes the long-term dependency of time series data through the gating mechanism and captures the dynamic changes in the time series through the transmission of hidden states. Model training and optimization: Use mean square error to evaluate the gap between the predicted results and the actual results, optimize parameters through methods such as back propagation and gradient descent, repeat the process of forward propagation, loss calculation, back propagation and parameter update during training until the model performance reaches the expected level or the stopping condition is met, evaluate the model performance on the test set, adjust hyperparameters or use regularization methods to further optimize the model.
Citation Information
Cited By
Process quality time sequence prediction method based on graph neural network
CN116307528A
Basin flood real-time early warning method and system based on multi-source data fusion
CN120974242A
Water quality monitoring method and device for secondary water supply system
CN121258292A
Multi-mode traditional opera motion capture method, device and equipment and storage medium
CN121354219A
Multimodal methods, devices, equipment and storage media for capturing motion in traditional Chinese opera
CN121354219B