A water quality prediction method based on machine learning

By constructing a third-order tensor dataset and a multi-scale fusion machine learning model, the problems of data preprocessing robustness and insufficient feature mining in water quality prediction were solved, and high-precision water quality prediction was achieved.

CN122220745APending Publication Date: 2026-06-16ECOLOGICAL ENVIRONMENT MONITORING & SCI RES CENT OF THE HAIHE RIVER BASIN & BEIHAI SEA ECOLOGICAL ENVIRONMENT SUPERVISION & ADMINISTRATION BUREAU OF THE MINISTRY OF ECOLOGY & ENVIRONMENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ECOLOGICAL ENVIRONMENT MONITORING & SCI RES CENT OF THE HAIHE RIVER BASIN & BEIHAI SEA ECOLOGICAL ENVIRONMENT SUPERVISION & ADMINISTRATION BUREAU OF THE MINISTRY OF ECOLOGY & ENVIRONMENT
Filing Date
2026-03-03
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing water quality prediction methods lack robustness in data preprocessing, have insufficient feature mining, and cannot effectively capture the multi-dimensional coupling characteristics of water quality data. This results in high noise in model training data and low prediction accuracy, making them particularly difficult to apply in small and medium-sized watersheds or areas with missing data.

Method used

By adopting a multi-dimensional innovative coupling design, a third-order tensor dataset is constructed. Temporally weighted isolated forest anomaly detection and Bayesian tensor decomposition algorithms are used to repair missing values. Features are selected by combining maximum information coefficient and grey relational analysis. A multi-scale fusion machine learning model is constructed to capture temporal and spatial correlation features and achieve dynamic deep fusion of spatiotemporal features.

Benefits of technology

It improves the quality of data preprocessing, reduces the computational complexity of the model, and enhances the accuracy and robustness of water quality prediction, especially in scenarios with high missing data and high noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122220745A_ABST
    Figure CN122220745A_ABST
Patent Text Reader

Abstract

The application discloses a water quality prediction method based on machine learning, comprising the following steps: collecting multi-source data of a target river basin, standardizing data samples, screening abnormal values and filling in missing values; extracting time sequence characteristics and spatial characteristics to obtain an initial feature set, and fusing a maximum information coefficient MIC and a grey relational analysis GRA to construct a double-criterion feature screening framework, calculate a comprehensive correlation degree, and screen out a key feature set; constructing a multi-scale fusion machine learning water quality prediction model; a space-time attention fusion module is used for fusing a spatial feature matrix and a time sequence feature matrix, and outputting fused space-time features; a prediction output module is provided with a multi-scale prediction head, and the prediction values of target water quality indexes of each detection station at multiple future time steps are respectively output through a full connection layer. The application solves the core defect that the existing model cannot simultaneously consider time sequence and spatial characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water quality prediction, and more specifically to a water quality prediction method based on machine learning. Background Technology

[0002] With the continuous improvement of the water environment governance system, accurate water quality forecasting has become a core technical support for watershed water environment management, water pollution early warning and source tracing, and water ecological protection and restoration. Water quality time series data has characteristics such as strong nonlinearity, non-stationarity, multi-source heterogeneity, and strong spatiotemporal correlation. Its changes are affected by multiple dimensions of factors such as hydrology, meteorology, and human activities. Traditional forecasting methods are difficult to achieve high-precision, long-term, and robust water quality forecasting.

[0003] In existing technologies, water quality prediction methods are mainly based on hydrodynamic-water quality coupling models with physical mechanisms. These models rely on detailed watershed hydrogeological parameters and pollution source emission data, resulting in high modeling complexity, high computational costs, and extremely high requirements for data completeness. They are difficult to promote and apply in small and medium-sized watersheds or areas with missing data.

[0004] Existing data-driven water quality prediction technologies have the following significant drawbacks: 1. Insufficient robustness in data preprocessing: Most existing methods use simple linear interpolation and KNN interpolation to handle missing values ​​and the 3σ criterion to detect outliers. They do not consider the multi-dimensional coupling characteristics of water quality data in terms of time, space and indicators. This results in low accuracy in repairing high proportions of missing values ​​and sudden outliers, which can easily lead to noise in the model training data and reduce prediction accuracy.

[0005] 2. Insufficient feature mining and failure to simultaneously consider temporal and spatial correlations: Most existing deep learning models only model the time series of a single site, ignoring the spatial propagation characteristics of water quality between upstream and downstream areas and between adjacent monitoring sites within the watershed, and are unable to capture the nonlinear spatial correlations between multiple sites; at the same time, a single time series network is difficult to capture both the short-term local mutation features and long-term trend features of the water quality sequence, resulting in a severe decrease in the accuracy of long-term time series predictions. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a water quality prediction method based on machine learning. Through multi-dimensional innovative coupling design, it solves the defects of existing water quality prediction methods, such as poor robustness of data preprocessing and insufficient mining of spatiotemporal features.

[0007] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A machine learning-based water quality prediction method is provided, comprising: Step S1: Collect data within the target watershed N Each testing site T Each time step,M We constructed a third-order tensor dataset using multi-source data on water quality and related indicators, standardized the data samples, filtered out outliers and filled in missing values, and output a complete third-order tensor standard dataset. Step S2: Extract the temporal and spatial features from the complete third-order tensor standard dataset to obtain the initial feature set. Combine the maximum information coefficient (MIC) and grey relational analysis (GRA) to construct a dual-criteria feature selection framework. Calculate the comprehensive correlation between each feature in the initial feature set and the target prediction index to select the key feature set. Step S3: Construct a multi-scale fusion machine learning water quality prediction model, including a temporal feature extraction module, a spatial correlation feature extraction module, a spatiotemporal attention fusion module, and a prediction output module; The temporal feature extraction module extracts the temporal feature matrix from the key feature set, the spatial correlation feature extraction module extracts the spatial feature matrix from the key feature set, and the spatiotemporal attention fusion module is used to fuse the spatial feature matrix and the temporal feature matrix to output the fused spatiotemporal feature. The prediction output module is equipped with a multi-scale prediction head. The fused spatiotemporal features are input into the prediction output module, and the predicted values ​​of the target water quality indicators of each detection station at multiple future time steps are output through a fully connected layer.

[0008] Further, step S1 includes: Step S11: Collect data within the target watershed N Each testing site T Each time step, M A third-order tensor dataset was constructed from multi-source data of water quality and related indicators. , It is the set of real numbers; Step S12: Process the third-order tensor dataset Data samples Standardize the data samples Convert to standardized value The standard dataset of third-order tensors is obtained. ; ; in, t For time, n For the testing site number, m Number the water quality and related indicators. For the first n The first testing site m The average of each water quality and related indicator over time. For the first n The first testing site m Standard deviation of water quality and related indicators at time step; Step S13: For the standard dataset of third-order tensors Standardized values ​​within Perform time-weighted isolated forest outlier detection, filter out outliers, mark them as missing values, and output a standard dataset of third-order tensors. ; Step S14: Employ Bayesian tensor decomposition to simultaneously mine multi-dimensional correlation information across time, space, and indicators in the third-order tensor standard dataset. Missing values ​​were repaired to obtain the complete standard dataset of third-order tensors. .

[0009] Furthermore, step S13 specifically includes: Step S131: Set the sliding window length during outlier detection. k Based on the standard dataset of third-order tensors Time series data volume calculation time step T Average path length of a binary search tree under certain conditions ; ; in, For harmonic numbers, It is the Euler-Macheroni constant; Step S132: Treat a sliding window as an isolated search tree and obtain the standardized values. In length of k The path length required for successful isolation on an isolated search tree And calculate the standardized value. Expected path length in all isolated search trees ; Step S133: Using the expected value and average path length Calculate data samples Basic abnormal score ; ; Step S134: Obtain the standardized value within each sliding window the median of Interquartile distance Calculate the time-series weighting factor Through time-series weighting factors Constraining the interference of temporal fluctuations on anomaly detection; ; in, This is the median calculation function. This is a function for calculating the interquartile range of data. Step S135: Utilize time-series weighting factors Calculate data samples Time-weighted outlier score ; ; Step S136: Set the outlier score threshold for the data samples. ; like Then determine the data sample These are outliers, and their corresponding standardized values ​​are... Mark as missing value; Otherwise, determine the data sample. Normal values, maintain standardized values constant.

[0010] Furthermore, step S14 specifically includes: Step S141: For the standard dataset of third-order tensors Perform Tucker decomposition and construct a tensor generation model; ; in, For the core tensor, The decomposition ranks are respectively for time, detection sites, and indicator dimensions. The factor matrices are for time, detection sites, and indicator dimensions, respectively. , These are the n-mode product operations of tensors, The noise tensor follows a Gaussian distribution. , For noise accuracy, Unit tensor; Step S142: For the core tensor Factor matrix and noise accuracy By applying a conjugate prior distribution and maximizing the lower bound of evidence using a variational Bayesian inference method, the posterior distribution of each parameter is iteratively solved, ultimately yielding a standard dataset of a third-order tensor with missing values ​​filled in. ; ; in, These are the precision hyperparameters of the prior distributions of each parameter. It is a gamma distribution. These are the initial values ​​for the hyperparameters.

[0011] Furthermore, a dual-criteria feature selection framework is constructed by integrating the Maximum Information Coefficient (MIC) and Grey Relational Analysis (GRA). This framework calculates the comprehensive correlation between each feature in the initial feature set and the target prediction index, thereby selecting the key feature set. Specifically, this includes: Step S21: Calculate the first feature in the initial feature set. u One feature to be screened and target prediction index Y Grey correlation coefficient between ; ; in, For the standardized sequence of the target prediction index, For the first u A standardized sequence of features to be screened. The resolution coefficient; Step S22: Utilize the grey relational coefficient Calculate grey relational degree ; ; Step S23: Calculate the first... u One feature to be screened and target prediction index Y Maximum information coefficient between ; ; in, For the first u The feature sequence of each feature to be screened at each time step. B The number of grid divisions on the feature axis to be filtered. C The number of grid divisions on the target prediction index axis. Represents a grid The next u One feature to be screened and target prediction index Y Mutual information between them Represents the normalization factor. U This represents the total number of features to be selected in the initial feature set; Step S24: Based on the maximum information coefficient and grey relational degree Calculate the overall correlation degree ; ; in, These are the weight coefficients for the maximum information coefficient and the grey relational degree, respectively. Step S25: Set the correlation threshold ;like Then determine the first u One feature to be screened and target prediction index YThe correlation between them is strong; otherwise, determine the first one. u One feature to be screened and target prediction index Y The correlation between them is small; Step S26: Select highly correlated features from the initial feature set and construct the target prediction index. Y Corresponding key feature set , F The feature dimension is the key feature set.

[0012] Furthermore, the temporal feature extraction module employs a fused temporal convolutional network (TCN), which expands the receptive field through dilated convolutions to output features from the key feature set. Extracted temporal feature matrix Specifically, it includes the following steps: The formula for a single-layer dilated causal convolution is: ; in, To extract key feature sets The temporal features of the input fusion temporal convolutional network (TCN) are used to input the temporal features of the input network. l The number of convolutional layers. k Number the convolution kernel. K The number of convolution kernels, For convolution kernel weights, For bias terms, This is the output of the convolutional layer. It is a convolution activation function; The output of the convolutional layer Residual connections are used to avoid gradient vanishing, and the output is derived from the key feature set. Extracted temporal feature matrix ; ; in, For layer normalization operation, This involves a two-layer dilated causal convolution operation. d t Time series feature matrix The temporal feature dimension in the data.

[0013] Furthermore, the spatial correlation feature extraction module employs an improved multi-head graph attention network (GAT) with spatial weight constraints to capture the nonlinear spatial correlations between monitoring stations in the target watershed, outputting key feature sets. The spatial feature matrix extracted from , The dimension of the spatial feature matrix is ​​given; the specific steps include: Extracting key feature sets The Middle nTemporal characteristics of each site For time series characteristics Perform linear transformation to enhance feature representation and output the transformed temporal features. ; ; in, The weight matrix is ​​a linear transformation matrix; One of the testing sites i For the central station, calculate the original attention coefficients of two neighboring detection stations. ; ; in, This is the spatial attention weight vector. Indicates testing site i Temporal characteristics With testing sites j Temporal characteristics splicing operation, It is a linear rectification activation function; Introducing inverse distance space weight constraints for the original attention coefficients Normalization is performed to obtain the normalized attention weights. ; ; in, To detect the elements of the inverse distance weight matrix between sites, For testing sites i The set of neighboring detection sites; use V The head-attention mechanism concatenates the outputs of each attention head to obtain the final spatial feature matrix, outputting the spatial feature matrix of all detection stations. ; ; in, v Number the attention head. Let V be the linear transformation weight matrix for the v-th attention head. No. j The temporal characteristics of each site For the first v Attention weights for each attention head This is the Sigmoid activation function.

[0014] Furthermore, the spatiotemporal attention fusion module dynamically allocates fusion weights for temporal and spatial features through a gating attention mechanism, and applies these weights to the spatial feature matrix. and time series feature matrix Perform fusion and output fused spatiotemporal features , To fuse feature dimensions, the specific steps include: For spatial characteristic matrix and time series feature matrix A linear transformation is performed to align the feature dimensions, resulting in a temporal feature matrix. Spatial feature matrix ; Calculate the gated activation values ​​and adaptive fusion weights of temporal and spatial features. ; ; in, This is the gating activation value. For the gated weight matrix, This is the bias term for the gated activation function; Using adaptive fusion weights Output fused spatiotemporal features ; ; in, This is an element-wise product operation.

[0015] The beneficial effects of this invention are as follows: This invention proposes a time-weighted isolated forest anomaly detection algorithm, which reduces the misjudgment rate of normal fluctuations by combining the correlation between water quality time series; it adopts a Bayesian tensor completion algorithm, and at the same time utilizes the multi-dimensional coupling characteristics of time, space and indicators to repair missing values, which greatly improves the data preprocessing quality in high missing value and high noise scenarios.

[0016] This invention integrates the maximum information coefficient (MIC) and grey relational analysis (GRA) to simultaneously capture the nonlinear correlation and trend similarity between features and target water quality indicators, thereby achieving accurate removal of redundant features, reducing model computational complexity while improving prediction accuracy.

[0017] This invention constructs a spatiotemporal attention multi-scale fusion model architecture to simultaneously capture multi-scale temporal features, captures nonlinear spatial correlation features through an improved GAT with spatial weight constraints, and designs an adaptive gated attention mechanism to achieve dynamic deep fusion of spatiotemporal features, thus solving the core defect of existing models that cannot simultaneously take into account both temporal and spatial characteristics. Attached Figure Description

[0018] Figure 1 This is a flowchart of a machine learning-based water quality prediction method. Detailed Implementation

[0019] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0020] like Figure 1 As shown, a water quality prediction method based on machine learning includes: Step S1: Collect data within the target watershed N Each testing site T Each time step, M Multi-source data on water quality and related indicators were used to construct a third-order tensor dataset. The data samples were standardized, outliers were filtered, and missing values ​​were imputed, resulting in a complete standard third-order tensor dataset. Step S1 specifically includes: Step S11: Collect data within the target watershed N Each testing site T Each time step, M A third-order tensor dataset was constructed from multi-source data of water quality and related indicators. , It is the set of real numbers; The time step in this embodiment T This represents the total length of the time series, with units matching the sampling step size (hours / day). N This represents the total number of water quality monitoring stations within the basin. M This refers to the total number of water quality and related indicators, including core water quality indicators (pH, dissolved oxygen DO, permanganate index CODMn, ammonia nitrogen NH3-N, total phosphorus TP, total nitrogen TN, turbidity, water temperature), hydrological indicators (flow rate, flow velocity, water level), meteorological indicators (rainfall, air temperature, wind speed, air pressure, sunshine duration), and sewage outlet load indicators, etc.

[0021] Step S12: Process the third-order tensor dataset Data samples Standardize the data samples Convert to standardized value The standard dataset of third-order tensors is obtained. ; ; in, t For time, n For the testing site number, m Number the water quality and related indicators. For the first n The first testing site mThe average of each water quality and related indicator over time. For the first n The first testing site m The standard deviation of each water quality and related indicator over time step, and setting To avoid the denominator being 0; Step S13: For the standard dataset of third-order tensors Standardized values ​​within Perform time-weighted isolated forest outlier detection, filter out outliers, mark them as missing values, and output a standard dataset of third-order tensors. Specifically, it includes the following steps: Step S131: Set the sliding window length during outlier detection. k Based on the standard dataset of third-order tensors Time series data volume calculation time step T Average path length of a binary search tree under certain conditions ; ; in, For harmonic numbers, It is the Euler-Macheroni constant; Step S132: Treat a sliding window as an isolated search tree and obtain the standardized values. In length of k The path length required for successful isolation on an isolated search tree And calculate the standardized value. Expected path length in all isolated search trees ; Path length The longer the length, the more difficult it is for the data sample to be isolated. The more normal, the shorter the path length. The shorter the sample, the easier it is for the data sample to be isolated. The more abnormal.

[0022] Step S133: Using the expected value and average path length Calculate data samples Basic abnormal score ; ; Step S134: Obtain the standardized value within each sliding window the median of Interquartile distance Calculate the time-series weighting factor Through time-series weighting factors Constraining the interference of temporal fluctuations on anomaly detection; ; in, This is the median calculation function. This is a function for calculating the interquartile range of data. Step S135: Utilize time-series weighting factors Calculate data samples Time-weighted outlier score ; ; Step S136: Set the outlier score threshold for the data samples. ; like Then determine the data sample These are outliers, and their corresponding standardized values ​​are... Mark as missing value; Otherwise, determine the data sample. Normal values, maintain standardized values constant.

[0023] Step S14: Employ Bayesian tensor decomposition to simultaneously mine multi-dimensional correlation information across time, space, and indicators in the third-order tensor standard dataset. Missing values ​​were repaired to obtain the complete standard dataset of third-order tensors. Specifically, it includes the following steps: Step S141: For the standard dataset of third-order tensors Perform Tucker decomposition and construct a tensor generation model; ; in, For the core tensor, The decomposition ranks are respectively for time, detection sites, and indicator dimensions. The factor matrices are for time, detection sites, and indicator dimensions, respectively. , These are the n-mode product operations of tensors, The noise tensor follows a Gaussian distribution. , For noise accuracy, Unit tensor; Step S142: For the core tensor Factor matrix and noise accuracy By applying a conjugate prior distribution and maximizing the lower bound of evidence using a variational Bayesian inference method, the posterior distribution of each parameter is iteratively solved, ultimately yielding a standard dataset of a third-order tensor with missing values ​​filled in. ; ; in, These are the precision hyperparameters of the prior distributions of each parameter. It is a gamma distribution. These are the initial values ​​for the hyperparameters.

[0024] This invention proposes a time-weighted isolated forest anomaly detection algorithm, which reduces the false positive rate of normal fluctuations by combining the correlation between water quality time series; it also adopts a Bayesian tensor completion algorithm and utilizes the multi-dimensional coupling characteristics of time, space and indicators to repair missing values, which greatly improves the data preprocessing quality in high missing value and high noise scenarios.

[0025] Step S2: For the standard dataset of third-order tensors Standardized values ​​in Temporal and spatial features are extracted to obtain an initial feature set. The maximum information coefficient (MIC) and grey relational analysis (GRA) are then integrated to construct a dual-criteria feature selection framework. The comprehensive correlation between each feature in the initial feature set and the target prediction index is calculated to select the key feature set.

[0026] The temporal features in this embodiment include a standard dataset of third-order tensors. Standardized values ​​in The corresponding sine and cosine position coding features, sliding window statistical features (mean, variance, range, skewness, kurtosis, first / second difference mean), and STL seasonal-trend decomposition features (trend term, seasonal term, residual term).

[0027] The spatial feature in this embodiment is a spatial adjacency matrix constructed based on the geographical coordinates of the monitoring stations and using the inverse distance weighting method. Spatial adjacency matrix Middle elements The calculation method is as follows: ; in, These are the numbers of two neighboring detection stations within the target watershed. The Euclidean distance between two neighboring detection sites. For distance attenuation coefficient, this embodiment takes... .

[0028] A dual-criteria feature selection framework is constructed by integrating the Maximum Information Coefficient (MIC) and Grey Relational Analysis (GRA). This framework calculates the comprehensive correlation between each feature in the initial feature set and the target prediction index, thereby selecting the key feature set. Specifically, this includes: Step S21: Calculate the first feature in the initial feature set. u One feature to be screened and target prediction index Y Grey correlation coefficient between ; ; in, For the standardized sequence of the target prediction index, For the first u A standardized sequence of features to be screened. For the resolution coefficient, this embodiment takes... ; Step S22: Utilize the grey relational coefficient Calculate grey relational degree ; ; Step S23: Calculate the first... u One feature to be screened and target prediction index Y Maximum information coefficient between ; ; in, For the first u The feature sequence of each feature to be screened at each time step. B The number of grid divisions on the feature axis to be filtered. C The number of grid divisions on the target prediction index axis. Represents a grid The next u One feature to be screened and target prediction index Y Mutual information between them This represents the normalization factor, which eliminates the influence of grid division scale. U This represents the total number of features to be selected in the initial feature set; Step S24: Based on the maximum information coefficient and grey relational degree Calculate the overall correlation degree ; ; in, These are the weighting coefficients for the maximum information coefficient and the grey relational degree, respectively. In this embodiment, we take... By default, the maximum information coefficient and the grey relational degree are of equal importance. Step S25: Set the correlation threshold ;like Then determine the first u One feature to be screened and target prediction index Y The correlation between them is strong; otherwise, determine the first one. u One feature to be screened and target prediction index Y The correlation between them is small; Step S26: Select highly correlated features from the initial feature set and construct the target prediction index. Y Corresponding key feature set , F The feature dimension is the key feature set.

[0029] This invention integrates the maximum information coefficient (MIC) and grey relational analysis (GRA) to simultaneously capture the nonlinear correlation and trend similarity between features and target water quality indicators, thereby achieving accurate removal of redundant features, reducing model computational complexity while improving prediction accuracy.

[0030] Step S6: Construct a multi-scale fusion machine learning water quality prediction model, including a temporal feature extraction module, a spatial correlation feature extraction module, a spatiotemporal attention fusion module, and a prediction output module; The temporal feature extraction module is used to extract key features from the set of features. Extracting the temporal feature matrix The spatial correlation feature extraction module is used to extract key feature sets. Extracting spatial feature matrix The spatiotemporal attention fusion module is used to integrate the spatial feature matrix. and time series feature matrix Perform fusion and output fused spatiotemporal features The fused spatiotemporal features of each detection site are obtained. ; The prediction output module is equipped with a multi-scale prediction head that fuses spatiotemporal features. In the input prediction output module, multiple future time steps are output through fully connected layers. Predicted values ​​of target water quality indicators at each monitoring station ; ; in, These are the weights and biases for the fully connected layer.

[0031] In this embodiment, the temporal feature extraction module employs a fused temporal convolutional network (TCN), which expands the receptive field through dilated convolutions to output key feature sets. Extracted temporal feature matrix Specifically, it includes the following steps: The formula for a single-layer dilated causal convolution is: ; in, To extract key feature sets The temporal features of the input fusion temporal convolutional network (TCN) are used to input the temporal features of the input network. l The number of convolutional layers. k Number the convolution kernel. KThe number of convolution kernels, For convolution kernel weights, For bias terms, This is the output of the convolutional layer. It is a convolution activation function; The output of the convolutional layer Residual connections are used to avoid gradient vanishing, and the output is derived from the key feature set. Extracted temporal feature matrix ; ; in, For layer normalization operation, This involves a two-layer dilated causal convolution operation. d t Time series feature matrix The temporal feature dimension in the data.

[0032] The spatial correlation feature extraction module in this embodiment employs an improved multi-head graph attention network (GAT) with spatial weight constraints to capture the nonlinear spatial correlations between monitoring stations in the target watershed and outputs features derived from the key feature set. The spatial feature matrix extracted from , The dimension of the spatial feature matrix is ​​given; the specific steps include: Extracting key feature sets The Middle n Temporal characteristics of each site For time series characteristics Perform linear transformation to enhance feature representation and output the transformed temporal features. ; ; in, The weight matrix is ​​a linear transformation matrix; One of the testing sites i For the central station, calculate the original attention coefficients of two neighboring detection stations. ; ; in, This is the spatial attention weight vector. Indicates testing site i Temporal characteristics With testing sites j Temporal characteristics splicing operation, It is a linear rectification activation function; Introducing inverse distance space weight constraints for the original attention coefficients Normalization is performed to obtain the normalized attention weights. ; ; in, To detect the elements of the inverse distance weight matrix between sites, For testing sites i The set of neighboring detection sites; use V The head-attention mechanism concatenates the outputs of each attention head to obtain the final spatial feature matrix, outputting the spatial feature matrix of all detection stations. ; ; in, v Number the attention head. Let V be the linear transformation weight matrix for the v-th attention head. No. j The temporal characteristics of each site For the first v Attention weights for each attention head This is the Sigmoid activation function.

[0033] The spatiotemporal attention fusion module dynamically allocates fusion weights for temporal and spatial features through a gating attention mechanism, and applies these weights to the spatial feature matrix. and time series feature matrix Perform fusion and output fused spatiotemporal features , To fuse feature dimensions, the specific steps include: For spatial characteristic matrix and time series feature matrix A linear transformation is performed to align the feature dimensions, resulting in a temporal feature matrix. Spatial feature matrix ; Calculate the gated activation values ​​and adaptive fusion weights of temporal and spatial features. ; ; in, This is the gating activation value. For the gated weight matrix, This is the bias term for the gated activation function; Using adaptive fusion weights Output fused spatiotemporal features ; ; in, This is an element-wise product operation.

[0034] This invention constructs a spatiotemporal attention multi-scale fusion model architecture to simultaneously capture multi-scale temporal features, captures nonlinear spatial correlation features through an improved GAT with spatial weight constraints, and designs an adaptive gated attention mechanism to achieve dynamic deep fusion of spatiotemporal features, thus solving the core defect of existing models that cannot simultaneously take into account both temporal and spatial characteristics.

Claims

1. A water quality prediction method based on machine learning, characterized in that, include: Step S1: Collect data within the target watershed N Each testing site T Each time step, M We constructed a third-order tensor dataset using multi-source data on water quality and related indicators, standardized the data samples, filtered out outliers and filled in missing values, and output a complete third-order tensor standard dataset. Step S2: Extract the temporal and spatial features from the complete third-order tensor standard dataset to obtain the initial feature set. Combine the maximum information coefficient (MIC) and grey relational analysis (GRA) to construct a dual-criteria feature selection framework. Calculate the comprehensive correlation between each feature in the initial feature set and the target prediction index to select the key feature set. Step S3: Construct a multi-scale fusion machine learning water quality prediction model, including a temporal feature extraction module, a spatial correlation feature extraction module, a spatiotemporal attention fusion module, and a prediction output module; The temporal feature extraction module extracts a temporal feature matrix from the key feature set, the spatial correlation feature extraction module extracts a spatial feature matrix from the key feature set, and the spatiotemporal attention fusion module is used to fuse the spatial feature matrix and the temporal feature matrix to output fused spatiotemporal features; The prediction output module is equipped with a multi-scale prediction head, which inputs the fused spatiotemporal features into the prediction output module and outputs the predicted values ​​of the target water quality indicators of each detection station at multiple future time steps through a fully connected layer.

2. The water quality prediction method based on machine learning according to claim 1, characterized in that, Step S1 includes: Step S11: Collect data within the target watershed N Each testing site T Each time step, M A third-order tensor dataset was constructed from multi-source data of water quality and related indicators. , It is the set of real numbers; Step S12: Process the third-order tensor dataset Data samples Standardize the data samples Convert to standardized value The standard dataset of third-order tensors is obtained. ; ; in, t For time, n For the testing site number, m Number the water quality and related indicators. For the first n The first testing site m The average of each water quality and related indicator over time. For the first n The first testing site m Standard deviation of water quality and related indicators at time step; Step S13: For the standard dataset of third-order tensors Standardized values ​​within Perform time-weighted isolated forest outlier detection, filter out outliers, mark them as missing values, and output a standard dataset of third-order tensors. ; Step S14: Employ Bayesian tensor decomposition to simultaneously mine multi-dimensional correlation information across time, space, and indicators in the third-order tensor standard dataset. Missing values ​​were repaired to obtain the complete standard dataset of third-order tensors. .

3. The water quality prediction method based on machine learning according to claim 2, characterized in that, Step S13 specifically includes: Step S131: Set the sliding window length during outlier detection. k Based on the standard dataset of third-order tensors Time series data volume calculation time step T Average path length of a binary search tree under certain conditions ; ; in, For harmonic numbers, It is the Euler-Macheroni constant; Step S132: Treat a sliding window as an isolated search tree and obtain the standardized values. In length k The path length required for successful isolation on an isolated search tree And calculate the standardized value. Expected path length in all isolated search trees ; Step S133: Using the expected value and average path length Calculate data samples Basic abnormal score ; ; Step S134: Obtain the standardized value within each sliding window the median of Interquartile distance Calculate the time-series weighting factor Through time-series weighting factors Constraining the interference of temporal fluctuations on anomaly detection; ; in, This is the median calculation function. This is a function for calculating the interquartile range of data. Step S135: Utilize time-series weighting factors Calculate data samples Time-weighted outlier score ; ; Step S136: Set the outlier score threshold for the data samples. ; like Then determine the data sample These are outliers, and their corresponding standardized values ​​are... Mark as missing value; Otherwise, determine the data sample. Normal values, maintain standardized values constant.

4. The water quality prediction method based on machine learning according to claim 2, characterized in that, Step S14 specifically includes: Step S141: For the standard dataset of third-order tensors Perform Tucker decomposition and construct a tensor generation model; ; in, For the core tensor, The decomposition ranks are respectively for time, detection sites, and indicator dimensions. The factor matrices are for time, detection sites, and indicator dimensions, respectively. , These are the n-mode product operations of tensors, The noise tensor follows a Gaussian distribution. , For noise accuracy, Unit tensor; Step S142: For the core tensor Factor matrix and noise accuracy By applying a conjugate prior distribution and maximizing the lower bound of evidence using a variational Bayesian inference method, the posterior distribution of each parameter is iteratively solved, ultimately yielding a standard dataset of a third-order tensor with missing values ​​filled in. ; ; in, These are the precision hyperparameters of the prior distributions of each parameter. It is a gamma distribution. These are the initial values ​​for the hyperparameters.

5. The water quality prediction method based on machine learning according to claim 2, characterized in that, The framework for constructing a dual-criteria feature selection method by fusing the Maximum Information Coefficient (MIC) and Grey Relational Analysis (GRA) calculates the comprehensive correlation between each feature in the initial feature set and the target prediction index, thereby selecting the key feature set; specifically including: Step S21: Calculate the first feature in the initial feature set. u One feature to be screened and target prediction index Y Grey correlation coefficient between ; ; in, For the standardized sequence of the target prediction index, For the first u A standardized sequence of features to be screened. The resolution coefficient; Step S22: Utilize the grey relational coefficient Calculate grey relational degree ; ; Step S23: Calculate the first... u One feature to be screened and target prediction index Y Maximum information coefficient between ; ; in, For the first u The feature sequence of each feature to be screened at each time step. B The number of grid divisions on the feature axis to be filtered. C The number of grid divisions on the target prediction index axis. Represents a grid The next u One feature to be screened and target prediction index Y Mutual information between them Represents the normalization factor. U This represents the total number of features to be selected in the initial feature set; Step S24: Based on the maximum information coefficient and grey relational degree Calculate the overall correlation degree ; ; in, These are the weight coefficients for the maximum information coefficient and the grey relational degree, respectively. Step S25: Set the correlation threshold ;like Then determine the first u One feature to be screened and target prediction index Y The correlation between them is strong; otherwise, the first one is judged to be strong. u One feature to be screened and target prediction index Y The correlation between them is small; Step S26: Select highly correlated features from the initial feature set and construct the target prediction index. Y Corresponding key feature set , F The feature dimensions are defined within the key feature set.

6. The water quality prediction method based on machine learning according to claim 5, characterized in that, The temporal feature extraction module employs a fused temporal convolutional network (TCN), which expands the receptive field through dilated convolutions to output key feature sets. Extracted temporal feature matrix ; Specifically, the following steps are included: The formula for a single-layer dilated causal convolution is: ; in, To extract key feature sets The temporal features of the input fusion temporal convolutional network (TCN) are used to input the temporal features of the input network. l The number of convolutional layers. k Number the convolution kernel. K The number of convolution kernels, For convolution kernel weights, For bias terms, This is the output of the convolutional layer. It is the convolution activation function; The output of the convolutional layer Residual connections are used to avoid gradient vanishing, and the output is derived from the key feature set. Extracted temporal feature matrix ; ; in, For layer normalization operation, This involves a two-layer dilated causal convolution operation. d t Time series feature matrix The temporal feature dimension in the data.

7. The water quality prediction method based on machine learning according to claim 6, characterized in that, The spatial correlation feature extraction module employs an improved multi-head graph attention network (GAT) with spatial weight constraints to capture the nonlinear spatial correlations between monitoring stations in the target watershed, and outputs key feature sets. The spatial feature matrix extracted from , is the dimension of the spatial feature matrix; Specifically, the following steps are included: Extracting key feature sets The Middle n Temporal characteristics of each site For time series characteristics Perform linear transformation to enhance feature representation and output the transformed temporal features. ; ; in, The weight matrix is ​​a linear transformation matrix; One of the testing sites i For the central station, calculate the original attention coefficients of two neighboring detection stations. ; ; in, This is the spatial attention weight vector. Indicates testing site i Temporal characteristics With testing sites j Temporal characteristics splicing operation, It is a linear rectification activation function; Introducing inverse distance space weight constraints for the original attention coefficients Normalization is performed to obtain the normalized attention weights. ; ; in, To detect the elements of the inverse distance weight matrix between sites, For testing sites i The set of neighboring detection sites; use V The head-attention mechanism concatenates the outputs of each attention head to obtain the final spatial feature matrix, outputting the spatial feature matrix of all detection stations. ; ; in, v Number the attention head. Let V be the linear transformation weight matrix for the v-th attention head. No. j The temporal characteristics of each site For the first v Attention weights for each attention head This is the Sigmoid activation function.

8. The water quality prediction method based on machine learning according to claim 7, characterized in that, The spatiotemporal attention fusion module dynamically allocates fusion weights for temporal and spatial features through a gating attention mechanism, and applies these weights to the spatial feature matrix. and time series feature matrix Perform fusion and output fused spatiotemporal features , To fuse feature dimensions, the specific steps include: For spatial characteristic matrix and time series feature matrix A linear transformation is performed to align the feature dimensions, resulting in a temporal feature matrix. Spatial feature matrix ; Calculate the gated activation values ​​and adaptive fusion weights of temporal and spatial features. ; ; in, This is the gating activation value. For the gated weight matrix, This is the bias term for the gated activation function; Using adaptive fusion weights Output fused spatiotemporal features ; ; in, This is an element-wise product operation.