River water quality prediction method and system based on machine learning coupling hydrological model
By combining machine learning algorithms and hydrological models, key water quality parameters are screened, an LSTM-WQI model is constructed, and combined with a SWAT model, efficient and accurate prediction and management of river water quality are achieved. This solves the shortcomings of existing water quality assessment systems and promotes environmental protection and economic development.
Patent Information
- Application Number
- CN202510762228.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-10-31
AI Technical Summary
The existing water quality assessment system is inadequate to address the challenges of diverse pollution, and cannot provide targeted and effective guidance for river water quality management. Furthermore, existing models have numerous parameters, high input data requirements, and insufficient accuracy.
By coupling hydrological models with machine learning algorithms, a limit gradient boosting model and a long short-term memory network model are constructed. Key water quality parameters are screened, and combined with soil and water assessment tool models, efficient and accurate prediction of future water quality is achieved.
It enables comprehensive and efficient assessment of river water quality and dynamic prediction of future water quality, saves monitoring costs, provides accurate prediction results for water quality management and governance, and promotes environmental protection and sustainable economic development.
Smart Images

Figure CN120875113A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of water quality prediction technology, and in particular to a method and system for predicting river water quality based on machine learning coupled with a hydrological model. Background Technology
[0002] With the accelerating pace of urbanization, pollution of rivers and other surface waters is becoming increasingly severe, severely hindering sustainable socio-economic development. Scientific evaluation and accurate prediction of river water quality are indispensable components of environmental management, helping to provide reasonable river governance recommendations and proactively mitigate potential environmental pollution risks. However, existing water quality assessment systems struggle to adapt to the diverse challenges of pollution over time, and they lack differentiated assessment systems for water environments under varying natural conditions. This prevents current standards from providing targeted and effective guidance for water environment governance. Therefore, establishing a more site-specific water quality assessment system and achieving accurate prediction and analysis of future water quality is urgently needed.
[0003] The Water Quality Index (WQI) is a standardized tool for assessing water quality by integrating water quality parameters based on mathematical models. It generates a comprehensive index of 0-100 points through multi-dimensional calculations, providing a direct description of the overall water quality and aquatic environment. In recent years, with the development of machine learning, extreme gradient boosting (XGBoost) and long short-term memory (LSTM) models have been extensively studied and gradually applied in water quality monitoring and prediction. The XGBoost model accelerates the search process for split points by optimizing the training data using a block structure, and its feature parallelism strategy demonstrates excellent performance in data feature analysis. The LSTM model, through three gating mechanisms—forget gate, input gate, and output gate—selectively filters and updates the information flow using gating signals, ultimately achieving dynamic control of the unit state. Therefore, the LSTM model performs exceptionally well in processing long-term data series.
[0004] However, in the process of water environment management and governance, in order to achieve accurate prediction of river water quality, it is often necessary to conduct comprehensive and complete monitoring of river water quality parameters. However, an excessive number of water quality monitoring indicators inadvertently increases the investment and operating costs of water environment management. The Soil and Water Assessment Tool (SWAT) model, developed by the USDA Agricultural Research Service, aims to simulate hydrological processes, pollution loads, and the migration and transformation of pollutants in rivers under different land uses and agricultural activities, thereby hoping to reduce the cost of water environment management and monitoring through model simulation. However, due to the large number of model parameters, high input data requirements, and relatively simplified simulation of the underlying hydraulic structure, its accuracy still needs to be supplemented by other models. Summary of the Invention
[0005] This application provides a method and system for predicting river water quality based on a machine learning coupled hydrological model. The method uses machine learning algorithms to couple a hydrological model to construct a water quality assessment model that is easy to use and adaptable to local conditions, thereby achieving efficient and accurate prediction of future water quality.
[0006] To address the aforementioned technical problems, in a first aspect, embodiments of this application provide a river water quality prediction method based on a machine learning-coupled hydrological model, comprising the following steps: First, acquiring a dataset and dividing the dataset into a training set and a test set; the dataset includes water quality index (WQI) values and a set of water quality parameters; then, based on the dataset, constructing a limit gradient boosting model and optimizing the model's hyperparameters to obtain a trained limit gradient boosting model; based on the trained limit gradient boosting model, performing correlation analysis on the water quality parameter set and the WQI values to screen out key water quality parameters affecting the WQI values; then... Next, using the selected key water quality parameters as input variables and the Water Quality Index (WQI) value as the output variable, an LSTM model was constructed, and hyperparameter optimization was performed on the model to obtain a trained LSTM-WQI model. Then, spatial and attribute data of rivers in the study area were acquired to construct a soil and water assessment tool model. Based on the soil and water assessment tool model, future hydrological and water quality processes were simulated, and future water quality parameter results were calculated and output. Finally, the future water quality parameter results were input into the trained LSTM-WQI model to predict future river WQI values, and the prediction results were obtained. The prediction results were then validated based on existing data.
[0007] In some exemplary embodiments, acquiring the dataset includes the following steps: obtaining a water quality parameter set by collecting long-term water quality monitoring data of rivers in the study area; after acquiring the water quality parameter set, organizing the water quality parameter data, removing missing values and outliers, and performing statistical analysis on the water quality parameter data to calculate the water quality index (WQI) value; the statistical analysis includes statistically analyzing the minimum, maximum, average, and standard deviation of each water quality parameter; using the comprehensive water quality index (WQI) as the river water quality evaluation index, calculating the WQI value for all monitoring nodes of the river; the WQI value calculation formula is:
[0008]
[0009] Where n is the number of water quality parameters, W i SI is the weight of the i-th water quality parameter. i is a sub-index of the i-th water quality parameter.
[0010] In some exemplary embodiments, the water quality parameters in the water quality parameter set include water temperature, dissolved oxygen, biochemical oxygen demand, chemical oxygen demand, ammonia nitrogen, suspended solids, pH value, permanganate index, total phosphorus, total nitrogen, zinc, arsenic, cadmium, lead, fluoride, cyanide, volatile phenols, petroleum hydrocarbons, fecal coliforms, conductivity, and redox potential.
[0011] In some exemplary embodiments, based on the dataset, a limiting gradient boosting model is constructed, and hyperparameter optimization is performed on the model to obtain a trained limiting gradient boosting model. This includes: setting up an Anaconda environment, using Python 3.9.18 software, using various water quality parameters as feature values, and WQI values as output values to construct an XGBoost model; the XGBoost model is a limiting gradient boosting model; hyperparameter optimization is performed on the XGBoost model, and statistical indicators of the XGBoost model on the test set are evaluated; the statistical indicators include root mean square error, mean absolute error, coefficient of determination, and Pearson correlation coefficient, to obtain a trained XGBoost model.
[0012] In some exemplary embodiments, the formula for calculating the statistical indicator is as follows:
[0013]
[0014] Where RMSE is the root mean square error, MAE is the mean absolute error, and R0 is the mean square error. 2 WQI is the coefficient of determination, r is the Pearson correlation coefficient, and WQI is the coefficient of determination. pred It is the WQI prediction value fitted by the model, WQI true It is the true WQI value of the input model. It is the average of the predicted WQI values. It is the average of the true WQI values.
[0015] In some exemplary embodiments, constructing an LSTM model and optimizing its hyperparameters to obtain a trained LSTM-WQI model includes: constructing an LSTM model based on a training set and a test set, using selected key water quality parameters as input variables and the water quality index (WQI) value as the output variable; the LSTM model is a long short-term memory network model; standardizing the input water quality parameters, randomly selecting 20% of the training set data as a validation set to assist in LSTM model training, and optimizing the hyperparameters of the LSTM model; evaluating the model performance using statistical indicators on the test set to obtain the trained LSTM-WQI model; the statistical indicators include: root mean square error, mean absolute error, coefficient of determination, and Pearson correlation coefficient.
[0016] In some exemplary embodiments, the calculation formula for standardizing the input water quality parameters is as follows:
[0017]
[0018] Where X represents the original data, μ represents the mean of the original dataset, σ represents the standard deviation of the original dataset, and X′ represents the standardized water quality data.
[0019] In some exemplary embodiments, the process involves acquiring spatial and attribute data of rivers in the study area, constructing a soil and water assessment tool model, simulating future hydrological and water quality processes based on the soil and water assessment tool model, and calculating and outputting future water quality parameter results. This includes: acquiring spatial and attribute data of rivers in the study area and preprocessing the data; the spatial data includes digital elevation model data, land use data, and soil type distribution data; the attribute data includes meteorological data, water quality data, and soil data; constructing a SWAT model based on the preprocessed data; the SWAT model is a soil and water assessment tool model; dividing the study area into sub-basins and hydrological response units according to the digital elevation model data and land use types, and initializing the model using SCS runoff curves and evapotranspiration functions; calculating runoff and water quality parameter loads by inputting meteorological data, water quality data, and soil data into the SWAT model, and validating the results based on the acquired data; and simulating future hydrological and water quality processes using the SWAT model and calculating and outputting future water quality parameter results.
[0020] In some exemplary embodiments, the formula for calculating the SCS runoff curve is:
[0021]
[0022] Among them, Q surf Let R be the surface runoff on day i. dayLet S be the precipitation on day i, S be the maximum possible soil retention, and CN be the runoff curve number on day i; the evapotranspiration function is calculated as follows:
[0023]
[0024] Where S is the retention parameter on day i, S max is the maximum retention parameter on day i, SW is the effective soil water content, w1 is the first form coefficient, w2 is the second form coefficient, FC is the field water holding capacity of the soil profile, and SAT is the saturated water content of the soil profile.
[0025] Secondly, this application also provides a river water quality prediction system based on a machine learning coupled hydrological model, comprising: a dataset module, an XGBoost model building module, an LSTM model building module, a SWAT model building module, and a prediction module connected in sequence; the dataset module is used to acquire a dataset and divide the dataset into a training set and a test set; the dataset includes water quality index (WQI) values and a set of water quality parameters; the XGBoost model building module is used to construct a limit gradient boosting model based on the dataset and optimize the hyperparameters of the model to obtain a trained limit gradient boosting model; based on the trained limit gradient boosting model, correlation analysis is performed on the water quality parameter set and the water quality index (WQI) values to select... The key water quality parameters affecting the Water Quality Index (WQI) value are identified. The LSTM model building module uses these selected key water quality parameters as input variables and the WQI value as the output variable to construct an LSTM model. Hyperparameter optimization is then performed on the model to obtain a trained LSTM-WQI model. The SWAT model building module acquires spatial and attribute data of rivers in the study area and constructs a soil and water assessment tool model. Based on this model, future hydrological and water quality processes are simulated, and future water quality parameters are calculated and output. The prediction module inputs these future water quality parameters into the trained LSTM-WQI model to predict future river WQI values, obtains the prediction results, and validates the prediction results using existing data.
[0026] The technical solution provided in this application has at least the following advantages:
[0027] This application provides a method and system for predicting river water quality based on a machine learning coupled hydrological model. The method includes the following steps: First, acquiring a dataset and dividing it into a training set and a test set; the dataset includes water quality index (WQI) values and a set of water quality parameters; then, based on the dataset, constructing a limit gradient boosting model and optimizing the model's hyperparameters to obtain a trained limit gradient boosting model; based on the trained limit gradient boosting model, performing correlation analysis on the water quality parameter set and the WQI values to screen out key water quality parameters affecting the WQI values; next, using... Key water quality parameters selected through screening were used as input variables, and the Water Quality Index (WQI) value was used as the output variable. An LSTM model was constructed, and hyperparameters were optimized to obtain a trained LSTM-WQI model. Then, spatial and attribute data of rivers in the study area were acquired to construct a soil and water assessment tool model. Based on the soil and water assessment tool model, future hydrological and water quality processes were simulated, and future water quality parameters were calculated and output. Finally, the future water quality parameters were input into the trained LSTM-WQI model to predict future river WQI values, and the prediction results were obtained. The prediction results were then validated based on existing data.
[0028] On the one hand, this application utilizes machine learning models to analyze and select key water quality parameters affecting river water quality. Simultaneously, it constructs machine learning WQI models for these key water quality parameters, enabling comprehensive and efficient assessment of river water quality. This also helps local authorities rationally adjust water quality monitoring plans, focusing on the centralized management of key water quality parameters affecting river water quality, thus saving on personnel, equipment, and related investment and operating costs. On the other hand, this application constructs a SWAT-WQI coupled prediction model, achieving dynamic prediction of future river water quality. This effectively reduces the number of river water quality parameters monitored, while the accurate prediction results of future river water quality can provide suggestions for water quality management and pollution control, achieving high-level, high-efficiency, and high-precision water environment management and governance, and promoting sustainable environmental protection and socio-economic development. Attached Figure Description
[0029] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments, and unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0030] Figure 1 This is a flowchart of a river water quality prediction method based on a machine learning coupled hydrological model, provided in an embodiment of this application.
[0031] Figure 2 This is a schematic diagram illustrating the execution process of the river water quality prediction method based on machine learning coupled with a hydrological model provided in this application embodiment.
[0032] Figure 3 This is a ranking chart of the importance of water quality parameter features obtained based on the XGBoost model, provided in an embodiment of this application.
[0033] Figure 4 This is a heatmap showing the correlation between all monitored water quality parameters and WQI values provided in the embodiments of this application.
[0034] Figure 5 This is a graph showing the fitting results between the WQI predicted value and the actual WQI value based on the LSTM model provided in the embodiments of this application.
[0035] Figure 6 This is a graph showing the SWAT model results for determining the flow rate at the mouth of the Shenzhen River, provided in this embodiment of the application.
[0036] Figure 7 This is a graph showing the WQI verification results of the LSTM-based SWAT model provided in this application embodiment.
[0037] Figure 8 This is a schematic diagram of a river water quality prediction system based on a machine learning coupled hydrological model, provided in an embodiment of this application. Detailed Implementation
[0038] As can be seen from the background technology, the accuracy of existing water quality assessment models still needs to be supplemented by other models due to the large number of parameters, high input data requirements, and relatively simplified simulation of the underlying hydraulic structure.
[0039] To address the aforementioned technical problems, this application provides a river water quality prediction method based on a machine learning-coupled hydrological model. The method includes the following steps: First, acquiring a dataset and dividing it into a training set and a test set; the dataset includes Water Quality Index (WQI) values and a set of water quality parameters; then, based on the dataset, constructing a limit gradient boosting model and optimizing its hyperparameters to obtain a trained limit gradient boosting model; based on the trained limit gradient boosting model, performing a correlation analysis on the water quality parameter set and the WQI values to identify key water quality parameters affecting the WQI values; then… Next, using the selected key water quality parameters as input variables and the Water Quality Index (WQI) value as the output variable, an LSTM model is constructed, and hyperparameter optimization is performed to obtain a trained LSTM-WQI model. Then, spatial and attribute data of rivers in the study area are acquired to construct a soil and water assessment tool model. Based on the soil and water assessment tool model, future hydrological and water quality processes are simulated, and future water quality parameters are calculated and output. Finally, the future water quality parameters are input into the trained LSTM-WQI model to predict future river WQI values, obtaining prediction results, and the prediction results are verified based on existing data. This application provides a river water quality prediction method and system based on machine learning coupled with a hydrological model. It uses machine learning algorithms coupled with a hydrological model to construct an easy-to-use and site-specific water quality assessment model, achieving efficient and accurate prediction of future water quality.
[0040] The embodiments of this application will now be described in detail with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.
[0041] See Figure 1 This application provides a method for predicting river water quality based on a machine learning-coupled hydrological model, comprising the following steps:
[0042] Step S1: Obtain the dataset and divide it into a training set and a test set; the dataset includes the Water Quality Index (WQI) values and a set of water quality parameters.
[0043] Step S2: Based on the dataset, construct the ultimate gradient boosting model and optimize the hyperparameters of the model to obtain the trained ultimate gradient boosting model; based on the trained ultimate gradient boosting model, perform correlation analysis on the water quality parameter set and the water quality index (WQI) value to screen out the key water quality parameters that affect the WQI value.
[0044] Step S3: Using the key water quality parameters obtained from the screening as input variables and the water quality index (WQI) value as the output variable, construct an LSTM model and optimize the hyperparameters of the model to obtain the trained LSTM-WQI model.
[0045] Step S4: Obtain spatial and attribute data of rivers in the study area, and construct a soil and water assessment tool model; simulate future hydrological and water quality processes based on the soil and water assessment tool model, and calculate and output future water quality parameter results.
[0046] Step S5: Input the future water quality parameter results into the trained LSTM-WQI model to predict the future WQI value of the river, obtain the prediction results, and verify the prediction results based on existing data.
[0047] In some embodiments, obtaining the dataset in step S1 includes the following steps: acquiring a water quality parameter set by collecting long-term water quality monitoring data of rivers in the study area; after acquiring the water quality parameter set, organizing the water quality parameter data, removing missing values and outliers, and performing statistical analysis on the water quality parameter data to calculate the water quality index (WQI); the statistical analysis includes statistically analyzing the minimum, maximum, average, and standard deviation of each water quality parameter; using the comprehensive water quality index (WQI) as the river water quality evaluation index, calculating the WQI values of all monitoring nodes in the river; the formula for calculating the WQI value is:
[0048]
[0049] Where n is the number of water quality parameters, W i SI is the weight of the i-th water quality parameter. i is a sub-index of the i-th water quality parameter.
[0050] In some embodiments, the water quality parameters collected include water temperature (WT), dissolved oxygen (DO), biochemical oxygen demand (BOD), chemical oxygen demand (COD), ammonia nitrogen (NH3-N), suspended solids (SS), pH value, and permanganate index (COD). Mn Total phosphorus (TP), total nitrogen (TN), zinc (Zn), arsenic (As), cadmium (Cd), lead (Pb), and fluoride (F) - ), cyanide (CN), volatile phenols (VP), petroleum hydrocarbons (TPH), fecal coliforms (E. coli), electrical conductivity (EC), and redox potential (ORP).
[0051] In some embodiments, step S2 involves constructing a limit gradient boosting model based on the dataset, optimizing the hyperparameters of the model, and obtaining a trained limit gradient boosting model. This includes: setting up an Anaconda environment, using Python 3.9.18 software, constructing an XGBoost model with various water quality parameters as feature values and WQI values as output values; the XGBoost model is a limit gradient boosting model; optimizing the hyperparameters of the XGBoost model; and evaluating the statistical indicators of the XGBoost model on the test set. The statistical indicators include root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 The trained XGBoost model is obtained by using the correlation coefficient (r) and Pearson correlation coefficient (r).
[0052] In some embodiments, the formula for calculating the statistical index in step S2 is:
[0053]
[0054]
[0055] Where RMSE is the root mean square error, MAE is the mean absolute error, and R0 is the mean square error. 2 WQI is the coefficient of determination, r is the Pearson correlation coefficient, and WQI is the coefficient of determination. pred It is the WQI prediction value fitted by the model, WQI true It is the true WQI value of the input model. It is the average of the predicted WQI values. It is the average of the true WQI values.
[0056] After calculating the statistical indicators, a correlation analysis was performed on the various water quality parameters and the WQI value. Based on the analysis results of the XGBoost model, the key water quality parameters affecting the WQI value of the studied river were selected. Then, step S3 was executed, using the selected key water quality parameters as input variables and the WQI value as output variables, to construct a Long Short-Term Memory (LSTM) network model based on the original training and test sets.
[0057] In some embodiments, step S3 involves constructing an LSTM model and optimizing its hyperparameters to obtain a trained LSTM-WQI model. This includes: constructing an LSTM model based on the training and test sets using selected key water quality parameters as input variables and the WQI value as the output variable; the LSTM model is a Long Short-Term Memory (LSTM) network model; standardizing the input water quality parameters; randomly selecting 20% of the training set data as a validation set to assist in LSTM model training; and optimizing the hyperparameters of the LSTM model; evaluating the model performance using statistical metrics on the test set to obtain the trained LSTM-WQI model; wherein the statistical metrics include: Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Determination (R²). 2 ) and Pearson correlation coefficient (r).
[0058] In some exemplary embodiments, the calculation formula for standardizing the input water quality parameters is as follows:
[0059]
[0060] Where X represents the original data, μ represents the mean of the original dataset, σ represents the standard deviation of the original dataset, and X′ represents the standardized water quality data.
[0061] The statistical indicators of the model on the test set are RMSE, MAE, and R. 2 The model performance is evaluated using r to obtain the trained LSTM-WQI model. Then, step S4 is executed to obtain spatial and attribute data of the study area, perform preprocessing, and construct the Soil and Water Assessment Tool (SWAT) model database.
[0062] In some embodiments, step S4 involves acquiring spatial and attribute data of rivers in the study area, constructing a soil and water assessment tool model, simulating future hydrological and water quality processes based on the soil and water assessment tool model, and calculating and outputting future water quality parameter results. This includes: acquiring spatial and attribute data of rivers in the study area and preprocessing the data; the spatial data includes digital elevation model (DEM) data, land use data, and soil type distribution data; the attribute data includes meteorological data, water quality data, and soil data; constructing a SWAT model based on the preprocessed data; the SWAT model is a soil and water assessment tool model; dividing the study area into sub-basins and hydrological response units (HRUs) according to the digital elevation model data and land use types, and initializing the model using SCS runoff curves and evapotranspiration functions; calculating runoff and water quality parameter loads by inputting meteorological data, water quality data, and soil data into the SWAT model, and validating the results based on the acquired data; and simulating future hydrological and water quality processes using the SWAT model and calculating and outputting future water quality parameter results.
[0063] In some exemplary embodiments, the formula for calculating the SCS runoff curve is:
[0064]
[0065] Among them, Q surf Let R be the surface runoff on day i. day Let S be the precipitation on day i, S be the maximum possible soil retention, and CN be the runoff curve number on day i; the evapotranspiration function is calculated as follows:
[0066]
[0067] Where S is the retention parameter on day i, S max is the maximum retention parameter on day i, SW is the effective soil water content, w1 is the first form coefficient, w2 is the second form coefficient, FC is the field water holding capacity of the soil profile, and SAT is the saturated water content of the soil profile.
[0068] Finally, in step S5, the water quality parameter results generated by the SWAT model are input into the constructed key parameter LSTM-WQI model. The prediction results are verified based on existing data. The future WQI value of the river is predicted using the future water quality parameters, thus achieving efficient prediction of the future river water quality.
[0069] The following detailed description of the river water quality prediction method based on machine learning coupled hydrological models provided in this application is illustrated through specific embodiments.
[0070] This application provides a method for predicting river water quality based on a machine learning coupled hydrological model, including the following steps:
[0071] Step 1: The Shenzhen River in Shenzhen City was selected as the study area. Monthly water quality monitoring was conducted at a total of 22 monitoring points along the Shenzhen River from 2021 to 2022. The monitored water quality parameters included water temperature (WT), dissolved oxygen (DO), biochemical oxygen demand (BOD), chemical oxygen demand (COD), ammonia nitrogen (NH3-N), suspended solids (SS), pH value, and permanganate index (COD). Mn Total phosphorus (TP), total nitrogen (TN), zinc (Zn), arsenic (As), cadmium (Cd), lead (Pb), and fluoride (F) - The water quality parameters included cyanide (CN), volatile phenols (VP), petroleum hydrocarbons (TPH), fecal coliforms (E. coli), electrical conductivity (EC), and oxidation-reduction potential (ORP). Water quality data were collected and organized, missing and outlier values were checked, and descriptive statistical analysis was used to analyze the river's water quality characteristics (see Table 1 for details).
[0072] Table 1. Statistical Table of Water Quality Parameters
[0073] Water quality parameters unit Minimum value Maximum value average value Standard deviation WT ℃ 11.1 33.3 25.34 4.52 DO mg / L 0.12 14.3 6.26 2.04 BOD mg / L 0.5 7.9 2.16 1.22 COD mg / L 4 31 12.36 5.13 <![CDATA[NH3-N]]> mg / L 0.01 4.82 0.70 0.79 SS mg / L 0.8 159 25.00 27.13 pH - 6.5 8.9 7.42 0.35 <![CDATA[COD Mn ]]> mg / L 0.5 7.6 3.18 1.19 TP mg / L 0.01 0.68 0.17 0.11 TN mg / L 0.16 12.76 5.87 2.66 Zn mg / L 0.004 0.225 0.017 0.028 As mg / L 0.0002 0.0038 0.0012 0.0006 Cd mg / L 0.00005 0.00278 0.00017 0.00027 Pb mg / L 0.0009 0.0082 0.0003 0.0009 <![CDATA[F - ]]> mg / L 0.05 0.58 0.323 0.097 CN mg / L 0.001 0.009 0.0017 0.0012 VP mg / L 0.0003 0.0024 0.0005 0.0003 TPH mg / L 0 0.16 0.021 0.025 E. coli Units / L 10 24000000 809932 2911810 EC mS / m 7.58 2980 158.83 417.97 ORP mV 175 552 416.22 67.33
[0074] Based on the reviewed high-level domestic and international journal articles, and utilizing the obtained water quality parameter data, including DO, BOD, COD, and NH3-N, the formula was used to... The Water Quality Index (WQI) is calculated, where n is the number of water quality parameters W. i SI is the weight of the i-th water quality parameter. i Let WQI be the sub-index of the i-th water quality parameter. The calculation method and weight allocation coefficients of the sub-index are shown in Table 2. A higher weight indicates that the water quality parameter has a greater impact on river water quality, and a higher WQI value indicates better overall river water quality. The WQI values of all monitoring nodes in the river are calculated and combined with the water quality parameter set to construct a river water quality dataset. The dataset is then divided into an 80% training set and a 20% test set to achieve the best machine learning fit.
[0075] Table 2. Calculation methods and weighting coefficients for WQI sub-indices
[0076]
[0077] Step Two: This application sets up an Anaconda environment and uses Python 3.9.18. Pandas is used for reading and processing the dataset from Step One, and functions from sklearn are used to evaluate model performance. The Extreme Gradient Boosting (XGBoost) model is employed to analyze the feature importance of different water quality parameters relative to the WQI value. A controlled experiment is designed using the controlled variable method to optimize hyperparameters. The optimal robust XGBoost model is obtained based on the model's statistical indicators on the test set. Finally, n_estimators = 160, learning_rate = 0.1, max_depth = 6, and random_state = 50 are selected as the optimal hyperparameters. The output model performance parameters are RMSE = 1.368, MAE = 1.068, and R0 = 50. 2 =0.967 and r=0.984. The XGBoost model analysis yielded the following ranking of the importance of water quality parameters: Figure 3 As shown, the top five most important indicators are, in order: DO, NH3-N, COD, etc. Mn Among the 18 water quality parameters (TP and SS), DO has the greatest impact on water quality. Correlation analysis was performed on all 18 water quality parameters and their corresponding WQI values, and a heatmap showing the correlation between water quality parameters and WQI was generated (e.g., TP and SS). Figure 4 As shown in the figure, the analysis indicates that the water quality parameters with the most significant impact on WQI are DO, TP, NH3-N, and As. Based on the results of machine learning algorithms and correlation analysis, the key water quality parameters were selected as DO, NH3-N, and TP, and these three key water quality parameters were used for subsequent predictive analysis.
[0078] Step 3: Using WQI as the output variable and DO, NH3-N, and TP obtained in Step 2 as input variables for key water quality parameters, a Long Short-Term Memory (LSTM) network model is constructed. After standardizing the input data, 20% of the training set data is randomly selected as the model validation set to assist model training. A two-layer LSTM network with 256 and 128 neurons is constructed based on the sequential function in Tensorflow.Keras, with tanh as the activation function. Subsequently, two fully connected layers are added, with 64 neurons with ReLU activation and 1 output unit. Adam is used as the optimizer for training. The optimal hyperparameter-optimized LSTM model is obtained based on the statistical indicators of the model on the test set. Finally, epochs=300, batch_size=16, and learning_rate=0.001 are selected as the optimal hyperparameter results. The output model performance parameters are RMSE=2.358, MAE=1.875, R 2 =0.901 and r=0.957, the training results are as follows Figure 5 As shown, the results indicate that an LSTM-WQI model for key water quality parameters was successfully constructed, and the model has high reliability.
[0079] Step 4: Acquire spatial and attribute data for the Shenzhen River Basin. Spatial data includes Digital Elevation Model (DEM) data, land use data, and soil type distribution data. Attribute data includes meteorological data, water quality data, and soil data. All data are preprocessed to construct a SWAT model database. The DEM data used in this application originates from 10m resolution raster data from the China Geographic Information Network. Spatial linear information, including contour lines, water systems, and administrative boundaries, is obtained through Digital Line Mapping (DLG) and converted into vector format for recognition by the Geographic Information System (GIS). ArcGIS is used to spatially interpolate the contour lines, generating continuous raster elevation data. Simultaneously, combined with a 3D extension module, a Digital Elevation Irregular Triangular Grid (TIN) model of the Shenzhen River Basin is constructed based on the extracted watershed information elements. The acquired land use data is cropped and reclassified according to the Land Use Status Classification Standard (GB / T 21010-2007), including cultivated land, grassland, wetland, building land, forest land, shrubland, water bodies, and bare land. The soil data used in this application is soil type vector data provided by the Resource and Environment Data Cloud Platform of the Chinese Academy of Sciences. Based on this, a soil type attribute database for the Shenzhen River Basin was re-established, including both physical and chemical attributes. Physical attributes focus on the numerical characterization of soil hydrological behavior, while chemical attributes record the background concentrations of pollutants such as nitrogen and phosphorus. The meteorological data used in this application comes from daily observation sequences of national standard meteorological stations within the basin from 2015 to 2022, covering four basic elements: temperature (maximum / minimum), precipitation, relative humidity, and evaporation. To enhance the simulation accuracy of the evapotranspiration module, solar radiation and wind speed data from the Shenzhen National Climate Observatory were additionally incorporated to compensate for the parameter sensitivity differences of the SWAT built-in database under subtropical monsoon climate conditions. The SUFI-2 optimization algorithm of the SWAT-CUP platform was used to calibrate the SWAT model for the Shenzhen River Basin. Figure 6 As shown, the runoff rate R in the SWAT model is periodically... 2 =0.91, verification period R 2 =0.89, indicating that the SWAT model can effectively characterize the hydrological cycle mechanism of the Shenzhen River Basin. Subsequently, the SWAT model was used to calculate and output future water quality parameters, which were then compiled into a prediction dataset.
[0080] Step 5: Use the output results of the calibrated and validated SWAT model as input variables, and input them into the LSTM-WQI model constructed in Step 3 with DO, NH3-N, and TP as key water quality parameters to further validate the output results of the SWAT model. Figure 7As shown, by comparing with the true WQI value, the R-value of the LSTM-WQI model is... 2 =0.86, indicating that the prediction fitting based on the SWAT model has high reliability, and the performance accuracy of the model simulation reaches 95.6%. The future water quality parameter prediction dataset calculated by the SWAT model obtained in step four is imported into the LSTM-WQI model to predict the future WQI value of the Shenzhen River.
[0081] See Figure 8 This application also provides a river water quality prediction system based on a machine learning coupled hydrological model, comprising: a dataset module 101, an XGBoost model building module 102, an LSTM model building module 103, a SWAT model building module 104, and a prediction module 105 connected in sequence; the dataset module 101 is used to acquire a dataset and divide the dataset into a training set and a test set; the dataset includes water quality index (WQI) values and a set of water quality parameters; the XGBoost model building module 102 is used to construct a limit gradient boosting model based on the dataset and optimize the hyperparameters of the model to obtain a trained limit gradient boosting model; based on the trained limit gradient boosting model, correlation analysis is performed on the water quality parameter set and the water quality index (WQI) values to screen... The key water quality parameters affecting the Water Quality Index (WQI) value are obtained; the LSTM model construction module 103 is used to construct an LSTM model with the selected key water quality parameters as input variables and the WQI value as output variable, and to optimize the hyperparameters of the model to obtain a trained LSTM-WQI model; the SWAT model construction module 104 is used to acquire spatial and attribute data of rivers in the study area and construct a soil and water assessment tool model; based on the soil and water assessment tool model, the future hydrological and water quality process is simulated, and the future water quality parameter results are calculated and output; the prediction module 105 is used to input the future water quality parameter results into the trained LSTM-WQI model to predict the future river WQI value, obtain the prediction results, and verify the prediction results based on existing data.
[0082] Based on the above technical solutions, this application provides a method and system for predicting river water quality based on a machine learning coupled hydrological model. The method includes the following steps: First, acquiring a dataset and dividing it into a training set and a test set; the dataset includes water quality index (WQI) values and a set of water quality parameters; then, based on the dataset, constructing a limit gradient boosting model and optimizing its hyperparameters to obtain a trained limit gradient boosting model; based on the trained limit gradient boosting model, performing correlation analysis on the water quality parameter set and the WQI values to screen out key water quality parameters affecting the WQI values; then… Next, using the selected key water quality parameters as input variables and the Water Quality Index (WQI) value as the output variable, an LSTM model was constructed, and hyperparameter optimization was performed on the model to obtain a trained LSTM-WQI model. Then, spatial and attribute data of rivers in the study area were acquired to construct a soil and water assessment tool model. Based on the soil and water assessment tool model, future hydrological and water quality processes were simulated, and future water quality parameter results were calculated and output. Finally, the future water quality parameter results were input into the trained LSTM-WQI model to predict future river WQI values, and the prediction results were obtained. The prediction results were then validated based on existing data.
[0083] On the one hand, this application utilizes machine learning models to analyze and select key water quality parameters affecting river water quality. Simultaneously, it constructs machine learning WQI models for these key water quality parameters, enabling comprehensive and efficient assessment of river water quality. This also helps local authorities rationally adjust water quality monitoring plans, focusing on the centralized management of key water quality parameters affecting river water quality, thus saving on personnel, equipment, and related investment and operating costs. On the other hand, this application constructs a SWAT-WQI coupled prediction model, achieving dynamic prediction of future river water quality. This effectively reduces the number of river water quality parameters monitored, while the accurate prediction results of future river water quality can provide suggestions for water quality management and pollution control, achieving high-level, high-efficiency, and high-precision water environment management and governance, and promoting sustainable environmental protection and socio-economic development.
[0084] Those skilled in the art will understand that the above-described embodiments are specific examples of implementing this application, and in practical applications, various changes in form and detail may be made without departing from the spirit and scope of this application. Any person skilled in the art can make their own modifications and alterations without departing from the spirit and scope of this application; therefore, the scope of protection of this application should be determined by the scope defined in the claims.
Claims
1. A method for predicting river water quality based on a machine learning-coupled hydrological model, characterized in that, Includes the following steps: Obtain the dataset and divide it into a training set and a test set; the dataset includes the Water Quality Index (WQI) values and a set of water quality parameters. Based on the dataset, a limit gradient boosting model is constructed, and the hyperparameters of the model are optimized to obtain the trained limit gradient boosting model. Based on the trained extreme gradient boosting model, a correlation analysis was conducted on the water quality parameter set and the water quality index (WQI) value to screen out the key water quality parameters that affect the WQI value. Using the key water quality parameters obtained through screening as input variables and the water quality index (WQI) value as output variable, an LSTM model is constructed, and the hyperparameters of the model are optimized to obtain the trained LSTM-WQI model. Acquire spatial and attribute data of rivers in the study area, construct a soil and water assessment tool model; simulate future hydrological and water quality processes based on the soil and water assessment tool model, and calculate and output future water quality parameters. The future water quality parameters are input into the trained LSTM-WQI model to predict the future WQI value of the river, and the prediction results are verified based on existing data.
2. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 1, characterized in that, Obtaining the dataset involves the following steps: A set of water quality parameters was obtained by collecting long-term water quality monitoring data of rivers in the study area; After obtaining the water quality parameter set, the water quality parameter data is organized. After removing missing values and outliers, the water quality parameter data is statistically analyzed to calculate the water quality index (WQI) value. The statistical analysis includes the minimum, maximum, average, and standard deviation of various water quality parameters; the comprehensive water quality index (WQI) is used as the river water quality evaluation index, and the WQI values of all monitoring nodes of the river are calculated. The formula for calculating WQI value is: Where n is the number of water quality parameters, W i SI is the weight of the i-th water quality parameter. i is a sub-index of the i-th water quality parameter.
3. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 1, characterized in that, The water quality parameters collected include water temperature, dissolved oxygen, biochemical oxygen demand, chemical oxygen demand, ammonia nitrogen, suspended solids, pH value, permanganate index, total phosphorus, total nitrogen, zinc, arsenic, cadmium, lead, fluoride, cyanide, volatile phenols, petroleum hydrocarbons, fecal coliforms, conductivity, and redox potential.
4. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 1, characterized in that, Based on the dataset, a limit gradient boosting model is constructed, and hyperparameters are optimized to obtain the trained limit gradient boosting model, including: An Anaconda environment was set up, and Python 3.9.18 software was used to construct an XGBoost model with various water quality parameters as feature values and WQI values as output values. The XGBoost model is an extreme gradient boosting model. The hyperparameters of the XGBoost model are optimized, and the statistical metrics of the XGBoost model on the test set are evaluated. The statistical metrics include root mean square error, mean absolute error, coefficient of determination, and Pearson correlation coefficient, to obtain the trained XGBoost model.
5. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 4, characterized in that, The formula for calculating statistical indicators is: Where RMSE is the root mean square error, MAE is the mean absolute error, and R0 is the mean square error. 2 WQI is the coefficient of determination, r is the Pearson correlation coefficient, and WQI is the coefficient of determination. pred It is the WQI prediction value fitted by the model, WQI true It is the actual WQI value input to the model. It is the average of the predicted WQI values. It is the average of the true WQI values.
6. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 1, characterized in that, Construct an LSTM model and optimize its hyperparameters to obtain a trained LSTM-WQI model, including: Using the key water quality parameters obtained through screening as input variables and the water quality index (WQI) value as output variable, an LSTM model is constructed based on the training and testing sets; the LSTM model is a long short-term memory network model. The input water quality parameters are standardized, and 20% of the training set data is randomly selected as the validation set to assist in the training of the LSTM model. The hyperparameters of the LSTM model are then optimized. The performance of the LSTM model is evaluated by statistical indicators on the test set to obtain the trained LSTM-WQI model. The statistical indicators include: root mean square error, mean absolute error, coefficient of determination, and Pearson correlation coefficient.
7. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 6, characterized in that, The calculation formula for standardizing the input water quality parameters is as follows: Where X represents the original data, μ represents the mean of the original dataset, σ represents the standard deviation of the original dataset, and X′ represents the standardized water quality data.
8. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 1, characterized in that, Spatial and attribute data of rivers in the study area were acquired, and a soil and water assessment tool model was constructed. Based on the soil and water assessment tool model, future hydrological and water quality processes were simulated, and future water quality parameters were calculated and output, including: Spatial and attribute data of rivers in the study area were acquired and preprocessed. The spatial data included digital elevation model data, land use data, and soil type distribution data. The attribute data included meteorological data, water quality data, and soil data. Based on the preprocessed data, a SWAT model is constructed; the SWAT model is a soil and water assessment tool model. Based on digital elevation model data and land use types, the study area is divided into sub-basins and hydrological response units, and the model is initialized using SCS runoff curves and evapotranspiration functions. By inputting meteorological, water quality, and soil data into the SWAT model, runoff and water quality parameter loads are calculated and validated based on the acquired data. The SWAT model is then used to simulate future hydrological and water quality processes, and the results of future water quality parameters are calculated and output.
9. The river water quality prediction method based on machine learning coupled with a hydrological model according to claim 8, characterized in that, The formula for calculating the SCS runoff curve is: Among them, Q surf Let R be the surface runoff on day i. day Let S be the precipitation on day i, S be the maximum possible soil retention, and CN be the runoff curve number on day i. The formula for calculating the evapotranspiration function is: Where S is the retention parameter on day i, S max is the maximum retention parameter on day i, SW is the effective soil water content, w1 is the first form coefficient, w2 is the second form coefficient, FC is the field water holding capacity of the soil profile, and SAT is the saturated water content of the soil profile.
10. A river water quality prediction system based on a machine learning coupled hydrological model, characterized in that, include: The dataset module, XGBoost model building module, LSTM model building module, SWAT model building module, and prediction module are connected sequentially. The dataset module is used to acquire a dataset and divide the dataset into a training set and a test set; the dataset includes the Water Quality Index (WQI) values and a set of water quality parameters. The XGBoost model building module is used to build an extreme gradient boosting model based on the dataset, and to optimize the hyperparameters of the model to obtain a trained extreme gradient boosting model. Based on the trained extreme gradient boosting model, a correlation analysis was conducted on the water quality parameter set and the water quality index (WQI) value to screen out the key water quality parameters that affect the WQI value. The LSTM model building module is used to construct an LSTM model with the selected key water quality parameters as input variables and the water quality index (WQI) value as output variable, and to optimize the hyperparameters of the model to obtain the trained LSTM-WQI model. The SWAT model building module is used to acquire spatial and attribute data of rivers in the study area, construct a soil and water assessment tool model, simulate future hydrological and water quality processes based on the soil and water assessment tool model, and calculate and output future water quality parameter results. The prediction module is used to input the future water quality parameter results into the trained LSTM-WQI model to predict the future WQI value of the river, obtain the prediction results, and verify the prediction results based on existing data.
Citation Information
Cited By
SAR (Synthetic Aperture Radar) sea surface flow velocity inversion method based on physical guidance and data driving fusion
CN121145736A
Method for predicting total number of bacterial colonies in water body based on multi-source data spatial heterogeneous random forest
CN121459952A
Water quality simulation and prediction method
CN121835448A