River total phosphorus concentration prediction method based on XGBOOST model
Through the prediction method of river total phosphorus concentration based on the XGBOOST model, the existing monitoring methods are solved with high cost, low frequency or insufficient spatial coverage, and high-precision and low-cost normalization of total phosphorus concentration in river basins is achieved.
Patent Information
- Application Number
- CN202510279090.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing river total phosphorus concentration monitoring methods have high costs, low frequency or insufficient spatial coverage, making it difficult to achieve normalized low-cost predictions.
The total phosphorus concentration prediction method of rivers based on the XGBOOST model is adopted, and the data set is pre-processed by obtaining online water quality monitoring site data, and the optimal hyperparameter combination is selected using the Bayesian optimization framework to construct the XGBOOST model for prediction.
High-precision prediction of total phosphorus concentration in river basins is achieved, monitoring costs are reduced, frequency and spatial coverage are improved, and generalization and interpretability of the model are enhanced.
Smart Images

Figure CN120220856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of water quality monitoring, and particularly to a method for predicting the total phosphorus concentration in rivers based on the XGBOOST model. Background Art
[0002] River basins are an important part of the water cycle, which involves the collection, storage, transmission, and distribution of water resources. Total phosphorus (TP) is an important indicator to measure the phosphorus content in water bodies, and phosphorus is one of the key nutrients required for plant growth. However, when the total phosphorus concentration in water bodies exceeds a certain limit, it will cause a series of negative effects such as water quality deterioration, dissolved oxygen decline, biodiversity reduction, and habitat degradation. Therefore, predicting the total phosphorus concentration in river basins is crucial for ensuring the sustainable use of water resources, protecting the ecological environment, and promoting social and economic development.
[0003] In the basin, common total phosphorus monitoring methods include on-site sampling combined with laboratory analysis, automatic online monitoring stations, and remote sensing technology. Although these traditional methods can provide accurate measurement data, they are often limited by high costs, low frequencies, or insufficient spatial coverage. Summary of the Invention
[0004] In view of this, in order to solve the problem of low-cost and normalized prediction of the existing river total phosphorus concentration monitoring method, the present invention proposes a method for predicting the total phosphorus concentration in rivers based on the XGBOOST model, and the method includes the following steps:
[0005] Obtain the water quality data of the river basin provided by the online water quality monitoring station, including five water quality parameters (pH value, dissolved oxygen, turbidity, conductivity, temperature), ammonia nitrogen, permanganate index, chlorophyll a, and total phosphorus.
[0006] Preprocess the water quality data of the river basin, use interpolation methods (such as linear interpolation) or delete samples with too many missing values, and eliminate duplicate data and outliers.
[0007] Construct a data set according to the preprocessed water quality data;
[0008] Use the XGBoost regression model to predict the total phosphorus concentration and construct a prediction model;
[0009] Initialize the basic parameters of the prediction model, train the prediction model using the data set, and introduce SHAP for interpretability analysis;
[0010] During the training process, define the objective function such as RMSE, R 2Scoring and other indicators are used as optimization targets, and the Bayesian optimization framework is used for hyperparameter optimization. The best parameter combination is selected through cross-validation. The early stopping mechanism is used to dynamically adjust the number of iterations to avoid overfitting. The training process is visualized through SHAP analysis.
[0011] Among them, the hyperparameters selected by the XGBOOST model include: the number of decision trees (n_estimators), the maximum depth of a single tree (max_depth), the step size (learning_rate), the minimum loss reduction (gamma), the sample ratio for training each tree (subsample), and the feature ratio used to build each tree (colsample_bytree).
[0012] Input the real-time data to be tested into the trained prediction model and output the prediction results.
[0013] Furthermore, the step of constructing a data set based on the preprocessed water quality data specifically includes:
[0014] After logarithmic transformation, standardization, feature extraction, feature screening and feature combination, the input water quality data is divided into training set and validation set to obtain the data set.
[0015] Based on the above scheme, the present invention provides a method for predicting the total phosphorus concentration in rivers based on the XGBOOST model. After fully preprocessing the collected water quality data, the Bayesian optimization method is used to select the optimal hyperparameter combination, and the XGBOOST model with the best effect is constructed to perform high-precision prediction of the total phosphorus concentration in the river basin. At the same time, the early stopping mechanism is combined to reduce the training time; the use of a large amount of data with wide temporal and spatial distribution for model training can fully improve the generalization ability of the model to adapt to efficient application in different scenarios; through SHAP analysis, the contribution of each feature to the prediction result is quantified to guide the optimization layout of monitoring sites and the tracking of pollution sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flow chart of the steps of a method for predicting total phosphorus concentration in a river based on the XGBOOST model of the present invention;
[0017] Figure 2 A residual comparison diagram of a method for predicting total phosphorus concentration in a river based on the XGBOOST model in an embodiment of the invention
[0018] Figure 3 A training process loss curve diagram of a method for predicting total phosphorus concentration in a river based on an XGBOOST model in an embodiment of the invention
[0019] Figure 4Feature importance plot of a total phosphorus concentration prediction method for rivers based on the XGBOOST model in an invention embodiment
[0020] Figure 5 SHAP feature impact summary plot of a total phosphorus concentration prediction method for rivers based on the XGBOOST model in an invention embodiment. Detailed implementation manners
[0021] The present invention applies the XGBOOST algorithm. The XGBOOST algorithm has high requirements for data, and appropriate preprocessing of the data can improve the prediction accuracy of the model. At the same time, the XGBOOST algorithm involves multiple hyperparameters, and considering how to select the optimal combination of hyperparameters has an important impact on the prediction accuracy and efficiency of the total phosphorus concentration in the basin.
[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0023] It should be noted that for the convenience of description, only parts related to the relevant invention are shown in the accompanying drawings. Without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0024] It should be understood that the "system", "device", "unit" and / or "module" used in the present application is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the word can be replaced by other expressions.
[0025] As shown in the present application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements. The element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0026] In the description of the embodiments of the present application, "a plurality" means two or more than two. The following terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0027] In addition, flowcharts are used in the present application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the operations before or after do not necessarily need to be executed precisely in sequence. On the contrary, the steps can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0028] Refer to Figure 1 , which is a schematic flowchart of an optional example of the total phosphorus concentration prediction method for rivers based on the XGBOOST model proposed by the present invention. This method can be applied to computer devices. The prediction method proposed in this embodiment may include but is not limited to the following steps:
[0029] Step S1: Collect water quality data of different online water quality monitoring stations in the river basin;
[0030] Step S2: Preprocess the collected water quality data;
[0031] Step S3: Construct a data set according to the preprocessed data;
[0032] Step S4: Construct an XGBOOST model suitable for predicting the total phosphorus concentration of the basin;
[0033] Step S5: Use the data set obtained in Step S3 to train the XGBOOST model in Step S4, perform hyperparameter optimization using the Bayesian optimization framework, and introduce SHAP for interpretability analysis;
[0034] Step S6: Input the data to be measured into the trained prediction model obtained in Step S6, and output the prediction result.
[0035] In some feasible embodiments, Step S1 specifically includes:
[0036] The water quality data covers diverse data in different time periods and different weather seasons.
[0037] The categories of water quality data include five water quality parameters (pH value, dissolved oxygen, turbidity, conductivity, temperature), ammonia nitrogen, permanganate index, chlorophyll a, and total phosphorus.
[0038] In some feasible embodiments, Step S2 specifically includes:
[0039] Perform outlier removal and duplicate data removal on the obtained water quality data; outlier removal includes removing blank values and values exceeding the detection limit of the monitoring equipment; duplicate data removal is to remove duplicate values in each sampling and detection cycle of the monitoring equipment.
[0040] In some feasible embodiments, step S3 specifically includes:
[0041] S3.1. Perform logarithmic transformation on the input water quality data x = [x1, x2, …, x i , and the calculation formula for the resulting array y = [y1, y2, …, y i is as follows:
[0042] y i = ln(x i + ∈)
[0043] where ln represents the natural logarithm function (with base e), ∈ is a very small positive number, usually used to avoid taking the logarithm of zero or negative numbers, because in mathematics, ln(0) and ln(x) (when x < 0) are undefined. To ensure that all elements of the input array are greater than zero, an appropriate positive number is added as an offset, and take ∈ = 1e -6 ;
[0044] S3.2. Standardize the input water quality data, and the calculation formula for standardization is as follows:
[0045]
[0046] where X is the original eigenvalue data, μ is the average of all sample sets of this feature, and σ is the standard deviation of all sample values of this feature;
[0047] S3.3. Divide it into a training set and a test set, with a ratio of 80%:20%.
[0048] In some feasible embodiments, step S4 specifically includes:
[0049] Construct an XGBOOST model, which includes: loss function definition, regularization mechanism, gradient-based split point selection, and an efficient parallel computing framework. In each iteration, XGBoost calculates the gradient and Hessian matrix of the prediction error of the current model to guide the growth of the new tree, ensuring that each tree can minimize the overall loss function, and at the same time controlling the model complexity through the regularization term to prevent overfitting.
[0050] The hyperparameters selected by the XGBoost model include: the number of decision trees (n_estimators), the maximum depth of a single tree (max_depth), the step size (learning_rate), the minimum loss reduction (gamma), the proportion of samples for training each tree (subsample), and the proportion of features used for constructing each tree (colsample_bytree).
[0051] In some feasible embodiments, in step S5:
[0052] Use the Bayesian optimization method to optimize the hyperparameters, evaluate and select the best hyperparameter combination through 10-fold cross-validation to maximize the R 2 score on the training set, and use the early stopping mechanism (early_stopping_rounds = 10) to dynamically adjust the number of iterations to avoid overfitting.
[0053] Conduct visual analysis on the model training process and prediction results, including: training process loss curve, residual comparison graph, feature importance graph, SHAP analysis.
[0054] In this embodiment, the residual distribution comparison graph ( Figure 2 ) shows the distribution characteristics of the residuals of the training set and the test set within the predicted value range. It can be seen from the figure that the distribution trends of the training set residuals (labeled as □5 to □1) and the test set residuals (in the same label range) are highly consistent. The absolute values of the residuals are generally controlled within the interval [-1.5, 1.0], and most of them are concentrated in the narrow range of [-0.5, 0.5], indicating that the model shows good stability and generalization ability in both the training and test stages. In particular, there is no significant deviation or discrete expansion in the residual distribution of each predicted value interval (□1 to □5), verifying the effectiveness of the model parameter optimization. This consistency further shows that through a unique algorithm design, this technology effectively suppresses the risk of overfitting while maintaining prediction accuracy in complex scenarios, providing a reliable theoretical support for practical applications.
[0055] The loss change curve of the model training involved ( Figure 3) clearly shows the optimization trajectory of the loss value with the number of iterations during the training process. As shown in the figure, in the initial stage (number of iterations < 50), the loss value rapidly drops from 0.6 to below 0.3, reflecting the algorithm's efficient parameter adjustment ability; as the number of iterations increases (50 - 200 times), the loss value continuously converges in a smooth trend and finally stabilizes in a low value range of around 0.1 (number of iterations > 200), indicating that the model has achieved sufficient convergence within a limited training cycle. There is no oscillation or rebound phenomenon in the whole curve, verifying the rationality of the learning rate strategy and the robustness of the training data. Compared with traditional methods, this technology significantly improves the convergence speed and stability through an innovative optimization mechanism, providing key technical support for the rapid deployment and high-performance output of complex models.
[0056] In this embodiment, the feature importance ranking graph ( Figure 4 ) reveals the differential contributions of various water quality parameters to the model prediction performance. As shown in the figure, the feature importance scores show a significant hierarchical distribution. Among them, conductivity ranks first with an importance score of 0.35, indicating its decisive influence on the prediction results; log_COD-Mn (logarithm of chemical oxygen demand) and pH follow with importance scores of 0.25 and 0.20 respectively, verifying the core role of pollutant concentration and acidity / alkalinity parameters. It is worth noting that by introducing logarithmic transformation (such as log_COD-Mn, log_turbidity, log_chlorophyll-a), the linear correlation of the original data is effectively improved, enabling the model to capture complex non-linear relationships more accurately. In addition, although features such as NH3-N (ammonia nitrogen) and temperature have relatively low contribution degrees, their synergistic effects still provide multi-dimensional information support for the model. This ranking result highlights the innovation of this technology in feature engineering and parameter optimization. By focusing on high-impact features and reasonably reducing dimensions, the interpretability and computational efficiency of the model are significantly improved, providing a scientific basis for the engineering application of water quality prediction.
[0057] In this embodiment, the SHAP feature impact summary graph ( Figure 5)The system quantifies the local and global contribution degrees of each water quality parameter to the model prediction results. As shown in the figure, the SHAP value distribution range of conductivity is the widest (from 0.2 to 0.8), indicating that its positive driving effect on the model output is significantly dominant. Especially in the extreme high value range (>0.6), its contribution is prominent, verifying the sensitivity of conductivity as a core water quality indicator. The SHAP values of log_COD-Mn (logarithm of chemical oxygen demand) show a two-way distribution (-0.2 to 0.6), indicating that the change in its concentration can not only significantly increase the predicted value (positive value area), but also inhibit the model output under specific conditions (negative value area), revealing the complex non-linear characteristics of the dynamic response of pollutants. The SHAP values of pH are concentrated in the range of 0.1 to 0.4, reflecting a robust positive correlation between acidity and alkalinity and the prediction results. In addition, the SHAP value ranges of NH3-N (ammonia nitrogen) and temperature are relatively narrow (-0.1 to 0.3), but their interaction effects with high-weight features (such as conductivity) can indirectly optimize the model decision boundary. It is worth noting that by introducing logarithmic transformations (such as log_turbidity, log_chlorophyll-a), not only the distribution balance of the original data is improved, but also the interpretability of the SHAP values is significantly enhanced, enabling the model to more accurately quantify the causal mechanism between features and prediction targets. This analysis result verifies the innovative design of this patented technology in feature selection, data preprocessing, and model transparency from the perspective of explainable artificial intelligence (XAI), providing a quantitative basis for the reliability assessment and engineering optimization of the water quality prediction system.
[0058] A river total phosphorus concentration prediction system based on the XGBOOST model, comprising:
[0059] A dataset construction module for collecting water quality data of different online water quality monitoring stations in the river basin; preprocessing the collected water quality data; and constructing a dataset based on the preprocessed data;
[0060] A model construction module for constructing an XGBOOST model applicable to predicting the total phosphorus concentration of the basin;
[0061] A model training module for training the XGBOOST model in step S4 using the dataset obtained in step S3, performing hyperparameter optimization using the Bayesian optimization framework, and introducing SHAP for interpretability analysis;
[0062] A model application module for inputting the data to be measured into the trained prediction model obtained in step S6 and outputting the prediction result.
[0063] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented in the system embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0064] A device for predicting the total phosphorus concentration in a river based on the XGBoost model:
[0065] At least one processor;
[0066] At least one memory for storing at least one program;
[0067] When the at least one program is executed by the at least one processor, the at least one processor implements the method for predicting the total phosphorus concentration in a river based on the XGBoost model as described above.
[0068] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented in the device embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0069] A storage medium storing instructions executable by a processor, where the instructions executable by the processor are used to implement the method for predicting the total phosphorus concentration in a river based on the XGBoost model as described above when executed by the processor.
[0070] The content in the above method embodiments is applicable to the storage medium embodiments of the present invention. The functions specifically implemented in the storage medium embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0071] The above is a specific description of the preferred embodiments of the present invention. However, the present invention is not limited to the described embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for predicting total phosphorus concentration in a river based on the XGBOOST model, characterized in that: The following steps are involved: Acquire water quality data and construct datasets; Build a prediction model based on the XGBOOST algorithm; The prediction model is trained based on the data set, and SHAP is introduced to perform interpretability analysis to obtain a trained prediction model; The data to be tested is input into the trained prediction model, and the prediction result is output.
2. A method for predicting total phosphorus concentration in a river based on the XGBOOST model according to claim 1, characterized in that: The categories of water quality data include five water quality parameters, ammonia nitrogen, permanganate index, chlorophyll a and total phosphorus.
3. A method for predicting total phosphorus concentration in a river based on the XGBOOST model according to claim 1, characterized in that: The step of obtaining water quality data and constructing a data set specifically includes: Collect water quality data from different water quality monitoring stations in the river basin; Eliminating duplicate data and outliers from the water quality data to obtain preprocessed data; A data set is constructed according to the preprocessed data.
4. A method for predicting total phosphorus concentration in a river based on the XGBOOST model according to claim 3, characterized in that: The step of constructing a data set based on the preprocessed data specifically includes: Performing logarithmic transformation on the preprocessed data to obtain transformed data; Performing standardization processing on the transformed data to obtain standardized data; The standardized data is divided into a training set and a test set.
5. A method for predicting total phosphorus concentration in a river based on the XGBOOST model according to claim 4, characterized in that: The step of constructing a data set based on the preprocessed data further includes: Performing feature extraction on the preprocessed data to obtain extracted features; Screening the extracted features based on a principal component analysis method; The extracted features are combined based on a correlation analysis method to construct interactive features.
6. A method for predicting total phosphorus concentration in a river based on the XGBOOST model according to claim 1, characterized in that: The step of training the prediction model based on the data set and introducing SHAP to perform interpretability analysis to obtain the trained prediction model specifically includes: The prediction model is trained based on the data set, the objective function is defined, the hyperparameters are optimized using the Bayesian optimization method, and the best hyperparameter combination is evaluated and selected through 10-fold cross validation to maximize the R of the model on the training set. 2 score; During the training process, the number of iterations is dynamically adjusted through the early stopping mechanism, and SHAP analysis is introduced to obtain the trained prediction model.
7. A method for predicting total phosphorus concentration in a river based on the XGBOOST model according to claim 6, characterized in that: The hyperparameters include the number of decision trees, the maximum depth of a single tree, the step size, the minimum loss reduction, the proportion of samples used to train each tree, and the proportion of features used to build each tree.
8. A river total phosphorus concentration prediction system based on the XGBOOST model, characterized in that: include: Dataset construction module, used to obtain water quality data and construct datasets; Model building module, building prediction models based on XGBOOST algorithm; A model training module trains the prediction model based on the data set and introduces SHAP to perform interpretability analysis to obtain a trained prediction model; The model application module is used to input the data to be tested into the trained prediction model and output the prediction result.
9. A device for predicting total phosphorus concentration in a river based on the XGBOOST model, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a method for predicting total phosphorus concentration in a river based on the XGBOOST model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Water environment treatment prediction system and method based on spatio-temporal data cleaning technology
CN115829143A
Water body total phosphorus integrated intelligent inversion method based on feature map layer screening
CN116026796A
Total nitrogen and total phosphorus inversion method based on sample migration and lasso regression
CN116092597A
Method for improving gastric cancer typing prognosis prediction precision based on gradient lifting depth feature selection algorithm
CN116417070A
Method and system for forecasting particulate matters and ozone based on automatic feature model
CN118395147A
Cited By
Freezing circle drainage basin total phosphorus concentration simulation method and system
CN120724872A
Method for constructing high-resolution atmospheric carbon dioxide concentration data set based on XGBoost-BO
CN120873607A