Pollutant control method and device in water quality monitoring process

By combining the basic prediction model and the residual model with a multi-objective optimization algorithm, the problem of pollutant prediction bias in water quality monitoring was solved, achieving precise pollutant concentration control and improving water quality monitoring efficiency, and providing a scientific pollutant optimization control scheme.

CN121545618APending Publication Date: 2026-02-17BEIJING WATER SCI & TECH INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511693879.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing technologies neglect the interaction between pollutants in water quality monitoring, resulting in one-sided predictions of multiple pollutants, significant deviations between the predictions and actual conditions, and poor pollution control effectiveness.

Method used

By acquiring historical multi-source characteristic data of the target water area, one-to-one pollutant prediction and residual correction are performed using basic prediction models and residual models. Combined with multi-objective optimization algorithms, a multi-objective optimization problem model is constructed to determine the optimal combination of environmental control variables and generate a pollutant control report.

Benefits of technology

It enables more accurate prediction and control of pollutant concentrations, improves the efficiency and effectiveness of water quality monitoring, provides a scientific basis for optimized control of pollutants, and supports environmental protection and resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545618A_ABST
    Figure CN121545618A_ABST
Patent Text Reader

Abstract

The invention relates to the field of water quality monitoring, and discloses a pollutant control method and device in a water quality monitoring process, and the method comprises the steps: obtaining historical multi-source feature data of a target water area; respectively inputting basic prediction models corresponding to different pollutants one by one, and predicting pollutant prediction concentrations corresponding to different pollutants one by one; inputting the pollutant predicted concentrations corresponding to the different pollutants one by one and the historical multi-source feature data into a pre-established residual model to obtain residual predicted values corresponding to the different pollutants one by one; correcting the corresponding pollutant predicted concentrations based on the residual predicted values corresponding to the different pollutants one by one to obtain pollutant target concentrations corresponding to the different pollutants one by one; and taking the pollutant target concentrations corresponding to different pollutants one by one as optimization targets, and determining a pollutant optimization target. According to the method, the prediction performance is improved by fusing the basic prediction model and the residual error model, and the accuracy of pollutant prediction and control is remarkably improved in combination with multi-objective optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water environment, specifically to a method and apparatus for controlling pollutants during water quality monitoring. Background Technology

[0002] With economic development, the water quality of reservoirs and watersheds has received widespread attention. Water quality is affected by a variety of factors, including natural and anthropogenic factors. Traditional single-objective modeling methods ignore the potential interactions between pollutants and cannot effectively reflect the interrelationships between multiple pollutants, leading to biased prediction results. This approach has limited effectiveness when dealing with multiple pollutants and complex water quality. Furthermore, because various prediction models are trained independently, they often fail to effectively predict the complex behavior of multiple pollutants interacting in wastewater, resulting in significant discrepancies between predictions and actual conditions, and consequently, poor pollution control effectiveness. Summary of the Invention

[0003] In view of this, the present invention provides a method and apparatus for controlling pollutants during water quality monitoring, in order to solve the problem of poor pollutant prediction and control in the prior art during wastewater treatment.

[0004] In a first aspect, the present invention provides a method for controlling pollutants during water quality monitoring, the method comprising: Acquire historical multi-source feature data of the target water area; Historical multi-source feature data are input into the basic prediction model corresponding to different pollutants to predict the predicted concentration of pollutants corresponding to different pollutants. By inputting the predicted concentrations of different pollutants and historical multi-source characteristic data into the pre-established residual model, the predicted residual values ​​corresponding to different pollutants are obtained. Based on the residual prediction values ​​corresponding to different pollutants, the predicted concentrations of the corresponding pollutants are corrected to obtain the target concentrations of the pollutants corresponding to different pollutants. The target concentrations of different pollutants are used as optimization targets to determine the pollutant optimization targets; these targets are then used to control the pollutants.

[0005] This invention effectively improves the accuracy of pollutant concentration prediction by combining historical multi-source feature data, a basic prediction model, and a residual correction mechanism. By performing one-to-one prediction and residual correction for different pollutants, more accurate prediction of target pollutant concentrations can be achieved, providing a scientific basis for subsequent optimized pollutant control. This method can precisely regulate pollutant concentrations, reduce environmental pollution, and improve the efficiency and effectiveness of water quality monitoring.

[0006] In one optional implementation, the target concentrations of different pollutants are used as optimization targets to determine the pollutant optimization targets, including: Obtain the environmental constraints corresponding to each pollutant; Determine the set of environmental control variables for the target water area; A multi-objective optimization problem model is established based on the optimization objective, environmental constraints, and set of environmental control variables. Based on a multi-objective optimization problem model, the Pareto front solution set is determined; the Pareto front solution set includes the optimal combination of environmental control variables for controlling pollutants; the optimal combination of environmental control variables is used as the objective for determining pollutant optimization.

[0007] This implementation method, while meeting environmental constraints, can accurately balance the control targets of various pollutants, achieve accurate pollutant control, and keep pollutant concentrations within a reasonable range.

[0008] In one alternative implementation, after determining the Pareto front solution set, the process includes: Based on the target concentrations of different pollutants and the Pareto frontier solution set, a pollutant control report is generated and visualized.

[0009] This implementation method, through the output of pollutant control reports, makes the entire water quality regulation process more refined and customized, providing stronger support for environmental protection and resource management.

[0010] In one optional implementation, before acquiring historical multi-source feature data of the target water area, the following steps are included: Obtain raw historical multi-source data for the target water area; raw historical multi-source data includes water quality monitoring data, hydrological and meteorological data, and land use information; Data preprocessing is performed on the original historical multi-source data. Data preprocessing includes: data cleaning, missing value handling, and outlier handling. Standardization and / or category feature processing are performed on the raw historical multi-source data after data preprocessing; Historical multi-source feature data is constructed based on the original historical multi-source data after standardization and / or categorical feature processing. The historical multi-source feature data includes features that are associated with pollutants.

[0011] In this embodiment, to ensure the accuracy and stability of the model prediction, the original multi-source data needs to be structured to construct a unified and standardized model input feature set, eliminate data noise and redundant information, and provide a reliable foundation for subsequent predictions.

[0012] In one alternative implementation, the base prediction model is pre-trained using historical multi-source data of the target water area; during training, the base prediction model is hyperparameter-tuned using time-series cross-validation.

[0013] In this embodiment, testing is performed in chronological order, which can better simulate the model prediction process in real-world application scenarios.

[0014] In one alternative implementation, the residual model is established through the following steps: Obtain the prediction residual of the basic prediction model on the training set. The prediction residual is the difference between the predicted value and the actual value of the basic prediction model. The initial regression model is trained based on the predicted residuals and historical multi-source data to obtain the residual model.

[0015] In this embodiment, by modeling the prediction error of the basic prediction model, missed nonlinear information and complex fluctuation trends can be captured, thereby correcting the residuals of the basic prediction model and optimizing the prediction effect, further improving the prediction accuracy.

[0016] In a second aspect, the present invention provides a pollutant control device for water quality monitoring, the device comprising: The acquisition module is used to acquire historical multi-source feature data of the target water area; The basic prediction model prediction module is used to input historical multi-source feature data into the basic prediction model corresponding to different pollutants, and predict the pollutant concentration corresponding to different pollutants. The residual model prediction module is used to input the predicted concentrations of pollutants corresponding to different pollutants and historical multi-source characteristic data into the pre-established residual model to obtain the residual prediction values ​​corresponding to different pollutants. The fusion module is used to correct the predicted concentration of the corresponding pollutants based on the residual prediction values ​​corresponding to different pollutants, so as to obtain the target concentration of the pollutants corresponding to different pollutants. The target optimization module is used to determine the pollutant optimization target by taking the target concentration of different pollutants as the optimization target; the pollutant optimization target is used to control the pollutants.

[0017] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the pollutant control method in the water quality monitoring process described in the first aspect or any corresponding embodiment.

[0018] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the pollutant control method in the water quality monitoring process described in the first aspect or any corresponding embodiment.

[0019] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the pollutant control method in the water quality monitoring process described in the first aspect or any corresponding embodiment.

[0020] It should be noted that the pollutant control device, computer equipment, computer-readable storage medium, and computer program product provided in this invention correspond to the pollutant control method in the water quality monitoring process described above. Therefore, for the beneficial effects of the pollutant control device, computer equipment, computer-readable storage medium, and computer program product in the water quality monitoring process, please refer to the description of the corresponding beneficial effects of the pollutant control method in the water quality monitoring process above, and will not be repeated here. Attached Figure Description

[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0022] Figure 1 This is a schematic flowchart of a pollutant control method during water quality monitoring according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating multi-objective optimization according to an embodiment of the present invention; Figure 3 This is a structural block diagram of a pollutant control device in the water quality monitoring process according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Water quality in different river basins is often affected by a variety of factors, including natural factors (such as rainfall, temperature, evaporation, dissolved oxygen, etc.) and anthropogenic factors (such as domestic emissions, industrial emissions, and agricultural non-point source pollution).

[0025] Current pollutant prediction models are mostly based on the independence assumption, neglecting the potential interactions between pollutants. Furthermore, existing water quality data is often from complex sources and of varying quality, posing a challenge to building a stable and highly accurate unified model. More importantly, there is a lack of intelligent integrated methods that deeply integrate pollutant prediction with pollution control optimization, resulting in a lack of effective connection between prediction results and actual control strategies. Moreover, the setting of model input parameters largely relies on human experience, lacking a systematic and automated parameter optimization mechanism, which reduces the generalization ability and practicality of the methods. Therefore, there is an urgent need to propose an intelligent prediction method that integrates multiple prediction models and simultaneously supports multi-pollutant target constraint optimization, comprehensively improving the accuracy, stability, and intelligence level of water quality control.

[0026] In view of this, according to an embodiment of the present invention, a method for controlling pollutants in a water quality monitoring process is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0027] This embodiment provides a method for controlling pollutants during water quality monitoring, which can be executed by devices such as servers, terminals, and mobile terminals. Figure 1 This is a flowchart of a pollutant control method during water quality monitoring according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain historical multi-source feature data of the target water area.

[0028] The target water body can be any water body, and the historical multi-source feature data refers to the feature data obtained after processing historical multi-source data. Historical multi-source data can include water quality monitoring data, such as pollutant content in the water; hydrological and meteorological data, such as precipitation, temperature, runoff, and evaporation; and land use information, such as land use types in different regions, including cultivated land, construction land, and forest land. This data should have a sufficient time span to cover factors such as seasonal changes and unforeseen events.

[0029] Step S102: Input the historical multi-source feature data into the basic prediction model corresponding to each pollutant to predict the predicted concentration of each pollutant.

[0030] In this embodiment, establishing independent basic prediction models for each pollutant indicator, such as total nitrogen, total phosphorus, ammonia nitrogen, and chemical oxygen demand, is a key step in improving model accuracy. Gradient Boosting Decision Tree (GBDT) algorithms, such as XGBoost, can be used to build these basic prediction models because they can effectively capture nonlinear relationships and handle complex feature interactions by integrating multiple weak prediction models (decision trees). The basic prediction model for each pollutant is trained based on its feature set, which includes historical water quality monitoring data, hydrological and meteorological data, land use information, etc. By training an independent model for each pollutant, the mapping relationship between pollutant concentration and various features can be accurately captured. In this embodiment, XGBoost is preferably used for model building. XGBoost has high computational efficiency and accuracy when processing large-scale data and can effectively avoid overfitting. Therefore, when building the prediction model for each pollutant, multiple rounds of training and validation can be used to continuously adjust the feature set and model parameters to improve the predictive ability of each model.

[0031] Of course, other models can also be chosen, such as random forest models, deep learning models, etc.

[0032] Step S103: Input the predicted concentrations of pollutants corresponding to different pollutants and historical multi-source characteristic data into the pre-established residual model to obtain the predicted residual values ​​corresponding to different pollutants.

[0033] In some alternative implementations, the residual model is established through the following steps: Obtain the prediction residual of the basic prediction model on the training set. The prediction residual is the difference between the predicted value and the actual value of the basic prediction model. The initial regression model is trained based on the predicted residuals and historical multi-source data to obtain the residual model.

[0034] Specifically, recording the prediction residuals of the base prediction model on the training set is a crucial first step in improving model accuracy when constructing a residual model. Prediction residuals are the differences between actual values ​​and model predictions, reflecting information the model failed to capture or existing biases. By calculating the residuals for each sample, systematic errors in the model at specific data points can be identified. For example, if the base prediction model's predictions deviate significantly under certain specific water quality monitoring data, meteorological conditions, or geographical locations, the residual predictions will highlight the prediction deficiencies in these areas. In-depth analysis of the residuals reveals the underlying patterns and nonlinear relationships that the base prediction model failed to accurately capture.

[0035] After recording the prediction residuals of the baseline prediction model, the next step is to construct a residual model. The goal of this residual model is to correct the prediction errors of the baseline prediction model by capturing nonlinear relationships and complex fluctuation trends that it fails to capture. In the residual model, input variables include the predicted values ​​from the baseline prediction model and the original input features (such as water quality monitoring data, hydrological and meteorological data, land use information, etc.). These input features provide richer information for residual modeling, enabling the model to learn the potential patterns of errors in the baseline prediction model. The residual prediction values, as the target variable, accurately describe the sources of error by training the residual model, thereby revealing complex factors that the baseline prediction model fails to handle. In this way, the residual model can not only correct systematic biases but also capture details and trends overlooked by the baseline prediction model, thus optimizing the final prediction results.

[0036] To further improve the accuracy of the prediction results, this embodiment uses the lightweight regression model XGBoost to model the residuals. XGBoost is a powerful gradient boosting tree model that can effectively capture the nonlinear relationship between features and the target variable. By using XGBoost to model the residuals, the biases of the base prediction model can be accurately corrected. The training data of the residual model consists of the predicted values ​​of the base prediction model and the original features. XGBoost, through its excellent feature selection ability and efficient fitting ability, can learn the pattern of the residuals and predict the corrected values ​​of the residuals. Adding this corrected value back to the predicted values ​​of the base prediction model yields a more accurate prediction result. This method can significantly improve the prediction accuracy of the base prediction model, correct systematic biases, ensure that the final prediction results are more consistent with the actual observation data, and improve the overall performance and reliability of the model.

[0037] By modeling the prediction error of the basic prediction model, we can capture the missed nonlinear information and complex fluctuation trends, thereby correcting the residuals of the basic prediction model and optimizing the prediction effect, and further improving the prediction accuracy.

[0038] Step S104: Based on the residual prediction values ​​corresponding to different pollutants, the predicted concentrations of the corresponding pollutants are corrected to obtain the target concentrations of the pollutants corresponding to different pollutants.

[0039] In this embodiment, the basic prediction model and the residual model together constitute a fusion prediction model (FusionModel). That is, the outputs of the basic prediction model and the residual model are combined to achieve more accurate pollutant concentration predictions. Specifically, the FusionModel first uses the basic prediction model to predict the input parameters, obtaining the predicted values ​​of the basic prediction model. Then, the residual model corrects these prediction results, predicting errors (i.e., residual prediction values) based on the predicted values ​​of the basic prediction model and the original features. Finally, the output of the FusionModel is the sum of the predicted values ​​of the basic prediction model and the residual model, i.e., final_prediction = main_model.predict(X) + resid_model.predict(X_with_main_pred). This fusion method effectively utilizes the advantages of both the basic prediction model and the residual model, eliminates the bias that may exist in a single model, and enhances the stability and accuracy of the prediction results. Through this integration method, the contributions of different models can be comprehensively considered, improving overall prediction performance and ensuring that the prediction results are more consistent with reality.

[0040] Using a fusion prediction model as the final pollutant prediction tool can significantly improve the accuracy and stability of predictions. While a single basic prediction model can provide a general trend, it often fails to capture all complex patterns and nonlinear relationships. By fusing the outputs of the residual model with those of the basic prediction model, the fusion prediction model can correct the errors of the basic prediction model and compensate for its shortcomings. This ensemble approach not only improves the model's fitting accuracy and reduces prediction bias but also enhances its stability. Because it combines the prediction results of different models, it can effectively address the overfitting or underfitting problems that may occur with single models. The advantages of fusion prediction models also lie in their scalability and generalization ability. With the introduction of more data or new features, they can be flexibly adjusted and expanded while maintaining high predictive performance. Therefore, fusion prediction models provide a more accurate, stable, and efficient solution for final pollutant prediction, adapting to the needs of different scenarios and effectively improving the accuracy and reliability of decision support systems.

[0041] This embodiment integrates the prediction results of the basic prediction model and the residual model to achieve a more accurate prediction output of pollutant concentration. The fusion method has scalability and generalization ability, effectively solving the problem of inconsistency among multiple models.

[0042] Step S105: The target concentrations of different pollutants are used as optimization targets to determine the pollutant optimization targets; the pollutant optimization targets are used to control pollutants.

[0043] In this step, multiple pollutant prediction results are used as optimization objectives, and pollutant environmental standards are introduced as constraint conditions to construct a multi-objective optimization problem model. This can comprehensively consider the impacts of different pollution sources on the water environment and balance the optimization objectives among different pollutants according to actual needs.

[0044] The present invention proposes a water quality prediction and regulation method combining fusion modeling and multi-objective optimization, which can effectively solve problems such as low pollutant prediction accuracy, insufficient model generalization ability, and lack of systematic optimization strategies for pollution control in related technologies. By integrating the basic prediction model and the residual model, the prediction performance is improved, and the multi-objective optimization algorithm is combined to achieve synchronous regulation of multiple water quality factors, thereby significantly improving the accuracy of pollutant prediction and control and enhancing the intelligent and scientific level of water quality management.

[0045] By combining historical multi-source feature data, the basic prediction model, and the residual correction mechanism, the accuracy of pollutant concentration prediction is effectively improved. Through one-to-one prediction and residual correction of different pollutants, more accurate prediction of the target concentration of pollutants can be achieved, providing a scientific basis for subsequent optimization control of pollutants. This method can precisely regulate pollutant concentration, reduce environmental pollution, and improve the efficiency of water quality monitoring.

[0046] In some optional embodiments, in the above step S105, that is, taking the target concentration of pollutants corresponding to different pollutants one by one as the optimization objective to determine the pollutant optimization objective, including: Step S1051, obtaining the environmental constraint conditions corresponding to each pollutant one by one.

[0047] To ensure that the pollutant concentration is within a reasonable range, it is necessary to set the environmental constraint conditions for each pollutant in advance. The environmental standard constraints can be based on local or international water quality standards and ecological requirements to set the maximum allowable concentration of pollutants. For example, it can be set that the concentration of ammonia nitrogen (NH4 + -N) should satisfy 0 < NH4 + -N ≤ 0.5 mg / L, the total phosphorus (TP) should satisfy 0 < TP ≤ 0.1 mg / L, the total nitrogen (TN) should satisfy 0 < TN ≤ 1.0 mg / L, and the chemical oxygen demand (COD) should satisfy 0 < COD ≤ 15 mg / L. These constraint conditions ensure that the pollutant concentration does not exceed the safety limit and avoid causing serious impacts on the ecological system. In the optimization process, these environmental standards not only act as constraint conditions to limit the optimization space but also provide the implementation boundary for the final pollution control plan. By introducing environmental constraints, the optimization problem becomes more realistic and operable, ensuring that the optimization solution can meet the water quality requirements.

[0048] Step S1052, determining the set of environmental control variables for the target water area.

[0049] In constructing multi-objective optimization problems, environmental control variables are key factors determining changes in pollutant concentrations. These environmental control variables include land use structure parameters (such as the proportion of arable land, forest land, and grassland), population density, hydrological and meteorological parameters (such as precipitation, temperature, and evaporation), and other factors that may affect pollutant concentrations. These control variables directly or indirectly influence the generation, distribution, and migration of pollutants. For example, agricultural activities may lead to more nitrogen and phosphorus flowing into water bodies, and urbanization may increase pollutant emissions. Therefore, rationally setting and optimizing these control variables will help reduce pollutant concentrations and achieve environmental goals. The optimization process will explore the optimal pollution control scheme by adjusting the combination of these input variables, thereby achieving globally optimal pollutant concentration control while meeting environmental standards.

[0050] Step S1053: Based on the optimization objective, environmental constraints, and set of environmental control variables, establish a multi-objective optimization problem model.

[0051] In constructing a multi-objective optimization problem, the optimization objective must first be clearly defined, namely, the predicted values ​​of pollutant concentrations (such as total nitrogen (TN), total phosphorus (TP), and ammonia nitrogen (NH4). + Pollutants such as nitrogen oxides (N₂) and chemical oxygen demand (COD) are key indicators affecting environmental quality. Therefore, the optimization objective is to minimize their concentrations. To achieve this, these pollutant concentrations are used as objective functions in a multi-objective optimization problem, where reducing the concentration of each pollutant is considered the objective to be maximized. By treating different pollutant concentrations as multiple objectives, the impact of different pollution sources on the aquatic environment can be comprehensively considered, and the optimization objectives among different pollutants can be weighed according to actual needs. This approach provides a theoretical basis and mathematical support for achieving the global optimum of pollution control schemes.

[0052] Step S1054: Based on the multi-objective optimization problem model, determine the Pareto front solution set; the Pareto front solution set includes the optimal combination of environmental control variables for controlling pollutants; use the optimal combination of environmental control variables as the objective for determining pollutant optimization.

[0053] In this embodiment, multi-objective evolutionary algorithms (such as NSGA-II or MOEA / D) can be used to handle multi-objective optimization problems. These algorithms can simultaneously optimize multiple objective functions and find the optimal solution set in a multi-dimensional space. In this embodiment, a multi-objective evolutionary algorithm will be used to iteratively search the space of control variables, such as land use structure, population density, and hydro-meteorological parameters, to find the optimal solution set that can simultaneously satisfy multiple pollutant concentrations. Through iterative evolution, these algorithms can evaluate the quality of the current solution in each generation and generate new solutions through operations such as selection, crossover, and mutation, thereby continuously optimizing pollutant concentrations. In the final solution set, a set of optimal solutions that balance multiple objectives can be obtained. These solutions provide different combinations of control variables, enabling pollutant concentrations to meet environmental standards while achieving the best pollution control effect.

[0054] The Pareto front solution set, derived from a multi-objective optimization algorithm, represents a set of solutions that cannot be further optimized, signifying the optimal trade-off among multiple objectives. In this step, optimal control configurations are extracted from the Pareto front solution set of the optimization results, serving as the basis for decision-making. These optimal solutions correspond to the best control variable settings (such as land use structure combinations, control parameter settings, etc.) under different combinations of pollutant concentrations. For example, in some scenarios, reducing total nitrogen (TN) concentration may require increasing arable land area, while in others, it may be achieved by expanding forest land or increasing wetland cover to control pollutant concentrations. These control configurations represent different land management and pollutant control strategies that not only minimize pollutant concentrations while meeting environmental standards but also allow for flexible adjustments based on different local and environmental needs. Therefore, extracting these Pareto optimal control configurations provides diverse and feasible decision-making schemes for practical applications, helping decision-makers select the most appropriate pollution control strategies in different scenarios.

[0055] This embodiment, while meeting environmental constraints, can accurately balance the control targets of various pollutants, achieve accurate pollutant control, and keep the pollutant concentration within a reasonable range.

[0056] In some alternative implementations, after determining the Pareto front solution set, the following steps are included: Based on the target concentrations of different pollutants and the Pareto frontier solution set, a pollutant control report is generated and visualized.

[0057] After obtaining the Pareto optimal solution, the next step is to configure the predicted pollutant values ​​with the corresponding input variables, generate a pollutant control report, and provide a visual output, such as a tabular file or graphical interface, for water quality managers to analyze and select. This serves as the basis for formulating water quality governance strategies and water quality compliance control plans. Each optimal solution provides a corresponding predicted pollutant concentration and clarifies how to achieve this goal by adjusting input variables (such as land use structure, population density, hydrological and meteorological conditions). By outputting this information, decision-makers can formulate differentiated governance strategies based on the specific circumstances of different watersheds. For example, in some watersheds, it may be necessary to increase vegetation cover and wetland restoration to control nitrogen and phosphorus pollution, while in other watersheds, it may be more necessary to rely on adjustments to agricultural activities or optimization of wastewater treatment facilities. This output based on the optimal solution set provides clear guidance for water quality compliance control plans, ensuring that pollution control measures are not only scientific but also practically operable, helping decision-makers ensure that water quality meets standards. Through this method of outputting pollutant control reports, the entire water quality control process becomes more refined and customized, providing stronger support for environmental protection and resource management.

[0058] In some optional implementations, prior to acquiring historical multi-source feature data of the target water area, the following steps are included: Step a1: Obtain the original historical multi-source data of the target water area; the original historical multi-source data includes water quality monitoring data, hydrological and meteorological data, land use information, etc.

[0059] It can acquire historical water quality monitoring data, such as pollutant content in the water; hydrological and meteorological data, such as precipitation, temperature, runoff, and evaporation; and land use information, such as land use types in different regions, including cultivated land, forest land, and building area. This data should have a sufficient time span to cover factors such as seasonal changes and unforeseen events.

[0060] Step a2 involves preprocessing the original historical multi-source data, including data cleaning, missing value handling, and outlier handling.

[0061] After obtaining the raw historical multi-source data, data cleaning, including removing missing and outlier values, is a crucial step. Missing values ​​can be handled by imputation using the mean, interpolation, or deletion of records with a large number of missing values; while outliers can be detected and removed using methods such as box plots and Z-scores. Finally, ensuring that all data is formatted consistently, such as timestamp format and unit standardization, guarantees data quality and consistency, providing a reliable foundation for subsequent modeling.

[0062] Step a3 involves standardizing and / or performing category feature processing on the preprocessed original historical multi-source data.

[0063] Specifically, numerical features, such as precipitation, evaporation, and water temperature, are standardized; categorical variables, such as land use types, are either one-hot encoded or converted to categorical representations. For numerical features, standardization is a crucial step in ensuring model stability and accuracy. Standardization methods, such as Z-score standardization or Min-Max standardization, can transform features with different dimensions to the same scale, avoiding unnecessary influence from features with large values ​​on the model. For example, for features like precipitation, evaporation, and water temperature, standardization transforms them into zero mean and unit standard deviation, or maps them to the [0,1] interval, making the model easier to train and converge. For categorical variables, such as land use types and climate zones, one-hot encoding is required to convert them into binary matrices, eliminating the order relationship between categories and preventing the model from misinterpreting the data. For ordered categories, such as land use classification, label encoding can be used to convert different categories into integer representations. Furthermore, multiple categorical variables can be combined through feature cross-validation to generate new features, improving the model's performance.

[0064] Step a4: Construct historical multi-source feature data based on the original historical multi-source data after standardization and / or category feature processing. The historical multi-source feature data includes features that are associated with pollutants.

[0065] Constructing a complete input feature set is a core step in ensuring model performance. In this embodiment, domain knowledge was incorporated into the feature set construction to create features that are strongly correlated with the target variable and with pollutants. For example, features such as water temperature and precipitation can be interacted to generate new features, capturing the nonlinear relationships between them; features such as pH value are also strongly correlated with pollutants. Furthermore, if the data contains time-series information, time features such as months and seasonal trends can be extracted to enhance the model's ability to identify periodic fluctuations.

[0066] For datasets with a large number of features, feature dimensionality reduction may be necessary. For example, PCA (Principal Component Analysis) can be used to reduce the data from high to low dimensions, preserving important information while reducing noise and redundant features, thus improving the efficiency of model training. In terms of feature selection, correlation analysis and feature importance ranking can be used to select features highly correlated with the target variable, reducing unnecessary feature inputs, thereby avoiding overfitting and improving the model's predictive ability.

[0067] In addition, the same data processing procedure is required when training the model. Finally, the data is divided into training set, validation set and test set to ensure that the model can be optimized and its generalization ability can be verified during the training process, so as to obtain an efficient and robust model.

[0068] In this embodiment, to ensure the accuracy and stability of the model prediction, the original multi-source data needs to be structured to construct a unified and standardized model input feature set, eliminate data noise and redundant information, and provide a reliable foundation for subsequent predictions.

[0069] In some alternative implementations, the base prediction model is pre-trained using historical multi-source data of the target water area; during training, the base prediction model is hyperparameter-tuned using time-series cross-validation.

[0070] To ensure the consistency and generalization ability of the established prediction model over time, this embodiment employs time-series cross-validation to fine-tune the model's hyperparameters. Traditional cross-validation methods randomly divide data into training and validation sets, while this embodiment trains and tests sequentially over time. This approach better simulates the model's prediction process in real-world application scenarios. Specifically, a sliding window method or an expanded training set method can be chosen to train and validate the model through multiple time windows, avoiding information leakage and ensuring good generalization ability. Simultaneously, hyperparameter tuning, such as adjusting the learning rate, tree depth, and regularization parameters, optimizes the model and improves its prediction accuracy. The optimal hyperparameter settings obtained through cross-validation enable the model not only to adapt to past data but also to effectively predict future pollutant concentration changes, thus possessing higher practical value and reliability.

[0071] The following is a complete implementation example using a monitoring section at the inflow point of the Guanting Reservoir as an example.

[0072] Continuous monthly monitoring data were collected from a monitoring section at the inflow point of the Guanting Reservoir from 2012 to 2022. This data covers water quality indicators such as COD, TN, TP, and NH4. + The data include hydrological parameters (such as DO, flow rate, water temperature, pH, etc.), meteorological factors (such as rainfall, evaporation, air temperature, etc.), and socio-economic indicators (such as population size, land use structure, agricultural arable land area, forest coverage, etc.). These data provide comprehensive input information for water quality prediction models.

[0073] During the data preparation phase, pandas tools were used for data cleaning and preprocessing, including missing value imputation, outlier removal, time series alignment, and format standardization to ensure data consistency and integrity. Missing value imputation employed appropriate methods (such as mean imputation or interpolation), while outlier removal was accomplished through statistical analysis methods (such as box plots or standard deviation testing). Time series alignment ensured consistency in timestamps across different data sources, avoiding data bias introduced by time inconsistencies. After these processes, a complete dataset with no missing values, no outliers, and a uniform format was ultimately formed.

[0074] For numerical variables (such as flow rate, water temperature, and precipitation), standardization is used to scale them to zero mean and unit variance to avoid bias in model training caused by variables with different dimensions. Standardization ensures that each feature contributes equally during model training, preventing certain features with larger values ​​from dominating the training process. Furthermore, for categorical variables (such as land use type), OneHotEncoder is used for one-hot encoding, converting categorical information into dummy variables usable for model computation. One-hot encoding eliminates the order relationship between categorical variables by converting each category into a binary variable of 0 and 1, enabling the model to effectively handle categorical data. After these preprocessing steps, an input matrix containing multi-source features is generated, providing a solid data foundation for subsequent model training and validation.

[0075] Finally, to ensure fairness and accuracy in model training and validation, the training and validation sets were divided chronologically. This division of time-series data avoids information leakage, ensuring that the data in the training set is not influenced by information from the validation set, thereby improving the model's generalization ability. This data processing flow provides high-quality data support for subsequent modeling, laying the foundation for the development of water quality prediction models.

[0076] Furthermore, a basic prediction model based on the XGBoost regression algorithm was constructed to predict COD, TN, TP, and NH4, respectively. + The concentration levels of four typical water pollutants (-N) were measured. XGBoost, as an efficient gradient boosting tree model, possesses excellent nonlinear modeling capabilities and automatic feature selection, making it suitable for handling multi-source heterogeneous features in complex environmental data. During model training, the objective function was set to minimize the squared error to measure the deviation between the model's predictions and actual observations.

[0077] To improve the model's generalization ability and stability, a time-series cross-validation method was employed for parameter optimization. This validation strategy fully considers the time dependence of the data, avoids future information leakage, and effectively simulates real-world prediction scenarios. During hyperparameter optimization, a search space was defined, including key parameters such as the maximum tree depth, learning rate, subsampling rate, and minimum leaf node sample weight. Grid search was then used to select the optimal parameter combination based on cross-validation.

[0078] After completing the parameter optimization, various pollutants (COD, TN, TP, NH4) were used as the basis for their selection. +Using (-N) as the target variable and incorporating standardized multi-source feature inputs, four basic prediction models are independently trained. Each model accurately characterizes the complex relationship between the corresponding pollutant and influencing factors, outputting stable and interpretable prediction results. Simultaneously, the model outputs serve as crucial inputs for subsequent residual modeling and fusion modeling, laying a solid foundation for achieving high-precision water quality prediction. Through this strategy, the accuracy, generalization ability, and robustness of the models are effectively improved, meeting the needs of practical water quality monitoring and early warning systems.

[0079] To further improve the accuracy of pollutant concentration prediction, a residual correction-based modeling mechanism is introduced to effectively capture complex nonlinear relationships that the basic prediction model fails to learn. The specific method is as follows: First, the residual sequence between the predicted values ​​of the basic prediction model and the actual observed values ​​is calculated in the training set, i.e., residual = actual value - basic prediction model predicted value. Next, the output of the basic prediction model is used as a new input feature and added to the original feature set to form an expanded feature set. This expanded feature set retains the original input information while incorporating the prediction behavior of the basic prediction model, thus providing a richer context for subsequent residual learning. Using the residual as a new target variable, an XGBoost regressor is constructed for modeling and training to obtain the residual correction model. This model can learn the high-order feature interactions and nonlinear patterns that the basic prediction model failed to effectively characterize during the modeling process, thereby achieving accurate correction of prediction errors. Finally, in practical applications, by adding the predicted values ​​of the basic prediction model to the output of the residual model, a more accurate pollutant concentration prediction result is obtained, significantly improving the overall model performance and the reliability of environmental prediction.

[0080] In the final fusion stage, the predicted values ​​of the basic prediction model and the residual model are superimposed to obtain the final prediction result, that is: fused prediction value = basic prediction model prediction value + residual model prediction value.

[0081] This two-tiered modeling structure maintains the basic prediction model's grasp of the overall trend while utilizing the residual model to enhance the fitting ability to complex disturbances, thereby significantly improving prediction accuracy and robustness.

[0082] Experiments show that the fusion model of this invention has a high coefficient of determination R for the COD index. 2 It improves to 0.99 without causing overfitting.

[0083] Based on the water quality prediction model, a multi-objective optimization problem is further constructed. Decision variables include land use type, population density, precipitation, evaporation, runoff, and pH. The optimization objective is to minimize the overall pollution risk while ensuring that the concentrations of each pollutant do not exceed the national Class II surface water standard. By establishing a multi-constraint, multi-objective mathematical model, it is ensured that the optimization results are both environmentally significant and policy-feasible.

[0084] This implementation of NSGA-II (Non-dominated sorting genetic algorithm-II) is based on the Python pymoo library, referencing... Figure 2 As shown, the initial population, crossover rate, mutation rate, and number of iterations are set. In each generation of evolution, superior individuals are selected through non-dominated sorting and crowding distance calculation, and crossover and mutation operations are performed to generate new solutions. After iterative convergence, a Pareto front solution set is obtained, generating a set of optimal solutions that weigh multiple objectives.

[0085] The final optimization yielded 22 Pareto solutions, corresponding to different combinations of land use and water quality improvement schemes. The predicted results and optimized solution set were then integrated and output to an Excel file for management analysis and decision-making. Ultimately, this effectively controlled COD, TN, TP, and NH4+. + -N concentration meets the standard.

[0086] This invention integrates a basic prediction model and a residual model to achieve accurate prediction of the concentrations of multiple pollutants in water bodies (such as COD, TN, TP, and NH4). + This invention provides high-precision predictions for water quality management systems (including pollutants such as nitrogen, phosphorus, and nitrogen) and incorporates environmental standards as optimization constraints to achieve synergistic optimization control of multiple pollutants. Compared to traditional single-modeling methods, this invention enhances the intelligence, stability, and practicality of water quality management systems, making it suitable for applications such as water environment prediction, pollution control, and intelligent scheduling, and possesses promising engineering application prospects.

[0087] This embodiment also provides a pollutant control device for water quality monitoring, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0088] This embodiment provides a pollutant control device for water quality monitoring, such as... Figure 3 As shown, the device includes: The acquisition module 201 is used to acquire historical multi-source feature data of the target water area; The basic prediction model prediction module 202 is used to input historical multi-source feature data into the basic prediction model corresponding to different pollutants, and predict the pollutant prediction concentration corresponding to different pollutants. The residual model prediction module 203 is used to input the predicted concentrations of pollutants corresponding to different pollutants and historical multi-source characteristic data into the pre-established residual model to obtain the residual prediction values ​​corresponding to different pollutants. The fusion module 204 is used to correct the predicted concentration of the corresponding pollutant based on the residual prediction value corresponding to each pollutant, so as to obtain the target concentration of the pollutant corresponding to each pollutant. The target optimization module 205 is used to determine the pollutant optimization target by taking the target concentration of different pollutants as the optimization target; the pollutant optimization target is used to control the pollutants.

[0089] In some optional implementations, the target optimization module 205 is specifically used for: Obtain the environmental constraints corresponding to each pollutant; Determine the set of environmental control variables for the target water area; A multi-objective optimization problem model is established based on the optimization objective, environmental constraints, and set of environmental control variables. Based on a multi-objective optimization problem model, the Pareto front solution set is determined; the Pareto front solution set includes the optimal combination of environmental control variables for controlling pollutants; the optimal combination of environmental control variables is used as the objective for determining pollutant optimization.

[0090] In some alternative embodiments, the apparatus further includes: The output module is used to generate pollutant control reports and visualize them based on the target concentrations of different pollutants and the Pareto front solution set.

[0091] The data processing module is used to acquire raw historical multi-source data of the target water area. The raw historical multi-source data includes water quality monitoring data, hydrological and meteorological data, and land use information. Data preprocessing is performed on the raw historical multi-source data, including data cleaning, missing value handling, and outlier handling. The raw historical multi-source data after data preprocessing is then standardized and / or subjected to category feature processing. Based on the raw historical multi-source data after standardization and / or category feature processing, historical multi-source feature data is constructed, which includes features that are correlated with pollutants.

[0092] The training module is used to obtain the prediction residuals of the basic prediction model on the training set. The prediction residuals are the differences between the predicted values ​​and the actual values ​​of the basic prediction model. Based on the prediction residuals and historical multi-source data, the initial regression model is trained to obtain the residual model.

[0093] In this embodiment, the pollutant control device for water quality monitoring is presented in the form of a functional unit. Here, a unit refers to an ASIC circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0094] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0095] This invention also provides a computer device having the above-described features. Figure 4 The pollutant control device shown is used in the water quality monitoring process.

[0096] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 4 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 4 Take a processor 10 as an example.

[0097] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.

[0098] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0099] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0100] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0101] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0102] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0103] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0104] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method of contaminant control in a water quality monitoring process, characterized by, The method comprises: acquiring historical multi-source feature data of a target water area; inputting the historical multi-source feature data into different pollutant one-to-one corresponding basic prediction models respectively, to predict pollutant prediction concentrations of different pollutants one-to-one corresponding; inputting the pollutant prediction concentrations of different pollutants one-to-one corresponding and the historical multi-source feature data into a pre-established residual model, to obtain residual prediction values of different pollutants one-to-one corresponding; correcting the pollutant prediction concentrations corresponding based on the residual prediction values of different pollutants one-to-one corresponding, to obtain pollutant target concentrations of different pollutants one-to-one corresponding; taking the pollutant target concentrations of different pollutants one-to-one corresponding as optimization targets, to determine pollutant optimization targets; the pollutant optimization targets are used for controlling pollutants.

2. The method of claim 1, wherein, The taking the pollutant target concentrations of different pollutants one-to-one corresponding as optimization targets, to determine pollutant optimization targets, comprises: acquiring environmental constraint conditions corresponding to each pollutant one-to-one; determining an environmental control variable set of the target water area; based on the optimization targets, the environmental constraint conditions, and the environmental control variable set, establishing a multi-objective optimization problem model; based on the multi-objective optimization problem model, determining a Pareto frontier solution set; the Pareto frontier solution set comprises optimal environmental control variable combinations for controlling pollutants; the optimal environmental control variable combinations are taken as the determination of the pollutant optimization targets.

3. The method of claim 2, wherein, After determining the Pareto frontier solution set, comprising: based on the pollutant target concentrations of different pollutants one-to-one corresponding and the Pareto frontier solution set, generating a pollutant control report and performing visual output.

4. The method of claim 1, wherein, Before the acquiring historical multi-source feature data of a target water area, comprising: acquiring original historical multi-source data of the target water area; the original historical multi-source data comprises water quality monitoring data, hydro-meteorological data, and land use information; performing data preprocessing on the original historical multi-source data; the data preprocessing comprises data cleaning, missing value processing, and abnormal value processing; performing standardization processing and / or category feature processing on the original historical multi-source data after data preprocessing; based on the original historical multi-source data after standardization processing and / or category feature processing, constructing historical multi-source feature data; the historical multi-source feature data comprises features associated with pollutants.

5. The method of claim 1, wherein, The basic prediction model is obtained by pre-training the historical multi-source data of the target water area; when training, a time series cross-validation method is used to perform hyperparameter tuning on the basic prediction model.

6. The method of claim 1, wherein, The residual model is established by the following steps: acquiring prediction residuals of the basic prediction model on a training set; the prediction residuals are differences between prediction values and actual values of the basic prediction model; based on the prediction residuals and historical multi-source data, training an initial regression model, to obtain the residual model.

7. A water quality monitoring process for controlling pollutants, characterized by, The device comprises: an acquisition module, configured to acquire historical multi-source feature data of a target water area; The basic prediction model prediction module is configured to input the historical multi-source feature data into a one-to-one corresponding basic prediction model of different pollutants respectively, and predict a pollutant prediction concentration of the one-to-one corresponding different pollutants; The residual model prediction module is configured to input the pollutant prediction concentration of the one-to-one corresponding different pollutants and the historical multi-source feature data into a pre-established residual model, and obtain a residual prediction value of the one-to-one corresponding different pollutants; The fusion module is configured to correct the pollutant prediction concentration of the one-to-one corresponding different pollutants based on the residual prediction value of the one-to-one corresponding different pollutants, and obtain a pollutant target concentration of the one-to-one corresponding different pollutants. The target optimization module is configured to determine a pollutant optimization target by taking the pollutant target concentration of the one-to-one corresponding different pollutants as an optimization target; and the pollutant optimization target is used for controlling the pollutants.

8. A computer device, comprising: The memory and the processor are communicatively connected, and the memory stores computer instructions. The processor executes the computer instructions to perform the water quality monitoring process pollutant control method of any one of claims 1-6. The computer readable storage medium stores computer instructions for causing the computer to perform the water quality monitoring process pollutant control method of any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer instructions are used to cause the computer to perform the water quality monitoring process pollutant control method of any one of claims 1-6.

10. A computer program product, characterised in that, ​