Drainage basin heavy metal content and load prediction method based on SWAT and machine learning coupling
By combining the SWAT distributed hydrological model with machine learning, time series of multi-dimensional physical characteristic parameters are generated, which solves the limitations of traditional watershed heavy metal monitoring methods and achieves efficient and accurate prediction of heavy metal content and load, as well as dynamic risk warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TECH CENT FOR SOIL AGRI & RURAL ECOLOGY & ENVIRONMENT MINIST OF ECOLOGY & ENVIRONMENT
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-15
AI Technical Summary
Traditional watershed heavy metal monitoring methods struggle to capture the peak and complete process of pollution events, resulting in large deviations in load estimation. This fails to meet the needs of risk warning and proactive management, and high-frequency monitoring is costly, making it difficult to apply to complex and extensive watersheds.
By combining the SWAT distributed hydrological model with machine learning, time series of multi-dimensional physical characteristic parameters are generated, a machine learning regression model is trained to predict heavy metal content and load, and dynamic risk warnings are provided.
It improves the physical interpretability and mechanistic consistency of the prediction model, reduces monitoring costs, enhances the accuracy of prediction and spatiotemporal extrapolation capabilities, and enables forward-looking risk warning.
Smart Images

Figure CN122050591A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of water environment simulation and pollution prediction technology, and more specifically, to a method for predicting watershed heavy metal content and load based on the coupling of SWAT and machine learning. Background Technology
[0002] Heavy metal pollution in watersheds is one of the most serious challenges facing global water environment governance. Heavy metals such as mercury, cadmium, lead, and arsenic, due to their persistence, bioaccumulation, and high toxicity, can amplify through the food chain even at extremely low concentrations, posing a long-term and irreversible threat to the health of aquatic ecosystems and the safety of human drinking water. Accurate monitoring of heavy metal content within watersheds and reliable prediction of their migration fluxes are the cornerstones for implementing precise source tracing, effective control, and scientific early warning systems.
[0003] Traditional watershed heavy metal monitoring primarily relies on "on-site sampling-laboratory analysis." While this method yields accurate results, it has significant limitations. First, heavy metal migration is strongly driven by hydrological processes such as rainfall and runoff, exhibiting instantaneous and highly dynamic characteristics. Discrete, low-frequency sampling (e.g., monthly sampling) struggles to capture the peak and complete process of pollution events, leading to large biases in load estimation. Comprehensive, high-frequency water quality monitoring requires substantial investment in equipment, manpower, and funding, making it unsustainable, especially in complex and extensive watersheds. Furthermore, it cannot predict areas without monitoring points or future periods, failing to meet the needs of risk warning and proactive management.
[0004] Therefore, it is necessary to design a watershed heavy metal content and load prediction method based on the coupling of SWAT and machine learning to solve the problems existing in the current technology. Summary of the Invention
[0005] In view of this, the present invention proposes a method for predicting watershed heavy metal content and load based on the coupling of SWAT and machine learning, aiming to solve the problem of low efficiency of current aggregated payment.
[0006] This invention proposes a watershed heavy metal content and load prediction method based on SWAT coupled with machine learning, including: Collect measured data from the target watershed, preprocess it, and construct a basic database; Based on the aforementioned basic database, a watershed distributed hydrological model is constructed and calibrated, wherein the watershed distributed hydrological model is run to generate time series of multi-dimensional physical characteristic parameters, including at least hydrological variables and water quality-related variables. A machine learning regression model is trained based on the time series of the multi-dimensional physical feature parameters and the basic database. The new feature data generated by the watershed distributed hydrological model is input into the machine learning regression model to predict the heavy metal content of the target time period or location, and the heavy metal load is determined based on the heavy metal content and the multi-dimensional physical feature parameters. Early warning is issued based on the heavy metal load and the heavy metal content.
[0007] Furthermore, the watershed distributed hydrological model is a SWAT model; the multi-dimensional physical characteristic parameters include hydrological variables and water quality-related variables, wherein the hydrological variables are selected from at least one of surface runoff, groundwater runoff, lateral flow, actual evapotranspiration, and soil moisture content, and the water quality-related variables are selected from at least one of sediment load, nitrate concentration, and organic nitrogen load.
[0008] Furthermore, when collecting and preprocessing measured data from the target watershed and constructing the basic database, the following steps are included: The measured data includes at least spatial data of the target watershed, meteorological and hydrological data, pollution source data, and measured data of heavy metal concentrations. The preprocessing includes at least unifying the spatial coordinate system of all data to the same reference and unifying the time reference of all time series data with the simulation step size; Based on the preprocessed data, a spatiotemporally consistent basic database is constructed. In constructing the basic database, the time series of multi-dimensional hydrological and water quality parameters generated by the watershed distributed hydrological model are aligned with the sampling time points of the measured heavy metal concentration data to form a feature matrix and label vector for training the machine learning regression model.
[0009] Furthermore, when constructing and calibrating a distributed hydrological model for a watershed based on the aforementioned basic database, the following steps are included: Based on the spatial data, the watershed space is discretized, and model calculation units are defined; Load the meteorological and hydrological data and pollution source data, and configure the model driving conditions; Using the measured river flow data from the meteorological and hydrological data, the model is subjected to parameter sensitivity analysis, automatic optimization calibration and performance verification until the evaluation index of simulated flow and measured flow reaches the preset accuracy standard. After the model is calibrated and validated, hydrological variables and water quality-related variables related to the migration and transformation of heavy metals are extracted to generate the time series of the multidimensional hydrological and water quality parameters.
[0010] Furthermore, when training a machine learning regression model based on the time series of the multi-dimensional physical feature parameters and the basic database, the process includes: The machine learning regression model is an algorithm suitable for small sample regression, selected from at least one of support vector regression, long short-term memory network, Transformer model, XGBoost model and combination thereof; During or after model training, the SHAP value analysis method is used to evaluate the contribution of each input feature to the prediction results.
[0011] Furthermore, when inputting the new feature data generated by the watershed distributed hydrological model into the machine learning regression model to predict the heavy metal content at a target time period or location, the following steps are included: The multi-dimensional hydrological and water quality parameters generated by the watershed distributed hydrological model for the target time period or location are used as input feature data. The input feature data is input into the trained machine learning regression model for calculation; Obtain the predicted heavy metal content output by the machine learning regression model, corresponding to the target time period or location, and perform uncertainty analysis on the predicted heavy metal content.
[0012] Furthermore, when performing uncertainty analysis on the predicted heavy metal content, the following steps are included: Based on the feasible set of parameters of the watershed distributed hydrological model, multiple sets of different multi-dimensional hydrological and water quality parameters are generated. Each set of parameters is input into an ensemble model composed of multiple machine learning regression models constructed by the bootstrap sampling method, and corresponding predicted values are obtained to form a set of predicted values. Statistical analysis is performed on the predicted value set to calculate and output the confidence interval of the predicted heavy metal content.
[0013] Furthermore, when determining the heavy metal load based on the heavy metal content and the multi-dimensional physical characteristic parameters, the following steps are included: Obtain hydrological flow data simulated by the watershed distributed hydrological model for the same target time period or location; The hydrological flow data is coupled with the predicted heavy metal content to calculate the heavy metal load at the target time period or location.
[0014] Furthermore, when issuing an early warning based on the heavy metal load and the heavy metal content, it includes: The predicted heavy metal content is compared with a preset concentration threshold, and / or the predicted heavy metal load is compared with a preset load threshold; Based on the comparison results, generate and send corresponding early warning information; Specifically, when the predicted value of the heavy metal content exceeds the first concentration threshold and / or the predicted value of the heavy metal load exceeds the first load threshold, a first-level early warning message is generated and sent. When the predicted value of the heavy metal content exceeds the second concentration threshold but does not exceed the first concentration threshold, and / or the predicted value of the heavy metal load exceeds the second load threshold but does not exceed the first load threshold, a secondary warning message is generated and sent. Wherein, the first concentration threshold is higher than the second concentration threshold, and the first load threshold is higher than the second load threshold.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention improves the physical interpretability and mechanism consistency of the prediction model. Traditional pure data-driven machine learning models are often regarded as "black boxes", and their prediction results lack clear physical process support, making it difficult for environmental managers to understand and trust them. The present invention introduces a rigorously calibrated SWAT model as a physical feature generator, and uses a series of parameters with clear hydrological, hydrodynamic and environmental geochemical significance, such as surface runoff, groundwater runoff, soil moisture content, and sediment load, as inputs to the machine learning model. This makes the final prediction model no longer a simple fitting of data curves, but based on the indirect characterization of the basic physicochemical process of "how water flow carries sediment and how sediment adsorbs and transports heavy metals". The prediction results therefore have a solid physical mechanism background, the credibility and acceptability of the model are fundamentally enhanced, and decision-makers can understand the dominant environmental factors behind the prediction.
[0016] (2) This invention reduces the over-reliance on high-frequency, high-density field monitoring data, thereby reducing the economic cost and manpower investment of long-term monitoring. Precise analysis of heavy metals is costly and complex, resulting in sparse and spatially limited monitoring data for most watersheds. This invention utilizes the SWAT model to generate continuous, spatiotemporally complete "surrogate features." While these features do not directly measure heavy metals, they precisely characterize the key environmental conditions driving their migration and transformation. The machine learning model only needs to use limited measured concentration data collected at key spatiotemporal nodes as "labels" to learn the complex relationship between these continuous physical features and the target pollutant. Without the need for costly encrypted monitoring, state estimation can be achieved for periods and locations without monitoring, allowing limited monitoring resources to be used for the most critical model validation and calibration stages, thus optimizing cost-effectiveness.
[0017] (3) This invention comprehensively improves the accuracy, robustness, and spatiotemporal extrapolation capability of prediction. Machine learning models trained solely on a small amount of measured data are prone to overfitting, and their prediction performance drops sharply when encountering hydrological and meteorological conditions that have not been experienced. The multidimensional features provided by the SWAT model constitute a comprehensive and stable environmental state description system, which can more fully reflect the complex factors affecting the fate of heavy metals (such as hydrodynamic intensity, sediment transport flux, soil moisture conditions, etc.). This enables the machine learning model to learn more essential and universal laws, rather than just memorizing limited samples. Therefore, the model performs more stably in watersheds with scarce data, and also shows stronger adaptability and reliability in predictions under extreme rainfall events or long-term climate change scenarios, reducing the risk of prediction failure due to insufficient model extrapolation capability.
[0018] (4) This invention enables dynamic risk early warning, whereas traditional monitoring can only report pollution conditions after they have occurred. By coupling prediction and early warning mechanisms, this invention can proactively calculate heavy metal flux loads for future periods and, combined with preset environmental quality standards and risk assessment thresholds, achieve tiered early warning. This directly transforms technical output into management action instructions, thereby shifting environmental management from passive response to proactive prevention and control, and improving the timeliness and effectiveness of risk management. Attached Figure Description
[0019] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 The flowchart illustrates a watershed heavy metal content and load prediction method based on SWAT coupled with machine learning, provided in an embodiment of the present invention. Detailed Implementation
[0020] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specified, embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0021] In some embodiments of this application, see Figure 1 As shown, a watershed heavy metal content and load prediction method based on SWAT coupled with machine learning includes: Collect measured data from the target watershed, preprocess the data, and construct a basic database.
[0022] A distributed hydrological model of the watershed is constructed and calibrated based on a basic database. The distributed hydrological model of the watershed is run to generate time series of multi-dimensional physical characteristic parameters, including at least hydrological variables and water quality-related variables.
[0023] Machine learning regression models are trained based on time series data and a basic database of multi-dimensional physical feature parameters.
[0024] The new feature data generated by the watershed distributed hydrological model are input into the machine learning regression model to predict the heavy metal content of the target time period or location, and the heavy metal load is determined based on the heavy metal content and multi-dimensional physical characteristic parameters.
[0025] Early warning is based on heavy metal load and heavy metal content.
[0026] Specifically, heavy metal pollution poses a serious threat to aquatic ecosystems and human health due to its persistence, bioaccumulation, and high toxicity. Accurate prediction of heavy metal content and load within watersheds is crucial for effective pollution management and control. Currently, major prediction methods have the following limitations: 1. Limitations of traditional mechanistic models: Models like the SWAT model, based on physical mechanisms, can effectively simulate hydrological cycles and the migration and transformation of conventional pollutants (such as nitrogen and phosphorus). However, their simulation of trace pollutants like heavy metals is typically complex, requiring high precision in basic data and computational resources. Furthermore, parameter calibration is difficult in data-scarce watersheds, making it hard to guarantee prediction accuracy. 2. Limitations of purely data-driven models: Machine learning models (such as neural networks, random forests, and XGBoost) excel at handling nonlinear, high-dimensional data. However, these "black box" models heavily rely on large amounts of high-quality monitoring data. Given that heavy metal monitoring data is typically sparse and discontinuous, these models suffer from insufficient generalization ability and poor interpretability, making it difficult to reveal the physical processes of pollutant migration and transformation. In recent years, there has been a research trend of coupling SWAT models with machine learning, but existing research has mostly focused on runoff simulation or prediction of conventional water quality indicators. Applying these coupling ideas to the prediction of heavy metal pollution, which has more complex mechanisms and scarcer data, remains an under-explored and challenging field. This invention aims to fill this gap by using the SWAT model as a feature generator for physical information, and using its output of multi-dimensional hydrological and water quality variables with clear physical meaning as input features for a machine learning model. This constructs a "physical mechanism-guided" machine learning regression model, ultimately achieving accurate prediction of heavy metals. For example, predicting the mercury (Hg) load in the estuary of a river basin with frequent mining activities: 1. Data Preparation: Collect watershed DEM, land use, and soil data. Meteorological data. Location and emission data of major mercury-related enterprises' wastewater discharge outlets. Monthly measured mercury concentration data at the estuary section (for two consecutive years).
[0027] 2. SWAT Modeling: Construct a SWAT model of the watershed and calibrate and validate it using historical runoff data. Incorporate industrial mercury emission sources as point source loads into the model.
[0028] 3. Feature generation: Run SWAT to output time series data of 10 key features at the daily scale, including surface runoff, groundwater runoff, sediment load, and nitrate concentration.
[0029] 4. Machine Learning Training: The SWAT features mentioned above were mapped to measured mercury concentration values over 24 months to construct a dataset. SVR or XGBoost models were used for training. SHAP analysis showed that "surface runoff" and "sediment load" were the two most important features affecting mercury concentration prediction, consistent with the understanding that heavy metals easily bind to particulate matter.
[0030] 5. Prediction and Validation: The trained model was used to predict the daily mercury concentration in the third year and validated with the newly added measured data for that year. The expected Nash-Sutcliffe efficiency coefficient (Ens) is above 0.75, showing good predictive performance.
[0031] Understandably, this invention enhances the physical interpretability and mechanistic consistency of the predictive model. Traditional purely data-driven machine learning models are often considered "black boxes," lacking clear physical process support for their predictions, making them difficult for environmental managers to understand and trust. This invention introduces a rigorously calibrated SWAT model as a physical feature generator, using parameters with clear hydrological, hydrodynamic, and environmental geochemical significance, such as surface runoff, groundwater runoff, soil moisture content, and sediment load, as inputs to the machine learning model. This ensures that the final predictive model is no longer simply a fit to data curves, but rather built upon an indirect characterization of the fundamental physicochemical processes of "how water flow carries sediment and how sediment adsorbs and transports heavy metals." The prediction results thus possess a solid physical mechanism background, fundamentally enhancing the model's credibility and acceptability, enabling decision-makers to understand the dominant environmental factors behind the predictions.
[0032] This invention reduces over-reliance on high-frequency, high-density field monitoring data, thereby lowering the economic costs and manpower required for long-term monitoring. Precise analysis of heavy metals is costly and complex, resulting in sparse and spatially limited monitoring data for most watersheds. This invention utilizes a SWAT model to generate continuous, spatiotemporally complete "surrogate features." While these features do not directly measure heavy metals, they finely characterize the key environmental conditions driving their migration and transformation. The machine learning model only needs limited measured concentration data collected at key spatiotemporal nodes as "labels" to learn the complex relationships between these continuous physical features and target pollutants. It eliminates the need for costly, intensive monitoring, enabling state estimation for periods and locations without monitoring. This allows limited monitoring resources to be used for the most critical model validation and calibration, achieving cost-effectiveness optimization.
[0033] This invention comprehensively improves the accuracy, robustness, and spatiotemporal extrapolation capability of predictions. Machine learning models trained solely on limited measured data are prone to overfitting, and their predictive performance drops sharply when encountering hydrological and meteorological conditions they have not experienced. The multidimensional features provided by the SWAT model constitute a comprehensive and stable environmental state description system, which can more fully reflect the complex factors influencing heavy metal fate (such as hydrodynamic intensity, sediment transport flux, and soil moisture conditions). This allows the machine learning model to learn more fundamental and universal laws, rather than simply memorizing limited samples. Therefore, the model performs more stably in data-scarce watersheds and exhibits stronger adaptability and reliability in predictions under extreme rainfall events or long-term climate change scenarios, reducing the risk of prediction failure due to insufficient model extrapolation capability.
[0034] This invention enables dynamic risk early warning, whereas traditional monitoring can only report pollution conditions after they have occurred. By coupling prediction and early warning mechanisms, this invention can proactively calculate heavy metal flux loads for future periods and, combined with preset environmental quality standards and risk assessment thresholds, achieve tiered early warning. This directly transforms technological output into management action instructions, thereby shifting environmental management from passive response to proactive prevention and control, improving the timeliness and effectiveness of risk management.
[0035] Finally, this invention broadens the application scenarios and scientific value of technical tools. It not only serves routine pollution status assessments, but its high spatiotemporal resolution content and load data can also deeply support a series of advanced analytical needs, such as pollution source tracing and analysis, quantification of the contribution rate of key pollution sources, simulation and evaluation of the effectiveness of pollution control projects, and long-term ecological risk trend assessment. It provides an unprecedented dynamic quantitative tool for formulating precise watershed pollution control plans, optimizing the allocation of pollution control resources, and assessing the environmental impacts of land use or climate change, possessing profound scientific significance and broad application prospects.
[0036] In some embodiments of this application, the watershed distributed hydrological model is a SWAT model. The multidimensional physical characteristic parameters include hydrological variables and water quality-related variables, wherein the hydrological variables are selected from at least one of surface runoff, groundwater runoff, lateral flow, actual evapotranspiration, and soil moisture content, and the water quality-related variables are selected from at least one of sediment load, nitrate concentration, and organic nitrogen load.
[0037] In some embodiments of this application, when collecting measured data from the target watershed, preprocessing it, and constructing a basic database, the process includes: The measured data should include at least spatial data of the target watershed, meteorological and hydrological data, pollution source data, and measured data of heavy metal concentrations.
[0038] Preprocessing includes at least unifying the spatial coordinate system of all data with the same reference and unifying the time reference of all time series data with the simulation step size.
[0039] Based on the preprocessed data, a spatiotemporally consistent basic database is constructed.
[0040] In constructing the basic database, the time series of multi-dimensional hydrological and water quality parameters generated by the distributed hydrological model of the watershed are aligned with the sampling time points of the measured heavy metal concentration data to form a feature matrix and label vector for training the machine learning regression model.
[0041] Specifically, the following data will be collected for the target watershed: Spatial data: Digital Elevation Model (DEM), Land Use / Cover (LUCC), Soil Type Map.
[0042] Meteorological and hydrological data: daily meteorological data such as precipitation, temperature, wind speed, and sunshine. River runoff data (used for calibrating the SWAT model).
[0043] Pollution source data: Inventory of heavy metal emissions from point sources and non-point sources, such as industrial sewage outlets, urban domestic sewage, and mining areas.
[0044] Measured data: Water samples were collected periodically from representative sections within the watershed, and the concentrations of heavy metals were measured (used as training labels for machine learning models).
[0045] The above data is preprocessed, including formatting, missing value imputation, coordinate unification and normalization, to build a spatiotemporally consistent basic database.
[0046] Digital elevation models (DEMs) are used to define watershed boundaries, drainage systems, and topographic slopes. Land use maps (LUCCs) and soil type maps are used to delineate hydrological response units (HRUs). To ensure the model can correctly identify spatial relationships, all spatial data must completely cover the target watershed geographically and achieve strict spatial alignment after preprocessing.
[0047] Use Geographic Information System (GIS) software (such as ArcGIS, QGIS) or a programming library (such as GDAL). First, specify the same projected coordinate system (e.g., WGS84 / UTMzone50N) for all raster data (DEM, LUCC, soil type). Then, using the DEM as the spatial reference, adjust the raster resolution of the LUCC and soil type maps to match the DEM through resampling techniques, and ensure that the boundaries of all layers completely coincide with the watershed boundaries through clipping operations, generating a spatial data stack with perfect pixel-to-pixel matching.
[0048] Meteorological data drives the hydrological cycle. Measured river flow data is used to calibrate the model. Pollution source data serves as the boundary condition for heavy metal inputs. All time-series data must be standardized to the time step required for model operation (typically daily) by creating a unified time index table, such as a continuous daily series from the start year-month-day to the end year-month-day. Collected meteorological and hydrological data, which may be hourly, daily, or irregular, are mapped to this daily index through aggregation (e.g., accumulating hourly rainfall into daily rainfall) or allocation (e.g., distributing monthly emissions evenly across days). For point source pollution, the geographical locations in the emission inventory need to be spatially matched with the river network nodes generated by the SWAT model to determine the specific sub-basins and river segments into which the pollution originates.
[0049] Measured heavy metal concentration data serves as the sole "true value" label for training and validating machine learning models. Its core characteristic is sparsity, meaning the sampling frequency is far lower than the model's simulation time step (e.g., monthly vs. daily). For this type of data, in principle, no time interpolation aimed at increasing data points (e.g., interpolating monthly data to daily data) is performed to avoid introducing unverifiable errors. Its value lies in its accuracy at specific time points. Preprocessing only needs to ensure the accuracy of the sampling time and convert it to a uniform time format. A spatiotemporally consistent foundational database is constructed, where spatiotemporal consistency refers to spatial consistency: all data points to the same geographic entity (target watershed). After preprocessing, the attribute information of any geographic coordinate point on the DEM, LUCC, and soil layers is clear and consistent. Temporal consistency: all processes follow the same timeline. The SWAT model runs on a daily basis, and the output feature sequence is daily values. Meteorological driving data is daily values. Measured heavy metal data is labeled with specific sampling days. The entire system's time base is synchronized, allowing for alignment operations, which are precise relational queries and extraction processes based on time keys. The inputs, processing, and outputs are as follows: Input A: The output file after the SWAT model has been successfully calibrated (e.g., output.rch represents the river channel output, output.sub represents the sub-basin output). Extract the required hydrological and water quality variables (e.g., SURQ_GEN surface runoff, SYLD sediment load) from this file to form a feature summary table. Each row in this table represents a simulation day, and each column represents a different feature variable. Input B: Measured heavy metal concentration data table. Each row represents a sampling event, including the sampling date and concentration value. Processing procedure: Iterate through each row of input B (i.e., each sampling record), read the sampling date of that record, find the row in input A (feature table) that exactly matches the sampling date, extract all pre-selected feature variable values (e.g., surface runoff, sediment load, soil moisture content, etc., a total of N) from that row to form an N-dimensional feature vector, pair this feature vector with the heavy metal concentration value corresponding to this sampling date to form a complete training sample (feature vector, concentration label), output: Repeat the above process until all valid sampling records are processed, finally generating a machine learning dataset with a sample size equal to the number of valid samplings. In this dataset, the features of each sample accurately reflect the physical state of the watershed on the sampling day.
[0050] Understandably, this invention enhances the physical interpretability and mechanistic reliability of the prediction system. Traditional single-function machine learning predictions are like "black boxes," with conclusions difficult to trace back to specific environmental processes, resulting in a lack of solid scientific basis for management decisions. This solution uses a rigorously calibrated SWAT model as its core physical engine. Its output variables, such as surface runoff, soil moisture content, and sediment load, all possess clear hydrological, hydrodynamic, and environmental geochemical connotations. For example, sediment load is directly related to the particulate adsorption and transport potential of heavy metals, while runoff along different pathways reveals possible surface erosion or underground leaching sources of pollutants. The machine learning model, based on this foundation, essentially establishes a complex mapping relationship between these interpretable physical driving factors and the concentration of target pollutants. Therefore, the final prediction result is no longer an abstract data fitting curve, but a credible inference rooted in the physical picture of "water-sediment-pollutant" migration and transformation, strengthening the foundation of the model output at the level of scientific understanding and decision-making trust. This invention reduces the absolute dependence on high-frequency, high-cost on-site monitoring, thereby optimizing the resource allocation efficiency of environmental monitoring. Precise chemical analysis of heavy metals is time-consuming, labor-intensive, and costly, resulting in highly sparse temporal and unevenly distributed measured data in most watersheds, constituting a major bottleneck for accurate prediction. This solution creatively utilizes the SWAT model to simulate and generate continuous, complete, and spatiotemporally covered physical characteristic sequences across the entire watershed. While these sequences do not directly measure heavy metals, they meticulously characterize all the key environmental conditions driving their migration. The machine learning model only needs to use a limited number of valuable measured concentrations obtained at key time points as "calibration anchors" to learn a robust correlation between continuous physical states and instantaneous pollution responses. This means that state estimation and comprehensive reconstruction of historical sequences for dates without monitoring can be achieved without implementing extremely costly intensive monitoring programs. Limited human and financial resources can be focused on the most critical control sections and validation stages, achieving an optimal balance between monitoring costs and model performance. This invention improves the model's predictive robustness, accuracy, and spatiotemporal extrapolation capabilities in data-scarce scenarios. Machine learning models trained purely on sparse measured data are prone to overfitting and exhibit highly unstable prediction performance for conditions not experienced by meteorological and hydrological data. In this approach, the multi-dimensional features provided by SWAT constitute a comprehensive, stable, and physically consistent environmental state description system, which can more fundamentally reflect the complex mechanisms affecting heavy metal fate, such as hydrodynamic intensity, carrier transport flux, and soil chemical environment. This enables the machine learning model to capture more universal patterns, rather than simply memorizing limited sample points. Therefore, the model exhibits stronger adaptability and generalization ability when facing watersheds with limited data, extreme rainfall events, or future climate scenarios, reducing the risk of prediction failure due to insufficient training data or conditional extrapolation, and making the prediction results more reliable.
[0051] In some embodiments of this application, the process of constructing and calibrating a distributed watershed hydrological model based on a basic database includes: Based on spatial data, the watershed is spatially discretized, and model computational units are defined.
[0052] Load meteorological and hydrological data and pollution source data, and configure model-driven conditions.
[0053] Using measured river flow data from meteorological and hydrological data, parameter sensitivity analysis, automatic optimization calibration, and performance verification are performed on the model until the evaluation indicators of simulated flow and measured flow reach the preset accuracy standards.
[0054] After the model is calibrated and validated, hydrological variables and water quality-related variables related to the migration and transformation of heavy metals are extracted to generate time series of multi-dimensional hydrological and water quality parameters.
[0055] Specifically, the process begins with the delineation of the river network and sub-basins: Load the DEM (Diagram of the river system) into a graphical interface such as ArcSWAT, QSWAT, or the QGIS SWAT plugin. First, specify a catchment area threshold (e.g., 5000 hectares). Based on this threshold, the software will automatically generate a simulated river network by calculating flow direction and accumulation. Then, the user needs to manually define or have the software automatically identify the watershed outlet at the target river cross-section (usually a heavy metal monitoring section or watershed outlet). The model will then automatically delineate sub-basins using this outlet as the endpoint. Each sub-basin is the smallest hydrogeographic unit with an independent outlet, whose internal flow ultimately converges to that outlet, resulting in a set of sub-basin polygons and their corresponding river network vector lines.
[0056] Hydrological Response Unit (HRU) Delineation: After completing the sub-basin delineation, the LUCC map and soil type map are loaded sequentially. The model will read the area distribution of different land use types and soil types within each sub-basin. Users need to set the area percentage threshold for HRU delineation (e.g., a certain land use type accounts for ≥5% of the sub-basin area, and a certain soil type accounts for ≥10% of the sub-basin area). The model will further delineate each sub-basin based on the combination of "sub-basin—land use—soil type," generating multiple HRUs. Each HRU has uniform land use and soil attributes and is the most basic spatial calculation unit for the SWAT model to simulate water balance, sediment, and pollutants, thus obtaining an HRU distribution list and attribute table, which records the sub-basin, land use code, soil code, and area percentage of each HRU.
[0057] Next, configure the model-driven conditions: save the pre-processed daily-scale meteorological data files (such as precipitation PCP, maximum and minimum temperatures TMP, etc.) into the designated folder according to the model's requirements. In the model interface, spatially associate the meteorological data with each sub-basin by defining the meteorological station locations or using grid data. A common method is the "nearest distance method," where each sub-basin uses the meteorological station data closest to its centroid to complete the meteorological generator configuration, ensuring that each sub-basin has corresponding meteorological driving data every day during the simulation period. Then, input the pollution source data: Point source input: In the point source loading module of the model, spatially match the geographical location (latitude and longitude) of point sources such as industrial sewage outlets with the river network generated by SWAT to determine the specific river segment into which they discharge (located in which sub-basin). Then, according to the time step (day or month) required by the model, create an input file containing flow rate and pollutant (here, heavy metals, but in SWAT, they are often treated as conserved substances or pollutants attached to sediment) concentration, and associate it with the river segment. Non-point source input: For non-point sources such as mining areas and urban areas, their spatial distribution maps are overlaid with sub-basins or HRUs for analysis to determine the affected areas. Then, the input of non-point source load is approximately characterized by modifying the management operation files of the corresponding HRUs or the initial soil pollutant concentrations.
[0058] Finally, calibration is performed: In SWAT-CUP, select the parameters to be analyzed, such as the number of runoff curves, effective soil moisture content, and groundwater reevaporation coefficient, and set a reasonable range of physical variation for each parameter. Select a sensitivity analysis method (such as the global LH-OAT method). Then calculate the sensitivity index of each parameter's change on the simulation results (such as daily runoff). Obtain a parameter sensitivity ranking list. Highly sensitive parameters are the focus of subsequent calibration.
[0059] Next, automatic optimization and calibration are performed: an optimization algorithm (such as SUFI-2, PSO) is selected. The simulation period is divided into calibration periods. An objective function is defined, typically using the Nash efficiency coefficient (NSE) as the core indicator. Calibration parameters and their ranges are set in SWAT-CUP, and automatic optimization is initiated. The algorithm iterates continuously, adjusting parameter combinations, running the SWAT model, and calculating the NSE value between simulated and measured flow rates to find the optimal parameter set that maximizes NSE. In addition to NSE, the coefficient of determination and percentage deviation are also considered. Generally, when the calibration period satisfies NSE > 0.7, coefficient of determination > 0.6, and percentage deviation < 15%, the hydrological simulation performance of the model reaches an acceptable standard. The model can then be validated. After the model passes validation, the finally configured SWAT model is run to generate comprehensive output results, from which features for machine learning prediction are extracted.
[0060] Understandably, this invention enhances the model's ability to accurately and precisely represent complex physical processes within a watershed. By spatially discretizing the watershed, decomposing the continuous geographic space into sub-watersheds and hydrological response units (HRUs), the model can accurately characterize the impact of spatial heterogeneity of topography, soil, and land use on the hydrological cycle. Loading and precisely configuring meteorological and pollution source-driven data ensures that the external boundary conditions simulated by the model conform to reality. In particular, the spatial matching and quantitative characterization of point source and non-point source pollution inputs provide a clear physical input framework for subsequent analysis of the "source-sink" relationship of heavy metals. This entire process transforms the final model from an abstract set of mathematical equations into a "digital twin" that faithfully reflects the unique underlying surface characteristics and hydrological response mechanisms of the target watershed. The simulated flow paths, runoff generation and sinking sequences, and pollutant transport environment exhibit high geographical realism and process reliability. Through systematic calibration and verification, this invention reduces the subjective uncertainty and arbitrariness of the model simulation results, enhancing the objectivity, credibility, and scientific rigor of its output results. In traditional modeling processes, parameter selection often relies on experience, resulting in significant subjectivity. This invention introduces parameter sensitivity analysis, which scientifically identifies the core parameters most critical to the simulation results, making the calibration work more focused and targeted. An automatic optimization algorithm is employed, along with internationally recognized indicators such as the Nash efficiency coefficient, for quantitative calibration and verification, establishing an objective and transparent model performance evaluation standard. Only when simulated flow and long-term measured sequences simultaneously reach preset high accuracy standards across multiple statistical indicators is the model considered reliable. This rigorous "examination" process ensures that the final model used to generate features is a high-quality model with stable performance, independently verified by data, thus fundamentally reducing the risk of systematic biases in subsequent machine learning predictions due to errors in the model's structure or parameters. This invention reduces reliance on the long-term experience and repeated trial and error of professional modelers, improving the standardization and reproducibility of complex hydrological model construction. The steps of model construction, discretization, driver configuration, calibration, and verification are clearly decomposed and logically described, forming a standardized operating procedure. By using integrated graphical interface tools and automated calibration software (such as SWAT-CUP), even environmental researchers or data analysts who are not hydrological modeling experts can build a reliable hydrological model under clear step-by-step guidance. This standardized and tool-based process lowers the technical threshold and improves the transferability and reproducibility of the method when applied in different watersheds and by different teams, which is conducive to the promotion and standardized application of this technology. The core output of this invention is a multi-dimensional time series of hydrological and water quality parameters, which provides high-quality, high-information-density input features for machine learning models, improving the efficiency of subsequent data-driven analysis.These features are not raw, coarse meteorological observation data, but rather secondary variables generated after being digested and transformed by physical models. They directly characterize key environmental states affecting heavy metal migration. For example, they provide not raw rainfall data, but surface runoff and soil moisture content transformed from rainfall. They directly output sediment load, rather than erosion dynamic indicators that require indirect calculation. These features have clear physical meaning and are highly correlated with the environmental behavior of heavy metals (such as dissolution, adsorption, and influx transport mechanisms). Therefore, they provide highly valuable and interpretive "nutrients" for machine learning models, enabling them to learn the essential relationship between pollutant concentration and the driving environment more efficiently and accurately, avoiding the tedious process of inefficient feature engineering starting from raw data.
[0061] In some embodiments of this application, training a machine learning regression model based on time series data of multi-dimensional physical feature parameters and a basic database includes: The machine learning regression model is an algorithm suitable for small sample regression, selected from at least one of support vector regression, long short-term memory network, Transformer model, XGBoost model and combination thereof.
[0062] During or after model training, the SHAP value analysis method is used to evaluate the contribution of each input feature to the prediction results.
[0063] Specifically, suppose the dataset contains M samples (i.e., M valid heavy metal sampling events). The feature vector X of each sample... i It is an N-dimensional vector , corresponding to the values of N hydrological and water quality characteristic variables (such as surface runoff, sediment load, etc.) generated by the SWAT model at the i-th sampling time. The corresponding label Y i The measured heavy metal concentration value C in this sampling is... iThe dataset is divided into a test set and a training set. Support Vector Regression (SVR): Its core is to find a "gaps" that allows as many samples as possible to fall within this zone while minimizing the bias of samples outside the zone. For small sample data, SVR maps the data to a high-dimensional space using kernel functions (such as radial basis functions, RBF), effectively capturing nonlinear relationships and reducing the risk of overfitting. XGBoost: A gradient boosting tree algorithm. It sequentially builds multiple decision trees, with each tree learning to correct the residuals of the previous tree. Its built-in regularization term effectively prevents overfitting under small sample conditions and automatically handles the interactions between features, typically resulting in high prediction accuracy. Long Short-Term Memory (LSTM) or Transformer: Suitable for scenarios where feature data has strong temporal dependencies. If heavy metal concentration is related not only to the environmental conditions of the day but also to the hydrological conditions of the previous few days (such as the previous week) (such as accumulated soil moisture and previous rainfall), these sequence models are suitable. They can automatically learn long-term dependency patterns in time series. Combination: Multiple models can be trained, and their predictions can be combined using averaging, weighted averaging, or stacking to integrate the advantages of different models, improving prediction stability and accuracy. The training and test sets are input into the training interface, automatically optimizing their internal parameters (such as weights in SVR, tree structure in XGBoost, and network weights in LSTM) to minimize the loss function (mean squared error) between predicted and measured concentrations. To avoid overfitting and improve performance, the model's hyperparameters (such as the penalty coefficient C and kernel coefficient γ in SVR, tree depth and learning rate in XGBoost, and hidden layer size in LSTM) need to be optimized. K-fold cross-validation combined with grid search or random search is used on the training set. The training set is randomly divided into K parts, with one part used as the validation set and the remaining K-1 parts as sub-training sets. This training and validation is repeated K times. Finally, the average of the K validation scores is used as an estimate of the hyperparameter performance of this group. This process ensures that the evaluation of hyperparameters is robust. In the predefined hyperparameter combination space, the hyperparameter combination that optimizes the cross-validation average score (such as negative mean square error) is found through grid search (traversing all combinations) or a more efficient random search. The optimal hyperparameter combination is then used to retrain the model using the entire training set to obtain the final trained machine learning regression model.
[0064] To demystify the "black box" of machine learning models and quantitatively reveal the contribution of each input feature (i.e., each hydrological and water quality variable generated by SWAT) to a single predicted sample and the overall model, the SHAP value analysis method is used to evaluate the contribution of each input feature to the prediction result. SHAP values are based on cooperative game theory, assigning an importance value to each feature. For any given sample's prediction result, its SHAP value explains the direction (positive / negative) and extent to which each feature value of that sample deviates from the final prediction value relative to the average value of that feature across all samples. The Python shap library is used, loading the trained model. SHAP values are calculated on the test set or the entire dataset. Global interpretation: Plotting a summary map of SHAP values for all samples (such as a bee colony diagram) visually shows which feature (such as "sediment load") has the largest average absolute SHAP value, indicating it has the greatest global influence. Local interpretation: For a specific high-risk prediction day (e.g., a very high predicted concentration), the SHAP value of that sample can be extracted. The results show that the extremely high "surface runoff" characteristic value on that day likely contributed a positive SHAP value (increasing the predicted concentration), while the lower "soil moisture content" contributed a negative SHAP value (suppressing the concentration). This makes the predictions transparent and in line with physical intuition. Through SHAP analysis, it not only verifies whether the characteristics generated by SWAT are as important as the physical mechanisms expect (e.g., confirming that "sediment loading" is a key factor), but also enhances the credibility of the entire coupled model in the minds of environmental decision-makers, because any prediction can be traced back to specific environmental drivers.
[0065] In some embodiments of this application, when inputting new feature data generated by a watershed distributed hydrological model into a machine learning regression model to predict the heavy metal content at a target time period or location, the following steps are included: Multi-dimensional hydrological and water quality parameters generated by the watershed distributed hydrological model for the target time period or location are used as input feature data.
[0066] The input feature data is fed into the trained machine learning regression model for calculation.
[0067] Obtain the predicted heavy metal content values corresponding to the target time period or location from the output of the machine learning regression model, and perform uncertainty analysis on the predicted heavy metal content values.
[0068] It is understandable that for the target prediction period (such as a future period or a period without historical monitoring data) or the target spatial location (such as the outlet of a specific sub-basin), the SWAT model that has been calibrated and validated is used for simulation.
[0069] When running the model, meteorological driving data and pollution source data for the corresponding time period are loaded. The model will output a daily series of multi-dimensional hydrological and water quality parameters for the target spatial range within that time period. Assuming the prediction period is T days and the number of features is N, the generated feature data is a matrix X of size T x N. new Each line X new(t) The vector of N eigenvalues on day t is [F1(t),F2(t),...,FN(t)].
[0070] To ensure consistency with the data scale during model training, X must be... new Perform the exact same preprocessing transformation as during the training phase, using the scaler parameters saved during training, on X. new Transform each column of features in the data.
[0071] Normalization is used during training: Among them, X new This is a matrix (size: T rows × N columns), i.e., the original feature matrix for the forecast period, where T is the total number of days (or time steps) in the forecast period, and N is the total number of feature variables (i.e., the number of hydrological and water quality parameters extracted from the SWAT model, such as surface runoff, sediment load, etc.). This represents the original simulated value of the j-th feature variable on day t, where j is the feature index, indicating that we are processing the j-th feature variable. The value of j ranges from 1 to N, and : represents all rows. This is the minimum value of the j-th feature variable in the entire training dataset used during the training phase. It is a constant calculated and stored during model training. This represents the maximum value of the j-th feature variable across the entire training dataset used during the training phase. It is also a constant calculated and stored during training. Let X be a matrix new The j-th column is a vector containing the original values of the j-th feature variable at all time points within the prediction period. Incorrectly using the data from the prediction period itself to calculate the new scaling parameter will lead to severe distortion of the prediction results.
[0072] Load the finally trained machine learning regression model from the storage medium. This model file contains all learned parameters (such as support vectors and coefficients of the support vector machine, the tree structure of XGBoost, and the weights of the neural network). The preprocessed feature matrix (size T x N) is passed as input to the model's prediction interface. The model then performs forward computation on each row of its input. For regression tasks, the model output is typically a continuous scalar value. Therefore, the model will output T predicted values for T input vectors, forming a prediction sequence of length T. ,in, Let be the predicted value of heavy metal content on day t.
[0073] In some embodiments of this application, uncertainty analysis of the predicted heavy metal content includes: Based on the feasible set of parameters of the watershed distributed hydrological model, multiple sets of different multi-dimensional hydrological and water quality parameters are generated.
[0074] Each set of parameters is input into an ensemble model consisting of multiple machine learning regression models constructed using the bootstrap sampling method, and corresponding predicted values are obtained to form a set of predicted values.
[0075] Perform statistical analysis on the predicted value set, calculate and output the confidence interval of the predicted heavy metal content.
[0076] Specifically, to ensure a clear and complete reproduction of the entire process of uncertainty analysis for heavy metal content predictions, a systematic supplementary explanation of the above technical solution is provided. This analysis employs a two-stage Monte Carlo simulation framework, comprehensively quantifying the uncertainty of SWAT model parameters and the prediction uncertainty of the machine learning model, and ultimately outputting the confidence interval of the predicted values. During the SWAT model calibration process, a large number of parameter combinations are generated, each ensuring that the simulated runoff falls within the uncertainty range of the measured data. These parameter combinations constitute the "feasible set of parameters," denoted as... Each It is a vector containing the specific values of all the parameters being calibrated. During the initial training phase, M samples are randomly drawn with replacement from the training set containing M samples, forming a bootstrap sampling set. A base learner is trained using this set. This process is repeated multiple times to obtain multiple base learners, which together constitute an ensemble model, denoted as . Because the training data for each base learner is slightly different, their prediction results will vary. This difference quantifies the uncertainty inherent in the machine learning model itself due to the randomness of the training data and the randomness of model optimization, from the feasible set of parameters. In this model, L sets of parameters are randomly drawn with replacement (usually L is a large number, such as 500 or 1000, where L=K). These L sets of parameters represent a random sample of the uncertainty of the SWAT model parameters, denoted as . Then, SWAT is run to generate features, for each sampled parameter. Perform the following operations: using Replace the default parameter values of the SWAT model, use the same meteorological and pollution source driving data as the forecast period, run the SWAT model to simulate the entire target forecast period, and extract the required N hydrological and water quality feature variables from the model output. The final result is a T-row × N-column feature matrix. Finally, L characteristic matrices are obtained. For each day t in the prediction period, there are L possible distinct feature vectors. Then, a double-loop calculation is performed: traversing L groups of SWAT parameters (l=1 to L), for the feature matrix generated by the l-th group of parameters... Input it into a machine learning ensemble model Each base learner in In, each base learner For each row of the input matrix (each day), a predicted value is output. Therefore, for the l-th set of SWAT parameters and the b-th base learner, a predicted sequence of length T is obtained. After iterating through all L×B combinations, for each day t of the prediction period, we obtain a huge set containing L×B predicted values: This set It integrates all possible effects of SWAT parameter uncertainty and machine learning model uncertainty.
[0077] The set of predicted values for each day t Perform statistical analysis: [The sentence is incomplete and requires more context to be translated accurately.] All L×B predicted values are arranged in ascending order, and a pre-set confidence level, such as 95%, is then used. The lower limit of the confidence interval is the 2.5% quantile, and the upper limit is the 97.5% quantile. Given sorted data, the index of the p-th percentile (e.g., p=2.5) is: ,like If it is an integer, then the value at that position is taken directly. If the integer is not an integer (e.g., kd), then take the linear interpolation of the k-th and (k+1)-th bit values: ,in For the k-th value after sorting, calculate the upper and lower bounds of the confidence interval, and take... The median of the median is used as the final robust point prediction.
[0078] In some embodiments of this application, determining the heavy metal load based on heavy metal content and multi-dimensional physical characteristic parameters includes: Obtain hydrological flow data generated by a distributed hydrological model for the same target time period or location.
[0079] By coupling hydrological flow data with predicted heavy metal content, the heavy metal load for the target time period or location can be obtained.
[0080] In some embodiments of this application, when issuing an early warning based on heavy metal load and heavy metal content, the following methods are included: Compare the predicted heavy metal content with a preset concentration threshold, and / or compare the predicted heavy metal load with a preset load threshold.
[0081] Based on the comparison results, corresponding early warning information is generated and sent.
[0082] Specifically, when the predicted value of heavy metal content exceeds the first concentration threshold and / or the predicted value of heavy metal load exceeds the first load threshold, a first-level early warning message is generated and sent.
[0083] When the predicted heavy metal content exceeds the second concentration threshold but does not exceed the first concentration threshold, and / or the predicted heavy metal load exceeds the second load threshold but does not exceed the first load threshold, a level two warning message is generated and sent.
[0084] Among them, the first concentration threshold is higher than the second concentration threshold, and the first load threshold is higher than the second load threshold.
[0085] Specifically, the core of load calculation is "flow-weighted calculation," which means that the mass of pollutants passing through a cross-section per unit time is equal to the product of the flow rate and the concentration. ,in, Daily heavy metal load on day t The simulated daily average flow rate on day t. Predicted heavy metal concentration on day t. Unit conversion factor.
[0086] In summary, the beneficial effects of this invention are as follows: (1) This invention improves the physical interpretability and mechanistic consistency of the prediction model. Traditional pure data-driven machine learning models are often regarded as "black boxes," and their prediction results lack clear physical process support, making it difficult for environmental managers to understand and trust them. This invention introduces a rigorously calibrated SWAT model as a physical feature generator, using a series of parameters with clear hydrological, hydrodynamic, and environmental geochemical significance, such as surface runoff, groundwater runoff, soil moisture content, and sediment load, as inputs to the machine learning model. This makes the final prediction model no longer simply fit the data curve, but is based on the indirect characterization of the basic physicochemical process of "how water flow carries sediment and how sediment adsorbs and transports heavy metals." The prediction results thus have a solid physical mechanism background, the credibility and acceptability of the model are fundamentally enhanced, and decision-makers can understand the dominant environmental factors behind the prediction.
[0087] (2) This invention reduces the over-reliance on high-frequency, high-density field monitoring data, thereby reducing the economic cost and manpower investment of long-term monitoring. Precise analysis of heavy metals is costly and complex, resulting in sparse and spatially limited monitoring data for most watersheds. This invention utilizes the SWAT model to generate continuous, spatiotemporally complete "surrogate features." While these features do not directly measure heavy metals, they precisely characterize the key environmental conditions driving their migration and transformation. The machine learning model only needs to use limited measured concentration data collected at key spatiotemporal nodes as "labels" to learn the complex relationship between these continuous physical features and the target pollutant. Without the need for costly encrypted monitoring, state estimation can be achieved for periods and locations without monitoring, allowing limited monitoring resources to be used for the most critical model validation and calibration stages, thus optimizing cost-effectiveness.
[0088] (3) This invention comprehensively improves the accuracy, robustness, and spatiotemporal extrapolation capability of prediction. Machine learning models trained solely on a small amount of measured data are prone to overfitting, and their prediction performance drops sharply when encountering hydrological and meteorological conditions that have not been experienced. The multidimensional features provided by the SWAT model constitute a comprehensive and stable environmental state description system, which can more fully reflect the complex factors affecting the fate of heavy metals (such as hydrodynamic intensity, sediment transport flux, soil moisture conditions, etc.). This enables the machine learning model to learn more essential and universal laws, rather than just memorizing limited samples. Therefore, the model performs more stably in watersheds with scarce data, and also shows stronger adaptability and reliability in predictions under extreme rainfall events or long-term climate change scenarios, reducing the risk of prediction failure due to insufficient model extrapolation capability.
[0089] (4) This invention enables dynamic risk early warning, whereas traditional monitoring can only report pollution conditions after they have occurred. By coupling prediction and early warning mechanisms, this invention can proactively calculate heavy metal flux loads for future periods and, combined with preset environmental quality standards and risk assessment thresholds, achieve tiered early warning. This directly transforms technical output into management action instructions, thereby shifting environmental management from passive response to proactive prevention and control, and improving the timeliness and effectiveness of risk management.
[0090] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program goods. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program goods embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0091] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program goods according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0092] These computer program instructions may also be stored in a computer-readable storage device that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage device produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0093] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A watershed heavy metal content and load prediction method based on SWAT coupled with machine learning, characterized in that, include: Collect measured data from the target watershed, preprocess it, and construct a basic database; Based on the aforementioned basic database, a watershed distributed hydrological model is constructed and calibrated, wherein the watershed distributed hydrological model is run to generate time series of multi-dimensional physical characteristic parameters, including at least hydrological variables and water quality-related variables. A machine learning regression model is trained based on the time series of the multi-dimensional physical feature parameters and the basic database. The new feature data generated by the watershed distributed hydrological model is input into the machine learning regression model to predict the heavy metal content of the target time period or location, and the heavy metal load is determined based on the heavy metal content and the multi-dimensional physical feature parameters. Early warning is issued based on the heavy metal load and the heavy metal content.
2. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 1, characterized in that, The watershed distributed hydrological model is a SWAT model; the multi-dimensional physical characteristic parameters include hydrological variables and water quality-related variables, wherein the hydrological variables are selected from at least one of surface runoff, groundwater runoff, lateral flow, actual evapotranspiration, and soil moisture content, and the water quality-related variables are selected from at least one of sediment load, nitrate concentration, and organic nitrogen load.
3. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 2, characterized in that, When collecting measured data from the target watershed, preprocessing it, and constructing the basic database, the following steps are included: The measured data includes at least spatial data of the target watershed, meteorological and hydrological data, pollution source data, and measured data of heavy metal concentrations. The preprocessing includes at least unifying the spatial coordinate system of all data to the same reference and unifying the time reference of all time series data with the simulation step size; Based on the preprocessed data, a spatiotemporally consistent basic database is constructed. In constructing the basic database, the time series of multi-dimensional hydrological and water quality parameters generated by the watershed distributed hydrological model are aligned with the sampling time points of the measured heavy metal concentration data to form a feature matrix and label vector for training the machine learning regression model.
4. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 3, characterized in that, When constructing and calibrating a distributed hydrological model for a watershed based on the aforementioned basic database, the following steps are included: Based on the spatial data, the watershed space is discretized, and model calculation units are defined; Load the meteorological and hydrological data and pollution source data, and configure the model driving conditions; Using the measured river flow data from the meteorological and hydrological data, the model is subjected to parameter sensitivity analysis, automatic optimization calibration and performance verification until the evaluation index of simulated flow and measured flow reaches the preset accuracy standard. After the model is calibrated and validated, hydrological variables and water quality-related variables related to the migration and transformation of heavy metals are extracted to generate the time series of the multidimensional hydrological and water quality parameters.
5. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 4, characterized in that, When training a machine learning regression model based on the time series of the multi-dimensional physical feature parameters and the basic database, the following steps are included: The machine learning regression model is an algorithm suitable for small sample regression, selected from at least one of support vector regression, long short-term memory network, Transformer model, XGBoost model and combination thereof; During or after model training, the SHAP value analysis method is used to evaluate the contribution of each input feature to the prediction results.
6. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 5, characterized in that, When inputting the new feature data generated by the watershed distributed hydrological model into the machine learning regression model to predict the heavy metal content of a target time period or location, the following methods are included: The multi-dimensional hydrological and water quality parameters generated by the watershed distributed hydrological model for the target time period or location are used as input feature data. The input feature data is input into the trained machine learning regression model for calculation; Obtain the predicted heavy metal content output by the machine learning regression model, corresponding to the target time period or location, and perform uncertainty analysis on the predicted heavy metal content.
7. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 6, characterized in that, When performing uncertainty analysis on the predicted heavy metal content, the following steps are included: Based on the feasible set of parameters of the watershed distributed hydrological model, multiple sets of different multi-dimensional hydrological and water quality parameters are generated. Each set of parameters is input into an ensemble model composed of multiple machine learning regression models constructed by the bootstrap sampling method, and corresponding predicted values are obtained to form a set of predicted values. Statistical analysis is performed on the predicted value set to calculate and output the confidence interval of the predicted heavy metal content.
8. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 7, characterized in that, Determining the heavy metal load based on the heavy metal content and the multi-dimensional physical characteristic parameters includes: Obtain hydrological flow data simulated by the watershed distributed hydrological model for the same target time period or location; The hydrological flow data is coupled with the predicted heavy metal content to calculate the heavy metal load at the target time period or location.
9. The method for predicting watershed heavy metal content and load based on SWAT and machine learning coupling as described in claim 8, characterized in that, When issuing an early warning based on the heavy metal load and the heavy metal content, it includes: The predicted heavy metal content is compared with a preset concentration threshold, and / or the predicted heavy metal load is compared with a preset load threshold; Based on the comparison results, generate and send corresponding early warning information; Specifically, when the predicted value of the heavy metal content exceeds the first concentration threshold and / or the predicted value of the heavy metal load exceeds the first load threshold, a first-level early warning message is generated and sent. When the predicted value of the heavy metal content exceeds the second concentration threshold but does not exceed the first concentration threshold, and / or the predicted value of the heavy metal load exceeds the second load threshold but does not exceed the first load threshold, a secondary warning message is generated and sent. Wherein, the first concentration threshold is higher than the second concentration threshold, and the first load threshold is higher than the second load threshold.