A dynamic calculation method and system for watershed non-point source pollution load based on distributed process model

By combining distributed process models with machine learning methods, the problem of spatiotemporal quantitative measurement of non-point source pollution in a river basin under a dynamically changing environment was solved, and the accurate measurement of the non-point source pollution load in the river basin was achieved, supporting the refined management of water resources and water environment in the river basin.

CN120297771BActive Publication Date: 2025-09-09NANJING HYDRAULIC RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510758246.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve spatiotemporal quantitative measurement of non-point source pollution in a river basin under a dynamically changing environment, especially in the context of accelerated urbanization. Existing methods lack an explanation of the process mechanism, and the driving factors of socioeconomic development are separated from land use changes and the dynamic changes of terrestrial pollution sources, resulting in insufficient measurement accuracy.

Method used

A method based on a distributed process model is used, combined with meteorological, hydrological, soil, topographic and other data, to construct a runoff-sediment-nitrogen-phosphorus model. Combined with the driving factors of socioeconomic development and land use changes, a temporally and spatially consistent database is formed through geospatial interpolation, spatial resampling and frequency matching. Combined with machine learning methods, the dynamic changes of terrestrial pollution sources are calculated to achieve temporal and spatial quantitative measurement of non-point source pollution loads in the basin.

Benefits of technology

It realizes the spatiotemporal quantitative measurement of the non-point source pollution load in the basin under a dynamically changing environment, avoids the separation of the driving factors of social and economic development and land use changes, improves the measurement accuracy, and provides support for the refined management of water resources and water environment in the basin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297771B_ABST
    Figure CN120297771B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for dynamically calculating the non-point source pollution load in a watershed based on a distributed process model. First, meteorological, hydrological, water quality, soil and topography, land use, vegetation cover, atmospheric deposition and surface pollution source data of the target watershed are collected and a basic database of the target watershed is constructed. Then, geospatial interpolation, spatial resampling and frequency matching are used to integrate and process different types of data in the basic database of the target watershed. The present invention realizes the function of coupling the physical processes of runoff-sediment-nitrogen-phosphorus on the basis of a distributed hydrological model and combining the driving factors of socio-economic development, land use changes and the dynamic changes of terrestrial pollution sources to perform spatiotemporal quantitative calculation of the non-point source pollution load in the watershed under the changing scenarios. This method can not only better capture the dynamic changes of the predicted input data, but also avoid the separation of the driving factors of socio-economic development from the dynamic changes of land use changes and terrestrial pollution sources, and is suitable for wide promotion and use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hydrological pollution load calculation, and in particular to a method and system for dynamically calculating watershed non-point source pollution load based on a distributed process model. Background Art

[0002] Dynamic quantification of non-point source pollution loads in river basins is a core issue in water environment governance. Its complexity stems from the multi-scale coupling of hydrological processes, human activities, and the ecological environment. In river basins where control of industrial and domestic pollution sources has been relatively stable, atmospheric deposition brought on by climate change and changes in the scale and composition of agricultural production in the basin have become the primary sources of non-point source pollution. However, due to the significant spatiotemporal heterogeneity of industrial structure and pollution sources, as well as the complex dynamic relationships between them, methods for quantifying their contribution to non-point source pollution are still immature. This further increases uncertainty in the impact of non-point source pollution on river water quality, making it a challenge to manage water resources and the environment in the river basin.

[0003] At present, most of the methods for calculating non-point source pollution in watersheds are based on field monitoring, such as SCS-CN, USLE and RULSE formula methods. These methods can only quantify the different sources of non-point sources in the historical period with existing basic data, and it is difficult to achieve quantitative spatiotemporal measurement of non-point source pollution in watersheds under dynamic change scenarios. Non-point source pollution measurement methods based on machine learning generally only measure non-point source pollution in watersheds based on hydrological and meteorological data, pollutant time series data and spatial characteristics of remote sensing images, resulting in a lack of process mechanism explanation and a great influence of uncertain factors under a changing environment. At the same time, there is also a technical bottleneck of separating the driving factors of socio-economic development from the dynamic changes of land use changes and terrestrial pollution sources. In the context of accelerated urbanization, the spatiotemporal dynamic measurement accuracy of non-point source pollution is also insufficient. Therefore, it is necessary to design a dynamic measurement method and system for non-point source pollution load in watersheds based on a distributed process model. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings of the existing technology and to better and effectively solve the current basin non-point source pollution measurement methods, which are mostly based on field monitoring, such as SCS-CN, USLE and RULSE formula methods. These methods can only quantify the different sources of non-point sources in the historical period of existing basic data, and it is difficult to achieve spatiotemporal quantitative measurement of basin non-point source pollution under dynamic change scenarios. The non-point source pollution measurement methods based on machine learning are generally based on hydrological and meteorological, pollutant time series data and remote sensing image spatial features to measure basin non-point source pollution, resulting in a lack of process mechanism explanation, and the influence of uncertain factors is great under the changing environment. At the same time, there are also socio-economic development driving factors and land use. The paper provides a method and system for dynamic measurement of non-point source pollution load in watershed based on distributed process model, which realizes the function of quantitative measurement of non-point source pollution load in watershed under changing scenarios in terms of spatiotemporal information, coupling runoff-sediment-nitrogen and phosphorus physical processes on the basis of distributed hydrological model, and combining socio-economic development driving factors, land use change and dynamic change of terrestrial pollution sources. It can not only better capture the dynamic change of prediction input data, but also avoid the separation of socio-economic development driving factors from land use change and dynamic change of terrestrial pollution sources.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A method for dynamic calculation of watershed non-point source pollution load based on a distributed process model includes the following steps:

[0007] Step A: Collect meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, atmospheric deposition, and surface pollution source data for the target basin and construct a basic database for the target basin. Then, geospatial interpolation, spatial resampling, and frequency matching are used to integrate and process different types of data in the basic database for the target basin, thereby forming a processed basic database for the target basin with consistent temporal and spatial scales.

[0008] Step B: Construct a runoff-sediment-nitrogen-phosphorus model based on physical processes. Then, input the processed target watershed basic database into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff process, sediment scouring-migration-sedimentation process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model.

[0009] Step C: The future land use state transition matrix in the LUH2 dataset is used as the land use change constraint condition, and then the CA-Markov model is used to calculate the land use change result according to the land use change constraint condition;

[0010] Step D: Analyze the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding. Then, based on the quantitative analysis results of fertilizer and livestock and poultry breeding, integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models to measure the dynamic changes in land-based pollution sources of fertilizer and livestock and poultry breeding pollution.

[0011] In step E, the processed target basin basic database, land use change results, and land pollution source dynamic change results are input into the distributed process model to obtain the spatiotemporal quantitative results of the basin non-point source pollution load under the change scenario, thereby completing the dynamic measurement of the basin non-point source pollution load.

[0012] The aforementioned method for dynamic measurement of non-point source pollution load in a watershed based on a distributed process model includes step A, which collects meteorological, hydrological, water quality, soil and topography, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target watershed and constructs a basic database of the target watershed. Then, geospatial interpolation, spatial resampling, and frequency matching are used to integrate and process different types of data in the basic database of the target watershed, thereby forming a processed basic database of the target watershed with consistent temporal and spatial scales. The specific steps are as follows:

[0013] Step A1: collecting target basin meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, atmospheric deposition, and surface pollution source data and constructing a target basin basic database, wherein the target basin meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, and atmospheric deposition data are remote sensing grid data, and the surface pollution source data are yearbook statistical data. The target basin basic database also includes future climate scenario meteorological data and future atmospheric deposition scenario data, wherein the future climate scenario meteorological data are spatial grid data output by GCM, and the future atmospheric deposition scenario data are atmospheric deposition spatial grid data provided by the MRI-ESM2-0 model.

[0014] Step A2: Use geospatial interpolation, spatial resampling, and frequency matching to integrate different types of data in the target basin basic database, thereby forming a processed target basin basic database with consistent temporal and spatial scales. The specific steps are as follows:

[0015] Step A21, using geospatial interpolation to integrate different types of data in the target basin basic database. Specifically, the geospatial interpolation is to use the inverse distance weighted interpolation method (IDW) to interpolate the site data into space according to the spatial resolution of the remote sensing grid data and then enter it into the target basin basic database;

[0016] Step A22: spatial resampling is used to integrate different types of data in the target watershed basic database. Specifically, spatial resampling is to unify the spatial resolution of soil, topography, land use, vegetation cover, and atmospheric deposition grid data using the Majority type.

[0017] Step A23: Frequency matching is used to integrate and process different types of data in the target basin basic database. Specifically, frequency matching is applied to the meteorological data output by the GCM and the cumulative frequency curve of the measured climate data is used to correct the deviation of the cumulative frequency curve of the GCM simulation value, so that the cumulative frequency curve of the historical measured data remains consistent. The corrected data can then be entered into the basic database as the input data for future scenarios.

[0018] The aforementioned method for dynamic calculation of non-point source pollution load in a watershed based on a distributed process model, step B, constructs a runoff-sediment-nitrogen-phosphorus model based on physical processes, then inputs the processed target watershed basic database into the runoff-sediment-nitrogen-phosphorus model and obtains the river grid runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model. The specific steps are as follows:

[0019] Step B1: Construct a runoff-sediment-nitrogen-phosphorus model based on physical processes, wherein the runoff-sediment-nitrogen-phosphorus model adopts the GBNP model, which includes a hydrological module, a soil erosion module, and a nitrogen-phosphorus load module. The specific steps are as follows:

[0020] Step B11: Construct a hydrological module of the GBNP model. The hydrological module is based on the distributed hydrological model (GBHM), which uses the digital elevation model (DEM) to extract the watershed system and divide the hillside-river grid cells as the basic model units. Then, the hydrological processes of precipitation interception, evapotranspiration, soil water movement, and phreatic outflow are calculated in the hillside cells. The specific steps are as follows:

[0021] Step B111: Calculate the hydrological process of precipitation interception. Specifically, assuming that the environmental surface is covered with vegetation, the hydrological process of precipitation interception is as shown in formula (1):

[0022] (1);

[0023] in, for Moment The actual transpiration rate of soil water from the roots to the leaves of plants, is the vegetation coverage rate, is the crop coefficient, is the potential evaporation, is the plant root depth distribution function, is the soil moisture function, for Leaf area index at time It is the maximum value of leaf area index of vegetation in a year;

[0024] Step B112: Calculate the evapotranspiration hydrological process, specifically including the situation where the environmental surface has no vegetation cover and the situation where the environmental surface water does not meet the evaporation capacity. The specific steps are as follows:

[0025] Step B1121: Assuming that the environment surface is not covered by vegetation, the evapotranspiration hydrological process is as shown in formula (2):

[0026] (2);

[0027] in, for The actual surface evaporation rate at the moment, for The depth of surface water at that moment, is the time step;

[0028] Step B1122: Assuming that the surface water in the environment does not meet the evaporation capacity, the evaporation hydrological process is as shown in formula (3):

[0029] (3);

[0030] in, for The actual evaporation rate of the soil surface at the moment, is the soil moisture function;

[0031] Step B113: Calculate the soil water movement hydrological process, specifically including soil unsaturated hydraulic conductivity, unsaturated zone soil water movement process and soil moisture characteristic curve. The specific steps are as follows:

[0032] Step B1131: Calculate the unsaturated hydraulic conductivity of the soil, as shown in formula (4):

[0033] (4);

[0034] in, is the unsaturated hydraulic conductivity of the soil, is the saturated hydraulic conductivity of the soil, is the moisture content, are the fitting parameters of soil moisture characteristic curve;

[0035] Step B1132: Calculate the soil water movement process in the unsaturated zone. Specifically, the one-dimensional Richard equation is used to calculate the soil water movement in the unsaturated zone. The soil water movement process in the unsaturated zone is shown in formula (5).

[0036] (5);

[0037] in, is the soil depth, for Soil depth at any moment The moisture content at is the soil water flux, for Soil evaporation rate at any moment, is the soil moisture function;

[0038] Step B1132: Calculate the soil moisture characteristic curve. Specifically, the soil moisture characteristic curve represents the relationship between soil moisture content and soil suction. The soil moisture characteristic curve is shown in formula (6).

[0039] (6);

[0040] in, is the residual moisture content of the soil, is the saturated moisture content of the soil, and All are fitting parameters of soil moisture characteristic curve;

[0041] Step B114: Calculate the hydrological process of phreatic outflow, including the water exchange process between the hillside phreatic layer and the river channel and the river channel confluence process. The specific steps are as follows:

[0042] Step B1141: Calculate the water exchange process between the hillside aquifer and the river channel, as shown in formula (7):

[0043] (7);

[0044] in, is the rate of change of groundwater storage in the saturated aquifer over time, is the recharge rate between the saturated aquifer and the upper unsaturated aquifer, is the downward leakage rate of the saturated aquifer, is the exchange flow rate between groundwater and river channel per unit width, is the hillside unit area, is the saturated hydraulic conductivity of the aquifer, and is the groundwater level of the aquifer before and after water exchange, and The water level of the river before and after the water exchange, is the slope unit length;

[0045] Step B1142: Calculate the river confluence process, as shown in formula (8):

[0046] (8);

[0047] in, The total confluence of rivers, It is the lateral inflow of the river. is the cross-sectional area of ​​the river, is the distance along the river channel, is the river slope, is the river channel roughness, is the wet perimeter length;

[0048] Step B12: Construct the soil erosion module and nitrogen and phosphorus loading module of the GBNP model, specifically including the nutrient salt slope process, transformation process and river process. The specific steps are as follows:

[0049] Step B121: Calculate the nutrient salt slope process, where the nutrient salt slope process is specifically an accumulation-washing process, as shown in formula (9):

[0050] , (9);

[0051] in, is the dust accumulation amount, is the proportion coefficient of pollutants in dust, is the washout coefficient, is the runoff intensity, is the maximum cumulative amount, The number of days of dust accumulation since the last rainfall, The number of days required for dust accumulation to reach half of the maximum accumulation;

[0052] Step B122: Calculate the nutrient conversion process, where the nutrient conversion process is specifically a convection-diffusion process, and the convection-diffusion process includes nutrient adsorption, desorption, volatilization and conversion processes, and the compound form conversion process is described by a first-order kinetic reaction equation, as shown in formula (10).

[0053] ;

[0054] ;

[0055] ;

[0056] (10);

[0057] in, is the adsorption term concentration, is the water migration velocity, is the hydrodynamic dispersion coefficient, is the solute source and sink term, is the soil surface nitrogen content equivalent to the amount of fertilizer applied, is the inorganic phosphorus content, is the amount of transformation in different forms, is the amount absorbed by plants, is the loss caused by runoff, is the conversion coefficient, is the nitrogen and phosphorus absorption ratio of plants at different growth stages, The ratio of the cumulative growing days on that day to the growing period;

[0058] Step B123, calculate the nutrient salt river process, where the nutrient salt river process specifically includes the convection diffusion process, the slope inflow process, the morphological transformation process, the sediment adsorption process and the sediment release process, as shown in formula (11).

[0059] ;

[0060] (11);

[0061] in, is the cross-sectional area of ​​the river flow, is the concentration of nutrient solutes in the river channel, is the cross-sectional flow rate, is the diffusion coefficient along the river direction, is the source-sink term, is the river sediment concentration, is the flow that flows laterally into the river channel, is the concentration of pollutants entering laterally, is the concentration of lateral inflow sediment, is the sediment recovery saturation coefficient, is the sediment settling rate, is the river width, The maximum sediment-carrying capacity of a river. To flush the concentration of adsorbed pollutants in silt and sand;

[0062] Step B2: Input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model and obtain the river grid runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transfer-diffusion-transformation process, thereby obtaining a distributed process model. Specifically, the grid meteorological data in the processed target basin basic database is input into the runoff-sediment-nitrogen-phosphorus model as a driving force, the point source pollution data is input into the nearest river grid based on the latitude and longitude information of the sewage outlet, the domestic pollution source is input into the residential land grid, the solid waste and livestock and poultry breeding pollution source is input into the paddy field and dry land grid, the aquaculture pollution source is input into the pond water grid and the atmospheric deposition data is input into the corresponding spatial grid. The distributed process model is calibrated according to the runoff, sediment and nutrient concentration in turn and the Nash efficiency coefficient NSE, the deviation percentage PBIAS and the ratio coefficient RSR of the root mean square error to the observation standard deviation are used as model accuracy measurement indicators.

[0063] In the aforementioned method for dynamic estimation of watershed non-point source pollution load based on a distributed process model, step C uses the future land use state transition matrix in the LUH2 dataset as a land use change constraint, and then uses the CA-Markov model to estimate the land use change results based on the land use change constraint. The specific steps are as follows:

[0064] Step C1: Use the future land use state transition matrix in the LUH2 dataset as a constraint on land use change, where the future land use state transition matrix includes land use maps and land use type transition maps under different future scenarios;

[0065] Step C2, the CA-Markov model is used to calculate the land use change results according to the land use change constraints. Specifically, the Markov model is used to generate a land use type transfer probability map and the area transfer matrix relative to the base period is used as a constraint to modify the land use type transfer probability map. The CA model is then used to iterate so that the constraints and the area transfer matrix are satisfied and the land use change is calculated.

[0066] The aforementioned method for dynamically calculating the non-point source pollution load in a watershed based on a distributed process model, step D, analyzes the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding, and then integrates and cross-validates the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model based on the quantitative analysis results of the amount of fertilizer and livestock and poultry breeding, so as to calculate the dynamic changes in the land-based pollution sources of fertilizer and livestock and poultry breeding pollution. The specific steps are as follows:

[0067] Step D1, analyzing the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer applied and the amount of livestock and poultry raised, wherein the climate variables include precipitation and temperature, and the socioeconomic variables include agricultural population variables, gross domestic product (GDP), and primary industry output value;

[0068] Step D2, based on the quantitative analysis results of fertilization amount and livestock and poultry breeding amount, integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model, so as to measure the dynamic changes of terrestrial pollution sources of fertilization and livestock and poultry breeding pollution, the linear regression model includes linear regression, interaction effect linear regression, robust linear regression, stepwise linear regression, the regression tree model includes fine tree, medium tree and coarse tree, the support vector machine model includes linear support vector machine, quadratic support vector machine, cubic support vector machine, fine Gaussian support vector machine, medium Gaussian support vector machine and coarse Gaussian support vector machine, the Gaussian process regression model includes square exponential Gaussian process regression model, exponential Gaussian process regression model and rational quadratic Gaussian process regression model, the tree integration model includes boosted tree integration and bagged tree integration, the specific steps are as follows,

[0069] Step D21, integrating the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model, wherein the integration is specifically based on the ensemble learning method of the Boosting framework and optimizes the function space by the gradient descent method to obtain the total ensemble model, and the goal of the total ensemble model is to minimize the loss function , the total integrated model output is , the specific steps are as follows,

[0070] Step D211, the initial value is the target mean ;

[0071] Step D212, calculate the pseudo residual, as shown in formula (12),

[0072] (12);

[0073] in, is the pseudo residual, Output results for the current model;

[0074] Step D213, fitting the residuals, using regression tree Fit the pseudo residuals and then select the splitting features and thresholds by minimizing the squared error;

[0075] Step D214, multiply the prediction result of the new tree by the learning rate And accumulated to the current model, as shown in formula (13),

[0076] (13);

[0077] Step D215, set the termination condition, if the residual converges to the set threshold, stop sampling,

[0078] Step D22, cross-validating the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model, specifically adopting a systematic data partitioning strategy to balance computational efficiency and evaluation reliability. The specific steps are as follows:

[0079] Step D221, data partitioning, randomly divide the original data set into five equal and mutually exclusive subsets for stratified sampling and ensure that the category distribution in each subset is consistent with the original data;

[0080] Step D222, iterative training and validation, the first time using the first subset as the validation set and the second to fifth subsets as the training set, the second time using the second subset as the validation set and the remaining subsets as the training set, and so on to the fifth time, each time generating an independent model;

[0081] Step D223, performance evaluation, uses the root mean square error (RMSE) as the evaluation indicator of the validation set and takes the average and standard deviation of the five results to evaluate the stability of the overall integrated model.

[0082] A system for dynamically calculating watershed non-point source pollution loads based on a distributed process model includes a data acquisition and processing module, a distributed process model construction module, a land use change calculation module, a land pollution source dynamic change calculation module, and a watershed non-point source pollution load dynamic calculation module. The data acquisition and processing module is used to collect meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, atmospheric deposition, and surface pollution source data for a target watershed and construct a target watershed basic database. The system then integrates and processes different types of data in the target watershed basic database using geospatial interpolation, spatial resampling, and frequency matching, thereby forming a processed target watershed basic database with consistent temporal and spatial scales.

[0083] The distributed process model construction module is used to construct a runoff-sediment-nitrogen-phosphorus model based on physical processes, and then input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model;

[0084] The land use change calculation module is used to use the future land use state transfer matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition;

[0085] The land pollution source dynamic change calculation module is used to analyze the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding, and then integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model based on the quantitative analysis results of the amount of fertilizer and livestock and poultry breeding, so as to calculate the dynamic change results of the land pollution sources of fertilization and livestock and poultry breeding pollution;

[0086] The dynamic measurement module of the non-point source pollution load in the watershed is used to input the processed target watershed basic database, land use change results and dynamic change results of terrestrial pollution sources into the distributed process model and obtain the spatiotemporal quantitative results of the non-point source pollution load in the watershed under the change scenario, thereby completing the dynamic measurement operation of the non-point source pollution load in the watershed.

[0087] The beneficial effects of the present invention are as follows: a method and system for dynamically calculating the non-point source pollution load in a watershed based on a distributed process model of the present invention first collects meteorological, hydrological, water quality, soil and topography, land use, vegetation cover, atmospheric deposition and surface pollution source data of the target watershed and constructs a basic database of the target watershed, then integrates and processes different types of data in the basic database of the target watershed using geographic space interpolation, spatial resampling and frequency matching to form a processed basic database of the target watershed with consistent time and space scales, then constructs a runoff-sediment-nitrogen-phosphorus model based on physical processes, and then inputs the processed basic database of the target watershed into the runoff-sediment-nitrogen-phosphorus model to obtain the river flow. The runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transport-diffusion-transformation process of the road grid are used to obtain a distributed process model. Then, the future land use state transfer matrix in the LUH2 dataset is used as the land use change constraint condition. Then, the CA-Markov model is used to calculate the land use change results based on the land use change constraint condition. Then, the quantitative relationship between climate variables and socioeconomic variables and fertilizer application and livestock and poultry breeding capacity is analyzed. Then, based on the quantitative analysis results of fertilizer application and livestock and poultry breeding capacity, linear regression models, regression tree models, support vector machine models, Gaussian process regression models and tree integration models are integrated and cross-validated to calculate the relationship between fertilizer application and livestock and poultry breeding capacity. The dynamic change results of terrestrial pollution sources of livestock and poultry breeding pollution are obtained. Finally, the basic database of the target watershed after processing, the results of land use change and the dynamic change results of terrestrial pollution sources are input into the distributed process model to obtain the spatiotemporal quantitative results of the watershed non-point source pollution load under the change scenario, thereby completing the dynamic measurement of the watershed non-point source pollution load; effectively realizing the dynamic measurement method and system of the watershed non-point source pollution load with the function of spatiotemporal quantitative measurement of the watershed non-point source pollution load under the change scenario by coupling the runoff-sediment-nitrogen and phosphorus physical processes on the basis of the distributed hydrological model and combining the driving factors of social and economic development, land use change and the dynamic changes of terrestrial pollution sources, and through the spatial Interpolation and frequency matching obtain an input database with unified temporal and spatial resolution. By coupling the runoff-sediment-nitrogen and phosphorus physical processes based on a distributed hydrological model, the interface mismatch of the existing hydrological model nested in the water environment model and the model connection problems caused by the independence of the input and output processes are avoided. At the same time, the ensemble machine learning method can not only better capture the dynamic changes of the predicted input data, but also avoid the separation of the driving factors of socio-economic development from the dynamic changes of land use changes and terrestrial pollution sources. This provides a new method for accurately and quantitatively measuring the spatiotemporal changes of non-point source pollution loads in the basin under changing scenarios, and is also of great significance for guiding the refined management of water resources and water environment in the basin. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 This is an overall flow chart of a method for dynamically calculating watershed non-point source pollution loads based on a distributed process model according to the present invention;

[0089] Figure 2 It is a frequency curve matching schematic diagram of the present invention;

[0090] Figure 3 It is a schematic diagram of the total nitrogen and total phosphorus calculation results of the watershed under the changing scenarios in the example of the present invention. DETAILED DESCRIPTION

[0091] The present invention will be further described below with reference to the accompanying drawings.

[0092] like Figure 1 As shown, the present invention provides a method for dynamically calculating the non-point source pollution load in a watershed based on a distributed process model, comprising the following steps:

[0093] Step A: Collect meteorological, hydrological, water quality, soil and topography, land use, vegetation cover, atmospheric deposition and surface pollution source data of the target basin and build a basic database of the target basin. Then, use geospatial interpolation, spatial resampling and frequency matching to integrate different types of data in the basic database of the target basin, so as to form a processed basic database of the target basin with consistent temporal and spatial scales. The specific steps are as follows:

[0094] Step A1: collecting target basin meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, atmospheric deposition, and surface pollution source data and constructing a target basin basic database, wherein the target basin meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, and atmospheric deposition data are remote sensing grid data, and the surface pollution source data are yearbook statistical data. The target basin basic database also includes future climate scenario meteorological data and future atmospheric deposition scenario data, wherein the future climate scenario meteorological data are spatial grid data output by GCM, and the future atmospheric deposition scenario data are atmospheric deposition spatial grid data provided by the MRI-ESM2-0 model.

[0095] Step A2: Use geospatial interpolation, spatial resampling, and frequency matching to integrate different types of data in the target basin basic database, thereby forming a processed target basin basic database with consistent temporal and spatial scales. The specific steps are as follows:

[0096] Step A21, using geospatial interpolation to integrate different types of data in the target basin basic database. Specifically, the geospatial interpolation is to use the inverse distance weighted interpolation method (IDW) to interpolate the site data into space according to the spatial resolution of the remote sensing grid data and then enter it into the target basin basic database;

[0097] Step A22: spatial resampling is used to integrate different types of data in the target watershed basic database. Specifically, spatial resampling is to unify the spatial resolution of soil, topography, land use, vegetation cover, and atmospheric deposition grid data using the Majority type.

[0098] like Figure 2 As shown in FIG2 , step A23 uses frequency matching to integrate different types of data in the target basin basic database. Specifically, frequency matching is applied to the meteorological data output by the GCM and uses the cumulative frequency curve of the measured climate data to correct the deviation of the cumulative frequency curve of the GCM simulation value, so that the cumulative frequency curve of the historical measured data remains consistent. The corrected data can then be entered into the basic database as the input data for future scenarios.

[0099] Step B: Construct a runoff-sediment-nitrogen-phosphorus model based on physical processes. Then, input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff process, sediment scouring-migration-sedimentation process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model. The specific steps are as follows:

[0100] Step B1: Construct a runoff-sediment-nitrogen-phosphorus model based on physical processes, wherein the runoff-sediment-nitrogen-phosphorus model adopts the GBNP model, which includes a hydrological module, a soil erosion module, and a nitrogen-phosphorus load module. The specific steps are as follows:

[0101] Step B11: Construct a hydrological module of the GBNP model. The hydrological module is based on the distributed hydrological model (GBHM), which uses the digital elevation model (DEM) to extract the watershed system and divide the hillside-river grid cells as the basic model units. Then, the hydrological processes of precipitation interception, evapotranspiration, soil water movement, and phreatic outflow are calculated in the hillside cells. The specific steps are as follows:

[0102] Step B111: Calculate the hydrological process of precipitation interception. Specifically, assuming that the environmental surface is covered with vegetation, the hydrological process of precipitation interception is as shown in formula (1):

[0103] (1);

[0104] in, for Moment The actual transpiration rate of soil water from the roots to the leaves of plants, is the vegetation coverage rate, is the crop coefficient, is the potential evaporation, is the plant root depth distribution function, is the soil moisture function, for Leaf area index at time It is the maximum value of leaf area index of vegetation in a year;

[0105] Step B112: Calculate the evapotranspiration hydrological process, specifically including the situation where the environmental surface has no vegetation cover and the situation where the environmental surface water does not meet the evaporation capacity. The specific steps are as follows:

[0106] Step B1121: Assuming that the environment surface is not covered by vegetation, the evapotranspiration hydrological process is as shown in formula (2):

[0107] (2);

[0108] in, for The actual surface evaporation rate at the moment, for The depth of surface water at that moment, is the time step;

[0109] Step B1122: Assuming that the surface water in the environment does not meet the evaporation capacity, the evaporation hydrological process is as shown in formula (3):

[0110] (3);

[0111] in, for The actual evaporation rate of the soil surface at the moment, is the soil moisture function;

[0112] Step B113: Calculate the soil water movement hydrological process, specifically including soil unsaturated hydraulic conductivity, unsaturated zone soil water movement process and soil moisture characteristic curve. The specific steps are as follows:

[0113] Step B1131: Calculate the unsaturated hydraulic conductivity of the soil, as shown in formula (4):

[0114] (4);

[0115] in, is the unsaturated hydraulic conductivity of the soil, is the saturated hydraulic conductivity of the soil, is the moisture content, are the fitting parameters of soil moisture characteristic curve;

[0116] Step B1132: Calculate the soil water movement process in the unsaturated zone. Specifically, the one-dimensional Richard equation is used to calculate the soil water movement in the unsaturated zone. The soil water movement process in the unsaturated zone is shown in formula (5).

[0117] (5);

[0118] in, is the soil depth, for Soil depth at any moment The moisture content at is the soil water flux, for Soil evaporation rate at any moment, is the soil moisture function;

[0119] Step B1132: Calculate the soil moisture characteristic curve. Specifically, the soil moisture characteristic curve represents the relationship between soil moisture content and soil suction. The soil moisture characteristic curve is shown in formula (6).

[0120] (6);

[0121] in, is the residual moisture content of the soil, is the saturated moisture content of the soil, and All are fitting parameters of soil moisture characteristic curve;

[0122] Step B114: Calculate the hydrological process of phreatic outflow, including the water exchange process between the hillside phreatic layer and the river channel and the river channel confluence process. The specific steps are as follows:

[0123] Step B1141: Calculate the water exchange process between the hillside aquifer and the river channel, as shown in formula (7):

[0124] (7);

[0125] in, is the rate of change of groundwater storage in the saturated aquifer over time, is the recharge rate between the saturated aquifer and the upper unsaturated aquifer, is the downward leakage rate of the saturated aquifer, is the exchange flow rate between groundwater and river channel per unit width, is the hillside unit area, is the saturated hydraulic conductivity of the aquifer, and is the groundwater level of the aquifer before and after water exchange, and The water level of the river before and after the water exchange, is the slope unit length;

[0126] Step B1142: Calculate the river confluence process, as shown in formula (8):

[0127] (8);

[0128] in, The total confluence of rivers, It is the lateral inflow of the river. is the cross-sectional area of ​​the river, is the distance along the river channel, is the river slope, is the river channel roughness, is the wet perimeter length;

[0129] Step B12: Construct the soil erosion module and nitrogen and phosphorus loading module of the GBNP model, specifically including the nutrient salt slope process, transformation process and river process. The specific steps are as follows:

[0130] Step B121: Calculate the nutrient salt slope process, where the nutrient salt slope process is specifically an accumulation-washing process, as shown in formula (9):

[0131] , (9);

[0132] in, is the dust accumulation amount, is the proportion coefficient of pollutants in dust, is the washout coefficient, is the runoff intensity, is the maximum cumulative amount, The number of days of dust accumulation since the last rainfall, The number of days required for dust accumulation to reach half of the maximum accumulation;

[0133] Step B122: Calculate the nutrient conversion process, where the nutrient conversion process is specifically a convection-diffusion process, and the convection-diffusion process includes nutrient adsorption, desorption, volatilization and conversion processes, and the compound form conversion process is described by a first-order kinetic reaction equation, as shown in formula (10).

[0134] ;

[0135] ;

[0136] ;

[0137] (10);

[0138] in, is the adsorption term concentration, is the water migration velocity, is the hydrodynamic dispersion coefficient, is the solute source and sink term, is the soil surface nitrogen content equivalent to the amount of fertilizer applied, is the inorganic phosphorus content, is the amount of transformation in different forms, is the amount absorbed by plants, is the loss caused by runoff, is the conversion coefficient, is the nitrogen and phosphorus absorption ratio of plants at different growth stages, The ratio of the cumulative growing days on that day to the growing period;

[0139] Step B123, calculate the nutrient salt river process, where the nutrient salt river process specifically includes the convection diffusion process, the slope inflow process, the morphological transformation process, the sediment adsorption process and the sediment release process, as shown in formula (11).

[0140] ;

[0141] (11);

[0142] in, is the cross-sectional area of ​​the river flow, is the concentration of nutrient solutes in the river channel, is the cross-sectional flow rate, is the diffusion coefficient along the river direction, is the source-sink term, is the river sediment concentration, is the flow that flows laterally into the river channel, is the concentration of pollutants entering laterally, is the concentration of lateral inflow sediment, is the sediment recovery saturation coefficient, is the sediment settling rate, is the river width, The maximum sediment-carrying capacity of a river. To flush the concentration of adsorbed pollutants in silt and sand;

[0143] Step B2: Input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model and obtain the river grid runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transfer-diffusion-transformation process, thereby obtaining a distributed process model. Specifically, the grid meteorological data in the processed target basin basic database is input into the runoff-sediment-nitrogen-phosphorus model as a driving force, the point source pollution data is input into the nearest river grid based on the latitude and longitude information of the sewage outlet, the domestic pollution source is input into the residential land grid, the solid waste and livestock and poultry breeding pollution source is input into the paddy field and dry land grid, the aquaculture pollution source is input into the pond water grid and the atmospheric deposition data is input into the corresponding spatial grid. The distributed process model is calibrated according to the runoff, sediment and nutrient concentration in turn and the Nash efficiency coefficient NSE, the deviation percentage PBIAS and the ratio coefficient RSR of the root mean square error to the observation standard deviation are used as model accuracy measurement indicators.

[0144] Step C: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change results based on the land use change constraint condition. The specific steps are as follows:

[0145] Step C1: Use the future land use state transition matrix in the LUH2 dataset as a constraint on land use change, where the future land use state transition matrix includes land use maps and land use type transition maps under different future scenarios;

[0146] Step C2, the CA-Markov model is used to calculate the land use change results according to the land use change constraints. Specifically, the Markov model is used to generate a land use type transfer probability map and the area transfer matrix relative to the base period is used as a constraint to modify the land use type transfer probability map. The CA model is then used to iterate so that the constraints and the area transfer matrix are satisfied and the land use change is calculated.

[0147] Step D: Analyze the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding. Then, based on the quantitative analysis results of the amount of fertilizer and livestock and poultry breeding, integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model to measure the dynamic changes in the land-based pollution sources of fertilizer and livestock and poultry breeding pollution. The specific steps are as follows:

[0148] Step D1, analyzing the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer applied and the amount of livestock and poultry raised, wherein the climate variables include precipitation and temperature, and the socioeconomic variables include agricultural population variables, gross domestic product (GDP), and primary industry output value;

[0149] Step D2, based on the quantitative analysis results of fertilization amount and livestock and poultry breeding amount, integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model, so as to measure the dynamic changes of terrestrial pollution sources of fertilization and livestock and poultry breeding pollution, the linear regression model includes linear regression, interaction effect linear regression, robust linear regression, stepwise linear regression, the regression tree model includes fine tree, medium tree and coarse tree, the support vector machine model includes linear support vector machine, quadratic support vector machine, cubic support vector machine, fine Gaussian support vector machine, medium Gaussian support vector machine and coarse Gaussian support vector machine, the Gaussian process regression model includes square exponential Gaussian process regression model, exponential Gaussian process regression model and rational quadratic Gaussian process regression model, the tree integration model includes boosted tree integration and bagged tree integration, the specific steps are as follows,

[0150] Step D21, integrating the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model, wherein the integration is specifically based on the ensemble learning method of the Boosting framework and optimizes the function space by the gradient descent method to obtain the total ensemble model, and the goal of the total ensemble model is to minimize the loss function , the total integrated model output is , the specific steps are as follows,

[0151] Step D211, the initial value is the target mean ;

[0152] Step D212, calculate the pseudo residual, as shown in formula (12),

[0153] (12);

[0154] in, is the pseudo residual, Output results for the current model;

[0155] Step D213, fitting the residuals, using regression tree Fit the pseudo residuals and then select the splitting features and thresholds by minimizing the squared error;

[0156] Step D214, multiply the prediction result of the new tree by the learning rate And accumulated to the current model, as shown in formula (13),

[0157] (13);

[0158] Step D215, set the termination condition, if the residual converges to the set threshold, stop sampling,

[0159] Step D22, cross-validating the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model, specifically adopting a systematic data partitioning strategy to balance computational efficiency and evaluation reliability. The specific steps are as follows:

[0160] Step D221, data partitioning, randomly divide the original data set into five equal and mutually exclusive subsets for stratified sampling and ensure that the category distribution in each subset is consistent with the original data;

[0161] Step D222, iterative training and validation, the first time using the first subset as the validation set and the second to fifth subsets as the training set, the second time using the second subset as the validation set and the remaining subsets as the training set, and so on to the fifth time, each time generating an independent model;

[0162] Step D223, performance evaluation, uses the root mean square error (RMSE) as the evaluation indicator of the validation set and takes the average and standard deviation of the five results to evaluate the stability of the overall integrated model.

[0163] like Figure 3 As shown, in step E, the processed target watershed basic database, land use change results and terrestrial pollution source dynamic change results are input into the distributed process model to obtain the spatiotemporal quantitative results of the watershed non-point source pollution load under the change scenario, thereby completing the dynamic measurement of the watershed non-point source pollution load.

[0164] A system for dynamically calculating watershed non-point source pollution loads based on a distributed process model includes a data acquisition and processing module, a distributed process model construction module, a land use change calculation module, a land pollution source dynamic change calculation module, and a watershed non-point source pollution load dynamic calculation module. The data acquisition and processing module is used to collect meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, atmospheric deposition, and surface pollution source data for a target watershed and construct a target watershed basic database. The system then integrates and processes different types of data in the target watershed basic database using geospatial interpolation, spatial resampling, and frequency matching, thereby forming a processed target watershed basic database with consistent temporal and spatial scales.

[0165] The distributed process model construction module is used to construct a runoff-sediment-nitrogen-phosphorus model based on physical processes, and then input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model;

[0166] The land use change calculation module is used to use the future land use state transfer matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition;

[0167] The land pollution source dynamic change calculation module is used to analyze the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding, and then integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model based on the quantitative analysis results of the amount of fertilizer and livestock and poultry breeding, so as to calculate the dynamic change results of the land pollution sources of fertilization and livestock and poultry breeding pollution;

[0168] The dynamic measurement module of the non-point source pollution load in the watershed is used to input the processed target watershed basic database, land use change results and dynamic change results of terrestrial pollution sources into the distributed process model and obtain the spatiotemporal quantitative results of the non-point source pollution load in the watershed under the change scenario, thereby completing the dynamic measurement operation of the non-point source pollution load in the watershed.

[0169] In summary, the present invention provides a method and system for dynamically calculating the non-point source pollution load in a watershed based on a distributed process model. First, the target watershed meteorological, hydrological, water quality, soil and topography, land use, vegetation cover, atmospheric deposition and surface pollution source data are collected and a target watershed basic database is constructed. Then, geospatial interpolation, spatial resampling and frequency matching are used to integrate and process different types of data in the target watershed basic database to form a processed target watershed basic database with consistent spatiotemporal scales. Then, a runoff-sediment-nitrogen-phosphorus model based on physical processes is constructed. The processed target watershed basic database is then input into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff. The distributed process model is obtained by combining the flow process, sediment scouring-migration-sedimentation process and pollutant concentration transport-diffusion-transformation process. The future land use state transfer matrix in the LUH2 dataset is then used as the land use change constraint condition. The CA-Markov model is then used to calculate the land use change results based on the land use change constraint condition. The quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding is then analyzed. The linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model are then integrated and cross-validated based on the quantitative analysis results of the amount of fertilizer and livestock and poultry breeding, so as to calculate the relationship between fertilizer and livestock and poultry breeding. The dynamic change results of terrestrial pollution sources of aquaculture pollution are obtained. Finally, the basic database of the target watershed after processing, the results of land use change and the dynamic change results of terrestrial pollution sources are input into the distributed process model to obtain the spatiotemporal quantitative results of the watershed non-point source pollution load under the change scenario, thereby completing the dynamic measurement of the watershed non-point source pollution load; effectively realizing the dynamic measurement method and system of the watershed non-point source pollution load has the function of coupling the runoff-sediment-nitrogen and phosphorus physical processes on the basis of the distributed hydrological model and combining the driving factors of social and economic development, land use change and the dynamic changes of terrestrial pollution sources to perform spatiotemporal quantitative measurement of the watershed non-point source pollution load under the change scenario, and through spatial Interpolation and frequency matching obtain an input database with unified temporal and spatial resolution. By coupling the runoff-sediment-nitrogen and phosphorus physical processes based on a distributed hydrological model, the interface mismatch of the existing hydrological model nested in the water environment model and the model connection problems caused by the independence of the input and output processes are avoided. At the same time, the ensemble machine learning method can not only better capture the dynamic changes of the predicted input data, but also avoid the separation of the driving factors of socio-economic development from the dynamic changes of land use changes and terrestrial pollution sources. This provides a new method for accurately and quantitatively measuring the spatiotemporal changes of non-point source pollution loads in the basin under changing scenarios, and is also of great significance for guiding the refined management of water resources and water environment in the basin.

[0170] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the foregoing embodiments. The foregoing embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for dynamically calculating watershed non-point source pollution loads based on a distributed process model, characterized by: The following steps are included: Step A: Collect meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, atmospheric deposition, and surface pollution source data for the target basin and construct a basic database for the target basin. Then, geospatial interpolation, spatial resampling, and frequency matching are used to integrate and process different types of data in the basic database for the target basin, thereby forming a processed basic database for the target basin with consistent temporal and spatial scales. Step B: Construct a runoff-sediment-nitrogen-phosphorus model based on physical processes. Then, input the processed target watershed basic database into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff process, sediment scouring-migration-sedimentation process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model. Step C: The future land use state transition matrix in the LUH2 dataset is used as the land use change constraint condition, and then the CA-Markov model is used to calculate the land use change result according to the land use change constraint condition; Step D: Analyze the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding. Then, based on the quantitative analysis results of fertilizer and livestock and poultry breeding, integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models to measure the dynamic changes in land-based pollution sources of fertilizer and livestock and poultry breeding pollution. In step E, the processed target basin basic database, land use change results, and land pollution source dynamic change results are input into the distributed process model to obtain the spatiotemporal quantitative results of the basin non-point source pollution load under the change scenario, thereby completing the dynamic measurement of the basin non-point source pollution load.

2. The method for dynamically calculating watershed non-point source pollution load based on a distributed process model according to claim 1, characterized in that: Step A: Collect meteorological, hydrological, water quality, soil and topography, land use, vegetation cover, atmospheric deposition and surface pollution source data of the target basin and build a basic database of the target basin. Then, use geospatial interpolation, spatial resampling and frequency matching to integrate different types of data in the basic database of the target basin, so as to form a processed basic database of the target basin with consistent temporal and spatial scales. The specific steps are as follows: Step A1: collecting target basin meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, atmospheric deposition, and surface pollution source data and constructing a target basin basic database, wherein the target basin meteorological, hydrological, water quality, soil and topographic, land use, vegetation cover, and atmospheric deposition data are remote sensing grid data, and the surface pollution source data are yearbook statistical data. The target basin basic database also includes future climate scenario meteorological data and future atmospheric deposition scenario data, wherein the future climate scenario meteorological data are spatial grid data output by GCM, and the future atmospheric deposition scenario data are atmospheric deposition spatial grid data provided by the MRI-ESM2-0 model. Step A2: Use geospatial interpolation, spatial resampling, and frequency matching to integrate different types of data in the target basin basic database, thereby forming a processed target basin basic database with consistent temporal and spatial scales. The specific steps are as follows: Step A21, using geospatial interpolation to integrate different types of data in the target basin basic database. Specifically, the geospatial interpolation is to use the inverse distance weighted interpolation method (IDW) to interpolate the site data into space according to the spatial resolution of the remote sensing grid data and then enter it into the target basin basic database; Step A22: spatial resampling is used to integrate different types of data in the target watershed basic database. Specifically, spatial resampling is to unify the spatial resolution of soil, topography, land use, vegetation cover, and atmospheric deposition grid data using the Majority type. Step A23: Frequency matching is used to integrate and process different types of data in the target basin basic database. Specifically, frequency matching is applied to the meteorological data output by the GCM and the cumulative frequency curve of the measured climate data is used to correct the deviation of the cumulative frequency curve of the GCM simulation value, so that the cumulative frequency curve of the historical measured data remains consistent. The corrected data can then be entered into the basic database as the input data for future scenarios.

3. The method for dynamically calculating watershed non-point source pollution load based on a distributed process model according to claim 2, characterized in that: Step B: Construct a runoff-sediment-nitrogen-phosphorus model based on physical processes. Then, input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff process, sediment scouring-migration-sedimentation process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model. The specific steps are as follows: Step B1: Construct a runoff-sediment-nitrogen-phosphorus model based on physical processes, wherein the runoff-sediment-nitrogen-phosphorus model adopts the GBNP model, which includes a hydrological module, a soil erosion module, and a nitrogen-phosphorus load module. The specific steps are as follows: Step B11: Construct a hydrological module of the GBNP model. The hydrological module is based on the distributed hydrological model (GBHM), which uses the digital elevation model (DEM) to extract the watershed system and divide the hillside-river grid cells as the basic model units. Then, the hydrological processes of precipitation interception, evapotranspiration, soil water movement, and phreatic outflow are calculated in the hillside cells. The specific steps are as follows: Step B111: Calculate the hydrological process of precipitation interception. Specifically, assuming that the environmental surface is covered with vegetation, the hydrological process of precipitation interception is as shown in formula (1): Among them, E tr (t, j) is the actual transpiration rate of soil water in the jth layer from the roots to the leaves at time t, K v is the vegetation coverage rate, K c is the crop coefficient, E p is the potential evaporation, f1(Z j ) is the plant root depth distribution function, f2(θ j ) is the soil moisture function, LAI(t) is the leaf area index at time t, and LAI0 is the maximum leaf area index of vegetation in a year; Step B112: Calculate the evapotranspiration hydrological process, specifically including the situation where the environmental surface has no vegetation cover and the situation where the environmental surface water does not meet the evaporation capacity. The specific steps are as follows: Step B1121: Assuming that the environment surface is not covered by vegetation, the evapotranspiration hydrological process is as shown in formula (2): Among them, E surface (t) is the actual surface evaporation rate at time t, S s (t) is the depth of surface water at time t, and Δt is the time step; Step B1122: Assuming that the surface water in the environment does not meet the evaporation capacity, the evaporation hydrological process is as shown in formula (3): E s (t)=[(1-K v )E p -E surface (t)]f2(θ) (3); Among them, E s (t) is the actual evaporation rate of the soil surface at time t, and f2(θ) is the soil moisture function; Step B113: Calculate the soil water movement hydrological process, specifically including soil unsaturated hydraulic conductivity, unsaturated zone soil water movement process and soil moisture characteristic curve. The specific steps are as follows: Step B1131: Calculate the unsaturated hydraulic conductivity of the soil, as shown in formula (4): K(θ,z)=K s (z)θ 1 / 2 [1-(1-θ 1 / m ) m ] 2 (4); Where K(θ,z) is the unsaturated hydraulic conductivity of the soil, K s (z) is the soil saturated hydraulic conductivity, θ is the water content, and m is the fitting parameter of the soil moisture characteristic curve; Step B1132: Calculate the soil water movement process in the unsaturated zone. Specifically, use the one-dimensional Richard equation to calculate the soil water movement in the unsaturated zone. The soil water movement process in the unsaturated zone is shown in formula (5). Where z is the soil depth, θ(z,t) is the soil moisture content at the soil depth z at time t, and q v is the soil water flux, s(z,t) is the soil evaporation rate at time t, and Ψ(θ) is the soil moisture function; Step B1133: Calculate the soil moisture characteristic curve. Specifically, the soil moisture characteristic curve represents the relationship between soil moisture content and soil suction. The soil moisture characteristic curve is shown in formula (6): θ(Ψ)=θ r +(θ s -θ r )[1+(aΨ) n ] 1 / m (6); Among them, θ r is the residual moisture content of soil, θ s is the soil saturated moisture content, a and n are the fitting parameters of the soil moisture characteristic curve; Step B114: Calculate the hydrological process of phreatic outflow, including the water exchange process between the hillside phreatic layer and the river channel and the river channel confluence process. The specific steps are as follows: Step B1141: Calculate the water exchange process between the hillside aquifer and the river channel, as shown in formula (7): in, is the rate of change of groundwater storage in the saturated aquifer over time, rech(t) is the recharge rate between the saturated aquifer and the upper unsaturated aquifer, L(t) is the downward leakage rate of the saturated aquifer, q G (t) is the exchange flow rate between groundwater and river channel per unit width, A1 is the unit area of ​​hillside, K G is the saturated hydraulic conductivity of the phreatic layer, H1 and H2 are the groundwater levels of the phreatic layer before and after the water exchange, h1 and h2 are the river water levels before and after the water exchange, and l is the slope length of the hillside unit; Step B1142: Calculate the river confluence process, as shown in formula (8): Among them, Q is the total confluence of the river, q is the lateral inflow of the river, A2 is the cross-sectional area of ​​the river, x is the distance along the river, S0 is the slope of the river, and n r is the channel roughness, p is the wetted perimeter length; Step B12: Construct the soil erosion module and nitrogen and phosphorus loading module of the GBNP model, specifically including the nutrient salt slope process, transformation process and river process. The specific steps are as follows: Step B121: Calculate the nutrient salt slope process, where the nutrient salt slope process is specifically an accumulation-wash process, as shown in formula (9): Among them, D ac is the dust accumulation, c p is the proportion coefficient of pollutants in dust, c e is the scour coefficient, q r is the runoff intensity, D acm is the maximum cumulative amount, t ac is the number of days of dust accumulation since the last rainfall, t ha The number of days required for dust accumulation to reach half of the maximum accumulation; Step B122: Calculate the nutrient conversion process, wherein the nutrient conversion process is specifically a convection-diffusion process, and the convection-diffusion process includes nutrient adsorption, desorption, volatilization and conversion processes, and the compound form conversion process is described by a first-order kinetic reaction equation, as shown in formula (10), S s dt=c fer +c min -c tran -c uptake -c runoff ; c tran =-μ i c i ; Among them, c SS is the adsorption concentration, ρ is the water migration velocity, D is the hydrodynamic diffusion coefficient, S s is the solute source and sink term, c fer is the soil surface nitrogen content equivalent to the amount of fertilizer applied, c min is the inorganic phosphorus content, c tran is the amount of transformation of different forms, c uptake is the amount absorbed by plants, c runoff is the loss caused by runoff, μ i is the conversion coefficient, F up is the nitrogen and phosphorus absorption ratio of plants at different growth stages, f gs The ratio of the cumulative growing days on that day to the growing period; Step B123: Calculate the nutrient salt channel process, where the nutrient salt channel process specifically includes the convection diffusion process, the slope inflow process, the morphological transformation process, the sediment adsorption process, and the sediment release process, as shown in formula (11). Among them, A3 is the cross-sectional area of ​​the river, c rd is the concentration of nutrient solute in the river, Q is the cross-sectional flow, E x is the diffusion coefficient along the river, S is the source and sink term, c r is the sediment concentration in the river channel, q l is the lateral inflow into the river channel, c rsl is the concentration of pollutants entering laterally, c rl is the concentration of lateral inflow sediment, α is the sediment recovery saturation coefficient, ω is the sediment settling rate, B is the river channel width, is the maximum sediment carrying capacity of the river, c k To flush the concentration of adsorbed pollutants in silt and sand; Step B2: Input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model and obtain the river grid runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transfer-diffusion-transformation process, thereby obtaining a distributed process model. Specifically, the grid meteorological data in the processed target basin basic database is input into the runoff-sediment-nitrogen-phosphorus model as a driving force, the point source pollution data is input into the nearest river grid based on the latitude and longitude information of the sewage outlet, the domestic pollution source is input into the residential land grid, the solid waste and livestock and poultry breeding pollution source is input into the paddy field and dry land grid, the aquaculture pollution source is input into the pond water grid and the atmospheric deposition data is input into the corresponding spatial grid. The distributed process model is calibrated according to the runoff, sediment and nutrient concentration in turn and the Nash efficiency coefficient NSE, the deviation percentage PBIAS and the ratio coefficient RSR of the root mean square error to the observation standard deviation are used as model accuracy measurement indicators.

4. The method for dynamically calculating watershed non-point source pollution load based on a distributed process model according to claim 3, characterized in that: Step C: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change results based on the land use change constraint condition. The specific steps are as follows: Step C1: Use the future land use state transition matrix in the LUH2 dataset as a constraint on land use change, where the future land use state transition matrix includes land use maps and land use type transition maps under different future scenarios; Step C2, the CA-Markov model is used to calculate the land use change results according to the land use change constraints. Specifically, the Markov model is used to generate a land use type transfer probability map and the area transfer matrix relative to the base period is used as a constraint to modify the land use type transfer probability map. The CA model is then used to iterate so that the constraints and the area transfer matrix are satisfied and the land use change is calculated.

5. The method for dynamically calculating watershed non-point source pollution load based on a distributed process model according to claim 4, characterized in that: Step D: Analyze the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding. Then, based on the quantitative analysis results of the amount of fertilizer and livestock and poultry breeding, integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model to measure the dynamic changes in the land-based pollution sources of fertilizer and livestock and poultry breeding pollution. The specific steps are as follows: Step D1, analyzing the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer applied and the amount of livestock and poultry raised, wherein the climate variables include precipitation and temperature, and the socioeconomic variables include agricultural population variables, gross domestic product (GDP), and primary industry output value; Step D2, based on the quantitative analysis results of fertilization amount and livestock and poultry breeding amount, integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model, so as to measure the dynamic changes of terrestrial pollution sources of fertilization and livestock and poultry breeding pollution, the linear regression model includes linear regression, interaction effect linear regression, robust linear regression, stepwise linear regression, the regression tree model includes fine tree, medium tree and coarse tree, the support vector machine model includes linear support vector machine, quadratic support vector machine, cubic support vector machine, fine Gaussian support vector machine, medium Gaussian support vector machine and coarse Gaussian support vector machine, the Gaussian process regression model includes square exponential Gaussian process regression model, exponential Gaussian process regression model and rational quadratic Gaussian process regression model, the tree integration model includes boosted tree integration and bagged tree integration, the specific steps are as follows, Step D21, integrating the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model, wherein the integration is specifically based on the ensemble learning method of the Boosting framework and is optimized in the function space by the gradient descent method to obtain the total ensemble model. The goal of the total ensemble model is to minimize the loss function L(y,F(x)), and the output of the total ensemble model is F(x). The specific steps are as follows: Step D211, the initial value is the target mean F0(x); Step D212, calculate the pseudo residual, as shown in formula (12), r ti =y i -F t-1 (x i ) (12); Among them, r ti is the pseudo residual, F t-1 (x) is the output result of the current model; Step D213, fitting the residual, using the regression tree h t (x) Fit the pseudo residuals and then select the splitting features and thresholds by minimizing the squared error; Step D214, multiply the prediction result of the new tree by the learning rate ν and add it to the current model, as shown in formula (13), F(x)=F t-1 (x)+vh t (x) (13); Step D215, set the termination condition, if the residual converges to the set threshold, stop sampling, Step D22, cross-validating the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree ensemble model, specifically adopting a systematic data partitioning strategy to balance computational efficiency and evaluation reliability. The specific steps are as follows: Step D221, data partitioning, randomly divide the original data set into five equal and mutually exclusive subsets for stratified sampling and ensure that the category distribution in each subset is consistent with the original data; Step D222, iterative training and validation, the first time using the first subset as the validation set and the second to fifth subsets as the training set, the second time using the second subset as the validation set and the remaining subsets as the training set, and so on to the fifth time, each time generating an independent model; Step D223, performance evaluation, uses the root mean square error (RMSE) as the evaluation indicator of the validation set and takes the average and standard deviation of the five results to evaluate the stability of the overall integrated model.

6. A system for dynamically calculating watershed non-point source pollution loads based on a distributed process model, wherein the specific calculation process of the system is based on the method for dynamically calculating watershed non-point source pollution loads according to any one of claims 1 to 5, and is characterized by: It includes a data acquisition and processing module, a distributed process model construction module, a land use change calculation module, a land pollution source dynamic change calculation module and a basin non-point source pollution load dynamic calculation module. The data acquisition and processing module is used to collect meteorological, hydrological, water quality, soil and topography, land use, vegetation cover, atmospheric deposition and surface pollution source data of the target basin and build a target basin basic database. Then, geospatial interpolation, spatial resampling and frequency matching are used to integrate and process different types of data in the target basin basic database, so as to form a processed target basin basic database with consistent temporal and spatial scales. The distributed process model construction module is used to construct a runoff-sediment-nitrogen-phosphorus model based on physical processes, and then input the processed target basin basic database into the runoff-sediment-nitrogen-phosphorus model to obtain the river grid runoff process, sediment scouring-migration-sedimentation process and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model; The land use change calculation module is used to use the future land use state transfer matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition; The land pollution source dynamic change calculation module is used to analyze the quantitative relationship between climate variables and socioeconomic variables and the amount of fertilizer and livestock and poultry breeding, and then integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model and tree integration model based on the quantitative analysis results of the amount of fertilizer and livestock and poultry breeding, so as to calculate the dynamic change results of the land pollution sources of fertilization and livestock and poultry breeding pollution; The dynamic measurement module of the non-point source pollution load in the watershed is used to input the processed target watershed basic database, land use change results and dynamic change results of terrestrial pollution sources into the distributed process model and obtain the spatiotemporal quantitative results of the non-point source pollution load in the watershed under the change scenario, thereby completing the dynamic measurement operation of the non-point source pollution load in the watershed.

Citation Information

Patent Citations

  • Distributed simulation method of non-point source nitrogen and phosphorus loss form constitution in hilly regions

    CN107066808A

  • Non-point source pollution control method

    CN116993235A