Drainage basin non-point source pollution load dynamic measurement and calculation method and system based on distributed process model
Through distributed process models and machine learning methods, combined with meteorological and hydrological data, runoff-silt-nitrogen and phosphorus model is constructed, which solves the problem of insufficient accuracy of dynamic calculation of non-point source pollution in the basin, and realizes spatiotemporal quantitative calculation and refined management in a changing environment.
Patent Information
- Application Number
- CN202510758246.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-09
AI Technical Summary
It is difficult for the existing technology to realize spatiotemporal quantitative calculation of non-point source pollution in a dynamic changing environment, and the driving factors of social and economic development are separated from changes in land use and dynamic changes in land pollution sources, resulting in insufficient calculation accuracy.
A distributed process model is used, combining meteorological, hydrological, soil, terrain and other data to construct a runoff-silt-nitrogen-phosphorus model, combined with CA-Markov model and machine learning method, to calculate the dynamic changes of land use and land pollution sources, and realize spatiotemporal quantitative measurement.
The quantitative measurement of the non-point source pollution load in the basin under changing scenarios has been realized, which avoids the separation of social and economic development factors and land use changes, improves the calculation accuracy, and supports the refined management of basin water resources and water environment.
Smart Images

Figure CN120297771A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydrological pollution load measurement, and in particular to a dynamic measurement method and system for watershed non-point source pollution load based on a distributed process model. Background Art
[0002] The dynamic quantification of watershed non-point source pollution load is a core proposition in water environment governance, and its complexity stems from the multi-scale coupling effect of hydrological processes, human activities and ecological environment. In watersheds where the control effects of industrial and domestic pollution sources are relatively stable, atmospheric deposition brought about by climate change and changes in the scale and composition structure of watershed agriculture have become the main sources of non-point source pollution. Due to the significant spatio-temporal heterogeneity of the changes in industrial structure and pollution sources, and the complex dynamic relationship between the two, the quantification method of their contribution to non-point source pollution is not yet mature, further leading to an increase in the uncertainty of the impact of non-point source pollution on river water quality changes, which has become a difficult point in the management and control of watershed water resources and water environment.
[0003] At present, most of the methods for measuring watershed non-point source pollution are based on field monitoring, such as the SCS-CN, USLE and RULSE formula methods. These methods can only quantify the non-point source pollution from different sources in the historical period with existing basic data, and it is difficult to achieve the spatio-temporal quantitative measurement of watershed non-point source pollution under dynamic change scenarios. The non-point source pollution measurement methods based on machine learning generally only measure watershed non-point source pollution based on hydrometeorological, pollutant time series data and remote sensing image spatial characteristics, resulting in a lack of process mechanism explanation, and being greatly affected by uncertain factors in a changing environment. At the same time, there are also technical bottlenecks in the disconnection between social and economic development driving factors and land use changes and the dynamic changes of land-based pollution sources. In the context of the accelerating urbanization process, there is also a situation of insufficient dynamic measurement accuracy of non-point source pollution spatio-temporally; therefore, it is necessary to design a dynamic measurement method and system for watershed non-point source pollution load based on a distributed process model. Summary of the Invention
[0004] The objective of the present invention is to overcome the deficiencies of the prior art and, in order to better and effectively solve the problem that most of the current methods for calculating non-point source pollution in river basins are based on field monitoring, such as the SCS-CN, USLE, and RULSE formula methods. These methods can only quantify non-point source pollution from different sources in historical periods with existing basic data and are difficult to achieve spatio-temporal quantitative calculation of non-point source pollution in river basins under dynamic change scenarios. Moreover, the methods for calculating non-point source pollution based on machine learning generally only calculate non-point source pollution in river basins based on hydrometeorological, pollutant time series data, and spatial characteristics of remote sensing images, resulting in a lack of process mechanism explanation and being greatly affected by uncertain factors in a changing environment. At the same time, there is also a technical bottleneck in the disconnection between social and economic development driving factors, land use changes, and dynamic changes in land-based pollution sources. Under the background of accelerating urbanization, there is also a problem of insufficient dynamic calculation accuracy in the spatio-temporal non-point source pollution. The present invention provides a method and system for dynamically calculating non-point source pollution load in a river basin based on a distributed process model, which realizes the function of spatio-temporal quantitative calculation of non-point source pollution load in a river basin under changing scenarios by coupling runoff-sediment-nitrogen and phosphorus physical processes on the basis of a distributed hydrological model and combining social and economic development driving factors, land use changes, and dynamic changes in land-based pollution sources. It can not only better capture and predict the dynamic changes of input data but also avoid the disconnection between social and economic development driving factors, land use changes, and dynamic changes in land-based pollution sources.
[0005] In order to achieve the above objective, the technical solution adopted by the present invention is as follows: A method for dynamically calculating non-point source pollution load in a river basin based on a distributed process model, comprising the following steps: Step A: Collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target river basin and construct a basic database of the target river basin. Then, integrate and process different types of data in the basic database of the target river basin by using geospatial interpolation, spatial resampling, and frequency matching to form a processed basic database of the target river basin with consistent spatio-temporal scales. Step B: Construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. Then, input the processed basic database of the target river basin into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, sediment erosion-transport-settlement process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model. Step C: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then calculate the land use change result by using the CA-Markov model according to the land use change constraint condition. Step D: Analyze the quantitative relationships between climate variables, socioeconomic variables, fertilization amount, and livestock and poultry breeding amount. Then, based on the quantitative analysis results of fertilization amount and livestock and poultry breeding amount, integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models to calculate the dynamic change results of land-based pollution sources of fertilization and livestock and poultry breeding pollution. Step E: Input the processed basic database of the target watershed, land use change results, and dynamic change results of land-based pollution sources into a distributed process model to obtain the spatio-temporal quantitative results of non-point source pollution load in the watershed under the change scenario, thereby completing the dynamic calculation operation of non-point source pollution load in the watershed.
[0006] The foregoing method for dynamically calculating non-point source pollution load in a watershed based on a distributed process model, Step A: Collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target watershed and construct a basic database of the target watershed. Then, integrate and process different types of data in the basic database of the target watershed using geospatial interpolation, spatial resampling, and frequency matching to form a processed basic database of the target watershed with consistent spatio-temporal scales. The specific steps are as follows. Step A1: Collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target watershed and construct a basic database of the target watershed. Among them, the meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, and atmospheric deposition data of the target watershed are remote sensing grid data, the surface pollution source data are annual statistical data, the basic database of the target watershed also includes meteorological data of future climate scenarios and atmospheric deposition scenario data of the future. The meteorological data of the future climate scenario are spatial grid data output by GCM, and the atmospheric deposition scenario data of the future are atmospheric deposition spatial grid data provided by the MRI-ESM2-0 model. Step A2: Integrate and process different types of data in the basic database of the target watershed using geospatial interpolation, spatial resampling, and frequency matching to form a processed basic database of the target watershed with consistent spatio-temporal scales. The specific steps are as follows. Step A21: Integrate and process different types of data in the basic database of the target watershed using geospatial interpolation. Specifically, the geospatial interpolation uses the inverse distance weighted interpolation method (IDW) to interpolate the station data into space according to the spatial resolution of the remote sensing grid data and then input it into the basic database of the target watershed. Step A22: Integrate and process different types of data in the basic database of the target watershed using spatial resampling. Specifically, the spatial resampling uses the Majority type to unify the spatial resolution of soil, terrain, land use, vegetation cover, and atmospheric deposition grid data. Step A23: Integrate and process different types of data in the target watershed basic database using frequency matching. Specifically, the frequency matching is applied to the meteorological data output by GCM, and the cumulative frequency curve of the measured climate data is used to correct the deviation of the cumulative frequency curve of the GCM simulation values, so that the cumulative frequency curves of the historical measured data are consistent, and then the calibrated data can be entered into the basic database as future scenario input data.
[0007] For the aforementioned method for dynamically calculating the non-point source pollution load in a watershed based on a distributed process model, in step B, a runoff-sediment-nitrogen and phosphorus model based on physical processes is constructed, and then the processed target watershed basic database is input into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, sediment scouring-transportation-settlement process, and pollutant concentration transport-diffusion-transformation process, so as to obtain a distributed process model. The specific steps are as follows. Step B1: Construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. The runoff-sediment-nitrogen and phosphorus model adopts the GBNP model, and the GBNP model includes a hydrological module, a soil erosion module, and a nitrogen and phosphorus load module. The specific steps are as follows. Step B11: Construct the hydrological module of the GBNP model. The hydrological module extracts the watershed water system based on the distributed hydrological model GBHM using the digital elevation model DEM and divides the hillslope-river channel grid units as the basic units of the model. Then, the precipitation interception, evapotranspiration, soil water movement, and phreatic water outflow hydrological processes are calculated in the hillslope units. The specific steps are as follows. Step B111: Calculate the precipitation interception hydrological process. Specifically, assume that the environmental surface has vegetation cover, and the precipitation interception hydrological process is shown in formula (1). (1); Where, is the actual transpiration rate of the soil water in the th layer from the vegetation root system to the leaf surface at time is the vegetation coverage rate, is the crop coefficient, is the potential evaporation, is the plant root depth distribution function, is the soil moisture content function, is the leaf area index at time is the maximum value of the leaf area index of the vegetation in a year; Step B112: Calculate the evapotranspiration hydrological process, specifically including the case where the environmental surface has no vegetation cover and the case where the environmental surface water accumulation does not meet the evaporation capacity. The specific steps are as follows. Step B1121: Assume that the environmental surface has no vegetation cover, then the evapotranspiration hydrological process is shown in formula (2). (2); Wherein, is the actual evaporation rate of the ground surface at time is the accumulated water depth on the ground surface at time is the time step; Step B1122: Assume that the accumulated water on the environmental ground surface does not meet the evaporation capacity, then the evapotranspiration hydrological process is as shown in formula (3), (3); Wherein, is the actual evaporation rate of the soil surface at time is the soil moisture content function; Step B113: Calculate the soil water movement hydrological process, which specifically includes the soil unsaturated hydraulic conductivity, the soil water movement process in the unsaturated zone, and the soil water characteristic curve. The specific steps are as follows, Step B1131: Calculate the soil unsaturated hydraulic conductivity, as shown in formula (4), (4); Wherein, is the soil unsaturated hydraulic conductivity, is the soil saturated hydraulic conductivity, is the water content, is the fitting parameter of the soil water characteristic curve; Step B1132: Calculate the soil water movement process in the unsaturated zone. Specifically, the one-dimensional Richard equation is used to calculate the soil water movement in the unsaturated zone. The soil water movement process in the unsaturated zone is as shown in formula (5), (5); Wherein, is the soil depth, is the water content at the soil depth of time at is the soil water flux, is the soil evaporation rate at time is the soil moisture content function; Step B1132: Calculate the soil water characteristic curve. Specifically, the soil water characteristic curve represents the relationship between the soil water content and the soil suction. The soil water characteristic curve is as shown in formula (6), (6); Wherein, is the soil residual water content, is the soil saturated water content, and Both are fitting parameters of the soil water characteristic curve; Step B114: Calculate the groundwater outflow hydrograph, which specifically includes the water volume exchange process between the hillslope phreatic aquifer and the river channel and the river channel confluence process. The specific steps are as follows: Step B1141: Calculate the water volume exchange process between the hillslope phreatic aquifer and the river channel, as shown in formula (7): (7); Where: is the change rate of groundwater storage in the saturated aquifer with time; is the recharge rate between the saturated aquifer and the upper unsaturated aquifer; is the downward leakage rate of the saturated aquifer; is the single-width exchange flow of groundwater and the river channel; is the area of the hillslope unit; is the saturated hydraulic conductivity of the phreatic aquifer; and are the groundwater levels in the phreatic aquifer before and after the water volume exchange; and are the river channel water levels before and after the water volume exchange; is the slope length of the hillslope unit; Step B1142: Calculate the river channel confluence process, as shown in formula (8): (8); Where: is the total river channel confluence; is the lateral inflow of the river channel; is the cross-sectional area of the river channel; is the distance along the river channel direction; is the river channel slope; is the river channel roughness; is the wetted perimeter length; Step B12: Construct the soil erosion module and nitrogen and phosphorus load module of the GBNP model, which specifically includes the nutrient slope process, transformation process, and river channel process. The specific steps are as follows: Step B121: Calculate the nutrient slope process, where the nutrient slope process is specifically the accumulation-scouring process, as shown in formula (9): , (9); Where: is the dust accumulation amount; is the proportional coefficient of pollutants in the dust; is the scouring coefficient; is the runoff intensity; is the maximum accumulation amount; is the number of days of dust accumulation after the last rainfall, is the number of days required for the dust to accumulate to half of the maximum accumulation; Step B122, calculate the nutrient transformation process, where the nutrient transformation process is specifically a convection-diffusion process, and the convection-diffusion process includes nutrient adsorption, desorption, volatilization, and transformation processes, and the compound form transformation process is described by a first-order kinetic reaction equation, as shown in formula (10), ; ; ; (10); where, is the adsorption term concentration, is the moisture migration velocity, is the hydrodynamic dispersion coefficient, is the solute source-sink term, is the equivalent soil surface nitrogen content of the fertilization amount, is the inorganic phosphorus content, is the transformation amount of different forms, is the plant absorption amount, is the loss amount caused by runoff, is the transformation coefficient, is the nitrogen and phosphorus absorption ratio of plants at different growth stages, is the ratio of the daily cumulative growth days to the growth period; Step B123, calculate the nutrient river channel process, where the nutrient river channel process specifically includes a convection-diffusion process, a slope along-flow convergence process, a form transformation process, a sediment adsorption process, and a sediment release process, as shown in formula (11), ; (11); where, is the cross-sectional area of the river channel, is the nutrient solute concentration in the river channel, is the cross-sectional flow rate, is the dispersion coefficient along the river direction, is the source-sink term, is the sediment concentration in the river channel, is the lateral inflow discharge into the river channel, is the pollutant concentration of the lateral inflow, is the sediment concentration of the lateral inflow, is the sediment recovery saturation coefficient, is the sediment settlement rate, is the river channel width, is the maximum sediment-carrying capacity of the river, is the concentration of adsorbed pollutants in the sediment for erosion and deposition; Step B2: Input the processed target watershed basic database into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, sediment erosion-transport-settlement process, and pollutant concentration transport-diffusion-transformation process, so as to obtain a distributed process model. Specifically, input the grid meteorological data in the processed target watershed basic database into the runoff-sediment-nitrogen and phosphorus model as the driving force, input the point source pollution data into the river channel grid closest to the sewage outlet based on the longitude and latitude information of the sewage outlet, input the domestic pollution source into the residential land grid, input the solid waste and livestock and poultry breeding pollution sources into the paddy field and dry land grids, input the aquaculture pollution source into the pond water area grid, and input the atmospheric deposition data into the corresponding spatial grid. The distributed process model is calibrated in sequence according to runoff, sediment, and nutrient concentration, and the Nash-Sutcliffe efficiency coefficient NSE, percent bias PBIAS, and ratio coefficient RSR of root mean square error to standard deviation of observations are used as the model accuracy measurement indicators.
[0008] For the aforementioned method for dynamically measuring the non-point source pollution load of a watershed based on a distributed process model, step C: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to measure the land use change result according to the land use change constraint condition. The specific steps are as follows: Step C1: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, where the future land use state transition matrix includes land use maps and land use type transition maps under different future scenarios; Step C2: Use the CA-Markov model to measure the land use change result according to the land use change constraint condition. Specifically, use the Markov model to generate a land use type transfer probability map and correct the land use type transfer probability map with the area transfer matrix relative to the base period as the constraint condition, and then use the CA model to iterate to meet the constraint condition and the area transfer matrix and measure the land use change.
[0009] For the aforementioned method for dynamically measuring the non-point source pollution load of a watershed based on a distributed process model, step D: Analyze the quantitative relationship between climate variables and socioeconomic variables and fertilization amount and livestock and poultry breeding amount, and then integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models according to the quantitative analysis results of fertilization amount and livestock and poultry breeding amount, so as to measure the dynamic change results of land-based pollution sources of fertilization and livestock and poultry breeding. The specific steps are as follows: Step D1: Analyze the quantitative relationships between climate variables, socio-economic variables, fertilizer application rates, and livestock and poultry breeding amounts. The climate variables include precipitation and temperature, and the socio-economic variables include agricultural population variables, gross domestic product (GDP), and the output value of the primary industry. Step D2: Integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models based on the quantitative analysis results of fertilizer application rates and livestock and poultry breeding amounts, so as to measure the dynamic changes of land-based pollution sources of fertilizer application and livestock and poultry breeding. The linear regression models include linear regression, interactive effect linear regression, robust linear regression, and stepwise linear regression. The regression tree models include fine trees, medium trees, and rough trees. The support vector machine models include linear support vector machines, quadratic support vector machines, cubic support vector machines, fine Gaussian support vector machines, medium Gaussian support vector machines, and rough Gaussian support vector machines. The Gaussian process regression models include squared exponential Gaussian process regression models, exponential Gaussian process regression models, and rational quadratic Gaussian process regression models. The tree ensemble models include boosting tree ensembles and bagging tree ensembles. The specific steps are as follows: Step D21: Integrate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models. The integration is specifically an ensemble learning method based on the Boosting framework and is optimized in the function space through the gradient descent method to obtain the total ensemble model. The goal of the total ensemble model is to minimize the loss function. The output of the total ensemble model is The specific steps are as follows: Step D211: The initial value is the target mean. ; Step D212: Calculate the pseudo-residuals, as shown in formula (12). (12); Where is the pseudo-residual, is the output result of the current model; Step D213: Fit the residuals. Use the regression tree to fit the pseudo-residuals, and then select the splitting feature and threshold by minimizing the squared error. Step D214: Multiply the prediction result of the new tree by the learning rate and accumulate it to the current model, as shown in formula (13). (13); Step D215: Set the termination condition. If the residuals converge to the set threshold, stop sampling. Step D22, cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model, and tree ensemble model. Specifically, a systematic data partitioning strategy is adopted to balance computational efficiency and evaluation reliability. The specific steps are as follows: Step D221, data partitioning: Randomly partition the original dataset into five equal and mutually exclusive subsets for stratified sampling to ensure that the class distribution in each subset is the same as that of the original data. Step D222, iterative training and validation: For the first time, use the first subset as the validation set and the second to fifth subsets as the training set. For the second time, use the second subset as the validation set and the remaining subsets as the training set, and so on until the fifth time. Each time, an independent model is generated. Step D223, performance evaluation: Use the root mean square error (RMSE) as the evaluation metric for the validation set, and take the average and standard deviation of the five results to evaluate the stability of the overall ensemble model.
[0010] A dynamic calculation system for watershed non-point source pollution load based on a distributed process model, including a data acquisition and processing module, a distributed process model construction module, a land use change calculation module, a terrestrial pollution source dynamic change calculation module, and a watershed non-point source pollution load dynamic calculation module. The data acquisition and processing module is used to collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target watershed and construct a basic database for the target watershed. Then, different types of data in the basic database of the target watershed are integrated and processed by geospatial interpolation, spatial resampling, and frequency matching to form a processed basic database of the target watershed with consistent spatio-temporal scales. The distributed process model construction module is used to construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. Then, the processed basic database of the target watershed is input into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, sediment scouring-migration-settlement process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining the distributed process model. The land use change calculation module is used to use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition. The terrestrial pollution source dynamic change calculation module is used to analyze the quantitative relationships between climate variables and socioeconomic variables and fertilization amount and livestock and poultry breeding amount. Then, based on the quantitative analysis results of fertilization amount and livestock and poultry breeding amount, the linear regression model, regression tree model, support vector machine model, Gaussian process regression model, and tree ensemble model are integrated and cross-validated to calculate the dynamic change results of terrestrial pollution sources of fertilization and livestock and poultry breeding pollution. The dynamic calculation module for non-point source pollution load in the basin is used to input the processed basic database of the target basin, the results of land use change, and the dynamic change results of land-based pollution sources into a distributed process model and obtain the spatio-temporal quantitative results of non-point source pollution load in the basin under changing scenarios, so as to complete the dynamic calculation operation of non-point source pollution load in the basin.
[0011] The beneficial effects of the present invention are as follows: A dynamic calculation method and system for non-point source pollution load in a basin based on a distributed process model of the present invention first collect data on meteorology, hydrology, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution sources in the target basin and construct a basic database for the target basin. Then, geospatial interpolation, spatial resampling, and frequency matching are used to integrate and process different types of data in the basic database of the target basin to form a processed basic database of the target basin with consistent spatio-temporal scales. Next, a runoff-sediment-nitrogen and phosphorus model based on physical processes is constructed. Then, the processed basic database of the target basin is input into the runoff-sediment-nitrogen and phosphorus model, and the river channel grid runoff process, sediment erosion-transport-settlement process, and pollutant concentration transport-diffusion-transformation process are obtained to obtain a distributed process model. Subsequently, the future land use state transition matrix in the LUH2 dataset is used as a land use change constraint condition, and the CA-Markov model is used to calculate the land use change result according to the land use change constraint condition. Then, the quantitative relationships between climate variables and socio-economic variables and fertilizer application rates and livestock and poultry breeding amounts are analyzed. Then, based on the quantitative analysis results of fertilizer application rates and livestock and poultry breeding amounts, linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models are integrated and cross-validated to calculate the dynamic change results of land-based pollution sources of fertilizer application and livestock and poultry breeding. Finally, the processed basic database of the target basin, the land use change result, and the dynamic change result of land-based pollution sources are input into the distributed process model, and the spatio-temporal quantitative results of non-point source pollution load in the basin under the changing scenario are obtained, thus completing the dynamic calculation operation of non-point source pollution load in the basin; effectively realizing that the dynamic calculation method and system for non-point source pollution load in the basin have the function of coupling the runoff-sediment-nitrogen and phosphorus physical processes on the basis of a distributed hydrological model and combining socio-economic development driving factors, land use change, and dynamic change of land-based pollution sources to conduct spatio-temporal quantitative calculation of non-point source pollution load in the basin under the changing scenario. Moreover, a unified input database with spatio-temporal resolution is obtained through spatial interpolation and frequency matching. And by coupling the runoff-sediment-nitrogen and phosphorus physical processes on the basis of a distributed hydrological model, the interface mismatch between the existing hydrological model and the water environment model and the model connection problem caused by the independent input and output processes are avoided. At the same time, through the ensemble machine learning method, not only can the dynamic changes of prediction input data be better captured, but also the separation of socio-economic development driving factors from land use change and dynamic change of land-based pollution sources is avoided. This provides a new method for accurately and quantitatively calculating the spatio-temporal changes of non-point source pollution load in the basin under the changing scenario, and also has important significance for guiding the refined management of basin water resources and water environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is the overall flowchart of a dynamic calculation method for non-point source pollution load in a basin based on a distributed process model of the present invention; Figure 2 It is a schematic diagram of frequency curve matching of the present invention; Figure 3 It is a schematic diagram of the total nitrogen measurement result and total phosphorus measurement result in the basin under changing scenarios in the example of the present invention. Specific implementation manners
[0013] Next, the present invention will be further described in conjunction with the accompanying drawings of the specification.
[0014] As Figure 1 shown, a method for dynamically measuring non-point source pollution load in a basin based on a distributed process model of the present invention includes the following steps. Step A, collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target basin and construct a target basin basic database, and then use geospatial interpolation, spatial resampling, and frequency matching to integrate and process different types of data in the target basin basic database, so as to form a processed target basin basic database with consistent spatio-temporal scales. The specific steps are as follows. Step A1, collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target basin and construct a target basin basic database, where the meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, and atmospheric deposition data of the target basin are remote sensing grid data, the surface pollution source data are yearbook statistical data, the target basin basic database also includes meteorological data of future climate scenarios and atmospheric deposition scenario data of the future, the meteorological data of the future climate scenario are spatial grid data output by GCM, and the atmospheric deposition scenario data of the future are atmospheric deposition spatial grid data provided by the MRI-ESM2-0 model; Step A2, use geospatial interpolation, spatial resampling, and frequency matching to integrate and process different types of data in the target basin basic database, so as to form a processed target basin basic database with consistent spatio-temporal scales. The specific steps are as follows. Step A21, use geospatial interpolation to integrate and process different types of data in the target basin basic database. The specific geospatial interpolation is to use the inverse distance weighted interpolation method IDW to interpolate the station data into space according to the spatial resolution of the remote sensing grid data and then input it into the target basin basic database; Step A22, use spatial resampling to integrate and process different types of data in the target basin basic database. The specific spatial resampling is to use the Majority type to unify the spatial resolution of the soil, terrain, land use, vegetation cover, and atmospheric deposition grid data; As Figure 2As shown in the figure, in step A23, frequency matching is used to integrate different types of data in the basic database of the target basin. Specifically, frequency matching is applied to the meteorological data output by GCM, and the cumulative frequency curve of the GCM simulation values is corrected for deviation using the cumulative frequency curve of the measured climate data, so that the cumulative frequency curves of the historical measured data are kept consistent. Furthermore, the verified data can be used as input data for future scenarios and entered into the basic database.
[0015] Step B: Construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. Then, input the processed basic database of the target basin into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, sediment scouring-transportation-settlement process, and pollutant concentration transport-diffusion-transformation process, so as to obtain a distributed process model. The specific steps are as follows: Step B1: Construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. The runoff-sediment-nitrogen and phosphorus model uses the GBNP model, and the GBNP model includes a hydrological module, a soil erosion module, and a nitrogen and phosphorus load module. The specific steps are as follows: Step B11: Construct the hydrological module of the GBNP model. The hydrological module extracts the river system of the basin using the digital elevation model DEM based on the distributed hydrological model GBHM and divides the hillslope-river channel grid units as the basic units of the model. Then, calculate the precipitation interception, evapotranspiration, soil water movement, and phreatic water outflow hydrological processes in the hillslope units. The specific steps are as follows: Step B111: Calculate the precipitation interception hydrological process. Specifically, assume that the surface of the environment is covered with vegetation, and the precipitation interception hydrological process is shown in formula (1): (1); Where, is the actual transpiration rate of the soil water in the th layer from the vegetation root system to the leaf surface at time is the vegetation coverage rate, is the crop coefficient, is the potential evaporation, is the plant root depth distribution function, is the soil moisture content function, is the leaf area index at time is the maximum value of the leaf area index of the vegetation in a year; Step B112: Calculate the evapotranspiration hydrological process, specifically including the case where the surface of the environment has no vegetation coverage and the case where the surface water accumulation in the environment does not meet the evaporation capacity. The specific steps are as follows: Step B1121: Assume that the surface of the environment has no vegetation coverage, and the evapotranspiration hydrological process is shown in formula (2): (2); where is the actual evaporation rate of the land surface at time is the water depth on the land surface at time and is the time step; (3); where is the actual evaporation rate of the soil surface at time and is the soil water content function; Step B113. Calculate the hydrological process of soil water movement, which specifically includes the soil unsaturated hydraulic conductivity, the soil water movement process in the unsaturated zone, and the soil water characteristic curve. The specific steps are as follows. (4); where is the soil unsaturated hydraulic conductivity, is the soil saturated hydraulic conductivity, is the water content, and are the fitting parameters of the soil water characteristic curve; (5); where is the soil depth, is the water content at the soil depth at time is the soil water flux, is the soil evaporation rate at time and is the soil water content function; (6); where is the residual soil water content, is the saturated soil water content, and are both the fitting parameters of the soil water characteristic curve; Step B114: Calculate the hydrological process of phreatic outflow, specifically including the water volume exchange process between the hillside phreatic aquifer and the river channel and the river channel confluence process. The specific steps are as follows: Step B1141: Calculate the water volume exchange process between the hillside phreatic aquifer and the river channel, as shown in formula (7): (7); Where: is the change rate of groundwater storage in the saturated aquifer with time; is the recharge rate between the saturated aquifer and the upper unsaturated aquifer; is the downward leakage rate of the saturated aquifer; is the unit-width exchange flow of groundwater and the river channel; is the area of the hillside unit; is the saturated hydraulic conductivity of the phreatic aquifer; and are the phreatic water levels in the phreatic aquifer before and after the water volume exchange; and are the river channel water levels before and after the water volume exchange; is the slope length of the hillside unit; Step B1142: Calculate the river channel confluence process, as shown in formula (8): (8); Where: is the total river channel confluence; is the lateral inflow of the river channel; is the cross-sectional area of the river channel; is the distance along the river channel direction; is the river channel slope; is the river channel roughness; is the wetted perimeter length; Step B12: Construct the soil erosion module and nitrogen and phosphorus load module of the GBNP model, specifically including the nutrient slope process, transformation process, and river channel process. The specific steps are as follows: Step B121: Calculate the nutrient slope process, where the nutrient slope process is specifically the accumulation-scouring process, as shown in formula (9): , (9); Where: is the dust accumulation amount; is the proportional coefficient of pollutants in the dust; is the scouring coefficient; is the runoff intensity; is the maximum accumulation amount; is the number of days of dust accumulation after the previous rainfall; The number of days required for dust to accumulate to half of the maximum accumulation amount; Step B122: Calculate the nutrient transformation process. The nutrient transformation process is specifically a convection-diffusion process, and the convection-diffusion process includes nutrient adsorption, desorption, volatilization, and transformation processes. The compound form transformation process is described by a first-order kinetic reaction equation, as shown in formula (10). ; ; ; (10); Where, is the adsorption term concentration, is the moisture migration velocity, is the hydrodynamic dispersion coefficient, is the solute source-sink term, is the equivalent soil surface nitrogen content of the fertilization amount, is the inorganic phosphorus content, is the transformation amount of different forms, is the plant absorption amount, is the loss amount caused by runoff, is the transformation coefficient, is the nitrogen-phosphorus absorption ratio of plants at different growth stages, is the ratio of the cumulative growth days of the day to the growth period; Step B123: Calculate the nutrient river process. The nutrient river process specifically includes a convection-diffusion process, a process of lateral inflow along the slope, a form transformation process, a sediment adsorption process, and a sediment release process, as shown in formula (11). ; (11); Where, is the cross-sectional area of the river flow, is the nutrient solute concentration in the river, is the cross-sectional flow rate, is the dispersion coefficient along the river direction, is the source-sink term, is the sediment concentration in the river, is the lateral inflow into the river flow rate, is the pollutant concentration of the lateral inflow, is the sediment concentration of the lateral inflow, is the sediment recovery saturation coefficient, is the sediment settling rate, is the river width, is the maximum sediment-carrying capacity of the river, is the concentration of adsorbed pollutants in scoured and silted sediment; Step B2: Input the processed basic database of the target basin into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, sediment scouring-transportation-settlement process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model. Specifically, input the grid meteorological data in the processed basic database of the target basin into the runoff-sediment-nitrogen and phosphorus model as the driving force, input the point source pollution data into the river channel grid closest to the sewage outlet based on the longitude and latitude information of the sewage outlet, input the domestic pollution sources into the residential land grid, input the solid waste and livestock and poultry breeding pollution sources into the paddy field and dry land grids, input the aquaculture pollution sources into the pond water area grid, and input the atmospheric deposition data into the corresponding spatial grid. The distributed process model is calibrated in sequence according to runoff, sediment, and nutrient concentration, and the Nash-Sutcliffe efficiency coefficient NSE, percentage bias PBIAS, and ratio coefficient RSR of the root mean square error to the standard deviation of observations are used as the model accuracy measurement indicators.
[0016] Step C: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition. The specific steps are as follows. Step C1: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, where the future land use state transition matrix includes land use maps and land use type transition maps under different future scenarios. Step C2: Use the CA-Markov model to calculate the land use change result according to the land use change constraint condition. Specifically, use the Markov model to generate the land use type transition probability map and correct the land use type transition probability map with the area transition matrix relative to the base period as the constraint condition, and then use the CA model to iterate to satisfy the constraint condition and the area transition matrix and calculate the land use change.
[0017] Step D: Analyze the quantitative relationships between climate variables and socioeconomic variables and fertilization amount and livestock and poultry breeding amount, and then integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model, and tree ensemble model according to the quantitative analysis results of fertilization amount and livestock and poultry breeding amount, so as to calculate the dynamic change results of land-based pollution sources of fertilization and livestock and poultry breeding pollution. The specific steps are as follows. Step D1: Analyze the quantitative relationships between climate variables and socioeconomic variables and fertilization amount and livestock and poultry breeding amount, where the climate variables include precipitation and temperature, and the socioeconomic variables include agricultural population variables, gross domestic product GDP, and primary industry output value. Step D2: Integrate and cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model, and tree ensemble model based on the quantitative analysis results of the fertilization amount and livestock and poultry breeding amount, so as to measure the dynamic changes of the land-based pollution sources of fertilization and livestock and poultry breeding. The linear regression model includes linear regression, interactive effect linear regression, robust linear regression, and stepwise linear regression. The regression tree model includes fine tree, medium tree, and rough tree. The support vector machine model includes linear support vector machine, quadratic support vector machine, cubic support vector machine, fine Gaussian support vector machine, medium Gaussian support vector machine, and rough Gaussian support vector machine. The Gaussian process regression model includes squared exponential Gaussian process regression model, exponential Gaussian process regression model, and rational quadratic Gaussian process regression model. The tree ensemble model includes boosted tree ensemble and bagged tree ensemble. The specific steps are as follows: Step D21: Integrate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model, and tree ensemble model. Specifically, it is an ensemble learning method based on the Boosting framework and is optimized in the function space through the gradient descent method to obtain the total ensemble model. The goal of the total ensemble model is to minimize the loss function. The output of the total ensemble model is The specific steps are as follows: Step D211: The initial value is the target mean ; Step D212: Calculate the pseudo-residuals, as shown in formula (12): (12); where is the pseudo-residual, and is the output result of the current model; Step D213: Fit the residuals. Use the regression tree to fit the pseudo-residuals, and then select the splitting feature and threshold by minimizing the squared error; Step D214: Multiply the prediction result of the new tree by the learning rate and accumulate it to the current model, as shown in formula (13): (13); Step D215: Set the termination condition. If the residuals converge to the set threshold, stop sampling. Step D22: Cross-validate the linear regression model, regression tree model, support vector machine model, Gaussian process regression model, and tree ensemble model. Specifically, adopt a systematic data partitioning strategy to balance the computational efficiency and evaluation reliability. The specific steps are as follows: Step D221: Data partitioning. Randomly partition the original dataset into five equal and mutually exclusive subsets for stratified sampling and ensure that the class distribution in each subset is the same as that of the original data; Step D222: Iterative training and validation. For the first time, use the first subset as the validation set and the second to fifth subsets as the training sets. For the second time, use the second subset as the validation set and the remaining subsets as the training sets, and so on until the fifth time. An independent model is generated each time. Step D223: Performance evaluation. Use the root mean square error (RMSE) as the evaluation metric for the validation set and take the average value and standard deviation of the five results to evaluate the stability of the overall integrated model.
[0018] As Figure 3 shown, in Step E, input the processed basic database of the target watershed, the results of land use change, and the results of the dynamic changes of land-based pollution sources into the distributed process model to obtain the spatio-temporal quantitative results of the non-point source pollution load in the watershed under the changing scenario, thereby completing the dynamic calculation operation of the non-point source pollution load in the watershed.
[0019] A dynamic calculation system for the non-point source pollution load in a watershed based on a distributed process model, including a data acquisition and processing module, a distributed process model construction module, a land use change calculation module, a dynamic change calculation module for land-based pollution sources, and a dynamic calculation module for the non-point source pollution load in the watershed. The data acquisition and processing module is used to collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target watershed and construct a basic database of the target watershed, and then integrate and process different types of data in the basic database of the target watershed by using geospatial interpolation, spatial resampling, and frequency matching to form a processed basic database of the target watershed with consistent spatio-temporal scales; The distributed process model construction module is used to construct a runoff-sediment-nitrogen and phosphorus model based on physical processes, and then input the processed basic database of the target watershed into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, the sediment scouring-transportation-settlement process, and the pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model; The land use change calculation module is used to use the future land use state transfer matrix in the LUH2 dataset as the land use change constraint condition, and then calculate the land use change results by using the CA-Markov model according to the land use change constraint condition; The dynamic change calculation module for land-based pollution sources is used to analyze the quantitative relationships between climate variables and socioeconomic variables and the fertilizer application amount and livestock and poultry breeding amount, and then integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models according to the quantitative analysis results of the fertilizer application amount and livestock and poultry breeding amount to calculate the dynamic change results of land-based pollution sources of fertilizer application and livestock and poultry breeding pollution; The dynamic calculation module for non-point source pollution load in the basin is used to input the processed basic database of the target basin, the results of land use change, and the dynamic change results of land-based pollution sources into the distributed process model and obtain the spatio-temporal quantitative results of non-point source pollution load in the basin under the changed scenario, so as to complete the dynamic calculation operation of non-point source pollution load in the basin.
[0020] In summary, for the method and system for dynamically calculating the non-point source pollution load in a watershed based on a distributed process model of the present invention, first, meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target watershed are collected to construct a basic database of the target watershed. Then, different types of data in the basic database of the target watershed are integrated and processed by using geospatial interpolation, spatial resampling, and frequency matching to form a processed basic database of the target watershed with consistent spatio-temporal scales. Next, a runoff-sediment-nitrogen and phosphorus model based on physical processes is constructed. Then, the processed basic database of the target watershed is input into the runoff-sediment-nitrogen and phosphorus model, and the river channel grid runoff process, sediment erosion-transport-settlement process, and pollutant concentration transport-diffusion-transformation process are obtained to obtain a distributed process model. Subsequently, the future land use state transition matrix in the LUH2 dataset is used as a land use change constraint condition, and the CA-Markov model is used to calculate the land use change result according to the land use change constraint condition. Then, the quantitative relationships between climate variables and socio-economic variables and fertilizer application rates and livestock and poultry breeding amounts are analyzed. Then, based on the quantitative analysis results of fertilizer application rates and livestock and poultry breeding amounts, linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models are integrated and cross-validated to calculate the dynamic change results of land-based pollution sources of fertilizer application and livestock and poultry breeding. Finally, the processed basic database of the target watershed, the land use change result, and the dynamic change result of land-based pollution sources are input into the distributed process model, and the spatio-temporal quantitative results of the non-point source pollution load in the watershed under changing scenarios are obtained to complete the dynamic calculation operation of the non-point source pollution load in the watershed; effectively realizing that the method and system for dynamically calculating the non-point source pollution load in the watershed have the function of coupling the runoff-sediment-nitrogen and phosphorus physical processes on the basis of a distributed hydrological model and combining socio-economic development driving factors, land use changes, and dynamic changes in land-based pollution sources to perform spatio-temporal quantitative calculations of the non-point source pollution load in the watershed under changing scenarios. Moreover, a unified input database with spatio-temporal resolution is obtained through spatial interpolation and frequency matching. By coupling the runoff-sediment-nitrogen and phosphorus physical processes on the basis of a distributed hydrological model, the interface mismatch between the existing hydrological model and the water environment model and the model connection problem caused by the independence of the input and output processes are avoided. At the same time, through the ensemble machine learning method, not only can the dynamic changes of the prediction input data be better captured, but also the separation of socio-economic development driving factors from land use changes and dynamic changes in land-based pollution sources is avoided. This provides a new method for accurately and quantitatively calculating the spatio-temporal changes of the non-point source pollution load in the watershed under changing scenarios and is also of great significance for guiding the refined management of watershed water resources and water environment.
[0021] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A dynamic calculation method for non-point source pollution load in a watershed based on a distributed process model, characterized in that: including the following steps: Step A: Collect data on meteorology, hydrology, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution sources in the target basin, and construct a basic database for the target basin. Then, use geospatial interpolation, spatial resampling, and frequency matching to integrate and process different types of data in the basic database of the target basin, so as to form a processed basic database of the target basin with consistent spatio-temporal scales; Step B: Construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. Then, input the processed basic database of the target basin into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, sediment erosion-transport-settlement process, and pollutant concentration transport-diffusion-transformation process, so as to obtain a distributed process model; Step C: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition; Step D: Analyze the quantitative relationships between climate variables and socioeconomic variables and fertilizer application rates and livestock and poultry breeding amounts. Then, integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models according to the quantitative analysis results of fertilizer application rates and livestock and poultry breeding amounts, so as to calculate the dynamic change results of land-based pollution sources of fertilizer application and livestock and poultry breeding pollution; Step E: Input the processed basic database of the target basin, land use change results, and dynamic change results of land-based pollution sources into the distributed process model to obtain spatio-temporal quantitative results of non-point source pollution loads in the basin under changing scenarios, thus completing the dynamic calculation of non-point source pollution loads in the basin.
2. The dynamic measurement method for non-point source pollution load in a basin based on a distributed process model according to claim 1, characterized in that: Step A: Collect data on meteorology, hydrology, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution sources in the target basin, and construct a basic database for the target basin. Then, use geospatial interpolation, spatial resampling, and frequency matching to integrate and process different types of data in the basic database of the target basin, so as to form a processed basic database of the target basin with consistent spatio-temporal scales. The specific steps are as follows: Step A1: Collect data on meteorology, hydrology, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution sources in the target basin, and construct a basic database for the target basin. Among them, the meteorology, hydrology, water quality, soil and terrain, land use, vegetation cover, and atmospheric deposition data of the target basin are remote sensing grid data, the surface pollution source data is annual statistical data, the basic database of the target basin also includes meteorological data of future climate scenarios and atmospheric deposition scenario data of the future. The meteorological data of the future climate scenario is spatial grid data output by GCM, and the atmospheric deposition scenario data of the future uses the atmospheric deposition spatial grid data provided by the MRI-ESM2-0 model; Step A2: Use geospatial interpolation, spatial resampling, and frequency matching to integrate and process different types of data in the basic database of the target basin, so as to form a processed basic database of the target basin with consistent spatio-temporal scales. The specific steps are as follows: Step A21: Integrate and process different types of data in the target watershed basic database using geospatial interpolation. Specifically, the geospatial interpolation uses the Inverse Distance Weighting (IDW) method to interpolate the station data into the space according to the spatial resolution of the remote sensing grid data and then enter it into the target watershed basic database. Step A22: Integrate and process different types of data in the target watershed basic database using spatial resampling. Specifically, the spatial resampling uses the Majority type to unify the spatial resolutions of the soil, terrain, land use, vegetation cover, and atmospheric deposition grid data. Step A23: Integrate and process different types of data in the target watershed basic database using frequency matching. Specifically, the frequency matching is applied to the meteorological data output by the GCM and uses the cumulative frequency curve of the measured climate data to correct the bias of the cumulative frequency curve of the GCM simulation values, so that the cumulative frequency curves of the historical measured data are consistent, and then the calibrated data can be used as the input data for future scenarios and entered into the basic database.
3. A dynamic measurement method for non-point source pollution load in a watershed based on a distributed process model according to claim 2, characterized in that: Step B: Construct a runoff-sediment-nitrogen and phosphorus model based on physical processes, and then input the processed target watershed basic database into the runoff-sediment-nitrogen and phosphorus model to obtain the channel grid runoff process, sediment scour-transport-settlement process, and pollutant concentration transport-diffusion-transformation process, thereby obtaining a distributed process model. The specific steps are as follows. Step B1: Construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. The runoff-sediment-nitrogen and phosphorus model uses the GBNP model, and the GBNP model includes a hydrological module, a soil erosion module, and a nitrogen and phosphorus load module. The specific steps are as follows. Step B11: Construct the hydrological module of the GBNP model. The hydrological module extracts the watershed water system based on the distributed hydrological model GBHM using the Digital Elevation Model (DEM) and divides the hillslope-channel grid units as the basic units of the model. Then, calculate the precipitation interception, evapotranspiration, soil water movement, and groundwater outflow hydrological processes in the hillslope unit. The specific steps are as follows. Step B111: Calculate the precipitation interception hydrological process. Specifically, assume that the environmental surface is covered with vegetation, and the precipitation interception hydrological process is shown in Equation (1). (1); Among them, is the actual transpiration rate of soil water in the layer from the vegetation root system to the leaf surface at time is the vegetation coverage rate, is the crop coefficient, is the potential evaporation, is the plant root depth distribution function, is the soil water content function, is the leaf area index at time is the maximum value of the leaf area index of the vegetation in a year; Step B112: Calculate the evapotranspiration hydrological process, specifically including the situation where the environmental surface has no vegetation cover and the situation where the environmental surface water accumulation does not meet the evaporation capacity. The specific steps are as follows. Step B1121: Assume that the environmental surface has no vegetation cover, then the evapotranspiration hydrological process is shown in Equation (2). (2); Among them, is the actual evaporation rate of the land surface at time is the water depth of the surface water at time is the time step; Step B1122: Assume that the environmental surface water accumulation does not meet the evaporation capacity, then the evapotranspiration hydrological process is shown in Equation (3). (3); Among them, is the actual evaporation rate of the soil surface at a given time, which is a function of soil moisture content; Step B113: Calculate the soil water movement hydrological process, specifically including the soil unsaturated hydraulic conductivity, the soil water movement process in the unsaturated zone, and the soil water characteristic curve. The specific steps are as follows. Step B1131: Calculate the soil unsaturated hydraulic conductivity, as shown in Equation (4). (4); Among them, is the unsaturated hydraulic conductivity of the soil, is the saturated hydraulic conductivity of the soil, is the water content, is the fitting parameter of the soil water characteristic curve; Step B1132: Calculate the soil water movement process in the unsaturated zone. Specifically, use the one-dimensional Richard equation to calculate the soil water movement in the unsaturated zone, and the soil water movement process in the unsaturated zone is shown in Equation (5). (5); Among them, is the soil depth, is the water content of the soil at the soil depth at time , is the soil water flux, is the soil evaporation rate at time is the soil water content function; Step B1132: Calculate the soil water characteristic curve. Specifically, the soil water characteristic curve represents the relationship between soil water content and soil suction, and the soil water characteristic curve is as shown in Equation (6). (6); wherein, is the residual water content of the soil, is the saturated water content of the soil, and are both fitting parameters of the soil water characteristic curve; Step B114: Calculate the groundwater outflow hydrograph, which specifically includes the water volume exchange process between the hillslope phreatic aquifer and the river channel and the river channel confluence process. The specific steps are as follows. Step B1141: Calculate the water volume exchange process between the hillslope phreatic aquifer and the river channel, as shown in Equation (7). (7); Among them, is the change rate of the groundwater storage in the saturated aquifer with time, is the recharge rate between the saturated aquifer and the upper unsaturated aquifer, is the downward leakage rate of the saturated aquifer, is the unit-width exchange flow of groundwater and the river channel, is the area of the hillslope unit, is the saturated hydraulic conductivity of the phreatic aquifer, and are the groundwater levels in the phreatic aquifer before and after water volume exchange, and are the river channel water levels before and after water volume exchange, is the length of the hillslope unit; Step B1142: Calculate the river channel confluence process, as shown in Equation (8). (8); Among them, is the total river confluence, is the lateral inflow of the river, is the cross-sectional area of the river, is the distance along the river direction, is the river slope, is the roughness coefficient of the river, is the wetted perimeter length; Step B12: Construct the soil erosion module and nitrogen and phosphorus load module of the GBNP model, which specifically includes the nutrient slope process, transformation process, and river channel process. The specific steps are as follows. Step B121: Calculate the nutrient slope process, where the nutrient slope process is specifically the accumulation-scouring process, as shown in Equation (9). , (9); Among them, is the dust accumulation amount, is the proportionality coefficient of pollutants in the dust, is the scouring coefficient, is the runoff intensity, is the maximum accumulation amount, is the number of days of dust accumulation after the previous rainfall, is the number of days required for the dust to accumulate to half of the maximum accumulation amount; Step B122: Calculate the nutrient transformation process, where the nutrient transformation process is specifically the convection-diffusion process, and the convection-diffusion process includes the processes of nutrient adsorption, desorption, volatilization, and transformation. The compound form transformation process is described by a first-order kinetic reaction equation, as shown in Equation (10). ; ; ; (10); wherein, is the concentration of the adsorption term, is the moisture migration velocity, is the hydrodynamic dispersion coefficient, is the solute source-sink term, is the equivalent soil surface nitrogen content of the fertilization amount, is the inorganic phosphorus content, is the conversion amount of different forms, is the plant uptake amount, is the loss amount caused by runoff, is the conversion coefficient, is the nitrogen-phosphorus uptake ratio of plants at different growth stages, is the proportion of the cumulative growth days of the day in the growth period; Step B123: Calculate the nutrient river channel process, where the nutrient river channel process specifically includes the convection-diffusion process, the process of lateral inflow along the slope, the form transformation process, the sediment adsorption process, and the sediment release process, as shown in Equation (11). ; (11); Among them, is the cross-sectional area of the river channel for flow through, is the concentration of nutrient solutes in the river channel, is the cross-sectional flow rate, is the dispersion coefficient along the river direction, is the source-sink term, is the sediment concentration in the river channel, is the lateral inflow into the river channel flow rate, is the concentration of pollutants in the lateral inflow, is the sediment concentration in the lateral inflow, is the sediment recovery saturation coefficient, is the sediment settling rate, is the river channel width, is the maximum sediment-carrying capacity of the river, is the concentration of pollutants adsorbed in the scoured and silted sediment; Step B2: Input the processed target watershed basic database into the runoff-sediment-nitrogen and phosphorus model to obtain the river channel grid runoff process, the sediment scour-transport-settlement process, and the pollutant concentration transport-diffusion-transformation process, so as to obtain a distributed process model. Specifically, input the grid meteorological data in the processed target watershed basic database into the runoff-sediment-nitrogen and phosphorus model as the driving force, input the point source pollution data based on the longitude and latitude information of the sewage outfall into the nearest river channel grid, input the domestic pollution source into the residential land grid, input the solid waste and livestock and poultry breeding pollution sources into the paddy field and dry land grids, input the aquaculture pollution source into the pond water area grid, and input the atmospheric deposition data into the corresponding spatial grid. The distributed process model is calibrated in sequence according to runoff, sediment, and nutrient concentration, and the Nash efficiency coefficient NSE, the percentage bias PBIAS, and the ratio coefficient RSR of the root mean square error to the standard deviation of the observation are used as the model accuracy measurement indicators.
4. A dynamic measurement method for non-point source pollution load in a watershed based on a distributed process model according to claim 3, characterized in that: Step C: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition. The specific steps are as follows. Step C1: Use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, where the future land use state transition matrix includes the future land use maps and land use type transition maps under different scenarios. Step C2: Use the CA-Markov model to calculate the results of land use change according to the land use change constraint conditions. Specifically, use the Markov model to generate the land use type transfer probability map and correct the land use type transfer probability map with the area transfer matrix relative to the base period as the constraint condition, and then use the CA model to iterate to satisfy the constraint condition and the area transfer matrix and calculate the land use change.
5. A dynamic calculation method for non-point source pollution load in a watershed based on a distributed process model according to claim 4, characterized in that: Step D: Analyze the quantitative relationships between climate variables and socioeconomic variables and the fertilization amount and livestock and poultry breeding amount, and then integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models based on the quantitative analysis results of the fertilization amount and livestock and poultry breeding amount, so as to calculate the dynamic change results of land-based pollution sources of fertilization and livestock and poultry breeding. The specific steps are as follows. Step D1: Analyze the quantitative relationships between climate variables and socioeconomic variables and the fertilization amount and livestock and poultry breeding amount, where the climate variables include precipitation and temperature, and the socioeconomic variables include agricultural population variables, gross domestic product (GDP), and the output value of the primary industry. Step D2: Integrate and cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models based on the quantitative analysis results of the fertilization amount and livestock and poultry breeding amount, so as to calculate the dynamic change of land-based pollution sources of fertilization and livestock and poultry breeding. The linear regression models include linear regression, interactive effect linear regression, robust linear regression, and stepwise linear regression. The regression tree models include fine trees, medium trees, and rough trees. The support vector machine models include linear support vector machines, quadratic support vector machines, cubic support vector machines, fine Gaussian support vector machines, medium Gaussian support vector machines, and rough Gaussian support vector machines. The Gaussian process regression models include squared exponential Gaussian process regression models, exponential Gaussian process regression models, and rational quadratic Gaussian process regression models. The tree ensemble models include boosting tree ensembles and bagging tree ensembles. The specific steps are as follows. Step D21, integrate a linear regression model, a regression tree model, a support vector machine model, a Gaussian process regression model, and a tree ensemble model. The integration is specifically an ensemble learning method based on the Boosting framework and is optimized in the function space through the gradient descent method to obtain a total ensemble model. The goal of the total ensemble model is to minimize the loss function , and the output of the total ensemble model is . The specific steps are as follows Step D211, with the initial value being the target mean value ; Step D212: Calculate the pseudo-residuals as shown in formula (12). (12); Among them, is the pseudo-residual, is the output result of the current model; Step D213, fitting the residual, using a regression tree Fitting the pseudo-residual, and then selecting the splitting feature and threshold by minimizing the squared error; Step D214, multiply the prediction result of the new tree by the learning rate and accumulate it to the current model, as shown in formula (13). (13); Step D215: Set the termination condition. If the residuals converge to the set threshold, stop sampling. Step D22: Cross-validate linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models. Specifically, use a systematic data partitioning strategy to balance computational efficiency and evaluation reliability. The specific steps are as follows. Step D221: Data partitioning. Randomly partition the original dataset into five equal and mutually exclusive subsets for stratified sampling and ensure that the class distribution in each subset is the same as that of the original data. Step D222: Iterative training and validation. First, use the first subset as the validation set and the second to fifth subsets as the training set. Second, use the second subset as the validation set and the remaining subsets as the training set, and so on until the fifth time. Each time, an independent model is generated. Step D223: Performance evaluation. Use the root mean square error (RMSE) as the evaluation index for the validation set and take the average and standard deviation of the five results to evaluate the stability of the total ensemble model.
6. A dynamic calculation system for watershed non-point source pollution load based on a distributed process model, wherein the specific calculation process of the dynamic calculation system for watershed non-point source pollution load is based on the dynamic calculation method for watershed non-point source pollution load according to any one of claims 1-5, and is characterized in that: It includes a data acquisition and processing module, a distributed process model construction module, a land use change calculation module, a terrestrial pollution source dynamic change calculation module, and a watershed non-point source pollution load dynamic calculation module. The data acquisition and processing module is used to collect meteorological, hydrological, water quality, soil and terrain, land use, vegetation cover, atmospheric deposition, and surface pollution source data of the target watershed and construct a basic database of the target watershed. Then, geospatial interpolation, spatial resampling, and frequency matching are used to integrate and process different types of data in the basic database of the target watershed, so as to form a processed basic database of the target watershed with consistent spatio-temporal scales; The distributed process model construction module is used to construct a runoff-sediment-nitrogen and phosphorus model based on physical processes. Then, the processed basic database of the target watershed is input into the runoff-sediment-nitrogen and phosphorus model, and the river channel grid runoff process, sediment scour-migration-settlement process, and pollutant concentration transport-diffusion-transformation process are obtained, so as to obtain a distributed process model; The land use change calculation module is used to use the future land use state transition matrix in the LUH2 dataset as the land use change constraint condition, and then use the CA-Markov model to calculate the land use change result according to the land use change constraint condition; The terrestrial pollution source dynamic change calculation module is used to analyze the quantitative relationship between climate variables and socio-economic variables and the fertilization amount and livestock and poultry breeding amount. Then, according to the quantitative analysis results of the fertilization amount and livestock and poultry breeding amount, linear regression models, regression tree models, support vector machine models, Gaussian process regression models, and tree ensemble models are integrated and cross-validated, so as to calculate the dynamic change results of terrestrial pollution sources of fertilization and livestock and poultry breeding pollution; The watershed non-point source pollution load dynamic calculation module is used to input the processed basic database of the target watershed, the land use change result, and the terrestrial pollution source dynamic change result into the distributed process model and obtain the spatio-temporal quantitative results of the watershed non-point source pollution load under the change scenario, so as to complete the dynamic calculation operation of the watershed non-point source pollution load.
Citation Information
Patent Citations
Non-point source pollution calculation method based on remotely sensed image element
CN102867120A
Distributed simulation method of non-point source nitrogen and phosphorus loss form constitution in hilly regions
CN107066808A
Compilation method of surface water basin pollution source emission list (dynamic list)
CN110765213A
Land utilization pattern prediction and optimization method considering non-point source pollution control
CN114021829A
Non-point source pollution control method
CN116993235A
Cited By
Optimized regulation and control method for land utilization of multi-outlet watershed
CN120654901A
A method for optimizing land use in multi-outlet watersheds
CN120654901B
Planting and breeding collaborative optimization method based on spatial distribution characteristic analysis
CN122222145A