Agricultural non-point source pollution load assessment method based on XGBoost algorithm
By adopting the S-APPI model integrated based on XGBoost algorithm and APPI model in Pingyuanhe.com area, the problem that the existing technology is difficult to accurately evaluate agricultural non-point source pollution in the region is solved, and more efficient and accurate pollution load evaluation and simulation results are achieved.
Patent Information
- Application Number
- CN202411780608.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-02
AI Technical Summary
The existing agricultural non-point source pollution assessment method is difficult to accurately quantify the emission characteristics of regional pollutants in plain river network areas, especially in areas without data, which leads to deviations from the actual situation.
The agricultural non-point source pollution assessment model (S-APPI model) based on the integrated XGBoost algorithm and APPI model is adopted to learn laws and knowledge from a large amount of data through machine learning methods, improve prediction accuracy, and adapt to the characteristics of the plain river network area.
It improves the accuracy of agricultural non-point source pollution load assessment, can effectively simulate the generation, migration and transformation process of pollutants, reduces dependence on a large amount of measured data, reduces the cost and difficulty of data acquisition, and improves the calculation efficiency and adaptability of the model.
Smart Images

Figure CN119918984A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural non-point source pollution assessment, and in particular to an agricultural non-point source pollution load assessment method based on an XGBoost algorithm. Background Art
[0002] With the rapid development of agricultural modernization, agricultural non-point source pollution has become a key factor affecting water environment quality and ecological security. In plain river network areas, the polder area is usually used as a water conservancy management unit. The farmland and water system in the polder area are complex, and the water level (flow) is affected by tidal action, pump irrigation and drainage, and rainfall runoff. The way farmland drains into the river is diverse, and the impact process is particularly complex.
[0003] Existing assessment methods for agricultural non-point source pollution mostly rely on empirical models or physical models, which have limitations in data acquisition, processing and accuracy. Especially in plain river network areas with flat terrain and dense water networks, these methods are difficult to accurately quantify the emission characteristics of regional pollutants, and the assessment results deviate from the actual situation. For example, although the empirical model is easy to operate, its accuracy is limited; although the physical model has good simulation effects, it has high requirements for measured data and its application in plain river network areas is limited, especially in some areas where there is no data. Therefore, there is an urgent need for a new assessment model that can adapt to the characteristics of plain river network areas and accurately predict the status of agricultural non-point source pollution. At present, there is no report on the use of XGBoost algorithm to assess agricultural non-point source pollution load. Summary of the invention
[0004] In view of the technical difficulties in estimating agricultural non-point source pollution load in the plain river network area mentioned above, this application proposes an agricultural non-point source pollution assessment model (S-APPI model) based on the integration of XGBoost algorithm and APPI model. Through machine learning methods, by learning laws and knowledge from a large amount of data, the prediction accuracy is improved, the problem that traditional mechanism models cannot be applied in areas without data is solved, and nonlinear processes that cannot be explained by traditional evaluation models are fitted.
[0005] Specifically, the above invention object is achieved through the following technical solutions:
[0006] First, this application provides an agricultural non-point source pollution load assessment method based on the XGBoost algorithm, and its specific steps are as follows:
[0007] Step 1: Select the area to be evaluated and obtain the basic data required to build the model;
[0008] This step first needs to determine the specific area to be evaluated, and construct the main primary and secondary indicators of the basic data of the model. Their sources are all public (official) data; the primary indicators include six calculation indicators, namely surface runoff, sediment loss, planting loss load, aquaculture load, livestock and poultry breeding load, and rural domestic sewage load; the secondary indicators are the data required to calculate the six primary indicators, mainly a series of attribute data, including meteorological data, soil physical attribute data, agricultural production activities and other indicators; among them, meteorological data mainly include annual precipitation and monthly precipitation; soil physical attribute data include soil particle size distribution, soil organic carbon content, soil slope, and vegetation coverage; agricultural production activities include land use type, fertilizer application category / application amount, agricultural planting category / area / loss coefficient, livestock and poultry category / stock and output quantity / livestock and poultry excretion coefficient / livestock and poultry manure utilization rate, and aquatic product category / output / pollution coefficient / aquaculture wastewater treatment rate / treatment efficiency; other indicators include runoff curve number, rural permanent population / domestic sewage generation coefficient / domestic sewage treatment rate / treatment efficiency.
[0009] Step 2: Correct the crop loss coefficient and complete the construction of the basic data set of the S-APPI model:
[0010] 2.1) The original APPI model system includes four calculation indexes, namely, Runoff Index (RI), Sediment Production Index (SPI), Cultivating Loading Index (CLI), and People and Animal Loading Index (PALI), which involve relevant agricultural non-point source pollution parameter categories such as surface runoff, soil erosion, land use, agricultural planting, aquatic livestock and poultry breeding, and rural domestic sewage.
[0011] APPI i =RI i WF1+SPI i WF2+CLI i WF3+PALI i WF4 (1)
[0012] Where: APPI i is the comprehensive index of agricultural non-point source pollution potential in region i; RI i is the runoff index of region i, representing the surface runoff capacity of the study area; SPI i is the sediment generation index of region i, representing the sediment loss capacity of the study area; CLI iis the planting load index of region i, representing the contribution of agricultural planting to non-point source pollution load in the study area; PALI i is the human and livestock load index of region i, which evaluates the contribution of human and livestock pollution in the study area; WF is the corresponding weight of each index.
[0013] The basic model formula (1) is disclosed in the document “Application of Agricultural Non-point Source Pollution Potential Index System (APPI) in Typical Areas of Taihu Lake, Zhou Xuhai et al., Journal of Agricultural Environment Science, 2006”.
[0014] 2.2) The entropy weight method is introduced to correct the spatial heterogeneity of the pollution emission coefficient of the planting industry. By weighting the fertilizer application amount, pesticide application amount, cultivated land area and garden area of each unit in the statistical data, the weights of the four attribute indicators of fertilizer use index, pesticide application index, crop planting index and garden planting index are determined to ensure that the model can accurately reflect the pollution characteristics of the planting industry in different regions and improve the accuracy of load estimation. The calculation formula of the planting industry correction parameter is as follows:
[0015] Planting loss correction parameter = fertilizer use index × W1 + pesticide application index × W2 + crop planting index × W3 + garden planting index × W4 (2)
[0016] Where: the crop loss correction parameter is the correction value (dimensionless) of the crop nitrogen and phosphorus emission (loss) coefficient, ranging from 0 to 1; W1-W4 are the corresponding weights of each index.
[0017] The reference crop emission (loss) coefficients of each province provided in the Manual of Agricultural Source Production and Emission Accounting Methods and Coefficients are used to calculate the median value of the crop loss correction parameter of each unit's crop nitrogen and phosphorus emission (loss) coefficient, and the crop emission (loss) coefficient is corrected proportionally.
[0018] 2.3) After correcting the coefficient of crop production, the six indices required to construct the S-APPI model are calculated using the following formulas:
[0019] The surface runoff calculation formula is:
[0020]
[0021] Where Q is the runoff depth (mm); P is the rainfall depth (mm); S is the potential maximum retention or infiltration (mm); CN is the CN value (dimensionless) of different land use types.
[0022] The calculation formula for sediment loss is:
[0023] A=R×K×L×S×C×P (5)
[0024] Where A is the soil loss per unit area (t·ha- 1 ·yr- 1 )R is the rainfall and erosibility factor, unit is MJ·mm·ha- 1 ·h- 1 ·yr- 1 ; K is the soil erodibility factor, unit t·ha·h·ha- 1 ·MJ- 1 mm- 1 ; L is the slope length factor, S is the slope steepness factor, C is the coverage and management factor, and P is the protection practice factor;
[0025] Plant loss load:
[0026] Nitrogen and phosphorus loss from crop production (tons) = {crop planting area (hectares) × crop planting process loss coefficient (kg / hectare) × crop loss correction parameter + garden area (hectares) × garden loss coefficient (kg / hectare) × crop loss correction parameter} × {fertilizer usage per unit area for crop production in the survey year / 2017 (kg / hectare)} ÷ 1000 (6)
[0027] Rural domestic sewage load:
[0028] Rural domestic sewage pollutant discharge (tons) = rural permanent population (10,000 people) × pollutant generation intensity (g / person / day) × 365 (days) ÷ 100 × (1-rural domestic sewage treatment rate (%) × rural pollutant comprehensive removal rate (%)) (7)
[0029] Rural domestic sewage treatment rate (%) = rural domestic sewage design treatment scale (tons / day) × 80% ÷ (rural permanent population (person) × rural domestic sewage discharge coefficient (liters / person / day) ÷ 1000) (8)
[0030] Livestock and poultry breeding load:
[0031] Livestock and poultry breeding load (tons) = scale breeding wastewater discharge + breeder breeding wastewater discharge (tons) (9)
[0032] Pollution discharge from large-scale farming (tons) = {stock on hand + number of animals slaughtered in the current year (heads)} × scale farming ratio × livestock and poultry pollution discharge coefficient (kg / head) × (1 – scale farming resource utilization rate (%) × utilization efficiency (%)) ÷ 1000 (10)
[0033] Pollution discharge of livestock and poultry farmers = {stock on hand + number of livestock and poultry sold in the current year (heads)} × livestock and poultry farming ratio × livestock and poultry pollution discharge coefficient (kg / head) ÷ 1000 (11)
[0034] Aquaculture load:
[0035] Aquaculture load (tons) = aquatic product volume (tons) × aquaculture discharge coefficient (kg / ton) ÷ 1000 (12);
[0036] Step 3: Use the acquired basic data to train and verify the model so that the model meets the evaluation criteria;
[0037] The long-term and relatively continuous data series for training and verifying the model results are mainly: the six indicators of surface runoff, sediment loss, crop loss load, aquaculture load, livestock and poultry breeding load, and rural domestic sewage load are used as the driving data set for model training; the water quality monitoring data of chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus are used as the response data set.
[0038] 3.1) The XGBoost algorithm is used to build an integrated learning model. The negative gradient of the model is approximated by the high-order Taylor expansion of the loss function, and it is used as the residual of the previous model for learning, thereby realizing the serial iteration of multiple regression trees and improving the prediction accuracy and generalization ability of the model. The calculation formula is:
[0039]
[0040] In the formula, is the model prediction value, i is the number of samples, k is the number of integrated regression trees, x i is the data sample, F is the set space formed by the regression tree, f k Represents one of the regression trees.
[0041] 3.2) Combined with the traditional APPI model, a model-driven factor system with six specific indicators, including surface runoff, sediment loss, crop loss load, aquaculture load, livestock and poultry breeding load, and rural domestic sewage load, was constructed as the driving data set for model training. Water quality monitoring data was used as the response data set, and the model was trained and verified using the crop discharge load before correction.
[0042] 3.3) Use the corrected crop emission (loss) coefficient to simulate agricultural non-point source load. After the corrected crop pollutant loss load is calculated, the model is retrained and verified according to the constructed S-APPI model data set format. The trained model is used to simulate water quality, and the water quality prediction results are compared with the water quality monitoring data to complete the model simulation accuracy output.
[0043] 3.4) The model was validated using long-term water quality monitoring data and the coefficient of determination R 2 The root mean square error (RMSE) was used as the evaluation parameters, and chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus were selected as the water quality parameter indicators for the model simulation. 2Used to evaluate the model fitting effect, R 2 The value range is between 0 and 1. The closer the value is to 1, the better the fitting effect is. Otherwise, the fitting accuracy is not high. RMSE can compare the difference between the true value and the predicted value. The smaller the value, the better the prediction effect of the model. 2 The root mean square error RMSE calculation formula is as follows:
[0044]
[0045]
[0046] in, Indicates the actual value, Represents the predicted value, y a Indicates the average of actual values;
[0047] When the following is met during calibration and verification: 2 The root mean square error (RMSE) is not less than 0.6 and not higher than 0.07, and the coefficient of determination R between the training set and the validation set is 2 When the difference between the root mean square error (RMSE) and the training result is within 0.1, it indicates that the S-APPI model has good simulation effect in the selected watershed and high credibility, and the trained S-APPI model is obtained.
[0048] In this step, after the model accuracy is verified through machine learning and measured water quality data, the trained S-APPI model can be used to evaluate the agricultural non-point source pollution load in the target area.
[0049] Step 4: Input the basic data collected in step 1 into the trained S-APPI model obtained in step 3.2), and use the modified crop loss coefficient combined with the output coefficient method to calculate the pollution load of each research unit, thus completing the target agricultural non-point source pollution load assessment.
[0050] Secondly, the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method steps for assessing agricultural non-point source pollution load based on the XGBoost algorithm are implemented.
[0051] Third, the present application provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned agricultural non-point source pollution load assessment method steps based on the XGBoost algorithm are implemented.
[0052] Compared with the prior art, the model provided by the present application based on the XGBoost algorithm is suitable for the assessment of agricultural non-point source pollution load in plain river network areas, and can effectively simulate the generation, migration and transformation processes of pollutants, thereby improving the accuracy of the assessment; through the application of machine learning algorithms, the dependence on a large amount of measured data is reduced, the cost and difficulty of data acquisition are reduced, and the computational efficiency of the model is improved; the integrated learning and hyperparameter optimization techniques of the model make the S-APPI model advantageous in dealing with nonlinear relationships and complex data structures, and can better adapt to changes in different regions and conditions; the application of the entropy weight method enables the model to be adjusted according to the specific conditions of the study area, improves the adaptability and flexibility of the model, and makes the estimation results more in line with reality; it is helpful to achieve effective prevention and control of agricultural non-point source pollution, and promote green agricultural development and improvement of water environment quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a schematic diagram of the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2022;
[0054] Figure 2 This is a schematic diagram of the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2021;
[0055] Figure 3 This is a schematic diagram of the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2020;
[0056] Figure 4 This is a schematic diagram of the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2019;
[0057] Figure 5 This is a schematic diagram of the distribution of total phosphorus, ammonia nitrogen, total nitrogen and COD loads from agricultural non-point sources in Jiangsu Province in 2018;
[0058] Figure 6 This is a schematic diagram of the distribution of total phosphorus, ammonia nitrogen, total nitrogen and COD loads from agricultural non-point sources in Jiangsu Province in 2017;
[0059] Figure 7 The pollution trends of rural domestic sewage, livestock and poultry breeding, aquaculture, and crop farming in Jiangsu Province from 2017 to 2022;
[0060] Figure 8 This is the changing trend of agricultural non-point source pollution in Jiangsu Province from 2017 to 2022. DETAILED DESCRIPTION
[0061] The S-APPI model of the present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0062] Example 1
[0063] The agricultural non-point source pollution load assessment method based on the XGBoost algorithm in this embodiment is implemented in the following steps:
[0064] Step 1: Select the area to be evaluated and obtain the basic data required to build the model;
[0065] Jiangsu Province is a typical plain river network area, with dense and crisscross river networks, gentle hydraulic slopes, serious lack of hydrodynamics, and weak water self-purification capacity. At the same time, the water flow is affected by tides and the hydrological conditions are relatively complex. Rainfall runoff is the main driving force for agricultural non-point source pollution emissions and the carrier of non-point source pollution load transfer.
[0066] This embodiment conducts load assessment on agricultural non-point source pollution in the plain river network area of Jiangsu Province, and selects six specific indicators of surface runoff, sediment loss, planting loss load, aquaculture load, livestock and poultry breeding load, and rural domestic sewage load as the input features of the model to construct a water quality prediction model. Among them, the indicator system takes into account natural factors such as rainfall, soil, vegetation cover and terrain, as well as human factors such as agricultural planting, livestock and poultry aquaculture and population. The data required for model calculation include land use, annual vegetation cover, digital elevation map, monthly precipitation, soil particle size distribution, fertilizer application, aquatic livestock and poultry breeding and rural permanent population, etc. The specific data sources are shown in Table 1.
[0067] Table 1 Main data sources
[0068]
[0069] The four water quality parameters (total nitrogen, total phosphorus, ammonia nitrogen, and chemical oxygen demand) of the water quality monitoring data are used as output labels. The water quality dataset used is from the Jiangsu Provincial Environmental Monitoring Center. The water quality data includes water quality monitoring data from 373 monitoring points across Jiangsu Province (river, lake and reservoir monitoring data in rural counties of Jiangsu Province, agricultural non-point source control section monitoring data, and farmland water withdrawal monitoring section monitoring data).
[0070] After data collection is completed, it is cleaned and preprocessed, including removing redundant features of missing data exceeding 40% or repeated data, to form a complete machine learning training and test data set. The data set of each indicator is constructed according to the time and space sequence, that is, the values of each indicator are arranged in time and space order so that they correspond to the time and space sequence of water quality monitoring data one by one, and saved in the corresponding data format to construct the modeling data set.
[0071] Step 2: Correct the loss coefficient of the planting industry and complete the construction of the basic data set of the S-APPI model;
[0072] The original APPI model is:
[0073] APPI i =RI i WF1+SPI i WF2+CLI i WF3+PALI i WF4 (1)
[0074] Where: APPI i is the comprehensive index of agricultural non-point source pollution potential in region i; RI i is the runoff index of region i, representing the surface runoff capacity of the study area; SPI i is the sediment generation index of region i, representing the sediment loss capacity of the study area; CLI i is the planting load index of region i, representing the contribution of agricultural planting to non-point source pollution load in the study area; PALI i is the human and livestock load index of region i, which evaluates the contribution of human and livestock pollution in the study area; WF is the corresponding weight of each index.
[0075] This embodiment introduces the entropy weight method to correct the model. The specific steps are as follows:
[0076] Use the information entropy method to determine the following decision matrix f j (x i ), the weights of the four attribute indicators of fertilizer use index, pesticide application index, crop planting index and garden planting index, where the row title represents the evaluation object i and the column title represents the attribute indicator j. The calculated weights are shown in Table 2.
[0077] Table 2 Index weights of driving factors in different counties
[0078] index Pesticide application index Fertilizer use index Crop Planting Index Garden Planting Index Pukou District 17.993 2772 12.79 2828.00 Jiangning District 309.626 7665 66.47 8702.00 Liuhe District 246.746 20400 87.63 10256.00 Lishui District 296.47 8037 46.82 14048.00 Gaochun District 222.491 7735 29.65 6150.00 … … … … … Suyu District 409.384 25548 74.09 2850.00 Shuyang County 1823.279 135591 252.86 3240.00 Siyang County 494.505 31923 120.32 8050.00 Sihong County 2025.206 87986 201.51 18672.00
[0079] And list the original data as a matrix, namely:
[0080]
[0081] First, according to the entropy weight method formula The probabilities of each item are calculated and shown in Table 3.
[0082] Table 3 Weight probability of each driving factor index in different counties
[0083] index Pesticide application index Fertilizer use index Crop Planting Index Garden Planting Index Pukou District 0.00031 0.001163 0.001775 0.006756 Jiangning District 0.00539 0.003215 0.009227 0.020787 Liuhe District 0.00429 0.008556 0.012164 0.0245 Lishui District 0.00516 0.003371 0.006499 0.033558 Gaochun District 0.00387 0.003244 0.004116 0.014691 … … … … … Suyu District 0.00712 0.010715 0.010284 0.006808 Shuyang County 0.03172 0.056869 0.035099 0.00774 Siyang County 0.0086 0.013389 0.016701 0.01923 Sihong County 0.03523 0.036903 0.027971 0.044604
[0084] Secondly, according to the formula And k = 1 / ln(m) (18), we can get the entropy value of each attribute index j:
[0085] Table 4 Entropy values of various attribute indicators
[0086] Pesticide application index Fertilizer use index Crop Planting Index Garden Planting Index <![CDATA[E j ]]> 0.86792 0.854091 0.879296 0.80542
[0087] Note: m is the number of attribute index j; E1, E2, E3, and E4 represent the entropy values of fertilizer use index, pesticide application index, crop planting index, and garden planting index, respectively.
[0088] Finally, according to formula d j =1-E j (19) and Calculate the entropy weight:
[0089] Table 5 Entropy weight of each attribute index
[0090] Pesticide application index Fertilizer use index Crop Planting Index Garden Planting Index <![CDATA[d j ]]> 0.13208 0.145909 0.120704 0.19458 <![CDATA[W j ]]> 0.223 0.246 0.203 0.328
[0091] Note: d1, d2, d3 and d4 represent the d values of fertilizer use index, pesticide application index, crop planting index and garden planting index respectively; W1, W2, W3 and W4 represent the entropy weights of fertilizer use index, pesticide application index, crop planting index and garden planting index respectively.
[0092] According to the formula for crop industry loss correction parameters, and referring to the crop industry emission (loss) coefficient of Jiangsu Province in the Manual of Agricultural Source Production Pollution Accounting Methods and Coefficients, the median value of the crop industry loss correction parameters was taken, and the crop industry emission (loss) coefficients of 64 districts and counties were calculated proportionally. Then, the corrected crop industry emission load was calculated as the input data for subsequent algorithms.
[0093] Step 3: Calculation of other indicator data
[0094] The aquaculture load, livestock and poultry farming load, and rural domestic sewage load are calculated according to the following formula:
[0095] Rural domestic sewage load:
[0096] Rural domestic sewage pollutant discharge (tons) = rural permanent population (10,000 people) × pollutant generation intensity (g / person / day) × 365 (days) ÷ 100 × (1-rural domestic sewage treatment rate (%) × rural pollutant comprehensive removal rate (%)) (7)
[0097] Rural domestic sewage treatment rate (%) = rural domestic sewage design treatment scale (tons / day) × 80% ÷ (rural permanent population (person) × rural domestic sewage discharge coefficient (liters / person / day) ÷ 1000) (8)
[0098] Livestock and poultry breeding load:
[0099] Livestock and poultry breeding load (tons) = scale breeding wastewater discharge + breeder breeding wastewater discharge (tons) (9)
[0100] Pollution discharge from large-scale farming (tons) = {stock on hand + number of animals slaughtered in the current year (heads)} × scale farming ratio × livestock and poultry pollution discharge coefficient (kg / head) × (1 – scale farming resource utilization rate (%) × utilization efficiency (%)) ÷ 1000 (10)
[0101] Pollution discharge of livestock and poultry farmers = {stock on hand + number of livestock and poultry sold in the current year (heads)} × livestock and poultry farming ratio × livestock and poultry pollution discharge coefficient (kg / head) ÷ 1000 (11)
[0102] Aquaculture load:
[0103] Aquaculture load (tons) = aquatic product volume (tons) × aquaculture discharge coefficient (kg / ton) ÷ 1000 (12)
[0104] The surface runoff and sediment loss are calculated as follows:
[0105] The surface runoff (SCS-CN runoff curve model of the Soil Conservation Service of the United States Department of Agriculture (USDA)) is calculated as follows:
[0106]
[0107]
[0108] Where Q is the runoff depth (mm); P is the rainfall depth (mm); S is the potential maximum retention or infiltration (mm); CN is the CN value (dimensionless) of different land use types.
[0109] The source of P data is regional precipitation data, generally annual precipitation, and the time scale of P corresponds to the time scale of Q.
[0110] The formula for calculating sediment loss (Universal Soil Loss Equation (USLE) created by the United States Department of Agriculture) is:
[0111] A=R×K×L×S×C×P (5)
[0112] Where A is the soil loss per unit area (t·ha -1 ·yr -1 )R is the rainfall and erosibility factor, unit: MJ·mm·ha -1 ·h -1 ·yr -1 ; K is the soil erodibility factor, unit: t·ha·h·ha -1 ·MJ -1 mm -1 ; L is the slope length factor, S is the slope steepness factor, C is the coverage and management factor, and P is the protection practice factor;
[0113] The R value is calculated according to the following formula (21):
[0114]
[0115] Among them, P i is the monthly rainfall (mm).
[0116] The K value is calculated according to the following formula (22):
[0117]
[0118] In the formula: SAN: sand content (%); SIL: silt content (%); CLA: clay content (%); C: organic carbon content (%); SN1 = 1-SAN / 100.
[0119] The L value is calculated as follows:
[0120] L = (λ / 22.13) α (twenty three)
[0121] α=β / (β+1) (24)
[0122] β=(sinθ / 0.0896) / [3×(sinθ) 0.8 +0.56] (25)
[0123] λ=D i / cosθ (26)
[0124] Among them, λ is the sum of the slope lengths in the horizontal direction; α is the slope length coefficient; θ is the slope angle extracted based on digital elevation.
[0125] The S value is calculated as follows:
[0126]
[0127] Here, θ is the slope.
[0128] The C value is calculated as follows:
[0129]
[0130] Where c is the vegetation coverage rate.
[0131] The P value was obtained through the runoff plot measurement method and literature review method.
[0132] Step 4: Use the acquired basic data to train and verify the model so that the model meets the evaluation criteria;
[0133] The data set required for constructing the S-APPI model uses six specific indicators, namely, surface runoff, sediment loss, corrected planting loss load, aquaculture load, livestock and poultry breeding load, and rural domestic sewage load, as the driving data set for model training, and water quality monitoring data as the response data set.
[0134] Combined with the XGBoost algorithm, the simulation results of the S-APPI model are trained and calibrated. In this example, Jiangsu Province is divided into 64 districts and counties for simulation. During the model verification process, the determination coefficient R between the simulated water quality data and the monitoring data is output. 2 Only when the RMSE and RMS meet the requirements can the model be applied to the calculation and prediction of non-point source pollution load. The prediction results are shown in Table 6.
[0135] Table 6 Prediction performance of the S-APPI model based on the XGBoost algorithm
[0136]
[0137] When the hyperparameters obtained by random search are applied to the model to simulate water quality parameters (chemical oxygen demand, ammonia nitrogen, total nitrogen, and total phosphorus), the R 2 They are 0.59, 0.59, 0.59 and 0.64 respectively. 2 The values are 0.59, 0.53, 0.55 and 0.54 respectively. In the test phase, R 2 The difference is small compared with the training period, indicating that the model fitting effect is good. The RMSE of the four indicators in the training stage were 2.23, 0.20, 0.50, and 0.04, and the RMSE in the test stage were 2.22, 0.21, 0.51, and 0.04, respectively. In general, the RMSE of the training and prediction stages of the four pollutants is small, indicating that the difference between the water quality prediction value simulated by the model and the monitoring value is small, which shows that when the hyperparameters generated by random search are used in the model, the S-APPI model based on the XGBoost algorithm can better fit the model to the data.
[0138] In this example, the trained model is used to predict the four water quality parameters in the flood season, dry season and normal water season. 2 All of them are above 0.7, which indicates that the S-APPI model has a high simulation accuracy for water quality parameters in the flood season. 2 The range can reach 0.72-0.99, but the R of the test set and the training set for different pollutants during these two water periods is 2 The maximum difference can reach 0.29 and 0.46 respectively, which indicates that the S-APPI model may overfit when simulating water quality parameters for the training set data in the normal water season and the dry water season. Therefore, the simulation effect of this model in the dry and normal water seasons is poor, which is also in line with the characteristic law that agricultural non-point source pollution mainly occurs in periods of large rainfall.
[0139] Table 7 Prediction performance of S-APPI model under different water periods
[0140]
[0141] Step 4: Use the trained model to calculate the agricultural non-point source pollution load in the area to be assessed
[0142] The specific steps are as follows:
[0143] After the model is calibrated and verified, the collected basic data are imported into the model to calculate the pollutant emissions from the crop industry, livestock and poultry farming, aquaculture and rural domestic sewage. Finally, the above estimated results are summed up by pollutant type to obtain the total amount of agricultural non-point source pollutant emissions in Jiangsu Province. The simulated pollutant load is shown in the figure below. Figure 1-6 shown.
[0144] in, Figure 1 AD respectively represent the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2022, and the corresponding original data are shown in Table 8 below.
[0145] Table 8 Distribution of total nitrogen, ammonia nitrogen, total phosphorus and COD loads from agricultural non-point sources in 2022
[0146]
[0147]
[0148] Figure 2 AD respectively represent the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2021, and the corresponding original data are shown in Table 9 below.
[0149] Table 9 Distribution of total phosphorus, ammonia nitrogen, total nitrogen and COD loads from agricultural non-point sources in 2021
[0150]
[0151]
[0152] Figure 3 AD respectively represent the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2020, and the corresponding original data are shown in Table 10 below.
[0153] Table 10 Distribution of total phosphorus, ammonia nitrogen, total nitrogen and COD loads from agricultural non-point sources in 2020
[0154]
[0155]
[0156]
[0157] Figure 4 AD represents the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2019, and the corresponding original data are shown in Table 11 below.
[0158] Table 11 Distribution of total phosphorus, ammonia nitrogen, total nitrogen and COD loads from agricultural non-point sources in 2019
[0159]
[0160]
[0161]
[0162] Figure 5 AD respectively represent the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2018, and the corresponding original data are shown in Table 12 below.
[0163] Table 12 Distribution of total phosphorus, ammonia nitrogen, total nitrogen and COD loads from agricultural non-point sources in 2018
[0164]
[0165]
[0166]
[0167] Figure 6 AD respectively represent the distribution of agricultural non-point source total phosphorus, ammonia nitrogen, total nitrogen and COD loads in Jiangsu Province in 2017. The corresponding original data are shown in Table 13 below.
[0168] Table 13 Distribution of total phosphorus, ammonia nitrogen, total nitrogen and COD loads from agricultural non-point sources in 2017
[0169]
[0170]
[0171]
[0172] The overall simulated pollutant change trend is as follows Figure 7 As shown, Figure 7 In the data, AD represents the changing trends of total phosphorus, ammonia nitrogen, total nitrogen and COD in rural domestic sewage, livestock and poultry breeding, aquaculture and planting industry respectively. Figure 8 This is a schematic diagram of the total load trend of agricultural non-point source pollution in 64 districts and counties in Jiangsu Province.
[0173] The above embodiments are applied to 64 districts and counties in Jiangsu Province, and the total amount of agricultural non-point source pollutant emissions in Jiangsu Province from 2017 to 2022 is estimated. The pollution load from different agricultural sources can also be estimated, breaking through the technical bottleneck of agricultural non-point source pollution load estimation and governance effect evaluation in plain river network areas.
[0174] This method has low implementation requirements and is simple to operate for the increasingly severe agricultural non-point source pollution. It can also effectively identify the priority pollution types to be controlled in the study area, thereby accurately managing and preventing agricultural non-point source pollution.
Claims
1. A method for assessing agricultural non-point source pollution load based on XGBoost algorithm, characterized in that: The specific steps are as follows: Step 1: Select the area to be evaluated and obtain the basic data required to build the model; The basic data consists of primary indicators and secondary indicators; the primary indicators are surface runoff, sediment loss, planting loss load, aquaculture load, livestock and poultry breeding load, rural domestic sewage load; the secondary indicators include meteorological data, soil physical property data, agricultural production activities and other indicators; Step 2: Correct the loss coefficient of the planting industry and complete the construction of the basic data set of the S-APPI model; 2.1) S-APPI model basic model: APPI i =RI i WF1+SPI i WF2+CLI i WF3+PALI i WF4 (1) Where: APPI i is the comprehensive index of agricultural non-point source pollution potential in region i; RI i is the runoff index of region i, representing the surface runoff capacity of the study area; SPI i is the sediment generation index of region i, representing the sediment loss capacity of the study area; CLI i is the planting load index of region i, representing the contribution of agricultural planting to non-point source pollution load in the study area; PALU i is the human and livestock load index of region i, which evaluates the contribution of pollution generated by humans and livestock in the study area; WF is the corresponding weight of each index. 2.2) The entropy weight method is introduced to correct the spatial heterogeneity of the pollution discharge coefficient of the planting industry, and the weights of the four attribute indicators, namely, the fertilizer use index, the pesticide application index, the crop planting index and the garden planting index, are determined. The calculation formula of the planting industry correction parameter is: Planting loss correction parameter = fertilizer use index × W1 + pesticide application index × W2 + crop planting index × W3 + garden planting index × W4 (2) Where: the crop loss correction parameter is the correction value of the crop nitrogen and phosphorus discharge (loss) coefficient, ranging from 0 to 1; W is the corresponding weight of each index. Then, the median value of the crop loss correction parameter of the crop nitrogen and phosphorus emission / loss coefficient of each unit is calculated according to the crop emission / loss coefficient, and the crop emission / loss coefficient is corrected in proportion; 2.3) Calculate the parameters required to build the S-APPI model, including: Surface runoff: Where, Q is the runoff depth, in mm; P is the rainfall depth, in mm; S is the potential maximum retention or infiltration, in mm; CN is the CN value of different land use types; Sediment loss: A=R×K×L×S×C×P (5) Where A is the soil loss per unit area (t·ha -1 ·yr -1 )R is the rainfall and erosibility factor, unit: MJ·mm·ha -1 ·h -1 ·yr -1 ; K is the soil erodibility factor, unit: t·ha·h·ha -1 ·MJ -1 mm -1 ; L is the slope length factor, S is the slope steepness factor, C is the coverage and management factor, and P is the protection practice factor; Nitrogen and phosphorus loss in the crop industry (tons) = {crop sowing area (hectares) × crop sowing process loss coefficient (kg / hectare) × crop loss correction parameter + garden area (hectares) × garden loss coefficient (kg / hectare) × crop loss correction parameter} × {survey year / 2017 fertilizer usage per unit area for crop industry (kg / hectare)} ÷ 1000; (6) Rural domestic sewage pollutant discharge (tons) = rural permanent population (10,000 people) × pollutant generation intensity (g / person / day) × 365 (days) ÷ 100 × (1-rural domestic sewage treatment rate (%) × rural pollutant comprehensive removal rate (%)); (7) Rural domestic sewage treatment rate (%) = rural domestic sewage design treatment scale (tons / day) × 80% ÷ (rural permanent population (person) × rural domestic sewage discharge coefficient (liters / person / day) ÷ 1000); (8) Livestock and poultry breeding load (tons) = wastewater discharge from large-scale breeding + wastewater discharge from breeding households (tons); (9) Pollution discharge from large-scale farming (tons) = {stock on hand + number of livestock and poultry slaughtered in the current year (heads)} × scale farming ratio × livestock and poultry pollution discharge coefficient (kg / head) × (1-scale farming resource utilization rate (%) × utilization efficiency (%)) ÷ 1000; (10) Pollution discharge of livestock and poultry farmers = {stock number + number of livestock and poultry sold in the current year (heads)} × livestock and poultry farming ratio × livestock and poultry pollution discharge coefficient (kg / head) ÷ 1000; (11) Aquaculture load (tons) = aquatic product volume (tons) × aquaculture discharge coefficient (kg / ton) ÷ 1000; (12) Step 3: Use the acquired basic data to train and verify the model so that the model meets the evaluation criteria. The specific steps are as follows: 3.1) Use the XGBoost algorithm to build an integrated learning model, and its calculation formula is: In the formula, is the model prediction value, i is the number of samples, k is the number of integrated regression trees, x i is the data sample, F is the set space formed by the regression tree, f k represents one of the regression trees; 3.2) Combined with the traditional APPI model, a model driving factor system with six specific indicators including surface runoff, sediment loss, crop loss load, aquaculture load, livestock and poultry breeding load, and rural domestic sewage load was constructed as the driving data set for model training. Water quality monitoring data was used as the response data set, and the crop discharge load before correction was used for model training and verification; 3.3) Use the corrected crop emission / loss coefficient to simulate agricultural non-point source load. After the corrected crop pollutant loss load is calculated, the S-APPI model data set format is constructed. After the corrected prediction data set is completed, the model is retrained and verified. The trained model is used to simulate water quality, and the water quality prediction results are compared with the water quality monitoring data to complete the model simulation accuracy output; 3.4) The model was validated using long-term water quality monitoring data and the coefficient of determination R 2 The root mean square error (RMSE) was used as the evaluation parameters, and four indicators, namely, chemical oxygen demand, ammonia nitrogen, total nitrogen and total phosphorus, were selected as the water quality parameter indicators for the model simulation; R 2 Used to evaluate the model fitting effect, R 2 The value range is between 0 and 1; the determination coefficient R 2 The root mean square error RMSE calculation formula is as follows: in, Indicates the actual value, Represents the predicted value, y a Indicates the average of actual values; When the following is met during calibration and verification: 2 The root mean square error (RMSE) is not less than 0.6 and not higher than 0.07, and the coefficient of determination R between the training set and the validation set is 2 When the difference between the RMSE value and the RMSE value is within 0.1, the trained S-APPI model is obtained; Step 4: Input the basic data collected in step 1 into the trained S-APPI model obtained in step 3.4), and use the modified crop loss coefficient combined with the output coefficient method to calculate the pollution load of each research unit, thus completing the target agricultural non-point source pollution load assessment.
2. The agricultural non-point source pollution load assessment method based on the XGBoost algorithm according to claim 1 is characterized in that: The meteorological data include annual and monthly precipitation; the soil physical property data include soil particle size distribution, soil organic carbon content, soil slope and vegetation coverage; the agricultural production activities include land use type, fertilizer application category / application amount, agricultural planting category / area / loss coefficient, livestock and poultry category / stock and output number / livestock and poultry excretion coefficient / livestock and poultry manure utilization rate and aquatic product category / output / pollution coefficient / aquaculture wastewater treatment rate / treatment efficiency; the other indicators include runoff curve number, rural permanent population / domestic sewage generation coefficient / domestic sewage treatment rate / treatment efficiency.
3. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for assessing agricultural non-point source pollution load based on the XGBoost algorithm as described in claim 1 or 2 are implemented.
4. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for assessing agricultural non-point source pollution load based on the XGBoost algorithm as described in claim 1 or 2 are implemented.