Coastal zone agricultural non-point source pollution prediction method and system

By constructing a random forest regression model and a SHAP interpretable model, and screening the main influencing factors, the problems of low model efficiency and low accuracy in the prediction of agricultural non-point source pollution in coastal zones were solved, achieving high-precision and interpretable pollution flux prediction and supporting targeted governance.

CN122491948APending Publication Date: 2026-07-31SHANDONG JIANZHU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG JIANZHU UNIV
Filing Date
2026-07-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In coastal areas, existing technologies struggle to effectively utilize basic data for high-precision prediction of agricultural non-point source pollution, especially total phosphorus (TP) non-point source pollution. The problems include low model efficiency, low accuracy, and a lack of interpretability in machine learning models.

Method used

By constructing a random forest regression model, extracting pollutant generation and transport factors from soil, meteorological, DEM, and remote sensing image data, and combining the SHAP interpretable model to screen the main influencing factors, a coastal zone agricultural non-point source pollution prediction model is constructed to achieve high-resolution pollution flux prediction.

Benefits of technology

It improves the accuracy and efficiency of agricultural non-point source pollution prediction in coastal areas, reduces reliance on large-scale mechanistic models, enhances the generalization ability in data-scarce regions, provides causal explanation chains, and supports targeted governance solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122491948A_ABST
    Figure CN122491948A_ABST
Patent Text Reader

Abstract

This invention relates to the field of agricultural environmental monitoring and prediction technology, and discloses a method and system for predicting agricultural non-point source pollution in coastal zones. The method involves acquiring basic data of coastal watersheds; extracting pollutant production and transport factors at the sub-watershed scale, including at least source TP concentration, farmland area ratio, erosion, and distance attenuation factors; obtaining tidal-corrected measured TP flux at the estuary; constructing and training a random forest model based on this data; using SHAP analysis to screen dominant factors; retraining the model based on the dominant factors to obtain a prediction model; inputting dominant raster data; and outputting pollution flux prediction results. This invention provides a method and system for predicting agricultural non-point source pollution in coastal zones, enabling rapid prediction of the risk of agricultural TP non-point source pollution discharge into the sea based on existing shared basic data and supplementary monitoring and sampling data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural environmental monitoring and prediction technology, and in particular to a method and system for predicting agricultural non-point source pollution in coastal zones. Background Technology

[0002] In recent years, agricultural non-point source pollution has gradually become a key bottleneck in the continuous and stable improvement of the marine environment quality of my country's coastal waters. After years of research both domestically and internationally, a good understanding of the key influencing factors of agricultural non-point source pollution has been achieved. Soil erosion is generally considered the main driving factor for agricultural total phosphorus (TP) non-point source pollution. Current simulation methods for agricultural non-point source pollution mainly fall into two categories: empirical models and mechanistic models. However, both types of models largely rely on traditional empirical parameters. Empirical models are simple and require less data, but their accuracy is relatively low; mechanistic models have complex operating parameters and often require long-term spatial data support. Generally speaking, coastal areas have strong land-sea interactions, and the influencing factors have greater uncertainty. Monitoring basic data related to agricultural non-point source pollution is more difficult in coastal areas compared to inland areas, and the corresponding empirical model parameters are relatively insufficient. Therefore, in assessing agricultural non-point source pollution in coastal areas, the lack of basic data often leads to low operating efficiency of large-scale mechanistic models and low prediction accuracy of various models.

[0003] Soil erosion processes are typically quantified using the Modified Universal Soil Loss Equation (RUSLE). However, due to the complex mechanisms of agricultural non-point source pollution, the contribution of soil erosion to agricultural non-point source pollution output is unclear, and RUSLE is rarely used directly for agricultural non-point source pollution risk analysis of typical pollutants, especially total phosphorus (TP). With the integration of artificial intelligence technologies, the application of machine learning models for predicting agricultural non-point source pollution emissions has made good progress, significantly improving the efficiency and accuracy of agricultural non-point source pollution analysis and prediction. However, since most classic machine learning models are data-driven black-box models, their interpretability for agricultural non-point source pollution processes is relatively poor. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a method and system for predicting agricultural non-point source pollution in coastal zones, which can rapidly predict the risk of agricultural non-point source pollution (TP) discharge into the sea based on existing shared basic data and supplementary monitoring and sampling data.

[0005] This invention provides a method for predicting agricultural non-point source pollution in coastal zones, comprising the following steps: Acquire basic data for the target coastal watershed, including soil sampling data, meteorological data, digital elevation model (DEM) data, remote sensing image data, and land use data; Based on basic data, pollutant generation and transport factors at the sub-basin scale are extracted. The pollutant generation and transport factors include: soil source factors that characterize the content and distribution characteristics of soil-source pollutants, erosion factors that characterize soil erosion dynamics and erosion resistance, and confluence attenuation factors that characterize the degree of attenuation of pollutants during their migration from grid cells to receiving water bodies. Obtain measured TP pollution flux data at the estuary of the sub-basin, that is, obtain measured TP pollution flux data after correction based on tidal patterns; Using pollutant production and transport factors as input features and measured TP pollution flux data as output labels, a random forest regression model is constructed and trained. After the random forest regression model is trained, the Shapley value of each pollutant production and transport factor is calculated using the SHAP interpretable model to obtain the average marginal contribution of each factor under all feature combinations, which is used as the global average absolute contribution of the factor. The factors are sorted from largest to smallest based on their global average absolute contribution, and the top m factors with a cumulative contribution greater than 95% are selected as the main influencing factors. Based on the main influencing factors, the random forest regression model was retrained or optimized to obtain a prediction model for agricultural non-point source pollution in the coastal zone. The raster data of the main influencing factors for the period to be predicted are input into the prediction model, and the prediction results of the agricultural non-point source pollution flux in the coastal zone are output.

[0006] Specifically, before being input into the random forest regression model, the basic data are uniformly converted into 3km×3km target raster data with uniform spatial resolution through spatial aggregation or resampling; among them, continuous variables are resampled using bilinear interpolation, count variables are aggregated using the mode method, and proportional variables are aggregated using the mean method.

[0007] Specifically, soil source factors include soil TP concentration factor SC. TP SC TP Based on regional surface soil sampling data, the spatial interpolation method using geographic weighted regression (GWR) was employed.

[0008] Specifically, soil source factors also include the farmland area ratio factor P. al P al Obtained by extracting the proportion of farmland area at the raster scale; SC TP Factors and P al The factors are input as independent feature variables into the random forest regression model.

[0009] Specifically, the erosion factors include one or more of the following: The precipitation erosibility factor R is obtained by calculating the cumulative rainfall from the start of monitoring to the current monitoring day based on daily rainfall data monitored by regional meteorological stations. The soil erodibility factor K was calculated based on the soil sand content (SAN), soil silt content (SIL), soil clay content (CLA), and soil organic carbon content (OC). The slope factor S is obtained from the slope data obtained by performing terrain analysis on the DEM. The maximum upstream flow path length L is extracted based on DEM hydrological analysis and is counted by the cumulative number of grid cells to reflect the possibility of pollutants carried by upstream water being deposited at the grid cells. The upstream cumulative runoff F1 is obtained from the cumulative raster data of the runoff calculated based on DEM hydrological analysis, and is used to reflect the impact of runoff scouring on pollutant production. The vegetation cover factor C was calculated based on the Normalized Difference Vegetation Index (NDVI) of remote sensing images.

[0010] Specifically, the confluence attenuation factor is the confluence migration attenuation factor D. D is obtained by hydrological analysis based on DEM, calculating the distance from each grid cell to the outlet of the sub-basin, and is expressed in terms of the number of grid cells.

[0011] Specifically, the hyperparameters of the random forest regression model are set as follows: 800 decision trees, no limit on the maximum tree depth, 2 minimum samples required for internal node splits, 1 minimum sample number for leaf nodes, and 42 random seeds. During training, the hyperparameters are optimized by combining k-fold cross-validation with grid search.

[0012] Specifically, measured TP pollution flux data were collected after correction based on tidal patterns, including: Obtain tidal forecast data for the estuary of the target area; Determine the low tide point within the sampling date; A preset time window centered on the low tide point was set as the optimal sampling period to avoid interference from seawater upwelling on land-based runoff monitoring; Water samples were collected and analyzed within a preset time window to obtain TP concentration, and the TP pollution flux into the sea was calculated in combination with the runoff during the corresponding time period.

[0013] A coastal agricultural non-point source pollution prediction system includes: The data acquisition module is used to acquire basic data for the target coastal watershed. The factor extraction module is used to extract pollutant generation and transport factors at the sub-basin scale based on basic data. The pollutant generation and transport factors include at least: soil source factors characterizing the content and distribution characteristics of soil-source pollutants, erosion factors characterizing soil erosion dynamics and erosion resistance, and confluence attenuation factors characterizing the degree of attenuation of pollutants during their migration from the grid unit to the receiving water body. The data correction module is used to acquire and store measured TP pollution flux data at the estuary after correction based on tidal patterns; The model building and training module is used to construct and train a random forest regression model by taking pollutant production and transport factors as input and measured TP pollution flux data as output. After training, the Shapley value of each factor is calculated using the SHAP interpretable model to obtain the global average absolute contribution of each factor, and the main influencing factors with a cumulative contribution greater than 95% are screened. Based on the screened main influencing factors, the random forest regression model is retrained or optimized to obtain the coastal zone agricultural non-point source pollution prediction model. The prediction execution module is used to call the coastal zone agricultural non-point source pollution prediction model, calculate the data to be predicted, and output the pollution flux prediction results.

[0014] Specifically, the system provides a multimodal data input interface; among which, The input interface is configured to read directly if it receives existing Normalized Difference Vegetation Index (NDVI) raster data input by the user. If the system receives raw remote sensing imagery data in the near-infrared and infrared bands from the user, it calls the built-in band calculation tool to generate NDVI raster data in order to obtain the vegetation cover factor.

[0015] The technical solution provided by this invention has the following advantages compared with the prior art: This invention transforms basic data such as soil, meteorology, DEM, remote sensing images, and land use into pollutant generation and transport factors with clear physical meaning. Using these factors as input features and measured TP pollution flux corrected based on tidal patterns as the output label, a random forest regression model is constructed to fit the complex nonlinear relationship between multiple factors. Then, a SHAP interpretable model is introduced to quantitatively strip away the marginal contribution of each factor to the prediction results, transforming the black-box decision-making logic into a verifiable quantitative explanation, and thereby selecting the main influencing factors whose cumulative contribution meets preset conditions. Finally, based on the master control factors, the random forest model is reconstructed or optimized, redundant features are removed, and the master control factor raster data for the period to be predicted is input to output high-resolution pollution flux prediction results. Structurally, erosion factors and distance attenuation factors with rigorous physical definitions are embedded into the model input, making the decision-making constrained by the scientific mechanism framework of pollutant generation, stripping, and migration. Post-hoc SHAP analysis accurately quantifies the percentage contribution of each factor to the prediction results, clearly tracing the degree of drive by soil erosion and pollution source intensity distribution, thus providing a causal explanation chain for the assessment results. This method significantly improves the prediction accuracy and operational efficiency in data-scarce coastal areas. By identifying a few dominant factors, high-precision predictions can be driven by focusing on key monitoring, greatly reducing reliance on massive amounts of full data. It overcomes the problems of low efficiency and insufficient accuracy associated with large-scale mechanistic models. The simplified structure also reduces the risk of overfitting and enhances the generalization ability for data-scarce regions. The entire framework achieves complementary advantages between mechanistic and data models: mechanistic models are used for high-quality feature extraction and knowledge injection, data models achieve high-precision nonlinear fitting, and interpretable models are used for knowledge discovery and verification feedback. This overcomes the shortcomings of pure data models (ignoring physical laws) and pure mechanistic models (numerous parameters and poor adaptability). Furthermore, managers can flexibly adjust the cumulative contribution threshold of the selected dominant factors to achieve the best economic balance between prediction accuracy and monitoring costs, forming a quantifiable targeted governance plan. The method has good transferability and can be extended to different coastal watersheds or other non-point source pollution assessment and prediction. Attached Figure Description

[0016] Figure 1 The principle of a random forest model provided in this embodiment of the invention; Figure 2 A technical roadmap for a method to predict agricultural non-point source pollution in coastal zones, provided by an embodiment of the present invention; Figure 3 A honeycomb diagram of global SHAP values ​​for quantifying the contributions of each factor in a SHAP model, provided in an embodiment of the present invention; Figure 4 This invention provides an embodiment of the SHAP model that quantifies the average absolute contribution distribution of each factor. Detailed Implementation

[0017] The following detailed description of a specific embodiment of the present invention is provided in conjunction with the accompanying drawings. However, it should be understood that the scope of protection of the present invention is not limited to the specific embodiment.

[0018] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the technical solution of this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0019] The present invention will be described below through several specific embodiments. To keep the following description of the embodiments clear and concise, detailed descriptions of known functions and components may be omitted. When any component of an embodiment of the present invention appears in more than one drawing, the component may be represented by the same reference numerals in each drawing.

[0020] Figure 1 This invention provides a principle for a random forest model. Figure 2 This invention provides a technical roadmap for a method to predict agricultural non-point source pollution in coastal zones. Figure 3 This invention provides a honeycomb diagram of global SHAP values ​​for quantifying the contributions of each factor in a SHAP model. Figure 4 This invention provides an embodiment of the SHAP model that quantifies the average absolute contribution distribution of each factor.

[0021] like Figure 1 , Figure 2 , Figure 3 and Figure 4As shown in the figure, this invention provides a method and system for predicting agricultural non-point source pollution in coastal zones. The specific process is as follows: Screening of soil-source pollutant generation and transport factors in watersheds. Extracting and quantifying three types of influencing factors in terrestrial watersheds: soil source TP concentration, soil erosion, and runoff attenuation. Conducting simulations of agricultural non-point source pollution discharge into the sea based on RF model regression analysis. Combining pollutant flux monitoring data into the sea, applying the RF machine learning model to achieve regression simulation and verification of the influence of terrestrial generation and transport factors on pollutant discharge into the sea at the sub-watershed scale; quantifying the contribution of each factor and screening the main factors. Constructing a precise and rapid prediction system for coastal zone agricultural TP non-point source pollution. Based on this, system development is carried out, setting the screened pollutant generation and transport influencing factors as the input, setting the RF model training optimization parameters and the SHAP model as the analysis execution end, and using the predicted flux into the sea and the verification of measured data as the output, thus constructing a precise and rapid prediction system for coastal zone agricultural TP non-point source pollution. The specific technical route is as follows: Watershed soil source pollutant generation and transport factors and data acquisition: All factors are uniformly converted into 3km×3km raster data. The calculation methods and basic data acquisition process for each factor are as follows: Soil source pollutant factors: 1) Soil TP concentration factor SC TP Based on regional surface soil sampling (0-20cm) data, spatial interpolation is performed using soil sampling density to achieve regional rasterization processing of soil indicators such as TP. The example data provided in this invention is obtained by using the geographically weighted regression (GWR) spatial interpolation method, with a spatial resolution of 3km×3km.

[0022] The Gross Regression (GWR) model is a type of spatially variable coefficient model. Compared to the traditional linear regression model (ordinary least squares, OLS), GWR employs the concept of local smoothing and incorporates the idea of ​​relevant spatial location; that is, its regression coefficients are all functions of spatial location variables. The basic equation of the GWR model is as follows: (1), In the formula, y i Let be the index value of the i-th prediction point; β 0 ( u i ,v i) The polynomial intercept; β j ( u i ,v i) Let be the regression coefficient of the j-th influencing variable at prediction point i, which is the position ( u i ,vi) The function; x ij Let be the index value of the j-th influencing variable at point i; ε i Let be the regression residual term at point i.

[0023] Farmland proportion factor P al To quantify the impact of agricultural production activities on soil erosion and agricultural TP non-point source pollution output, this invention uses a grid-scale approach to measure the percentage of farmland (%, P). al) The impact characteristics of agricultural production activities are quantified.

[0024] Soil erosion factors: Soil erosion factors are quantified based on the classical RUSLE model equation (Equation (1)): (2), In the formula, A is the annual soil erosion output modulus, t·ha -1 ·yr -1 R represents the rainfall erosivity factor, in MJ·mm·ha. -1 ·h -1 ·yr -1 K represents the soil erodibility factor, t·h·MJ. -1 ·mm -1 L and S represent the topographic slope length and slope factor, respectively; C represents the vegetation cover and management factor; and P represents the soil and water conservation measures factor, which is not considered in the examples of this invention.

[0025] The RUSLE factors were simplified and regionally adjusted as follows: Rainfall erodibility factor R: To match the monitoring of TP pollution flux at the estuary and to eliminate the time delay effect of land-based pollutants entering the sea, this invention mainly uses daily rainfall data (mm / d) monitored by regional meteorological stations to obtain the corresponding cumulative rainfall (mm, accumulated from the start of monitoring time) during the TP pollution flux monitoring period at the estuary, thereby quantifying the rainfall erodibility factor: (3), In the formula, R is the rainfall erosibility factor of this invention, R i Let n be the cumulative rainfall in the i-th time period, and n be the time period for monitoring TP pollution flux at the estuary.

[0026] Soil erodibility factor K: The basic calculation formula for the K factor is shown in equation (4).

[0027] (4), In the formula, K is the influencing factor of soil erodibility due to non-point source pollution, t·h·MJ -1 ·mm -1SAN represents soil sand content (%); SIL represents soil silt content (%); CLA represents soil clay content (%); OC represents soil organic carbon content (%).

[0028] The SAN, SIL, and CLA data involved in the K-factor calculation were mainly obtained from the basic soil spatial database of relevant websites (such as the Resource and Environmental Science and Data Platform of the Chinese Academy of Sciences); OC was derived from the spatial interpolation results of soil sampling point data using the GWR method.

[0029] Topographic factors: Unlike the slope length and slope factor in the original RUSLE, the topographic factors are characterized by slope factor (S), upstream cumulative runoff (Fl), and upstream maximum flow path length (M).

[0030] In this invention, since the study area has relatively flat terrain, θ is directly used as the slope factor S. The slope data obtained from terrain analysis using a DEM (Digital Elevation Model) (the data in this invention has a spatial resolution of 30m × 30m and is sourced from the National Earth System Science Data Center) based on ArcGIS software is used to characterize the impact of terrain slope on pollutant production. The cumulative runoff (Fl, the cumulative raster count calculated using the DEM) is used to represent the influence of upstream water flow, reflecting the impact of runoff scouring on pollutant production. Furthermore, the maximum upstream flow path length (M) is added to characterize the slope length factor. Generally, when the upstream cumulative runoff of a sub-basin is constant, the narrower the basin, the lower the risk of upstream pollutants migrating to it and increasing the risk of pollution output from that source point. Therefore, we added the maximum upstream flow path length (M) to characterize the secondary pollution output characteristics carried by upstream water flow.

[0031] Vegetation cover factor C: The vegetation cover factor C is mainly calculated based on the normalized vegetation index (NDVI) of remote sensing images, as shown in equations (5) and (6): (5), (6), In the formula, C is the quantified vegetation cover factor in RUSLE; α and β are dimensionless parameters, with empirical values ​​of 2 and 1 respectively; NDVI is the normalized vegetation index. The NDVI data used in this study were obtained by inversion from Landsat series remote sensing images. For near-infrared reflectivity, This refers to the reflectivity in the red light band.

[0032] Drainage migration attenuation factor D: Considering the migration and attenuation of soil-source TP pollutants after they are output from each grid cell and during the runoff process, this invention directly performs hydrological analysis based on DEM and calculates the distance from each grid cell to the outlet of the sub-basin (in terms of the number of grid cells) to present the characteristics of the drainage migration attenuation factor.

[0033] Sub-basin zoning: Since agricultural non-point source pollution processes converge according to watersheds, the above factors need to be zoned according to sub-basins flowing into the sea. Sub-basin data are generally obtained through hydrological analysis using a DEM (Digital Elevation Model). This invention specifies that the data should be directly input in vector format.

[0034] Data spatial scale unification processing: After the above factors are quantized in the region, they all need to be uniformly converted into 3km×3km (soil sampling density) raster data; spatial data interpolation, slope length and slope, hydrological analysis and spatial analysis of raster data of each factor all need to call the geoprocessing toolbox (ArcToolbox) based on the ArcGIS platform; for NDVI index input, existing data products can be used, or remote sensing images (such as Landsat) near-infrared and infrared bands can be input, and the ENVI band calculation tool can be called for calculation.

[0035] TP pollution flux at the estuary of the sub-basin flux Monitoring and data acquisition: It is necessary to obtain the actual TP flux into the sea from the sub-basin as the data sample for model regression training output simulation and result verification. The flux into the sea from the sub-basin needs to be sorted into a unified format in the early stage.

[0036] Because the monitoring locations at the estuary are significantly affected by ocean tides, to ensure that the water flow at the monitoring points is a land-based water discharge process that is unaffected or minimally affected by tides, and to reduce the impact of seawater upwelling on the monitoring results, thereby further improving prediction accuracy, the TP inflow data required by this invention needs to be obtained by analyzing tidal patterns and selecting an appropriate time for sampling surveys at low tide. This invention, based on daily and hourly tidal statistics and prediction data from relevant websites, analyzed data from the Xinhui East Coast tidal monitoring point in Kenli District, Dongying City, to obtain the low tide time distribution for each sampling date. The optimal time for estuary hydrological and water quality monitoring was set at 2.5 hours before and after the low tide point. Based on this, sampling surveys were conducted, and continuous time-scale inflow flux data for TP and other related indicators were obtained. An example is the inflow flux data of the estuary in the covered area, collected every 5 days during the concentrated summer rainfall period from June to October 2020.

[0037] RF Regression Training and Result Validation: For each time period, based on the selected influencing factors and the results calculated according to the corresponding methods, the independent variables of each factor data were organized into RF model inputs by sub-basin partitioning and raster cells; the TP pollution flux at the estuary of each sub-basin was used as the dependent variable output of the RF model and the simulation validation data (corresponding to the sub-basin raster cells). Finally, the spatiotemporal dataset of RF regression simulation training for coastal agricultural TP non-point source pollution of this invention was obtained (in Excel data table format), as shown in Table 1.

[0038] RF enhances model robustness by integrating multiple decision trees and is suitable for handling complex nonlinear relationships between multiple source factors.

[0039] Regression Prediction of Agricultural TP Non-point Source Pollution Discharge into the Sea Based on RF Model: Feature Data Processing: For each time period, based on the selected influencing factors and the results calculated according to the corresponding methods, the independent variables of each factor data are processed into RF model inputs by sub-basin partitioning and raster cells; the TP pollution flux at the sea mouth of each sub-basin is used as the dependent variable output of the RF model and simulation validation data (corresponding to the sub-basin raster cells). Finally, the spatiotemporal dataset for RF regression simulation training of coastal agricultural TP non-point source pollution of this invention is obtained (in Excel data table format).

[0040] The Randomized Regression (RF) model was applied, assigning values ​​to the aforementioned influencing factors at the watershed grid scale. Continuous time-scale flux data from each sub-watershed were used as the regression prediction output for simulation (the main output variables involved in the regression prediction included sub-watershed zoning and the flux monitored at each time point). Furthermore, to ensure the accuracy and scientific applicability of the simulation results, sub-watershed zoning and the flux monitoring time were used as the basis for random sampling, with 80% of the data used for regression training and 20% for validation. To reduce the instability of validation results caused by random sampling, the regression simulation employed k-fold cross-validation (5-fold or 10-fold, depending on the data sample size) for model training and parameter optimization.

[0041] Table 1: Input / Output Variables of the RF Model

[0042] In this embodiment of the invention, the main hyperparameters of the random forest model are set as follows: 800 decision trees, no limit on maximum tree depth to allow for full growth of decision trees, a minimum number of samples required for internal node splits of 2, a minimum number of samples required for leaf nodes of 1, and 42 random seeds. During training, these parameters are optimized using k-fold cross-validation combined with grid search.

[0043] Model validation: using the coefficient of determination (r) 2 Quantifying the accuracy of RF regression predictions: (7), In the formula, n is the number of samples in the validation set (test set); y i It is the true value of the i-th sample; y is the model's prediction for the i-th sample; ave It is the mean of all true values ​​in the validation set.

[0044] While the trained RF model can predict TP pollution fluxes into the sea relatively accurately, its internal decision-making process is opaque as an ensemble learning model, and the contribution of each input factor to the prediction result is unclear. To address this issue, this invention embeds the SHAP (SHapley Additive exPlanations) interpretable model into the aforementioned trained RF model, quantifies the contribution of each factor, and performs factor selection accordingly.

[0045] The specific steps are as follows: The core idea of ​​SHAP originates from the Shapley value in game theory. It treats each input feature of the model as a "player" in a game, and the deviation between the model's predicted output value and the mean of all sample predictions as the "payoff". By traversing all possible feature combinations, it calculates the average marginal contribution (i.e., the Shapley value) of each feature across all input orders, thereby achieving a fair quantification of feature importance. Because the calculation of the Shapley value considers the interaction effects between features, SHAP can more accurately reflect the true contribution of each factor under the coupling effect compared to traditional feature importance ranking methods.

[0046] In this invention, SHAP analysis is introduced after the RF model training is completed. Using the trained RF model as the object of interpretation and the same training dataset as the computational basis, the contributions of nine pollution generation and transport factors are calculated. This invention sets a cumulative contribution greater than 95% as the screening criterion, selecting the top m factors whose cumulative contribution first exceeds 95% of the total contribution as the main influencing factors for predicting regional agricultural TP non-point source pollution. The remaining unselected factors are discarded in the subsequent prediction system construction due to their small contributions, thereby simplifying and optimizing the model input features.

[0047] Construction of a precise and rapid prediction system for TP non-point source pollution in coastal agriculture: The main influencing factors selected based on the SHAP model and the TP pollution flux monitored at the estuary over a continuous time scale are used again for regression training using the RF model, and these are used as the final parameters for system construction.

[0048] Simultaneously, a multimodal data input interface is designed. For example, NDVI data can be selected from existing NDVI data products or obtained by inputting remote sensing image data and calling the ENVI band analysis tool. At the same time, the SHAP interpretable model is nested to execute system development and build a precise and rapid prediction system for TP non-point source pollution in coastal agriculture.

[0049] The embodiments of the present invention have been applied and verified in the Yellow River Delta region. The coefficient of determination r² of the simulation prediction results can reach 0.95, indicating that the method of the present invention has high prediction accuracy.

[0050] The above inventions are merely a few specific embodiments of the present invention. However, the embodiments of the present invention are not limited thereto, and any variations that can be conceived by those skilled in the art should fall within the protection scope of the present invention.

Claims

1. A method for predicting agricultural non-point source pollution in coastal zones, characterized in that, Includes the following steps: Acquire basic data for the target coastal watershed, including soil sampling data, meteorological data, digital elevation model (DEM) data, remote sensing image data, and land use data; Based on the aforementioned basic data, pollutant production and transport factors at the sub-basin scale are extracted; The pollutant generation and transport factors include: soil source factors characterizing the content and distribution characteristics of soil-source pollutants, erosion factors characterizing soil erosion dynamics and erosion resistance, and confluence attenuation factors characterizing the degree of attenuation of pollutants during their migration from grid units to receiving water bodies. Obtain measured TP pollution flux data at the estuary of the sub-basin, that is, obtain measured TP pollution flux data after correction based on tidal patterns; Using the pollutant production and transport factors as input features and the measured TP pollution flux data as output labels, a random forest regression model is constructed and trained. After the random forest regression model is trained, the Shapley value of each pollutant production and transport factor is calculated using the SHAP interpretable model to obtain the average marginal contribution of each factor under all feature combinations, which is used as the global average absolute contribution of the factor. The factors are sorted from largest to smallest according to the global average absolute contribution, and the top m factors with a cumulative contribution greater than 95% are selected as the main influencing factors. Based on the main influencing factors, the random forest regression model is retrained or optimized to obtain a coastal zone agricultural non-point source pollution prediction model. The raster data of the main influencing factors for the period to be predicted are input into the prediction model, and the prediction results of the coastal agricultural non-point source pollution flux are output.

2. The method for predicting agricultural non-point source pollution in coastal zones as described in claim 1, characterized in that, Before being input into the random forest regression model, the basic data are uniformly converted into 3km×3km target raster data with uniform spatial resolution through spatial aggregation or resampling; among them, continuous variables are resampled using bilinear interpolation, count variables are aggregated using the mode method, and proportional variables are aggregated using the mean method.

3. The method for predicting agricultural non-point source pollution in coastal zones as described in claim 2, characterized in that, The soil source factors include soil TP concentration factor SC. TP The SC TP Based on regional surface soil sampling data, the spatial interpolation method using geographic weighted regression (GWR) was employed.

4. The method for predicting agricultural non-point source pollution in coastal zones as described in claim 3, characterized in that, The soil source factors also include the farmland area ratio factor P. al The P al The SC is obtained by extracting the proportion of farmland area at the raster scale; TP Factors and P al The factors are input as independent feature variables into the random forest regression model.

5. The method for predicting agricultural non-point source pollution in coastal zones as described in claim 2, characterized in that, The erosion factor includes one or more of the following: The precipitation erosibility factor R is obtained by calculating the cumulative rainfall from the start of monitoring to the current monitoring day based on daily rainfall data monitored by regional meteorological stations. The soil erodibility factor K was calculated based on the soil sand content (SAN), soil silt content (SIL), soil clay content (CLA), and soil organic carbon content (OC). The slope factor S is obtained from the slope data obtained by performing terrain analysis on the DEM. The maximum upstream flow path length L is extracted based on DEM hydrological analysis and is counted by the cumulative number of grid cells to reflect the possibility of pollutants carried by upstream water being deposited at the grid cells. The upstream cumulative runoff F1 is obtained from the cumulative raster data of the runoff calculated based on DEM hydrological analysis, and is used to reflect the impact of runoff scouring on pollutant production. The vegetation cover factor C was calculated based on the Normalized Difference Vegetation Index (NDVI) of remote sensing images.

6. The method for predicting agricultural non-point source pollution in coastal zones as described in claim 2, characterized in that, The confluence attenuation factor is the confluence migration attenuation factor D, which is obtained by hydrological analysis based on DEM and by calculating the distance from each grid cell to the outlet of the sub-basin, in terms of the number of grid cells.

7. The method for predicting agricultural non-point source pollution in coastal zones as described in claim 1, characterized in that, The hyperparameters of the random forest regression model are set as follows: 800 decision trees, no limit on maximum tree depth, 2 minimum samples required for internal node splits, 1 minimum sample number for leaf nodes, and 42 random seeds. During training, the hyperparameters are optimized by combining k-fold cross-validation with grid search.

8. The method for predicting agricultural non-point source pollution in coastal zones as described in claim 1, characterized in that, The measured TP pollution flux data obtained after correction based on tidal patterns specifically includes: Obtain tidal forecast data for the estuary of the target area; Determine the low tide point within the sampling date; A preset time window centered on the low tide point is set as the optimal sampling period to avoid interference from seawater upwelling on land-based runoff monitoring; Water samples are collected and analyzed within the preset time window to obtain TP concentration, and the TP pollution flux into the sea is calculated in combination with the runoff during the corresponding time period.

9. A coastal zone agricultural non-point source pollution prediction system, characterized in that, include: The data acquisition module is used to acquire basic data for the target coastal watershed. The factor extraction module is used to extract pollutant production and transport factors at the sub-basin scale based on the aforementioned basic data; The pollutant generation and transport factors include at least: soil source factors characterizing the content and distribution characteristics of soil-source pollutants, erosion factors characterizing soil erosion dynamics and erosion resistance, and confluence attenuation factors characterizing the degree of attenuation of pollutants during their migration from grid units to receiving water bodies. The data correction module is used to acquire and store measured TP pollution flux data at the estuary after correction based on tidal patterns; The model building and training module is used to construct and train a random forest regression model by taking the pollutant production and transport factors as input and the measured TP pollution flux data as output. After training, the Shapley value of each factor is calculated using the SHAP interpretable model to obtain the global average absolute contribution of each factor, and the main influencing factors with a cumulative contribution greater than 95% are screened. Based on the screened main influencing factors, the random forest regression model is retrained or optimized to obtain the coastal zone agricultural non-point source pollution prediction model. The prediction execution module is used to call the coastal agricultural non-point source pollution prediction model, calculate the data to be predicted, and output the pollution flux prediction results.

10. The coastal zone agricultural non-point source pollution prediction system as described in claim 9, characterized in that, The system provides a multimodal data input interface; wherein... The input interface is configured to directly read existing Normalized Difference Vegetation Index (NDVI) raster data input by the user. If the system receives raw remote sensing imagery data in the near-infrared and infrared bands from the user, it calls the built-in band calculation tool to generate NDVI raster data in order to obtain the vegetation cover factor.