Construction method of pollutant non-point source migration prediction model based on gated circulation unit neural network and application thereof

By applying the GRU-based pollutant non-source migration prediction model in large-scale urban watersheds, the problems of high cost of non-source pollution research, strong data dependence and limited model applicability in the existing technology are solved, and efficient and accurate pollutant migration prediction and urban pollution management support are achieved.

CN120217918APending Publication Date: 2025-06-27SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510135248.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology has problems such as high cost, strong data dependence and limited model applicability in large-scale urban watershed area pollution research.

Method used

By optimizing data acquisition, model structure and parameterization scheme, a pollutant surface source migration prediction model based on gated recurrent unit neural network (GRU) is proposed, combining hydrological simulation, hydraulic simulation and pollutant trend modules to achieve efficient prediction of pollutant migration in large-scale watersheds.

Benefits of technology

This method improves the applicability and accuracy of the simulation of non-point source pollution, reduces the cost of laboratory testing and large-scale sampling, and provides comprehensive data to support the monitoring and management of urban non-point source pollution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217918A_ABST
    Figure CN120217918A_ABST
Patent Text Reader

Abstract

The invention relates to a construction method and application of a pollutant non-point source migration prediction model based on a gated circulation unit neural network, and belongs to the urban non-point source pollution prediction technology, and the construction method is characterized by comprising the following steps: S1, collecting historical data of a to-be-detected region, dividing the historical data into a plurality of sub-basin data, and carrying out the standardized integration to form a feature data set; s2, establishing a pollutant non-point source migration prediction model through the feature data set; the pollutant non-point source migration prediction model comprises a hydrological simulation module, a hydraulic simulation module and a pollutant fate module; and S3, training the hydraulic simulation module and the pollutant fate module, and optimizing module parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technology of urban non-point source pollution prediction, and specifically relates to a method for constructing a pollutant non-point source migration prediction model based on a gated recurrent unit neural network and its application. Background Technique

[0002] Traditionally, the study of pollutant concentrations in urban rainfall runoff has used tracking monitoring and multi-point sampling methods. These methods require researchers to collect rainwater runoff samples at different time points and evaluate the temporal variation and cumulative effects of pesticide concentrations through laboratory analysis. Although the results of this method are accurate, it is costly and complex to operate in large-scale urban non-point source pollution studies. To overcome these deficiencies, the non-point source pollution model (NPS) has gradually become an important tool for studying urban non-point source pollution.

[0003] In recent years, mechanistic models have attracted extensive attention in the study of urban non-point source pollution. This type of model quantifies the processes of hydrology, soil erosion, and pollutant transport through mathematical expressions, thereby achieving a comprehensive simulation from the pollution source to the convergence point. For example, SWMM (Storm Water Management Model), as a mature urban non-point source model, has been widely used to simulate rainfall runoff and pollutant migration in cities around the world. However, traditional urban non-point source models still have significant limitations in the application of large-scale basins. On the one hand, these models are usually applicable to runoff analysis in small-scale basins. On the other hand, physical models rely on detailed and accurate local data (such as pipeline layouts) to truly reflect the runoff process, but in large-scale studies, this data is often difficult to obtain.

[0004] In small-scale basins, physical models usually perform better. For example: Zhao et al. (Zhao D Q, Chen J N, Wang H Z, Tong Q Y, Cao S B, Sheng Z. GIS-based urban rainfall-runoff modeling using an automatic catchment-discretization approach: a case study in Macau, China [J]. ENVIRONMENTAL EARTHSCIENCES, 2009, 59(2): 465-472) obtained a relatively high NSE value (0.86) in the runoff simulation of the internal catchment in Macau.

[0005] Li et al. (Li et al., Effects of Urban Non-Point Source Pollution from Baoding City on Baiyangdian Lake, China, 2017) used the SWMM model to simulate the runoff of the Baoding catchment area (21.18 km²), and the relative error (RE) ranged from -36.6% to -32.7%.

[0006] In the study of Yu et al. (Yu Y, Chen L, Xiao Y C, Chang C C, Zhi X S, Shen Z Y. New framework for assessing urban stormwater management measures in the context of climate change [J]. SCIENCE OF THE TOTAL ENVIRONMENT, 2022, 813: 13) in the Beijing catchment area (312.3 km²), the NSE values of the SWMM model fluctuated between 0.605 and 0.887.

[0007] However, in the application of large-scale basins, the performance of physical models is relatively limited. For example, Mahdi et al. (Mahdi N, Pagilla K R. Continuous Simulation of Highly Urbanized Watershed to Quantify Nutrients' Loadings [J]. WATER, 2021, 13(20): 17.) used the HSPF model to simulate the runoff of the Chicago basin (1670 km²), and its NSE value was only -0.10 and R² was 0.61. This indicates that with the increase of the basin area, the difficulty of runoff simulation increases significantly.

[0008] In addition, traditional non-point source pollution research also has problems such as high costs and strong data dependence in terms of methods. Taking the research by Pettigrove et al. (Pettigrove, et al. Catchment sourcing urban pesticide pollution using constructed wetlands in Melbourne, Australia. Sci Total Environ 2023;863: 160556.) as an example, although evaluating pesticide pollution through manual sampling and laboratory analysis is precise, it lacks the feasibility of large-scale promotion. These problems have given rise to the need for new methods. Especially in the field of prediction and management of non-point source pollution in large basins, it is necessary to construct a more efficient and applicable non-point source pollution model. Summary of the Invention

[0009] Based on the technical gaps and needs in the above fields, the present invention proposes a method for predicting the non-point source migration of pollutants based on a gated recurrent unit neural network (GRU) by optimizing data collection, model structure, and parameterization scheme to meet the needs of non-point source pollution research in large-scale urban basins.

[0010] The specific solution is as follows: A method for constructing a non-point source migration prediction model of pollutants based on a gated recurrent unit neural network, characterized by comprising the following steps: S1. Collect historical data of the area to be measured and divide it into complex sub-basin data, and then perform standardized integration to form a feature data set; The sub-basin data is comprehensive information on the geography, land cover, terrain, climate, and environmental characteristics of each sub-basin; The historical data includes: land use type, rainfall, pollutant usage, and elevation data; S2. Establish a non-point source migration prediction model of pollutants through the feature data set; The non-point source migration prediction model of pollutants includes: a hydrological simulation module, a hydraulic simulation module, and a pollutant fate module; The hydrological simulation module is used to calculate the total surface runoff time series of the sub-basin according to rainfall and land use type; The hydraulic simulation module is used to calculate the time series of the outlet diameter flow of the sub-basin pipe network according to the total surface runoff of the sub-basin; The pollutant fate module is used to calculate the time series of pollutant concentrations at the outlet of the sub-basin pipe network according to the outlet diameter flow of the sub-basin pipe network and pollutant usage; S3. Train the hydraulic simulation module and the pollutant fate module to optimize the module parameters.

[0011] In the above construction method, the land use types include: farmland, forest, grassland, shrubbery, wetland, water area, tundra, impervious surface, bare land, and ice and snow; The standardized integration of land use types includes the following steps: S111. Convert the land use type data into land use vector data; S112. Output the forest, grassland, and shrubbery in the vector data as green land data; Output the impervious surface in the vector data as impervious surface data; S113. Use the sub-basin data as the input area data, and the green land data and impervious surface data as the input type data, and input them into the spatial intersection calculation tool respectively to obtain the percentage data of each land use type in the sub-basin; S114. Merge the percentage data with the sub-basin data to obtain the proportion of land use types in the sub-basin, and complete the standardized integration.

[0012] In the above construction method, the standardized integration of the rainfall includes the following steps: S121. Use the linear interpolation method to convert the rainfall data recorded every hour into the data recorded every 10 minutes. The specific formula is as follows:

[0013] In formula (1), p is the rainfall at the current time, P0 is the rainfall at the previous time point, P1 is the rainfall at the next time point, t is the current time, t1 is the time of the previous time point, and t2 is the time of the next time point; S122. Match the rainfall at the current time with the corresponding sub-basin data to obtain the sub-basin rainfall time series, and complete the standardized integration.

[0014] In the above construction method, the standardized integration of the pollutant usage includes the following steps: S131. Use the green land data as the input to the geographic information software. According to the pollutant application characteristics, define the main pollutant application areas in the green land, collect the total pollutant amount, and generate the vector data of the green land buffer area through spatial analysis; S132. Use the sub-basin data as the input area, and the vector data of the green land buffer area as the input type, and obtain the proportion of the green land buffer area in each sub-basin through the spatial intersection calculation tool; S133. Incorporate the area proportion of the green land buffer area into the sub-basin data to update each sub-basin database; S134. Calculate the pollutant usage data for each sub - watershed based on the proportion of the green space buffer area in the sub - watershed, and complete the standardization integration of pollutant usage.

[0015] In the above construction method, the sub - watershed data includes: the number of the sub - watershed, the name of the region to which the sub - watershed belongs, the area of the region to which the sub - watershed belongs, the area of the sub - watershed, the proportion of the green space area in the sub - watershed area, the proportion of the impervious surface area in the sub - watershed area, the average slope of the sub - watershed, the rainfall of the sub - watershed, and the pollutant usage of the sub - watershed. The sub - watershed is divided by the following method: S141. Through polygonizing the vector data of roads and water bodies in the area to be measured for linear features, preliminarily divide the sub - watersheds and generate the planar areas for analysis. S142. Calculate the shortest distance between the preliminarily divided sub - watersheds and the main rivers, and merge the sub - watersheds with the shortest distance into the corresponding main river basins to optimize the sub - watershed division and reduce the computational complexity. S143. Identify the mountainous areas by using the terrain difference extraction tool and raster calculator of elevation data, eliminate them, retain the urban areas as the target sub - watersheds, and calculate the average slope of the sub - watersheds to complete the division.

[0016] In the above construction method, the hydrological simulation module includes an infiltration module and a depression storage module. The infiltration module is used to calculate the runoff of green space data. The specific method is:

[0017] In formula (2), f t is the infiltration ratio at time t, f c is the stable infiltration ratio, f0 is the initial infiltration rate, k is the infiltration rate decay coefficient, and t is the time.

[0018] In formula (3), R g is the runoff of green space data at time t, i t is the rainfall at time t, a is the area of the region occupied by the green space, f t is the infiltration ratio at time t. The depression storage module is used to calculate the runoff of impervious surface data. The specific method is:

[0019] In formula (4), R i is the runoff of impervious surface data, W is the width of the sub - catchment area, S is the average slope of the sub - catchment area, n is the Manning roughness coefficient, d is the water depth of the land plane, dp is the initial surface water depth; The total surface runoff calculation formula for the sub - watershed is as follows:

[0020] In formula (5), R s is the total surface runoff of the sub - watershed, and the data form is a time series.

[0021] 7. The construction method according to claim 6, wherein the hydraulic simulation module is a gated recurrent unit neural network, and the calculation method for the time series of the outlet diameter flow of the sub - watershed pipe network is:

[0022] In formulas (6), (7) and (8), W r is the reset gate, W z is the update gate, r z is the output of the reset gate at time t, z t is the output of the update gate at time t, [h t-1 ,x t and [r z ,x t are both concatenation operators for two vectors, x t is the total surface runoff input at time t, h t is the hidden state at time t, and additionally, it is also the outlet diameter flow of the sub - watershed pipe network at time t, and the data form is a time series.

[0023] In the above - mentioned construction method, the pollutant fate module includes a pollutant accumulation module and a pollutant scouring module; The calculation method of the pollutant accumulation module is:

[0024] In formula (9), P is the accumulated amount of pollutants, t is the previous drought time, K B is the pollutant accumulation rate per unit area, and d is the pollution removal rate; In formula (13), C0 is the initial accumulation days, W c is the total amount of pollutants used in the urban area where the sub - watershed belongs, G c is the green area of the urban area where the sub - watershed belongs, G sub is the green area of the sub - watershed; The pollutant scouring module is a Logistic scouring model, and the specific calculation method is:

[0025] In formulas (10), (11) and (12), Wt is the pollutant concentration at the outlet at time t, and the data form is a time series; Q t is the outlet flow rate of the sub - basin pipe network at time t, and δ t is the ratio of the pollution load available for scouring at time t. B1 and B2 are both coefficients of the logistic curve, C1 and C2 are both coefficients of the exponential scouring curve, and V t is the cumulative runoff at time t. P0 is the initial cumulative pollutant amount in the catchment area before rainfall, and P wt is the cumulative pollutant load washed away at time t.

[0026] In the above construction method, the hydrological simulation module and the hydraulic simulation module are encapsulated into a non - point source runoff simulation module; The training method of the non - point source runoff simulation module includes the following steps: S311: Using the sub - basin rainfall time series as input data and the sub - basin pipe network outlet flow rate time series data as the target variable, perform data cleaning, remove outliers, and divide them into a training set and a validation set according to a ratio of 7:3; S312: Calculate surface runoff through the Horton infiltration formula and the Manning formula as the input features of the model. Use the rectified linear unit as the activation function and predict the sub - basin pipe network outlet flow rate through a gated recurrent unit neural network; S313: Using the mean squared error as the loss function and the stochastic gradient descent method as the optimizer, set training parameters for the non - point source runoff simulation module, perform model training, and optimize module parameters through the backpropagation algorithm; The training parameters of the non - point source runoff simulation module are: f0 is defaulted to 0.9%; f c is defaulted to 0.5%; k is defaulted to 0.2; n is defaulted to 0.9; d p is defaulted to 0.0001 m; W r is the weight of the reset gate, defaulted to 1.0; W z is the weight of the update gate, defaulted to 1.0; The learning rate is defaulted to 0.00001, and the number of training epochs is defaulted to 60; S314: Based on the validation set, evaluate the performance of the model, use the Nash - Sutcliffe efficiency coefficient and the coefficient of determination as evaluation indicators to verify the fitting ability of the module; When NSE > 0.36 and R 2 > 0.5, it is determined that the model meets the requirements and runs normally;​​​​ When NSE ≤ 0.36 or R 2 ≤ 0.54, check whether there are missing values and outliers in the model input and output values. During model training, increase the internal parameters of the model and repeat the training until the probability space that the model can represent covers all possibilities corresponding to the internal runoff characteristics of the basin.

[0027] In the above construction method, the training of the pollutant fate module includes the following steps: S321: Use the pollutant usage data and the time series data of the outlet diameter flow rate of the sub - basin pipe network in each sub - basin as input data, and the time series of the pollutant concentration at the outlet of the sub - basin pipe network as the target variable. Perform data cleaning, remove outliers, and divide them into a training set and a validation set according to a ratio of 7:3; S322: Calculate the cumulative amount of pollutants through the linear accumulation formula, and calculate the pollutant concentration at the outlet through the Logistic scouring model; S323: Use the mean absolute error as the loss function and the stochastic gradient descent method as the optimizer. Set the training parameters for the pollutant fate module and perform model training, and optimize the module parameters through the backpropagation algorithm; The training parameters of the pollutant fate module are: The initial value of C0 is 3.04219055811343 mg / m³; K B The initial value is 0.907313539587267% / d; The initial value of d is 0.895373541685726% / d; The initial value of C1 is 0.227983188296579; The initial value of C2 is 0.262647971835829; The initial value of B1 is 0.40876438184221; The initial value of B2 is 0.630350367766635; The learning rate is defaulted to 0.00001, and the number of training epochs is defaulted to 1; S314: Based on the validation set, evaluate the performance of the model, use the logarithmic difference between the predicted value and the actual value as the evaluation index, and verify the fitting ability of the module; When LD < 100%, it is determined that the model meets the requirements and runs normally; When LD ≥ 100%, check whether there are missing values and outliers in the model input and output values. During model training, increase the internal parameters of the model and repeat the training until the probability space that the model can represent covers all possibilities corresponding to the internal runoff characteristics of the basin.

[0028] The present invention also provides a method for predicting the migration of pollutant non-point sources based on a gated recurrent unit neural network, which is characterized by comprising the following steps: S1. Collect the current data of the area to be measured, and then perform standardized integration; The current data of the area to be measured includes: the current land use type, rainfall, pollutant usage, and elevation data of the area to be measured; S2. Input the current data of the area to be measured after standardized integration into the pollutant non-point source migration prediction model constructed by using the method according to any one of claims 1 to 10, obtain the prediction result, and complete the prediction; The prediction result includes: the runoff time series at the outlet of the sub-basin pipe network, the pollutant concentration time series at the outlet of the sub-basin pipe network, and the total pollutant migration amount at the outlet of the sub-basin pipe network.

[0029] The present application has at least the following beneficial technical effects: The present invention proposes a non-point source pollution simulation and prediction method with strong applicability and adaptable to diverse scenarios, which precisely solves the problem of simulating urban non-point source pollution through modular design. This method combines historical data such as land use type, rainfall, and pollutant usage to construct a hydrological simulation module, a hydraulic simulation module, and a pollutant fate module, and can flexibly meet the characteristic requirements of different urban areas and pollutants. For example, through standardized integration and targeted training, the model can be adjusted according to different land use types, pollutant concentrations, and regional characteristics, making it widely applicable in complex urban scenarios.

[0030] The present invention innovatively combines a physical model with a data-driven model, calculates surface runoff through the Horton infiltration formula and the Manning formula, and simultaneously introduces a gated recurrent unit (GRU) neural network to predict the runoff at the outlet diameter of the pipe network. This hybrid modeling method that combines physical basis and data-driven not only overcomes the dependence of traditional physical models on complex pipeline layout data but also can efficiently simulate the pollutant migration process in a large-scale basin, improving the accuracy and calculation efficiency of the model and being applicable to the study of non-point source pollution in urban areas of different scales.

[0031] The present invention adopts modular design, supports efficient expansion and optimization. The hydrological simulation module and the hydraulic simulation module are encapsulated into a non-point source runoff simulation module, and the pollutant fate module is combined to achieve accurate prediction of pollutant concentration. By adjusting the pollutant accumulation formula and the scouring formula, this method can adapt to the pollutant behavior characteristics of different environmental fate processes. At the same time, the optimized module training scheme significantly improves the convergence effect and expansion ability of the model parameters, providing strong technical support for the applicability of the model in a large-scale urban area.

[0032] The present invention has the advantages of high efficiency and low cost, and is convenient for wide promotion. By optimizing the acquisition and training of calibration data for key sub-watersheds, this method effectively reduces the costs of laboratory testing and large-scale sampling. The prediction results include the runoff time series, pollutant concentration time series, and total pollutant migration volume, providing comprehensive data support for the monitoring and management of urban non-point source pollution. This method is suitable for popularization and application in urban environmental pollution control, providing an efficient and reliable solution for non-point source pollution research. Description of the Drawings

[0033] Figure 1 It is a sub-watershed division map of Guangzhou City adopted in the embodiment of the model constructed by the present invention; Figure 2 It is a map of pesticide usage data for each region of Guangzhou City adopted in the embodiment of the model constructed by the present invention; Figure 3 It is a simulation effect diagram of the non-point source runoff module adopted in the embodiment of the model constructed by the present invention; Figure 4 It is a performance diagram of the pollutant fate model in the training period and verification period in the embodiment of the model constructed by the present invention; Figure 5 It is a distribution map of pesticide (target pollutant) usage and emissions in the urban sub-watersheds of Guangzhou City generated in the embodiment of the present invention; Figure 6 It is a runoff time series diagram of the outlet of the sub-watershed pipe network in the application example of the present invention; Figure 7 It is a pollutant concentration time series diagram of the outlet of the sub-watershed pipe network in the application example of the present invention; Figure 8 It is the land use type data of Guangzhou City in the application example of the present invention; Figure 9 It is the elevation data of Guangzhou City in the application example of the present invention; Figure 10 It is the 1h time series data of the weather in the sub-watersheds of Guangzhou City in the application example of the present invention; Figure 11 It is the original pesticide usage data of Panyu District, Guangzhou City in the application example of the present invention; Figure 12 It is the sub-watershed characteristic data of Guangzhou City in the application example of the present invention; Figure 13 It is the 10min time series data of the rainfall in the sub-watersheds of Guangzhou City in the application example of the present invention; Figure 14 It is the data of pesticide usage per unit green area for each region of Guangzhou City in the application example of the present invention. Detailed Embodiments

[0034] To make the purpose, technical solutions, and advantages of the present application more clear, the technical solutions in the embodiments of the present application will be described in more detail below in conjunction with the accompanying drawings in the embodiments of the present application. In the drawings, the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The described embodiments are some, but not all, of the embodiments of the present application. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as a limitation to the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application. The embodiments of the present application will be described in detail below in conjunction with the accompanying drawings.

[0035] The present application discloses a method for constructing a pollutant non-point source migration prediction model based on a gated recurrent unit neural network (GRU), which specifically includes the following steps: S1. Collect historical data of the area to be measured and divide it into complex sub-basin data, and then perform standardized integration to form a feature data set.

[0036] The sub-basin data is comprehensive information that includes the geographical, land cover, terrain, climate, and environmental characteristics of each sub-basin.

[0037] The historical data includes: land use type, rainfall, pollutant usage, and elevation data.

[0038] For example Figure 1 , taking Guangzhou City as an example, it is divided into 3,722 sub-basins in total. The sub-basin data includes: the number of the sub-basin, the name of the area to which the sub-basin belongs, the area of the area to which the sub-basin belongs, the area of the sub-basin, the proportion of the green area in the area of the sub-basin, the proportion of the impervious surface area in the area of the sub-basin, the average slope of the sub-basin, the rainfall of the sub-basin, and the pollutant usage of the sub-basin. The historical data includes: land use type, rainfall, and pollutant usage.

[0039] Further, the sub-basins are divided by the following method: S141. Use the vector data of roads and water bodies to divide the target basin, that is: perform the operation of "line to surface" on the road vector data and water body vector data in the study area at the same time, initially divide the sub-basins, and generate a planar area for analysis.

[0040] S142. Calculate the shortest distance between the initially divided sub-basins and the main river. If there are too many sub-regions divided for the first time and the subsequent calculation amount needs to be reduced, merge them according to the distance between the sub-basins divided for the first time and the main river, that is: judge the main river closest to each sub-basin and merge the sub-basins with the same closest main river.

[0041] S143. Identify mountainous areas using the terrain difference extraction tool of elevation data (specifically, digital elevation model, hereinafter referred to as DEM) and the Raster Calculator, exclude them, retain the urban area as the target sub-basin, and calculate the average slope of the sub-basin to complete the division.

[0042] To screen the usage locations of target pollutants, that is, urban areas, other irrelevant areas, that is, mountainous areas, need to be deleted. The specific method is to use the "Focal Statistics" tool to extract the terrain difference within 900 meters using the "Range" statistical type with a radius of 30 pixels for the DEM layer, and then use the "Raster Calculator" tool to convert the raster with an elevation of more than 100 meters to 1 as the mountainous area, and the remaining part to 0 as the urban area.

[0043] Use the "Convert Raster To Polygon" to convert the area with values into a polygon vector layer, start data editing, select the area with a value of 1 according to "Select Layer By Attribute", and then use the "Erase" toolbox to remove this part of the mountainous area. The remaining area is the sub-basin of the urban area.

[0044] Based on the DEM data, use a slope calculation tool (such as the "Slope" tool) to extract the slope value of each grid cell, and count the grid slopes within each sub-basin to calculate the average slope of the sub-basin. The result is stored in the sub-basin database with the field named "slope".

[0045] The data of land use types is from the Finer Resolution Observation and Monitoring-Global Land Cover (FROM-GLC) dataset (Gong et al., 2019). This dataset is a high-resolution global land cover data with a spatial resolution of 10 meters and a file type of.TIF raster data format. The land use types in the dataset are divided into the following 10 categories: farmland, forest, grassland, shrubland, wetland, water area, tundra, impervious surface, bare land, and ice and snow.

[0046] The standardization and integration of land use types include the following steps: S111. Convert raster data to vector data. In the ArcMap 10.6.1 software, use the Raster toPolygon tool to convert the FROM-GLC dataset from raster data format to vector data. After conversion, the generated land use vector data contains the boundary information and its classification attributes of each land use type, facilitating more convenient screening of regional types and spatial calculations in subsequent processing.

[0047] S112. Extract key land use type data.

[0048] In the attribute table of the generated land use vector data, according to the classification information of land use types, select the types relevant to this study: forests, grasslands, and shrubs are defined as the "green land" type (i.e., green land data), and exported as a new vector data set. Impervious surfaces are separately extracted as another group of vector data sets and defined as the "impervious" type (i.e., impervious surface data), representing the impervious surface features in this area.

[0049] Other land types can be classified as "others" and are not used, so they can be deleted. This step extracts the target types by filtering the classification fields in the attribute table.

[0050] S113. Calculate the proportion of land use types in the sub - watersheds. Use the Tabulate Intersection tool, with the sub - watershed data as the input area data (Input Zone Features), and the "green land data" and "impervious surface data" generated in S112 as the input class features respectively. This tool calculates the intersection area of each land use type in each sub - watershed and generates a percentage data table, recording the proportion of the areas of green land and impervious surfaces in each sub - watershed.

[0051] S114. Integrate the proportion data of land use types. Merge the percentage data table of land use types obtained in S113 with the attribute table of the sub - watershed data. Through the merge operation, directly associate the calculated proportion information of green land and impervious surfaces to the sub - watershed dataset, thus completing the standardized integration of land use type data. The purpose of this step is to ensure that each record in the sub - watershed contains the corresponding land use type proportion data, facilitating direct calls in the hydrological simulation module and subsequent analyses.

[0052] The rainfall data is sourced from the kilometer - level grid weather historical data of Xinzhi Weather. This data has a spatial resolution of 1×1 km, a time step of 1 hour, a rainfall unit of millimeters, and a file format of.csv. Each.csv file records the rainfall time - series data of a rainfall point. To correspond the rainfall data with the sub - watersheds in the study area, use the Create Fishnet tool in ArcMap10.6.1 to generate a grid with appropriate density for the study area, and use the center points of the grids as the representative positions of the rainfall points.

[0053] Subsequently, through spatial analysis, each sub - watershed is matched with the nearest rainfall point. In the attribute table of the sub - watershed, the corresponding rainfall point number is recorded to establish the association between the rainfall point and the sub - watershed. In this way, rainfall data can be directly called through the rainfall point number and used for subsequent calculations.

[0054] The standardization and integration of rainfall amount includes the following steps: S121. Use the linear interpolation method to convert the rainfall data recorded every hour into data recorded every 10 minutes. The specific formula is as follows:

[0055] In formula (1), p is the rainfall amount at the current time (mm), P0 is the rainfall amount at the previous time point (mm), P1 is the rainfall amount at the next time point (mm), t is the current time (min), t1 is the time at the previous time point (min), and t2 is the time at the next time point (min).

[0056] For each rainfall point, convert the old.csv file containing hourly rainfall amounts into a new.csv file containing rainfall amounts every 10 minutes. The file is named with the rainfall point number. Through the rainfall point number in the sub - watershed attribute table, call the corresponding rainfall amount.csv file to obtain the rainfall amount data.

[0057] S122. Match the rainfall amount at the current time with the corresponding sub - watershed data to obtain the sub - watershed rainfall time series and complete the standardization and integration.

[0058] The above process can be implemented through the following Python programming language.

[0059] Code #1 from _parameters import parameters from _method import write_statistic_data from _algorithm import get_missing_rainfall import pandas as pd import os import datetime # Write csv data def write_statistic_data(path: str, lists: list): with open(path, "a+") as f: f.write(",".join([str(i) for i in lists]) + "\n") f.close() # Calculate missing rainfall data based on previous and subsequent data def get_missing_rainfall(time, file_data): interval = datetime.timedelta(hours = 1) # Calculate backward interval_pre_num = 0 time_pre = time while True: interval_pre_num += 1 time_pre = time_pre - interval if str(time_pre).replace(" ", "T") in file_data.index: rainfall_pre = file_data.loc[str(time_pre).replace(" ", "T"), " precip"] break elif interval_pre_num>3: rainfall_pre = 0 break # Calculate forward interval_next_num = 0 time_next = time while True: interval_next_num += 1 time_next = time_next + interval if str(time_next).replace(" ", "T") in file_data.index: rainfall_next = file_data.loc[str(time_next).replace(" ", "T"), " precip"] break elif interval_next_num>3: rainfall_next = 0 break # Output the result proportionally return rainfall_pre * interval_next_num / ( interval_next_num + interval_pre_num) + rainfall_next * interval_pre_ num / ( interval_next_num + interval_pre_num) # 2 - Decompose the time into 10 min if __name__ == "__main__": rainfall_output_dir = parameters.rainfall_output_dir rainfall_output_1_dir = os.path.join(rainfall_output_dir, "1_combine_ files") rainfall_output_2_dir = os.path.join(rainfall_output_dir, "2_split_ time_files") files_id = range(1, 217) for file_id in files_id: rainfall_file_path = os.path.join(rainfall_output_1_dir, "guangzhou_ {}.csv".format(str(file_id))) rainfall_file_data = pd.read_csv(rainfall_file_path, delimiter=",", encoding="GBK", index_col="last_update") current_time = datetime.datetime.fromisoformat(rainfall_file_ data.index[0]) end_time = datetime.datetime.fromisoformat(rainfall_file_data.index[- 1]) rainfall_output_2_path = os.path.join(rainfall_output_2_dir, "split_ guangzhou_{}.csv".format(str(file_id))) write_statistic_data(rainfall_output_2_path, ["time", "prec"]) while True: # Current rainfall current_time_str = str(current_time).replace(" ", "T") if current_time_str in rainfall_file_data.index: current_prec = rainfall_file_data.loc[current_time_str, "precip"] # If rainfall is empty if str(current_prec) == "nan": current_prec = 0 else: current_prec = get_missing_rainfall(time=current_time, file_data= rainfall_file_data) # Rainfall in the previous hour last_time = current_time - datetime.timedelta(hours=1) last_time_str = str(last_time).replace(" ", "T") if last_time_str in rainfall_file_data.index: last_prec = rainfall_file_data.loc[last_time_str, "precip"] # If rainfall is empty if str(last_prec) == "nan": last_prec = 0 else: last_prec = get_missing_rainfall(time=last_time, file_data=rainfall_ file_data) # Rainfall in the next hour next_time = current_time + datetime.timedelta(hours=1) next_time_str = str(next_time).replace(" ", "T") if next_time_str in rainfall_file_data.index: next_prec = rainfall_file_data.loc[next_time_str, "precip"] # If rainfall is empty if str(next_prec) == "nan": next_prec = 0 else: next_prec = get_missing_rainfall(time=next_time, file_data=rainfall_ file_data) # Calculate rainfall trend last_trend = current_prec - last_prec / 6 current_prec_10_trend = last_prec + last_trend * 4 current_prec_20_trend = last_prec + last_trend * 5 current_prec_30_trend = last_prec + last_trend * 6 next_trend = next_prec - current_prec / 6 current_prec_40_trend = current_prec + next_trend * 1 current_prec_50_trend = current_prec + next_trend * 2 current_prec_60_trend = current_prec + next_trend * 3 trend_sum = sum([current_prec_10_trend, current_prec_20_trend, current_prec_30_trend, current_prec_40_trend, current_prec_50_trend, current_prec_60_trend]) if trend_sum == 0: trend_sum = 1 current_prec_10 = current_prec * current_prec_10_trend / trend_sum current_prec_20 = current_prec * current_prec_20_trend / trend_sum current_prec_30 = current_prec * current_prec_30_trend / trend_sum current_prec_40 = current_prec * current_prec_40_trend / trend_sum current_prec_50 = current_prec * current_prec_50_trend / trend_sum current_prec_60 = current_prec * current_prec_60_trend / trend_sum write_statistic_data(rainfall_output_2_path, [str(last_time + datetime.timedelta(minutes=10)), str(current_prec_10)]) write_statistic_data(rainfall_output_2_path, [str(last_time + datetime.timedelta(minutes=20)), str(current_prec_20)]) write_statistic_data(rainfall_output_2_path, [str(last_time + datetime.timedelta(minutes=30)), str(current_prec_30)]) write_statistic_data(rainfall_output_2_path, [str(last_time + datetime.timedelta(minutes=40)), str(current_prec_40)]) write_statistic_data(rainfall_output_2_path, [str(last_time + datetime.timedelta(minutes=50)), str(current_prec_50)]) write_statistic_data(rainfall_output_2_path, [str(last_time + datetime.timedelta(minutes=60)), str(current_prec_60)]) current_time += datetime.timedelta(hours=1) if (end_time.timestamp() - current_time.timestamp())<0: break The sources of pollutant usage data vary depending on the types of target pollutants. Taking urban mosquito repellents as an example, users can apply for relevant data through the government's information disclosure system upon request (such as the information disclosure system in Guangdong Province). The data usually obtained includes the total usage mass (kg) of mosquito repellents and the proportion of active ingredients (%). By multiplying the total usage mass by the proportion of active ingredients, the total usage amount of the target ingredient can be calculated. The calculation formula is as follows: Mass of target ingredient (kg) = Total usage mass (kg) * Proportion of active ingredients (%).

[0060] In addition, for the usage amount of neonicotinoid insecticides, users can calculate the total sales volume (kg) in the target area by researching the regional sales data of major manufacturers and multiply it by the proportion of active ingredients to obtain the usage amount of the target ingredient. The above methods are applicable to the calculation of the usage amount of various types of pollutants and can provide accurate input data.

[0061] The standardized integration of pollutant usage amounts includes the following steps (in this embodiment, it is assumed that the unit of the initially obtained pollutant usage amount is kg): S131. Define the main application areas and collect the total amount of pollutants. Define the main application areas according to specific pollutants. For example, for mosquito repellents and insecticides, assume that their main application areas are within a 5-meter range outside urban green spaces. Through spatial analysis, using the green space vector data as input and the Buffer tool in ArcMap 10.6.1, set the buffer range to 5 meters to generate the vector data of the green space buffer area, and only retain the outer part of the green space.

[0062] S132. Obtain the area proportion of the main application areas in the sub-basins through the spatial intersection calculation tool. Use the sub-basin data as the input area data (Input Zone Features) and the generated vector data of the green space buffer area as the input type data (Input Class Features). Through the Tabulate Intersection tool, calculate the area proportion of the green space buffer area in each sub-basin and generate a percentage data table.

[0063] S133. Incorporate the area proportion of the green space buffer area in the sub-basins into the sub-basin database. Merge the percentage data table obtained in step S132 with the sub-basin data table to update the sub-basin database so that it contains the area proportion information of the main application areas in each sub-basin.

[0064] S134. Calculate the pollutant usage of each sub - basin according to the area proportion. According to the area proportion of the green buffer area in the sub - basin database, allocate the total pollutant usage to each sub - basin to complete the standardized integration of pollutant usage. Use the pollutant usage data of each sub - basin as the input data for subsequent pollutant concentration calculation.

[0065] Specifically, in this embodiment, as Figure 2 shown, the usage database of pesticides (pollutants) in each region of Guangzhou is made into a table, where the abscissa is the name of each region in Guangzhou; the ordinate is the name of pesticides, and the unit is μg / m 2 / day.

[0066] S2. Establish a pollutant non - point source migration prediction model through the characteristic data set. This model consists of a hydrological simulation module, a hydraulic simulation module, and a pollutant fate module. In addition, the hydrological simulation module and the hydraulic simulation module are jointly encapsulated as a non - point source runoff simulation module.

[0067] Specifically, the hydrological simulation module is used to simulate the complete process of rainfall from landing to flowing into the drainage pipe system, with data such as rainfall, land use area proportion, and average slope as inputs. To more accurately reflect the response of different surface types to rainfall, the module divides the land use types of sub - basins into two major categories: green land and impervious surface, and conducts modeling based on the following scenarios. For green land, part of the rainwater will infiltrate into the soil interior, affected by the soil infiltration capacity; when the rainfall intensity exceeds the infiltration capacity of the green land, the excess rainwater will overflow and flow along the terrain to the inlet of the drainage pipe. For the impervious surface, the rainwater is first temporarily stored, and when the rainfall exceeds the storage capacity, the remaining rainwater will form runoff and flow along the terrain to the inlet of the drainage pipe.

[0068] To achieve a detailed simulation of the above - mentioned process, the hydrological simulation module is divided into an infiltration module and a depression storage module. The infiltration module uses the Horton infiltration formula to dynamically describe the change of soil infiltration capacity by calculating the hourly infiltration rate. This module uses the Horton infiltration capacity curve formula to calculate the instantaneous infiltration rate during the rainfall process and calculates the cumulative infiltration amount of the soil based on the cumulative infiltration curve formula to determine the surface runoff of the green land. The calculation formula is as follows:

[0069] In formula (2), f t is the infiltration ratio at time t (%), f c is the stable infiltration ratio (%), f0 is the initial infiltration rate, k is the infiltration rate decay coefficient, which is related to the physical properties of the soil, and t is the time (h).

[0070] In the discontinuous model designed in this application with time points as steps, the original Horton infiltration is difficult to apply. Therefore, the infiltration rate in the original Horton infiltration formula is changed to an infiltration ratio here.

[0071] After obtaining the infiltration ratio at the current time point, the runoff generated by the green space can be expressed as:

[0072] In formula (3), R g is the runoff of the green space data at time t (m 3 ), i t is the rainfall at time t (mm), a is the area of the region occupied by the green space (m 2 ), f t is the infiltration ratio at time t (%).

[0073] The depression storage module is mainly designed for impervious surfaces and is used to simulate the accumulation and runoff process of rainfall on impervious surfaces. By calculating the storage capacity, the module can estimate the rainwater storage effect on impervious surfaces at the beginning of rainfall. When the rainfall exceeds the storage capacity, the module generates runoff according to the terrain analysis and determines the inlet flow rate into the drainage pipe. The depression storage module regards the impervious surface as a non-linear reservoir, so the Manning formula is used to calculate the surface runoff of the impervious surface:

[0074] In formula (4), R i is the runoff of the impervious surface data (m 3 ), W is the width of the sub-catchment area (m), S is the average slope of the sub-catchment area (m / m), n is the Manning roughness coefficient, d is the water depth of the land plane (m), d p is the initial surface water depth (m).

[0075] The runoff from the two major land use types is jointly discharged into the street gutter and pipe through the inlet of the drainage pipe. As the total surface runoff of the sub-basin, it is input into the hydraulic simulation module, and the calculation formula is as follows:

[0076] In formula (5), R s is the total surface runoff of the sub-basin, and the data form is a time series (i.e., the total surface runoff time series of the sub-basin). The above process can be implemented through the following Python language programming.

[0077] Code #2 import datetime import os import numpy as np import pandas as pd import torch from torch import nn torch.set_default_tensor_type(torch.DoubleTensor) # Data normalization - [0, 1] class BatchNorm1d(nn.Module):# Only normalize one input item, does not support dynamic modification

n, 1, 1

[0078] During the model development process, due to limited training data, the traditional GRU failed to achieve the expected results in terms of training efficiency and result accuracy. To solve this problem, this method optimizes the GRU structure, removes some parameters and activation functions of the reset gate in the traditional GRU, thereby simplifying the model structure, improving the calculation efficiency and enhancing the training effect.

[0079] The input data of the hydraulic simulation module mainly comes from the runoff volume calculated by the hydrological simulation module. After being processed by the simplified GRU model, these runoff data generate the predicted runoff volume at the outlet of the sub - watershed pipe network. These prediction results not only improve the understanding of the pipeline transportation process but also provide important input data for the subsequent water quality simulation module, ensuring the accuracy and efficiency of the overall simulation system. In the model, the specific calculation formula is as follows:

[0080] In formulas (6), (7) and (8), W r is the reset gate, W z is the update gate, initially set as a one - dimensional tensor containing the value 1.0, set to requires_grad = True, indicating that the value of this tensor will participate in the backpropagation, so the gradient will be calculated. r z is the output of the reset gate at time t, z t is the output of the update gate at time t, [h t-1 ,x t and [r z ,x t are both concatenation operators for two vectors. x t is the total surface runoff at time t (m 3 ), h t is the hidden state at time t. Additionally, it is also the runoff volume at the outlet of the pipe network at time t, in the form of a time series (m 3, namely: the time series of the outlet diameter flow of the sub - basin pipe network).

[0081] The pollutant fate module consists of a pollutant accumulation module and a pollutant wash - off module, which is used to simulate the whole process of pollutants in the sub - basin from generation, accumulation to being washed off by rainwater and finally transported to the pipe network outlet with runoff. By modeling the behavior of pollutants under different conditions, this module provides key data support for the dynamic changes of pollutant concentration and flow direction.

[0082] The pollutant fate module is designed based on the following scenario: Pollutants are released and accumulated on the ground surface at a certain frequency during daily activities. During dry weather, pollutants gradually accumulate due to the lack of runoff; when rainfall occurs, these accumulated pollutants are washed off by rainwater, enter the pipe system with surface runoff, and finally are discharged with runoff at the pipe network outlet.

[0083] Specifically, the pollutant accumulation module is responsible for calculating the accumulation amount of pollutants under different weather conditions, considering the duration of drought and the pollutant release characteristics. The pollutant accumulation module can adopt different functions according to the actual situation. In this embodiment, a simple linear function is adopted as the pollutant accumulation method. Using the pollutant usage data in the database as input, the output pollutant accumulation load is used for the calculation of the pollutant wash - off module. The specific calculation formula is as follows:

[0084] In formula (9), P is the accumulation amount of pollutants (g), t is the previous drought time (day), K B is the pollutant accumulation rate per unit area (g / d), and d is the pollution removal rate (%).

[0085] In formula (13), C0 is the initial accumulation days (day), W c is the total usage amount of pollutants in the urban area where the sub - basin belongs (μg / m 2 / day), G c is the green area of the urban area where the sub - basin belongs (m 2 ), G sub is the green area of the sub - basin (m 2 ).

[0086] Furthermore, as Figure 5 shown, taking the distribution map of the usage and emission amounts of pesticides (target pollutants) in the sub - basins of the urban area of Guangzhou generated in this embodiment as an example, where: Usage amount (kg) = the total unit usage amount obtained from the investigation (μg / m 2 / day) * the area of 5m outside the green area of the sub - basin (m 2 ) * the number of drought days (day) * 10 -9 ; Emission (kg) = Concentration obtained from model calculation (ng / L) * Outlet flow rate of the effluent (m 3 ) * 10 -12 .

[0087] The pollutant wash-off module simulates the process of pollutants being washed from the surface into the pipe network based on rainfall intensity, surface type, and flow velocity. The pollutant wash-off module can adopt different methods according to the actual situation. In this embodiment, the Logistic wash-off model is adopted, and the output data of the hydraulic simulation module and the pollutant accumulation module are used as the input data of this module to calculate the pollutant concentration and complete a non-point source migration prediction. The specific formula is as follows:

[0088] In formulas (10), (11), and (12), W t is the wash-off pollution concentration at time t, and the data form is a time series (g / L, that is: the time series of pollutant concentration at the outlet of the sub-basin pipe network), Q t is the outlet flow rate of the sub-basin pipe network at time t (m 3 / s), δ t is the ratio of the pollution load available for wash-off at time t (g / g), B1 and B2 are both coefficients of the logistic curve, C1 and C2 are both coefficients of the exponential wash-off curve, V t is the cumulative runoff at time t (m 3 ), P0 is the initial pollutant accumulation in the catchment area before rainfall (g), and P wt is the cumulative pollutant load washed off at time t (g).

[0089] With the synergistic effect of the above two modules, the water quality simulation module can comprehensively evaluate the path and impact of pollutants from generation to emission, providing a scientific basis for water quality management and pollution control.

[0090] The above process can be implemented through the following Python language programming.

[0091] Code #3 import datetime import os import numpy as np import pandas as pd import torch from torch import nn torch.set_default_tensor_type(torch.DoubleTensor) # Limit range def clamp_value(value, min_value, max_value): value_return = value if value<min_value: value_return = min_value elif value>max_value: value_return = max_value return value_return # torch - Pesticide mosquito repellent prediction class PesticideTotalPredict(nn.Module): def __init__(self): super(PesticideTotalPredict, self).__init__() # Pollutant accumulation # self.accumulate_max = torch.nn.Parameter(torch.tensor ([40.440791974599927], requires_grad=True))# mg / m3 # self.register_parameter("accumulate_max", self.accumulate_max) # self.accumulate_n = torch.nn.Parameter(torch.tensor ([4.672340705329424], requires_grad=True))# day # self.register_parameter("accumulate_n", self.accumulate_n) # self.accumulate_m0_day = torch.nn.Parameter(torch.tensor ([3.05018056500963], requires_grad=True))# mg / m3 self.register_parameter("accumulate_m0_day", self.accumulate_m0_day) self.accumulate_degrade = torch.nn.Parameter(torch.tensor ([0.904118422535808], requires_grad=True))# % / d self.register_parameter("accumulate_degrade", self.accumulate_ degrade) # Spraying correction factor self.accumulate_usage_plus_870 = torch.nn.Parameter(torch.tensor ([1.08440819110883], requires_grad=True)) self.register_parameter("accumulate_usage_plus_870", self.accumulate_ usage_plus_870) self.accumulate_usage_plus_3718 = torch.nn.Parameter(torch.tensor ([1.58718310554308], requires_grad=True)) self.register_parameter("accumulate_usage_plus_3718", self.accumulate_usage_plus_3718) self.accumulate_usage_plus_3716 = torch.nn.Parameter(torch.tensor ([2.1140102328541], requires_grad=True)) self.register_parameter("accumulate_usage_plus_3716", self.accumulate_usage_plus_3716) # Pollutant scour self.scour_C1 = torch.nn.Parameter(torch.tensor([0.472591131935476], requires_grad=True)) self.register_parameter("scour_C1", self.scour_C1) self.scour_C2 = torch.nn.Parameter(torch.tensor([0.407377747374919], requires_grad=True)) self.register_parameter("scour_C2", self.scour_C2) self.scour_B1 = torch.nn.Parameter(torch.tensor([0.445718540597239], requires_grad=True)) self.register_parameter("scour_B1", self.scour_B1) self.scour_B2 = torch.nn.Parameter(torch.tensor ([0.00667679290725365], requires_grad=True)) self.register_parameter("scour_B2", self.scour_B2) def forward(self, runflow_data, usage_data, pesticide_name, params): # Read the incoming parameters if len(params) != 0: self.accumulate_m0_day.data = torch.tensor(params[0]) self.accumulate_degrade.data = torch.tensor(params[1]) self.accumulate_usage_plus_870.data = torch.tensor(params[2]) self.accumulate_usage_plus_3716.data = torch.tensor(params[3]) self.accumulate_usage_plus_3718.data = torch.tensor(params[4]) self.scour_C1.data = torch.tensor(params[5]) self.scour_C2.data = torch.tensor(params[6]) self.scour_B1.data = torch.tensor(params[7]) self.scour_B2.data = torch.tensor(params[8]) # Record data total_pesticide_output_tensor = torch.tensor([]) # Output pollutant concentration # Extract fixed area parameters area_name = runflow_data.loc[runflow_data.index[0], "name"] area_id = runflow_data.loc[runflow_data.index[0], "area_id"] greenland_5m = torch.tensor([runflow_data.loc[runflow_data.index[0], "greenland_5m"]]) water = torch.tensor([runflow_data.loc[runflow_data.index[0], " water"]]) impervious = torch.tensor([runflow_data.loc[runflow_data.index[0], " impervious"]]) area = torch.tensor([runflow_data.loc[runflow_data.index[0], " area"]]) slope = torch.tensor([runflow_data.loc[runflow_data.index[0], " slope"]]) classification = str(runflow_data.loc[runflow_data.index[0], " class"]) pesticide_usage = torch.tensor([usage_data.loc[pesticide_name, area_ name]]) greenland = greenland_5m * area # Initial time time = datetime.datetime.strptime(runflow_data.loc[runflow_data.index [0], "time"], "%Y-%m-%dT%H:%M") last_time_begin = datetime.datetime.strptime(runflow_data.loc [runflow_data.index[0], "time"], "%Y-%m-%dT%H:%M") # Modify the spraying volume # print(time) datetime.timedelta() # Modify the spraying volume if area_id == 870: pesticide_usage = pesticide_usage * self.accumulate_usage_plus_870 elif area_id == 3718: pesticide_usage = pesticide_usage * self.accumulate_usage_plus_3718 elif area_id == 3716: pesticide_usage = pesticide_usage * self.accumulate_usage_plus_3716 # Initial accumulation of pollutants ug / m3 accumulation = self.accumulate_m0_day * pesticide_usage * \ torch.pow(self.accumulate_degrade, self.accumulate_m0_day) * \ torch.pow(self.accumulate_retain, self.accumulate_m0_day) # Screen rain events 0 -? rain_num = runflow_data.loc[runflow_data.index[-1], "id"]# The last rainfall rain for rain_th in range(rain_num + 1):# For each rainfall runflow_data_th = runflow_data[runflow_data["id"] == rain_th]# Information for each rainfall (with many time points) v_t = 0# Flow accumulation # Calculate the time interval since the last rainfall time_begin_str = runflow_data_th.loc[runflow_data_th.index[0], " time"]# The first time point of each rainfall time_begin = datetime.datetime.strptime(time_begin_str, "%Y-%m-%dT% H:%M") time_interval = torch.tensor(# The time difference between the first time point of this rainfall and the first time point in total of the difference (time_begin.timestamp() - last_time_begin.timestamp()) / 60 / 60 / 24)# day # Pollutant accumulation (ug / m3) B = C1 * (1 - e^(-C2 * t)) accumulation = accumulation + pesticide_usage * time_interval accumulation = accumulation * torch.pow(self.accumulate_degrade, time_interval)# Degradation loss accumulation_ug = accumulation * greenland accumulation_sub = 0# Pollutants lost during this rainfall # Record the start time of this rainfall for use in calculating the time for the next rainfall last_time_begin_str = runflow_data_th.loc[runflow_data_th.index[0], " time"] last_time_begin = datetime.datetime.strptime(last_time_begin_str, "% Y-%m-%dT%H:%M") # Step in for i in runflow_data_th.index:# Each time point of each rainfall # Current time time_str = runflow_data_th.loc[i, "time"] time = datetime.datetime.strptime(time_str, "%Y-%m-%dT%H:%M") runflow = torch.tensor([runflow_data_th.loc[i, "runflow"]])# m3 per 10 minutes within # Pollutant scour (ug / L) W = C1 * Qt^C2 * (P0 * delta_t - Pwt) *** delta_t = 1 / (1 + B1 * e^ (-B2 * Vt)) v_t = v_t + runflow# m3 delta_t = 1 / (1 + self.scour_B1 * torch.pow(np.e, -self.scour_B2 * v_t)) if greenland != 0: scouring_ug = self.scour_C1 * torch.pow(runflow / 10 / 60, self.scour_C2) * ( (accumulation_ug * delta_t - accumulation_sub) / greenland) else: scouring_ug = torch.tensor([0.0]) scouring_ug = clamp_value(scouring_ug, torch.tensor([0.0]), np.inf) # clamp_value(value, min_value, max_value) scouring = scouring_ug# ug / L # Record the lost pollutants accumulation_sub = accumulation_sub + scouring_ug * 1000 accumulation_sub = clamp_value(accumulation_sub, torch.tensor([0.0]), np.inf) # Record data total_pesticide_output_tensor = torch.cat((total_pesticide_output_ tensor, scouring), 0) # print("accumulation_ug:", str(accumulation_ug.item())) # print("accumulation_sub:", str(accumulation_sub.item())) time_rainfall_length = (time.timestamp() - time_begin.timestamp()) / 60 / 60 / 24 # Remove the degraded pollutants during this rainfall time accumulation_ug = (accumulation_ug - accumulation_sub) * torch.pow (self.accumulate_degrade, time_rainfall_length) accumulation = accumulation_ug / greenland total_pesticide_output_np = total_pesticide_output_tensor.detach() .numpy() return total_pesticide_output_np S3. Train the hydraulic simulation module and the pollutant fate module, and optimize the module parameters.

[0092] Specifically, the training method of the non-point source runoff simulation module includes the following steps: S311. Use the sub - basin rainfall time series as input data and the sub - basin pipe network outlet diameter flow time series data as the target variable. Clean the data, remove outliers, and divide it into a training set and a validation set at a ratio of 7:3.

[0093] S312. Calculate surface runoff through the Horton infiltration formula and the Manning formula as the input features of the model. Use the rectified linear unit (ReLU) as the activation function and predict the sub - basin pipe network outlet diameter flow through the gated recurrent unit neural network (GRU). S313. Use the mean squared error (MSE) as the loss function and the stochastic gradient descent method (SGD) as the optimizer. Set the training parameters for the non - point source runoff simulation module and perform model training. During training, calculate the gradient through the backpropagation algorithm and gradually optimize the module parameters to reduce the mean squared error.

[0094] Among them, the training parameters of the non - point source runoff simulation module are: f0 is defaulted to 0.9%; f c is defaulted to 0.5%; k is defaulted to 0.2; n is defaulted to 0.9; d p is defaulted to 0.0001 m; W r and W z When resetting the weight of the reset gate, it is defaulted to 1.0; W r and W z When resetting the weight of the update gate, it is defaulted to 1.0; The learning rate is defaulted to 0.00001, and the number of training epochs is defaulted to 60.

[0095] In this embodiment, setting the weights of W r and W z to 1.0 is not the only parameter, but only a reference value that can make the model achieve the expected effect. Users can adopt strategies such as random initialization during actual use and flexibly adjust the initial parameter values according to needs.

[0096] S314. Based on the validation set, evaluate the performance of the module. Use the Nash - Sutcliffe efficiency coefficient (NSE) and the coefficient of determination (R 2 ) as evaluation indicators to verify the fitting ability of the module.

[0097] When NSE > 0.36 and R 2When it is > 0.54, according to the literature "Green C H, Tomer M D, Di Luzio M, Arnold J G. Hydrologic evaluation of The Soil and Water Assessment Tool for a large tile-drained watershed in Iowa [J]. TRANSACTIONS OF THE ASABE, 2006, 49(2): 413-422." and "Van Liew M W, Garbrecht J. Hydrologic simulation of the Little Washita River Experimental Watershed using SWAT [J]. Journal of the American Water Resources Association, 2003, 39(2): 413-426.", it can be known that the model meets the requirements and runs normally.

[0098] When NSE ≤ 0.56 or R 2 ≤ 0.54, check whether there are missing values and outliers in the model input value (rainfall) and output value (outlet flow rate). If so, reasonably fill or remove the missing positions, and remove or smooth the outliers. During model training, increase the internal parameters of the model and repeat the training until the probability space that the model can represent covers all possibilities corresponding to the internal runoff characteristics of the watershed.

[0099] As Figure 3 shown, this chart further shows the simulation effect of the non-point source runoff module. The abscissa of the chart is time, the data above the chart is rainfall, and the data below the chart is runoff. Among them, the gray line is the actual runoff, and the dark gray line is the simulated runoff. The left side of the dotted line in the figure is the simulation effect during the model training (calibration) stage, and R 2 and NSE are 0.520 and 0.516 respectively. The right side of the dotted line is the simulation effect during the model verification stage, and R 2 and NSE are 0.557 and 0.554 respectively.

[0100] The training of the pollutant fate module includes the following steps: S321. Use the pollutant usage data in each sub-watershed and the time series data of the outlet flow rate of the sub-watershed pipe network as input data, and the time series of the pollutant concentration at the outlet of the sub-watershed pipe network as the target variable. Clean the data and remove outliers, and divide them into a training set and a verification set according to a ratio of 7:3.

[0101] S322. Calculate the cumulative amount of pollutants using the linear accumulation formula and calculate the pollutant concentration at the outlet using the Logistic flushing model.

[0102] S323. Use the mean absolute error (MAE) as the loss function and the stochastic gradient descent (SGD) as the optimizer to set the training parameters for the pollutant fate module and perform model training. During the training process, calculate the gradient through the backpropagation algorithm and gradually optimize the module parameters to reduce the mean square error.

[0103] Among them, the training parameters of the pollutant fate module are: The initial value of C0 is 3.04219055811343 mg / m³; K B The initial value is 0.907313539587267% / d; The initial value of d is 0.895373541685726% / d; The initial value of C1 is 0.227983188296579; The initial value of C2 is 0.262647971835829; The initial value of B1 is 0.40876438184221; The initial value of B2 is 0.630350367766635; The learning rate is defaulted to 0.00001 and the number of training epochs is defaulted to 1.

[0104] S314. Evaluate the performance of the module based on the validation set, and use the logarithmic difference (LD) between the predicted value and the actual value as the evaluation index to verify the fitting ability of the module.

[0105] As Figure 4 shown, when LD < 50%, it represents the proportion of the logarithmic difference less than 50% in all pollutant data; when LD < 70%, it represents the proportion of the logarithmic difference less than 70% in all pollutant data; when LD < 100%, it represents the proportion of the logarithmic difference less than 100% in all pollutant data.

[0106] When LD < 100%, according to the literature "Cao H Y, Tao S, Xu F L, Coveney R M, Cao J, Li B G, Liu W X, Wang X J, Hu J Y, Shen W R, Qin B P, Sun R. Multimedia fatemodel for hexachlorocyclohexane in Tianjin, China [J]. Environmental Science & Technology, 2004, 38(7): 2126-2132.", it indicates that the error between the model prediction value and the actual value is within the acceptable range; when 50% < LD < 70%, it means that the error between the model prediction value and the actual value is negligible, and the module has general applicability and operates normally in most cases.

[0107] When LD ≥ 100%, check whether there are missing values and outliers in the model input value (effluent diameter flow) and output value (concentration). If so, reasonably fill or remove the missing positions, and remove or smooth the outliers. During model training, increase the internal parameters of the model and repeat the training until the probability space that the model can represent covers all possibilities corresponding to the internal runoff characteristics of the basin.

[0108] This application also provides a method for predicting the migration of pollutant non-point sources using the above model, including the following steps: S1. Collect the current data of the area to be measured and then perform standardized integration. The current data of the area to be measured includes: the current land use type, rainfall, pollutant usage, and elevation data of the area to be measured.

[0109] S2. Input the current data of the area to be measured that has been standardized and integrated into the pollutant non-point source migration model to obtain the prediction result and complete the prediction.

[0110] The prediction results include: the time series of runoff volume at the outlet of the sub-basin pipe network, the time series of pollutant concentration at the outlet of the sub-basin pipe network, and the total pollutant migration volume at the outlet of the sub-basin pipe network.

[0111] Application Example In this experiment, taking multiple rainfall events in the urban area of Guangzhou as an example, verify the actual prediction ability of the pollutant non-point source migration prediction model constructed by the above method. The specific steps are as follows: S1. Collect the current data of the urban area of Guangzhou. The current data includes: land use type ( Figure 8 ), rainfall ( Figure 10 ), pollutant usage ( Figure 11 ), and elevation data ( Figure 9). It covers the application amounts of eleven target pesticides. Meanwhile, the test data cover different rainfall intensities and sub - basin characteristics.

[0112] Standardize and integrate the current data. The specific integration results are as Figures 12 - 14 shown.

[0113] S2. Input the current data that has been standardized and integrated into the pollutant non - point source migration model to obtain the prediction results and complete the prediction.

[0114] The direct output results of the prediction are: the time series of runoff volume at the outlet of the sub - basin pipe network ( Figure 6 ) and the time series of pollutant concentration ( Figure 7 ).

[0115] From Figure 6 , the runoff response of the model after a rainfall event can be observed. For example, when the rainfall increased from 0.00 mm to 0.22 mm at 2022 - 06 - 29T13:10 - 14:00, the runoff volume began to increase from 0 m³ to 4.06 m³. Subsequently, with the increase in rainfall, the runoff volume gradually rose and reached a peak of 76.96 m³ at 2022 - 06 - 29T15:20. After the rainfall ended, the runoff volume also gradually decreased to 0 m³. This trend is consistent with the actual hydrological process, indicating that the model can capture the dynamic relationship between rainfall and runoff volume and shows good accuracy in its response to runoff when the rainfall intensity changes.

[0116] From Figure 7 , it can be observed that the pollutant concentration fluctuates with the change of runoff volume. For example, at 2022 - 06 - 29T14:10 with a rainfall of 0.45 mm and a runoff volume of 8.93 m³, the THM concentration is 13.12 ng / L, while the CLO concentration is 2.07 ng / L. As the runoff volume further increases (such as at 2022 - 06 - 29T15:20 with a runoff volume of 76.96 m³), the THM and CLO concentrations rise to 40.48 ng / L and 5.56 ng / L respectively. This change trend of pollutant concentration is highly consistent with the peak of runoff volume, indicating that the model can accurately reflect the pollutant migration characteristics and capture the impact of runoff volume changes on the concentrations of the same or different pollutants. This consistency further proves the reliable prediction ability of the model for complex hydrological pollution processes. On this basis, for each time step, multiply the corresponding runoff volume by the pollutant concentration to calculate the pollutant migration amount within that time step. Subsequently, sum up the migration amounts within all time steps to obtain the total pollutant migration amount at the outlet of the sub - basin pipe network.

[0117] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. The software used for Code #1-3, sub-basin division, and data standardization integration, as well as the form of code writing, are only shown as a method to achieve the necessary functions. If technicians implement the same functions through other software or software integration plugins according to their own needs, it should be regarded as within the technical scope disclosed in the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for constructing a pollutant non-point source migration prediction model based on a gated recurrent unit neural network, characterized in that: The following steps are involved: S1. Collect historical data of the area to be measured and divide it into multiple sub-basin data, and then standardize and integrate them to form a feature data set; The sub-basin data is comprehensive information on the geography, land cover, topography, climate and environmental characteristics of each sub-basin; The historical data include: land use type, rainfall, pollutant usage and elevation data; S2. Establishing a pollutant surface source migration prediction model through the characteristic data set; The pollutant surface source migration prediction model includes: a hydrological simulation module, a hydraulic simulation module and a pollutant fate module; The hydrological simulation module is used to calculate the total surface runoff time series of the sub-basin based on rainfall and land use type; The hydraulic simulation module is used to calculate the flow time series of the outlet diameter of the sub-basin pipe network according to the total surface flow of the sub-basin; The pollutant fate module is used to calculate the pollutant concentration time series of the sub-basin pipe network outlet according to the sub-basin pipe network outlet flow and pollutant usage; S3, training the hydraulic simulation module and pollutant fate module, and optimizing module parameters.

2. The construction method according to claim 1, characterized in that: The land use types described include: cropland, forest, grassland, shrubland, wetland, water area, tundra, impervious surface, bare land, and ice and snow; The standardized integration of land use types includes the following steps: S111, converting land use type data into land use vector data; S112, outputting the forest, grassland and bushes in the vector data as green space data; Outputting the impervious surface in the vector data as impervious surface data; S113, using the sub-basin data as input area data, and the green space data and impervious surface data as input type data, respectively input into a spatial intersection calculation tool to obtain percentage data of each land use type in the sub-basin; S114. Merge the percentage data with the sub-basin data to obtain the proportion of land use types in the sub-basin, and complete the standardized integration.

3. The construction method according to claim 1, characterized in that: The normalized integration of the rainfall includes the following steps: S121. Use linear interpolation to convert rainfall data recorded once every hour into rainfall data recorded every 10 minutes. The specific formula is as follows: In formula (1), p is the rainfall at the current time, P0 is the rainfall at the previous time point, P1 is the rainfall at the next time point, t is the current time, t1 is the time at the previous time point, and t2 is the time at the next time point; S122. Match the rainfall at the current time with the corresponding sub-basin data to obtain the sub-basin rainfall time series and complete the standardized integration.

4. The construction method according to claim 3, characterized in that: The standardized integration of the pollutant usage includes the following steps: S131, using the green space data as input to geographic information software, defining the main pollutant application areas in the green space according to the pollutant application characteristics, collecting the total pollutant amount, and generating green space buffer area vector data through spatial analysis; S132, using the sub-watershed data as the input area and the green space buffer area vector data as the input type, and obtaining the proportion of the green space buffer area in each sub-watershed through a spatial intersection calculation tool; S133, incorporating the area ratio of the green space buffer zone into the sub-watershed data, and updating each sub-watershed database; S134. Calculate the pollutant usage data in each sub-basin according to the proportion of the green space buffer area in the sub-basin, and complete the standardized integration of pollutant usage.

5. The construction method according to claim 4, characterized in that: The sub-basin data include: the number of the sub-basin, the name of the region to which the sub-basin belongs, the area of ​​the region to which the sub-basin belongs, the area of ​​the sub-basin, the proportion of the green area to the sub-basin area, the proportion of the impervious surface area to the sub-basin area, the average slope of the sub-basin, the rainfall in the sub-basin, and the pollutant usage of the sub-basin; The sub-basins are divided in the following way: S141, polygonizing the linear features of the vector data of roads and water bodies in the area to be tested, preliminarily dividing the sub-basins, and generating a planar area for analysis; S142, calculating the shortest distance between the initially divided sub-basins and the main river, and merging the sub-basins with the shortest distance into the corresponding main river basin, so as to optimize the sub-basin division and reduce the calculation complexity; S143. The mountainous area is identified by using the elevation data terrain difference extraction tool and the raster calculator, and is removed, while the urban area is retained as the target sub-basin, and the average slope of the sub-basin is calculated to complete the division.

6. The construction method according to claim 5, characterized in that: The hydrological simulation module includes an infiltration module and a depression storage module; The infiltration module is used to calculate the runoff of green space data. The specific method is: In formula (2), f t is the infiltration rate at time t, f c is the stable infiltration ratio, f0 is the initial infiltration rate, k is the infiltration rate attenuation coefficient, and t is the time; In formula (3), R g is the runoff of green space data at time t, i t is the rainfall at time t, a is the area occupied by the green space, f t is the infiltration rate at time t; The depression storage module is used to calculate the runoff of impervious surface data. The specific method is: In formula (4), R i is the runoff from the impervious surface data, W is the width of the subcatchment, S is the average slope of the subcatchment, n is the Manning roughness coefficient, d is the water depth at the land level, and d p is the initial surface water depth; The total surface runoff calculation formula for the sub-basin is as follows: In formula (5), R s is the total surface runoff of the sub-basin, and the data is in the form of time series.

7. The construction method according to claim 6, characterized in that: The hydraulic simulation module is a gated recurrent unit neural network, and the calculation method of the sub-basin pipe network outlet flow time series is: In formulas (6), (7) and (8), W r is the reset gate, W z is the update gate, r z is the output of the reset gate at time t, z t is the output of the update gate at time t, [h t-1 ,x t ] and [r z ,x t ] are both concatenation operators of two vectors, x t is the total surface runoff input at time t, h t It is the hidden state at time t. In addition, it is the outlet flow rate of the sub-basin pipe network at time t. The data is in the form of time series.

8. The construction method according to claim 1, characterized in that: The pollutant fate module includes a pollutant accumulation module and a pollutant flushing module; The pollutant accumulation module calculation method is: In formula (9), P is the accumulated amount of pollutants, t is the previous drought time, K B is the pollutant accumulation rate per unit area, and d is the pollution removal rate; In formula (13), C0 is the initial cumulative number of days, W c is the total amount of pollutants used in the urban area of ​​the sub-basin, G c is the green area of ​​the urban area of ​​the sub-basin, G sub is the green area of ​​the sub-basin; The pollutant flushing module is a Logistic flushing model, and the specific calculation method is: In formulas (10), (11) and (12), W t is the pollutant concentration at the outlet at time t, and the data is in the form of a time series; Q t is the outlet flow rate of the sub-basin pipe network at time t, δ t is the ratio of the pollution load available for flushing at time t, B1 and B2 are the coefficients of the logistic curve, C1 and C2 are the coefficients of the exponential flushing curve, V t is the cumulative runoff at time t, P0 is the initial pollutant accumulation in the catchment area before rainfall, and P wt is the cumulative pollutant load washed away at time t.

9. The construction method according to claim 8, characterized in that: Encapsulating the hydrological simulation module and the hydraulic simulation module into a non-point source runoff simulation module; The training method of the non-point source runoff simulation module comprises the following steps: S311, using the sub-basin rainfall time series as input data and the sub-basin pipe network outlet flow time series data as the target variable, performing data cleaning and removing outliers, and dividing into a training set and a validation set in a ratio of 7:3; S312, the surface runoff is calculated by Horton's infiltration formula and Manning's formula as the input feature of the model, the modified linear unit is used as the activation function, and the gated recurrent unit neural network is used to predict the sub-basin pipe network outlet flow; S313, using mean square error as a loss function and stochastic gradient descent method as an optimizer, setting training parameters for the non-point source runoff simulation module, performing model training, and optimizing module parameters through a back propagation algorithm; The training parameters of the non-point source runoff simulation module are: The default value of f0 is 0.9%; f c The default value is 0.5%; The default value of k is 0.2; The default value of n is 0.9; d p The default value is 0.0001m; W r To reset the gate weight, the default value is 1.0; W z To update the weight of the gate, the default value is 1.0; The default learning rate is 0.00001, and the default number of training rounds is 60; S314, performing performance evaluation on the model based on the validation set, using the Nash efficiency coefficient and the determination coefficient as evaluation indicators to verify the fitting ability of the module; When NSE>0.36 and R 2 >0.5, the model is judged to meet the requirements and operate normally; When NSE≤0.36 or R 2 When ≤0.5, check whether there are missing values ​​and outliers in the model input and output values, increase the internal parameters of the model during model training, and repeat the training until the probability space that the model can represent covers all possibilities corresponding to the runoff characteristics within the basin.

10. The construction method according to claim 8, characterized in that: The training of the pollutant fate module includes the following steps: S321, using the pollutant consumption data in each sub-basin and the sub-basin pipe network outlet flow time series data as input data, and the pollutant concentration time series at the sub-basin pipe network outlet as the target variable, perform data cleaning and remove outliers, and divide the data into a training set and a validation set in a ratio of 7:3; S322, calculating the accumulation of pollutants by a linear accumulation formula, and calculating the pollutant concentration at the outlet by a Logistic flushing model; S323, using mean absolute error as the loss function and stochastic gradient descent method as the optimizer, setting training parameters for the pollutant fate module, performing model training, and optimizing module parameters through back propagation algorithm; The training parameters of the pollutant fate module are: The initial value of C0 is 3.04219055811343 mg / m³; K B The initial value is 0.907313539587267% / d; The initial value of d is 0.895373541685726% / d; The initial value of C1 is 0.227983188296579; The initial value of C2 is 0.262647971835829; The initial value of B1 is 0.40876438184221; The initial value of B2 is 0.630350367766635; The default learning rate is 0.00001, and the default number of training rounds is 1; S314, performing performance evaluation on the model based on the validation set, using the logarithmic difference between the predicted value and the actual value as an evaluation index to verify the fitting ability of the module; When LD < 100%, the model is judged to meet the requirements and operate normally; When LD ≥ 100%, check whether there are missing values ​​and outliers in the model input and output values. Increase the internal parameters of the model during model training and repeat the training until the probability space that the model can represent covers all possibilities corresponding to the runoff characteristics within the basin.

11. A method for predicting the migration of pollutants from non-point sources based on a gated recurrent unit neural network, characterized in that: The following steps are involved: S1. Collect the current data of the area to be tested and then standardize and integrate it; The current data of the area to be tested includes: the current land use type, rainfall, pollutant usage and elevation data of the area to be tested; S2, inputting the standardized and integrated current data of the area to be tested into the pollutant surface source migration prediction model constructed by the construction method described in any one of claims 1 to 10, obtaining the prediction result, and completing the prediction; The prediction results include: the time series of runoff at the outlet of the sub-basin pipe network, the time series of pollutant concentration at the outlet of the sub-basin pipe network, and the total migration of pollutants at the outlet of the sub-basin pipe network.

Citation Information

Cited By

  • Non-point source pollution regulation and rain flood stagnating storage method and system

    CN120806685A

  • Non-point source pollution treatment and rain flood storage method and system

    CN120806685B