Data-driven water pollution source and pollution emission period prediction method

By establishing a spatiotemporal and spatial characteristic database, meteorological-hydrological coupled feature analysis and dynamic monitoring network, a pollution warning map of the entire basin was built, which solved the lag and inaccurate prediction problems of traditional water pollution monitoring methods, and realized the accurate traceability and dynamic transmission simulation of pollutants in the entire basin, improving the scientificity and real-time nature of pollution risk warning and governance.

CN120409793APending Publication Date: 2025-08-01HYDROGEN-OXYCARBON (NINGBO) TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510495358.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional water pollution monitoring methods are difficult to detect pollution risks in a timely manner and make accurate predictions, and cannot effectively solve the problem of water pollution.

Method used

Establish a spatiotemporal feature database, conduct meteorological-hydrological coupled feature analysis, quantify the pollution transport contribution of upstream and downstream sub-basins, build a knowledge spectrum, use DBSCAN clustering and LSTM and Bayesian networks to analyze pollution sources and predict emission cycles, dynamically adjust the monitoring network, and build a pollution warning map of the entire basin.

Benefits of technology

It has realized accurate traceability, dynamic transmission simulation and long-term monitoring of pollutants in the entire basin, improved the scientificity and real-time nature of pollution risk warning and environmental governance, captured pollution changes in a timely manner, and improved the timeliness and accuracy of data collection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409793A_ABST
    Figure CN120409793A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of water pollution and sewage discharge, in particular to a data-driven water pollution source and pollution discharge period prediction method. Comprising the following steps: step 1, establishing a spatial-temporal feature database, and unifying spatial-temporal references of data sources; 2, meteorological-hydrological coupling characteristics are obtained; step 3, quantifying the pollution transportation contribution of the upstream and downstream sub-basins to the downstream; step 4, automatically associating pollution point sources with sub-basins, feature monitoring points and the like based on spatial topology, and constructing a knowledge spectrogram; 5, pollution source analysis and emission period prediction; step 6, dynamically adjusting the monitoring network to cover the pollution high-risk area; step 7, constructing a whole watershed pollution early warning map; through data driving, map construction, pollution / pollution source analysis and prediction, monitoring dynamic adjustment and a whole-basin pollution early warning map, accurate traceability and dynamic transmission simulation of whole-basin pollutants, remote monitoring of a river leader and intelligent operation management of a river channel are realized, and the safety and health of a water ecological environment are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water pollution and sewage discharge, and specifically to a data-driven method for predicting water pollution sources and pollution discharge cycles. Background Technique

[0002] Water pollution monitoring refers to continuously monitoring and analyzing water bodies to promptly detect and evaluate the presence of pollutants and their concentration changes, so as to ensure the safety of water resources and the health of the ecological environment; water pollution monitoring plays an important role in protecting human health, maintaining ecological balance, controlling pollution sources, preventing water pollution incidents, and evaluating treatment effects, and is a key measure for protecting water resources and the ecological environment;

[0003] With the acceleration of urbanization and industrialization, the problem of water pollution has become increasingly prominent. Traditional monitoring methods are difficult to promptly detect pollution risks and make accurate predictions. Therefore, it is urgent to construct a data-driven water pollution source analysis and discharge cycle prediction model based on multi-source heterogeneous data fusion and artificial intelligence technology. Summary of the Invention

[0004] The purpose of the present invention is to provide a data-driven method for predicting water pollution sources and pollution discharge cycles to solve the problems in the above background technique.

[0005] The purpose of the present invention can be achieved through the following technical solutions: A data-driven method for predicting water pollution sources and pollution discharge cycles includes:

[0006] Step 1, establish a spatio-temporal feature database: The spatio-temporal feature database includes environmental monitoring data, geospatial data, meteorological and hydrological data, and remote sensing data, and unify the spatio-temporal benchmarks of each data source;

[0007] Step 2, meteorological-hydrological coupling characteristics to determine the lag time Δt and pollution discharge intensity of each sub-basin;

[0008] Step 3, quantify the pollution transport contribution of the upstream and downstream sub-basins to the downstream to output the output index of each pollutant for each pair of upstream and downstream sub-basins;

[0009] Step 4, based on spatial topology, automatically associate pollution point sources with sub-basins, characteristic monitoring points, etc. to construct a knowledge spectrum;

[0010] Step 5, pollution source analysis and discharge cycle prediction;

[0011] Step 6, dynamically adjust the monitoring network to cover high-risk pollution areas; specifically:

[0012] 1-1: Extract the emission intensity and transmission index within each sub-basin, mark the existing monitoring points on the map of each sub-basin, connect adjacent monitoring points to grid each sub-basin, and calculate the area of each grid;

[0013] 1-2: Take any one sub-basin as the termination node, and other sub-basins that are upstream of this sub-basin as related nodes. Extract the transmission index of each related node and the termination node for each pollutant, and calculate the sum according to the type of pollutant to obtain the transmission value of each pollutant at the termination node. Then calculate the average value to obtain the transmission average value; perform normalization calculation on the sewage discharge intensity and transmission average value Z of the termination node to obtain the pollution judgment value Q of the sub-basin. If the sewage discharge judgment value is greater than the maximum value in the set judgment interval, then fill this sub-basin with red and match the monitoring coverage rate F1 and sampling frequency L1; if the sewage discharge judgment value is within the set judgment interval, then fill this sub-basin with yellow and match the monitoring coverage rate F2 and sampling frequency L2; if the sewage discharge judgment value is less than the minimum value in the set judgment interval, then fill this sub-basin with blue and match the monitoring coverage rate F3 and sampling frequency L3; F1 > F2 > F3 > 0, L1 > L2 > L3 > 0; where F3 is the basic coverage rate;

[0014] 1-3: Set the effective coverage radius of each monitoring point, calculate the coverage area of the monitoring point based on the effective coverage radius, then sum up the coverage areas of all monitoring points within the sub-basin to obtain the total effective coverage area, and divide the total effective coverage area by the sub-basin area to obtain the actual monitoring coverage rate. If the actual monitoring coverage rate is less than the matched monitoring coverage rate, then additional monitoring points need to be arranged; otherwise, no additional arrangement is required;

[0015] 1-4: Extract the area with the largest grid area within the sub-basin as the monitoring point layout point, set up a monitoring point at the middle position of this grid, re-grid the sub-basin after arranging the monitoring points, and recalculate the actual monitoring coverage rate until the actual monitoring coverage rate is no longer less than the matched monitoring coverage rate;

[0016] Step Seven, integrate the optimized monitoring points, real-time data and prediction results to construct a full-basin pollution warning map.

[0017] Preferably, the specific process of the meteorological - hydrological coupling characteristics is as follows:

[0018] Obtain elevation data with a resolution ≤ 30m, import the river centerline, hydrological station location, and sensitive area boundary. Based on the digital elevation model and hydrological analysis, divide the river basin into N sub-regions, and control the area of each sub-basin within 50 - 200 km 2, monitoring points are evenly distributed in each area, with a hydrological station or a sensitive area as the outlet point to generate the sub-basin boundary; taking the sub-basin as the unit, align the rainfall data and runoff data, the initial value of the time resolution is T1, and it increases according to the rainfall intensity, and the maximum is encrypted to 10 minutes during heavy rain;

[0019] Calculate the lag time Δt of the rainfall sequence and runoff sequence for each sub-basin separately, specifically:

[0020] Obtain rainfall data, use cross-correlation analysis to determine the lag time between the rainfall peak and the runoff peak, calculate the Pearson correlation coefficient between the rainfall sequence and the runoff sequence at different lag times, and select the lag time corresponding to the maximum correlation coefficient;

[0021] A single rainfall event is defined as continuous rainfall ≥ h1 and no rainfall for an interval ≥ h2, where h1 and h2 are constants; calculate the runoff coefficient C for each rainfall, and thus obtain the runoff coefficients of each sub-basin:

[0022] Analyze the emission intensity of a certain pollutant in each sub-basin to obtain the emission intensity.

[0023] Preferably, the emission intensity analysis is specifically:

[0024] Mark the pollution discharge outlets at the river channel boundary according to their corresponding positions, and extract the emissions of a certain pollutant at each pollution discharge outlet and mark it as the point source emission A point ;

[0025] Non-point source pollution includes agricultural non-point source pollution and residential non-point source pollution. The calculation and analysis process for non-point source pollution emissions is as follows:

[0026] Use remote sensing satellites or GIS data to extract the farmland distribution within the sub-basin, obtain the fertilization amount per unit area of different crops from agricultural statistical yearbooks or real-time surveys, and use SWAT or HSPF to simulate runoff and pollutant scouring amounts;

[0027] Calculation of agricultural non-point source pollution emissions, the calculation formula is: where g = 1, 2, 3... G, G represents the types of crops within the sub-basin, g represents any one type of crop, α g represents the fertilizer utilization rate of crop g, F g represents the planting area of crop g, R g represents the runoff coefficient of crop g, M pollutant represents the content of the target pollutant in the chemical fertilizer applied to crop g;

[0028] Calculation of residential non-point source pollution emissions, the calculation formula is: where p = 1, 2, 3 …… P, P represents the total number of surface types in the sub - watershed, and p is any one of the surface types; D p represents the cumulative rate of pollutants of surface type p, W p represents the impervious area of type p, E p represents the removal rate of pollutants by sewage treatment facilities;

[0029] Sum up the point - source emissions A point of a certain pollutant at each pollution discharge outlet in the sub - watershed, the agricultural non - point source pollution emissions A agri and the residential non - point source pollution emissions A urban to obtain the emission intensity A of a certain pollutant in the sub - watershed.

[0030] Preferably, quantify the pollution transport contribution of upstream and downstream sub - watersheds to the downstream, specifically:

[0031] 4 - 1, Topological relationship construction: Based on D8 flow direction data, generate a sub - watershed topological relationship network and identify the upstream - downstream relationship;

[0032] 4 - 2, Calculate the transfer index S ij using the formula: where i and j represent any two sub - watersheds in an upstream - downstream relationship, A i represents the emission intensity of the upstream sub - watershed i, t ij represents the water flow time from sub - watershed i to j, which is calculated from the river channel length and flow velocity, Slope j represents the average slope of the downstream sub - watershed j, k represents the attenuation coefficient of the pollutant, D ij represents the river channel migration distance from sub - watershed i to j;

[0033] 4 - 3, Update t ij according to the real - time rainfall intensity and river channel flow velocity per hour, and store it in the spatio - temporal feature database; Take the rainfall - runoff coupling feature, spatial correlation feature, temporal periodicity feature, and the pollution emissions, emission intensities and transfer indices of each sub - watershed as the multi - dimensional feature vectors of each sub - watershed.

[0034] Preferably, based on spatial topology, automatically associate pollution point sources with sub - watersheds to construct a knowledge spectrum diagram; specifically:

[0035] Determine entity categories, which include sub - watershed nodes, pollution sources, transmission paths, and sensitive areas; taking sub - watersheds as nodes, the attributes include area, emission intensity, runoff coefficient, rainfall - runoff lag time; pollution sources include point sources, agricultural non - point sources, residential non - point sources, etc., and each pollution discharge port or area is regarded as an entity, with attributes including emissions and geographical location; the transmission path represents the water flow path between sub - watersheds, and the attributes include flow direction, water flow time, migration distance; sensitive areas such as drinking water sources and ecological protection areas are regarded as important nodes and are prominently marked with colors different from other nodes, with attributes including population density and ecological sensitivity;

[0036] Build a spatial topological network, where the edges represent the water flow transmission paths, and the relationships between nodes reflect the possibility of pollutant migration within the watershed; import the multi - dimensional feature vectors of each sub - watershed as structured information, and use geographical location to automatically match the relationships between entities;

[0037] Using automated rules and statistical methods, extract the relationships between entities from the characteristic data of each sub - watershed, import the extracted entities and relationships into the graph database, and construct a whole - watershed knowledge graph of "sub - watershed - pollution source - transmission path - sensitive area".

[0038] Preferably, source analysis of pollution sources and prediction of emission cycles are as follows:

[0039] Identify the spatial aggregation areas of pollutant events, distinguish the high - incidence areas of point source and non - point source pollution, and obtain historical pollution datasets and geographical data. The historical pollution datasets include the coordinates, pollutant types, and concentration values of water quality monitoring exceeding the standard events for 3 years or more; the geographical data includes soil infiltration number, slope, and land use type;

[0040] Determine the neighborhood radius through the K - distance graph, set the minimum number of samples, use the neighborhood radius and the minimum number of samples as the parameters of the DBSCAN algorithm, run the DBSCAN algorithm, and output the clustering labels. If the clustering center is located in the industrial area and the soil infiltration coefficient > the preset coefficient V1, it is determined as a point source pollution area and output; if the clustering area satisfies the slope > the preset slope V2, and the land use is farmland and the soil infiltration coefficient < the preset coefficient V1, it is determined as a non - point source pollution area and output; where V1 and V2 are constants;

[0041] Construct a pollutant concentration matrix for multiple pollution indicators collected at each monitoring point, use the PMF model to decompose the pollutant concentration matrix into a source contribution matrix and a source characteristic matrix, analyze the contribution rate of each pollution source to the overall pollution, and identify the main pollution factors;

[0042] Adopt a long - short - term memory network, use the extracted spatio - temporal features as input, fuse the enterprise production cycle data and real - time meteorological data to train the model, and the model outputs the predicted values of the pollution load of each sub - watershed in the short term and identifies short - term abnormal fluctuations;

[0043] A multivariate time series forecasting model is constructed using a dynamic Bayesian network. The input data includes long-term monitoring data, seasonal meteorological changes, and historical emission data. The model outputs medium- and long-term emission trends to determine the stability of the pollutant emission cycle and potential abnormal events.

[0044] Preferably, a basin-wide pollution warning map is constructed, specifically:

[0045] Upload updated monitoring point information to the GIS platform, including the geographic coordinates, coverage, and measurement parameters of each monitoring point; access real-time water quality data from each monitoring point and automatically update the data through wireless transmission or the Internet of Things platform; introduce short-term and medium-term emission load data predicted by LSTM and dynamic Bayesian networks, and compare and integrate them with real-time monitoring data;

[0046] Using high-precision DEM and watershed boundary data as the base map, the boundaries of each sub-watershed, river direction and sensitive areas are marked on the map, and the monitoring point data are superimposed in layers to display the real-time data of each monitoring point. Using the results of the prediction model, a pollutant distribution heat map is generated on the map to reflect the changing trend of pollution load in the future.

[0047] Beneficial effects of the present invention:

[0048] 1. Divide the watershed into 50-200km based on DEM 2 The system accurately locates the pollution diffusion range in sub-regions, uses cross-correlation analysis to determine the lag time Δt, and improves the accuracy of pollutant migration simulation. Through sub-basin-level meteorological-hydrological coupling modeling, it realizes real-time dynamic prediction of pollution load. It uses D8 flow direction to generate sub-basin adjacency matrix, clarifies the upstream and downstream pollution transmission paths, and accurately quantifies the pollution transmission contribution of upstream and downstream sub-basins to downstream. It updates rainfall intensity and flow rate data every hour and adjusts the transmission index in real time. The transmission index accurately locates the pollution source and diffusion path, supports responsibility tracing and governance priority division.

[0049] 2. By regularly updating the knowledge graph, the dynamic and real-time nature of basin-wide information is maintained, enabling comprehensive monitoring and early warning of basin-wide pollution risks. The organic combination of spatial data, spatiotemporal characteristics, and pollution source data not only enables automatic association and knowledge fusion of basin-wide data, but also provides users with an intuitive and dynamic basin-wide pollutant query platform, greatly enhancing the scientific and real-time nature of pollution risk warnings and environmental governance.

[0050] 3. By using DBSCAN clustering to distinguish point sources from area sources (such as high COD emissions in industrial areas and ammonia nitrogen diffusion in farmland), and the PMF model to quantify the contribution rate, and using short-term prediction (LSTM) and medium- and long-term prediction (Bayesian network), not only the accurate analysis of pollution sources is achieved, but also the emission cycle and abnormal fluctuations can be predicted in advance, effectively solving the problems of lag and inaccurate prediction of traditional monitoring methods;

[0051] 4. By dynamically optimizing the monitoring network and adjusting the sampling frequency in real time, it is ensured that pollution changes can be captured in a timely manner in case of emergencies, greatly improving the timeliness and accuracy of data collection, making up for the deficiencies of traditional fixed monitoring networks, and achieving monitoring coverage across the entire basin; using an adaptive adjustment mechanism to increase the sampling frequency; integrating the optimized monitoring points, real-time water quality data with short-term and medium- and long-term prediction results to construct a GIS-based pollution early warning map for the entire basin. The map uses three-level warning color codes of red, yellow, and blue to intuitively display the pollution load, trend, and risk distribution in each sub-basin, facilitating managers to quickly locate high-risk areas;

[0052] In summary, through data-driven - atlas construction - pollution / pollution source analysis and prediction - dynamic adjustment of monitoring - pollution early warning map for the entire basin, the present invention realizes the accurate tracing of pollutants across the entire basin, dynamic transmission simulation, remote monitoring of river chiefs, and intelligent operation management of river channels, ensuring the safety and health of the water ecological environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The present invention will be further described below with reference to the accompanying drawings.

[0054] Figure 1 is a schematic flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] Please refer to Figure 1 as shown, the present invention is a data-driven water pollution source and pollution emission cycle prediction method, including the following steps:

[0057] Step 1: Establish a spatio-temporal feature database. Specifically, the spatio-temporal feature database includes environmental monitoring data, geospatial data, meteorological and hydrological data, and remote sensing data; specifically:

[0058] Environmental monitoring data: including historical and real-time water quality data (such as indicators like COD, ammonia nitrogen, total phosphorus, etc.), and pollution source monitoring data (industrial wastewater, agricultural non-point source, urban domestic sewage, etc.).

[0059] Geospatial data: Obtain high-precision DEM terrain data with a resolution ≤ 30m (such as ALOS PALSAR or SRTM), river channel water system distribution, land use type, soil permeability coefficient, etc.

[0060] Meteorological and hydrological data: Collect rainfall, wind speed, temperature, runoff, and hydrological station monitoring data.

[0061] Remote sensing data: Use multi-spectral satellite images (such as NDVI, water quality inversion parameters) and UAV aerial photography to obtain the pollution diffusion heat map;

[0062] Unify the spatio-temporal benchmarks of each data source (for example, adopt the WGS-84 coordinate system and align the time stamps of each data), use KNN interpolation or spatio-temporal Kriging method to fill in the missing data, and adopt the isolation forest algorithm to eliminate abnormal data (such as a 20-fold sudden increase in COD);

[0063] By integrating multi-source heterogeneous data (environmental monitoring, geospatial, meteorological and hydrological, remote sensing), break the traditional data islands, solve the problem of low traceability efficiency caused by data dispersion; unify the spatio-temporal benchmarks (WGS-84 coordinate system, time stamp alignment) to ensure data consistency, provide high-quality input for the subsequent model, and use the isolation forest algorithm to eliminate sensor abnormal data to improve data reliability.

[0064] Step 2, meteorological-hydrological coupling characteristics, specifically:

[0065] Obtain elevation data with a resolution ≤ 30m (such as ALOS PALSAR or SRTM data), import the river channel center line, hydrological station location, and boundaries of sensitive areas (such as water source areas), and based on the digital elevation model (DEM) and hydrological analysis, divide the river channel basin into N sub-regions (such as Sub, Sub, SubN), and control the area of each sub-basin within 50 - 200km 2 , and evenly distribute monitoring points in each area, take the hydrological station or sensitive area as the outlet point to generate the sub-basin boundary (Sub-basin); N is a positive integer representing the total number of sub-basins within the monitoring range;

[0066] Taking the sub-basin as the unit, align the rainfall data (rainfall data from meteorological stations / satellite inversion) and runoff data (from the flow meters at the monitoring points), the initial value of the time resolution is T1 (those skilled in the art set the initial value of the time resolution to 1h), and increase according to the rainfall intensity, and encrypt to 10min during the heavy rain period;

[0067] Calculate the lag time Δt of the rainfall sequence and the runoff sequence for each sub-basin separately, specifically as follows:

[0068] Obtain high-resolution rainfall data with a time interval ≤ 1h, synchronize the river channel runoff data, and align the timestamps with the rainfall data; use cross-correlation analysis to determine the lag time between the rainfall peak and the runoff peak. For example, if the runoff reaches its peak 3h after rainfall, calculate the Pearson correlation coefficient between the rainfall sequence and the runoff sequence at different lag times, and select the lag time corresponding to the maximum correlation coefficient;

[0069] Define a single rainfall event as continuous rainfall ≥ h1 (those skilled in the art set the value of h1 to 5mm) and no rainfall for an interval ≥ h2 (those skilled in the art set the value of h2 to 6h); calculate the runoff coefficient C for each rainfall event, and thus obtain the runoff coefficients of each sub-basin. The specific calculation formula is:

[0070] Where Rounoff represents the runoff volume, with the unit of cubic meters (m 3 ), which refers to the total amount of water passing through the river channel section during a single rainfall event; Rainfall represents the rainfall depth with the unit of millimeters (mm), which refers to the total precipitation of a single rainfall; Area represents the basin area, with the unit of square kilometers (km 2 ), which refers to the area range where rainwater converges to the monitoring point; calculate and group by rainfall intensity level to generate feature vectors:

[0071] Regarding the analysis method of the pollutant emission intensity of a certain pollutant in each sub-basin:

[0072] Mark the pollution discharge outlets at the river channel boundary according to their corresponding positions. Here, the pollution discharge outlets include the sewage outlets of each enterprise, etc., and extract the emissions of a certain pollutant at each pollution discharge outlet and mark them as the point source emissions A point (The data comes from the sewage ledgers of each enterprise, with the unit of kg / day);

[0073] Non-point source pollution includes agricultural non-point source pollution and residential (urban) non-point source pollution. The calculation and analysis process for non-point source pollution emissions is as follows:

[0074] Use remote sensing satellites or GIS data to extract the farmland distribution within the sub-basin (such as paddy fields, orchards, etc.), obtain the fertilization amount per unit area of different crops from agricultural statistical yearbooks or real-time surveys, and use SWAT (Soil and Water Assessment Tool) or HSPF (Hydrological Simulation Program - Fortran) to simulate the runoff volume and pollutant scouring amount;

[0075] Calculation of agricultural non-point source pollution emissions. The calculation formula is as follows: Where \(g = 1, 2, 3,\cdots,G\), \(G\) represents the types of crops in the sub-watershed, \(g\) represents any one type of crop, \(\alpha\) g represents the chemical fertilizer application rate of crop \(g\) (kg / ha·year), \(F\) g represents the planting area of crop \(g\) (ha), \(R\) g represents the runoff coefficient of crop \(g\) (the value is determined by those skilled in the art based on field experiments or literature, usually the proportion of chemical fertilizer loss after rainfall, with a value ranging from 0 to 1), \(M\) pollutant represents the content of target pollutants in the chemical fertilizer applied to crop \(g\) (for example, the ammonia nitrogen content in nitrogen fertilizer accounts for about 30%);

[0076] Example calculation: In a certain sub-watershed, the area of paddy fields \(F\) g = 100 ha, the amount of nitrogen fertilizer applied \(\alpha\) g = 200 kg / ha·year, \(R\) g = 0.2, \(M\) pollutant = 0.3. Then the pollution emissions from the paddy fields in the sub-watershed are \(200\times100\times0.2\times0.3 = 1200\) kg / year = 3.29 kg / day;

[0077] Calculation of non-point source pollution emissions from residents. The calculation formula is as follows: Where \(p = 1, 2, 3,\cdots,P\), \(P\) represents the total number of surface types in the sub-watershed, and \(p\) is any one of the surface types; \(D\) p represents the cumulative rate of pollutants of surface type \(p\) (g / m 2 ·day), \(W\) p represents the impervious area of type \(p\) (m 2 ), \(E\) p represents the removal rate of pollutants by rainwater treatment facilities (for example, the removal rate of rain gardens is about 60%);

[0078] Sum up the point source emissions \(A\) point of a certain pollutant at each pollution discharge outlet in the sub-watershed, the agricultural non-point source pollution emissions \(A\) agri and the non-point source pollution emissions \(A\) urban from residents to obtain the emission intensity \(A\) of a certain pollutant in the sub-watershed;

[0079] Based on DEM, the watershed is divided into sub-regions of 50 - 200 km 2 to accurately locate the pollution diffusion range, determine the lag time \(\Delta t\) through cross-correlation analysis, and improve the accuracy of pollutant migration simulation; through sub-watershed level meteorological-hydrological coupling modeling, realize real-time dynamic prediction of pollution load.

[0080] Step three, quantify the pollution transport contribution of upstream and downstream sub-watersheds to the downstream, specifically:

[0081] Topological relationship construction: Based on the D8 flow direction data, generate the topological relationship network (adjacency matrix) of sub-watersheds, and identify the upstream and downstream relationships (such as Sub-1→Sub-3→Sub-5);

[0082] Transportation index S ij Calculation, the formula is: where i and j represent any two sub-watersheds in the upstream and downstream relationships, and A i represents the emission intensity of the upstream sub-watershed i (the emission intensity here refers to the emission degree index of a certain pollutant), and t ij represents the water flow time from sub-watershed i to j (in hours), and the water flow time can be calculated through the river channel length and flow velocity. Slope j represents the average slope of the downstream sub-watershed j (the slope unit in the formula is in percentage and needs to be converted to a decimal, such as a slope of 15% is recorded as 0.15). The greater the slope, the less the pollutant deposition. k represents the attenuation coefficient (those skilled in the art usually take 0.1 for COD and 0.5 for ammonia nitrogen), and D ij represents the river channel migration distance from sub-watershed i to j. The greater the river channel migration distance, the more the influence of the upstream sub-region i on the downstream sub-region j attenuates with the increase of the distance; for example, if the river channel path from sub-watershed i to j consists of 3 segments with lengths of 2 km, 3 km, and 5 km respectively, then D ij = 2 + 3 + 5 = 10 km;

[0083] Update t according to the real-time rainfall intensity and river channel flow velocity per hour ij , and store it in the spatio-temporal feature database;

[0084] Integrate the aforementioned rainfall-runoff coupling characteristics, spatial correlation characteristics, temporal periodic characteristics, and the pollution emissions, emission intensities, and transportation indices of each sub-watershed to form the multi-dimensional feature vectors of each sub-watershed, which are used as the input data for the subsequent prediction model;

[0085] Generate the sub-watershed adjacency matrix based on the D8 flow direction, clarify the upstream and downstream pollution transmission paths, accurately quantify the pollution transportation contributions of the upstream and downstream sub-watersheds to the downstream, update the rainfall intensity and flow velocity data per hour, and adjust the transportation index in real time; accurately locate the pollution sources and diffusion paths through the transportation index, and support responsibility tracing and the division of governance priorities.

[0086] Step 4, based on the spatial topology, automatically associate pollution point sources with sub-watersheds, characteristic monitoring points, etc., and construct a knowledge spectrum diagram, specifically:

[0087] Determine the main entity categories, which include sub - watershed nodes, pollution sources, transmission paths, and sensitive areas; taking sub - watersheds as nodes, the attributes include area, emission intensity, runoff coefficient, rainfall - runoff lag time, etc.; pollution sources include point sources, agricultural non - point sources, domestic non - point sources, etc., with each pollution discharge outlet or area as an entity, and the attributes include emissions, geographical location, etc.; the transmission path represents the water flow path between sub - watersheds, and the attributes include flow direction, water flow time, migration distance, etc.; sensitive areas such as drinking water sources and ecological protection areas are important nodes, and the attributes include population density, ecological sensitivity, etc.;

[0088] Build a spatial topological network, where the edges represent the water flow transmission paths, and the relationships between nodes reflect the possibility of pollutant migration within the watershed (e.g., Sub - 1→Sub - 3→Sub - 5); import the characteristics of each sub - watershed (multi - dimensional feature vectors), pollution discharge data, and transmission index data calculated in the previous steps as structured information, and use geographical location (coordinates, sub - watershed boundaries) to automatically match the relationships between entities. For example, the sub - watershed node is associated with the pollution discharge data within it, and sub - watersheds are automatically associated through topological relationships (based on D8 flow direction), forming upstream - downstream transmission edges and an inclusion relationship between sensitive areas and the sub - watersheds where they are located;

[0089] Use automated rules (based on spatial proximity, flow - direction topology) and statistical methods to extract the relationships between entities from the characteristic data of each sub - watershed, such as "the sub - watersheds located upstream have an impact on the pollution transmission of downstream sub - watersheds"; import the extracted entities and relationships into a graph database (such as Neo4j) to construct a whole - watershed knowledge graph of "sub - watershed - pollution source - transmission path - sensitive area";

[0090] Based on the constructed knowledge graph, use a graph query language (such as Cypher) to achieve semantic queries.

[0091] For example, users can query "the emission intensity and propagation path of a certain pollutant in the whole watershed", and the system automatically extracts relevant sub - watershed node, transmission index, and sensitive area information from the graph, and generates a visualization report on the distribution and propagation dynamics of pollutants in the whole watershed.

[0092] By regularly updating the knowledge graph, maintain the dynamics and real - time nature of the whole - watershed information, and achieve comprehensive monitoring and early warning of the whole - watershed pollution risk. The organic combination of spatial data, spatio - temporal characteristics, and pollution source data not only realizes the automatic association and knowledge fusion of the whole - watershed data, but also provides users with an intuitive and dynamic whole - watershed pollutant query platform, greatly improving the scientific nature and real - time nature of pollution risk early warning and environmental governance.

[0093] Step Five, pollution source analysis and emission cycle prediction, specifically:

[0094] Identify the spatial aggregation areas of pollutant events, distinguish the high-incidence areas of point source and non-point source pollution, and obtain historical pollution datasets and geographical data. The historical pollution datasets include the coordinates (longitude, latitude), pollutant types (such as COD, ammonia nitrogen, etc.) and concentration values of water quality monitoring exceeding the standard events for 3 years or more; the geographical data includes soil infiltration number (unit: cm / h), slope (unit: %), and land use type (industrial area / farmland / residential area).

[0095] Determine the neighborhood radius through the K-distance graph. For example, select the neighborhood radius eps = 200 meters (so that the neighborhood of 80% of pollution events contains ≥5 similar events), and set the minimum number of samples (personnel in this field set it to 5, indicating that if a certain area exceeds the standard continuously 5 times, it is regarded as clustering); use the neighborhood radius and the minimum number of samples as the parameters of the DBSCAN algorithm, run the DBSCAN algorithm, and output clustering labels (core points, boundary points, and noise points). If the clustering center is located in the industrial area and the soil infiltration coefficient > the preset coefficient V1 (personnel in this field set the value of the preset coefficient V1 to 0.5 cm / h), it is determined as a point source pollution area and output; if the clustering area satisfies the slope > the preset slope V2, the land use is farmland and the soil infiltration coefficient < the preset coefficient V1, it is determined as a non-point source pollution area and output.

[0096] Construct a pollutant concentration matrix for multiple pollution indicators collected at each monitoring point, and use the PMF (Positive Matrix Factorization) model to decompose the pollutant concentration matrix into a source contribution matrix and a source characteristic matrix, analyze the contribution rate of each pollution source to the overall pollution, and identify the main pollution factors (for example: industrial wastewater may mainly contribute certain specific organic substances or heavy metals, while agricultural non-point sources contribute more to nutrients such as nitrogen and phosphorus).

[0097] Adopt the Long Short-Term Memory Network (LSTM), use the spatio-temporal features extracted in the early stage (including rainfall-runoff lag time, runoff coefficient, local emission intensity, etc.) as input, and fuse enterprise production cycle data (such as industrial sewage discharge time periods) and real-time meteorological data (rainfall, temperature, wind speed) to train the model. The model outputs the predicted pollution load values of each sub-basin in the short term and identifies possible short-term abnormal fluctuations.

[0098] Use the Dynamic Bayesian Network to construct a multivariate time series prediction model; the input data includes long-term monitoring data, seasonal meteorological changes, and emission history data; the model outputs the medium- and long-term emission trends to help judge the stability of the pollutant emission cycle and potential abnormal events.

[0099] Example:

[0100] Short-term prediction: Predict the pollutant load fluctuations (such as COD concentration changes) in the next 72 hours.

[0101] Implementation steps:

[0102] Input: Historical 72-hour rainfall, temperature, wind speed, enterprise production flag (0 / 1), river flow

[0103] Output: Future 72-hour COD concentration sequence (time resolution 1 hour).

[0104] LSTM network structure: Input layer: Time step = 72, number of features = 5. Hidden layer: 2 layers of LSTM, 128 units per layer, Dropout = 0.2. Output layer: Fully connected layer (72 nodes).

[0105] Training and validation: Loss function: Mean Squared Error (MSE); Optimizer: Adam (learning rate = 0.001);

[0106] Validation result: Test set MSE = 0.04, MAE = 0.08 mg / L.

[0107] Prediction example:

[0108] Input: Chemical plant production day (flag = 1), rainfall 10 mm / h, flow 50 m 3 / s;

[0109] Output: The peak value of future 72-hour COD concentration appears at the 48th hour (predicted value = 48 mg / L, actual value = 50 mg / L).

[0110] Medium- and long-term prediction (Dynamic Bayesian Network):

[0111] Implementation steps:

[0112] Multi-media migration modeling:

[0113] Migration path: Surface water → Groundwater → Soil → Surface water (closed-loop feedback);

[0114] Key parameters: COD degradation rate (0.1 d -1 )), soil adsorption coefficient (Freundlich equation Kf = 0.3). Example rule: If monthly rainfall > 200 mm, the contribution rate of agricultural non-point source increases by 30%.

[0115] Prediction output: High-risk period calendar: July 15 - 25 is the high-risk period for COD exceeding the standard (probability > 80%); Treatment suggestions: Deploy ecological intercepting ditches in advance and increase the treatment capacity of sewage treatment plants.

[0116] By using DBSCAN clustering to distinguish point sources from area sources (such as high COD emissions in industrial areas and ammonia nitrogen diffusion in farmland), and the PMF model to quantify the contribution rate, and using short-term prediction (LSTM) and medium- and long-term prediction (Bayesian network), not only the accurate analysis of pollution sources is achieved, but also the emission cycle and abnormal fluctuations can be predicted in advance, effectively solving the problems of lag and inaccurate prediction of traditional monitoring methods.

[0117] Step 6: Dynamically adjust the monitoring network to cover high-pollution-risk areas, specifically as follows:

[0118] 1-1: Extract the pollutant concentrations (such as COD, ammonia nitrogen, total phosphorus, etc.), historical anomaly frequencies, emission intensities, and transmission indices within each sub-basin, mark the existing monitoring points on the map of each sub-basin, connect adjacent monitoring points to grid each sub-basin, and calculate the area of each grid;

[0119] 1-2: Take any one sub-basin as the termination node, and other sub-basins that are upstream of this sub-basin as relevant nodes, extract the transmission indices of each relevant node and the termination node for each pollutant, sum them according to the type of pollutant to obtain the transmission value of each pollutant at the termination node, and then calculate its average value to obtain the transmission average value; normalize the sewage discharge intensity and transmission average value Z of the termination node to obtain the pollution judgment value Q of the sub-basin. The specific calculation formula is where Amax is the maximum allowable sewage discharge intensity set by the technical personnel, and Zmax is the maximum allowable transmission average value set by the technical personnel;

[0120] Compare and analyze the pollution judgment value of the sub-basin with the set judgment interval. If the pollution judgment value is greater than the maximum value in the set judgment interval, it indicates that the pollution in the sub-basin is relatively serious, then fill the sub-basin with red and match it with a monitoring coverage rate of F1 and a sampling frequency of L1; if the pollution judgment value is within the set judgment interval, then fill the sub-basin with yellow and match it with a monitoring coverage rate of F2 and a sampling frequency of L2; if the pollution judgment value is less than the minimum value in the set judgment interval, it indicates that the pollution in the sub-basin is within the normal range, then fill the sub-basin with blue and match it with a monitoring coverage rate of F3 and a sampling frequency of L3; F1 > F2 > F3 > 0, L1 > L2 > L3 > 0; where F3 is the basic coverage rate. Usually, the monitoring point coverage rate of each sub-basin at the initial stage of water pollution monitoring satisfies F3 (at the initial stage of monitoring, monitoring points will definitely be arranged at sensitive areas and sewage outlets in the sub-basin); therefore, when the coverage rate of F3 is matched, there is no need to increase the layout of monitoring points within the sub-basin.

[0121] 1-3: Set the effective coverage radius for each monitoring point (the effective coverage radius is usually the affected range determined based on sensor sensitivity and environmental characteristics). Calculate the coverage area of the monitoring point according to the effective coverage radius, then sum up the coverage areas of all monitoring points within the sub-watershed to obtain the total effective coverage area. Divide the total effective coverage area by the sub-watershed area to get the actual monitoring coverage rate. If the actual monitoring coverage rate is less than the matched monitoring coverage rate, additional monitoring points need to be arranged; otherwise, no additional arrangement is required.

[0122] 1-4: Extract the area with the largest grid area within the sub-watershed as the location for arranging the monitoring point. Set the monitoring point at the middle position of this grid, re-divide the sub-watershed with the arranged monitoring points into grids, and recalculate the actual monitoring coverage rate until the actual monitoring coverage rate is no longer less than the matched monitoring coverage rate.

[0123] By dynamically optimizing the monitoring network and adjusting the sampling frequency in real time, it is ensured that pollution changes can be captured in a timely manner in case of emergencies, greatly improving the timeliness and accuracy of data collection, making up for the deficiencies of traditional fixed monitoring networks, and achieving monitoring coverage across the entire watershed; using an adaptive adjustment mechanism to increase the sampling frequency.

[0124] Step 7: Integrate the optimized monitoring points, real-time data, and prediction results to construct a pollution warning map for the entire watershed. Specifically:

[0125] Upload the updated monitoring point information to the GIS platform, including the geographical coordinates, coverage range, and measurement parameters of each monitoring point; access the real-time water quality data of each monitoring point (such as COD, ammonia nitrogen, total phosphorus, etc.), and achieve automatic data update through wireless transmission or the Internet of Things platform; introduce short-term and medium- to long-term emission load data predicted based on LSTM and dynamic Bayesian networks, and compare and integrate them with real-time monitoring data.

[0126] Using high-precision DEM and watershed demarcation data as the base map, mark the boundaries of each sub-watershed, the river channel trends, and sensitive areas (such as drinking water sources) on the map. Overlay the monitoring point data in the form of layers, display the real-time data of each monitoring point (which can be intuitively represented by icons, colors, sizes, etc.), and use the results of the prediction model to generate a heat map of pollutant distribution on the map to reflect the change trend of pollution load in the future for a period of time.

[0127] Users can click on any sub-watershed or monitoring point on the map to pop up a detailed data window, displaying historical data, current monitoring values, and prediction trends; support layer switching, and users can choose to display information such as pollutant types, time series, or transmission paths.

[0128] By integrating the optimized monitoring points, real-time water quality data with short-term, medium- and long-term prediction results, a GIS-based pollution warning map for the entire basin is constructed. The map uses three-level warning color codes of red, yellow, and blue to visually display the pollution load, trend, and risk distribution within each sub-basin, facilitating managers to quickly locate high-risk areas.

[0129] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods for substitution. As long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should fall within the protection scope of the present invention.

Claims

1. A data-driven prediction method for water pollution sources and pollution discharge cycles, characterized in that, Including: Step 1: Establish a spatio-temporal feature database. The spatio-temporal feature database includes environmental monitoring data, geospatial data, meteorological and hydrological data, and remote sensing data, and unifies the spatio-temporal benchmarks of each data source. Step 2: Meteorological-hydrological coupling characteristics to determine the lag time and pollutant discharge intensity of each sub-basin. Step 3: Quantify the pollution transport contribution of upstream and downstream sub-basins to the downstream, and output the output index of each pollutant for sub-basins in the upstream and downstream relationship. Step 4: Automatically associate pollution point sources with sub-basins based on spatial topology to construct a knowledge spectrum diagram. Step 5: Source analysis of pollution sources and prediction of emission cycles. Step 6: Dynamically adjust the monitoring network to cover high-risk pollution areas. Step 7: Integrate the optimized monitoring points, real-time data and prediction results to construct a pollution warning map for the entire basin.

2. The data-driven water pollution source and pollution discharge cycle prediction method according to claim 1, wherein, The specific process of meteorological-hydrological coupling characteristics is as follows: Obtain elevation data with a resolution of ≤30m, import the centerline of the river channel, the location of hydrological stations, and the boundaries of sensitive areas. Based on the digital elevation model and hydrological analysis, divide the river basin into N sub-regions, with the area of each sub-basin controlled within 50 - 200 km 2 , and evenly distribute monitoring points in each area. Using the hydrological station or sensitive area as the outlet point, generate the sub-basin boundary; taking the sub-basin as the unit, align the rainfall data and runoff data, with the initial time resolution being T1, increasing according to the rainfall intensity, and encrypted to 10 minutes during heavy rainstorms at most; Calculate the lag time Δt of the rainfall sequence and runoff sequence for each sub-basin separately, specifically: Obtain rainfall data, use cross-correlation analysis to determine the lag time between the rainfall peak and runoff peak, calculate the Pearson correlation coefficient between the rainfall sequence and runoff sequence at different lag times, and select the lag time corresponding to the maximum correlation coefficient. A single rainfall event is defined as continuous rainfall ≥ h1 and no rainfall for an interval ≥ h2, where h1 and h2 are constants; calculate the runoff coefficient C for each rainfall, and thus obtain the runoff coefficient of each sub-basin: Analyze the emission intensity of a certain pollutant in each sub-basin to obtain the emission intensity.

3. A data-driven water pollution source and pollution discharge cycle prediction method according to claim 2, characterized in that, The specific analysis of emission intensity is as follows: Mark the pollution discharge outlets at the corresponding positions on the river boundary, and extract the discharge amount of a certain pollutant at each pollution discharge outlet and mark it as the point source discharge amount A point ; Non-point source pollution includes agricultural non-point source pollution and residential non-point source pollution. The calculation and analysis process of non-point source pollution emissions is as follows: Use remote sensing satellites or GIS data to extract the farmland distribution within the sub-basin, obtain the fertilizer application rate per unit area of different crops from agricultural statistical yearbooks or real-time surveys, and use SWAT or HSPF to simulate runoff and pollutant scouring amounts. Calculation of agricultural non-point source pollution emissions, the calculation formula is: where g = 1, 2, 3... G, G represents the types of crops in the sub-watershed, g represents any one of the crop types, and α g represents the chemical fertilizer application rate of crop g, F g represents the planting area of crop g, R g represents the runoff coefficient of crop g, M pollutant represents the content of the target pollutant in the chemical fertilizer applied to crop g; Calculation of the pollutant emissions from the residential non-point sources, with the calculation formula as follows: where p = 1, 2, 3... P, P represents the total number of surface types in the sub-watershed, and p is any one of the surface types; D p represents the cumulative rate of pollutants of surface type p, W p represents the impervious area of type p, E p represents the removal rate of pollutants by the sewage treatment facilities; Sum the point source emissions A of a certain pollutant at each pollution discharge outlet within the sub - watershed point , the agricultural non - point source pollution emissions A agri , and the residential non - point source pollution emissions A urban to obtain the emission intensity A of a certain pollutant in the sub - watershed through summation calculation.

4. A data-driven prediction method for water pollution sources and pollution discharge cycles according to claim 3, characterized in that Quantify the pollution transport contribution of upstream and downstream sub-basins to the downstream, specifically: 4-1. Topological relationship construction: Based on D8 flow direction data, generate a sub-basin topological relationship network to identify upstream and downstream relationships. 4-2, Transmission Index S ij Calculated by the formula: where i and j represent any two sub-watersheds in an upstream-downstream relationship, and A i represents the emission intensity of the upstream sub-watershed i, and t ij represents the water flow time from sub-watershed i to j, which is calculated from the channel length and flow velocity, and Slope j represents the average slope of the downstream sub-watershed j, k represents the decay coefficient of the pollutant, and D ij represents the channel migration distance from sub-watershed i to j; 4-3. Update t according to the real-time rainfall intensity and river channel flow velocity per hour ij , and store it in the spatio-temporal feature database; use the rainfall-runoff coupling feature, spatial correlation feature, temporal periodicity feature, and the pollution emissions, emission intensity, and transmission index of each sub-watershed as the multi-dimensional feature vectors of each sub-watershed.

5. A data-driven prediction method for water pollution sources and pollution discharge cycles according to claim 4, characterized in that Automatically associate pollution point sources with sub-basins based on spatial topology to construct a knowledge spectrum diagram. Specifically: Determine the entity categories, which include sub-basin nodes, pollution sources, transmission paths, and sensitive areas; use sub-basins as nodes, and the attributes include area, emission intensity, runoff coefficient, rainfall-runoff lag time; pollution sources include point sources, agricultural non-point sources, and residential non-point sources. Each pollution discharge port or area is used as an entity, and the attributes include emission volume and geographical location. The transmission path represents the water flow path between sub-basins, and the attributes include flow direction, water flow time, and migration distance; sensitive areas such as drinking water sources and ecological protection areas are used as important nodes and are prominently marked with colors different from other nodes, and the attributes include population density and ecological sensitivity. Establish a spatial topology network, where the edges represent the water flow transmission paths, and the relationships between nodes reflect the possibility of pollutant migration within the basin; import the multi-dimensional feature vectors of each sub-basin as structured information, and use geographical location to automatically match the relationships between entities. Using automated rules and statistical methods, extract the relationships between entities from the characteristic data of each sub-basin, import the extracted entities and relationships into a graph database, and construct a whole-basin knowledge graph of "sub-basin - pollution source - transmission path - sensitive area".

6. A data-driven prediction method for water pollution sources and pollution discharge cycles according to claim 5, characterized in that Pollution source analysis and emission cycle prediction, specifically: Identify the spatial aggregation areas of pollutant events, distinguish the high-incidence areas of point source and non-point source pollution, and obtain historical pollution datasets and geographical data. The historical pollution datasets include the coordinates, pollutant types, and concentration values of water quality monitoring exceeding the standard events for 3 years or more; the geographical data includes soil permeability numbers, slopes, and land use types. Determine the neighborhood radius through the K-distance graph, set the minimum number of samples, use the neighborhood radius and the minimum number of samples as the parameters of the DBSCAN algorithm, run the DBSCAN algorithm, and output the clustering labels. If the clustering center is located in the industrial area and the soil permeability coefficient > the preset coefficient V1, it is determined as a point source pollution area and output; if the clustering area satisfies the slope > the preset slope V2, and the land use is farmland and the soil permeability coefficient < the preset coefficient V1, it is determined as a non-point source pollution area and output; where V1 and V2 are constants. Construct a pollutant concentration matrix for multiple pollution indicators collected at each monitoring point, use the PMF model to decompose the pollutant concentration matrix into a source contribution matrix and a source characteristic matrix, analyze the contribution rate of each pollution source to the overall pollution, and identify the main pollution factors. Adopt a long short-term memory network, use the extracted spatio-temporal features as input, fuse the enterprise production cycle data and real-time meteorological data to train the model. The model outputs the predicted pollution load values of each sub-basin in the short term and identifies short-term abnormal fluctuations. Use a dynamic Bayesian network to construct a multivariate time series prediction model; the input data includes long-term monitoring data, seasonal meteorological changes, and emission history data; the model outputs the medium- and long-term emission trends, and judges the stability of the pollutant emission cycle and potential abnormal events.

7. A data-driven prediction method for water pollution sources and pollution discharge cycles according to claim 1, characterized in that Dynamically adjust the monitoring network to cover high-risk pollution areas; specifically: 1-1: Extract the emission intensity and transmission index within each sub-basin, mark the existing monitoring points on the map of each sub-basin, connect the adjacent monitoring points to grid each sub-basin, and calculate the area of each grid. 1-2: Take any one sub-basin as the termination node, and other sub-basins that are in an upstream relationship with this sub-basin as related nodes. Extract the transmission index of each related node and the termination node for each pollutant, sum them according to the type of pollutant to obtain the transmission value of each pollutant at the termination node, and then calculate its average value to obtain the transmission average value. Normalize the pollution discharge intensity and transmission mean value Z of the termination node to obtain the pollution judgment value Q of the sub-watershed. If the pollution discharge judgment value is greater than the maximum value in the set judgment interval, then fill the sub-watershed with red and match it with the monitoring coverage rate F1 and sampling frequency L1; if the pollution discharge judgment value is within the set judgment interval, then fill the sub-watershed with yellow and match it with the monitoring coverage rate F2 and sampling frequency L2; if the pollution discharge judgment value is less than the minimum value in the set judgment interval, then fill the sub-watershed with blue and match it with the monitoring coverage rate F3 and sampling frequency L3; where F3 is the basic coverage rate. 1-3: Set the effective coverage radius of each monitoring point, calculate the coverage area of the monitoring point according to the effective coverage radius, then sum up the coverage areas of all monitoring points in the sub-watershed to obtain the total effective coverage area, and then divide the total effective coverage area by the sub-watershed area to obtain the actual monitoring coverage rate. If the actual monitoring coverage rate is less than the matched monitoring coverage rate, then additional monitoring points need to be arranged. Otherwise, there is no need to add. 1-4: Extract the area with the largest grid area in the sub-watershed as the monitoring point layout point, set a monitoring point at the middle position of this grid, re-divide the grid of the sub-watershed after arranging the monitoring points, and re-calculate the actual monitoring coverage rate until the actual monitoring coverage rate is no longer less than the matched monitoring coverage rate.

8. A data-driven prediction method for water pollution sources and pollution discharge cycles according to claim 6, characterized in that Construct a pollution early warning map for the entire watershed, specifically: Upload the updated monitoring point information to the GIS platform, including the geographical coordinates, coverage range and measurement parameters of each monitoring point; access the real-time water quality data of each monitoring point, and realize automatic data update through wireless transmission or the Internet of Things platform; introduce short-term and medium- to long-term emission load data based on LSTM and dynamic Bayesian network prediction, and compare and fuse them with the real-time monitoring data. Use high-precision DEM and watershed demarcation data as the base map, mark the boundaries of each sub-watershed, the river channel trend and sensitive areas on the map, overlay the monitoring point data in the form of layers, display the real-time data of each monitoring point, and use the results of the prediction model to generate a pollutant distribution heat map on the map to reflect the change trend of the pollution load in the future period.

Citation Information

Cited By

  • Big data-based thermoelectric power generation emission pollutant detection system and method

    CN120974238A

  • Basin scale agricultural non-point source pollution simulation method

    CN120995944A

  • Sewage treatment monitoring method and system applied to water source allocation

    CN121253789A