A quantitative tracing method for river and lake water pollution combining knowledge graph and machine learning

By combining knowledge graphs and machine learning methods, a quantitative traceability model for river and lake water pollution is constructed, which solves the problems of insufficient accuracy and poor real-time performance of river pollution traceability technology in complex river basins, and achieves rapid and accurate pollution source positioning and quantification, improving the accuracy and response speed of traceability.

CN120106206BActive Publication Date: 2025-08-19NANJING HYDRAULIC RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510585342.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-19
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing river pollution traceability technology is insufficient in complex river basins and has poor real-time performance, low computational efficiency of traditional mechanism models, and lacks physical interpretability in data-driven methods, making it difficult to achieve rapid and accurate positioning and quantification of pollution sources.

Method used

Combining knowledge graphs and machine learning, by constructing a hydrodynamic-water quality model, a knowledge graph of source strength-time-concentration relationship is generated, a nonlinear mapping relationship is learned using the Bi-LSTM model, and a Monte Carlo sampling method is used to determine the pollution source discharge port.

Benefits of technology

It has achieved rapid and accurate traceability of river and lake water pollution, improved the accuracy and response speed of traceability, overcome the problems of low computing efficiency and poor interpretability of traditional methods, and supported the decision-making response of smart water systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106206B_ABST
    Figure CN120106206B_ABST
Patent Text Reader

Abstract

The present invention discloses a quantitative source tracing method for river and lake water pollution that combines knowledge graphs and machine learning. The present invention relates to the technical field of water pollution source tracing. The present invention collects and processes data from target river sections, constructs a hydrodynamic-water quality model, and based on the outlet location of the target river section, obtains the pollutant diffusion characteristics of the downstream river channel and the concentration time series characteristics of the target section under different discharge scenarios of each outlet through unit pulse response testing, constructs a "source strength-time-concentration" relationship knowledge graph, randomly extracts a sample set from the graph, uses a machine learning method to train the sample set, learns the nonlinear mapping relationship between the downstream section concentration time series and the source strength of multiple outlets, performs dynamic inversion of the river pollutant diffusion process, realizes the rapid positioning of pollution sources and quantification of their contributions, and finally uses the Monte Carlo sampling method to generate a probability distribution of the pollution source location, which helps to improve the priority of source tracing judgment, thereby improving the accuracy and response speed of river water pollution source tracing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water pollution source tracing, and in particular to a method for quantitatively tracing the source of river and lake water pollution that combines knowledge graphs and machine learning. Background Art

[0002] With the deepening of industrialization and urbanization, river water pollution incidents occur frequently. Pollutants are characterized by multi-source complexity, temporal and spatial interweaving, and variable diffusion paths. Rapidly and accurately locating pollution sources and quantifying their contributions are the prerequisites for cutting off the pollution chain and implementing targeted governance. This is of great significance to ensuring the safety of river water ecological environment, reducing public health risks, and improving the effectiveness of river basin management.

[0003] Traditional river pollution source tracing techniques primarily involve field surveys, sample analysis, and isotope analysis to determine pollutant sources and transmission pathways. These empirical methods rely heavily on pollutant signature libraries and can only achieve qualitative source tracing. In complex watersheds, they are prone to misjudgment due to similarities in pollution source characteristics. Furthermore, monitoring costs are high and applicable scenarios are limited. Mechanistic modeling methods based on the hydrodynamic-water quality equation, while capable of simulating physical processes, suffer from complex parameter calibration, low computational efficiency, and difficulty addressing nonlinear issues such as unsteady-state emissions and multi-source superposition, as well as insufficient timeliness. While simple data-driven machine learning can mine high-dimensional nonlinear relationships, it lacks prior knowledge constraints on pollution diffusion mechanisms, resulting in poor model interpretability and difficulty supporting dynamic inversion and quantification of multi-source contributions. Therefore, it is necessary to propose a quantitative source tracing method for river and lake water pollution that combines knowledge graphs and machine learning to address these issues. Summary of the Invention

[0004] The purpose of this invention is to provide a quantitative tracing method for river and lake water pollution that combines knowledge graphs and machine learning to solve the problems of insufficient accuracy and poor real-time performance of existing river pollution tracing technologies.

[0005] The present invention provides a method for quantitatively tracing the source of river and lake water pollution by combining knowledge graphs and machine learning, including:

[0006] Step 1: Determine the target river control section and monitoring river section length, collect river hydrological data, water quality data, hydraulic structure data, river channel topography data, meteorological data, and pollution outlet location data, and pre-process the collected data;

[0007] Step 2: Based on the river topography data, the hydrodynamic-water environment model grid is divided. The measured river hydrological data, water quality data, and meteorological data are used as input to obtain the degradation and diffusion characteristics of pollutant concentrations in the study river section.

[0008] Step 3: Conduct a single pulse emission test on the pollution outlet in the river to obtain the concentration-time series of the target section, generate the unit response function of each outlet, and construct a knowledge map of the source intensity-time-concentration relationship of the target section;

[0009] Step 4: Build a Bi-LSTM model, extract time series data from the source intensity-time-concentration relationship knowledge graph as a training sample set, and learn the nonlinear mapping relationship between the target section concentration time series and the multi-row source intensity;

[0010] Step 5: Use the Monte Carlo sampling method to statistically analyze the competitive contribution rate weights of multiple outlets, combine the emission time and outlet location data, generate the probability distribution of the pollution source outlet, and finally determine the pollution source outlet.

[0011] Furthermore, in step 1, determining the target river control section and monitoring river section length includes: collecting regional administrative divisions, water function zoning, land use types, and industrial structure geographical location information, comprehensively considering the river management subject, pollution source distribution, and runoff generation and confluence process, and determining the upstream monitoring simulation river section length based on the river control section;

[0012] The river hydrological data includes: flow rate, water level or water depth; the water quality data includes: target regulated pollutant concentration index; the meteorological data includes: precipitation, temperature, wind speed, atmospheric pressure, relative humidity, and solar radiation; the river topography data includes: river bottom elevation and river width; the hydraulic structure data includes: hydraulic structure type and operation status; the pollution outlet location data includes: the distance between the hydraulic structure and the pollution outlet and the target control section;

[0013] Preprocess the collected data, including cleaning, time series alignment and correction of the data, screening out and deleting outliers, erroneous values and missing values, and extracting valid time series data from the processed data.

[0014] Furthermore, in step 2, the open source model EFDC model is used to establish a hydrodynamic-water quality model; the model grid division is based on the topography of the target river section and the traceability accuracy requirements. The hydrodynamic model adopts the Navier-Stokes continuity equation and momentum equation based on the orthogonal curvilinear coordinate system, and the water quality model adopts the convection-diffusion equation including source and sink terms and reaction terms.

[0015] Furthermore, in step 3, by setting the unit pulse discharge intensity and duration of each outlet, the constructed hydrodynamic-water quality model is run to obtain the concentration-time series of the downstream target section, and the unit response function is generated for the outlet i ; Construct a three-dimensional matrix ;

[0016] in, is the outlet source strength; For the time series of emissions, the temporal resolution is set according to the required accuracy; is the pollutant concentration at the downstream target section;

[0017] ;

[0018] in, is the time series of pollutant concentration at the downstream target section; is the time series of source intensity concentration at outlet i; Represents convolution calculation;

[0019] The knowledge graph structure includes core entities, entity relationships, metadata, and attribute features;

[0020] from The sample set is extracted from the dataset, with pollution source, time, and concentration as the core entities of the knowledge graph, where pollution source includes the attributes ID, geographic location, and outlet type; time includes the attribute time resolution, which is set to hours or days according to the traceability accuracy requirements; concentration includes the attributes pollutant type, pollutant concentration value, and monitoring section location;

[0021] The entity relationship is based on emission events and response associations, where emission events connect the entity pollution source-time, including the relationship attributes source intensity and emission mode; response association connects time-concentration, including the relationship attributes time lag, peak ratio, and contribution weight.

[0022] The hydrodynamic parameters and model configuration are used as metadata and attribute features. The hydrodynamic parameters include roughness and diffusion coefficient, which are calibrated by the hydrodynamic-water quality model and stored in the knowledge graph attributes as static parameters; flow velocity and water depth are used as additional attributes.

[0023] The model configuration includes the EFDC grid division as an additional attribute of spatial resolution and the model time step as an additional attribute of temporal resolution;

[0024] Based on the above knowledge graph structure, a source intensity-time-concentration relationship knowledge graph was constructed. The above information was imported into Neo4j using the py2neo library in the Python script to perform structured storage of the knowledge graph.

[0025] Furthermore, in step 4, the Bi-LSTM model includes an input layer, a Bi-LSTM layer, and an output layer;

[0026] The input layer receives the time series of pollutant concentrations at the downstream target section , time characteristics, hydrodynamic parameters;

[0027] The Bi-LSTM layer consists of a forward LSTM layer and a reverse LSTM layer. The forward LSTM layer processes the input sequence in chronological order. , ,..., , from t=1 to t=T, get the forward hidden state ; The reverse LSTM layer processes the input sequence in reverse order , ,..., , from t=T to t=1, get the reverse hidden state ;

[0028] The output layer adopts the form of a fully connected layer. 、 The hidden states in two directions are weighted fused to generate the final output :

[0029] [ ; ]+

[0030] in, 、 The weights and biases set for the output layer;

[0031] Extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set for the Bi-LSTM model to learn, and extract the time series data of pollutant concentration at the downstream target section according to the time step. 、 ,..., The time window K is based on the time accuracy requirement of tracing the source, and the emission intensity of each outlet at the corresponding time point is obtained based on the source strength-time-concentration relationship knowledge graph query. 、 、 ...represents the source strength of outlets 1, 2, 3... at time t. Based on the time-lag relationship attribute in the source strength-time-concentration relationship knowledge graph, the outlet discharge time is time-series aligned with the concentration response time of the downstream target section;

[0032] Perform Min-Max normalization on the aligned concentration and source intensity time series:

[0033]

[0034] Where, represents the original time series data, represents normalized time series data, 、 Respectively represent the maximum and minimum values of the original time series data;

[0035] The standardized target section concentration time series samples are input into the Bi-LSTM model as a training set, and the mean square error (MSE) is used as the loss function:

[0036]

[0037] in, is the sample prediction value, is the measured value, is the sample size;

[0038] The Adam optimizer is used, the learning rate is set to 0.001, and the model is optimized and trained using the L2 regularization method.

[0039] Furthermore, in step five, the Monte Carlo sampling method is used to statistically analyze the weights of the competitive contribution rates of multiple outlets. Based on historical data and emission permit limits of different types of outlets, N groups of source strength combinations are generated through simple random sampling for pollution source emission disturbances.

[0040] Through the source intensity-time-concentration relationship knowledge map, the disturbance concentration time series of the downstream target section is queried:

[0041] ;

[0042] in, Indicates the pollutant concentration at the downstream target section affected by the disturbance; Indicates the time series of pollutant concentration at the downstream target section; represents the disturbance term;

[0043] Run the Bi-LSTM model to output the source strength estimation for each outlet , calculate its contribution weight by time integration:

[0044]

[0045] in, represents the weight of row i; Indicates the estimated source strength of each outlet; represents the total source strength estimate of j outlets;

[0046] Calculate the mean, variance, and 95% confidence interval of each outlet's contribution weight:

[0047]

[0048]

[0049] 95% confidence interval:

[0050] in, 、 Represent the mean and variance of contribution weights, Represents the value of the kth sampling;

[0051] Statistical analysis of the contribution rate of each outlet as the main contribution source in N simulations:

[0052]

[0053] in, represents the contribution rate of outlet i, where 、 Represent the weights of outlets i and j respectively;

[0054] Based on the contribution rate value, the pollution source outlet that causes the water quality of the target section of the river to exceed the standard is finally determined.

[0055] The present invention has the following beneficial effects: a quantitative source tracing method for river and lake water pollution that combines knowledge graphs and machine learning, through the deep integration of mechanism models and data mining, is based on the hydrodynamic-water quality mechanism model, and integrates multi-source heterogeneous data by introducing knowledge graph technology to provide structured prior knowledge for machine learning. It also uses machine learning to learn the nonlinear mapping relationship between knowledge graph entities to achieve fast and accurate source tracing, making up for the defects of low computational efficiency and difficult multi-source quantification of traditional mechanism models, and overcoming the problem of weak physical interpretability of pure data-driven methods. It provides a new paradigm for solving the problem of real-time and accurate source tracing in complex river pollution scenarios, and is of great significance to improving the decision-making response capabilities of smart water systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0057] Figure 1 This is a flow chart of a quantitative source tracing method for river and lake water pollution that combines knowledge graphs and machine learning, provided by the present invention. DETAILED DESCRIPTION

[0058] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. The technical solutions provided by each embodiment of the present invention are described in detail below in conjunction with the drawings.

[0059] See also Figure 1 The embodiment of the present invention provides a method for quantitatively tracing the source of river and lake water pollution by combining knowledge graph and machine learning, including the following steps:

[0060] Step 1: Determine the target river control section and the length of the monitored river section, collect river hydrological data, water quality data, hydraulic structure data, river channel topography data, meteorological data, and pollution outlet location data, and pre-process the collected data.

[0061] Specifically, in step one, determining the target river control section and the length of the monitored river section includes: collecting regional administrative divisions, water function zoning, land use types, and geographical location information of the industrial structure, comprehensively considering the river management subject, pollution source distribution, and runoff production and convergence process, and determining the upstream monitoring simulation river length based on the river control section.

[0062] The river hydrological data includes: flow rate, water level or water depth; the water quality data includes: target regulated pollutant concentration index; the meteorological data includes: precipitation, temperature, wind speed, atmospheric pressure, relative humidity, and solar radiation; the river topography data includes: river bottom elevation and river width; the hydraulic structure data includes: hydraulic structure type and operation status; the pollution outlet location data includes: the distance between the hydraulic structure and the pollution outlet and the target control section;

[0063] Preprocess the collected data, including cleaning, time series alignment and correction of the data, screening out and deleting outliers, erroneous values and missing values, and extracting valid time series data from the processed data.

[0064] Step 2: Divide the hydrodynamic-water environment model grid based on the river topography data, and use the measured river hydrological data, water quality data, and meteorological data as input to obtain the degradation and diffusion characteristics of pollutant concentrations in the study river section.

[0065] Specifically, in step 2, the open source EFDC model is used to establish a hydrodynamic-water quality model; the model grid division is based on the topography of the target river section and the traceability accuracy requirements. The hydrodynamic model adopts the Navier-Stokes continuity equation and momentum equation based on the orthogonal curvilinear coordinate system, and the water quality model adopts the convection-diffusion equation including source-sink terms and reaction terms.

[0066] Construction of hydrodynamic model: The continuity equation and momentum equation based on the orthogonal curvilinear coordinate system are adopted, and the vertical direction is processed using the coordinate system.

[0067] Continuity equation:

[0068]

[0069] Where, is the water depth, in m; u, v, w are the velocity components along the x, y, and z directions, in units .

[0070] Horizontal momentum equation:

[0071]

[0072] Where u, v, and w are the horizontal flow velocities in the x, y, and z directions, respectively. ; is the reference plane Above water level, unit ; is the acceleration due to gravity, in m / s 2 ; is the water density, unit is kg / m 3 ; is the turbulent stress term on different planes, kg / (m·s²); is the Coriolis force coefficient, unit ;

[0073] Vertical momentum equation: consistent with the hydrostatic pressure assumption:

[0074]

[0075] Where P is the hydrostatic pressure, unit: Pa (kg / (m·s²)). is the acceleration due to gravity, in m / s 2 ; is the water density, unit is kg / m 3 ;

[0076] Calibration of the hydrodynamic model: The initial value range of the relevant parameters of the hydrodynamic model is set through relevant literature, and the hydrodynamic model parameters are calibrated and verified according to the measured water depth and flow velocity of the study river section.

[0077] Water environment model construction: Adopting the convection-diffusion equation including source-sink and reaction terms:

[0078]

[0079] Where c is the concentration of the substance, in kg / m 3 ; t is time, unit is s; D x , D y is the diffusion coefficient in the x and y directions, in m 2 / s; S is the source and sink term, unit is kg / m 3 / s;f R is the reaction term, unit is kg / m 3 / s.

[0080] Water environment model calibration: The initial value range of the relevant parameters of the water environment model is set through relevant literature, and the water environment model parameters are further calibrated and verified based on the measured water quality indicators of the study river section.

[0081] Step three: conduct a single pulse emission test on the pollution outlets in the river channel to obtain the concentration-time series of the target section, generate the unit response function of each outlet, and construct a knowledge graph of the source strength-time-concentration relationship of the target section.

[0082] Specifically, in step 3, by setting the unit pulse discharge intensity and duration of each outlet, running the constructed hydrodynamic-water quality model, obtaining the concentration-time series of the downstream target section, and generating the unit response function for the outlet i ; Construct a three-dimensional matrix ;

[0083] in, is the source intensity at outlet i, in kg / s; For the time series of emissions, the temporal resolution is set according to the required accuracy; is the pollutant concentration at the downstream target section, in mg / L;

[0084] ;

[0085] in, is the time series of pollutant concentration at the downstream target section; is the time series of source intensity concentration at outlet i; Represents convolution calculation;

[0086] The knowledge graph structure includes core entities, entity relationships, metadata, and attribute features;

[0087] from The sample set is extracted from the data, with pollution source, time and concentration as the core entities of the knowledge graph, among which: the pollution source includes attribute ID, geographical location and outlet type; the attribute ID is the unique identifier, the geographical location is the distance from the target section, and the outlet type includes rainwater, industry, agriculture and life.

[0088] Time includes the attribute time resolution, which can be set to hours or days according to the traceability accuracy requirements; concentration includes the attributes pollutant type, pollutant concentration value, and monitoring section location; attribute pollutant type such as ammonia nitrogen, chemical oxygen demand (COD), etc., and the unit of pollutant concentration value is mg / L.

[0089] The entity relationship is based on emission events and response associations. Emission events connect the entity pollution source to time, including the relationship attributes source strength and emission mode. Response associations connect time to concentration, including the relationship attributes time lag, peak ratio, and contribution weight. The relationship attribute source strength is expressed in kg / s, and emission modes include continuous and pulsed emissions. The relationship attribute time lag is the time difference between the peak emission intensity of the pollution source and the peak value at the monitoring section. The peak ratio is the ratio of the peak value at the monitoring section to the source intensity.

[0090] The hydrodynamic parameters and model configuration are used as metadata and attribute features. The hydrodynamic parameters include roughness and diffusion coefficient, which are calibrated by the hydrodynamic-water quality model and stored in the knowledge graph attributes as static parameters; flow velocity and water depth are used as additional attributes.

[0091] The model configuration includes the Environmental Fluid Dynamics Code (EFDC) meshing as an additional attribute for spatial resolution and the model time step as an additional attribute for temporal resolution.

[0092] Based on the above knowledge graph structure, a knowledge graph of source intensity, time, and concentration relationships was constructed. The above information was imported into Neo4j using the py2neo library in the Python script for structured storage of the knowledge graph.

[0093] Step 4: Build a Bi-LSTM model, extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set, and learn the nonlinear mapping relationship between the target section concentration time series and the multi-row source strength.

[0094] Specifically, in step 4, the bidirectional long short-term memory model (Bi-LSTM) includes an input layer, a Bi-LSTM layer, and an output layer;

[0095] The input layer receives the time series of pollutant concentrations at the downstream target section , time characteristics, hydrodynamic parameters;

[0096] The Bi-LSTM layer consists of a forward long short-term memory (LSTM) layer and a reverse LSTM layer. The forward LSTM layer processes the input sequence in chronological order. , ,..., , from t=1 to t=T, get the forward hidden state ; The reverse LSTM layer processes the input sequence in reverse order , ,..., , from t=T to t=1, get the reverse hidden state ;

[0097] The output layer adopts the form of a fully connected layer. 、 The hidden states in two directions are weighted fused to generate the final output :

[0098] [ ; ]+

[0099] in, 、 The weights and biases set for the output layer;

[0100] Extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set for the Bi-LSTM model to learn, and extract the time series data of pollutant concentration at the downstream target section according to the time step. 、 ,..., The time window K is based on the time accuracy requirement of tracing the source, and the emission intensity of each outlet at the corresponding time point is obtained based on the source strength-time-concentration relationship knowledge graph query. 、 、 ...represents the source strength of outlets 1, 2, 3... at time t. Based on the time-lag relationship attribute in the source strength-time-concentration relationship knowledge graph, the outlet discharge time is time-series aligned with the concentration response time of the downstream target section;

[0101] Perform Min-Max normalization on the aligned concentration and source intensity time series:

[0102]

[0103] Where, represents the original time series data, represents normalized time series data, 、 Respectively represent the maximum and minimum values of the original time series data;

[0104] The standardized target section concentration time series samples are input into the Bi-LSTM model as a training set, and the mean square error (MSE) is used as the loss function:

[0105]

[0106] in, is the sample prediction value, is the measured value, is the sample size;

[0107] The Adam optimizer is used, the learning rate is set to 0.001, and the model is optimized and trained using the L2 regularization method.

[0108] Step 5: Use the Monte Carlo sampling method to statistically analyze the competitive contribution rate weights of multiple outlets, combine the emission time and outlet location data, generate the probability distribution of the pollution source outlet, and finally determine the pollution source outlet.

[0109] Specifically, in step 5, the Monte Carlo sampling method is used to calculate the weights of the competitive contribution rates of multiple outlets. Based on historical data and emission permit limits of different types of outlets, N groups of source strength combinations are generated through simple random sampling for pollution source emission disturbances.

[0110] Through the source intensity-time-concentration relationship knowledge map, the disturbance concentration time series of the downstream target section is queried:

[0111] ;

[0112] in, Indicates the pollutant concentration at the downstream target section affected by the disturbance; Indicates the time series of pollutant concentration at the downstream target section; represents the disturbance term;

[0113] Run the Bi-LSTM model to output the source strength estimation for each outlet , calculate its contribution weight by time integration:

[0114]

[0115] in, represents the weight of row i; Indicates the estimated source strength of each outlet; represents the total source strength estimate of j outlets;

[0116] Calculate the mean, variance, and 95% confidence interval of each outlet's contribution weight:

[0117]

[0118]

[0119] 95% Confidence Interval (CI):

[0120] in, 、 Represent the mean and variance of contribution weights, Represents the value of the kth sampling;

[0121] Statistical analysis of the contribution rate of each outlet as the main contribution source in N simulations:

[0122]

[0123] in, represents the contribution rate of outlet i, where 、 Represent the weights of outlets i and j respectively;

[0124] Based on the contribution rate value, the pollution source outlet that causes the water quality of the target section of the river to exceed the standard is finally determined.

[0125] Table 1 Results of tracing the source of the example river

[0126]

[0127] It can be seen from the above embodiments that the present invention provides a quantitative source tracing method for river pollution that combines knowledge graphs and machine learning. By collecting and processing data of target river sections, a hydrodynamic-water quality model is constructed. Based on the outlet position of the target river section, the unit pulse response test is performed to obtain the pollutant diffusion characteristics of the downstream river channel and the concentration time series characteristics of the target section under different discharge scenarios of each outlet. A "source strength-time-concentration" relationship knowledge graph is constructed, and a sample set is randomly extracted from the graph. The sample set is trained using a machine learning method to learn the nonlinear mapping relationship between the downstream section concentration time series and the source strength of multiple outlets. The river pollutant diffusion process is dynamically inverted to achieve rapid positioning of pollution sources and quantification of contributions. Finally, the Monte Carlo sampling method is used to generate a probability distribution of the pollution source location to assist in improving the priority of source tracing judgment, thereby improving the accuracy and response speed of river water quality pollution tracing.

[0128] The above disclosure is only a preferred embodiment of the present invention and cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A quantitative tracing method for river and lake water pollution combining knowledge graph and machine learning, characterized by: include: Step 1: Determine the target river control section and monitoring river section length, collect river hydrological data, water quality data, hydraulic structure data, river channel topography data, meteorological data, and pollution outlet location data, and pre-process the collected data; Step 2: Based on the river topography data, the hydrodynamic-water environment model grid is divided. The measured river hydrological data, water quality data, and meteorological data are used as input to obtain the degradation and diffusion characteristics of pollutant concentrations in the study river section. Step 3: Conduct a single pulse emission test on the pollution outlet in the river to obtain the concentration-time series of the target section, generate the unit response function of each outlet, and construct a knowledge map of the source intensity-time-concentration relationship of the target section; Step 4: Build a Bi-LSTM model, extract time series data from the source intensity-time-concentration relationship knowledge graph as a training sample set, and learn the nonlinear mapping relationship between the target section concentration time series and the multi-row source intensity; Step 5: Use the Monte Carlo sampling method to calculate the weights of the competitive contribution rates of multiple outlets, combine the emission time and outlet location data, generate the probability distribution of the pollution source outlet, and finally determine the pollution source outlet; Among them, the Monte Carlo sampling method is used to calculate the weights of the competitive contribution rates of multiple outlets. In response to the disturbance of pollution source emissions, based on historical data and emission permit limits of different types of outlets, N groups of source strength combinations are generated through simple random sampling; Through the source intensity-time-concentration relationship knowledge map, the disturbance concentration time series of the downstream target section is queried: ; in, Indicates the pollutant concentration at the downstream target section affected by the disturbance; Indicates the time series of pollutant concentration at the downstream target section; represents the disturbance term; Run the Bi-LSTM model to output the source strength estimation for each outlet , calculate its contribution weight by time integration: in, represents the weight of row i; Indicates the estimated source strength of each outlet; represents the total source strength estimate of j outlets; Calculate the mean, variance, and 95% confidence interval of each outlet's contribution weight: 95% confidence interval: in, 、 Represent the mean and variance of contribution weights, Represents the value of the kth sampling; Statistical analysis of the contribution rate of each outlet as the main contribution source in N simulations: in, represents the contribution rate of outlet i, where 、 Represent the weights of outlets i and j respectively; Based on the contribution rate value, the pollution source outlet that causes the water quality of the target section of the river to exceed the standard is finally determined.

2. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1 is characterized in that: In step 1, determining the target river control section and monitoring river section length includes: collecting regional administrative divisions, water function zoning, land use types, industrial structure and geographical location information, comprehensively considering the river management entity, pollution source distribution, and runoff generation and convergence process, and determining the upstream monitoring simulation river length based on the river control section; The river hydrological data includes: flow rate, water level or water depth; the water quality data includes: target regulated pollutant concentration index; the meteorological data includes: precipitation, temperature, wind speed, atmospheric pressure, relative humidity, and solar radiation; the river topography data includes: river bottom elevation and river width; the hydraulic structure data includes: hydraulic structure type and operation status; the pollution outlet location data includes: the distance between the hydraulic structure and the pollution outlet and the target control section; Preprocess the collected data, including cleaning, time series alignment and correction of the data, screening out and deleting outliers, erroneous values and missing values, and extracting valid time series data from the processed data.

3. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1 is characterized in that: In step 2, the open source EFDC model is used to establish a hydrodynamic-water quality model. The model grid division is based on the topography of the target river section and the traceability accuracy requirements. The hydrodynamic model uses the Navier-Stokes continuity equation and momentum equation based on the orthogonal curvilinear coordinate system, and the water quality model uses the convection-diffusion equation including source and sink terms and reaction terms.

4. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1 is characterized in that: In step 3, by setting the unit pulse discharge intensity and duration of each outlet, the constructed hydrodynamic-water quality model is run to obtain the concentration-time series of the downstream target section, and the unit response function is generated for the outlet i ; Construct a three-dimensional matrix ; in, is the outlet source strength; For the time series of emissions, the temporal resolution is set according to the required accuracy; is the pollutant concentration at the downstream target section; ; in, is the time series of pollutant concentration at the downstream target section; is the time series of source intensity concentration at outlet i; Represents convolution calculation; The knowledge graph structure includes core entities, entity relationships, metadata, and attribute features; from The sample set is extracted from the dataset, with pollution source, time, and concentration as the core entities of the knowledge graph, where pollution source includes the attributes ID, geographic location, and outlet type; time includes the attribute time resolution, which is set to hours or days according to the traceability accuracy requirements; concentration includes the attributes pollutant type, pollutant concentration value, and monitoring section location; The entity relationship is based on emission events and response associations, where emission events connect the entity pollution source-time, including the relationship attributes source intensity and emission mode; response association connects time-concentration, including the relationship attributes time lag, peak ratio, and contribution weight. The hydrodynamic parameters and model configuration are used as metadata and attribute features. The hydrodynamic parameters include roughness and diffusion coefficient, which are calibrated by the hydrodynamic-water quality model and stored in the knowledge graph attributes as static parameters; flow velocity and water depth are used as additional attributes. The model configuration includes the EFDC grid division as an additional attribute of spatial resolution and the model time step as an additional attribute of temporal resolution; Based on the above knowledge graph structure, a source intensity-time-concentration relationship knowledge graph was constructed. The above information was imported into Neo4j using the py2neo library in the Python script to perform structured storage of the knowledge graph.

5. The method for quantitatively tracing the source of river and lake water pollution by combining knowledge graph and machine learning as claimed in claim 1, characterized in that: In step 4, the Bi-LSTM model includes an input layer, a Bi-LSTM layer, and an output layer; The input layer receives the time series of pollutant concentrations at the downstream target section , time characteristics, hydrodynamic parameters; The Bi-LSTM layer consists of a forward LSTM layer and a reverse LSTM layer. The forward LSTM layer processes the input sequence in chronological order. , ,..., , from t=1 to t=T, get the forward hidden state ; The reverse LSTM layer processes the input sequence in reverse order , ,..., , from t=T to t=1, get the reverse hidden state ; The output layer adopts the form of a fully connected layer. 、 The hidden states in two directions are weighted fused to generate the final output : [ ; ]+ in, 、 The weights and biases set for the output layer; Extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set for the Bi-LSTM model to learn, and extract the time series data of pollutant concentration at the downstream target section according to the time step. 、 ,..., The time window K is based on the time accuracy requirement of tracing the source, and the emission intensity of each outlet at the corresponding time point is obtained based on the source strength-time-concentration relationship knowledge graph query. 、 、 ...represents the source strength of outlets 1, 2, 3... at time t. Based on the time-lag relationship attribute in the source strength-time-concentration relationship knowledge graph, the outlet discharge time is time-series aligned with the concentration response time of the downstream target section; Perform Min-Max normalization on the aligned concentration and source intensity time series: Where, represents the original time series data, represents normalized time series data, 、 Respectively represent the maximum and minimum values of the original time series data; The standardized target section concentration time series samples are input into the Bi-LSTM model as a training set, and the mean square error (MSE) is used as the loss function: in, is the sample prediction value, is the measured value, is the sample size; The Adam optimizer is used, the learning rate is set to 0.001, and the model is optimized and trained using the L2 regularization method.

Citation Information

Patent Citations

  • Method for calculating tracing contribution of sudden accidental water pollution source at a point source

    CN107563139A

  • Dynamic traceability analysis method and system for water pollution at discharge port

    CN114841601A