Quantitative tracing method for river and lake water pollution in combination with knowledge graph and machine learning

By combining knowledge graphs and machine learning, we build a hydrodynamic-water quality model and source strength-time-concentration relationship knowledge graph, and using Bi-LSTM model and Monte Carlo sampling method, we achieve rapid and accurate traceability of river and lake water pollution, solving the problems of insufficient accuracy and poor real-time performance in the existing technology.

CN120106206AActive Publication Date: 2025-06-06NANJING HYDRAULIC RES INST +2

Patent Information

Application Number
CN202510585342.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

The existing river pollution traceability technology has insufficient accuracy and poor real-time performance, making it difficult to quickly and accurately locate pollution sources and quantify their contributions.

Method used

Combining knowledge graphs and machine learning, by collecting and preprocessing river data, a knowledge graph of hydrodynamic-water quality model and source strength-time-concentration relationship is constructed, and the nonlinear mapping relationship between pollutant concentration timing and source strength of multiple outlets is learned by using the Bi-LSTM model, and the pollution source discharge port is determined by combining the Monte Carlo sampling method.

Benefits of technology

It has achieved rapid and accurate traceability of river and lake water pollution, made up for the shortcomings of low computing efficiency and difficulty in multi-source quantification, and improved the accuracy and response speed of pollution source positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106206A_ABST
    Figure CN120106206A_ABST
Patent Text Reader

Abstract

The invention discloses a quantitative source tracing method for river and lake water pollution in combination with a knowledge graph and machine learning. The invention relates to the technical field of water quality pollution traceability, and the method comprises the steps: collecting and processing target river section data, constructing a hydrodynamic force-water quality model, and obtaining river downstream pollutant diffusion characteristics and target section concentration time sequence characteristics under different discharge situations of each discharge port based on the discharge port position of a target river section through a unit pulse response test; the method comprises the following steps: constructing a source intensity-time-concentration relation knowledge graph, randomly extracting a sample set from the graph, training the sample set by adopting a machine learning method, learning a nonlinear mapping relation between a downstream section concentration time sequence and multi-outlet source intensity, carrying out dynamic inversion of a river pollutant diffusion process, and realizing rapid positioning and contribution quantification of a pollution source. And finally, generating probability distribution of pollution source positions by adopting a Monte Carlo sampling method to assist in improving the traceability judgment priority, thereby improving the accuracy and response speed of river water quality pollution traceability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water pollution source tracing, and in particular to a method for quantitatively tracing the source of river and lake water pollution by combining knowledge graph and machine learning. Background Art

[0002] With the deepening of industrialization and urbanization, river water pollution incidents occur frequently. Pollutants are characterized by multi-source complexity, spatial and temporal interweaving, and variable diffusion paths. Rapidly and accurately locating pollution sources and quantifying their contributions is a prerequisite for cutting off the pollution chain and implementing targeted governance. It is of great significance to ensuring the safety of river water ecological environment, reducing public health risks, and improving the effectiveness of river basin management.

[0003] Traditional river pollution source tracing technologies mainly include on-site investigations, sample analysis, and isotope analysis to determine the source and transmission path of pollutants. These empirical methods are highly dependent on pollutant feature libraries and can only achieve qualitative tracing. In complex watersheds, they are prone to misjudgment due to the similarity of pollution source characteristics, and the monitoring cost is high and the applicable scenarios are limited. Although the mechanism model method based on the hydrodynamic-water quality equation can simulate physical processes, the parameter calibration is complex and the calculation efficiency is low. It is difficult to deal with non-linear problems such as non-steady-state emissions and multi-source superposition, and the timeliness is insufficient. Although pure data-driven machine learning can mine high-dimensional nonlinear relationships, it lacks prior knowledge constraints on the pollution diffusion mechanism, the model has poor interpretability, and it is difficult to support dynamic inversion and multi-source contribution quantification needs. Therefore, it is necessary to propose a quantitative source tracing method for river and lake water pollution that combines knowledge graphs and machine learning to solve the above problems. Summary of the invention

[0004] The purpose of the present invention is to provide a quantitative tracing method for river and lake water pollution that combines knowledge graphs and machine learning to solve the problems of insufficient accuracy and poor real-time performance of existing river pollution tracing technologies.

[0005] The present invention provides a quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning, comprising: Step 1: determine the target river control section and monitoring river section length, collect river hydrological data, water quality data, hydraulic structure data, river topography data, meteorological data and pollution outlet location data, and pre-process the collected data; Step 2: divide the hydrodynamic-water environment model grid based on the river topography data, and use the measured river hydrological data, water quality data, and meteorological data as input to obtain the degradation and diffusion characteristics of pollutant concentrations in the study river section; Step 3: Conduct a single pulse emission test on the pollution outlets in the river, obtain the concentration-time series of the target section, generate the unit response function of each outlet, and construct a knowledge graph of the source strength-time-concentration relationship of the target section; Step 4: Build a Bi-LSTM model, extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set, and learn the nonlinear mapping relationship between the target section concentration time series and the multi-row source strength; Step five: Use the Monte Carlo sampling method to statistically analyze the competitive contribution rate weights of multiple outlets, combine the emission time and outlet location data, generate the probability distribution of the pollution source outlets, and finally determine the pollution source outlets.

[0006] Furthermore, in step 1, determining the target river control section and the length of the monitored river section includes: collecting regional administrative divisions, water function zoning, land use types, and geographical location information of industrial structure, comprehensively considering the river management subject, pollution source distribution, and runoff generation and convergence process, and determining the upstream monitoring simulation river channel length according to the river control section; The river hydrological data include: flow rate, water level or water depth; the water quality data include: target regulated pollutant concentration index; the meteorological data include: precipitation, temperature, wind speed, atmospheric pressure, relative humidity, solar radiation; the river terrain data include: river bottom elevation, river width; the hydraulic structure data include: hydraulic structure type, operation status; the pollution outlet location data include: hydraulic structure, pollution outlet distance from the target control section; Preprocess the collected data, including: cleaning, time series alignment and correction of the data, screening out and deleting outliers, erroneous values ​​and missing values, and extracting valid time series data from the processed data.

[0007] Furthermore, in step 2, the open source model EFDC model is used to establish a hydrodynamic-water quality model; the model grid division is based on the topography of the target river section and the traceability accuracy requirements, the hydrodynamic model adopts the Navier-Stokes continuity equation and momentum equation based on the orthogonal curvilinear coordinate system, and the water quality model adopts the convection-diffusion equation including source-sink terms and reaction terms.

[0008] Furthermore, in step 3, by setting the unit pulse discharge intensity and duration of each outlet, the constructed hydrodynamic-water quality model is run to obtain the concentration-time series of the downstream target section, and the unit response function is generated for the outlet i ; Construct a three-dimensional matrix ; in, is the outlet i source strength; For the time series of emissions, the temporal resolution is set according to the required accuracy; is the pollutant concentration in the downstream target section; ; in, It is the time series of pollutant concentration in the downstream target section; is the time series of source intensity concentration at outlet i; Represents convolution calculation; The knowledge graph structure includes core entities, entity relationships, metadata, and attribute features; from The sample set is extracted from the knowledge graph, with pollution source, time and concentration as the core entities of the knowledge graph, where: pollution source includes the attribute ID, geographical location, and outlet type; time includes the attribute time resolution, which is set to hours or days according to the traceability accuracy requirements; concentration includes the attribute pollutant type, pollutant concentration value, and monitoring section location; The entity relationship is based on emission events and response associations, where: emission events connect entity pollution source-time, including the relationship attributes source intensity and emission mode; response associations connect time-concentration, including the relationship attributes time lag, peak ratio, and contribution weight; The hydrodynamic parameters and model configuration are used as metadata and attribute features, where: the hydrodynamic parameters include roughness and diffusion coefficient, which are calibrated by the hydrodynamic-water quality model and stored in the knowledge graph attributes in the form of static parameters; flow velocity and water depth are used as additional attributes; The model configuration includes the EFDC gridding as an additional attribute for spatial resolution and the model time step as an additional attribute for temporal resolution; Based on the above knowledge graph structure, a source intensity-time-concentration relationship knowledge graph was constructed, and the above information was imported into Neo4j using the py2neo library in the Python script to perform structured storage of the knowledge graph.

[0009] Furthermore, in step 4, the Bi-LSTM model includes an input layer, a Bi-LSTM layer, and an output layer; The input layer receives the pollutant concentration time series of the downstream target section , time characteristics, hydrodynamic parameters; The Bi-LSTM layer consists of a forward LSTM layer and a reverse LSTM layer. The forward LSTM layer processes the input sequence in chronological order. , , ..., , from t=1 to t=T, get the forward hidden state ; The reverse LSTM layer processes the input sequence in reverse order , , ..., , from t=T to t=1, get the reverse hidden state ; The output layer adopts the form of a fully connected layer. , The hidden states in two directions are weighted fused to generate the final output : [ ; ]+

[0010] in, , Weights and biases set for the output layer; Extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set for the Bi-LSTM model to learn, and extract the pollutant concentration time series data of the downstream target section according to the time step , ,..., The time window K is based on the time accuracy requirement of source tracing, and the emission intensity of each outlet at the corresponding time point is obtained based on the source strength-time-concentration relationship knowledge graph query. , , ... represents the source strength of outlets 1, 2, 3 at time t. According to the time-lag relationship attribute in the source strength-time-concentration relationship knowledge graph, the outlet discharge time is aligned with the concentration response time of the downstream target section; Perform Min-Max normalization on the aligned concentration and source intensity time series:

[0011] In the formula, represents the original time series data, represents the normalized time series data, , Respectively represent the maximum and minimum values ​​of the original time series data; The standardized target section concentration time series samples are input into the Bi-LSTM model as training sets, and the mean square error MSE is used as the loss function:

[0012] in, is the sample prediction value, is the measured value, is the sample size; The Adam optimizer is used, the learning rate is set to 0.001, and the model is optimized and trained using the L2 regularization method.

[0013] Furthermore, in step five, the Monte Carlo sampling method is used to statistically analyze the weights of competitive contribution rates of multiple outlets, and N groups of source strength combinations are generated by simple random sampling based on historical data and emission allowance limits of different types of outlets for disturbances in pollution source emissions; Through the source intensity-time-concentration relationship knowledge graph, the disturbance concentration time series of the downstream target section is queried accordingly: ; in, Indicates the pollutant concentration of the downstream target section affected by the disturbance; Indicates the time series of pollutant concentration in the downstream target section; represents the disturbance term; Run the Bi-LSTM model to output the source strength estimation for each outlet , calculate its contribution weight by time integration:

[0014] in, represents the weight of row i; It indicates the estimated source strength of each outlet; represents the total source strength estimate of j outlets; Calculate the mean, variance, and 95% confidence interval of each outlet contribution weight:

[0015]

[0016] 95% Confidence Interval:

[0017] in, , Respectively represent the mean and variance of the contribution weight, represents the value of the kth sampling; Statistical analysis of the contribution rate of each outlet as the main contribution source in N simulations:

[0018] in, represents the contribution rate of outlet i, where , Represent the weights of outlets i and j respectively; Based on the contribution rate value, the pollution source outlet that causes the water quality of the target section of the river to exceed the standard is finally determined.

[0019] The present invention has the following beneficial effects: a quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning of the present invention, through the deep integration of mechanism model and data mining, is based on the hydrodynamic-water quality mechanism model, and integrates multi-source heterogeneous data by introducing knowledge graph technology, providing structured prior knowledge for machine learning, and learning the nonlinear mapping relationship between knowledge graph entities through machine learning to achieve fast and accurate source tracing, which makes up for the defects of low computational efficiency and difficult multi-source quantification of traditional mechanism models, and overcomes the problem of weak physical interpretability of pure data-driven methods, and provides a new paradigm for solving the problem of real-time and accurate source tracing in complex river pollution scenarios, which is of great significance to improving the decision-making response capability of smart water systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solution of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 This is a flow chart of a method for quantitatively tracing the source of river and lake water pollution that combines knowledge graph and machine learning, provided by the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. The technical solutions provided by the embodiments of the present invention are described in detail below in conjunction with the drawings.

[0023] See also Figure 1 The embodiment of the present invention provides a quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning, comprising the following steps: Step 1: Determine the target river control section and the length of the monitored river section, collect river hydrological data, water quality data, hydraulic structure data, river channel topography data, meteorological data and pollution outlet location data, and pre-process the collected data.

[0024] Specifically, in step one, determining the target river control section and the length of the monitored river section includes: collecting regional administrative divisions, water function zoning, land use types, and geographical location information of the industrial structure, comprehensively considering the river management entity, pollution source distribution, and runoff production and confluence process, and determining the upstream monitoring simulation river length based on the river control section.

[0025] The river hydrological data include: flow rate, water level or water depth; the water quality data include: target regulated pollutant concentration index; the meteorological data include: precipitation, temperature, wind speed, atmospheric pressure, relative humidity, solar radiation; the river terrain data include: river bottom elevation, river width; the hydraulic structure data include: hydraulic structure type, operation status; the pollution outlet location data include: hydraulic structure, pollution outlet distance from the target control section; Preprocess the collected data, including: cleaning, time series alignment and correction of the data, screening out and deleting outliers, erroneous values ​​and missing values, and extracting valid time series data from the processed data.

[0026] Step 2: divide the hydrodynamic-water environment model grid based on the river topography data, and use the measured river hydrological data, water quality data, and meteorological data as input to obtain the degradation and diffusion characteristics of pollutant concentrations in the study river section.

[0027] Specifically, in step 2, the open source model EFDC model is used to establish a hydrodynamic-water quality model; the model grid division is based on the topography of the target river section and the traceability accuracy requirements, the hydrodynamic model adopts the Navier-Stokes continuity equation and momentum equation based on the orthogonal curvilinear coordinate system, and the water quality model adopts the convection-diffusion equation including source-sink terms and reaction terms.

[0028] Construction of hydrodynamic model: The continuity equation and momentum equation based on orthogonal curvilinear coordinate system are adopted, and the vertical direction is processed by coordinate system.

[0029] Continuity equation:

[0030] In the formula, is the water depth, in m; u, v, w are the velocity components along the x, y, and z directions, in units .

[0031] Horizontal momentum equation:

[0032] Where u, v, and w are the horizontal flow velocities in the x, y, and z directions, respectively. ; is the reference plane Water level above, unit ; is the acceleration due to gravity, in m / s 2 ; is the water density, in kg / m 3 ; is the turbulence stress term on different planes, kg / (m·s²); is the Coriolis force coefficient, unit ; Vertical momentum equation: Conforms to the hydrostatic pressure assumption:

[0033] Where P is the hydrostatic pressure, unit: Pa (kg / (m·s²)); is the acceleration due to gravity, in m / s 2 ; is the water density, in kg / m 3 ; Calibration of hydrodynamic model: The initial value range of relevant parameters of the hydrodynamic model is set according to relevant literature, and the parameters of the hydrodynamic model are calibrated and verified according to the measured water depth and flow velocity of the study river section.

[0034] Water environment model construction: Adopting the convection-diffusion equation including source-sink term and reaction term:

[0035] Where c is the concentration of the substance, in kg / m 3 ; t is time, unit is s; D x , D y is the diffusion coefficient in the x and y directions, in m 2 / s; S is the source and sink term, unit: kg / m 3 / s;f R is the reaction term, unit: kg / m 3 / s.

[0036] Water environment model calibration: The initial value range of relevant parameters of the water environment model is set through relevant literature, and the water environment model parameters are further calibrated and verified according to the measured water quality indicators of the study river section.

[0037] Step three, conduct a single pulse emission test on the pollution outlets in the river, obtain the concentration-time series of the target section, generate the unit response function of each outlet, and construct a knowledge graph of the source strength-time-concentration relationship of the target section.

[0038] Specifically, in step 3, by setting the unit pulse discharge intensity and duration of each outlet, running the constructed hydrodynamic-water quality model, obtaining the concentration-time series of the downstream target section, and generating a unit response function for the outlet i ; Construct a three-dimensional matrix ; in, is the source intensity at outlet i, in kg / s; For the time series of emissions, the temporal resolution is set according to the required accuracy; is the pollutant concentration of the downstream target section, in mg / L; ; in, It is the time series of pollutant concentration in the downstream target section; is the time series of source intensity concentration at outlet i; Represents convolution calculation; The knowledge graph structure includes core entities, entity relationships, metadata, and attribute features; from The sample set is extracted from the knowledge graph, with pollution source, time and concentration as the core entities of the knowledge graph, where: the pollution source includes attribute ID, geographical location, and outlet type; the attribute ID is the unique identifier, the geographical location is the distance from the target section, and the outlet type includes rainwater, industry, agriculture, and life.

[0039] Time includes the attribute time resolution, which can be set to hours or days according to the traceability accuracy requirements; concentration includes the attribute pollutant type, pollutant concentration value, and monitoring section location; attribute pollutant type such as ammonia nitrogen, chemical oxygen demand (COD), etc., and the unit of pollutant concentration value, mg / L.

[0040] The entity relationship is based on emission events and response associations, where: emission events connect entity pollution source-time, including the relationship attributes source strength and emission mode; response association connects time-concentration, including the relationship attributes time lag, peak ratio, and contribution weight; the relationship attribute source strength is in kg / s; emission modes include continuous emission and pulse emission. The relationship attribute time lag is the time difference from the peak value of the pollution source emission source strength to the peak value of the monitoring section, and the peak ratio is the ratio of the peak value of the monitoring section to the source strength.

[0041] The hydrodynamic parameters and model configuration are used as metadata and attribute features, where: the hydrodynamic parameters include roughness and diffusion coefficient, which are calibrated by the hydrodynamic-water quality model and stored in the knowledge graph attributes in the form of static parameters; flow velocity and water depth are used as additional attributes; The model configuration includes the Environmental Fluid Dynamics Code (EFDC) meshing as an additional attribute for spatial resolution and the model time step as an additional attribute for temporal resolution. Based on the above knowledge graph structure, a source intensity-time-concentration relationship knowledge graph is constructed. The above information is imported into Neo4j using the py2neo library in the Python script to perform structured storage of the knowledge graph.

[0042] Step 4: Construct a Bi-LSTM model, extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set, and learn the nonlinear mapping relationship between the target section concentration time series and the multi-row source strength.

[0043] Specifically, in step 4, the bidirectional long short-term memory model (Bi-directional Long Short-Term Memory, Bi-LSTM) includes an input layer, a Bi-LSTM layer, and an output layer; The input layer receives the pollutant concentration time series of the downstream target section , time characteristics, hydrodynamic parameters; The Bi-LSTM layer consists of a forward long short-term memory (LSTM) layer and a reverse LSTM layer. The forward LSTM layer processes the input sequence in chronological order. , , ..., , from t=1 to t=T, get the forward hidden state ; The reverse LSTM layer processes the input sequence in reverse order , , ..., , from t=T to t=1, get the reverse hidden state ; The output layer adopts the form of a fully connected layer. , The hidden states in two directions are weighted fused to generate the final output : [ ; ]+

[0044] in, , Weights and biases set for the output layer; Extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set for the Bi-LSTM model to learn, and extract the pollutant concentration time series data of the downstream target section according to the time step , ,..., The time window K is based on the time accuracy requirement of source tracing, and the emission intensity of each outlet at the corresponding time point is obtained based on the source strength-time-concentration relationship knowledge graph query. , , ... represents the source strength of outlets 1, 2, 3 at time t. According to the time-lag relationship attribute in the source strength-time-concentration relationship knowledge graph, the outlet discharge time is aligned with the concentration response time of the downstream target section; Perform Min-Max normalization on the aligned concentration and source intensity time series:

[0045] In the formula, represents the original time series data, represents the normalized time series data, , Respectively represent the maximum and minimum values ​​of the original time series data; The standardized target section concentration time series samples are input into the Bi-LSTM model as training sets, and the mean square error (MSE) is used as the loss function:

[0046] in, is the sample prediction value, is the measured value, is the sample size; The Adam optimizer is used, the learning rate is set to 0.001, and the model is optimized and trained using the L2 regularization method.

[0047] Step five: Use the Monte Carlo sampling method to statistically analyze the competitive contribution rate weights of multiple outlets, combine the emission time and outlet location data, generate the probability distribution of the pollution source outlets, and finally determine the pollution source outlets.

[0048] Specifically, in step 5, the Monte Carlo sampling method is used to count the weights of competitive contribution rates of multiple outlets, and N groups of source strength combinations are generated by simple random sampling based on historical data and emission allowance limits of different types of outlets for disturbances in pollution source emissions; Through the source intensity-time-concentration relationship knowledge graph, the disturbance concentration time series of the downstream target section is queried accordingly: ; in, Indicates the pollutant concentration of the downstream target section affected by the disturbance; Indicates the time series of pollutant concentration in the downstream target section; represents the disturbance term; Run the Bi-LSTM model to output the source strength estimation for each outlet , calculate its contribution weight by time integration:

[0049] in, represents the weight of row i; It indicates the estimated source strength of each outlet; represents the total source strength estimate of j outlets; Calculate the mean, variance, and 95% confidence interval of each outlet contribution weight:

[0050]

[0051] 95% Confidence Interval (CI):

[0052] in, , Respectively represent the mean and variance of the contribution weight, represents the value of the kth sampling; Statistical analysis of the contribution rate of each outlet as the main contribution source in N simulations:

[0053] in, represents the contribution rate of outlet i, where , Represent the weights of outlets i and j respectively; Based on the contribution rate value, the pollution source outlet that causes the water quality of the target section of the river to exceed the standard is finally determined.

[0054] Table 1 Results of tracing the source of the example river

[0055] It can be seen from the above embodiments that the present invention provides a quantitative source tracing method for river pollution that combines knowledge graphs and machine learning. By collecting and processing data of target river sections, a hydrodynamic-water quality model is constructed. Based on the outlet position of the target river section, the unit pulse response test is used to obtain the pollutant diffusion characteristics of the downstream river channel and the concentration time series characteristics of the target section under different discharge scenarios of each outlet, and a "source strength-time-concentration" relationship knowledge graph is constructed. A sample set is randomly extracted from the graph, and a machine learning method is used to train the sample set. The nonlinear mapping relationship between the downstream section concentration time series and the source strength of multiple outlets is learned, and the river pollutant diffusion process is dynamically inverted to achieve rapid positioning of pollution sources and quantification of contributions. Finally, the Monte Carlo sampling method is used to generate a probability distribution of the pollution source location, which helps to improve the priority of source tracing, thereby improving the accuracy and response speed of river water quality pollution tracing.

[0056] The above disclosure is only a preferred embodiment of the present invention, which cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.

Claims

1. A quantitative tracing method for river and lake water pollution combining knowledge graph and machine learning, characterized in that: include: Step 1: determine the target river control section and monitoring river section length, collect river hydrological data, water quality data, hydraulic structure data, river topography data, meteorological data and pollution outlet location data, and pre-process the collected data; Step 2: divide the hydrodynamic-water environment model grid based on the river topography data, and use the measured river hydrological data, water quality data, and meteorological data as input to obtain the degradation and diffusion characteristics of pollutant concentrations in the study river section; Step 3: Conduct a single pulse emission test on the pollution outlets in the river, obtain the concentration-time series of the target section, generate the unit response function of each outlet, and construct a knowledge graph of the source strength-time-concentration relationship of the target section; Step 4: Build a Bi-LSTM model, extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set, and learn the nonlinear mapping relationship between the target section concentration time series and the multi-row source strength; Step five: Use the Monte Carlo sampling method to statistically analyze the competitive contribution rate weights of multiple outlets, combine the emission time and outlet location data, generate the probability distribution of the pollution source outlets, and finally determine the pollution source outlets.

2. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1 is characterized in that: In step 1, determining the target river control section and monitoring river section length includes: collecting regional administrative divisions, water function zoning, land use types, industrial structure and geographical location information, comprehensively considering the river management subject, pollution source distribution, and runoff generation and confluence process, and determining the upstream monitoring simulation river length according to the river control section; The river hydrological data include: flow rate, water level or water depth; the water quality data include: target regulated pollutant concentration index; the meteorological data include: precipitation, temperature, wind speed, atmospheric pressure, relative humidity, solar radiation; the river terrain data include: river bottom elevation, river width; the hydraulic structure data include: hydraulic structure type, operation status; the pollution outlet location data include: hydraulic structure, pollution outlet distance from the target control section; Preprocess the collected data, including: cleaning, time series alignment and correction of the data, screening out and deleting outliers, erroneous values ​​and missing values, and extracting valid time series data from the processed data.

3. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1 is characterized in that: In step 2, the open source model EFDC model is used to establish a hydrodynamic-water quality model; the model grid division is based on the topography of the target river section and the traceability accuracy requirements. The hydrodynamic model adopts the Navier-Stokes continuity equation and momentum equation based on the orthogonal curvilinear coordinate system, and the water quality model adopts the convection-diffusion equation including source-sink terms and reaction terms.

4. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1, characterized in that: In step 3, by setting the unit pulse discharge intensity and duration of each outlet, the constructed hydrodynamic-water quality model is run to obtain the concentration-time series of the downstream target section, and the unit response function is generated for the outlet i ; Construct a three-dimensional matrix ; in, is the outlet i source strength; For the time series of emissions, the temporal resolution is set according to the required accuracy; is the pollutant concentration in the downstream target section; ; in, It is the time series of pollutant concentration in the downstream target section; is the time series of source intensity concentration at outlet i; Represents convolution calculation; The knowledge graph structure includes core entities, entity relationships, metadata, and attribute features; from The sample set is extracted from the knowledge graph, with pollution source, time and concentration as the core entities of the knowledge graph, where: pollution source includes the attribute ID, geographical location, and outlet type; time includes the attribute time resolution, which is set to hours or days according to the traceability accuracy requirements; concentration includes the attribute pollutant type, pollutant concentration value, and monitoring section location; The entity relationship is based on emission events and response associations, where: emission events connect entity pollution source-time, including the relationship attributes source intensity and emission mode; response associations connect time-concentration, including the relationship attributes time lag, peak ratio, and contribution weight; The hydrodynamic parameters and model configuration are used as metadata and attribute features, where: the hydrodynamic parameters include roughness and diffusion coefficient, which are calibrated by the hydrodynamic-water quality model and stored in the knowledge graph attributes in the form of static parameters; flow velocity and water depth are used as additional attributes; The model configuration includes the EFDC gridding as an additional attribute for spatial resolution and the model time step as an additional attribute for temporal resolution; Based on the above knowledge graph structure, a source intensity-time-concentration relationship knowledge graph was constructed, and the above information was imported into Neo4j using the py2neo library in the Python script to perform structured storage of the knowledge graph.

5. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1, characterized in that: In step 4, the Bi-LSTM model includes an input layer, a Bi-LSTM layer, and an output layer; The input layer receives the pollutant concentration time series of the downstream target section , time characteristics, hydrodynamic parameters; The Bi-LSTM layer consists of a forward LSTM layer and a reverse LSTM layer. The forward LSTM layer processes the input sequence in chronological order. , , ..., , from t=1 to t=T, get the forward hidden state ; The reverse LSTM layer processes the input sequence in reverse order , , ..., , from t=T to t=1, get the reverse hidden state ; The output layer adopts the form of a fully connected layer. , The hidden states in two directions are weighted fused to generate the final output : [ ; ]+ in, , Weights and biases set for the output layer; Extract time series data from the source strength-time-concentration relationship knowledge graph as a training sample set for the Bi-LSTM model to learn, and extract the pollutant concentration time series data of the downstream target section according to the time step , ,..., The time window K is based on the time accuracy requirement of source tracing, and the emission intensity of each outlet at the corresponding time point is obtained based on the source strength-time-concentration relationship knowledge graph query. , , ... represents the source strength of outlets 1, 2, 3 at time t. According to the time-lag relationship attribute in the source strength-time-concentration relationship knowledge graph, the outlet discharge time is aligned with the concentration response time of the downstream target section; Perform Min-Max normalization on the aligned concentration and source intensity time series: In the formula, represents the original time series data, represents the normalized time series data, , Respectively represent the maximum and minimum values ​​of the original time series data; The standardized target section concentration time series samples are input into the Bi-LSTM model as training sets, and the mean square error MSE is used as the loss function: in, is the sample prediction value, is the measured value, is the sample size; The Adam optimizer is used, the learning rate is set to 0.001, and the model is optimized and trained using the L2 regularization method.

6. The quantitative source tracing method for river and lake water pollution combining knowledge graph and machine learning as claimed in claim 1, characterized in that: In step 5, the Monte Carlo sampling method is used to count the weights of competitive contribution rates of multiple outlets. For the disturbance of pollution source emissions, based on historical data and emission allowance limits of different types of outlets, N groups of source strength combinations are generated through simple random sampling; Through the source intensity-time-concentration relationship knowledge graph, the disturbance concentration time series of the downstream target section is queried accordingly: ; in, Indicates the pollutant concentration of the downstream target section affected by the disturbance; Indicates the time series of pollutant concentration in the downstream target section; represents the disturbance term; Run the Bi-LSTM model to output the source strength estimation for each outlet , calculate its contribution weight by time integration: in, represents the weight of row i; It indicates the estimated source strength of each outlet; represents the total source strength estimate of j outlets; Calculate the mean, variance, and 95% confidence interval of each outlet contribution weight: 95% Confidence Interval: in, , Respectively represent the mean and variance of the contribution weight, Represents the value of the kth sampling; Statistical analysis of the contribution rate of each outlet as the main contribution source in N simulations: in, represents the contribution rate of outlet i, where , Represent the weights of outlets i and j respectively; Based on the contribution rate value, the pollution source outlet that causes the water quality of the target section of the river to exceed the standard is finally determined.

Citation Information

Patent Citations

  • Method for calculating tracing contribution of sudden accidental water pollution source at a point source

    CN107563139A

  • Water environment pollution analysis system and method based on big data

    CN112417788A

  • Dynamic traceability analysis method and system for water pollution at discharge port

    CN114841601A

  • Water quality pollution tracing method and terminal

    CN116644889A

  • River pollution tracing method

    CN117391463A

Cited By

  • Locating and tracing method for heavy metal waste residues in hidden area

    CN120910587A

  • Agricultural watershed pollution tracing method and system, storage medium and electronic equipment

    CN121122463A

  • Tracing method based on coupling hydrodynamics and pollutant degradation equation

    CN121389886A

  • Water pollutant tracing method and system coupled with pollution trend

    CN121859263A