Surface water multi-point source pollution source tracing method and system

Through deep learning algorithms and cluster analysis technology, a pollutant label library and knowledge base are formed, which solves the problems of insufficient accuracy and inefficiency of water quality monitoring and pollution source traceability in the existing technology, and achieves the effect of quickly identifying pollution types and source locations.

CN120217089AActive Publication Date: 2025-06-27SUN YAT SEN UNIV

Patent Information

Application Number
CN202510284877.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing technology has problems of insufficient accuracy and inefficiency in water quality monitoring and pollution source traceability, which is difficult to meet the needs of environmental protection and water resource management.

Method used

Deep learning algorithms are used to learn and classify historical pollution source data to form a pollutant label library and knowledge base, and the pollution incidents are segmented and labeled through cluster analysis to quickly identify new pollution incidents.

Benefits of technology

It improves the efficiency and accuracy of pollution traceability, can quickly identify the type of pollution and the location of pollution sources, shorten the pollution traceability time, and improve work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217089A_ABST
    Figure CN120217089A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of pollution traceability, and discloses a surface water multi-point source pollution traceability method, which comprises the following steps: preprocessing water quality time sequence data and environmental parameter information; performing feature extraction on the preprocessed water quality time sequence data to obtain water quality key features; the water quality key features are input into a pre-constructed deep learning model to be recognized so as to obtain a corresponding pollution recognition result, the deep learning model comprises a time sequence processing module, and the time sequence processing module is used for processing water quality time sequence features in the water quality key features; and outputting a corresponding pollution identification result, wherein the pollution identification result comprises pollution type information, pollution source position information and confidence degree information. According to the method, deep learning and classified calibration are performed on historical pollution source data through a deep learning algorithm, a pollutant label library and a knowledge library are formed, and support is provided for rapid identification of new pollution events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pollution source tracing, and particularly relates to a method and system for multi-point source pollution tracing of surface water. Background Art

[0002] Currently, there are problems of insufficient accuracy and low efficiency in water quality monitoring and pollution source tracing technologies. Traditional surface water pollution source tracing technologies have many deficiencies in monitoring means, tracing methods, data processing and analysis capabilities, emergency response mechanisms, and costs, and it is difficult to meet the current environmental protection and water resource management requirements. The methods or solutions that the industry may adopt for the above-mentioned disadvantages or deficiencies include: First, increasing the number and distribution of fixed monitoring stations to improve the monitoring coverage, but this method may be limited by factors such as cost, terrain, and weather; Second, adopting more advanced chemical analysis technologies or instruments to improve the tracing accuracy and accuracy, but this method may be costly and technically complex; Third, strengthening data processing and analysis capabilities, integrating data from different sources, and improving tracing efficiency, but this requires professional data processing tools and methods, as well as cross-departmental data sharing and collaboration mechanisms; Fourth, establishing a more perfect emergency response mechanism, strengthening communication and coordination between departments, and improving tracing and treatment effects, but this requires the support and cooperation of the government and relevant departments. However, these methods or solutions often can only solve some problems and have many limitations and challenges. Summary of the Invention

[0003] In view of the above defects, an embodiment of the present invention discloses a method for multi-point source pollution tracing of surface water, which can quickly identify new pollution events and improve the pollution source tracing efficiency.

[0004] A first aspect of an embodiment of the present invention discloses a method for multi-point source pollution tracing of surface water, including:

[0005] When it is detected that the tracing trigger condition is satisfied, obtain the water quality time series data and environmental parameter information of the corresponding monitoring points within a set time range, and preprocess the water quality time series data and environmental parameter information;

[0006] Extract features from the preprocessed water quality time series data to obtain key water quality features, where the key water quality features include water quality statistical features, water quality time series features, and water quality frequency domain features;

[0007] Input the key water quality features into a pre-constructed deep learning model for identification to obtain corresponding pollution identification results, where the deep learning model includes a time series processing module, and the time series processing module is used to process the water quality time series features in the key water quality features;

[0008] Output corresponding pollution identification results, where the pollution identification results include pollution type information, pollution source location information, and confidence information.

[0009] As an alternative implementation, in the first aspect of the embodiments of the present invention, after extracting features from the preprocessed water quality time series data to obtain key water quality features, it further includes:

[0010] Perform feature matching with the pollution event features in the knowledge base pre-constructed according to the water quality statistical features, water quality time series features, and water quality frequency domain features to obtain a set of model parameters associated with the current pollution event, and update the model parameters in the deep learning model according to the set of model parameters; wherein, the knowledge base includes a pollutant label library and a model parameter library, the pollutant label library includes historical pollution type labels and pollution feature descriptions; the model parameter library includes the optimal model parameters corresponding to different pollution types.

[0011] As an alternative implementation, in the first aspect of the embodiments of the present invention, the pollutant label library is determined through the following steps:

[0012] Obtain a historical pollution data set, where the historical pollution data set includes historical pollution events, historical water quality time series data, and historical environmental parameters, and preprocess the historical pollution data set;

[0013] Extract features from the preprocessed historical pollution data to obtain historical key features, and use the PCA algorithm or the t-SNA algorithm to perform dimensionality reduction processing on the historical key features to obtain dimensionality reduction key features with a contribution rate exceeding a set contribution rate;

[0014] Divide the historical pollution data set into multiple clusters, and use the dimensionality reduction key features as input;

[0015] Randomly select a data point in the historical pollution data set as the first cluster center, and select subsequent cluster centers according to the principle of maximum distance;

[0016] For each data point, calculate its distance from each cluster center, and assign it to the cluster represented by the nearest cluster center; for each cluster, calculate the mean of all data points in the cluster, and use it as the new cluster center;

[0017] Continuously update the cluster assignment of the data points and the cluster centers until the change in the cluster centers is less than a certain threshold or reaches a preset number of iterations, then the clustering is completed;

[0018] After clustering is completed, statistics are calculated for each cluster to determine the cluster characteristics, and the cluster characteristics are defined according to the set label rules to determine the historical pollution type label and pollution characteristic description, and they are saved in the pollutant label library.

[0019] As an alternative implementation manner, in the first aspect of the embodiments of the present invention, the deep learning model is constructed through the following steps:

[0020] Obtain a historical pollution training set, and construct a three-dimensional input tensor according to the historical pollution training set, where the three-dimensional input tensor includes the number of samples, the time step, and the feature dimension;

[0021] Input the three-dimensional input tensor into a pre-constructed initial deep learning model for training until the set training requirements are met. Among them, the initial deep learning model includes an input layer, a first LSTM layer, a Dropout layer, a second LSTM layer, a Dense layer, and an output layer connected in sequence; if the model output is a classification task, the cross-entropy loss function is used as the loss function, and if the model output is a regression task, the mean square error loss function is selected as the loss function;

[0022] Save the hyperparameter combination with the optimal performance as the best parameters of the model to obtain the deep learning model.

[0023] As an alternative implementation manner, in the first aspect of the embodiments of the present invention, the obtaining of the historical pollution training set and the construction of the three-dimensional input tensor according to the historical pollution training set include:

[0024] Obtain a historical pollution training set, and perform SMOTE oversampling on the data with the corresponding pollution type data in the historical pollution training set less than the set value. The historical pollution training set includes training pollution events, training water quality time series data, and training environmental parameters;

[0025] Perform preprocessing on the historical pollution training set, and the preprocessing includes missing value processing, spatial interpolation, and noise filtering;

[0026] Extract features from the preprocessed historical pollution training to obtain training key features, and perform dimensionality reduction processing on the training key features using the PCA algorithm or the t-SNA algorithm to obtain dimensionality reduction training features with a contribution rate exceeding the set contribution rate;

[0027] Construct a three-dimensional input tensor according to the dimensionality reduction training features;

[0028] And / or, after inputting the three-dimensional input tensor into a pre-constructed initial deep learning model for training until the set training requirements are met, it further includes:

[0029] For each sample, the SHAP algorithm traverses all possible feature combinations to calculate the impact on the model output when each feature is added to the feature combination; among them, the SHAP algorithm includes the Monte Carlo simulation algorithm or the Kernel SHAP algorithm;

[0030] Calculate the Shapley value of each feature according to the impact on the model output;

[0031] Sort the Shapley values of all features by absolute value to obtain the global feature importance, and output the global feature importance;

[0032] And / or, the satisfaction of the traceability trigger condition includes: obtaining the water quality indicators of each station in the monitoring area, and if the water quality indicators obtained at the corresponding station in the monitoring area exceed the set water quality parameters, the traceability trigger condition is satisfied.

[0033] As an optional implementation manner, in the first aspect of the embodiments of the present invention, after outputting the corresponding pollution identification result, it further includes: if the confidence information is less than the set confidence parameter, execute the next step;

[0034] Obtain the pollution monitoring data of the corresponding monitoring point within the set time range, and the pollution monitoring data includes water quality time series data and environmental parameter information;

[0035] Construct a water quality evolution model, and the water quality evolution model includes the analytical solution of the permanganate index pollution control equation, the analytical solution of the ammonia nitrogen pollution control equation, the analytical solution of the total phosphorus pollution control equation, and the analytical solution of the dissolved oxygen pollution control equation;

[0036] Determine an n×3 matrix as the population, where n is the population size, and each population individual includes the emission location, the total emission amount, and the emission time. According to the preset parameter value range, randomly generate values and fill them into the population matrix to form an initial population;

[0037] Define a fitness function, and calculate the sum of the squares of the relative errors between the theoretical concentration and the actual monitored concentration of each individual at all monitoring points and monitoring times. The fitness function is:

[0038]

[0039] Among them, N s is the number of pollution sources, N r is the number of monitoring points, N t is the number of monitoring time times, i is the population individual, j is the monitoring moment, is the theoretical pollutant concentration value generated by pollution source k at monitoring point i at moment j. The pollution discharge characteristics of pollution source k are (M j ,X j ,Tj ) is the monitoring concentration information at the i-th monitoring point at time j;

[0040] Calculate the result of the fitness function, determine the individuals with a smaller sum of squared relative errors as the better solutions, and retain them. Repeat the above screening, evaluation, and optimization process until the preset number of iterations is reached or the result of the fitness function meets the stopping condition;

[0041] Output the optimal parameter values corresponding to the optimal individual. The optimal parameter values are the inversion results of the pollution sources, and the optimal parameter values include the emission location, total emission volume, and emission time.

[0042] As an alternative implementation, in the first aspect of the embodiments of the present invention, the analytical solution of the permanganate index pollution control equation includes:

[0043]

[0044] where is the mass of the pollutant instantaneously released at the i-th time, X i is the location of the i-th pollution occurrence, T i is the initial emission time of the i-th pollution source, A is the average cross-sectional area of the target river section, D x is the longitudinal diffusion coefficient of the pollutant, u is the average flow velocity of the target river section, is the pollutant decay coefficient, K DO is the half-saturation constant of dissolved oxygen, C DO ( t ) is the C at time t DO concentration, is the concentration at location x and time t;

[0045] The analytical solution of the ammonia nitrogen pollution control equation includes:

[0046]

[0047]

[0048]

[0049] where is the mass of the pollutant instantaneously released at the i-th time, X i is the location of the i-th pollution occurrence, T i is the initial emission time of the i-th pollution source, A is the average cross-sectional area of the target river section, D xis the longitudinal diffusion coefficient of pollutants, u is the average flow velocity of the target river section, K1 is the rate constant for the oxidation of ammonia nitrogen to nitrite, K2 is the rate constant for the oxidation of nitrite to nitrate, and τ is an integration variable representing a certain time point between the initial emission time T of the pollution source i and the current time t, which is used to calculate the total amount of converted from during this time period; τ1 and τ2 are integration variables. τ1 represents a certain time point between the initial emission time T of the pollution source i and the current time t, while τ2 represents another time point between T i and τ1 during this time period. These two variables are jointly used to calculate the total amount of first converted to and then from converted to ;

[0050] The analytical solution of the total phosphorus pollution control equation includes:

[0051]

[0052] Among them, is the mass of pollutants instantaneously released at the i-th time, X i is the location of the i-th pollution occurrence, T i is the initial emission time of the i-th pollution source, Q k (τ) is the emission rate of the k-th external phosphorus source, X ext,k is the emission location of the k-th external phosphorus source, τ is the initial emission time of the k-th external phosphorus source, and the integration interval [τ start,k , τ end,k represents the emission period of the k-th external phosphorus source, A is the average cross-sectional area of the target river section, D x is the longitudinal diffusion coefficient of pollutants, u is the average flow velocity of the target river section, K TP is the decay rate constant of phosphorus, K r is the sediment release rate constant of phosphorus, K ext,k is the decay rate constant of the k-th external phosphorus source;

[0053] The analytical solution of the dissolved oxygen pollution control equation includes:

[0054]

[0055] Among them, is the mass of pollutants instantaneously released at the i-th time, X i is the location of the i-th pollution occurrence, T i is the initial emission time of the i-th pollution source, A is the average cross-sectional area of the target river section, D xis the longitudinal diffusion coefficient of pollutants, u is the average flow velocity of the target river section, K p is the photosynthesis rate constant, K re is the respiration rate constant, and Chla is the chlorophyll a concentration.

[0056] As an alternative implementation, in the first aspect of the embodiments of the present invention, the method for pollution source tracing further includes:

[0057] Obtain the current state of the monitoring device, where the current state of the monitoring device includes device location information, environmental parameter information, and current concentration information;

[0058] According to the obtained current state of the monitoring device, use the pollution diffusion model to calculate the pollution concentration at the current location, splice the device location and concentration data into a state vector, and input the state vector into the reinforcement learning model for identification to obtain the updated device location and updated detection time;

[0059] Determine the distance between the new device location and the pollution source and the energy consumption of the action according to the updated device location and updated detection time, and calculate the reward value according to the reward formula; if the reward meets the termination condition, output the corresponding dynamic pollution source information.

[0060] The second aspect of the embodiments of the present invention discloses a system for multi-point source pollution tracing of surface water, including:

[0061] An acquisition module: used to obtain the water quality time series data and environmental parameter information of the corresponding monitoring points within a set time range when it is detected that the tracing trigger condition is met, and preprocess the water quality time series data and environmental parameter information;

[0062] A feature extraction module: used to extract features from the preprocessed water quality time series data to obtain key water quality features, where the key water quality features include water quality statistical features, water quality time series features, and water quality frequency domain features;

[0063] An identification module: used to input the key water quality features into a pre-constructed deep learning model for identification to obtain the corresponding pollution identification result, where the deep learning model includes a time series processing module, and the time series processing module is used to process the water quality time series features in the key water quality features;

[0064] An output module: used to output the corresponding pollution identification result, where the pollution identification result includes pollution type information, pollution source location information, and confidence information.

[0065] A third aspect of an embodiment of the present invention discloses an electronic device, including: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory for executing the method for tracing the source of surface water multi-point source pollution disclosed in the first aspect of the embodiment of the present invention.

[0066] A fourth aspect of an embodiment of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program causes a computer to execute the method for tracing the source of surface water multi-point source pollution disclosed in the first aspect of the embodiment of the present invention.

[0067] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0068] The method for tracing the source of surface water multi-point source pollution in the embodiments of the present invention deeply learns and classifies and calibrates historical pollution source data through a deep learning algorithm to form a pollutant label library and a knowledge base, providing support for the rapid identification of new pollution events. Using clustering analysis to subdivide and define labels for pollution events helps to more precisely capture the differences in pollution events and improve the accuracy and efficiency of pollution source identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 is a flowchart of the method for tracing the source of surface water multi-point source pollution disclosed in the embodiment of the present invention;

[0071] Figure 2 is a flowchart of the pollution label clustering process disclosed in the embodiment of the present invention;

[0072] Figure 3 is a flowchart of the construction process of the deep learning model disclosed in the embodiment of the present invention;

[0073] Figure 4 is a flowchart of the source inversion calculation of the pollution source disclosed in the embodiment of the present invention;

[0074] Figure 5 is a diagram of the optimized layout scheme of the monitoring equipment disclosed in the embodiment of the present invention;

[0075] Figure 6 is a flowchart of the determination of pollution events disclosed in the embodiment of the present invention;

[0076] Figure 7It is an example diagram for judging pollution incidents disclosed in the embodiments of the present invention;

[0077] Figure 8 It is a schematic structural diagram of a system for tracing the sources of multi-point pollution in surface water provided by the embodiments of the present invention;

[0078] Figure 9 It is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. Detailed implementation manners

[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0080] It should be noted that the terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention are used to distinguish different objects, rather than to describe a specific order. The terms "including" and "having" in the embodiments of the present invention and any of their deformations are intended to cover non-exclusive inclusion. Exemplarily, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0081] Currently, there are problems of insufficient accuracy and low efficiency in water quality monitoring and pollution source tracing technologies. Traditional surface water pollution source tracing technologies have many deficiencies in monitoring means, tracing methods, data processing and analysis capabilities, emergency response mechanisms, and costs, and it is difficult to meet the current requirements of environmental protection and water resource management. Based on this, the embodiments of the present invention disclose methods, systems, electronic devices, and storage media for tracing multi-point pollution sources in surface water. Through deep learning algorithms, in-depth learning and classification calibration are performed on historical pollution source data to form a pollutant label library and a knowledge base, providing support for the rapid identification of new pollution incidents. Using cluster analysis to subdivide and define pollution incidents helps to more precisely capture the differences in pollution incidents and improve the accuracy and efficiency of pollution source identification.

[0082] Embodiment 1

[0083] Please refer to Figure 1 , Figure 1It is a schematic flowchart of the method for tracing the sources of multi-point surface water pollution disclosed in the embodiments of the present invention. Among them, the execution subject of the method described in the embodiments of the present invention is an execution subject composed of software or / and hardware, which can receive relevant information through wired or / and wireless means and can send certain instructions. Of course, it can also have certain processing functions and storage functions. The execution subject can control multiple devices, such as remote physical servers or cloud servers and related software, or it can also be a local host or server and related software that performs relevant operations on devices placed somewhere. In some scenarios, it can also control multiple storage devices, and the storage devices can be placed in the same place or different places as the devices. As Figures 1 to 6 shown, the method for tracing the sources of multi-point surface water pollution includes the following steps:

[0084] S101: When it is detected that the tracing trigger condition is met, obtain the water quality time-series data and environmental parameter information of the corresponding monitoring points within the set time range, and preprocess the water quality time-series data and environmental parameter information;

[0085] S102: Extract features from the preprocessed water quality time-series data to obtain key water quality features, and the key water quality features include water quality statistical features, water quality time-series features, and water quality frequency-domain features;

[0086] S103: Input the key water quality features into a pre-constructed deep learning model for identification to obtain corresponding pollution identification results. Among them, the deep learning model includes a time-series processing module, and the time-series processing module is used to process the water quality time-series features in the key water quality features;

[0087] S104: Output the corresponding pollution identification results, and the pollution identification results include pollution type information, pollution source location information, and confidence information. Through the automated data collection, processing, and analysis process, the embodiments of the present invention can quickly identify the pollution type and pollution source location from a large amount of monitoring data. Compared with traditional methods such as manual investigation, the time for pollution source tracing is greatly shortened, the work efficiency is improved, and corresponding pollution control and treatment measures can be taken in a shorter time.

[0088] The embodiment of the present invention takes into account the multi-dimensional characteristics of water quality and uses a deep learning model for processing, so that the method can better cope with the complexity and uncertainty of multi-point source pollution of surface water. It can handle complex situations such as the mixing of various types of pollutants and the discharge of different pollution sources at different times and spaces, accurately identify each pollution source and its contribution, and provide a basis for formulating a comprehensive and effective pollution control strategy. In the process of continuous pollution tracing, a large amount of water quality data, characteristic data, pollution identification results and other information will be accumulated. These data can be further used to improve and optimize the deep learning model and improve the accuracy and reliability of the model. At the same time, it also provides rich data resources and practical experience for subsequent surface water environmental protection and pollution prevention and control research, which helps to promote the continuous development and progress of related technologies.

[0089] In the work of tracing the source of multi-point pollution in surface water, most cases involve the investigation of historical pollution events. The traditional inversion tracing method is often complicated to implement because it needs to consider many complex historical environmental factors, the diffusion changes of pollutants in different periods of time, etc. It is not only time-consuming and labor-intensive, but also requires extremely high technical and data accuracy, and is prone to deviations, resulting in low efficiency in tracing. This method uses a deep learning model, which has significant advantages in tracing the source of this type of historical pollution. With its powerful feature learning and pattern recognition capabilities, the deep learning model can deeply mine a large amount of historical water quality time series data and environmental parameter information. By extracting and analyzing water quality statistical characteristics, water quality time series characteristics, and water quality frequency domain characteristics, the model can quickly capture characteristic patterns related to different types of pollution and pollution source locations. Compared with traditional inversion tracing, there is no need to tediously simulate and calculate the environmental conditions at every moment. For example, when faced with abnormal water quality in a certain area, the deep learning model can directly analyze the massive data accumulated at the monitoring point at that time, quickly provide information on the type of pollution, such as industrial wastewater pollution, domestic sewage pollution, or agricultural non-point source pollution, and roughly determine the location of the pollution source, greatly improving the speed of tracing the source. In actual application scenarios, assuming that a river basin monitors water quality deterioration for unknown reasons, if traditional inversion tracing is used, it will take a lot of manpower and material resources to collect and analyze various types of data in different seasons and different hydrological conditions over many years, and due to problems such as missing or inaccurate data, tracing the source is extremely difficult. The deep learning model of this method can quickly process these data, quickly identify possible sources and locations of pollution, and save a lot of time for the formulation of subsequent governance measures, thereby significantly improving the efficiency and accuracy of historical pollution tracing, and effectively overcoming the complexity and limitations of traditional inversion tracing methods.

[0090] More preferably, after extracting the features of the pre-processed water quality time series data to obtain the key features of water quality, the method further includes:

[0091] Feature matching is performed on the pollution event features in the knowledge base pre-constructed according to the water quality statistical features, water quality time series features, and water quality frequency domain features to obtain a set of model parameters associated with the current pollution event, and the model parameters in the deep learning model are updated according to the set of model parameters; wherein, the knowledge base includes a pollutant label library and a model parameter library, the pollutant label library includes historical pollution type labels and pollution feature descriptions; the model parameter library includes the optimal model parameters corresponding to different pollution types.

[0092] In the embodiment of the present invention, by matching the currently extracted key water quality features with the pollution event features in the knowledge base, a set of model parameters most suitable for the current pollution situation can be found and used to update the deep learning model. This can make the model better fit the current pollution data, improve the recognition accuracy of pollution types, pollution source locations, etc., and reduce misjudgment and missed judgment situations. Using the optimal model parameters in the existing knowledge base as the initialization or update basis can provide a better starting point for the training of the deep learning model, reduce the number of training iterations, accelerate the model convergence speed, thereby improving the efficiency of pollution source tracing and obtaining accurate pollution recognition results faster.

[0093] The knowledge base composed of the pollutant label library and the model parameter library is the accumulation of historical pollution events and the corresponding optimal model parameters. New pollution event data is continuously added, enriching the content of the knowledge base, realizing the inheritance of knowledge, avoiding the loss of knowledge, and providing valuable materials for subsequent pollution source tracing and research. The historical pollution type labels and pollution feature descriptions in the pollutant label library enable staff to quickly query and understand the situations of past similar pollution events, providing references for the analysis and decision-making of the current pollution event. The model parameter library facilitates directly obtaining the optimal model parameters for different pollution types, improving work efficiency and reducing repetitive labor. The pollutant label library records pollution events from dimensions such as pollution types and feature descriptions, and the model parameter library stores from the dimension of model parameters. The combination of the two supports multi-dimensional analysis of pollution events, helps to understand pollution problems more comprehensively and deeply, and provides a basis for formulating more effective pollution prevention and control strategies. As Figure 7 shown, it is the specific pollution event judgment result and the corresponding label content.

[0094] More preferably, the pollutant label library is determined through the following steps:

[0095] S100a: Obtain a historical pollution data set, the historical pollution data set includes historical pollution events, historical water quality time series data, and historical environmental parameters, and preprocess the historical pollution data set;

[0096] S100b: Extract features from the preprocessed historical pollution data to obtain historical key features, and use the PCA algorithm or the t-SNE algorithm to perform dimensionality reduction on the historical key features to obtain dimensionality-reduced key features with a contribution rate exceeding the set contribution rate;

[0097] S100c: Divide the historical pollution data set into multiple clusters and use the dimensionality-reduced key features as input; when implementing specifically, the multiple clusters here can be configured accordingly according to the actual situation;

[0098] S100d: Randomly select a data point in the historical pollution data set as the first cluster center, and select subsequent cluster centers according to the principle of maximum distance;

[0099] S100e: For each data point, calculate its distance from each cluster center, and assign it to the cluster represented by the nearest cluster center; for each cluster, calculate the mean of all data points in the cluster and use it as the new cluster center;

[0100] S100f: Continuously update the cluster assignment of data points and the cluster centers until the change in the cluster centers is less than a certain threshold or the preset number of iterations is reached, then the clustering is completed;

[0101] S100g: After the clustering is completed, calculate the internal statistics of each cluster to determine the cluster features, and define the cluster features according to the set label rules to determine the historical pollution type labels and pollution feature descriptions, and save them to the pollutant label library.

[0102] After the clustering is completed, calculating the internal statistical features of each cluster to determine the cluster features, and determining the historical pollution type labels and pollution feature descriptions according to the set label rules can convert the clustering results into pollution type labels and feature descriptions with practical significance, which is convenient for staff to quickly understand and identify different types of pollution. Saving this information to the pollutant label library realizes the effective storage and management of historical pollution knowledge, provides an important reference basis for subsequent pollution source tracing, analysis and decision-making, helps improve the efficiency and accuracy of pollution source tracing, and provides support for formulating targeted pollution prevention and control measures.

[0103] In the embodiments of the present invention, combined with a mathematical model, the historical pollution source data is learned through deep learning algorithms, clustering analysis, and label definition to form a knowledge base and establish a mathematical model, providing support for the rapid identification of new pollution events. Clustering analysis is performed on the pollution source data within 24 hours of each pollution event. Through the clustering algorithm, events with similar pollution characteristics are grouped into one category to form a pollutant label library. This step helps to identify common pollution patterns in complex datasets and provides a structured data basis for subsequent analysis. Before clustering analysis and label definition, feature selection and dimensionality reduction are performed to remove redundant features, retain the most critical information for pollution source identification, and reduce computational complexity.

[0104] In the embodiments of the present invention, the number of clustering features mentioned can be designed according to actual needs. Since the situations in different basins are different, the number of clustering features and labels here can be designed according to the actual situation.

[0105] During the clustering process, key features such as the concentration and exceeding standard degree of pollutants are considered, and the events are subdivided into different groups. For example, taking the total phosphorus exceeding the standard as an example, it is divided into intervals such as Class IV water (concentration > 0.2mg / L, ≤ 0.3mg / L), Class V water (concentration > 0.3mg / L, ≤ 0.4mg / L), and inferior Class V water (concentration > 0.4mg / L) according to the exceeding standard degree, so as to more precisely capture the differences in pollution events. Secondly, on the basis of clustering, label definitions are made for various pollution events. Label definition is based on the similarity and difference between events, and specific descriptive labels are assigned to each group of events. These labels reflect the key features of pollution events. The definition of labels not only considers the numerical changes of pollutants, but also combines statistical indicators such as change curvature, average value, variance, and extreme value difference to ensure the accuracy and comprehensiveness of labels. With the continuous accumulation of artificial traceability events, the label library is continuously enriched and improved, providing a more detailed reference for subsequent pollution event analysis.

[0106] In specific implementation, explore the adaptability of different genetic algorithms in the traceability tasks of single-point source, double-point source, and three-point source pollution events and subsequent parameter optimization and adjustment. Select five advanced genetic algorithms, including Strengthen ElitistGAAlgorithm (Enhanced Elitist Retention Genetic Algorithm), Elitist Reservation GAAlgorithm (Elitist Retention Genetic Algorithm), Polysomy Steady State GAAlgorithm (Multi-chromosome Steady State Genetic Algorithm), Generational Gap Simple GAAlgorithm (Simple Genetic Algorithm with Generation Gap), and Polysomy StudGAAlgorithm (Multi-chromosome Stud Genetic Algorithm) to conduct relevant experiments.

[0107] To simulate real pollution events and provide sufficient test data for the algorithm, 10 sets of different monitoring data were generated for each type of point source (single point source, double point source, triple point source). During the data generation process, random noise was added to the original data, and the noise level was strictly controlled between 0% and 5% to ensure the authenticity and diversity of the data. A total of 30 sets of monitoring data provided a solid data foundation for subsequent experiments. Population size adjustment and test process section: The population size was set between 100 and 300 and incremented in steps of 10, resulting in a total of 21 different population sizes. For each type of point source, the following test process was carried out for the following experiments: ① Set the population size to the initial value (such as 100). ② Use five genetic algorithms to test 10 sets of point source monitoring data in sequence. Each algorithm repeated the test 30 times for each set of data, and the source tracing results were recorded. ③ Increment the population size to the next set value (such as 110) and repeat the above test process. ④ Continue to increment the population size until the maximum value of 300 is reached to complete all tests for this type of point source. ⑤ Organize and analyze the test data, and record key information such as the number of point sources in each test, parameter settings (population size), and the group of monitoring data used. In summary, this experiment was carried out a total of 94,500 times (5 algorithms × 21 population sizes × 10 sets of monitoring data × 30 repetitions × 3 types of point sources), aiming to comprehensively and deeply evaluate the source tracing capabilities of different genetic algorithms for single point source, double point source, and triple point source pollution events at different population sizes, providing a solid data support for subsequent algorithm performance comparison and optimization.

[0108] Specifically, the algorithm selection for clustering analysis: DBSCAN (suitable for noisy data) or K-means++ (the number of categories needs to be preset). Input data: The feature matrix after dimensionality reduction (such as 10-dimensional data after PCA). Key parameters: DBSCAN: Neighborhood radius (eps = 0.5), minimum number of samples (min_samples = 5). K-means: In the embodiment of the present invention, the optimal number of clusters is 4.

[0109] The label definition rules can be based on statistical features: Label 1: Peak-type pollution (such as a sudden increase in concentration at a certain monitoring point > 50%), Label 2: Continuous pollution (concentration continuously higher than the threshold > 6 hours). Based on environmental associations: Label 3: Flow velocity sensitive type (pollution concentration is strongly correlated with flow velocity, r > 0.7). Label 4: Temperature sensitive type (pollution concentration is strongly correlated with temperature, r > 0.6).

[0110] Knowledge base architecture design, module composition: Pollutant label library: stores historical event labels and feature descriptions. Model parameter library: records the optimal model parameters corresponding to different pollution types (such as the number of LSTM units, learning rate). Rule library: expert experience rules (such as "total phosphorus exceeds the standard and the flow rate is low → the probability of agricultural non-point source pollution is greater than 70%").

[0111] More preferably, the deep learning model is constructed by the following steps:

[0112] S1031: Obtain a historical pollution training set, and construct a three-dimensional input tensor according to the historical pollution training set, wherein the three-dimensional input tensor includes the number of samples, the time step, and the feature dimension;

[0113] S1032: Input the three-dimensional input tensor into a pre-built initial deep learning model for training until the set training requirements are met, wherein the initial deep learning model includes an input layer, a first LSTM layer, a Dropout layer, a second LSTM layer, a Dense layer, and an output layer connected in sequence; if the model output is a classification task, a cross entropy loss function is used as the loss function; if the model output is a regression task, a mean square error loss function is selected as the loss function;

[0114] S1033: Save the hyperparameter combination with the best performance as the optimal parameters of the model to obtain a deep learning model.

[0115] By defining the time step and feature dimension, the embodiment of the present invention can capture the relationship between time series information and different features in the data at the same time, which helps the model learn the change pattern of pollution data over time and the impact of each feature on pollution, and improves the model's understanding and prediction capabilities of pollution conditions.

[0116] The first and second LSTM layers in the initial deep learning model can process sequence data well, effectively extract long-term dependent features in contaminated data, and capture complex dynamic changes in the contamination process. The Dropout layer can prevent the model from overfitting, improve the generalization ability of the model, and enable the model to perform well when facing new and unseen data. The Dense layer can further combine and weight the extracted features, explore the potential relationship between features, and provide a more representative feature representation for the final output.

[0117] Select the cross-entropy loss function or the mean squared error loss function according to the type of the model output task, which can optimize the model for different task objectives. For classification tasks, the cross-entropy loss function can effectively measure the difference between the model prediction result and the true label, guiding the model to learn the correct classification boundary; for regression tasks, the mean squared error loss function can measure the error between the model prediction value and the true value, enabling the model to be trained in the direction of reducing the error and improving the prediction accuracy.

[0118] More preferably, the obtaining of the historical pollution training set and the construction of the three-dimensional input tensor according to the historical pollution training set include:

[0119] Obtain the historical pollution training set, perform SMOTE oversampling on the data of the corresponding pollution type in the historical pollution training set that is less than the set value, and the historical pollution training set includes training pollution events, training water quality time series data, and training environmental parameters;

[0120] Perform preprocessing on the historical pollution training set, and the preprocessing includes missing value processing, spatial interpolation, and noise filtering;

[0121] Extract features from the preprocessed historical pollution training to obtain training key features, and perform dimensionality reduction processing on the training key features using the PCA algorithm or the t-SNA algorithm to obtain dimensionality reduction training features with a contribution rate exceeding the set contribution rate;

[0122] Construct a three-dimensional input tensor according to the dimensionality reduction training features;

[0123] And / or, after inputting the three-dimensional input tensor into the pre-constructed initial deep learning model for training until the set training requirements are met, it further includes:

[0124] For each sample, the SHAP algorithm traverses all possible feature combinations to calculate the impact on the model output after each feature is added to the feature combination; wherein, the SHAP algorithm includes the Monte Carlo simulation algorithm or the Kernel SHAP algorithm;

[0125] Calculate the Shapley value of each feature according to the impact on the model output;

[0126] Sort the Shapley values of all features by the absolute value to obtain the global feature importance, and output the global feature importance;

[0127] And / or, the satisfaction of the traceability trigger condition includes: obtaining the water quality indicators of each station in the monitoring area, and if the water quality indicators obtained by the corresponding station in the monitoring area exceed the set water quality parameters, the traceability trigger condition is satisfied.

[0128] Specifically, the data sources include historical pollution event data: including the location of the pollution source (X), the emission volume (M), the emission time (T), and the pollutant concentration (COD, ammonia nitrogen, total phosphorus, etc.); water quality time series data: the time series concentration data of each monitoring point within 24 hours (sampling frequency ≥ 1 time / hour). Environmental parameters: auxiliary data such as flow velocity, water depth, temperature, dissolved oxygen, and chlorophyll a.

[0129] Data cleaning and enhancement, missing value processing: Use time series interpolation (such as cubic spline interpolation) to fill in the missing values. Noise filtering: Use wavelet denoising (such as Daubechies wavelet) to eliminate high-frequency noise. Data enhancement: Oversample the minority pollution events (such as SMOTE algorithm) to balance the dataset.

[0130] Feature extraction and dimensionality reduction, key feature extraction: Statistical features: mean, variance, extreme value, change rate (such as the concentration change per hour). Time series features: Sliding window statistics (such as the average concentration in the past 3 hours), and Fourier transform to extract frequency domain features. Environment-related features: Pearson correlation coefficients between pollutant concentration and flow velocity, temperature. Dimensionality reduction processing: Use PCA or t-SNE to reduce the dimensionality of high-dimensional features, and retain the principal components with a cumulative contribution rate > 90%. Feature examples: 1. Statistical features: Determine mutation type and persistent type; 2. Time series features: Determine the positive and negative correlations between concentration and time changes; 3. Environment-related features: Total phosphorus and turbidity, flow velocity, dissolved oxygen and temperature, chlorophyll a, pH, ammonia nitrogen and permanganate index.

[0131] When implementing specifically, the input format: three-dimensional tensor (number of samples × time step × feature dimension), such as (1000 × 24 × 8). Model training and hyperparameter tuning in the embodiments of the present invention, loss function: cross-entropy loss for classification tasks, and mean square error (MSE) for regression tasks. Optimizer: Adam (learning rate = 0.001, decay rate = 1e-6). Hyperparameter tuning: Use grid search or Bayesian optimization to adjust the number of LSTM units, Dropout rate, batch size (recommended range: 32 - 128). Feature importance analysis method: Use SHAP values (Shapley Additive Explanations) to quantify the contribution of each feature to the prediction result. Example: If the SHAP value of the dissolved oxygen concentration is the highest, it indicates that it has a significant impact on the identification of the pollution source.

[0132] More preferably, if the confidence information is less than the set confidence parameter, then perform the next step;

[0133] S105: Obtain the pollution monitoring data of the corresponding monitoring points within the set time range, and the pollution monitoring data includes water quality time series data and environmental parameter information;

[0134] S106: Construct a water quality evolution model, where the water quality evolution model includes the analytical solution of the permanganate index pollution control equation, the analytical solution of the ammonia nitrogen pollution control equation, the analytical solution of the total phosphorus pollution control equation, and the analytical solution of the dissolved oxygen pollution control equation;

[0135] S107: Determine an n×3 matrix as the population, where n is the population size, and each population individual includes the discharge location, the total discharge amount, and the discharge time. According to the preset parameter value range, randomly generate values and fill them into the population matrix to form an initial population;

[0136] S108: Define a fitness function, calculate the sum of the squares of the relative errors between the theoretical concentration and the actual monitored concentration of each individual at all monitoring points and monitoring times. The fitness function is:

[0137]

[0138] where, N s is the number of pollution sources, N r is the number of monitoring points, N t is the number of monitoring time instances, i is the population individual, j is the monitoring time, is the theoretical pollutant concentration value generated by pollution source k at the j-th moment and the i-th monitoring point. The pollution discharge characteristics of pollution source k are (M j , X j , T j ); is the monitoring concentration information at the j-th moment and the i-th monitoring point;

[0139] S109: Calculate the result of the fitness function, determine the individual with a smaller sum of the squares of the relative errors as the better solution, and retain it. Repeat the above screening evaluation and optimization process until the preset number of iterations is reached or the result of the fitness function meets the stop condition;

[0140] S1010: Output the optimal parameter values corresponding to the optimal individual. The optimal parameter values are the inversion results of the pollution sources, and the optimal parameter values include the discharge location, the total discharge amount, and the discharge time.

[0141] More preferably, the analytical solution of the permanganate index pollution control equation includes:

[0142]

[0143] where, is the mass of the pollutant instantaneously discharged at the i-th time, X i is the i-th pollution occurrence location, T i is the initial discharge moment of the i-th pollution source, A is the average cross-sectional area of the target river section, D x is the longitudinal diffusion coefficient of the pollutant, and u is the average flow velocity of the target river section. is the pollutant decay coefficient, K DO is the half-saturation constant of dissolved oxygen, C DO ( t ) is the C at time t DO concentration, is the concentration at position x and time t;

[0144] The analytical solution of the ammonia nitrogen pollution control equation includes:

[0145]

[0146]

[0147]

[0148]

[0149] where, is the mass of the pollutant instantaneously released at the i-th time, X i is the i-th pollution occurrence position, T i is the initial emission time of the i-th pollution source, A is the average cross-sectional area of the target river section, D x is the longitudinal diffusion coefficient of the pollutant, u is the average flow velocity of the target river section, K1 is the rate constant for the oxidation of ammonia nitrogen to nitrite, K2 is the rate constant for the oxidation of nitrite to nitrate, τ is an integration variable representing a certain time point between the initial emission time T i of the pollution source and the current time t, which is used to calculate the total amount of converted from during this time period; τ1 and τ2 are integration variables. τ1 represents a certain time point between the initial emission time T i of the pollution source and the current time t, while τ2 represents another time point between T i and τ1 during this time period. These two variables are jointly used to calculate the total amount of first converted to and then from converted to ;

[0150] The analytical solution of the total phosphorus pollution control equation includes:

[0151]

[0152] where, is the mass of the pollutant instantaneously released at the i-th time, X i is the i-th pollution occurrence position, T i is the initial emission time of the i-th pollution source, Qk (τ) is the emission rate of the k-th external phosphorus source, X ext,k is the emission location of the k-th external phosphorus source, τ is the initial moment of the k-th external phosphorus source emission, and the integration interval [τ start,k , τ end,k represents the emission period of the k-th external phosphorus source, A is the average cross-sectional area of the target river section, D x is the longitudinal diffusion coefficient of the pollutant, u is the average flow velocity of the target river section, K TP is the decay rate constant of phosphorus, K r is the sediment release rate constant of phosphorus, K ext,k is the decay rate constant of the k-th external phosphorus source;

[0153] The analytical solution of the dissolved oxygen pollution control equation includes:

[0154]

[0155] Among them, is the mass of the pollutant instantaneously released at the i-th time, X i is the pollution occurrence location at the i-th time, T i is the initial moment of the i-th pollution source emission, A is the average cross-sectional area of the target river section, D x is the longitudinal diffusion coefficient of the pollutant, u is the average flow velocity of the target river section, K p is the photosynthesis rate constant, K re is the respiration rate constant, and Chla is the chlorophyll a concentration.

[0156] In the embodiments of the present invention, source identification is essentially a process of finding the optimal solution by comparing the calculation results of the mass transfer model with the real monitoring data, that is, an optimization problem. In this process, it aims to find a set of source parameters to minimize the error between the concentration predicted by the model and the actually monitored concentration. Specifically, the monitored concentration is defined as Cmes (mg / L), while Cth is the calculated concentration simulated by the surface water quality model based on the source information X (including the emission location X (m), the total emission amount M (g), and the emission time T (min)). The goal of the embodiments of the present invention is to find the set of X that minimizes the error between Cth and Cmes. To achieve this goal, an inversion algorithm coupling the surface water quality model and the genetic algorithm is adopted. The surface water quality model is used to describe the change of pollutant concentration with time and space in a pollution accident, while the genetic algorithm is used to find the optimal solution in this complex multi-dimensional space.

[0157] Based on the analytical solution of the classical water quality diffusion equation and combined with the hydrodynamic code (EFDC), select the water quality data categories that can be obtained by on-line monitoring equipment, compile the analytical solutions of different water quality indicators for multi-point source pollutants, and construct an optimized water quality evolution model. The water quality evolution model: Based on the water quality diffusion equation and the hydrodynamic code, combined with the data of small water quality monitoring stations, construct a water quality evolution model. This model provides an optimized analytical solution, accurately reflects the water quality status, and provides a scientific basis for pollution source tracing.

[0158] Using the genetic algorithm and the multi-point source tracing spatio-temporal inversion algorithm, the genetic algorithm is used to transform the pollution source tracing problem into an iterative optimization problem of population individuals, and continuously iteratively optimize to approximate the real pollution source information. At the same time, combined with the water quality evolution model of sub-indicators, construct a multi-point source tracing spatio-temporal inversion algorithm to achieve accurate inversion of instantaneous emissions from multi-point sources and improve the accuracy of pollution source tracing. Through comparative test experiments, a more suitable genetic algorithm - the genetic algorithm with enhanced elite retention is selected and the parameters are adjusted. This algorithm enhances the exploration ability of new solutions while retaining excellent individuals, is more suitable for the processing and analysis of water quality monitoring data, and further optimizes the effect of pollution source tracing. In the subsequent introduction of reinforcement learning elements, the algorithm can dynamically adjust the search strategy according to the feedback and converge to the optimal solution faster. At the same time, design an adaptive mechanism to dynamically adjust algorithm parameters such as population size and crossover mutation rate according to the characteristics of pollution data.

[0159] More preferably, the method for pollution source tracing further includes:

[0160] Obtain the current state of the monitoring equipment, and the current state of the monitoring equipment includes equipment location information, environmental parameter information, and current concentration information;

[0161] According to the obtained current state of the monitoring equipment, use the pollution diffusion model to calculate the pollution concentration at the current location, splice the equipment location and concentration data into a state vector, and input the state vector into the reinforcement learning model for identification to obtain the updated equipment location and updated detection time;

[0162] Determine the distance between the new equipment location and the pollution source and the energy consumption of the action according to the updated equipment location and updated detection time, and calculate the reward value according to the reward formula; if the reward meets the termination condition, output the corresponding dynamic pollution source information.

[0163] Such as Figure 5As shown in the figure, the system integrates unmanned water quality monitoring vessels, drones, distributed monitoring instruments, and small water quality monitoring stations to form a full-life cycle monitoring system for "source-network-plant-river-basin". The integration of multi-source information greatly improves the flexibility and coverage of monitoring, enabling more comprehensive capture of water quality changes. Optimizing the layout of monitoring points: Using unmanned monitoring technology to optimize the layout of monitoring points to ensure efficient monitoring at the lowest cost is incomparable to traditional monitoring methods. The objective of the embodiment of the present invention: Dynamically adjust the position through mobile monitoring devices (such as unmanned vessels and drones) to quickly locate pollution sources. Dynamic pollution sources: Assume that pollution sources may move (such as illegal dumping vehicles) or the diffusion path is affected by environmental changes.

[0164] In the embodiment of the present invention, a suitable reward function is designed, and the accuracy of pollution source identification, identification speed, etc. are used as reward indicators. When the model makes a correct identification decision, a positive reward is given; when the identification is incorrect or inefficient, a negative reward is given. The model continuously interacts with the environment and dynamically adjusts the search strategy according to the reward feedback. For example, in the process of searching for the optimal solution, according to the magnitude and direction of the reward, it is decided whether to adjust the parameters of the model or change the search path, so as to converge to the optimal solution faster.

[0165] Specifically, the definition of reinforcement learning elements, State: Current monitoring data (COD, ammonia nitrogen concentration, etc.), monitoring device position coordinates (x, y), environmental parameters (flow velocity, wind direction, temperature). Action: Movement direction (Δx, Δy), range limit (such as no more than 10 meters per movement), monitoring frequency adjustment (optional). Reward: Positive reward: The reward increases when approaching the pollution source (such as the square of the distance reduction), negative penalty: energy consumption cost (linear function of the movement distance), time cost; formula example: R = -||d current -d previous ||2 - λ, where d is the distance between the device and the pollution source, and λ is the energy consumption weight.

[0166] Pollution diffusion model, using the classical water quality diffusion equation to simulate the pollution concentration field:

[0167]

[0168] Among them, (x0, y0) is the pollution source position, u x , u y is the flow velocity component.

[0169] In the embodiments of the present invention, the current state of the monitoring device including device location information, environmental parameter information, and current concentration information is obtained, which can provide a multi-dimensional data basis for pollution source tracing. The device location information can clarify the spatial distribution of the monitoring points, facilitating the determination of the location of pollution in space; the environmental parameter information helps to understand the external conditions for pollution diffusion, such as wind speed, water flow speed, etc., which will affect the diffusion direction and speed of pollutants; the current concentration information directly reflects the degree of pollution and provides a basis for analyzing the development trend of pollution. It is possible to grasp the state changes of the monitoring device in real time, timely discover abnormal fluctuations in pollution concentration and changes in environmental parameters, provide real-time data support for subsequent pollution diffusion analysis and source tracing, and help to take measures in a timely manner to respond to pollution incidents.

[0170] Using a pollution diffusion model to calculate the pollution concentration at the current location based on the current state of the monitoring device can simulate the diffusion process of pollutants in the environment. Considering the influence of environmental factors on pollution diffusion, it can more accurately predict the spread range and concentration distribution changes of pollution, providing more accurate pollution concentration data for source tracing. Concatenating the device location and concentration data into a state vector and inputting it into a reinforcement learning model for identification to obtain an updated device location and an updated detection time, this process realizes intelligent decision-making. The reinforcement learning model can automatically learn and determine the optimal device location and detection time according to the current pollution situation and environmental information to improve the tracking efficiency of pollution sources, optimize the monitoring strategy, and make the monitoring more scientific and reasonable. It is possible to adjust the location and detection time of the monitoring device in real time according to the changing pollution situation and environmental conditions, enabling the entire source tracing system to have an adaptive ability, being able to better cope with complex and changeable pollution scenarios, and improving the accuracy and reliability of source tracing.

[0171] The embodiments of the present invention have a dynamic tracking mechanism: the monitoring device is autonomously moved through RL to adapt to the change of the pollution source location. An energy consumption-accuracy trade-off is adopted: an energy consumption penalty is introduced into the reward function to avoid ineffective movement. Environmental interaction simulation: combining a mechanism model to simulate pollution diffusion to enhance the authenticity of training.

[0172] During the model training and prediction process of the embodiments of the present invention, an adaptive mechanism is introduced, enabling the model to automatically adjust parameters and strategies according to different input data and environmental conditions. For example, according to the complexity and data characteristics of a new pollution incident, parameters such as the learning rate and batch size of the deep learning model are dynamically adjusted to improve the adaptability and performance of the model.

[0173] In the embodiments of the present invention, an unmanned monitoring ship, an unmanned aerial vehicle, distributed monitoring stations and Internet of Things technology are integrated to construct a full-life-cycle monitoring system of "source-network-factory-river-basin", breaking through the limitations of traditional fixed monitoring stations and improving the monitoring coverage and flexibility. Existing technologies mostly rely on single monitoring means (such as fixed stations or manual sampling). Through multi-source data fusion and optimized layout, this research significantly reduces costs and improves real-time performance.

[0174] In the embodiments of the present invention, a multi-point source tracing spatio-temporal inversion algorithm is developed by combining a water quality diffusion equation and a genetic algorithm, transforming the source pollution inversion into an iterative optimization problem, and optimizing the algorithm performance through an enhanced elite retention strategy.

[0175] In the embodiments of the present invention, deep learning (a variant of RNN) is introduced to process time series data, and a knowledge base is constructed by combining cluster analysis and label definition to achieve intelligent learning of historical data and rapid matching of new events. Existing research mostly uses mechanism models (such as diffusion models) or data-driven models (such as machine learning) alone. Through the collaboration of the two, this research takes into account physical laws and data characteristics, significantly improving the source tracing accuracy.

[0176] In the embodiments of the present invention, analytical solutions for the diffusion of pollutants are derived for multiple water quality indicators such as total phosphorus, ammonia nitrogen, and dissolved oxygen, and dynamic parameters (such as sediment release rate, photosynthesis rate) are introduced to optimize the model adaptability. Traditional models mostly target single pollutants or simplified parameters, and this model is more suitable for the complex environment of actual water bodies.

[0177] In the embodiments of the present invention, the parameters of the genetic algorithm (such as population size, crossover rate) are dynamically adjusted through reinforcement learning, and a pollutant label library is formed by combining cluster analysis, realizing the self-learning and knowledge accumulation of the algorithm. Existing algorithms mostly rely on fixed parameters and lack dynamic optimization capabilities. This research is more robust in complex multi-source pollution scenarios.

[0178] Specifically, the implementation steps for cross-basin generalization source tracing driven by transfer learning in this embodiment are as follows: Data collection, source basin data: historical pollution event data, water quality monitoring data (COD, ammonia nitrogen, etc.), geographical information (flow velocity, water depth), source pollution labels. Target basin data: a small amount of monitoring data (without source pollution labels or partial labels), geographical information. Data preprocessing, standardization processing: perform Z-score standardization on monitoring data (such as concentration, flow velocity). Spatio-temporal alignment: unify timestamps and spatial coordinates (such as converting to the UTM coordinate system).

[0179] The specific implementation of the solution in the embodiment of the present invention is as follows: First, analyze the historical data of the monitoring sites. The main indicators of the historical data include total phosphorus, ammonia nitrogen, permanganate index, and dissolved oxygen. When performing data scraping, it is necessary to consider the data 24 hours before and 23 hours after the exceeding standard moment, output all monitoring factors within the corresponding time period, and independently organize the single warning event of each site into corresponding table data. Perform cluster analysis on the above independent table data and configure corresponding label classifications for it.

[0180] Second, in specific implementation, through the logical classification sorted out by on-site manual pollution source tracing experience and in cooperation with the clustering algorithm, improve the setting of various labels of existing pollution events, and on the basis of label definition, directly define the possible sources of pollution of this pollution event in combination with the expert experience library. For example: expert experience rules (such as "total phosphorus exceeds the standard and the flow rate is low → the probability of agricultural non-point source pollution > 70%"), and identify pollution events with higher occurrence frequencies (such as sewage treatment plant overflow, rainfall, agricultural non-point source irrigation, etc.) in the above way.

[0181] Third, if a black-box pollution event occurs (that is, an unknown and sudden pollution event) and the result cannot be obtained through the existing mathematical model, then enable the mechanism model, and obtain the approximate pollution location through the formula and algorithm of the mechanism model, reduce the scope of manual investigation, and finally determine the cause of the pollution event through manual source tracing. Fourth, add the confirmed pollution event data to the training set, and retrain the model regularly later to update the knowledge base labels and rules.

[0182] The method for tracing multi-point source pollution of surface water in the embodiment of the present invention deeply learns and classifies and calibrates the historical pollution source data through a deep learning algorithm to form a pollutant label library and a knowledge base, providing support for the rapid identification of new pollution events. Using cluster analysis to subdivide and define labels for pollution events helps to more precisely capture the differences of pollution events and improve the accuracy and efficiency of pollution source identification.

[0183] Embodiment 2

[0184] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the system for tracing multi-point source pollution of surface water disclosed in the embodiment of the present invention. As Figure 8 shown, the system for tracing multi-point source pollution of surface water may include:

[0185] An acquisition module 21: configured to, when it is detected that the tracing trigger condition is satisfied, acquire the water quality time series data and environmental parameter information of the corresponding monitoring points within the set time range, and preprocess the water quality time series data and environmental parameter information;

[0186] Feature extraction module 22: It is used to extract features from the preprocessed water quality time series data to obtain key water quality features, where the key water quality features include water quality statistical features, water quality time series features, and water quality frequency domain features;

[0187] Recognition module 23: It is used to input the key water quality features into a pre-constructed deep learning model for recognition to obtain corresponding pollution recognition results. Among them, the deep learning model includes a time series processing module, and the time series processing module is used to process the water quality time series features in the key water quality features;

[0188] Output module 24: It is used to output corresponding pollution recognition results, and the pollution recognition results include pollution type information, pollution source location information, and confidence information.

[0189] In the embodiment of the present invention, the method for tracing the sources of multi-point pollution in surface water deeply learns and classifies and calibrates historical pollution source data through a deep learning algorithm to form a pollutant label library and a knowledge base, providing support for the rapid identification of new pollution events. Using cluster analysis to subdivide and define pollution events helps to more precisely capture the differences in pollution events and improve the accuracy and efficiency of pollution source identification.

[0190] Embodiment III

[0191] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of an electronic device disclosed in the embodiment of the present invention. The electronic device can be a computer, a server, etc. Of course, in certain cases, it can also be an intelligent device such as a mobile phone, a tablet computer, and a monitoring terminal, as well as an image acquisition device with processing functions. As Figure 9 shown, the electronic device may include:

[0192] A memory 510 storing executable program code;

[0193] A processor 520 coupled to the memory 510;

[0194] Among them, the processor 520 calls the executable program code stored in the memory 510 to execute some or all of the steps in the method for tracing the sources of multi-point pollution in surface water in Embodiment I.

[0195] The embodiment of the present invention discloses a computer-readable storage medium, which stores a computer program, where the computer program enables a computer to execute some or all of the steps in the method for tracing the sources of multi-point pollution in surface water in Embodiment I.

[0196] An embodiment of the present invention also discloses a computer program product. When the computer program product runs on a computer, the computer is caused to execute some or all of the steps in the method for tracing the source of multi-point source pollution of surface water in the first embodiment.

[0197] An embodiment of the present invention also discloses an application publishing platform. The application publishing platform is used to publish a computer program product. When the computer program product runs on a computer, the computer is caused to execute some or all of the steps in the method for tracing the source of multi-point source pollution of surface water in the first embodiment.

[0198] In various embodiments of the present invention, it should be understood that the magnitude of the serial numbers of the various processes does not necessarily mean the inevitable sequence of execution. The execution sequence of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0199] The unit described as a separate component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0201] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several requests for causing a computer device (which may be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute some or all of the steps of the methods described in the various embodiments of the present invention.

[0202] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0203] Those of ordinary skill in the art can understand that some or all of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, magnetic tape memories, or any other medium that can be used to carry or store data and is readable by a computer.

[0204] The method, system, electronic device, and storage medium for tracing the source of surface water multi-point source pollution disclosed in the embodiments of the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A method for tracing the source of multi-point source pollution of surface water, characterized in that: include: When it is detected that the traceability trigger condition is met, the water quality time series data and environmental parameter information of the corresponding monitoring point within the set time range are obtained, and the water quality time series data and environmental parameter information are preprocessed; Extracting features from the preprocessed water quality time series data to obtain key water quality features, wherein the key water quality features include water quality statistical features, water quality time series features, and water quality frequency domain features; Inputting the key water quality features into a pre-built deep learning model for identification to obtain corresponding pollution identification results, wherein the deep learning model includes a time series processing module, and the time series processing module is used to process the water quality time series features in the key water quality features; The corresponding pollution identification result is output, and the pollution identification result includes pollution type information, pollution source location information and confidence information.

2. The method for tracing the source of multi-point source pollution of surface water according to claim 1, characterized in that: After extracting the features of the pre-processed water quality time series data to obtain the key features of water quality, the method further includes: The pollution event features in the knowledge base pre-constructed according to the water quality statistical features, water quality time series features and water quality frequency domain features are matched to obtain a model parameter set associated with the current pollution event, and the model parameters in the deep learning model are updated according to the model parameter set; wherein the knowledge base includes a pollutant label library and a model parameter library, the pollutant label library includes historical pollution type labels and pollution feature descriptions; the model parameter library includes optimal model parameters corresponding to different pollution types.

3. The method for tracing the source of multi-point source pollution of surface water according to claim 2, characterized in that: The pollutant label library is determined by the following steps: Acquire a historical pollution data set, the historical pollution data set including historical pollution events, historical water quality time series data and historical environmental parameters, and preprocess the historical pollution data set; Perform feature extraction on the pre-processed historical pollution data to obtain historical key features, and use the PCA algorithm or the t-SNA algorithm to perform dimensionality reduction processing on the historical key features to obtain dimensionality reduction key features whose contribution rate exceeds the set contribution rate; Dividing the historical pollution data set into a plurality of clusters and using the dimension reduction key features as input; Randomly select a data point in the historical contaminated data set as the first cluster center, and select subsequent cluster centers according to the maximum distance principle; For each data point, calculate its distance from each cluster center and assign it to the cluster represented by the cluster center closest to it; For each cluster, calculate the mean of all data points in the cluster and use it as the new cluster center; The clustering assignment and cluster center of the data points are continuously updated until the change of the cluster center is less than a certain threshold or the preset number of iterations is reached, and then the clustering is completed; After clustering is completed, statistics are calculated for each cluster to determine the cluster characteristics, and the cluster characteristics are defined according to the set label rules to determine the historical pollution type labels and pollution feature descriptions, and saved to the pollutant label library.

4. The method for tracing the source of multi-point source pollution of surface water according to claim 2, characterized in that: The deep learning model is constructed through the following steps: Obtain a historical pollution training set, and construct a three-dimensional input tensor according to the historical pollution training set, wherein the three-dimensional input tensor includes the number of samples, the time step, and the feature dimension; Input the three-dimensional input tensor into a pre-built initial deep learning model for training until the set training requirements are met, wherein the initial deep learning model includes an input layer, a first LSTM layer, a Dropout layer, a second LSTM layer, a Dense layer, and an output layer connected in sequence; if the model output is a classification task, a cross entropy loss function is used as the loss function, and if the model output is a regression task, a mean square error loss function is selected as the loss function; The hyperparameter combination with the best performance is saved as the optimal parameters of the model to obtain a deep learning model.

5. The method for tracing the source of multi-point source pollution of surface water according to claim 4, characterized in that: The obtaining of the historical pollution training set and constructing a three-dimensional input tensor according to the historical pollution training set includes: Obtain a historical pollution training set, and perform SMOTE oversampling processing on data in the historical pollution training set whose corresponding pollution type data is less than a set value, wherein the historical pollution training set includes training pollution events, training water quality time series data, and training environmental parameters; Preprocessing the historical pollution training set, wherein the preprocessing includes missing value processing, spatial interpolation and noise filtering; Perform feature extraction on the pre-processed historical pollution training to obtain key training features, and use the PCA algorithm or the t-SNA algorithm to perform dimensionality reduction processing on the key training features to obtain dimensionality reduction training features whose contribution rate exceeds the set contribution rate; Constructing a three-dimensional input tensor based on the dimension reduction training features; And / or, after inputting the three-dimensional input tensor into a pre-built initial deep learning model for training until set training requirements are met, the method further includes: For each sample, the SHAP algorithm traverses all possible feature combinations to calculate the impact of each feature added to the feature combination on the model output; wherein the SHAP algorithm includes a Monte Carlo simulation algorithm or a Kernel SHAP algorithm; Calculate the Shapley value of each feature based on its impact on the model output; Sort the Shapley values ​​of all features by absolute value to obtain the global feature importance, and output the global feature importance; And / or, satisfying the tracing trigger condition includes: obtaining water quality indicators of each station in the monitoring area. If the water quality indicators obtained at the corresponding station in the monitoring area exceed the set water quality parameters, the tracing trigger condition is satisfied.

6. The method for tracing the source of multi-point source pollution of surface water according to claim 1, characterized in that: After outputting the corresponding pollution identification result, the method further includes: If the confidence information is less than the set confidence parameter, execute the next step; Acquire pollution monitoring data of corresponding monitoring points within a set time range, wherein the pollution monitoring data includes water quality time series data and environmental parameter information; Constructing a water quality evolution model, wherein the water quality evolution model includes an analytical solution to a permanganate index pollution control equation, an analytical solution to an ammonia nitrogen pollution control equation, an analytical solution to a total phosphorus pollution control equation, and an analytical solution to a dissolved oxygen pollution control equation; Determine an n×3 matrix as the population, where n is the population size. Each population individual contains the emission location, total emission amount and emission time. According to the pre-set parameter value range, randomly generate values ​​and fill them into the population matrix to form an initial population. Define the fitness function to calculate the relative error sum of the theoretical concentration and the actual monitoring concentration of each individual at all monitoring points and monitoring times. The fitness function is: Among them, N s is the number of pollution sources, N r is the number of monitoring points, N t is the number of monitoring times, i is the population individual, j is the monitoring time, is the theoretical pollutant concentration value generated by pollution source k at time j and monitoring point i. The pollution discharge characteristic of pollution source k is (M j ,X j ,T j ); It is the monitoring concentration information at the i monitoring point at the j time; The fitness function result is calculated, and the individual with the smaller relative error square sum is determined as the better solution and retained. The above screening, evaluation and optimization process are repeated until the preset number of iterations is reached or the fitness function result meets the stopping condition; The optimal parameter value corresponding to the optimal individual is output, and the optimal parameter value is the inversion result of the pollution source, and the optimal parameter value includes the emission location, the total emission amount and the emission time.

7. The method for tracing the source of multi-point source pollution of surface water according to claim 6, characterized in that: The analytical solution of the permanganate index pollution control equation includes: in, is the mass of pollutants released at the ith instant, X i is the i-th pollution location, T i is the initial time of emission of the i-th pollution source, A is the average cross-sectional area of ​​the target river section, D x is the longitudinal diffusion coefficient of pollutants, u is the average flow velocity of the target river section, is the pollutant attenuation coefficient, K DO is the half-saturation constant of dissolved oxygen, C DO ( t ) is C at time t DO concentration, is at position x and time t concentration; The analytical solution of the ammonia nitrogen pollution control equation includes: in, is the mass of pollutants released at the ith instant, X i is the i-th pollution location, T i is the initial time of emission of the i-th pollution source, A is the average cross-sectional area of ​​the target river section, D x is the longitudinal diffusion coefficient of pollutants, u is the average flow velocity of the target river section, K1 is the rate constant of ammonia nitrogen oxidation to nitrite, K2 is the rate constant of nitrite oxidation to nitrate, and τ is an integral variable, which indicates the initial emission time T from the pollution source. i To a certain time point between the current time t, it is used to calculate within this time period, by Transformed The total amount; τ1, τ2 are integral variables, τ1 represents the initial emission time T from the pollution source i to a certain time point between T and the current time t, and τ2 represents the time period from T to the current time t. i to another time point between τ and τ1. These two variables are used together to calculate First convert to Then by Convert to Total amount; The analytical solution of the total phosphorus pollution control equation includes: in, is the mass of pollutants released at the ith instant, X i is the i-th pollution location, T i is the initial emission moment of the i-th pollution source, Q k (τ) is the emission rate of the kth external phosphorus source, X ext,k is the emission position of the kth external phosphorus source, τ is the initial emission time of the kth external phosphorus source, and the integral interval [τ start,k , τ end,k ] represents the discharge period of the kth external phosphorus source, A is the average cross-sectional area of ​​the target river section, D x is the longitudinal diffusion coefficient of pollutants, u is the average flow velocity of the target river section, K TP is the decay rate constant of phosphorus, K r is the sediment release rate constant of phosphorus, K ext,k is the decay rate constant of the kth external phosphorus source; The analytical solution of the dissolved oxygen pollution control equation includes: Among them, M i DO is the mass of pollutants released at the ith instant, X i is the i-th pollution location, T i is the initial time of emission of the i-th pollution source, A is the average cross-sectional area of ​​the target river section, D x is the longitudinal diffusion coefficient of pollutants, u is the average flow velocity of the target river section, K p is the photosynthesis rate constant, K re is the respiration rate constant, and Chla is the chlorophyll a concentration.

8. The method for tracing the source of multi-point source pollution of surface water according to claim 1, characterized in that: The method for tracing the source of pollution also includes: Obtain the current status of the monitoring device, which includes device location information, environmental parameter information, and current concentration information; According to the current state of the acquired monitoring equipment, the pollution concentration at the current location is calculated using the pollution diffusion model, the equipment location and concentration data are spliced ​​into a state vector, and the state vector is input into the reinforcement learning model for identification to obtain the updated equipment location and updated detection time; The distance between the new device location and the pollution source and the energy consumption of the action are determined according to the updated device location and the updated detection time, and the reward value is calculated according to the reward formula; if the reward satisfies the termination condition, the corresponding dynamic pollution source information is output.

9. A system for tracing the source of multi-point source pollution of surface water, characterized in that: include: Acquisition module: used to acquire the water quality time series data and environmental parameter information of the corresponding monitoring point within the set time range when it is detected that the traceability trigger condition is met, and pre-process the water quality time series data and environmental parameter information; Feature extraction module: used to extract features from the pre-processed water quality time series data to obtain key water quality features, which include water quality statistical features, water quality time series features and water quality frequency domain features; Identification module: used for inputting the water quality key features into a pre-built deep learning model for identification to obtain corresponding pollution identification results, wherein the deep learning model includes a time series processing module, and the time series processing module is used for processing the water quality time series features in the water quality key features; Output module: used to output corresponding pollution identification results, which include pollution type information, pollution source location information and confidence information.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program enables a computer to execute the method for tracing the source of multi-point source pollution of surface water as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Atmospheric pollution monitoring method, device, equipment, storage medium and product

    CN119086818A

  • Water quality pollution source reverse tracking method based on LSTM model and pollution scene database

    CN119167034A

Cited By

  • Atmospheric pollution analysis method and device based on big data portrait, equipment and medium

    CN120832539A

  • Atmospheric pollution analysis method and device based on big data portrait, equipment and medium

    CN120832539B

  • Sewage treatment water quality intelligent prediction method and system based on discharge characteristics

    CN121030241A

  • Bioloading kinetic analysis method for sludge recycling filter material

    CN121210919A

  • A method for analyzing the bioload kinetics of sludge resource utilization filter media

    CN121210919B