A method and system for watershed water environment monitoring and emergency pollution rapid tracing based on water quality pollution detection

By constructing a multilayer perceptron (MLP) and LSTM/GRU model, combined with distributed sampling points and sampling stations, collaborative monitoring and rapid source tracing of multi-source data in the watershed water environment were realized. This solved the problems of insufficient data fusion and model robustness in existing technologies, improved emergency response and evidentiary validity, and met the needs of environmental law enforcement.

CN120744675BActive Publication Date: 2025-12-16SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511133893.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-12-16
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

Existing technologies are insufficient for rapid and accurate multi-source data fusion and pollution source tracing in watershed environments. In particular, they are slow to respond to emergency pollution incidents, and their models lack robustness and evidentiary value, making it difficult to meet the needs of environmental law enforcement.

Method used

Employing a multilayer perceptron (MLP) model and an LSTM/GRU model, combined with distributed sampling points and sampling stations, we collect and fuse real-time and time-division multi-source heterogeneous data. Through batch normalization, Dropout, and SMOTE oversampling techniques, we achieve multi-parameter collaborative monitoring and automated source tracing. Combined with validation markers and cross-validation mechanisms, we improve the model's classification and source tracing capabilities.

Benefits of technology

It enables rapid source tracing of large-scale water pollution in the watershed, improves emergency response efficiency and evidentiary value, directly supports environmental law enforcement, reduces costs, and improves the accuracy and robustness of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744675B_ABST
    Figure CN120744675B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of environmental water quality pollution monitoring, and particularly relates to a watershed water environment monitoring and emergency pollution rapid tracing method and system based on water quality pollution detection, which comprises the following steps: constructing a multi-parameter collaborative monitoring model and training and deploying; collecting real-time data and time-sharing data of water quality pollution detection of watershed water environment through distributed sampling points and sampling stations; using MLP and LSTM / GRU models to fuse EEM spectrum data and environmental characteristics, and to fuse and analyze multi-source heterogeneous data; automatically identifying whether an emergency pollution event occurs according to preset conditions, and automatically outputting the pollution type and tracing result of water environment monitoring and emergency pollution, which is further checked by artificial for environmental law enforcement check. The present application can complete the monitoring of watershed water environment pollution condition and the rapid tracking and tracing of emergency pollution on a large scale, at low cost and with high efficiency, so as to support environmental law enforcement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of environmental water quality pollution monitoring, and particularly relates to a watershed water environment monitoring and emergency pollution rapid tracing method and system based on water quality pollution detection, multi-source data and artificial intelligence model fusion analysis. BACKGROUND

[0002] Watershed water pollution monitoring and tracing, especially real-time monitoring and rapid and accurate tracing in emergency water pollution events, is one of the core problems in environmental regulation and law enforcement activities, and is of great significance to the development of effective watershed water pollution control strategies and the protection of ecological safety and human health. Watershed water environment monitoring and emergency pollution rapid tracing technology is a key means to address water environmental pollution problems, and its importance is reflected in multiple dimensions such as ecological protection, public safety and economic development. Current environmental monitoring and law enforcement activities have higher requirements for watershed water environment monitoring and emergency pollution rapid tracing technology, which needs to consider technical feasibility, legal compliance, emergency response efficiency and evidence effectiveness, etc. However, the existing technology generally cannot efficiently meet these requirements.

[0003] With the intensification of industrialization, urbanization and agricultural activities, pollutants can enter water bodies through various pathways, leading to water quality deterioration and threatening ecosystems and human health. Traditional watershed water pollution monitoring and emergency pollution tracing methods, such as isotope tracing and chemical fingerprinting, while providing accurate results in specific scenarios, often require complex sample processing procedures, costly instruments and long time, limiting their application in large-scale watershed water environment and real-time, emergency pollution monitoring and law enforcement. In addition, existing methods have limited comprehensive analysis capabilities for multi-source pollutants, making it difficult to cope with the complexity of modern water pollution. Real-time sensors can quickly obtain basic physicochemical data of flowing water (such as pH, conductivity, DO, turbidity, flow rate, etc.) and concentrations of some pollutants (such as residual chlorine, heavy metal pollutant concentrations), which have practical value for pollution existence judgment and large category classification. In the emergency monitoring of sudden leakage events, when the indicators are abnormal, the alarm can be triggered through pH, conductivity, heavy metal concentration, etc., which helps to quickly locate the pollution area and preliminarily predict the diffusion range combined with flow rate data; In drinking water source early warning monitoring, DO, residual chlorine, turbidity, etc. sensors are used to monitor water sources in real time, and when the indicators are abnormal, an alarm can be triggered (such as algae outbreak leading to DO fluctuation). However, these sensors have limited ability to identify specific pollutant components, low-concentration pollutants and new pollutants, and the data and preliminary analysis results obtained independently are often insufficient to directly serve as effective evidence for environmental law enforcement, and still need to be verified by laboratory analysis and watershed monitoring network.

[0004] In recent years, spectral analysis techniques, especially fluorescence excitation-emission matrix (EEM) spectroscopy, have received widespread attention in water quality monitoring due to their non-destructive, rapid, and high sensitivity characteristics. EEM spectroscopy can characterize the "fingerprint" characteristics of organic pollutants in water samples, providing a new technical means for pollution tracing. However, the multi-dimensionality and complexity of EEM data make traditional statistical methods or simple machine learning methods inadequate in feature extraction and source classification, making it difficult to fully utilize the information contained therein. At the same time, environmental characteristics (such as geographic location, type of industrial activity, type of river, etc.) as important reference information for pollution sources have high discriminability, but existing technologies rarely effectively integrate them with multi-source heterogeneous data such as spectral data and other data (such as real-time sensor data) to improve the accuracy and reliability of the tracing, avoid disputes over responsibility identification due to single technology errors, and make the analysis process and results more legally binding.

[0005] Specifically, the existing water environment monitoring and pollution tracing methods have the following shortcomings, which are difficult to meet the needs of environmental law enforcement:

[0006] 1) Single data type dependence and insufficient data fusion: Most methods focus on spectral data or chemical parameters, and rarely integrate real-time and time-series data, environmental characteristics and spectral characteristics, and other multi-source heterogeneous data, resulting in insufficient information utilization and affecting comprehensive understanding and accurate tracing of complex pollution. For example, CN 118430684 A discloses a method for predicting the concentration of ICM in surface water, which only predicts specific pollutants and relies on traditional chemical analysis, and fails to achieve deep integration of multi-source heterogeneous data and rapid identification of a wide range of pollutants.

[0007] 2) Model architecture optimization and coordination deficiency: Traditional machine learning models (such as standard multilayer perceptron) have limited modeling capabilities when dealing with high-dimensional complex data, making it difficult to capture deep relationships between features and failing to adequately address the multi-dimensionality and statistical methods of EEM data processing. Existing neural network applications, such as the method disclosed in CN 118425458 A, emphasize performance optimization, but the model architecture is not fully optimized for specific challenges in water environment monitoring and tracing (such as real-time, multi-source data fusion, and class imbalance), and the collaborative working mechanism of multiple AI models is not explicitly stated.

[0008] 3) Class imbalance problem and poor robustness: There is often a class distribution imbalance phenomenon in water sample data (such as fewer wastewater samples), which affects the classification performance and tracing ability of the model for minority class pollution. Existing methods lack effective technical means (such as SMOTE oversampling technology) to address this issue, resulting in poor robustness of the model in identifying and tracing minority classes.

[0009] 4) Emergency response and lack of evidence support capability: existing methods often respond slowly when facing sudden emergency pollution events, and the tracing process relies on manual analysis, with low automation and intelligence. At the same time, the data collection and analysis process lacks necessary verification markers and cross-validation mechanisms, resulting in insufficient evidence effectiveness and difficulty in directly supporting environmental law enforcement. SUMMARY

[0010] In view of the above problems of the prior art, the purpose of the present application is to provide a watershed water environment monitoring and emergency pollution rapid tracing method and system based on water quality pollution detection, which combines large watershed distributed active sampling and passive sampling, real-time sampling and time-sharing sampling, software and hardware, and performs water quality pollution detection, objective data collection and multi-parameter collaborative analysis, combined with scalable models, large-scale multi-source data processing and automatic alarm mechanism, to solve the contradiction between large watershed area, large-scale sampling, multi-source data processing and timeliness and accuracy in water quality pollution detection, and to meet the needs of emergency pollution rapid tracing and routine monitoring in water quality pollution detection, and to meet the requirements of technical feasibility, legal compliance, emergency response efficiency and evidence effectiveness in environmental monitoring and law enforcement activities.

[0011] The technical solution provided by the present application to solve the above problems is:

[0012] A watershed water environment monitoring and emergency pollution rapid tracing method based on water quality pollution detection, which is based on water quality pollution detection, multi-source data collection and artificial intelligence model fusion analysis, for watershed water environment monitoring and emergency pollution rapid tracing, comprising the following steps: first, constructing a multi-parameter collaborative monitoring model, including a multi-layer perceptron MLP model and an LSTM / GRU model, and training and deploying; then collecting real-time data and time-sharing data of verification markers of watershed water environment and water quality pollution detection through distributed sampling points and sampling stations, using AI-driven pollution identification MLP model and LSTM / GRU model, combining batch normalization, Dropout and SMOTE oversampling technology, and fusing EEM spectral data and environmental characteristics to fuse and analyze multi-source heterogeneous data; then automatically identifying whether an emergency pollution event occurs according to the preset conditions, and if an emergency pollution event is preliminarily judged to occur, starting pollution type judgment, tracing, and cross-validation, and automatically outputting the pollution type and tracing result of water environment monitoring and emergency pollution, which is further checked by artificial, environmental law enforcement checking, completing watershed water environment pollution monitoring and emergency pollution rapid tracing, which specifically includes the following steps:

[0013] S1, constructing a multi-parameter collaborative monitoring model

[0014] A multi-parameter collaborative monitoring model for real-time monitoring of water pollution and rapid pollution tracing in emergency pollution in the entire river basin water environment area is constructed, specifically a multi-layer perceptron MLP model and an LSTM / GRU model are constructed, and environmental characteristics, pollutants and pollution source information of key node grids in the entire river basin water environment area are imported; a known historical data set is used to train the constructed multi-layer perceptron MLP model, and the model parameters are adjusted to optimize the model parameters to improve the classification and tracing performance, the training specifically uses a weighted cross-entropy loss function (the weight is calculated by the reciprocal of the number of samples in each class), uses an Adam optimizer, sets the learning rate scheduler to Lambda LR or Cosine Annealing LR, and introduces an early stopping mechanism (stops training according to the validation loss); the trained multi-layer perceptron MLP model and LSTM / GRU model are deployed to a network server;

[0015] S2, collecting water data

[0016] On-site sampling points and sampling stations are set up in the key node grids of the river basin, water quality pollution detection and water environment data are collected, and verification marks such as data collection equipment ID, GIS, time stamp, encryption information and tamper-proof identifiers are added to the data; the sensors of the on-site sampling points automatically collect water data in real time, obtain GIS, time, flow rate, color, transparency and other information of the water body (including turbidity, pH, conductivity, DO, residual chlorine, heavy metal pollutant concentration), obtain real-time data sets, and transmit the data sets to the network server after adding verification marks; the sampling instruments of the sampling stations regularly or irregularly sample (start sampling at any time in the emergency monitoring mode), detect the water samples, obtain data including the time, location, environmental characteristics and fluorescence excitation-emission matrix (EEM) spectral data of the water samples, obtain time-sharing data sets of water quality pollution detection, and transmit the data sets to the network server after adding verification marks;

[0017] S3, data feature extraction and preprocessing

[0018] A multi-element correlation dataset is constructed by analyzing and processing the real-time data group and the time-sharing data group of water quality pollution detection in the grid of each key node in the basin, and a multi-element correlation dataset with verification marks is constructed. The correlation parameters in the dataset include: device ID-time-place-pollutant type-pollution source-verification mark; then feature extraction and preprocessing are performed; data fusion and preprocessing are performed, and the EEM spectrum data is subjected to grid integration (FRI) processing, and then the FRI features and environmental feature levels are spliced. Data standardization: use sklearn preprocessing Label Encoder to convert the water sample source label into an integer label; feature importance analysis, identify the dominant features in the data through the feature importances attribute of the Random Forest Classifier.

[0019] S4, model running for abnormal monitoring

[0020] The LSTM / GRU model combined with the MLP model is used to classify and monitor the real-time data of the field sampling points in real time in the conventional monitoring mode; if an index anomaly is found, the emergency monitoring mode is entered, and the cross-validation and pollution tracing mechanism are started simultaneously, and the downstream sampling points and sampling stations are simultaneously sampled; the trained multi-layer perceptron MLP model is used for pollution classification and tracing, and the real-time data group and the time-sharing data group are compared and analyzed, cross-validated, and multiple parameter collaborative judgments are made to determine whether a water emergency pollution event has occurred; if a water emergency pollution event has occurred, the pollutants are further classified, and the pollution source and the predicted diffusion range information are preliminarily judged and output, which are further checked and processed by artificial, and the artificial checking result data is input into the network server;

[0021] S5, model updating and continuous monitoring

[0022] The verified multi-element correlation dataset is used to train the multi-layer perceptron model and the LSTM / GRU model again, and the new model is deployed to the network server for running, and steps S2-S4 are repeated to analyze and process the real-time data group and the time-sharing data group, and the monitoring results are output.

[0023] A basin water environment monitoring and rapid pollution tracing system for realizing the basin water environment monitoring and rapid pollution tracing method, comprising:

[0024] A network server, a management terminal, a plurality of distributed sampling points, and sampling stations connected to each other through a network;

[0025] The management terminal is used for input and output operations.

[0026] A plurality of distributed sampling points are respectively arranged at key nodes of the river basin to perform spot and timing sampling, collect real-time environmental characteristic data, and upload to the network server;

[0027] A plurality of distributed sampling stations are respectively arranged at key monitoring nodes of the river basin to collect EEM spectrum data and environmental characteristics at regular or irregular intervals, and upload to the network server;

[0028] The network server is used to receive monitoring data, analyze and process, and output the processing results. The network server has a built-in river basin water environment monitoring and emergency pollution rapid tracing program, which includes the following modules:

[0029] The monitoring station management module is used to connect and manage the distributed field sampling points and sampling stations, and receive real-time data groups returned by the field sampling points and time-sharing data groups returned by the sampling stations.

[0030] The data processing module is used to process real-time data groups and time-sharing data groups, construct a multivariate correlation data set, and divide the known multivariate correlation data set for model training. The data processing module includes calculating grid region integral FRI of spectrum data through the calculate grid integrals function, fusing multi-dimensional features, performing standardization and SMOTE oversampling processing.

[0031] The classification and anomaly monitoring module: through the trained multilayer perceptron MLP model, combined with the AI-driven pollution identification LSTM / GRU model, the real-time data of the field sampling points are classified and index anomaly monitored in real time in the conventional monitoring mode. If an index anomaly is found, the emergency monitoring mode is entered, the cross-validation and pollution tracing mechanism are started synchronously, the downstream sampling points and sampling stations are synchronously sampled, and the real-time data groups and time-sharing data groups are compared and analyzed, cross-validated to determine whether a water body emergency pollution event occurs.

[0032] The tracing and prediction module: used for further processing of data preliminarily determined by the model as a water body emergency pollution event, classifying and preliminarily determining and outputting the pollution source and predicting the diffusion range information.

[0033] The notification and visualization display module: used to output the monitoring data analysis results, including conventional monitoring information, emergency pollution event alarm information, pollution classification and positioning, pollution source, and predicted diffusion range information.

[0034] The pollution monitoring law enforcement management module: law enforcement personnel manually check and process according to the output monitoring data analysis results, and input the manual check results data into the network server.

[0035] The watershed water environment monitoring and emergency pollution rapid tracing method and system provided by the present application have the beneficial effects of at least including the following compared with the prior art:

[0036] 1. The present application focuses on large-area and large-scale water quality pollution monitoring and rapid tracing of emergency pollution in watershed water environment. Real-time and time-sharing water quality pollution data of distributed sampling points and sampling stations are collected and verified with markers (including non-tamperable identifiers), breaking through the complexity and high cost limitations of traditional tracing methods. Through the collaborative analysis of multi-source heterogeneous data fusion (EEM spectrum, environmental characteristics, real-time, time-sharing) and AI-driven identification models (MLP, LSTM / GRU), combined with batch normalization, Dropout, and SMOTE oversampling technology to process data and train models, the data collection, analysis process, and output results (including verification markers, manual verification result input, and log records) can all be directly used as effective evidence for environmental monitoring and law enforcement, significantly improving the evidence effectiveness.

[0037] 2. The present application fuses EEM spectrum data and environmental characteristics, combines real-time and time-sharing data collection of water quality pollution, multi-parameter collaborative monitoring model (MLP), and AI-driven pollution identification algorithm model LSTM / GRU, designs a multi-layer perceptron MLP model (PyTorch framework, total parameter quantity not less than 1.5M, including Batch Norm, Dropout 0.3 optimization) for water quality pollution classification, and uses SMOTE oversampling (k neighbors=3) and other technologies to handle class imbalance problems, achieving high-precision classification (5 categories) of water sample sources and water quality pollutants and pollution tracing. The present application combines large watershed distributed active sampling and passive sampling for water quality pollution detection, real-time sampling and time-sharing sampling, and constructs an extensible model and large-scale multi-source data processing framework, which can meet the needs of emergency pollution rapid tracing and routine monitoring of water quality pollution, and realize the consideration of both normal and emergency situations. The system runs smoothly and reliably, reduces costs through optimization technology and automated processes, is easy to promote, and provides an efficient and accurate tool for environmental protection and sustainable development.

[0038] 3、The AI model and other machine learning technologies adopted by the application provide a new idea for processing EEM data and environmental characteristics. Through the construction of a water pollution classification model, the source of water samples and pollution types can be automatically identified, thereby achieving efficient pollution tracing. Unlike early water pollution monitoring technologies based on statistical models and simple machine learning algorithms, the water sample source classification and pollution tracing method based on water quality pollution detection data and improved multilayer perceptron proposed by the application can, on the one hand, fuse fluorescence excitation-emission matrix (EEM) spectral data and environmental characteristics to construct a multi-source data fusion framework (generate a 225-dimensional feature vector, 56-dimensional environment, and 169-dimensional FRI), realize efficient processing and accurate analysis of large-scale, multi-source, and high-dimensional water quality detection data, and overcome the limitations of traditional methods in data processing capability. At the same time, the generalization ability and adaptability, timeliness of the model can be significantly improved. Specifically, the application improves the monitoring and tracing ability through a combination of real-time and time-sharing data collection. Real-time data collection realizes the immediate monitoring of sudden pollution events through fast-response sensors, and time-sharing data collection provides historical data support, laying a foundation for model training and long-term trend analysis. At the same time, the MLP is constructed through a multi-parameter collaborative monitoring model, which fuses EEM spectral data, environmental characteristics, and other water quality parameters such as pH, conductivity, dissolved oxygen (DO), residual chlorine, turbidity, heavy metals, etc., to comprehensively reflect the water pollution situation and improve the accuracy of tracing. In addition, the application adopts two types of pollution identification algorithm models LSTM / GRU and MLP driven by AI, combined with batch normalization, Dropout, and SMOTE oversampling (kneighbors=3) technology, through machine learning to fuse and analyze multi-source heterogeneous data, automatically identify pollution types and improve classification accuracy, which can accurately distinguish five types of water sample sources, including fish pond aquaculture water, groundwater, surface water, wastewater, and simulated emergency pollution wastewater. Especially in the case of uneven sample distribution, it still maintains excellent performance. This method does not require tedious sample pretreatment (compared to traditional chemical analysis), combined with the rapid computing power of machine learning, can realize the rapid classification, tracking, and tracing of water quality pollution status at low cost and large scale.

[0039] 4、The application adopts a variety of mutually coordinated water quality pollution monitoring and classification, intelligent model and algorithm of tracing, combines water body index, spatial characteristics, spectral characteristics and other multi-source data as the input of the prediction tracing model, uses a neural network model (especially an improved MLP, LSTM / GRU), realizes high-precision positioning of the pollution source and accurate identification of the pollution type, adds a dynamic correction factor (similar effects are realized by using an Adam optimizer and Lambda LR or CosineAnnealing LR learning rate scheduling strategy) in the neural network model, so that the model can adjust the learning rate in real time according to the prediction performance, combines an early stopping mechanism, makes the model more flexible in the training process, can quickly adapt to data changes and noise, improves the accuracy and robustness in water quality pollution classification and tracing, and further improves the efficiency and accuracy of water quality pollution monitoring.

[0040] 5、The application has significant advantages in application scenarios, can quickly locate the pollution area by combining pH, conductivity, heavy metal sensors in water quality pollution emergency monitoring such as sudden leakage events, and can predict the diffusion range by combining flow rate data; in drinking water source early warning monitoring, DO, residual chlorine, turbidity sensors can be used for real-time monitoring of water sources, and an alarm is triggered when the index is abnormal. For technical difficulties such as data fusion and intelligent analysis challenges, the limitations of single indicators in comprehensively reflecting pollution conditions are overcome.

[0041] 6、The application improves the accuracy of pollution type identification by integrating multi-parameter data through machine learning algorithm, and proposes a solution of constructing a large-scale multi-element correlation historical database with verification labels and using model updating and continuous monitoring to solve the problem of requiring a large amount of historical data for model training, so that routine and emergency can be considered, and a set of system and method can be used to consider routine monitoring and emergency pollution rapid tracing of water quality pollution, which significantly improves real-time performance and operability, and reflects careful consideration of practical application scenarios.

[0042] 7、The real-time and time-sharing data collection of the water quality pollution detection, combined with the fast response sensor (pH, conductivity, heavy metal, flow rate, etc.) with verification labels and time-sharing data (including EEM, environmental characteristics) with verification labels, can realize immediate monitoring and long-term trend analysis of sudden pollution events. In emergency monitoring, the downstream sampling points and sampling stations are simultaneously started to encrypt sampling, multi-element data of water quality pollution detection is obtained, the diffusion range of pollution is predicted by flow rate data, and cross-validation of real-time and time-sharing data is performed to enhance the accuracy of the monitoring and analysis results.

[0043] 8、The present application can comprehensively reflect the pollution status of water body through multi-parameter collaborative monitoring model construction (MLP, LSTM / GRU), fusion of 225-dimensional features (including 169-dimensional EEM spectrum FRI features and 56-dimensional environmental features), and other water quality parameters such as DO, residual chlorine, turbidity, etc. In the early warning monitoring of drinking water sources, when the index is abnormal (judged by AI-driven LSTM / GRU and improved MLP model), the alarm is triggered, which can avoid misjudgment and false alarm.

[0044] 9、The AI-driven pollution identification algorithm model constructed by the present application includes an improved pollutant classification MLP model (including Batch Norm, Dropout 0.3, total parameters not less than 1.5M, 5-class classification), and an LSTM / GRU model combined with SMOTE oversampling (k=3), batch normalization and Dropout technology, which improves the classification accuracy and traceability of pollutants. Through machine learning to integrate multi-parameter data, the limitations of single index are overcome, abnormal judgment is made through learning of unpolluted water body data, and when the classification result deviates from the preset threshold, an abnormal alarm is triggered, which can timely find and correct the problems of the model itself, and improve its running stability.

[0045] 10、The present application provides an automatic process from data loading with verification mark (including non-tamperable identifier), preprocessing (FRI calculation, fusion, Standard Scaler standardization, SMOTE oversampling), model training (Adam optimizer, LR scheduling strategy, early stopping mechanism), evaluation (classification report, confusion matrix, Random ForestClassifier feature importance analysis) to result output (visual display), notification (alarm sending), record (log file), manual data checking and data entry, model updating, full-process automation or semi-automation, reducing manual intervention; The log file records the classification report, confusion matrix and feature importance analysis, providing data support for event management.

[0046] 11、The application has wide application prospects, the method and the system provide a new tool of high efficiency, accuracy, robustness and supporting law enforcement for pollution tracing and water quality monitoring, and have important significance for watershed water pollution detection, environmental protection and sustainable development. When water pollution leakage events occur, the application can quickly locate the pollution area through multi-parameter sensor data and flow rate data, combine the predicted diffusion range information, guide the on-site evidence collection and emergency treatment. For early warning monitoring of drinking water sources, DO, residual chlorine and turbidity sensors can be used to monitor water sources in real time, and an alarm mechanism is triggered when an anomaly occurs; for monitoring and tracing of industrial wastewater, EEM spectrum data and environmental characteristics can be integrated to quickly identify the source and type of wastewater, and the pollution source can be preliminarily judged. At the same time, the visual display module of the system, the information notification (prompting law enforcement personnel to collect evidence on site, sending an alarm), and the pollution monitoring law enforcement management module (manual verification result input) enable it to be directly used as a basis for environmental monitoring law enforcement. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 is a module composition structure diagram of a watershed water environment monitoring and emergency pollution rapid tracing system based on water pollution detection according to an embodiment of the application;

[0048] Figure 2 is a confusion matrix diagram of a comparison of classification effects of different learning rate scheduling strategies in the embodiment of the application, that is, a comparison diagram of model predicted categories and actual categories;

[0049] Figure 3 is a training loss and validation loss curve diagram of each round of model training of the Cosine Annealing LR learning rate scheduling strategy in the embodiment of the application.

[0050] Figure 4 is a training accuracy and validation accuracy curve diagram of each round of model training of the Cosine Annealing LR learning rate scheduling strategy in the embodiment of the application.

[0051] Figure 5 is a learning rate change curve diagram of each round of model training of the Cosine Annealing LR learning rate scheduling strategy in the embodiment of the application.

[0052] Figure 6 is a training loss and validation loss curve diagram of each round of model training of the Lambda LR learning rate scheduling strategy in the embodiment of the application.

[0053] Figure 7 is a training accuracy and validation accuracy curve diagram of each round of model training of the Lambda LR learning rate scheduling strategy in the embodiment of the application.

[0054] Figure 8is a learning rate change curve diagram of each round of model training of different learning rate scheduling strategies of Lambda LR in the embodiment of the application. DETAILED DESCRIPTION

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the application or the prior art. It should be understood that the drawings in the following introduction are only for the convenience of clearly expressing part of the embodiments of the technical solutions of the application, and other drawings can be obtained by those skilled in the art without creative labor on the premise that there is no need to pay creative labor.

[0056] The experimental supplies, instruments and technologies used in the following embodiments of the application are all commercially available and conventional technologies unless otherwise specified.

[0057] Embodiment 1

[0058] Referring to Figure 1 The improved watershed water environment monitoring and emergency pollution rapid tracing method and system based on water quality pollution detection in the embodiments of the application are specifically a water quality pollution detection, monitoring and emergency pollution rapid tracing method and a matching information network management system for the characteristics of large area, wide range, complex water flow and water system network of watershed water environment. The method is based on multi-source data acquisition and artificial intelligence model fusion analysis, based on multi-source water quality pollution detection data, and performs watershed water quality pollution detection, water environment monitoring and rapid classification and tracing of emergency pollution, including the following steps: first, a multi-parameter collaborative monitoring model is constructed, including a multi-layer perceptron MLP model and an LSTM / GRU model, and is trained and deployed; then, through distributed sampling points and sampling stations, real-time data and time-sharing data of watershed water environment and water quality pollution detection with verification labels are collected, an AI-driven pollution identification MLP model, an LSTM / GRU model, batch normalization, Dropout and SMOTE oversampling technology are combined, EEM spectrum data and environmental characteristics are fused, and multi-source heterogeneous data are fused and analyzed; then, whether an emergency pollution event occurs is automatically identified according to a preset condition, if it is preliminarily judged that an emergency pollution event occurs, pollution type judgment, tracing and cross-validation are started, and the pollution type and tracing result of water environment monitoring and emergency pollution are automatically output, which are further checked by artificial, environmental law enforcement is checked, and the monitoring of watershed water environment pollution and the rapid tracing of emergency pollution are completed; the method of the embodiments of the application is suitable for various scenes such as daily routine monitoring, emergency monitoring and sudden leakage events, and drinking water source early warning monitoring of watershed water environment. The method specifically includes the following steps:

[0059] S1, constructing a multi-parameter collaborative monitoring model

[0060] Constructing an AI model for multi-parameter collaborative monitoring, training and deploying

[0061] Constructing an AI model for multi-parameter collaborative monitoring, training and deploying

[0062] Using known historical data sets, the constructed multi-layer perceptron MLP model and LSTM / GRU are collaboratively trained, and the model parameters are adjusted to optimize the model parameters to improve classification and tracing performance.

[0063] Deploying the trained multi-layer perceptron MLP model and LSTM / GRU model to a network server.

[0064] S2, collecting water data

[0065] Setting up field sampling points and sampling stations in the key node grid of the river basin, collecting water quality pollution detection and water environment data, and adding data collection device ID, GIS and timestamp, encryption information and tamper-proof identifier verification markers to the data.

[0066] The sensors of the field sampling points automatically collect water data in real time, obtain GIS, time, flow rate, color, transparency and other information of the water body, and get real-time data sets of water quality pollution detection, and then transmit them to the network server after adding verification markers.

[0067] The sampling instruments of the sampling stations regularly or irregularly sample and detect the water samples, obtain data including time, location, environmental characteristics and fluorescence excitation-emission matrix EEM spectral data of the water samples, and get time-sharing data sets of water quality pollution detection, and then transmit them to the network server after adding verification markers.

[0068] S3, data feature extraction and preprocessing

[0069] Constructing a multi-element correlation data set, analyzing and processing the real-time data sets and time-sharing data sets of water quality pollution detection in each key node grid in the river basin, and constructing a multi-element correlation data set with verification markers, the correlation parameters of which include device ID-time-location-pollutant type-pollution source-verification marker; then performing feature extraction and preprocessing.

[0070] S4, model running for abnormal monitoring

[0071] The trained multi-layer perceptron (MLP) model, combined with the LSTM / GRU model, adopts a conventional monitoring mode for real-time data of on-site sampling points, performs real-time classification and index anomaly monitoring; if an index anomaly is found, it switches to an emergency monitoring mode, simultaneously starts cross-validation and pollution tracing mechanisms, and simultaneously starts synchronous sampling of downstream sampling points and sampling stations, and then compares and analyzes real-time data groups and time-sharing data groups, performs cross-validation, determines whether a water body emergency pollution event has occurred, further classifies the pollutants, preliminarily determines and outputs the pollution source and the predicted diffusion range information, and hands over to manual further checking and processing, performs environmental law enforcement checking, and enters the manual checking result data into the network server;

[0072] The MLP model and the LSTM / GRU model are coordinated, a conventional monitoring mode is adopted for real-time data of on-site sampling points, real-time classification and index anomaly monitoring are performed, when partial index anomalies are preliminarily determined by either model, data comparison and cross-validation analysis are performed by the two models, for multiple same indexes, if both models determine that the indexes are abnormal, it is determined that a water body emergency pollution event has occurred, an alarm is triggered, the system further classifies the pollutants, preliminarily determines and outputs the pollution source and the predicted diffusion range information;

[0073] S5, model updating and continuous monitoring

[0074] A multi-element correlation data set verified by environmental law enforcement checking is used to retrain the multi-layer perceptron (MLP) model and the LSTM / GRU model, and the new models are deployed to the network server for operation, steps S2-S4 are repeated, real-time data groups and time-sharing data groups are analyzed and processed, and monitoring results are output.

[0075] A watershed water environment monitoring and rapid pollution tracing system for realizing the foregoing watershed water environment monitoring and rapid pollution tracing method, comprising:

[0076] A network server, a management terminal, a plurality of distributed sampling points, and sampling stations are connected to each other through a network;

[0077] The management terminal is used for input and output operations;

[0078] The plurality of distributed sampling points are respectively deployed at key nodes of the watershed, used to perform fixed-point and fixed-time sampling of water pollution, collect real-time environmental characteristic data, and upload the data to the network server;

[0079] The plurality of distributed sampling stations are respectively deployed at key monitoring nodes of the watershed, used to collect EEM spectrum data and environmental characteristics of water pollution at regular or irregular intervals, and upload the data to the network server;

[0080] The network server is used to receive monitoring data, analyze and output processing results, and has a built-in river basin water environment monitoring and emergency pollution rapid tracing program, which includes the following modules:

[0081] The monitoring station management module is used to connect and manage each distributed field sampling point and sampling station, and receive real-time data sets returned by the field sampling points and time-sharing data sets returned by the sampling stations;

[0082] The data processing module is used to process real-time data sets and time-sharing data sets, construct a multi-element correlation data set, and divide known multi-element correlation data sets for model training; which includes calculating grid region integral FRI through the calculate grid integrals function for spectral data, fusing multi-dimensional features, standardizing and SMOTE oversampling processing;

[0083] The classification and anomaly monitoring module: through the trained multi-layer perceptron MLP model, combined with the AI-driven pollution identification LSTM / GRU model (in series), the real-time data of the field sampling points are classified and index anomaly monitored in real time in the conventional monitoring mode; if the LSTM / GRU model finds index anomalies of water pollution, it will switch to the emergency monitoring mode, simultaneously start the cross-validation and pollution tracing mechanism, and simultaneously start the synchronous sampling of downstream sampling points and sampling stations, and then compare and analyze the real-time data sets and time-sharing data sets, cross-verify, and determine whether a water emergency pollution event has occurred;

[0084] Among them, the LSTM / GRU model is responsible for modeling time series features: learning the normal fluctuation pattern of indicators over time (such as the periodic change of dissolved oxygen during the day and night), supporting trend-based classification (such as "dissolved oxygen has been decreasing for 3 hours" corresponding to "mild pollution to severe pollution transformation") and "dynamic time series anomaly" detection (such as the mutation amplitude of the indicator exceeding the historical trend range); the MLP model is responsible for modeling static features: learning the association between "indicator value combination" and water quality state (such as "COD>50mg / L and ammonia nitrogen>10mg / L" corresponding to "severe pollution"), supporting real-time classification and "instantaneous static anomaly" detection (such as a certain moment when the indicator combination has never appeared in the historical normal sample).

[0085] The tracing and prediction module is used to further process the data preliminarily judged by one of the models as a water emergency pollution event, classify and preliminarily judge and output the pollution source and the predicted diffusion range information;

[0086] The notification and visualization display module is used to output the monitoring data analysis results, including conventional monitoring information, emergency pollution event alarm information, pollutant classification and positioning, pollution source, and predicted diffusion range information;

[0087] Pollution monitoring law enforcement management module: law enforcement personnel manually check according to the output monitoring data analysis results, and manually check the result data into the network server.

[0088] The daily operation process of the river basin water environment monitoring and rapid tracing system for emergency pollution provided by the embodiment is as follows:

[0089] 1) Collecting river basin water environment monitoring data by multiple distributed sampling points and sampling stations and uploading;

[0090] 2) The monitoring site management module receives data and forwards them to the data processing module for processing;

[0091] 3) The data processing module performs feature extraction and preprocessing;

[0092] 4) The classification and anomaly monitoring module performs real-time classification and anomaly detection;

[0093] 5) If an anomaly is detected, the tracing and prediction module starts the pollution tracing mechanism, locates the pollution area, predicts the pollution source and diffusion range, and outputs the preliminary monitoring analysis results;

[0094] 6) The notification and visualization display module outputs the display content according to the monitoring analysis results, automatically initiates an emergency response alarm when an emergency pollution event occurs, and presents the monitoring, tracing, processing results and the whole process to the environmental administrative law enforcement personnel for manual checking;

[0095] 7) The pollution monitoring law enforcement management module records pollution monitoring law enforcement data, manual checking, law enforcement evidence and processing results, and supports the legal, reasonable and compliant environmental administrative law enforcement work;

[0096] 8) The system uses the multi-element correlation data set verified by law enforcement and newly obtained data set, trains the multilayer perceptron model again, and deploys the new model to the network server for running, realizing continuous updating of the model and optimization of the system performance.

[0097] It should be noted that the various real-time sensors used by the distributed sampling points can quickly obtain the basic physicochemical data and part of the pollutant concentration of flowing water, and have practical value for pollution existence judgment and large category classification (such as organic matter, heavy metal, nutrient salt pollution), but the identification ability for specific pollutant composition, low concentration pollution and new type of pollutants is limited, and the laboratory analysis of the sampling station and the river basin monitoring network need to be combined to realize accurate monitoring and tracing, and to avoid misjudgment and false reporting.

[0098] The embodiment of the present application adopts an AI identification model for water pollution based on the combination of sensor data and laboratory data. The LSTM / GRU model in the model mainly processes time series data (such as pH value, dissolved oxygen, heavy metal concentration, etc.) of water quality sensors to capture the dynamic changes of pollution indicators and is suitable for analyzing the evolution law of pollution over time (such as the periodic fluctuations of industrial wastewater discharge) and performing abnormal detection: by learning the normal water quality mode, sudden pollution events can be identified. However, there are problems such as sensor noise sensitivity, and missing or noise of sensor data may lead to misjudgment of the model, so cross-validation is performed by laboratory data of sampling stations to ensure the objectivity and accuracy of the data.

[0099] In the real-time monitoring of the water environment of a river basin, the MLP and LSTM / GRU models of the embodiment of the present application are used in series (time series feature extraction → static feature classification). The LSTM / GRU first extracts features from the time series data of real-time monitoring, and then inputs the time series features into the MLP to complete classification or abnormality judgment in combination with static features. It is suitable for scenarios where time series dynamics dominate but static attributes need to be combined for decision-making (such as joint judgment of pollution sources based on “time series abnormal patterns + sampling point location”). Real-time classification (such as water quality classification) and index abnormality monitoring are achieved by using MLP (multilayer perceptron) and LSTM / GRU model to process real-time data of sampling points. The core is to take advantage of the complementary advantages of the two types of models: MLP captures the correlation between static features and indicators, and LSTM / GRU captures the time series dynamic trend and long-term dependence. Through the collaborative logic of “feature separation - separate modeling - output fusion”, the monitoring accuracy is improved. The core of the cooperation between MLP and LSTM / GRU is the dual-dimensional modeling of “static snapshot + dynamic trend”: MLP captures the correlation between indicators to improve the accuracy of classification and static abnormality, and LSTM / GRU captures the time series dependence to improve the sensitivity of trend classification and dynamic abnormality. Through the framework of feature separation, parallel reasoning, and output fusion, efficient monitoring can be achieved under the condition of “the least number of sampling points” while ensuring real-time, reducing false positives and false negatives, and providing rapid decision support for river basin pollution emergency tracing.

[0100] Embodiment 2

[0101] Referring to Figure 1 The river basin water environment monitoring and rapid pollution emergency tracing method and system of the embodiment of the present application are a specific implementation based on embodiment 1. The focus is to improve the efficiency and coverage of monitoring and management through a systematic new solution, so that it can be applied to water pollution detection, real-time monitoring, and pollution emergency response in large areas (such as the Pearl River Basin) to reduce the overall cost, making it easy to implement and providing strong support for environmental law enforcement at all stages. The difference lies in that:

[0102] Monitoring site management module, for connecting and managing distributed field sampling points and sampling stations deployed at key nodes and key monitoring nodes in the river basin, receiving real-time data sets (including GIS, time, flow rate, color, turbidity, pH, conductivity, DO, residual chlorine, heavy metal pollutant concentration, etc.) marked with verification labels returned by the distributed field sampling points, and time-sharing data sets (including time, place, environmental characteristics such as latitude and longitude, industrial activity related parameters, river type, and fluorescence excitation-emission matrix (EEM) spectral data, excitation / emission range 250-500nm, step 5nm, generating a 52x52 matrix) marked with verification labels returned by the distributed sampling stations; the sampling station has the ability to periodically sample in the conventional monitoring mode and to start sampling at any time in the emergency monitoring mode.

[0103] Data processing module: for processing real-time data sets and time-sharing data sets, and constructing a multi-element correlation data set marked with verification labels containing device ID-time-place-pollutant type-pollution source-verification label and other associated parameters through analysis and processing. The known historical data set is divided into 8:1:1 layers for model training. For the EEM spectral data in the time-sharing data, the calculate grid integrals function is used to calculate the grid area integral (FRI), and then the FRI features (169 dimensions) and environmental features (56 dimensions) are horizontally spliced and fused into a 225-dimensional feature vector. Standard Scaler is used for feature standardization, and sklearn preprocessing Label Encoder is used to convert water sample source labels into integers. The minority class in the training set is subjected to SMOTE oversampling processing (k neighbors=3). Feature importance analysis is performed through the feature importances attribute of the Random Forest Classifier.

[0104] Classification and anomaly monitoring module: through the trained multilayer perceptron MLP model (PyTorch construction, total parameter quantity not less than 1.5M, architecture containing Batch Norm and Dropout 0.3, etc.), and LSTM / GRU model, the real-time data of the field sampling point is classified and index anomaly monitored in the conventional monitoring mode. The MLP model classifies the water sample into 5 categories (fish pond breeding water, groundwater, surface water, wastewater, and simulated emergency pollution wastewater). This module learns the data of unpolluted water bodies to identify normal patterns. The classification and anomaly monitoring module classifies pollutants by pre-learning the classification data of unpolluted water bodies through the LSTM / GRU model and the MLP model, and then jointly completes the classification through the classification function of the MLP model.

[0105] When the MLP model classifies water samples, if the classification result deviates from the preset threshold, including an increase in the probability of abnormal categories, an anomaly alarm is triggered. The multilayer perceptron MLP model and the LSTM / GRU model work together to perform real-time classification and anomaly monitoring of real-time data from on-site sampling points using a conventional monitoring mode. The MLP model classifies water samples into five categories (fishpond aquaculture water, groundwater, surface water, wastewater, and simulated emergency polluted wastewater). If real-time monitoring detects anomalies in indicators, or if the MLP classification result deviates from the preset threshold (increased probability of abnormal categories), an anomaly alarm is automatically triggered and the system switches to emergency monitoring mode. Simultaneously, cross-validation and pollution source tracing mechanisms are initiated, and downstream sampling points and stations are simultaneously intensified for sampling. The simulated emergency polluted wastewater is prepared by mixing wastewater containing emergency pollutants with river water at ratios of 1:1, 1:2...1:100, and allowing it to stand at 25°C for 2 hours before use.

[0106] The aforementioned source tracing and prediction module is used to further analyze data initially judged to be from a water body emergency pollution event. By integrating a pollution source tracing model based on environmental characteristics and three-dimensional fluorescence, it locates and classifies pollutants in water samples, and initially judges and outputs information such as pollution sources and predicted diffusion ranges. The results are then verified by the calculation and validation model. Upon detecting an anomaly, the system immediately activates the emergency response mode, using multi-parameter sensor data (such as heavy metals, pH, etc.) and flow velocity data to quickly locate the pollution source and predict the diffusion range, providing law enforcement personnel with verification objects, targets, and pollution data for manual verification.

[0107] The notification and visualization module outputs monitoring data analysis results, system status, and emergency information. This includes routine monitoring information, emergency pollution event alarm information, pollutant classification and location, pollution sources, and predicted diffusion range information. In the event of an emergency pollution event, the system automatically sends an alarm to relevant law enforcement personnel or departments, including information such as the location of the pollution source and the type of pollutant, prompting law enforcement personnel to conduct on-site source tracing and evidence collection based on the source tracing analysis results, and guiding emergency response. The visualization module displays content including environmental characteristics, FRI feature heatmaps, training history curves, confusion matrices, and FRI box plots. The notifications in the notification and visualization module include prompts for law enforcement personnel to conduct on-site source tracing and evidence collection based on the source tracing analysis results; in the event of an emergency pollution event, the system automatically sends an alarm to relevant law enforcement personnel or departments, including information such as the location of the pollution source and the type of pollutant, guiding on-site evidence collection and emergency response; the visualization displays content including: environmental characteristics, FRI feature heatmaps, training history curves, confusion matrices, and FRI box plots.

[0108] The pollution monitoring law enforcement management module specifically manages various pollution events, is connected with the notification and visual display module, manages archived pollution events, and performs statistical analysis to generate a log file; the log file is used to record classified reports, confusion matrices, and feature importance analysis corresponding to various pollution events, thereby providing data support for environmental pollution event management and environmental law enforcement.

[0109] The watershed water environment monitoring and rapid tracing of emergency pollution method provided by the embodiment more specifically includes the following steps:

[0110] S1, constructing a multi-parameter collaborative monitoring model

[0111] S1-1 model construction

[0112] A water quality pollution classification multi-layer perceptron MLP model and an LSTM / GRU model covering the entire watershed water environment area are constructed, trained, and deployed, and then the environmental features, pollutants, and pollution source information of the watershed key node grid are imported. The model is built using the PyTorch framework, with a total parameter quantity of no less than 1.5M, and the architecture includes a 225-dimensional input layer, multiple hidden layers (including Linear + Batch Norm + ReLU + Dropout) with dimensions of [512, 256, 128], and a classification layer (corresponding to five categories of water sample sources, i.e., fish pond aquaculture, groundwater, surface water, wastewater, and simulated emergency pollution wastewater) outputting five categories. The dropout rate of the Dropout layer in the hidden layer is set to 0.3.

[0113] The multi-layer perceptron MLP model is constructed, including designing the model architecture, including designing the input layer, multiple hidden layers, and the classification layer in sequence, according to the following steps:

[0114] 1) input layer (225 dimensions) -> 512-dimensional hidden layer (Linear + Batch Norm + ReLU + Dropout);

[0115] 2) -> 256-dimensional hidden layer (Linear + Batch Norm + ReLU + Dropout);

[0116] 3) -> 128-dimensional hidden layer (Linear + Batch Norm + ReLU + Dropout);

[0117] 4) the classification layer outputs classification probabilities of five categories, corresponding to five categories of water sample sources, i.e., fish pond aquaculture, groundwater, surface water, wastewater, and simulated emergency pollution wastewater;

[0118] The design of the hidden layer further includes:

[0119] 1) Linear transformation layer: sequentially transform input features into 512 dimensions, 256 dimensions and 128 dimensions, realize the step-by-step dimension reduction and abstraction of features;

[0120] 2) Batch normalization layer, used to accelerate convergence and stabilize training;

[0121] 3) ReLU activation function, introducing nonlinear transformation;

[0122] 4) Dropout layer, set the dropout rate to 0.3, to prevent model overfitting.

[0123] Build LSTM / GRU model for time series feature extraction, its structure is:

[0124] Input: real-time monitoring time series data window (such as pH, dissolved oxygen, COD and other indicators of a sampling point in the past 1 hour, divided by time step, shape [time step, index dimension]).

[0125] Processing: capture time series dependence through the recurrent layer of LSTM / GRU (such as time correlation of index sudden rise / sudden drop, periodic fluctuation rule), output time series feature vector (such as the hidden state of the last time step, or the global pooling feature of all time steps, shape [1, time series feature dimension]).

[0126] Build MLP module for classification / anomaly judgment, its structure is:

[0127] Input: LSTM / GRU output time series feature vector + sampling point static features (such as sampling point coordinates, surrounding pollution source type (factory / farmland), river flow rate, etc., shape [1, static feature dimension]).

[0128] Processing: MLP performs nonlinear mapping on fused features through 2-3 layers of fully connected layers (activation function uses ReLU), and the output layer uses Softmax (classification task, such as pollution type: industrial pollution / agricultural non-point source pollution) or Sigmoid (anomaly monitoring, such as "normal / abnormal" binary classification).

[0129] S1-2 model training

[0130] Model training is performed using known historical data sets, which are divided into training, validation, and test sets in an 8:1:1 ratio. The SMOTE technique is used to oversample the minority class in the training set (k neighbors = 3). The feature vector is standardized using StandardScaler. The loss function is weighted cross-entropy loss (weights are calculated according to the inverse of the class sample number), the optimizer is Adam (initial learning rate 0.0005, weight decay 0.0001), Lambda LR or Cosine Annealing LR learning rate scheduler is set, and early stopping mechanism (patience value 10, stop according to validation loss) is introduced to prevent overfitting.

[0131] Collaborative training of MLP and LSTM / GRU models: Joint training of two types of modules, with a loss function of classification cross-entropy (classification task) or MSE + cross-entropy (abnormal monitoring, considering error and classification).

[0132] S1-3 Model evaluation and pollution source verification

[0133] Model performance is evaluated on the test set, and classification report is used to calculate accuracy, precision, recall, and F1 score, and a confusion matrix is drawn to analyze the classification effect. Random Forest Classifier is used to calculate feature importance using feature importances to identify key features; according to the classification results and feature importance, infer the pollution source, generate a source report, and then verify the accuracy of the report;

[0134] The classification results of water pollution detection are compared with the known actual source categories to analyze the classification accuracy and pollution source tracing ability, including: calculating the accuracy, precision, recall, and F1 score on the test set; generating a confusion matrix to analyze the misclassification between categories; using a random forest model to evaluate feature importance to identify the most important environmental and spectral features for classification. The trained model is deployed to a network server.

[0135] S2, Collect water data

[0136] Setting up field sampling points and sampling stations in the grid of key nodes in the river basin. The distribution of points is based on DEM grid division, taking into account the hydrological and hydrodynamic characteristics (rapid flow, slow flow, inlet), the surrounding of fixed and mobile pollution sources, and environmental sensitive areas and risk hotspots. Field sampling points automatically collect real-time data sets (including geographic information GIS, time, flow rate, color, turbidity, pH, conductivity, DO, residual chlorine, heavy metal pollutant concentration, etc.) with verification marks, and transmit them with verification marks. Sampling station instruments take regular or irregular samples (emergency mode starts at any time), detect water samples to obtain time, location, environmental characteristics (including latitude and longitude, industrial activity related parameters, river type), and fluorescence excitation-emission matrix (EEM) spectral data, and form time-sharing data sets with verification marks for transmission. The specific steps include the following:

[0137] S2-1 Setting up field sampling points and sampling stations in the grid of key nodes in the river basin

[0138] First, based on DEM (Digital Elevation Model), the river basin is divided into equal-area grids (such as 1 km x 1 km), and each grid contains at least one representative sampling point. Then, the Kriging interpolation method is used to analyze the spatial distribution of water quality parameters to ensure that the sampling points can reflect the average water quality in the grid and avoid data bias caused by insufficient points (such as COD concentration spatial variation coefficient > 15%, which requires additional distribution of points);

[0139] Second, set up field sampling points at key nodes of water quality changes

[0140] According to the hydrological and hydrodynamic characteristics, considering the water flow pattern and pollutant diffusion law, the tidal and runoff periodicity, set up sampling points at the places where the water flow state changes suddenly, such as waterfalls, shoals, slow-flow areas (such as river bays, reservoirs), and tributary inlets, because these areas are prone to pollution mixing or deposition;

[0141] Third, set up field sampling points and sampling stations around fixed and mobile pollution sources, focusing on industrial wastewater discharge outlets, urban sewage treatment plant effluent outlets, farmland drainage outlets, and ship navigation intensive areas (mobile sources) to ensure that pollution sources are directly monitored;

[0142] Fourth, according to environmental sensitive areas and risk hotspots, additional distribution of points is required in sensitive areas such as drinking water source protection areas, nature reserves, and aquaculture areas. Historical pollution events high-risk areas (such as river sections where heavy metal leakage has occurred) are considered as long-term monitoring priorities;

[0143] S2-2 Collecting river basin water pollution and water environment data at field sampling points

[0144] Collecting the water environment data of the basin, and adding the data collection device ID, GIS (geographic information), and time stamp, encryption information, and unalterable identifier verification marks in the data;

[0145] Among them, the sensors of the field sampling points automatically collect water body data in real time, obtain GIS, time, flow rate, color, turbidity, pH, conductivity, DO, residual chlorine, heavy metal pollutant concentration, and other information of the water body, obtain real-time data groups of water pollution, and transmit the data groups to the network server after adding verification marks;

[0146] S2-3 Sampling station collects water environment data of the basin

[0147] The sampling instrument of the sampling station is periodically sampled in the conventional monitoring mode to detect water pollution of the water sample, and is irregularly sampled in the emergency monitoring mode to start sampling at any time and detect water pollution of the water sample;

[0148] The sampling station obtains the environmental characteristics and the fluorescence excitation-emission matrix (EEM) spectral data of the water sample through water sample collection and data analysis, specifically as follows:

[0149] S2-3-1 Collects the water sample and records the environmental characteristics, the environmental characteristics including longitude and latitude, industrial activity related parameters, and river type; and collects the EEM spectral data of the water sample using a fluorescence spectrophotometer, the excitation wavelength range being 250-500 nm, the emission wavelength range being 250-500 nm, and the step length being 5 nm, to generate a 52x52 fluorescence intensity matrix;

[0150] S2-3-2 Fuses the environmental characteristics and the EEM spectral characteristics processed by grid integration (FRI), to generate a 225-dimensional feature vector, wherein the environmental characteristics account for 56 dimensions, and the FRI spectral characteristics account for 169 dimensions;

[0151] S2-3-3 Obtains a time-sharing data group including the time, location, environmental characteristics, and fluorescence excitation-emission matrix (EEM) spectral data of the water sample, and transmits the data group to the network server after adding verification marks.

[0152] S3, data feature extraction and preprocessing

[0153] S3-1 Constructing a multivariate correlation data set

[0154] The real-time data set and the split-time data set are analyzed and processed to construct a multi-element correlation data set with a verification mark. The correlation parameters include device ID-time-place-pollutant type-pollution source-verification mark. After analyzing and processing the real-time data set and the split-time data set of water quality pollution in the grid of each key node in the basin, a multi-element correlation data set with a verification mark is constructed. The correlation parameters in the multi-element correlation data set include device ID-time-place-pollutant type-pollution source-verification mark, so that the multi-element correlation data set meets the data evidence specification requirements of administrative law enforcement.

[0155] S3-2 Feature extraction and preprocessing

[0156] Data fusion and preprocessing are performed, and EEM spectral data are subjected to grid integration processing. Then, FRI features and environmental features are horizontally spliced. Data standardization: the water sample source label is converted into an integer label using sklearn preprocessing Label Encoder; that is, the water sample source label is converted into an integer using Label Encoder, and the feature vector is standardized using StandardScaler; then, feature extraction and preprocessing are performed, EEM spectral data are subjected to grid integration (FRI) processing, FRI features and environmental features are horizontally spliced to generate a 225-dimensional feature vector (environmental features account for 56 dimensions, and FRI spectral features account for 169 dimensions).

[0157] S3-3 Feature importance analysis

[0158] The feature importances attribute of the Random Forest Classifier is used to identify the dominant features in the data.

[0159] S4, model running for anomaly monitoring and traceability response

[0160] In the conventional monitoring mode, the AI-driven pollution identification LSTM / GRU model uses the conventional monitoring mode for real-time data of water quality pollution at the sampling point to perform real-time classification and index anomaly monitoring. If an index anomaly or a deviation of the MLP model classification result from the preset threshold (an increase in the abnormal category probability) is found, it is preliminarily determined that an emergency pollution event has occurred, the system immediately switches to the emergency monitoring mode, simultaneously starts the cross-validation and pollution traceability mechanism, and simultaneously starts the encryption sampling of downstream sampling points and sampling stations.

[0161] S4-2 uses the trained MLP model to classify and trace pollutants, and compares and verifies real-time data sets and time-sharing data sets, makes multi-parameter collaborative judgments, and determines whether a water body emergency pollution event has occurred. If it is confirmed, further classify the pollutants, preliminarily determine and output the pollution source and predict the diffusion range information, and hand over to law enforcement personnel for further manual verification and processing, and the manual verification result data is entered into the network server.

[0162] In this step, the two types of models can be connected in series and cooperated, or parallel inference can be performed. The MLP branch: input static features, learn "index combination patterns" under different water quality categories through fully connected layers, and output static classification probabilities (such as "the probability of the current static feature corresponding to'mild pollution' is 80%"). Advantages: quickly capture nonlinear relationships between indicators (such as "low pH often accompanied by high ammonia nitrogen"), fast inference speed, and suitable for real-time response. The LSTM / GRU branch: input time series features, capture long-term dependencies through the gating mechanism (such as "COD usually rises and then falls in the rainy season"), and output time series classification probabilities (such as "the probability of the current time series trend corresponding to'moderate pollution' is 75%"). Advantages: avoid the problem that traditional time series models (such as ARIMA) cannot handle nonlinear trends, especially suitable for periodic and sudden changes in water quality indicators. The two types of models can be fused through output to improve classification robustness.

[0163] The outputs of the two types of models are integrated through a fusion layer to solve the limitations of a single model (such as MLP ignoring trends and LSTM being sensitive to instantaneous anomalies), including:

[0164] Basic fusion: use weighted average (weights can be determined through historical data verification, such as static feature weight 0.4 and time series feature weight 0.6), and take the highest probability category as the final result.

[0165] Advanced fusion: introduce attention mechanism, dynamically adjust weights - for example, when the index value fluctuates gently (weak time series features), increase the MLP weight; when the index changes dramatically (strong time series features), increase the LSTM / GRU weight.

[0166] Conflict processing: if the highest probability categories output by the two types of models are different (such as MLP "good" and LSTM "mild pollution"), and the confidence is higher than the threshold (such as > 60%), trigger the "ambiguous classification" label (such as "good - mild pollution transition"), and push it to the law enforcement end for auxiliary decision-making.

[0167] The two models work together to monitor index anomalies, including identifying "values outside the normal range" or "changes inconsistent with trends" (such as sudden leaks leading to a sharp increase in COD). The collaborative logic process is as follows:

[0168] (1) Define anomalies and model division:

[0169] Static anomaly: The value or combination of indicators at a certain time deviates significantly from the historical normal sample (such as pH = 4.0, far below the historical normal range of 6.5-8.5), detected by MLP through "normal sample distribution learning" (such as calculating the reconstruction error of the current feature and the normal sample).

[0170] Time series anomaly: The change trend of the index deviates from the historical law (such as dissolved oxygen decreasing from 5mg / L to 1mg / L in 10 minutes, far exceeding the historical maximum change rate), detected by LSTM / GRU through "time series prediction error" (such as the residual error between the predicted value and the actual value at the next time).

[0171] (2) Parallel detection of models

[0172] Among them, the MLP model is responsible for static anomaly detection. During training, MLP learns the static feature distribution of normal samples (such as reconstructing normal samples through autoencoders, and the reconstruction error of abnormal samples is large); during real-time monitoring, the reconstruction error of the current static feature is calculated, and if it exceeds the threshold (such as the 95th percentile), it is marked as a "static anomaly candidate".

[0173] The LSTM / GRU model is responsible for time series anomaly detection. During training, LSTM predicts future values based on historical time series data (such as predicting the value at t+1 using data from t-5 to t), and learns the normal trend; during real-time monitoring, the value at t+1 is predicted using data from t-5 to t, and the residual error between the predicted value and the actual value at t+1 is calculated, and if the residual error exceeds the threshold (such as 3σ based on historical residual error distribution), it is marked as a "time series anomaly candidate".

[0174] (3) Anomaly confirmation and linkage

[0175] The anomaly label of a single model may be false (such as MLP misjudging occasional fluctuations as static anomalies, and LSTM misjudging short-term noise as time series anomalies), which needs to be confirmed by the two models:

[0176] Double verification: When "static anomaly candidate" and "time series anomaly candidate" appear at the same time, it is determined as "real anomaly" (such as a sharp increase in COD accompanied by a deviation from normality in its combined features, and an abnormal change trend), which immediately triggers an alarm.

[0177] Hierarchical response: If only static abnormality (such as transient value exceeding threshold but normal trend, which may be sensor error), mark as "low priority abnormality", wait for next time step data verification; if only time series abnormality (such as abnormal trend but value not exceeding threshold, which may be initial stage of pollution), mark as "medium priority abnormality", enhance the monitoring frequency of this indicator (such as from 1 time per 5 minutes to 1 time per 1 minute).

[0178] S5, model updating and continuous monitoring: using the multi-correlation dataset verified by environmental law enforcement and the newly obtained multi-correlation dataset in the running process, re-training the multi-layer perceptron MLP model and the LSTM / GRU model, and deploying the new model to the network server for running, repeating steps S2-S4 (i.e. collecting water body data, data feature extraction and preprocessing, model running abnormality monitoring and traceability response), analyzing and processing real-time data groups and time-sharing data groups, and outputting monitoring results, which is periodically performed (usually once a week), forming a continuous optimization closed loop.

[0179] Embodiment 3

[0180] Referring to Figures 2-8 , based on embodiments 1 and 2, this embodiment further proposes a specific method and system for watershed water environment monitoring and rapid emergency pollution traceability based on water quality pollution detection, aiming at the large area, wide range, complex water flow and network of the Pearl River Basin, through the cooperative collection of real-time data and time-sharing data by distributed sampling points and sampling stations, the application of multi-layer perceptron MLP model for multi-parameter collaborative monitoring and classification, and the combination of AI-driven pollution identification LSTM / GRU model for real-time abnormality detection. This method and system fuse EEM spectral data and environmental characteristics and other multi-source heterogeneous data, and use batch normalization, Dropout and SMOTE oversampling technology for data processing and model training, and automatically identify whether an emergency pollution event occurs according to the preset conditions. If it is preliminarily judged that an emergency pollution event occurs, the system will quickly start pollution type judgment, traceability and cross verification, automatically output the results and further check by manual, finally realize rapid monitoring, emergency traceability and support environmental law enforcement of watershed water environment, meet the demand of large area and wide range application, and optimize efficiency and cost through network technology.

[0181] The method for rapid traceability of Pearl River Basin water environment monitoring and emergency pollution provided in this embodiment has the following specific steps:

[0182] Step 1: To build a classification model suitable for watershed water pollution detection and water environment monitoring, a water sample dataset was collected and expanded. This step corresponds to data collection, feature extraction, and preprocessing of methods S2 and S3. 311 water samples were collected, with a class distribution of: fish pond farming (41), groundwater (32), surface water (129), wastewater (17), and simulated emergency pollution wastewater (92). The water sample information collection is shown in Table 1. Specifically, in addition to fish pond farming water, groundwater, surface water, and wastewater, which are field-collected water sample samples with verification marks, simulated emergency pollution wastewater is laboratory-prepared wastewater that simulates enterprise wastewater leakage or illegal discharge and other emergency water environmental pollution events, and diffuses in surface water. The preparation method of simulated emergency pollution wastewater is: 500 mL of wastewater is collected from the enterprise wastewater collection tank and transported back to the laboratory. A large amount of surface water is collected near the enterprise location. After returning to the laboratory, the original enterprise wastewater and river water are prepared according to the ratios of 1:1, 1:2, 1:3, 1:4, 1:6, 1:8, 1:10, 1:20, 1:50, 1:100, etc. For each water sample, environmental characteristics including 56-dimensional environmental characteristics (including latitude and longitude, regional information, industrial activity-related parameters, river type, and FluI, FreI, BIX, HIX, etc. are collected), and EEM spectrum data (excitation / emission wavelength range 250-500 nm, step size 5 nm, generating a 52x52 fluorescence intensity matrix) are collected using a fluorescence spectrophotometer. Python scripts are used to read the data of each sample.

[0183] Table 1

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190] Subsequently, feature extraction was performed on the EEM spectral data: the 52×52 matrix was integrated into a 13x13 grid using the `calculate gridintegrals` function, generating 169-dimensional FRI features. During feature extraction, it was ensured that no NaN or infinite values ​​would interfere with data quality. Finally, the 56-dimensional environmental features and the 169-dimensional FRI features were horizontally concatenated and fused to form a 225-dimensional feature vector, which served as the dataset for modeling. Data preprocessing included standardization using the Standard Scaler and oversampling of the training set using SMOTE (k neighbors=3). Labels were transformed using the Label Encoder.

[0191] Step 2: Construct a distributed water environment monitoring and pollution source tracing system to process, analyze, and output results for each specific monitoring and source tracing step. To achieve distributed sampling, the system includes multiple distributed field sampling points and stations. The network server has a built-in rapid source tracing program for watershed water environment monitoring and emergency pollution control.

[0192] Step 3: Model Training and Optimization. This step corresponds to the training and partial evaluation in step S1 of the aforementioned embodiment. On the web server, the model from step S1 is trained using the data processed in step S3. The MLP model parameters and training hyperparameters are as follows:

[0193] Input dimension: 225; Hidden layer structure: [512, 256, 128]; Activation function: ReLU; Regularization: BatchNorm + Dropout (0.3); Output layer: 5 classes.

[0194] Training hyperparameters: Optimizer: Adam (lr=0.0005, weight decay=1e-4); Batch size: 64; Scheduler: Cosine Annealing LR (T max=100) or Lambda LR (segmented decay). Early stopping mechanism (patience value 10).

[0195] Training results show that Cosine Annealing LR outperforms Lambda LR (96.83% test accuracy) in both test accuracy (98.41%) and validation loss (0.0224). This is evident in the Cosine Annealing LR learning rate scheduler model. Figure 3 As shown, both curves decrease rapidly with increasing training epochs, converging to a very low level close to 0 after approximately 20 epochs. The curves for training loss and validation loss are very similar, further demonstrating the model's good generalization ability and stable convergence. Meanwhile, Figure 4It is shown that both curves rise rapidly with the increase of training rounds (horizontal axis), and after about 15-20 rounds, both training accuracy and validation accuracy reach a level close to or equal to 100% and tend to be stable. The two curves are highly coincident, indicating that the model learns quickly without overfitting. Figure 5 It is shown that the learning rate starts from the initial 0.0005 and presents a smooth cosine function decline curve, gradually and slowly attenuating to 0. This smooth attenuation feature helps the model to fine-tune more precisely in the later training period and find a better local minimum. This is fully consistent with the data in column F of the training log table, verifying the successful implementation of the cosine annealing learning rate strategy.

[0196] Figure 6 The model loss value change curve of the Lambda learning rate scheduler is shown, which contains two curves: training loss (blue) and validation loss (orange). Both curves quickly decrease with the increase of training rounds and converge to a very low level close to 0 after about 15-20 rounds. The curve trend of training loss and validation loss is very close, further proving that the model has good generalization ability. At the same time, Figure 7 It is shown that both curves rise rapidly with the increase of training rounds, and after about 15-20 rounds, both training accuracy and validation accuracy reach a level close to 100% and tend to be stable, indicating that the model learns very quickly and there is no obvious overfitting. Figure 8 It is shown that during 0 to 30 training rounds, the learning rate is constant at 0.0005. At the 31st training round, the learning rate undergoes a stepwise decrease, suddenly dropping to 0.00005, and remains unchanged in the subsequent training. This is fully consistent with the data in column F of the training log table, verifying the successful implementation of the piecewise decay strategy.

[0197] Step 4: Data processing and classification result display. This step corresponds to part of the functions of steps S3 and S4 in the foregoing embodiments. As shown in FIG. 8, the data processing and classification result display interface is shown. Figure 2As shown, in the embodiments of the present application, when using two different learning rate scheduling strategies of LambdaLR and Cosine Annealing LR respectively, the classification effect confusion matrix diagram obtained on the test set by using the improved multilayer perceptron MLP model. The test set contains a total of 63 samples. The vertical axis of the confusion matrix represents the real classification of the sample, and the horizontal axis represents the predicted classification of the model. The values on the diagonal line represent the number of samples correctly predicted by the model, and the values on the non-diagonal line represent the number of samples incorrectly predicted by the model. The performance of the two strategies is compared in the figure. When using the Lambda LR learning rate scheduling strategy, the model has a total of 2 misclassifications on the test set: for the "fish pond breeding water" category: of the total 8 real samples, 7 are correctly classified, but 1 is incorrectly predicted as "groundwater". For the "surface water" category: of the total 26 real samples, 25 are correctly classified, but 1 is incorrectly predicted as "fish pond breeding water". For the "groundwater", "waste water" and "simulated emergency pollution waste water" three categories, the model achieves 100% correct classification. Under this strategy, the model correctly predicts a total of 61 samples, and the total classification accuracy is 96.83% (61 / 63). When using the Cosine Annealing LR learning rate scheduling strategy, the performance of the model is further improved, and there is only 1 misclassification: for the "fish pond breeding water" category: all 8 real samples are correctly classified. For the "surface water" category: of the total 26 real samples, 25 are correctly classified, and 1 is still incorrectly predicted as "fish pond breeding water". For the "groundwater", "waste water" and "simulated emergency pollution waste water" three categories, the model also achieves 100% correct classification. Under this strategy, the model correctly predicts a total of 62 samples, and the total classification accuracy is improved to 98.41% (62 / 63). It shows that the model has high classification accuracy and can quickly and accurately identify the source of water samples.

[0198] Step 5: Abnormality detection and rapid tracing process. This step corresponds to the core process of step S4 in the previous embodiment. When the AI-driven LSTM / GRU model detects real-time data anomalies or MLP classification results deviate from the preset threshold (such as abnormal category probability increases), the system triggers an abnormal alarm and starts the rapid tracing mechanism: simultaneously starting cross-validation and pollution tracing mechanism, and simultaneously downstream encrypted sampling. Combined with real-time data set and time-sharing data set comparison analysis, cross-validation, multi-parameter collaborative judgment is carried out, if it is confirmed that pollution event occurs, the pollution type identification, geographical location positioning, diffusion range estimation (combined with water flow model and historical data) are carried out, and the pollution source is preliminarily judged, the pollution enterprise database is matched, and the tracing result is verified by the accounting verification model. The tracing result is pushed to the relevant departments through the notification and visualization display module and the alarm (including pollution source location, pollution type and other information) is automatically sent, prompting law enforcement personnel to collect evidence on site according to the tracing analysis result, guiding rapid handling of pollution events.

[0199] MLP and LSTM / GRU model can also use parallel collaborative (multi-source feature parallel fusion) mode to run, which processes different types of input features (static features vs time series features) by MLP and LSTM / GRU in parallel, and then integrates the outputs of the two through a fusion layer to realize joint decision making; suitable for scenarios where static features and time series features are equally important (such as "pollution source intensity around sampling point + real-time index fluctuation" jointly determining abnormal level).

[0200] Step 6: System evaluation and optimization. This step corresponds to part of the evaluation and continuous monitoring of steps S1 and S5 in the previous embodiment. The system regularly verifies the model performance using the test set, analyzes the classification performance of the model on the test set, and evaluates the accuracy and reliability. The system manages the multi-element correlation data set that has been manually checked and verified through the pollution event management module, and uses these data to retrain and improve the MLP, LSTM / GRU model as described in step S5, to realize continuous optimization of the model and improvement of the system performance. The system follows the step-by-step process of steps S1-S5 in daily operation, forming a closed loop of continuous monitoring, analysis, response and optimization.

[0201] The watershed water environment monitoring and emergency pollution rapid tracing system provided by the embodiment can implement any combination of the steps in the above method, and there is no strict sequence between the steps implemented in the system, which can be combined and adjusted as needed. For example, the equipment system loads the fluorescence excitation-emission matrix (EEM) spectrum data verified with the mark and the environmental characteristics through the data processing module, calculates the grid region integral (FRI) characteristics through the calculate grid integrals function, fuses the FRI characteristics and the environmental characteristics into a 225-dimensional feature vector, performs Standard Scaler standardization and SMOTE oversampling processing, and finally uses the improved multilayer perceptron (MLP) model trained to realize high-precision 5-class classification of the water sample source in the classification and anomaly monitoring module, and combines the output results of the model to provide support for pollution tracing in the tracing and prediction module.

[0202] The above embodiments of the present application realize efficient and accurate water sample and pollutant classification and pollution tracing by means of real-time and time-sharing data acquisition, multi-parameter collaborative monitoring and AI-driven pollution identification algorithm model, and take into account daily monitoring and emergency needs. The system is stable, efficient, has a wide coverage, and has a low operation cost, and has a wide popularization value. It has important significance for large watershed water environment protection and sustainable development. For example, in the emergency monitoring of sudden leakage events, the pH, conductivity, heavy metal concentration and other sensor data are analyzed, and when the indicators are abnormal, an alarm is triggered to quickly locate the pollution area, and the diffusion range is predicted in combination with the flow rate data to quickly determine the emergency disposal and environmental law enforcement scheme. In the early warning monitoring of drinking water sources, DO, residual chlorine, turbidity and other sensors are used for real-time monitoring of water sources, and when the indicators are abnormal (such as DO fluctuation caused by algae outbreak), an alarm is triggered to quickly locate the abnormal area, and the diffusion range is predicted in combination with the flow rate data to quickly determine the emergency disposal and environmental law enforcement scheme.

[0203] It should be noted that in other embodiments of the present application, other different schemes obtained by making specific choices within the scope of the steps, instruments, algorithms, models and process parameters and conditions described in the present application can achieve the technical effects described in the present application, and therefore the present application will not list them one by one.

[0204] The above description is only a preferred embodiment of the present application, and does not limit the present application in any form. Any person skilled in the art can make many possible changes and modifications to the technical solutions of the present application, or modify equivalent embodiments, without departing from the scope of the technical solutions of the present application, by using the methods and technical contents disclosed above. Any equivalent changes made according to the components, proportions and processes of the present application shall be covered within the protection scope of the present application.

Claims

1. A method for watershed water environment monitoring and emergency pollution rapid tracing based on water quality pollution detection, characterized in that, It is based on water quality pollution detection, multi-source data collection and artificial intelligence model fusion analysis, and is used for rapid tracing of basin water environment monitoring and emergency pollution, including the following steps: first, a multi-parameter collaborative monitoring model is built, including a multi-layer perceptron MLP model and an LSTM / GRU model, and is trained and deployed; then, through distributed sampling points and sampling stations, real-time data and time-sharing data of basin water environment and water quality pollution detection are collected, and an AI-driven pollution identification MLP model and an LSTM / GRU model are used to fuse and analyze multi-source heterogeneous data by combining batch normalization, Dropout and SMOTE oversampling technology, and by fusing EEM spectrum data and environmental characteristics; then, whether an emergency pollution event occurs is automatically identified according to preset conditions, and if it is preliminarily judged that an emergency pollution event occurs, the pollution type judgment, tracing and cross-validation are started, and the pollution type and tracing result of water environment monitoring and emergency pollution are automatically output, which are further checked by artificial, environmental law enforcement is checked, and the monitoring of basin water environment pollution and the rapid tracing of emergency pollution are completed. The multi-layer perceptron MLP model is built, including the following steps: S1-1 model building: a multi-layer perceptron MLP model of water quality pollution is built by using a PyTorch framework, and the total parameter quantity is not less than 1.5M, the multi-layer perceptron model architecture includes an input layer, multiple hidden layers and a classification layer; The hidden layer includes a linear transformation layer, a batch normalization layer, a ReLU activation function and a Dropout layer; The classification layer outputs five categories of water sample classification probabilities, corresponding to five categories of water sample sources; S1-2 model training: the multi-layer perceptron MLP model is trained by using a known historical data set, a weighted cross-entropy loss function is used, the weight is calculated by the inverse of the number of samples in each category to balance the influence of minority classes; an Adam optimizer is used, a learning rate scheduler is set as Lambda LR or Cosine Annealing LR; an early stopping mechanism is introduced to stop training according to the validation loss; and the model parameters are adjusted to optimize the model parameters to improve the classification and tracing performance; S1-3 model evaluation and pollution tracing verification: the model performance is evaluated on the test set, the accuracy, precision, recall and F1 score are calculated using classificationreport, and the classification effect is analyzed by drawing a confusion matrix; the feature importance is calculated by using a RandomForest Classifier to identify key features; according to the classification result and the feature importance, the pollution source is inferred, a tracing report is generated, and the accuracy of the report is verified; The classification result of water quality pollution detection is compared with the known actual source category to analyze the classification accuracy and pollution tracing ability, including: calculating the accuracy, precision, recall and F1 score on the test set; generating a confusion matrix to analyze the misclassification between categories; and using a random forest model to evaluate the feature importance to identify the environmental characteristics and spectral characteristics that contribute most to classification.

2. The method according to claim 1, wherein, Specifically comprising the following steps: S1, constructing a multi-parameter collaborative monitoring model Constructing a multi-layer perceptron MLP model and an LSTM / GRU model for multi-parameter collaborative monitoring, training and deploying them; Constructing a multi-layer perceptron MLP model and an LSTM / GRU model for multi-parameter collaborative monitoring and rapid pollution tracing in the entire river basin water environment area, and importing the environmental characteristics, pollutants and pollution source information of the key node grid in the entire river basin water environment area; Using known historical data sets to train the constructed multi-layer perceptron MLP model and LSTM / GRU model, and adjusting the model parameters to optimize the model parameters and improve the classification and tracing performance; Deploying the trained multi-layer perceptron MLP model and LSTM / GRU model to the network server; S2, collecting water data Setting up field sampling points and sampling stations in the key node grid of the river basin, collecting water quality pollution detection and water environment data, and adding data collection device ID, GIS and timestamp, encryption information and tamper-proof identifier verification marks to the data; The sensors at the field sampling points automatically collect water data in real time, obtain GIS, time, flow rate, color and transparency information of the water body, and get real-time data sets of water quality pollution detection, and then transmit them to the network server after adding verification marks; The sampling instruments at the sampling stations periodically or irregularly sample and detect the water samples, obtain data including time, location, environmental characteristics and fluorescence excitation-emission matrix EEM spectral data of the water samples, and get time-sharing data sets of water quality pollution detection, and then transmit them to the network server after adding verification marks; S3, data feature extraction and preprocessing Constructing a multi-element correlation data set, analyzing and processing the real-time data sets and time-sharing data sets of water quality pollution detection in each key node grid in the river basin, and constructing a multi-element correlation data set with verification marks, wherein the correlation parameters include device ID-time-location-pollutant type-pollution source-verification mark; then performing feature extraction and preprocessing; S4, model running for abnormal monitoring Using the LSTM / GRU model combined with the multi-layer perceptron MLP model, the real-time data of the field sampling points are monitored in the conventional monitoring mode for real-time classification and index abnormality monitoring; if an index abnormality is found, the emergency monitoring mode is switched in, the cross-validation and pollution tracing mechanism are started simultaneously, the downstream sampling points and sampling stations are simultaneously sampled, the real-time data sets and time-sharing data sets are compared and analyzed, cross-validated, and it is judged whether a water emergency pollution event has occurred; if a water emergency pollution event has occurred, the pollutants are further classified, the pollution source is preliminarily judged and output, the diffusion range information is predicted, and the artificial further checking and processing are performed for environmental law enforcement checking, and the artificial checking result data is input into the network server; S5, model updating and continuous monitoring The multi-element correlation data set verified by environmental law enforcement verification is used to retrain the multi-layer perceptron MLP model and the LSTM / GRU model, and the new model is deployed into the network server to run, and steps S2-S4 are repeated to analyze and process the real-time data set and the split-time data set, and output the monitoring result.

3. The method according to claim 1, wherein, The step S1-2 model training specifically includes the following steps: The known historical data set is used to train the constructed model, including: S1-2-1 divides the data set into a training set, a validation set and a test set, with a ratio of 8:1:1, and uses stratified sampling to ensure consistent class proportions; S1-2-2 oversamples the minority class in the training set using the SMOTE technique, sets k neighbors=3, generates synthetic samples to balance the class distribution; S1-2-3 uses Standard Scaler to standardize the 225-dimensional feature vector, so that the mean is 0 and the variance is 1; uses a weighted cross-entropy loss function, where the weight is calculated according to the inverse of the class sample number, uses the Adam optimizer for training, the initial learning rate is 0.0005, and the weight decay is 0.0001; S1-2-4 sets the Cosine Annealing LR or Lambda LR learning rate scheduling strategy, and combines the early stopping mechanism with a patience value of 10 to prevent overfitting.

4. The method according to claim 2, wherein, The step S2 collects water body data, specifically including the following steps: S2-1 sets field sampling points and sampling stations in the key node grid of the watershed First, the watershed is divided into equal-area grids based on the digital elevation model DEM, and each grid contains at least one representative sampling point, and then the Kriging interpolation method is used to analyze the spatial distribution of water quality parameters to ensure that the sampling points can reflect the average water quality in the grid and avoid data deviation due to insufficient points; Second, set field sampling points at key nodes of water quality changes According to the hydrological and hydrodynamic characteristics, considering the flow pattern and pollutant diffusion law, the tide and runoff periodicity, set sampling points at places where the flow is turbulent, slow, and the flow state changes suddenly at the entrance of the tributary; Third, set field sampling points and sampling stations around fixed and mobile pollution sources, focusing on industrial wastewater discharge outlets, urban sewage treatment plant outlets, farmland drainage outlets, and ship navigation intensive areas to ensure that pollution sources are directly monitored; Fourth, according to environmental sensitive areas and risk hotspots, increase the number of points in sensitive areas such as drinking water source protection areas, nature reserves, and aquaculture areas, and consider historical pollution events in high-risk areas as long-term monitoring priorities; S2-2 Collects watershed water pollution and water environment data at field sampling points Collect watershed water environment data, and add data collection device ID, GIS and timestamp, encryption information, and tamper-proof identifier verification markers to the data; The sensors of the field sampling points automatically collect water body data in real time to obtain information of GIS, time, flow rate, color, turbidity, pH, conductivity, DO, residual chlorine, and heavy metal pollutant concentration of the water body, obtain a real-time data group of water pollution, and transmit the data group to the network server after adding a verification mark; The S2-3 sampling station collects the water environment data of the basin The sampling instrument of the sampling station is used for periodic sampling in the conventional monitoring mode, water quality pollution detection is performed on the water sample, and in the emergency monitoring mode, the sampling station is used for irregular sampling, sampling is started at any time, and water quality pollution detection is performed on the water sample; the sampling station obtains the environmental characteristics and the fluorescence excitation-emission matrix EEM spectral data of the water sample through water sample collection and data analysis, and the environmental characteristics and the fluorescence excitation-emission matrix EEM spectral data of the water sample are specifically as follows: S2-3-1 collects the water sample and records the environmental characteristics, the environmental characteristics include longitude and latitude, industrial activity related parameters and river type; the EEM spectral data of the water sample is collected by using a fluorescence spectrophotometer, the excitation wavelength range is 250-500 nm, the emission wavelength range is 250-500 nm, the step length is 5 nm, and a 52*52 fluorescence intensity matrix is generated; S2-3-2 fuses the environmental characteristics and the EEM spectral characteristics processed by the grid integration FRI to generate a 225-dimensional feature vector, wherein the environmental characteristics account for 56 dimensions and the FRI spectral characteristics account for 169 dimensions; S2-3-3 obtains a time-sharing data group including the time, place, environmental characteristics and fluorescence excitation-emission matrix EEM spectral data of the water sample, adds a verification mark, and transmits the data group to the network server.

5. The method according to claim 2, wherein, The step S3 data feature extraction and preprocessing specifically includes the following steps: S3-1 constructing a multivariate correlation data set After analyzing and processing the real-time data group and the time-sharing data group of the water quality pollution classification of the grid in each key node in the basin, a multivariate correlation data set with a verification mark is constructed, and the correlation parameters in the multivariate correlation data set include device ID-time-place-pollutant type-pollution source-verification mark; S3-2 feature extraction and preprocessing Data fusion and preprocessing are performed, the EEM spectral data is processed by grid integration, and then the FRI features and the environmental features are horizontally spliced; data standardization: the water sample source label is converted into an integer label by using sklearn preprocessing Label Encoder; S3-3 feature importance analysis, the dominant features in the data are identified by using the feature importances attribute of the Random Forest Classifier.

6. The method according to claim 2, wherein, The step S4 model running for abnormal monitoring specifically includes the following steps: S4-1 using the LSTM / GRU model, the water quality pollution real-time data of the field sampling point is monitored in the conventional monitoring mode, real-time classification and index abnormal monitoring are performed; if an index abnormality is found, the emergency monitoring mode is entered, the cross-validation and pollution tracing mechanisms are started synchronously, and the downstream sampling points and the sampling station are started synchronously. S4-2 uses a multi-layer perceptron (MLP) model to classify and trace the pollutants of water pollution, and to compare and analyze the real-time data set and the time-sharing data set, cross-verify, make a multi-parameter collaborative judgment, determine whether a water emergency pollution event has occurred, if so, further classify the pollutants, and preliminarily determine and output the pollution source and the predicted diffusion range information, and hand over to artificial further checking and processing, environmental law enforcement checking, and artificial checking result data entry into the network server.

7. The method according to claim 2, wherein, The step S5 model updating and continuous monitoring specifically includes the following steps: S5-1 uses the multi-element correlation data set verified by environmental law enforcement checking and the newly obtained multi-element correlation data set in the running process to retrain the multi-layer perceptron model, and deploy the new model to the network server for running; S5-2 repeats steps S2-S4 to analyze and process the real-time data set and the time-sharing data set, and outputs the monitoring result.

8. A river basin water environment monitoring and emergency pollution rapid tracing system based on water quality pollution detection, characterized in that, It is used to realize the watershed water environment monitoring and rapid pollution source tracing method of claim 1 to 7, which includes: A network server, a management terminal, a plurality of distributed sampling points, and a sampling station connected by a network; The management terminal is used for input and output operations; The plurality of distributed sampling points are respectively deployed at key nodes of the watershed to perform fixed-point and fixed-time sampling of water pollution, collect real-time water pollution and environmental characteristic data, and upload to the network server; The plurality of distributed sampling stations are respectively deployed at key monitoring nodes of the watershed to collect EEM spectrum data and environmental characteristics of water pollution at regular or irregular intervals, and upload to the network server; The network server is used to receive monitoring data, analyze and process, and output processing results, and has a watershed water environment monitoring and rapid pollution source tracing program built-in, which includes the following modules: A monitoring station management module is used to connect and manage each distributed field sampling point and sampling station, and receive real-time data sets returned by field sampling points and time-sharing data sets returned by sampling stations; A data processing module is used to process real-time data sets and time-sharing data sets, construct a multi-element correlation data set, and divide the known multi-element correlation data set for model training; which includes calculating the grid region integral FRI of the spectrum data through the calculate gridintegrals function, fusing multi-dimensional features, and performing standardization and SMOTE oversampling processing; A classification and anomaly monitoring module: through the trained multi-layer perceptron (MLP) model and LSTM / GRU model, the real-time data of the field sampling point is monitored in a conventional mode, and real-time classification and index anomaly monitoring are performed; if an index anomaly of water pollution is found, it is transferred to an emergency monitoring mode, and a cross-verification and pollution source tracing mechanism is started simultaneously, and the downstream sampling points and sampling stations are simultaneously sampled, and the real-time data set and the time-sharing data set are compared and analyzed, cross-verified, to determine whether a water emergency pollution event has occurred. The traceability and prediction module is used for further processing of data preliminarily judged by the model as water body emergency pollution event, positioning and classifying the pollutants, and preliminarily judging and outputting the pollution source and predicting the diffusion range information; The notification and visualization display module is used for outputting the monitoring data analysis results, including the routine monitoring information, the emergency pollution event alarm information, the pollutant classification and positioning, the pollution source and the predicted diffusion range information; The pollution monitoring law enforcement management module is used for artificial checking and processing by law enforcement personnel according to the output monitoring data analysis results, and the artificial checking result data is input into the network server.

9. The watershed water environment monitoring and emergency pollution rapid traceability system based on water quality pollution detection according to claim 8, characterized in that, The classification and anomaly monitoring module is pre-learned by the LSTM / GRU model and the MLP model for the classification data of the unpolluted water body, and is jointly completed by combining the classification function of the MLP model; when the MLP model classifies the water sample, if the classification result deviates from the preset threshold, including the increase of the abnormal category probability, the anomaly alarm is triggered; The traceability and prediction module analyzes the water sample by fusing the pollution traceability model based on the environmental characteristics and the three-dimensional fluorescence, and outputs the traceability analysis result by verifying the model verification result; after detecting the anomaly, the system immediately starts the emergency response mode, quickly locates the pollution source and predicts the diffusion range through the multi-parameter sensor data and the flow rate data; The information notification in the notification and visualization display module includes the information prompting the law enforcement personnel to perform the on-site traceability evidence according to the traceability analysis result; in the case of emergency pollution event, the system automatically sends the alarm to the relevant law enforcement personnel or department, including the pollution source position and the pollutant type information, guiding the on-site evidence and emergency treatment; The contents of the visualization display include the environmental characteristics, the FRI characteristic heat map, the training history curve, the confusion matrix and the FRI box plot; The pollution monitoring law enforcement management module specifically manages various pollution events, is connected with the notification and visualization display module, archives the pollution event management and performs statistical analysis, and generates the log file; the log file is used for recording the classification report, the confusion matrix and the feature importance analysis corresponding to various pollution events, and provides data support for the environmental pollution event management and the environmental law enforcement.

Citation Information

Patent Citations

  • Water pollution accurate tracing method based on neural network model

    CN118425458A

  • Surface water ICM concentration prediction method and system based on machine learning

    CN118430684A

  • Urban water environment small-scale traceability method based on multi-layer sensor model

    CN118098442A

  • Water body environment abnormal change detection method based on multi-source data space-time fusion network

    CN120451808A